Zyverna

API reference

OpenAI-compatible. One line to switch.

Use the official OpenAI SDKs in any language — just change the base URL and key. Pay in credits from the same balance as the workspace.

Authentication

Create a key in Account → API keys. Send it as a bearer token on every request. Keys start with zy_live_. Each key can have its own monthly credit limit; revoke it at any time.

header
Authorization: Bearer zy_live_…
Never expose keys in client-side code. Proxy requests through your backend.

Quickstart

Base URL: https://zyverna.com/api/v1

curl
curl https://zyverna.com/api/v1/chat/completions \
  -H "Authorization: Bearer $ZYVERNA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "zyverna-core",
    "messages": [{"role": "user", "content": "Explain tail-call optimization in 3 sentences."}]
  }'
node.ts
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.ZYVERNA_API_KEY,
  baseURL: "https://zyverna.com/api/v1",
});

const res = await client.chat.completions.create({
  model: "zyverna-core",
  messages: [{ role: "user", content: "Write a Zod schema for a User." }],
});

console.log(res.choices[0].message.content);
console.log("credits:", res.zyverna.credits);
main.py
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["ZYVERNA_API_KEY"],
    base_url="https://zyverna.com/api/v1",
)

stream = client.chat.completions.create(
    model="zyverna-flash",
    stream=True,
    messages=[{"role": "user", "content": "Summarize this PR description…"}],
)

for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

POST /chat/completions

Accepts the standard OpenAI request body. Supported parameters: model, messages, stream, temperature, top_p, max_tokens / max_completion_tokens, tools, tool_choice, response_format, stop, user. Responses are standard chat.completion objects with an extra zyverna field.

FieldTypeDescription
zyverna.creditsnumberCredits charged for this request (also in the X-Zyverna-Credits header).
modelstringThe Zyverna model id you requested (upstream model ids are never exposed).
usageobjectPrompt, completion and total tokens, including cached prompt tokens when available.

Streaming

Set stream: true to receive Server-Sent Events. The final chunk before [DONE] carries usage and zyverna.credits.

Tool calling

All models support parallel-safe function calling with JSON Schema parameters.

tools.ts
const res = await client.chat.completions.create({
  model: "zyverna-core",
  messages: [{ role: "user", content: "What's the weather in Toronto?" }],
  tools: [{
    type: "function",
    function: {
      name: "get_weather",
      description: "Get current weather for a city",
      parameters: {
        type: "object",
        properties: { city: { type: "string" } },
        required: ["city"],
      },
    },
  }],
});

const call = res.choices[0].message.tool_calls?.[0];
// → { function: { name: "get_weather", arguments: '{"city":"Toronto"}' } }

Models & pricing

GET /models returns the list below with live pricing. Prices are in credits per 1,000 tokens; 100 credits = 1 USD.

ModelTierContextInput / 1K tokensOutput / 1K tokens≈ USD per 1M out
Zyverna Flash
zyverna-flash
Fastest. Great for chat, quick edits and autocomplete.
Fast1M tokens0.05 cr0.2 cr$2.00
Zyverna Core
zyverna-core
Balanced. The default for agent runs and refactors.
Balanced1M tokens0.25 cr1 cr$10.00
Zyverna Ultra
zyverna-ultra
Deep reasoning for hard bugs, architecture and planning.
Frontier200K tokens0.5 cr2 cr$20.00

100 credits = 1 USD. Tool calls and agent runs are billed at the same per-token rates — there is no surcharge for agent mode. Cached prompt tokens are billed at 50% of the input rate.

Credits & billing

  • API usage draws from the same credit balance as the workspace. Included plan credits are used first, then purchased credits.
  • Cached prompt tokens are billed at 50% of the input rate.
  • Requests that fail with a 4xx error are not charged. Requests cancelled mid-stream are charged for tokens generated.
  • When your balance reaches zero, requests return 402 insufficient_credits. Top up in USD, CAD, EUR or GBP.

Rate limits

PlanRequests / minTokens / minConcurrent streams
Pay-as-you-go60100K3
Starter300400K10
Pro1,0002M25
Business3,0006M100
Scale10,00020M500 (custom on request)

Limits are returned in X-RateLimit-Limit-Requests, X-RateLimit-Remaining-Requests and Retry-After headers. Exceeding them returns 429.

Errors

Errors use the OpenAI envelope so SDK error handling works unchanged.

401 Unauthorized
{
  "error": {
    "message": "Invalid API key format.",
    "type": "invalid_request_error",
    "code": "invalid_api_key",
    "param": null
  }
}
StatusCodeMeaning
400invalid_requestMalformed body or unsupported parameter.
401invalid_api_keyMissing, malformed or revoked key.
402insufficient_creditsBalance is zero or the key's monthly limit was reached.
404model_not_foundUnknown model id.
429rate_limit_exceededSlow down; honour Retry-After.
500server_errorSomething went wrong on our side. Retry with backoff.