API reference
OpenAI-compatible. One line to switch.
Use the official OpenAI SDKs in any language — just change the base URL and key. Pay in credits from the same balance as the workspace.
Authentication
Create a key in Account → API keys. Send it as a bearer token on every request. Keys start with zy_live_. Each key can have its own monthly credit limit; revoke it at any time.
Authorization: Bearer zy_live_…Never expose keys in client-side code. Proxy requests through your backend.
Quickstart
Base URL: https://zyverna.com/api/v1
curl https://zyverna.com/api/v1/chat/completions \
-H "Authorization: Bearer $ZYVERNA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "zyverna-core",
"messages": [{"role": "user", "content": "Explain tail-call optimization in 3 sentences."}]
}'import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.ZYVERNA_API_KEY,
baseURL: "https://zyverna.com/api/v1",
});
const res = await client.chat.completions.create({
model: "zyverna-core",
messages: [{ role: "user", content: "Write a Zod schema for a User." }],
});
console.log(res.choices[0].message.content);
console.log("credits:", res.zyverna.credits);from openai import OpenAI
client = OpenAI(
api_key=os.environ["ZYVERNA_API_KEY"],
base_url="https://zyverna.com/api/v1",
)
stream = client.chat.completions.create(
model="zyverna-flash",
stream=True,
messages=[{"role": "user", "content": "Summarize this PR description…"}],
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")POST /chat/completions
Accepts the standard OpenAI request body. Supported parameters: model, messages, stream, temperature, top_p, max_tokens / max_completion_tokens, tools, tool_choice, response_format, stop, user. Responses are standard chat.completion objects with an extra zyverna field.
| Field | Type | Description |
|---|---|---|
zyverna.credits | number | Credits charged for this request (also in the X-Zyverna-Credits header). |
model | string | The Zyverna model id you requested (upstream model ids are never exposed). |
usage | object | Prompt, completion and total tokens, including cached prompt tokens when available. |
Streaming
Set stream: true to receive Server-Sent Events. The final chunk before [DONE] carries usage and zyverna.credits.
Tool calling
All models support parallel-safe function calling with JSON Schema parameters.
const res = await client.chat.completions.create({
model: "zyverna-core",
messages: [{ role: "user", content: "What's the weather in Toronto?" }],
tools: [{
type: "function",
function: {
name: "get_weather",
description: "Get current weather for a city",
parameters: {
type: "object",
properties: { city: { type: "string" } },
required: ["city"],
},
},
}],
});
const call = res.choices[0].message.tool_calls?.[0];
// → { function: { name: "get_weather", arguments: '{"city":"Toronto"}' } }Models & pricing
GET /models returns the list below with live pricing. Prices are in credits per 1,000 tokens; 100 credits = 1 USD.
| Model | Tier | Context | Input / 1K tokens | Output / 1K tokens | ≈ USD per 1M out |
|---|---|---|---|---|---|
Zyverna Flash zyverna-flash Fastest. Great for chat, quick edits and autocomplete. | Fast | 1M tokens | 0.05 cr | 0.2 cr | $2.00 |
Zyverna Core zyverna-core Balanced. The default for agent runs and refactors. | Balanced | 1M tokens | 0.25 cr | 1 cr | $10.00 |
Zyverna Ultra zyverna-ultra Deep reasoning for hard bugs, architecture and planning. | Frontier | 200K tokens | 0.5 cr | 2 cr | $20.00 |
100 credits = 1 USD. Tool calls and agent runs are billed at the same per-token rates — there is no surcharge for agent mode. Cached prompt tokens are billed at 50% of the input rate.
Credits & billing
- API usage draws from the same credit balance as the workspace. Included plan credits are used first, then purchased credits.
- Cached prompt tokens are billed at 50% of the input rate.
- Requests that fail with a 4xx error are not charged. Requests cancelled mid-stream are charged for tokens generated.
- When your balance reaches zero, requests return
402 insufficient_credits. Top up in USD, CAD, EUR or GBP.
Rate limits
| Plan | Requests / min | Tokens / min | Concurrent streams |
|---|---|---|---|
| Pay-as-you-go | 60 | 100K | 3 |
| Starter | 300 | 400K | 10 |
| Pro | 1,000 | 2M | 25 |
| Business | 3,000 | 6M | 100 |
| Scale | 10,000 | 20M | 500 (custom on request) |
Limits are returned in X-RateLimit-Limit-Requests, X-RateLimit-Remaining-Requests and Retry-After headers. Exceeding them returns 429.
Errors
Errors use the OpenAI envelope so SDK error handling works unchanged.
{
"error": {
"message": "Invalid API key format.",
"type": "invalid_request_error",
"code": "invalid_api_key",
"param": null
}
}| Status | Code | Meaning |
|---|---|---|
| 400 | invalid_request | Malformed body or unsupported parameter. |
| 401 | invalid_api_key | Missing, malformed or revoked key. |
| 402 | insufficient_credits | Balance is zero or the key's monthly limit was reached. |
| 404 | model_not_found | Unknown model id. |
| 429 | rate_limit_exceeded | Slow down; honour Retry-After. |
| 500 | server_error | Something went wrong on our side. Retry with backoff. |