B2B AI API — pricing and reference

One endpoint for Claude, GPT, Gemini, image and video models — below each vendor's official rates (the saving per model is computed on the model pages). No subscription: you top up a wallet and pay as you go at the prices below — text per token (input and output each at its own price per 1M tokens), images per picture, video per second at the requested resolution. Two text formats are supported: Anthropic Messages and OpenAI Chat Completions, so existing SDKs work with a changed base URL (Anthropic SDK: baseURL …/api; OpenAI SDK: baseURL …/api/v1 — see the snippets below).

Live prices

ModelTypePrice
claude-opus-5 — Claude Opus 5textinput $2.20 / output $11.00-56%(input 198 ₽ / output 990 ₽) / 1M tokens
claude-sonnet-5 — Claude Sonnet 5textinput $1.27 / output $6.34-36%(input 114 ₽ / output 570 ₽) / 1M tokens
gpt-5.4 — GPT-5.4textinput $1.10 / output $6.60-56%(input 99 ₽ / output 594 ₽) / 1M tokens
gpt-6-luna — GPT-6 Lunatextinput $0.024 / output $0.122-75%(input 2.2 ₽ / output 11 ₽) / 1M tokens
nano-banana-2 — Nano Banana 2image$0.027-60%(2.4 ₽) / request
nano-banana — Nano Bananaimage$0.011-72%(0.96 ₽) / request
gpt-image-2 — GPT Image 2image$0.027 (2.4 ₽) / request
seedream-5.0-lite — Seedream 5.0 Liteimage$0.053 (4.8 ₽) / request
seedance-2.0-fast — Seedance 2.0 Fastvideofrom $0.118 (from 10.6 ₽) / second of video
seedance-2.0-mini — Seedance 2.0 Minivideofrom $0.095 (from 8.54 ₽) / second of video
kling-v3 — Kling V3video$0.081-35%(7.32 ₽) / second of video
veo-3.1-fast — Veo 3.1 Fastvideo$0.467-22%(42 ₽) / request
veo-3.1-lite — Veo 3.1 Litevideo$0.267-33%(24 ₽) / request
minimax-3-hailuo — MiniMax 3 Hailuovideo$0.166-58%(14.9 ₽) / request
wan/2-7-text-to-video — Wan 2.7 text-to-videovideo$0.160-68%(14.4 ₽) / request

Prices update automatically. Full machine-readable list for every model: GET https://zerocoder.com/api/v1/models

«-NN%» — the saving vs the vendor's official price for the same unit (second of video, image, 1M tokens), official prices checked 4 October 2026. Sources and assumptions are on the model pages: Claude API · ChatGPT API · Nano Banana API · Kling API · Veo API · MiniMax Hailuo API · Wan API

Quick start

  1. Sign in and open the Studio — your API keys live in the account section.
  2. Top up the API wallet (separate from studio credits; minimum first top-up applies).
  3. Call the API — it is drop-in compatible with the Anthropic Messages format:
curl https://zerocoder.com/api/v1/messages \
  -H "x-api-key: YOUR_KEY" \
  -H "content-type: application/json" \
  -d '{"model": "claude-sonnet", "max_tokens": 300,
       "messages": [{"role": "user", "content": "Hello!"}]}'

Model aliases: claude-opus, claude-sonnet, gpt, gpt-luna (→ GPT-6 Luna), smart, nano-banana, gpt-image and more — see the models endpoint above. claude-haiku is retired: the provider removed Claude Haiku, so it returns 400 with a hint — use gpt-luna or claude-sonnet-5.

Get an API key

Authentication

Every request carries the key either in the x-api-key header (Anthropic style) or as Authorization: Bearer (OpenAI style). Keys start with zc-sk-; we store only a SHA-256 hash, so a lost key cannot be recovered — issue a new one in the account section.

A missing or revoked key returns 401 authentication_error; a suspended account returns 403 permission_error. Keys are per account: all keys of one account share the same wallet, while the 30 req/min limit is counted per key (an account can hold up to 10 active keys).

x-api-key: zc-sk-…
# или / or
Authorization: Bearer zc-sk-…

Limits and billing

Rate limit: 30 requests per minute per key, one shared bucket for every generating endpoint (messages, chat completions, images, edits, videos). Above that you get 429 rate_limit_error — back off for a few seconds and retry. GET /v1/models does not count; GET /v1/balance has its own soft bucket of 120 requests per minute per key.

Billing is in rubles from a dedicated API wallet (separate from studio credits). Text models are billed per token, input and output separately: GET /v1/models gives input_price_rub (per 1M input tokens) and output_price_rub (per 1M output tokens, usually 5–8× the input price) with unit "1M_tokens"; price_rub equals the input price and stays for older clients. A call costs input tokens × input price + output tokens × output price, with a minimum of 0.5 ₽ per call (min_rub). Before generation an upper estimate is reserved — the input estimate at the input price plus max_tokens (default 8192, at most 32768 — a larger value is clamped) at the output price; as soon as the answer is ready the actual usage is charged and the difference is refunded, in the same request. If the balance is below the reserve you get 402 billing_error quoting the reserve and your balance — lower max_tokens or top up. The reserve is computed for at least 512 output tokens (the fallback floor may answer with up to that many regardless of a smaller max_tokens). On gateway models the answer is capped at max_tokens (default 8192) — send a larger max_tokens for long answers; when an answer comes through our Claude/ChatGPT subscriptions the output is not hard-capped by max_tokens, and tokens above the reserve are charged after the reply at the output price. Images are billed per picture and video per second × seconds at the requested resolution — the price is charged upfront, as before. Any charge is refunded automatically if the upstream provider fails (you get 502 with “Charge refunded”).

Where you see the charge: the billing block is returned by /v1/chat/completions, /v1/images/edits, vision replies of /v1/messages, the final event billing of a /v1/messages stream and the last chunk of a /v1/chat/completions stream. For text models it is { charged_rub, reserved_rub, balance_rub, unit: "1M_tokens", rub_per_mtok_in, rub_per_mtok_out, rub_per_mtok } — what was actually charged, what was reserved at start, the balance after settlement and the per-1M input and output prices used (rub_per_mtok repeats the input price for older clients); for images it stays { charged_rub, balance_rub }. Non-streaming text replies of /v1/messages and /v1/images/generations carry no billing block (their shape is frozen for existing clients) and /v1/videos reports price_rub / balance_rub — use GET /v1/balance when you need the wallet state.

Worked example with gpt-luna (GPT-6 Luna): input 2.2 ₽ / output 11.02 ₽ per 1M tokens. A call with 1,000 input and 200 output tokens costs 0.0044 ₽ → the 0.5 ₽ minimum applies; 200,000 input + 50,000 output → 0.9915 ₽; 1,000,000 input + 200,000 output → 4.41 ₽. A long answer costs more than a long question: output is 5× the input price. Streams reserve the same way and settle after the last delta, before the billing event; if a stream breaks mid-way you pay only for the tokens actually produced (minimum applies) and the rest of the reserve is refunded.

Before the first call the wallet must be topped up at least once (minimum first top-up below); an empty wallet returns 402 billing_error with the minimum amount in the message. Auto top-up from a saved card is available in the account section.

Minimum first top-up: 1000 ₽

Models

GET/api/v1/models

GET /v1/models returns every model available to your key with its current price. Text models have endpoint /v1/messages (and /v1/chat/completions), image models /v1/images/generations and /v1/images/edits, video models /v1/videos. Aliases (claude-sonnet, gpt-luna, …) and full ids (claude-sonnet-4.6, gpt-6-luna) are both accepted. Retired ids (claude-haiku, and the removed claude-haiku-4.5, gpt-5.4-mini, gpt-5.4-nano) return 400 with the replacement to use.

Units: text models have unit "1M_tokens" — input_price_rub and output_price_rub are rubles per 1M input and output tokens (price_rub = the input price, kept for compatibility) and min_rub is the per-call minimum; image models have unit "request" (price per picture), video models "second" (price per second of video) — models priced per resolution also carry price_rub_by_resolution, and POST /v1/videos bills the resolution it sends to the provider. The other fields (id, aliases, type, endpoint, provider) are unchanged.

{
  "object": "list", "currency": "RUB", "balance_rub": 1840.5,
  "data": [
    { "id": "smart", "aliases": ["auto"], "type": "text", "endpoint": "/v1/messages", "unit": "1M_tokens", "price_rub": null, "provider": "router",
      "description": "Automatic model selection by task complexity (lite/standard/heavy). Billed at the routed model's price; …" },
    { "id": "claude-sonnet", "aliases": ["claude-sonnet-4.6", "claude-custom-default"], "type": "text", "endpoint": "/api/v1/messages", "unit": "1M_tokens", "price_rub": …, "input_price_rub": …, "output_price_rub": …, "min_rub": 0.5, "provider": "Anthropic" },
    { "id": "gpt-6-luna", "aliases": ["gpt-luna", "gpt-nano"], "type": "text", "endpoint": "/api/v1/messages", "unit": "1M_tokens", "price_rub": …, "input_price_rub": …, "output_price_rub": …, "min_rub": 0.5, "provider": "nexus" },
    { "id": "nano-banana-2", "aliases": ["gemini-3-flash"], "type": "image", "endpoint": "/api/v1/images/generations", "unit": "request", "price_rub": …, "provider": "gemini-image" },
    { "id": "seedance-2.0", "aliases": [], "type": "video", "endpoint": "/api/v1/videos", "unit": "second", "price_rub": …, "price_rub_by_resolution": { "480p": …, "720p": …, "1080p": … }, "provider": "nexus" }
  ]
}

Smart routing — model: "smart"

Don't want to pick a model per request? Pass "smart" (alias "auto") and the router chooses one for each call based on task complexity: simple work goes to fast, inexpensive models, hard work goes to the top tier — so you never overpay for easy calls. Direct model ids keep working exactly as before.

The router is a deterministic zero-latency heuristic — no extra cost, no extra network hop, and the same request always lands in the same tier. It scores each request by: input length, code blocks and programming keywords, reasoning/analysis markers, math, long-form writing, strict output formats (JSON/schema), and dialog length; simple mechanical tasks (translate, fix typos, summarize, classify, extract) lower the score. The total maps to one of three tiers:

TierModelTypical tasks
liteGPT-6 Luna (gpt-6-luna)translations, cleanups, extraction, classification, short answers
standardGPT-5.4 (gpt-5.4)everyday content: emails, product copy, posts, summaries with nuance
heavyClaude Opus 5 (claude-opus-5)code, deep analysis, long-form writing, multi-step reasoning
curl https://zerocoder.com/api/v1/messages \
  -H "x-api-key: YOUR_KEY" \
  -H "content-type: application/json" \
  -d '{"model": "smart", "max_tokens": 400,
       "messages": [{"role": "user", "content": "Refactor this function and explain the changes: …"}]}'
{
  "id": "msg_…", "type": "message", "role": "assistant",
  "model": "gpt-6-luna",
  "smart_routing": { "tier": "lite", "score": 1, "signals": ["short", "translate"] },
  "content": [{ "type": "text", "text": "…" }],
  "stop_reason": "end_turn", "stop_sequence": null,
  "usage": { "input_tokens": 12, "output_tokens": 8 }
}

Billing is always at the routed model's price (see the table above). The response tells you exactly what ran: the model field carries the actual model id, and smart_routing carries { tier, score, signals } so you can log and audit routing decisions. Works in both /v1/messages and /v1/chat/completions.

Balance

GET/api/v1/balance

GET /v1/balance — wallet state for the key's account, no rate-limit charge. spent_30d_rub sums the charges of the last 30 days (refunds already netted out). key.requests_count and key.last_used_at describe the key you called with.

curl https://zerocoder.com/api/v1/balance -H "x-api-key: YOUR_KEY"
curl https://zerocoder.com/api/v1/models  -H "x-api-key: YOUR_KEY"
{
  "currency": "RUB",
  "balance_rub": 1840.5,
  "topped_up_total_rub": 5000,
  "spent_30d_rub": 3159.5,
  "min_topup_rub": 1000,
  "auto_topup": { "enabled": true, "amount_rub": 1000, "threshold_rub": 300 },
  "key": { "name": "prod-backend", "requests_count": 4213, "last_used_at": "2026-09-26T10:41:07.000Z" }
}

Text: /v1/messages

POST/api/v1/messages

Anthropic Messages-compatible. Body: model, messages[] (roles user / assistant; content is a string or an array of blocks), optional system, max_tokens, stream. The last user message becomes the prompt, earlier messages are passed as history.

{
  "model": "claude-sonnet",
  "system": "You are a concise assistant.",
  "max_tokens": 300,
  "messages": [
    { "role": "user", "content": "Hi! What's the capital of Portugal?" },
    { "role": "assistant", "content": "Lisbon." },
    { "role": "user", "content": [{ "type": "text", "text": "And its population?" }] }
  ]
}

Response mirrors Anthropic: id msg_…, content[0].text, stop_reason "end_turn", usage with input_tokens / output_tokens. For gateway models usage is the provider's real count; for Claude/ChatGPT proxy models it is an estimate (~4 characters per token). smart_routing appears only when model was "smart". max_tokens caps the answer on gateway models (default 8192, at most 32768 — a larger value is clamped; answers through our Claude/ChatGPT subscriptions are not hard-capped) and sets the size of the token reserve (see Billing), so keep it realistic. usage is what the token bill is computed from.

{
  "id": "msg_1f0c9b1e5a2b4c7d9e8f0a1b2c3d4e5f",
  "type": "message",
  "role": "assistant",
  "model": "claude-sonnet-4.6",
  "content": [{ "type": "text", "text": "About 550 thousand in the city, ~2.9 million in the metro area." }],
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": { "input_tokens": 42, "output_tokens": 19 }
}

Text: /v1/chat/completions

POST/api/v1/chat/completions

OpenAI Chat Completions-compatible — point the OpenAI SDK's baseURL at our /api/v1 (the Anthropic SDK, which appends /v1/messages itself, uses /api instead — snippets below) and use your zc-sk- key. Body: model, messages[] with roles system / user / assistant (content is a string or an array of {type:"text"} / {type:"image_url"} parts), optional stream, max_tokens; temperature is accepted and ignored. system → system prompt, last user message → prompt, everything else → history.

// OpenAI SDK (Node) — only baseURL and the key change
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://zerocoder.com/api/v1", apiKey: "zc-sk-…" });
const r = await client.chat.completions.create({ model: "gpt-luna", messages: [{ role: "user", content: "Hello!" }] });
# OpenAI SDK (Python)
from openai import OpenAI
client = OpenAI(base_url="https://zerocoder.com/api/v1", api_key="zc-sk-…")
r = client.chat.completions.create(model="gpt-luna", messages=[{"role": "user", "content": "Hello!"}])
{
  "model": "gpt-luna",
  "messages": [
    { "role": "system", "content": "Answer in one sentence." },
    { "role": "user", "content": "Why is the sky blue?" }
  ],
  "max_tokens": 200
}
{
  "id": "chatcmpl_8c2f1d0e9a7b4c5d",
  "object": "chat.completion",
  "created": 1790419267,
  "model": "gpt-6-luna",
  "choices": [
    { "index": 0, "message": { "role": "assistant", "content": "Sunlight scatters off air molecules, and blue scatters most." }, "finish_reason": "stop" }
  ],
  "usage": { "prompt_tokens": 21, "completion_tokens": 14, "total_tokens": 35 },
  "billing": { "charged_rub": 0.5, "reserved_rub": 0.5, "balance_rub": 1840.0, "unit": "1M_tokens", "rub_per_mtok_in": …, "rub_per_mtok_out": …, "rub_per_mtok": … }
}

Same key, rate limit, billing, routing (including "smart") and vision as /v1/messages — it is the same engine with a different envelope. Errors use the OpenAI envelope (see Errors).

Streaming

Pass stream: true to either text endpoint and the answer arrives as Server-Sent Events (content-type text/event-stream). /v1/messages emits the Anthropic event sequence — message_start, content_block_start, N × content_block_delta with text_delta, content_block_stop, message_delta with stop_reason and output_tokens, message_stop — plus one extra final event billing with { charged_rub, reserved_rub, balance_rub, unit, rub_per_mtok_in, rub_per_mtok_out, rub_per_mtok } (already settled to the real usage). Official Anthropic SDKs ignore the unknown event.

event: message_start
data: {"type":"message_start","message":{"id":"msg_…","type":"message","role":"assistant","model":"gpt-5.4","content":[],"stop_reason":null,"usage":{"input_tokens":21,"output_tokens":0}}}

event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Sunlight "}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"scatters…"}}

event: content_block_stop
data: {"type":"content_block_stop","index":0}

event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"output_tokens":14}}

event: message_stop
data: {"type":"message_stop"}

event: billing
data: {"charged_rub":0.5,"reserved_rub":0.5,"balance_rub":1840.0,"unit":"1M_tokens","rub_per_mtok_in":…,"rub_per_mtok_out":…,"rub_per_mtok":…}

/v1/chat/completions emits OpenAI chunks: a first role-only chunk {delta:{role:"assistant", content:""}}, then data: {object:"chat.completion.chunk", choices:[{delta:{content}}]} …, a final chunk with finish_reason "stop", usage and billing { charged_rub, reserved_rub, balance_rub, unit, rub_per_mtok_in, rub_per_mtok_out, rub_per_mtok }, then data: [DONE].

curl -N https://zerocoder.com/api/v1/chat/completions \
  -H "authorization: Bearer YOUR_KEY" \
  -H "content-type: application/json" \
  -d '{"model": "gpt", "stream": true,
       "messages": [{"role": "user", "content": "Write a haiku about rain"}]}'
data: {"id":"chatcmpl_…","object":"chat.completion.chunk","created":1790419267,"model":"gpt-6-luna","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}

data: {"id":"chatcmpl_…","object":"chat.completion.chunk","created":1790419267,"model":"gpt-6-luna","choices":[{"index":0,"delta":{"content":"Sunlight "},"finish_reason":null}]}

data: {"id":"chatcmpl_…","object":"chat.completion.chunk","created":1790419267,"model":"gpt-6-luna","choices":[{"index":0,"delta":{"content":"scatters…"},"finish_reason":null}]}

data: {"id":"chatcmpl_…","object":"chat.completion.chunk","created":1790419267,"model":"gpt-6-luna","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":21,"completion_tokens":14,"total_tokens":35},"billing":{"charged_rub":0.5,"reserved_rub":0.5,"balance_rub":1840.0,"unit":"1M_tokens","rub_per_mtok_in":…,"rub_per_mtok_out":…,"rub_per_mtok":…}}

data: [DONE]

Honest note: streaming is real token-by-token only for gateway models (GPT, Gemini and other models served through the gateway). Claude and ChatGPT served through our subscriptions are generated in full first and then delivered at once, split into chunks without artificial delay — the stream is a compatibility format there, not a latency win.

Billing for streams: the price is reserved upfront exactly as for a non-streaming call. A failure before the first delta is refunded and reported as an error event; a failure mid-stream is not refunded (partial text was delivered) and is reported as an error event; a client disconnect aborts the upstream request without a refund. The error event is event: error with the Anthropic envelope on /v1/messages, and a data: {"error":{message,type,code}} chunk followed by data: [DONE] on /v1/chat/completions (the official OpenAI SDK raises APIError on it).

// /v1/messages
event: error
data: {"type":"error","error":{"type":"api_error","message":"Upstream generation failed: … Charge refunded."}}

// /v1/chat/completions
data: {"error":{"message":"Upstream generation failed: … Charge refunded.","type":"api_error","code":"upstream_error"}}

data: [DONE]

Image input (vision, beta)beta

Both text endpoints accept up to 4 images in the user message: Anthropic blocks {type:"image", source:{type:"base64", media_type, data}} and {type:"image", source:{type:"url", url}}, or OpenAI parts {type:"image_url", image_url:{url}} (https URL or data URI). Same rules as for edits: only https URLs on public hosts (private and loopback addresses are rejected), redirects are followed only to public https hosts, and each image is at most 8 MB — otherwise 400 before any charge. The reply is a normal text response with a billing block (per-token, like any text call); usage.input_tokens includes an estimate per image.

{
  "model": "gpt",
  "max_tokens": 300,
  "messages": [{
    "role": "user",
    "content": [
      { "type": "text", "text": "What is on this photo? Answer in Russian." },
      { "type": "image", "source": { "type": "url", "url": "https://example.com/photo.jpg" } },
      { "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "iVBORw0KGgo…" } }
    ]
  }]
}
{ "role": "user", "content": [
  { "type": "text", "text": "Describe the image." },
  { "type": "image_url", "image_url": { "url": "https://example.com/photo.jpg" } }
] }
{
  "id": "msg_7c1e2d3a-4b5c-4d6e-8f90-a1b2c3d4e5f6",
  "type": "message", "role": "assistant",
  "model": "gemini-flash-latest",
  "content": [{ "type": "text", "text": "На фото — рыжий кот на подоконнике…" }],
  "stop_reason": "end_turn",
  "usage": { "input_tokens": 1112, "output_tokens": 18 },
  "billing": { "charged_rub": 0.5, "reserved_rub": 0.5, "balance_rub": 1840.0, "unit": "1M_tokens", "rub_per_mtok_in": …, "rub_per_mtok_out": …, "rub_per_mtok": … }
}

Vision is beta: requests with images are answered by a Gemini vision model (the model field in the response shows what actually ran), with the OpenAI vision model as a fallback when configured. If no vision backend is available you get 502 "Vision backend unavailable" and the charge is refunded. Billing is at the price of the model you requested.

Image generation

POST/api/v1/images/generations

POST /v1/images/generations — OpenAI Images shape. Body: prompt, optional model (default nano-banana), size (1024x1024 / 1792x1024 / 1024x1792) or aspect ("1:1", "16:9", "9:16", "4:3", "3:4"); only n = 1. Response: { created, data:[{ url }] }. The URL is a temporary link — download the file promptly.

{ "model": "nano-banana-2", "prompt": "A watercolor fox in a birch forest", "aspect": "16:9" }
{ "created": 1790419267, "data": [{ "url": "https://…/out.png" }] }

Image editing

POST/api/v1/images/edits

POST /v1/images/edits — edit or combine reference images by a text instruction (NanoBanana edit model by default). JSON body: prompt, optional model, aspect, image (one https URL or data URI) or images[] (up to 4). Or multipart/form-data with fields prompt, model, aspect and file fields image / image[] — files up to 8 MB each, the whole request body up to 48 MB (413 above that). Only https URLs are fetched; private and loopback hosts are rejected with 400. model must be an image model from GET /v1/models (type "image") — video and text ids are rejected with 400.

{
  "model": "nano-banana",
  "prompt": "Replace the background with a sunset beach, keep the person unchanged",
  "images": ["https://example.com/portrait.jpg"],
  "aspect": "3:4"
}
# JSON
curl https://zerocoder.com/api/v1/images/edits \
  -H "x-api-key: YOUR_KEY" -H "content-type: application/json" \
  -d '{"prompt": "Make it a night scene", "image": "https://example.com/photo.jpg", "aspect": "16:9"}'

# multipart
curl https://zerocoder.com/api/v1/images/edits \
  -H "x-api-key: YOUR_KEY" \
  -F prompt="Make it a night scene" -F aspect=16:9 \
  -F "image[]=@photo1.jpg" -F "image[]=@photo2.jpg"

Response: { created, model, data:[{ url }], billing:{ charged_rub, balance_rub } }. The model's price is charged upfront and refunded if the provider fails; if you disconnect before the result arrives the charge is kept (the gateway job has already run), the same rule as for text streams.

Video

POST/api/v1/videos

POST /v1/videos starts a generation: model (kling-v3, seedance-2.0-fast, … — see /v1/models), prompt, optional seconds (4–10, default 5), size or aspect (default 16:9), image (https URL of a reference frame for image-to-video). Response: { id, object:"video", model, status:"queued", seconds, price_rub, balance_rub, size, resolution, billing:{ charged_rub, balance_rub, unit, unit_price_rub, price_option } }. Billing is the per-second price × seconds, charged at start; for models priced per resolution (price_rub_by_resolution in /v1/models) it is the price of the resolution sent to the provider — from resolution or size (default 1920x1080, i.e. 1080p or the nearest tier the model has). A malformed size or resolution is rejected with 400 before any charge.

{ "model": "kling-v3", "prompt": "A drone shot over a foggy pine forest at dawn", "seconds": 5, "aspect": "16:9",
  "image": "https://example.com/first-frame.jpg" }
{ "id": "vid_7a1…", "object": "video", "model": "kling-v3", "status": "queued", "seconds": 5, "price_rub": …, "balance_rub": …,
  "size": "1920x1080", "resolution": null, "billing": { "charged_rub": …, "balance_rub": …, "unit": "second", "unit_price_rub": … }, "poll": "/api/v1/videos?id=vid_7a1…" }

GET/api/v1/videos?id=…

Poll GET /v1/videos?id=… every few seconds: { id, status:"queued"|"processing"|"completed"|"failed", progress, url }. On failure the response carries error and refunded: true — the charge is returned automatically; tasks that never finish are refunded by a sweeper within a couple of hours. When video is unavailable on our side the start returns 503.

{ "id": "vid_7a1…", "status": "processing", "progress": 40, "url": null }
{ "id": "vid_7a1…", "status": "completed", "progress": 100, "url": "https://…/video.mp4" }
{ "id": "vid_7a1…", "status": "failed", "error": "generation failed", "refunded": true }
curl https://zerocoder.com/api/v1/videos -H "x-api-key: YOUR_KEY" -H "content-type: application/json" \
  -d '{"model": "kling-v3", "prompt": "A drone shot over a foggy pine forest", "seconds": 5}'
# → {"id":"vid_…","status":"queued",…}
curl "https://zerocoder.com/api/v1/videos?id=vid_…" -H "x-api-key: YOUR_KEY"

Errors

/v1/messages, /v1/models, /v1/balance, /v1/images/* and /v1/videos use the Anthropic envelope; /v1/chat/completions uses the OpenAI envelope. The message field is always human-readable and safe to log.

StatustypeWhen
400invalid_request_errormalformed JSON, missing messages/prompt, unknown model or a model of the wrong type for the endpoint, unsupported/private image source
401authentication_errormissing, malformed or revoked key
402billing_errorwallet never topped up (message names the minimum) or balance below the call price (for text — below the reserve for max_tokens; the message quotes both)
403permission_errorthe key's account is suspended
429rate_limit_errormore than 30 requests per minute on this key
413invalid_request_errorrequest body over 48 MB (/v1/images/edits)
502api_errorthe upstream provider failed — the charge is refunded (except mid-stream)
503api_errorvideo or image editing temporarily unavailable on our side
// Anthropic envelope — /v1/messages, /v1/models, /v1/balance, /v1/images/*, /v1/videos
{ "type": "error", "error": { "type": "billing_error", "message": "Not enough API balance: this call reserves up to 4.7522 ₽ (≈2 input tokens at 118.8 ₽ and up to 8000 output tokens at 594 ₽ per 1M tokens; the unused part is refunded after the reply), you have 0.2 ₽." } }

// OpenAI envelope — /v1/chat/completions
{ "error": { "message": "Rate limit exceeded (30 requests/minute).", "type": "rate_limit_error", "code": "rate_limit_exceeded" } }

OpenAI envelope: error.code mirrors the status so SDK code can switch on it — 400 invalid_request, 401 invalid_api_key, 402 insufficient_balance, 403 account_suspended, 429 rate_limit_exceeded, 502 upstream_error; error.type is the same string as in the Anthropic envelope.

Code samples

Official SDKs

The OpenAI SDK appends /chat/completions to its baseURL, the Anthropic SDK appends /v1/messages — so they need different bases:

// Anthropic SDK (Node) — baseURL WITHOUT /v1: the SDK appends /v1/messages itself
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({ baseURL: "https://zerocoder.com/api", apiKey: "zc-sk-…" });
const msg = await client.messages.create({ model: "claude-sonnet", max_tokens: 300, messages: [{ role: "user", content: "Hello!" }] });
# Anthropic SDK (Python)
from anthropic import Anthropic
client = Anthropic(base_url="https://zerocoder.com/api", api_key="zc-sk-…")
msg = client.messages.create(model="claude-sonnet", max_tokens=300, messages=[{"role": "user", "content": "Hello!"}])
// OpenAI SDK (Node) — only baseURL and the key change
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://zerocoder.com/api/v1", apiKey: "zc-sk-…" });
const r = await client.chat.completions.create({ model: "gpt-luna", messages: [{ role: "user", content: "Hello!" }] });
# OpenAI SDK (Python)
from openai import OpenAI
client = OpenAI(base_url="https://zerocoder.com/api/v1", api_key="zc-sk-…")
r = client.chat.completions.create(model="gpt-luna", messages=[{"role": "user", "content": "Hello!"}])

curl

curl https://zerocoder.com/api/v1/messages \
  -H "x-api-key: YOUR_KEY" \
  -H "content-type: application/json" \
  -d '{"model": "claude-sonnet", "max_tokens": 300,
       "messages": [{"role": "user", "content": "Hello!"}]}'
curl -N https://zerocoder.com/api/v1/chat/completions \
  -H "authorization: Bearer YOUR_KEY" \
  -H "content-type: application/json" \
  -d '{"model": "gpt", "stream": true,
       "messages": [{"role": "user", "content": "Write a haiku about rain"}]}'

Python (requests)

import json, requests

API = "https://zerocoder.com/api/v1"
H = {"x-api-key": "YOUR_KEY", "content-type": "application/json"}

# 1. Chat Completions (non-streaming)
r = requests.post(f"{API}/chat/completions", headers=H, json={
    "model": "gpt-luna",
    "messages": [{"role": "system", "content": "Answer briefly."},
                 {"role": "user", "content": "Why is the sky blue?"}],
})
r.raise_for_status()
data = r.json()
print(data["choices"][0]["message"]["content"], data["billing"])

# 2. Messages with streaming (Anthropic SSE)
with requests.post(f"{API}/messages", headers=H, stream=True, json={
    "model": "gpt", "stream": True, "max_tokens": 400,
    "messages": [{"role": "user", "content": "Tell a short story about a lighthouse"}],
}) as s:
    s.raise_for_status()
    event = None
    for line in s.iter_lines(decode_unicode=True):
        if line.startswith("event: "):
            event = line[7:]
        elif line.startswith("data: "):
            payload = json.loads(line[6:])
            if event == "content_block_delta":
                print(payload["delta"]["text"], end="", flush=True)
            elif event == "billing":
                print("\n", payload)  # {'charged_rub': …, 'reserved_rub': …, 'balance_rub': …, 'unit': '1M_tokens', 'rub_per_mtok_in': …, 'rub_per_mtok_out': …, 'rub_per_mtok': …}
            elif event == "error":
                raise RuntimeError(payload["error"]["message"])

# 3. Image edit (multipart)
with open("photo.jpg", "rb") as f:
    r = requests.post(f"{API}/images/edits", headers={"x-api-key": "YOUR_KEY"},
                      data={"prompt": "Make it a night scene", "aspect": "16:9"},
                      files={"image": f})
print(r.json()["data"][0]["url"])

Node (fetch)

const API = "https://zerocoder.com/api/v1";
const H = { "x-api-key": process.env.ZC_API_KEY, "content-type": "application/json" };

// 1. Messages (non-streaming)
const r = await fetch(`${API}/messages`, {
  method: "POST", headers: H,
  body: JSON.stringify({ model: "claude-sonnet", max_tokens: 300,
    messages: [{ role: "user", content: "Hello!" }] }),
});
if (!r.ok) throw new Error((await r.json()).error.message);
const msg = await r.json();
console.log(msg.content[0].text); // non-streaming /v1/messages carries no billing block — see GET /v1/balance

// 2. Chat Completions with streaming (OpenAI SSE)
const s = await fetch(`${API}/chat/completions`, {
  method: "POST", headers: H,
  body: JSON.stringify({ model: "gpt", stream: true,
    messages: [{ role: "user", content: "Write a haiku about rain" }] }),
});
if (!s.ok) throw new Error((await s.json()).error.message);
const reader = s.body.getReader();
const dec = new TextDecoder();
let buf = "";
for (;;) {
  const { value, done } = await reader.read();
  if (done) break;
  buf += dec.decode(value, { stream: true });
  const lines = buf.split("\n"); buf = lines.pop() ?? "";
  for (const line of lines) {
    if (!line.startsWith("data: ")) continue;
    const data = line.slice(6);
    if (data === "[DONE]") break;
    const chunk = JSON.parse(data);
    if (chunk.error) throw new Error(chunk.error.message);
    process.stdout.write(chunk.choices[0].delta.content ?? "");
  }
}

// 3. Video: start, then poll
const v = await (await fetch(`${API}/videos`, { method: "POST", headers: H,
  body: JSON.stringify({ model: "kling-v3", prompt: "A drone shot over a foggy forest", seconds: 5 }) })).json();
let st;
do {
  await new Promise((res) => setTimeout(res, 5000));
  st = await (await fetch(`${API}/videos?id=${v.id}`, { headers: H })).json();
} while (st.status !== "completed" && st.status !== "failed");
console.log(st.url ?? st.error);

What's new (06.10.2026)

  • 06.10.2026: output tokens are billed at their own price — /v1/models gained input_price_rub and output_price_rub (price_rub stays = input price), the billing block gained rub_per_mtok_in and rub_per_mtok_out, and the reserve takes max_tokens at the output price. Before, input and output were billed together at the input price. Gateway answers are now capped at the max_tokens the reserve covers (default 8192, at most 32768).
  • 06.10.2026: claude-haiku is retired (the provider removed Claude Haiku) and returns 400 with a hint; the new alias gpt-luna points at GPT-6 Luna, gpt-nano → GPT-6 Luna, gpt-mini → GPT-5.6 Luna.
  • 06.10.2026: POST /v1/videos bills the resolution sent to the provider (price_rub_by_resolution in /v1/models) and echoes size, resolution and billing in the response.
  • 26.09.2026: text is billed per token with a 0.5 ₽ minimum per call; an upper estimate (by max_tokens) is reserved at start and settled to the real usage right after the answer. The billing block gained reserved_rub, unit and rub_per_mtok; /v1/models gained unit "1M_tokens" and min_rub for text models.
  • Streaming (stream: true) on /v1/messages — Anthropic SSE with an extra billing event.
  • New endpoint POST /v1/chat/completions — OpenAI Chat Completions format, including streaming; the OpenAI SDK works with baseURL …/api/v1.
  • New endpoint GET /v1/balance — wallet, 30-day spend, auto top-up settings and key usage.
  • New endpoint POST /v1/images/edits — editing by reference images (JSON or multipart), up to 4 images.
  • Vision (beta) now runs on Gemini first with an OpenAI fallback; when no backend is available the charge is refunded.
  • Both text endpoints share one engine: identical auth, rate limit, billing, routing and error codes.

FAQ

Why is it cheaper than going direct?

We pool the volume of many customers into one contract, so prices are below the vendors' official list prices. The saving for each model is shown on its page.

Is there a minimum contract?

No. Wallet top-up, pay as you go (text per token, images per picture, video per second). The wallet is separate from studio credits so your bill reconciles in one currency.

What about rate limits and uptime?

Requests fail over between providers where the model allows it; your key's usage and spend are visible in the account section.

Do you store my prompts?

Requests are proxied for billing and abuse control; we can sign an NDA and discuss data handling for enterprise volumes.

Need volume pricing or an integration done for you?

Book a free consultation