For AI agents: step-by-step setup guide — /docs/agents.md, documentation index — /llms.txt.
Errors and limits
Rate limits, timeouts, and gateway error codes. For each response, we explain when it occurs and what to do: retry the request or fix it.
Limits#
The numbers come from the limits field of the GET /v1/capabilities response: values there are always current.
| Limit | Value | When exceeded |
|---|---|---|
| Requests per minute per key | 120, window 60 s from the first request | 429 rate_limit_exceeded and the Retry-After header; gateway 5xx failures don't consume quota |
| Concurrent account requests | capped | the excess waits in a queue; if it times out — 429 queue_timeout |
| Request body size | 16 MiB | 413 |
| Output length | by model — see the table below | above the model's cap — trimmed to the cap, no error |
| Child keys | up to 50 per managing key, up to 120 requests per minute each | to raise the caps — contact support |
Output length by model#
Without max_tokens, the gateway applies a default: without streaming — shorter, so the response fits within the timeouts; in a stream — the model's cap. max_completion_tokens is the same field.
model | Cap | No stream | In a stream |
|---|---|---|---|
MiniMaxAI/MiniMax-M2.7 | 8192 | 1500 | 8192 |
deepseek-ai/DeepSeek-V4-Flash-0731 | 32768 | 1500 | 32768 |
zai-org/GLM-5.3-Flash | 8192 | 3000 | 8192 |
Timeouts#
| Stage | Value | What happens |
|---|---|---|
| Waiting for a queue slot | 45 s | 429 queue_timeout with the Retry-After: 1 header |
| Start of network response, stream | 150 s | 504 upstream_timeout; an estimate of the input tokens is charged |
| Start of network response, no stream | 150 s | 504 upstream_timeout; an estimate of the input tokens is charged |
| Response generation | ≈ 300 s | the network cuts generation short: the response arrives with finish_reason: length — continue with the next request |
| Pause between stream chunks | 30 s | the stream closes: finish_reason: stop in the joingonka-stream-stalled chunk |
| Keepalive signal in the stream | 15 s | a : keep-alive comment — SSE clients skip it |
| Stream opening | 30 s | until this moment a failure comes as a response code, after — as the joingonka-error chunk |
| Early non-streaming response | 90 s | the gateway returns 200 and sends spaces every 15 s — the JSON stays valid; an error after that comes in the body with the error field, the status remains 200 |
What is charged on timeout
Once the network accepts a request, it cannot be cancelled. On 504 upstream_timeout, the input token estimate is charged, but not the output; a retry is a new charge. A stream cut off before the final usage is billed the same way. For long responses, use stream: true.
Error codes#
The error body is an error object with message, type, code, param; not every error has all fields. Rely on the status and type, and check code for details: the message text may change. The Anthropic format is in the Anthropic error format section.
{
"error": {
"message": "Model is currently overloaded in the Gonka network",
"type": "rate_limit_exceeded",
"code": "upstream_rate_limited"
}
}| Response | When | What to do |
|---|---|---|
400 invalid_request_error | Invalid body: no messages, a message isn't an object, the body isn't JSON; unknown model — with param: model and a list of available models in the text | Fix the request based on the error text |
400 invalid_request_error empty_content_after_normalization | The message is empty after normalization — for example, it contained only an image | Add text to the message |
400 invalid_request_error web_search_privacy_sanitization_not_supported | The web and privacy-sanitization plugins in one request | Keep only one of them |
400 invalid_request_error previous_response_id_not_supported conversation_not_supported item_reference_not_supported background_not_supported hosted_tool_choice_not_supported | OpenAI Responses: a reference to a stored response, conversation, or item; background mode; a built-in tool requirement | Send the full history in input |
400 api_error | The network rejected the parameters — for example, an reasoning_effort value outside its list | Fix the value based on the error text |
401 authentication_error | The key wasn't found, was revoked, or is in an unknown format | Check the key on the gate.joingonka.ai/keys page |
402 insufficient_funds | Insufficient balance to cover the request estimate; the remainder is in balance_ngonka | Top up your balance: gate.joingonka.ai/billing |
402 insufficient_funds | A request without a key, not from the website (is_demo: true) | Pass an API key |
402 child_key_limit_exceeded | The key's spending limit has been exceeded — daily, monthly, or total; the remainders are in daily_remaining, monthly_remaining, total_remaining | Raise the key's limit on the gate.joingonka.ai/keys page or wait for the reset |
403 forbidden | A management key gm- in a model request; an API key on a dashboard-only route | For requests, use the jg- or gc- key; for account management, use the dashboard |
404 invalid_request_error model_not_found | GET /v1/models/{model}: the model isn't in the catalog or is temporarily hidden | Take an id from GET /v1/models |
404 invalid_request_error not_found compact_not_supported | OpenAI Responses: /v1/responses/{id} and other state addresses, /v1/responses/compact | Keep the history on your side; for Codex CLI, set your own provider id |
404 invalid_request_error | Unknown path | Check the method, path, and base URL |
413 invalid_request_error | The request body exceeds the limit | Shorten the request |
415 invalid_request_error | The body isn't JSON according to the header: Content-Type: application/json is required | Send Content-Type: application/json |
429 rate_limit_exceeded | The per-key requests-per-minute limit has been exceeded; the body contains rate_limit with limit, remaining, reset | Wait for the number of seconds in Retry-After |
429 rate_limit_exceeded upstream_rate_limited | The model is overloaded on the Gonka network | Retry after Retry-After or pick another model — Models |
429 rate_limit_exceeded queue_timeout queue_full | All network slots are taken: the queue is full or the wait timed out | Retry after Retry-After |
500 server_error | Internal gateway error | Retry later; if it keeps happening, contact support and include x-request-id |
501 not_implemented | POST /v1/embeddings: there are no embedding models on the network | Use a different embedding service |
502 api_error upstream_unauthorized | The network provider rejected the gateway's credentials — your key is fine | Retry in a minute |
502 api_error | Gonka network error; code comes from the network if it sent one | Retry after a pause or pick another model |
503 model_unavailable model_outage model_initializing model_unstable model_not_served | The model is currently unavailable according to network probes: a crash, startup, instability, or nobody is serving it; rejected immediately, without waiting | Pick another model — the error text will suggest which; the list is at Models |
503 service_unavailable | No available nodes | Retry later |
504 timeout upstream_timeout | The network accepted the request but didn't respond in time; the input token estimate has been charged | For long responses, use stream: true; a retry is a new charge |
Errors in an open stream#
Until the stream is open, a rejection comes as a regular response code — just like without streaming. Once it's open, the status is already 200, and the error arrives like this:
- Chat Completions — a
joingonka-errorwith aerror, then a disconnect without[DONE]. - Without streaming after an early response — status
200and a body with aerror. - Anthropic Messages — an
event: error, then the stream closes. - OpenAI Responses — an
response.failed, with the reason inresponse.error.code. - Legacy Completions —
data: {"error": …}, then[DONE].
data: {"id":"joingonka-error","object":"chat.completion.chunk","model":"MiniMaxAI/MiniMax-M2.7","choices":[],"error":{"message":"Gonka network error","type":"api_error"}}Anthropic error format#
POST /v1/messagesresponds with the Anthropic envelope:{"type": "error", "error": {"type", "message"}}.- Gateway rejections preserve the
typefrom the table above (insufficient_funds,model_unavailable, and others); there's nocodefield in this envelope — the reason is in the text. - Request shape errors —
invalid_request_error: no messages ormax_tokens, a tool call without a name, unknown model. - Unknown path —
not_found_error, body too large —request_too_large.
event: error
data: {"type":"error","error":{"type":"timeout","message":"Upstream timeout"}}What to retry#
- After the pause from
Retry-After:429 - With increasing pause — 1, 2, 4 s, and so on:
500,502,503 service_unavailable,504 - With another model:
503 model_unavailable - Don't retry without changes — fix the request, key, or balance:
400,401,402,403,404,413,415,501
Every retry after 504 is a new input estimate charge; for long responses, enable stream: true.
If the error keeps happening, contact support and include the x-request-id from the response headers.