For AI agents: step-by-step setup guide — /docs/agents.md, documentation index — /llms.txt.
Models
Gonka network models available through the gateway: the ID for the model field, context window, output limit, price and status.
Models and pricing#
The model lineup changes through Gonka network votes, so the table is built from live gateway data. Prices are in dollars per 1M tokens at the current GNK rate; charges are in GNK — see the Pricing & Account section for details.
| Model | Context | Output | Input, $/1M | Output, $/1M | Status |
|---|---|---|---|---|---|
MiniMax MiniMaxAI/MiniMax-M2.7 | 200K | 8K | $0.0083 | $0.025 | Degraded |
DeepSeek deepseek-ai/DeepSeek-V4-Flash-0731 | 380K | 33K | $0.0083 | $0.025 | Degraded |
Z.ai zai-org/GLM-5.3-Flash | 390K | 8K | $0.0083 | $0.025 | Operational |
Paste the ID into the model field; letter case doesn't matter. Status is based on the share of successfully processed requests, as on the Network status page.
How to choose a model#
- Not sure where to start? Go with
MiniMaxAI/MiniMax-M2.7: the documentation examples and installer commands are built around it. - The longest context —
zai-org/GLM-5.3-Flash: 390K tokens. For large codebases and long documents. - The longest output —
deepseek-ai/DeepSeek-V4-Flash-0731: up to 33K tokens at a time. For generating large files in one go.
If the network stops serving a model, it disappears from GET /v1/models, and requests to it get a 503 response with the error type model_unavailable. The error text will name a model you can switch to.
Output length#
Output length is set by max_tokens — or max_completion_tokens, the gateway understands both fields. Each model has a cap: a larger max_tokens is silently trimmed to it. If max_tokens is not set, with stream: true the cap is used, and without streaming a smaller default applies.
Reasoning models spend part of max_tokens thinking before answering — leave some headroom.
model | Cap | Default, stream: true | Default, no stream |
|---|---|---|---|
MiniMaxAI/MiniMax-M2.7 | 8192 | 8192 | 1500 |
deepseek-ai/DeepSeek-V4-Flash-0731 | 32768 | 32768 | 1500 |
zai-org/GLM-5.3-Flash | 8192 | 8192 | 3000 |
The same numbers in machine-readable form are in the GET /v1/capabilities response, limits.models field.
Default model#
A request without the model field is sent by the gateway to the default model — currently MiniMaxAI/MiniMax-M2.7.
Claude Code and the Anthropic SDK send Anthropic model names (claude-*) in /v1/messages — the gateway replaces them with the same default model.
To work with another model, specify its ID from the table above; in the installer this is the --model flag.
Lineup and status via API#
The model lineup changes through Gonka network votes — don't hardcode the list, fetch it from the API. These requests don't need a key.
| Request | What it returns |
|---|---|
GET /v1/models | Models available right now: ID, context, output limit and token price in dollars. A model the network isn't currently serving won't be in the list. |
GET /v1/models/{model} | A single model in the same format. Pass the slash in the ID as is or as %2F. An unknown or hidden model returns 404 with code model_not_found, and a request to a hidden model gets 503 model_unavailable. |
GET /v1/capabilities | Gateway capabilities: protocols, request parameters and the limits field — key rate limit, timeouts and per-model max_tokens caps. |
GET /v1/network-status | Per-model status and incidents — the same data as on the Network status page. |
GET /health | Whether the gateway itself is responding: {"status":"ok"}. This request doesn't check model health. |