For AI agents: step-by-step setup guide — /docs/agents.md, documentation index — /llms.txt.

Models

Gonka network models available through the gateway: the ID for the model field, context window, output limit, price and status.

Models and pricing#

The model lineup changes through Gonka network votes, so the table is built from live gateway data. Prices are in dollars per 1M tokens at the current GNK rate; charges are in GNK — see the Pricing & Account section for details.

ModelContextOutputInput, $/1MOutput, $/1MStatus
MiniMax MiniMaxAI/MiniMax-M2.7200K8K$0.0083$0.025Degraded
DeepSeek deepseek-ai/DeepSeek-V4-Flash-0731380K33K$0.0083$0.025Degraded
Z.ai zai-org/GLM-5.3-Flash390K8K$0.0083$0.025Operational

Paste the ID into the model field; letter case doesn't matter. Status is based on the share of successfully processed requests, as on the Network status page.

How to choose a model#

  • Not sure where to start? Go with MiniMaxAI/MiniMax-M2.7: the documentation examples and installer commands are built around it.
  • The longest context — zai-org/GLM-5.3-Flash: 390K tokens. For large codebases and long documents.
  • The longest output — deepseek-ai/DeepSeek-V4-Flash-0731: up to 33K tokens at a time. For generating large files in one go.

If the network stops serving a model, it disappears from GET /v1/models, and requests to it get a 503 response with the error type model_unavailable. The error text will name a model you can switch to.

Output length#

Output length is set by max_tokens — or max_completion_tokens, the gateway understands both fields. Each model has a cap: a larger max_tokens is silently trimmed to it. If max_tokens is not set, with stream: true the cap is used, and without streaming a smaller default applies.

Reasoning models spend part of max_tokens thinking before answering — leave some headroom.

modelCapDefault, stream: trueDefault, no stream
MiniMaxAI/MiniMax-M2.7819281921500
deepseek-ai/DeepSeek-V4-Flash-073132768327681500
zai-org/GLM-5.3-Flash819281923000

The same numbers in machine-readable form are in the GET /v1/capabilities response, limits.models field.

Default model#

A request without the model field is sent by the gateway to the default model — currently MiniMaxAI/MiniMax-M2.7.

Claude Code and the Anthropic SDK send Anthropic model names (claude-*) in /v1/messages — the gateway replaces them with the same default model.

To work with another model, specify its ID from the table above; in the installer this is the --model flag.

Lineup and status via API#

The model lineup changes through Gonka network votes — don't hardcode the list, fetch it from the API. These requests don't need a key.

RequestWhat it returns
GET /v1/modelsModels available right now: ID, context, output limit and token price in dollars. A model the network isn't currently serving won't be in the list.
GET /v1/models/{model}A single model in the same format. Pass the slash in the ID as is or as %2F. An unknown or hidden model returns 404 with code model_not_found, and a request to a hidden model gets 503 model_unavailable.
GET /v1/capabilitiesGateway capabilities: protocols, request parameters and the limits field — key rate limit, timeouts and per-model max_tokens caps.
GET /v1/network-statusPer-model status and incidents — the same data as on the Network status page.
GET /healthWhether the gateway itself is responding: {"status":"ok"}. This request doesn't check model health.