> For AI agents: step-by-step setup guide — [`/docs/agents.md`](https://gate.joingonka.ai/docs/agents.md), documentation index — [`/llms.txt`](https://gate.joingonka.ai/llms.txt).

# Models

Gonka network models available through the gateway: the ID for the `model` field, context window, output limit, price and status.

## Models and pricing

The model lineup changes through Gonka network votes, so the table is built from live gateway data. Prices are in dollars per 1M tokens at the current GNK rate; charges are in GNK — see the [Pricing & Account](https://gate.joingonka.ai/docs/billing) section for details.

| Model | Context | Output | Input, `$/1M` | Output, `$/1M` | Status |
| --- | ---: | ---: | ---: | ---: | --- |
| **MiniMax** `MiniMaxAI/MiniMax-M2.7` | 200K | 8K | $0.0083 | $0.025 | Degraded |
| **DeepSeek** `deepseek-ai/DeepSeek-V4-Flash-0731` | 380K | 33K | $0.0083 | $0.025 | Degraded |
| **Z.ai** `zai-org/GLM-5.3-Flash` | 390K | 8K | $0.0083 | $0.025 | Operational |

Paste the ID into the `model` field; letter case doesn't matter. Status is based on the share of successfully processed requests, as on the [Network status](https://gate.joingonka.ai/status) page.

## How to choose a model

- Not sure where to start? Go with `MiniMaxAI/MiniMax-M2.7`: the documentation examples and installer commands are built around it.
- The longest context — `zai-org/GLM-5.3-Flash`: 390K tokens. For large codebases and long documents.
- The longest output — `deepseek-ai/DeepSeek-V4-Flash-0731`: up to 33K tokens at a time. For generating large files in one go.

If the network stops serving a model, it disappears from `GET /v1/models`, and requests to it get a 503 response with the error type `model_unavailable`. The error text will name a model you can switch to.

## Output length

Output length is set by `max_tokens` — or `max_completion_tokens`, the gateway understands both fields. Each model has a cap: a larger `max_tokens` is silently trimmed to it. If `max_tokens` is not set, with `stream: true` the cap is used, and without streaming a smaller default applies.

Reasoning models spend part of `max_tokens` thinking before answering — leave some headroom.

| `model` | Cap | Default, `stream: true` | Default, no stream |
| --- | ---: | ---: | ---: |
| `MiniMaxAI/MiniMax-M2.7` | 8192 | 8192 | 1500 |
| `deepseek-ai/DeepSeek-V4-Flash-0731` | 32768 | 32768 | 1500 |
| `zai-org/GLM-5.3-Flash` | 8192 | 8192 | 3000 |

The same numbers in machine-readable form are in the `GET /v1/capabilities` response, `limits.models` field.

## Default model

A request without the `model` field is sent by the gateway to the default model — currently `MiniMaxAI/MiniMax-M2.7`.

Claude Code and the Anthropic SDK send Anthropic model names (`claude-*`) in `/v1/messages` — the gateway replaces them with the same default model.

To work with another model, specify its ID from the table above; in the installer this is the `--model` flag.

## Lineup and status via API

The model lineup changes through Gonka network votes — don't hardcode the list, fetch it from the API. These requests don't need a key.

| Request | What it returns |
| --- | --- |
| `GET /v1/models` | Models available right now: ID, context, output limit and token price in dollars. A model the network isn't currently serving won't be in the list. |
| `GET /v1/models/{model}` | A single model in the same format. Pass the slash in the ID as is or as `%2F`. An unknown or hidden model returns `404` with code `model_not_found`, and a request to a hidden model gets `503 model_unavailable`. |
| `GET /v1/capabilities` | Gateway capabilities: protocols, request parameters and the `limits` field — key rate limit, timeouts and per-model `max_tokens` caps. |
| `GET /v1/network-status` | Per-model status and incidents — the same data as on the [Network status](https://gate.joingonka.ai/status) page. |
| `GET /health` | Whether the gateway itself is responding: `{"status":"ok"}`. This request doesn't check model health. |
