> For AI agents: step-by-step setup guide — [`/docs/agents.md`](https://gate.joingonka.ai/docs/agents.md), documentation index — [`/llms.txt`](https://gate.joingonka.ai/llms.txt).

# API reference

Everything about requests to the gateway: protocols, endpoints, keys, and parameters. Below — streaming, tool calling, reasoning, plugins, and the cost of a request in the response.

## Protocols and endpoints

The gateway accepts OpenAI and Anthropic formats. The base URL for the OpenAI SDK is `https://gate.joingonka.ai/v1`, for the Anthropic SDK — `https://gate.joingonka.ai`. All protocols run on the same key and the same balance: a request in any format follows the same path.

| Method and path | Format | Use case | Notes |
| --- | --- | --- | --- |
| `POST /v1/chat/completions` | OpenAI Chat Completions | Chat, agents, tool calling | The primary path: the gateway converts other formats into it. |
| `POST /v1/messages` | Anthropic Messages | Claude Code and the Anthropic SDK | Base URL without `/v1`; `claude-*` models are replaced with the recommended one; the `max_tokens` field is required. |
| `POST /v1/responses` | OpenAI Responses | Codex CLI and the new OpenAI SDKs | Stateless: send the full history with every request. |
| `POST /v1/completions` | OpenAI Completions (legacy) | Autocomplete and code editing in editors | `suffix` is passed to the model as a hint; the response includes `logprobs: null`. |
| `POST /v1/embeddings` | OpenAI Embeddings | Vector representations of text | Response `501`: the network has no embedding models. |

### Discovery endpoints

They respond without a key. A table of models with context and status is in the [Models](https://gate.joingonka.ai/docs/models) section.

| Method and path | Description |
| --- | --- |
| `GET /v1/models` | Model list: context, prices, supported parameters — fields in OpenRouter format. |
| `GET /v1/models/{model}` | A single model's card; a slash in the id works as-is or as `%2F`. A hidden or unknown model returns `404 model_not_found`. |
| `GET /v1/capabilities` | Gateway capabilities: parameters, protocols, plugins, cost fields, and limits (`limits`). |
| `GET /v1/plugins` | Plugins: id and name. |
| `GET /v1/network-status` | Status of network models: availability, latency, uptime. |
| `GET /v1/nodes` | Node pool summary: total, active, and quarantined. |
| `GET /v1/web-search/engines` | Whether web search is enabled and the state of its engines. |

### Anthropic Messages

- Base URL — `https://gate.joingonka.ai`: the SDK will append `/v1/messages` itself.
- The key goes in the `x-api-key` header (that's how the Anthropic SDK sends it) or in `Authorization: Bearer`.
- `claude-*` models are replaced by the gateway with the recommended one (`MiniMaxAI/MiniMax-M2.7`); the `model` field in the response keeps the name the client sent.
- `max_tokens` is required, as in the Anthropic API; above the model's ceiling it gets trimmed.
- Streaming uses Anthropic events; during pauses the gateway sends `event: ping`, and a failure arrives as an `event: error` event.
- The model's reasoning doesn't appear in the response: there are no `thinking` blocks.
- The built-in `web_search` is executed by the gateway's web search plugin — see the [Plugins](https://gate.joingonka.ai/docs/api#plugins) section.
- There's no token counting (`/v1/messages/count_tokens`) — response `404`.

The easiest way to set up Claude Code is with the installer — [connecting tools](https://gate.joingonka.ai/docs#connect). Manually, use environment variables; `ANTHROPIC_MODEL` pins the network model.

#### Claude Code

```bash
export ANTHROPIC_BASE_URL=https://gate.joingonka.ai
export ANTHROPIC_AUTH_TOKEN=$JOINGONKA_API_KEY
export ANTHROPIC_MODEL=MiniMaxAI/MiniMax-M2.7
claude
```

#### cURL

```bash
curl https://gate.joingonka.ai/v1/messages \
  -H "x-api-key: $JOINGONKA_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MiniMaxAI/MiniMax-M2.7",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "What is Gonka?"}]
  }'
```

### OpenAI Responses

- The gateway doesn't store responses: send the full history in `input`. The `previous_response_id` and `conversation` fields return error `400` with a code.
- `store` is accepted and changes nothing.
- Tools: `function` and `web_search` — the latter is executed by the web search plugin. Other built-in tools are skipped by the gateway and the request runs without them; requiring such a tool via `tool_choice` returns error `400`.
- The `input_image` and `input_file` parts return error `400`: network models work with text.
- State endpoints (`GET /v1/responses/{id}`, `DELETE /v1/responses/{id}`, `GET /v1/responses/{id}/input_items`, `POST /v1/responses/{id}/cancel`, `POST /v1/responses/compact`) respond `404` with a code — the gateway doesn't store responses.
- Codex CLI: set your own provider id in `model_provider` (not openai) — then Codex compresses history itself, without `/v1/responses/compact`.

### Legacy Completions

- `prompt` is a string or an array of one string; the response is in `choices[].text`. Multiple prompts or tokens instead of text return error `400`.
- `suffix` is passed to the model as a hint in the prompt: the network has no true middle-fill.
- The response includes `logprobs: null`; `best_of` is ignored; `echo` works.

## Keys and authorization

The key is passed in the `Authorization: Bearer jg-…` or `x-api-key: jg-…` header — on all endpoints. A key is created after signing up on the [gate.joingonka.ai/keys](https://gate.joingonka.ai/keys) page.

| Prefix | Key | Model requests |
| --- | --- | --- |
| `jg-` | Regular account key | yes |
| `gc-` | Child key: its own limits, spend comes from the owner's balance | yes |
| `gm-` | Management key: manages children only | no — `403 forbidden` |

- In the dashboard you can set a spend limit per key for day, month, and total; exceeding it returns `402 child_key_limit_exceeded`.
- The number of requests per minute per key is limited — see the [Limits](https://gate.joingonka.ai/docs/errors#limits) section for values.
- Only the demo chat on the website works without a key: a request without a key from your own code will get `402` with `is_demo`.
- Keys can only be managed in the dashboard: `/api/keys` with an API key is unavailable. Balance and spend per key — [Account API](https://gate.joingonka.ai/docs/billing#account-api).

> Your key is a secret: don't keep it in your repo or frontend code — pass it via environment variables.

## Examples

The same request across four SDKs. The model is the recommended one (`MiniMaxAI/MiniMax-M2.7`), and the key comes from the environment variable `JOINGONKA_API_KEY`.

### Python

```python
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://gate.joingonka.ai/v1",
    api_key=os.environ["JOINGONKA_API_KEY"],
)

response = client.chat.completions.create(
    model="MiniMaxAI/MiniMax-M2.7",
    messages=[{"role": "user", "content": "What is Gonka?"}],
)
print(response.choices[0].message.content)
```

### TypeScript

```typescript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://gate.joingonka.ai/v1",
  apiKey: process.env.JOINGONKA_API_KEY,
});

const response = await client.chat.completions.create({
  model: "MiniMaxAI/MiniMax-M2.7",
  messages: [{ role: "user", content: "What is Gonka?" }],
});
console.log(response.choices[0].message.content);
```

### cURL

```bash
curl https://gate.joingonka.ai/v1/chat/completions \
  -H "Authorization: Bearer $JOINGONKA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MiniMaxAI/MiniMax-M2.7",
    "messages": [{"role": "user", "content": "What is Gonka?"}]
  }'
```

### Anthropic SDK

```python
import os
import anthropic

client = anthropic.Anthropic(
    base_url="https://gate.joingonka.ai",
    api_key=os.environ["JOINGONKA_API_KEY"],
)

message = client.messages.create(
    model="MiniMaxAI/MiniMax-M2.7",
    max_tokens=1024,
    messages=[{"role": "user", "content": "What is Gonka?"}],
)
print(message.content[0].text)
```

### Streaming response

#### Python

```python
import os
from openai import OpenAI

client = OpenAI(base_url="https://gate.joingonka.ai/v1", api_key=os.environ["JOINGONKA_API_KEY"])

stream = client.chat.completions.create(
    model="MiniMaxAI/MiniMax-M2.7",
    messages=[{"role": "user", "content": "What is Gonka?"}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
```

#### TypeScript

```typescript
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://gate.joingonka.ai/v1", apiKey: process.env.JOINGONKA_API_KEY });

const stream = await client.chat.completions.create({
  model: "MiniMaxAI/MiniMax-M2.7",
  messages: [{ role: "user", content: "What is Gonka?" }],
  stream: true,
});
for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
```

#### cURL

```bash
curl -N https://gate.joingonka.ai/v1/chat/completions \
  -H "Authorization: Bearer $JOINGONKA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MiniMaxAI/MiniMax-M2.7",
    "messages": [{"role": "user", "content": "What is Gonka?"}],
    "stream": true
  }'
```

#### Anthropic SDK

```python
import os
import anthropic

client = anthropic.Anthropic(base_url="https://gate.joingonka.ai", api_key=os.environ["JOINGONKA_API_KEY"])

with client.messages.stream(
    model="MiniMaxAI/MiniMax-M2.7",
    max_tokens=1024,
    messages=[{"role": "user", "content": "What is Gonka?"}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
```

## Request parameters

The `POST /v1/chat/completions` parameters that the gateway guarantees itself are listed in the `supported_parameters` array of the capabilities response:

| Parameter | Description |
| --- | --- |
| `temperature` | Response randomness: the higher the value, the more varied the output. |
| `top_p` | Token selection by cumulative probability. |
| `top_k` | Selection from the k most likely tokens. |
| `min_p` | Cuts off unlikely tokens relative to the most likely one. |
| `frequency_penalty` | Penalty for frequent repetitions. |
| `presence_penalty` | Penalty for tokens already seen. |
| `repetition_penalty` | Multiplier against repetitions. |
| `stop` | Strings at which generation stops. |
| `seed` | Seed for reproducibility. |
| `max_tokens` | Response token limit; anything above the model ceiling is clipped to the ceiling. |
| `max_completion_tokens` | Another name for `max_tokens`: the gateway moves the value into it. |
| `tools` | Functions the model can call, in OpenAI format. |
| `tool_choice` | Whether to call a function: model's choice, never, required, or a specific one. |
| `response_format` | Structured response: `json_object` or `json_schema`. |

- Without `temperature`, the gateway substitutes `0.7`.
- Without `max_tokens`, the gateway substitutes the model default: shorter without streaming, the model ceiling when streaming. Per-model numbers are in the [Limits](https://gate.joingonka.ai/docs/errors#limits) section.

### Passed through to the network as-is

`reasoning_effort`, `reasoning`, `enable_thinking`, `chat_template_kwargs`, `thinking_token_budget`, `min_tokens`, `logit_bias`, `n`, `parallel_tool_calls`, `extra_body`. The gateway doesn't validate them: a value outside the network's list returns a `400` error of type `api_error`.

### Not passed through to the network

All other fields are accepted by the gateway but not passed to the network — for example, `user`, `metadata`, `store`, `logprobs`, `top_logprobs`, `thinking`, `stream_options`, `web_search_options`. `usage` always arrives in the stream.

## Streaming

- `stream: true` is a response of SSE events; the last event is `data: [DONE]`.
- Before finishing, a chunk with `usage` arrives — always, even without `stream_options`.
- During pauses, the gateway sends the comment `: keep-alive` every 15 s — SSE clients skip it.
- Before the stream is opened, a failure comes back as a regular response code; after it's opened, as a chunk with `joingonka-error`.
- In `delta.tool_calls`, one call per chunk: the gateway splits calls that the network glued together.

```text
data: {"object":"chat.completion.chunk","model":"MiniMaxAI/MiniMax-M2.7","choices":[{"index":0,"delta":{"content":"Hi"},"finish_reason":null}]}

data: {"object":"chat.completion.chunk","model":"MiniMaxAI/MiniMax-M2.7","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: {"object":"chat.completion.chunk","model":"MiniMaxAI/MiniMax-M2.7","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":2,"total_tokens":14}}

data: [DONE]
```

### Gateway service chunks

You can recognize them by the `id` field:

| `id` | When and what's inside |
| --- | --- |
| `joingonka-error` | Failure after the stream opened: the `error` field, then a break without `[DONE]`. |
| `joingonka-stream-stalled` | The network went silent longer than the allowed pause: the stream closes with `finish_reason: stop`. |
| `joingonka-stream-unfinished` | The network interrupted generation: `finish_reason: length` — continue with the next request. |
| `joingonka-citations` | Web search sources in `delta.annotations` — before finishing. |
| `joingonka-meta` | Cost and timings — only with the `x-joingonka-meta: 1` header. |

### Streaming in other protocols

- Anthropic Messages: events from `message_start` to `message_stop`, `event: ping` during pauses, failure as `event: error`.
- OpenAI Responses: `response.*` events, failure as `response.failed`.
- Legacy Completions: on failure, `data: {"error": …}`, then `[DONE]`.

## Tool calling

- OpenAI format: `tools` and `tool_choice`. The older `functions` and `function_call` format is also accepted — and the response comes back in it too.
- In streaming, one call per chunk: clients that only read the first element don't lose calls.
- The gateway repairs history that the network would reject with a `400` error: the `developer` role becomes `system`, empty and duplicate call ids get unique ones, an object `arguments` becomes a JSON string, a missing `type` is filled in, and a call without a name is removed along with its result.
- A call the model wrote as markup in the text is moved by the gateway into `tool_calls`; false calls in a response to a request without tools are removed.
- Generation broke off mid-arguments — you'll get `finish_reason: length`, not `tool_calls`: increase the response limit.

```json
{
  "model": "MiniMaxAI/MiniMax-M2.7",
  "messages": [{"role": "user", "content": "What is the weather in Paris?"}],
  "tools": [{
    "type": "function",
    "function": {
      "name": "get_weather",
      "description": "Current weather for a city",
      "parameters": {
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"]
      }
    }
  }]
}
```

### JSON-Schema restrictions

Tool schemas and `response_format` are compiled by the network into a grammar; regular expressions use the RE2 engine. The gateway normalizes the schema into a form the network accepts:

- `$ref` are expanded in place, the `$defs` and `definitions` sections are removed; a recursive reference becomes an unconstrained schema.
- `pattern` with constructs not supported in RE2 (lookahead and lookbehind, backreferences, atomic groups, possessive quantifiers) is dropped; repetitions above 1000 are reduced to 1000.
- `anyOf` and `oneOf` of constants are collapsed into `enum`; if there are more than 16 non-collapsible branches, the union is dropped.

> The schema may end up looser than the original — validate call arguments on your side.

## Structured response

`response_format`: `{"type": "json_object"}` returns valid JSON, `{"type": "json_schema", "json_schema": {"name": …, "schema": …}}` follows your schema with the restrictions above. Truncated JSON without streaming is fixed by the `response-healing` plugin.

```json
{
  "model": "MiniMaxAI/MiniMax-M2.7",
  "messages": [{"role": "user", "content": "Name three planets."}],
  "response_format": {
    "type": "json_schema",
    "json_schema": {
      "name": "planets",
      "schema": {
        "type": "object",
        "properties": {"planets": {"type": "array", "items": {"type": "string"}}},
        "required": ["planets"]
      }
    }
  }
}
```

## Reasoning

- The model's reasoning arrives separately from the response: `message.reasoning_content`, and in streaming, `delta.reasoning_content`. The gateway renames the `reasoning` field into this format.
- Reasoning spends `max_tokens`: with a small limit, the response cuts off (`finish_reason: length`) before any text.
- If there's no response text but there is reasoning, the gateway moves it into `content` — except for responses with a tool call.
- `reasoning_effort` and `reasoning.effort` are passed to the network. If the model has only two modes, the gateway maps the value to them: `none` and `minimal` → `low`, higher ones → default reasoning.
- If a node rejects the value, the gateway downgrades it (`max` and `xhigh` → `high`, `minimal` → `low`, otherwise drops the field) and retries the request.
- In `/v1/messages`, reasoning is not passed — there are no `thinking` blocks.

## Plugins

Plugins are enabled via the `plugins` field — an array of strings or objects with options. The list is `GET /v1/plugins`.

| Plugin | Description | Conditions |
| --- | --- | --- |
| `response-healing` | Fixes truncated JSON in the model's response. | Only without streaming, and only if the response starts with `{` or `[`. |
| `privacy-sanitization` | Masks emails, IPv4 addresses, card numbers, JWTs, 64-character hex keys, and keys of the form `sk-…`, `gw_…`, `gm-…`, `Bearer …` in text messages. | The mode is the `privacy_mode` field: `redact` (default) or `tokenize`. |
| `file-parser` | Extracts text from PDFs. | When the message text is entirely a base64 PDF: `data:application/pdf;base64,…` or without a prefix. |
| `web` | Web search: results are mixed into the request, and the response gets source links. | Together with `privacy-sanitization` — a `400` error. |

### Web search

- Options: `max_results` — from 1 to 10, default 5; `engine` — engine hint; `search_prompt` — custom text before results; `enabled: false` — disable search.
- Sources are in `message.annotations[].url_citation`; in streaming, in a `joingonka-citations` chunk before finishing.
- `mode: "agent"` — the model decides for itself whether and what to search; `max_searches` — from 1 to 5, default 3.
- Billing: in standard mode, tokens only (search results count as input tokens); in agent mode, tokens for every step plus 1000 nGNK per search performed (`x_joingonka.web_search_surcharge_ngonka`).
- In Anthropic Messages and OpenAI Responses, the built-in `web_search` tool runs this same plugin in agent mode.

#### plugins: web

```json
{
  "model": "MiniMaxAI/MiniMax-M2.7",
  "messages": [{"role": "user", "content": "What is new in the Gonka network?"}],
  "plugins": [{"id": "web", "max_results": 5}]
}
```

#### mode: agent

```json
{
  "model": "MiniMaxAI/MiniMax-M2.7",
  "messages": [{"role": "user", "content": "What is new in the Gonka network?"}],
  "plugins": [{"id": "web", "mode": "agent", "max_searches": 3}]
}
```

## Cost and service fields

A non-streaming response carries the request cost in `usage`:

| Field | Description |
| --- | --- |
| `usage.cost_gnk` | Request cost in GNK |
| `usage.platform_fee_gnk` | Of which platform markup, GNK |
| `usage.total_cost_gnk` | Total to be charged in GNK |
| `usage.total_cost_usd` | Total in dollars at the current GNK rate |

- In a stream, `usage` contains tokens only; the cost is in the `joingonka-meta` chunk.
- With the `x-joingonka-meta: 1` header, the `POST /v1/chat/completions` response gets a `x_joingonka` block: cost (`cost_ngonka`), balance after charging (`balance_ngonka`, non-streaming only) and timings (`ttft_ms`). Other protocols don't return this block.
- `x-request-id` is the request ID: include it when contacting support.
- `Retry-After` comes with `429`: wait this many seconds before retrying.
- The `X-Title` and `HTTP-Referer` headers (like OpenRouter's) help the gateway identify your app; their contents are not stored.

## Limits

- Images: `image_url` parts are replaced with a text placeholder — the model cannot see the picture (`vision: false` in capabilities).
- From a browser, the API is accessible only from JoinGonka domains (`Origin` check): call it from your own server, never put the key in the frontend.
- Embeddings: `POST /v1/embeddings` returns `501` — there are no embedding models on the network.
- Error codes, limits and timeouts are in the [Errors & Rate Limits](https://gate.joingonka.ai/docs/errors) section.
