For AI agents: step-by-step setup guide — /docs/agents.md, documentation index — /llms.txt.
API reference
Everything about requests to the gateway: protocols, endpoints, keys, and parameters. Below — streaming, tool calling, reasoning, plugins, and the cost of a request in the response.
Protocols and endpoints#
The gateway accepts OpenAI and Anthropic formats. The base URL for the OpenAI SDK is https://gate.joingonka.ai/v1, for the Anthropic SDK — https://gate.joingonka.ai. All protocols run on the same key and the same balance: a request in any format follows the same path.
| Method and path | Format | Use case | Notes |
|---|---|---|---|
POST /v1/chat/completions | OpenAI Chat Completions | Chat, agents, tool calling | The primary path: the gateway converts other formats into it. |
POST /v1/messages | Anthropic Messages | Claude Code and the Anthropic SDK | Base URL without /v1; claude-* models are replaced with the recommended one; the max_tokens field is required. |
POST /v1/responses | OpenAI Responses | Codex CLI and the new OpenAI SDKs | Stateless: send the full history with every request. |
POST /v1/completions | OpenAI Completions (legacy) | Autocomplete and code editing in editors | suffix is passed to the model as a hint; the response includes logprobs: null. |
POST /v1/embeddings | OpenAI Embeddings | Vector representations of text | Response 501: the network has no embedding models. |
Discovery endpoints#
They respond without a key. A table of models with context and status is in the Models section.
| Method and path | Description |
|---|---|
GET /v1/models | Model list: context, prices, supported parameters — fields in OpenRouter format. |
GET /v1/models/{model} | A single model's card; a slash in the id works as-is or as %2F. A hidden or unknown model returns 404 model_not_found. |
GET /v1/capabilities | Gateway capabilities: parameters, protocols, plugins, cost fields, and limits (limits). |
GET /v1/plugins | Plugins: id and name. |
GET /v1/network-status | Status of network models: availability, latency, uptime. |
GET /v1/nodes | Node pool summary: total, active, and quarantined. |
GET /v1/web-search/engines | Whether web search is enabled and the state of its engines. |
Anthropic Messages#
- Base URL —
https://gate.joingonka.ai: the SDK will append/v1/messagesitself. - The key goes in the
x-api-keyheader (that's how the Anthropic SDK sends it) or inAuthorization: Bearer. claude-*models are replaced by the gateway with the recommended one (MiniMaxAI/MiniMax-M2.7); themodelfield in the response keeps the name the client sent.max_tokensis required, as in the Anthropic API; above the model's ceiling it gets trimmed.- Streaming uses Anthropic events; during pauses the gateway sends
event: ping, and a failure arrives as anevent: errorevent. - The model's reasoning doesn't appear in the response: there are no
thinkingblocks. - The built-in
web_searchis executed by the gateway's web search plugin — see the Plugins section. - There's no token counting (
/v1/messages/count_tokens) — response404.
The easiest way to set up Claude Code is with the installer — connecting tools. Manually, use environment variables; ANTHROPIC_MODEL pins the network model.
export ANTHROPIC_BASE_URL=https://gate.joingonka.ai
export ANTHROPIC_AUTH_TOKEN=$JOINGONKA_API_KEY
export ANTHROPIC_MODEL=MiniMaxAI/MiniMax-M2.7
claudecurl https://gate.joingonka.ai/v1/messages \
-H "x-api-key: $JOINGONKA_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "MiniMaxAI/MiniMax-M2.7",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "What is Gonka?"}]
}'OpenAI Responses#
- The gateway doesn't store responses: send the full history in
input. Theprevious_response_idandconversationfields return error400with a code. storeis accepted and changes nothing.- Tools:
functionandweb_search— the latter is executed by the web search plugin. Other built-in tools are skipped by the gateway and the request runs without them; requiring such a tool viatool_choicereturns error400. - The
input_imageandinput_fileparts return error400: network models work with text. - State endpoints (
GET /v1/responses/{id},DELETE /v1/responses/{id},GET /v1/responses/{id}/input_items,POST /v1/responses/{id}/cancel,POST /v1/responses/compact) respond404with a code — the gateway doesn't store responses. - Codex CLI: set your own provider id in
model_provider(not openai) — then Codex compresses history itself, without/v1/responses/compact.
Legacy Completions#
promptis a string or an array of one string; the response is inchoices[].text. Multiple prompts or tokens instead of text return error400.suffixis passed to the model as a hint in the prompt: the network has no true middle-fill.- The response includes
logprobs: null;best_ofis ignored;echoworks.
Keys and authorization#
The key is passed in the Authorization: Bearer jg-… or x-api-key: jg-… header — on all endpoints. A key is created after signing up on the gate.joingonka.ai/keys page.
| Prefix | Key | Model requests |
|---|---|---|
jg- | Regular account key | yes |
gc- | Child key: its own limits, spend comes from the owner's balance | yes |
gm- | Management key: manages children only | no — 403 forbidden |
- In the dashboard you can set a spend limit per key for day, month, and total; exceeding it returns
402 child_key_limit_exceeded. - The number of requests per minute per key is limited — see the Limits section for values.
- Only the demo chat on the website works without a key: a request without a key from your own code will get
402withis_demo. - Keys can only be managed in the dashboard:
/api/keyswith an API key is unavailable. Balance and spend per key — Account API.
Your key is a secret: don't keep it in your repo or frontend code — pass it via environment variables.
Examples#
The same request across four SDKs. The model is the recommended one (MiniMaxAI/MiniMax-M2.7), and the key comes from the environment variable JOINGONKA_API_KEY.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://gate.joingonka.ai/v1",
api_key=os.environ["JOINGONKA_API_KEY"],
)
response = client.chat.completions.create(
model="MiniMaxAI/MiniMax-M2.7",
messages=[{"role": "user", "content": "What is Gonka?"}],
)
print(response.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://gate.joingonka.ai/v1",
apiKey: process.env.JOINGONKA_API_KEY,
});
const response = await client.chat.completions.create({
model: "MiniMaxAI/MiniMax-M2.7",
messages: [{ role: "user", content: "What is Gonka?" }],
});
console.log(response.choices[0].message.content);curl https://gate.joingonka.ai/v1/chat/completions \
-H "Authorization: Bearer $JOINGONKA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "MiniMaxAI/MiniMax-M2.7",
"messages": [{"role": "user", "content": "What is Gonka?"}]
}'import os
import anthropic
client = anthropic.Anthropic(
base_url="https://gate.joingonka.ai",
api_key=os.environ["JOINGONKA_API_KEY"],
)
message = client.messages.create(
model="MiniMaxAI/MiniMax-M2.7",
max_tokens=1024,
messages=[{"role": "user", "content": "What is Gonka?"}],
)
print(message.content[0].text)Streaming response#
import os
from openai import OpenAI
client = OpenAI(base_url="https://gate.joingonka.ai/v1", api_key=os.environ["JOINGONKA_API_KEY"])
stream = client.chat.completions.create(
model="MiniMaxAI/MiniMax-M2.7",
messages=[{"role": "user", "content": "What is Gonka?"}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://gate.joingonka.ai/v1", apiKey: process.env.JOINGONKA_API_KEY });
const stream = await client.chat.completions.create({
model: "MiniMaxAI/MiniMax-M2.7",
messages: [{ role: "user", content: "What is Gonka?" }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}curl -N https://gate.joingonka.ai/v1/chat/completions \
-H "Authorization: Bearer $JOINGONKA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "MiniMaxAI/MiniMax-M2.7",
"messages": [{"role": "user", "content": "What is Gonka?"}],
"stream": true
}'import os
import anthropic
client = anthropic.Anthropic(base_url="https://gate.joingonka.ai", api_key=os.environ["JOINGONKA_API_KEY"])
with client.messages.stream(
model="MiniMaxAI/MiniMax-M2.7",
max_tokens=1024,
messages=[{"role": "user", "content": "What is Gonka?"}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)Request parameters#
The POST /v1/chat/completions parameters that the gateway guarantees itself are listed in the supported_parameters array of the capabilities response:
| Parameter | Description |
|---|---|
temperature | Response randomness: the higher the value, the more varied the output. |
top_p | Token selection by cumulative probability. |
top_k | Selection from the k most likely tokens. |
min_p | Cuts off unlikely tokens relative to the most likely one. |
frequency_penalty | Penalty for frequent repetitions. |
presence_penalty | Penalty for tokens already seen. |
repetition_penalty | Multiplier against repetitions. |
stop | Strings at which generation stops. |
seed | Seed for reproducibility. |
max_tokens | Response token limit; anything above the model ceiling is clipped to the ceiling. |
max_completion_tokens | Another name for max_tokens: the gateway moves the value into it. |
tools | Functions the model can call, in OpenAI format. |
tool_choice | Whether to call a function: model's choice, never, required, or a specific one. |
response_format | Structured response: json_object or json_schema. |
- Without
temperature, the gateway substitutes0.7. - Without
max_tokens, the gateway substitutes the model default: shorter without streaming, the model ceiling when streaming. Per-model numbers are in the Limits section.
Passed through to the network as-is#
reasoning_effort, reasoning, enable_thinking, chat_template_kwargs, thinking_token_budget, min_tokens, logit_bias, n, parallel_tool_calls, extra_body. The gateway doesn't validate them: a value outside the network's list returns a 400 error of type api_error.
Not passed through to the network#
All other fields are accepted by the gateway but not passed to the network — for example, user, metadata, store, logprobs, top_logprobs, thinking, stream_options, web_search_options. usage always arrives in the stream.
Streaming#
stream: trueis a response of SSE events; the last event isdata: [DONE].- Before finishing, a chunk with
usagearrives — always, even withoutstream_options. - During pauses, the gateway sends the comment
: keep-aliveevery 15 s — SSE clients skip it. - Before the stream is opened, a failure comes back as a regular response code; after it's opened, as a chunk with
joingonka-error. - In
delta.tool_calls, one call per chunk: the gateway splits calls that the network glued together.
data: {"object":"chat.completion.chunk","model":"MiniMaxAI/MiniMax-M2.7","choices":[{"index":0,"delta":{"content":"Hi"},"finish_reason":null}]}
data: {"object":"chat.completion.chunk","model":"MiniMaxAI/MiniMax-M2.7","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"object":"chat.completion.chunk","model":"MiniMaxAI/MiniMax-M2.7","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":2,"total_tokens":14}}
data: [DONE]Gateway service chunks#
You can recognize them by the id field:
id | When and what's inside |
|---|---|
joingonka-error | Failure after the stream opened: the error field, then a break without [DONE]. |
joingonka-stream-stalled | The network went silent longer than the allowed pause: the stream closes with finish_reason: stop. |
joingonka-stream-unfinished | The network interrupted generation: finish_reason: length — continue with the next request. |
joingonka-citations | Web search sources in delta.annotations — before finishing. |
joingonka-meta | Cost and timings — only with the x-joingonka-meta: 1 header. |
Streaming in other protocols#
- Anthropic Messages: events from
message_starttomessage_stop,event: pingduring pauses, failure asevent: error. - OpenAI Responses:
response.*events, failure asresponse.failed. - Legacy Completions: on failure,
data: {"error": …}, then[DONE].
Tool calling#
- OpenAI format:
toolsandtool_choice. The olderfunctionsandfunction_callformat is also accepted — and the response comes back in it too. - In streaming, one call per chunk: clients that only read the first element don't lose calls.
- The gateway repairs history that the network would reject with a
400error: thedeveloperrole becomessystem, empty and duplicate call ids get unique ones, an objectargumentsbecomes a JSON string, a missingtypeis filled in, and a call without a name is removed along with its result. - A call the model wrote as markup in the text is moved by the gateway into
tool_calls; false calls in a response to a request without tools are removed. - Generation broke off mid-arguments — you'll get
finish_reason: length, nottool_calls: increase the response limit.
{
"model": "MiniMaxAI/MiniMax-M2.7",
"messages": [{"role": "user", "content": "What is the weather in Paris?"}],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}]
}JSON-Schema restrictions#
Tool schemas and response_format are compiled by the network into a grammar; regular expressions use the RE2 engine. The gateway normalizes the schema into a form the network accepts:
$refare expanded in place, the$defsanddefinitionssections are removed; a recursive reference becomes an unconstrained schema.patternwith constructs not supported in RE2 (lookahead and lookbehind, backreferences, atomic groups, possessive quantifiers) is dropped; repetitions above 1000 are reduced to 1000.anyOfandoneOfof constants are collapsed intoenum; if there are more than 16 non-collapsible branches, the union is dropped.
The schema may end up looser than the original — validate call arguments on your side.
Structured response#
response_format: {"type": "json_object"} returns valid JSON, {"type": "json_schema", "json_schema": {"name": …, "schema": …}} follows your schema with the restrictions above. Truncated JSON without streaming is fixed by the response-healing plugin.
{
"model": "MiniMaxAI/MiniMax-M2.7",
"messages": [{"role": "user", "content": "Name three planets."}],
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "planets",
"schema": {
"type": "object",
"properties": {"planets": {"type": "array", "items": {"type": "string"}}},
"required": ["planets"]
}
}
}
}Reasoning#
- The model's reasoning arrives separately from the response:
message.reasoning_content, and in streaming,delta.reasoning_content. The gateway renames thereasoningfield into this format. - Reasoning spends
max_tokens: with a small limit, the response cuts off (finish_reason: length) before any text. - If there's no response text but there is reasoning, the gateway moves it into
content— except for responses with a tool call. reasoning_effortandreasoning.effortare passed to the network. If the model has only two modes, the gateway maps the value to them:noneandminimal→low, higher ones → default reasoning.- If a node rejects the value, the gateway downgrades it (
maxandxhigh→high,minimal→low, otherwise drops the field) and retries the request. - In
/v1/messages, reasoning is not passed — there are nothinkingblocks.
Plugins#
Plugins are enabled via the plugins field — an array of strings or objects with options. The list is GET /v1/plugins.
| Plugin | Description | Conditions |
|---|---|---|
response-healing | Fixes truncated JSON in the model's response. | Only without streaming, and only if the response starts with { or [. |
privacy-sanitization | Masks emails, IPv4 addresses, card numbers, JWTs, 64-character hex keys, and keys of the form sk-…, gw_…, gm-…, Bearer … in text messages. | The mode is the privacy_mode field: redact (default) or tokenize. |
file-parser | Extracts text from PDFs. | When the message text is entirely a base64 PDF: data:application/pdf;base64,… or without a prefix. |
web | Web search: results are mixed into the request, and the response gets source links. | Together with privacy-sanitization — a 400 error. |
Web search#
- Options:
max_results— from 1 to 10, default 5;engine— engine hint;search_prompt— custom text before results;enabled: false— disable search. - Sources are in
message.annotations[].url_citation; in streaming, in ajoingonka-citationschunk before finishing. mode: "agent"— the model decides for itself whether and what to search;max_searches— from 1 to 5, default 3.- Billing: in standard mode, tokens only (search results count as input tokens); in agent mode, tokens for every step plus 1000 nGNK per search performed (
x_joingonka.web_search_surcharge_ngonka). - In Anthropic Messages and OpenAI Responses, the built-in
web_searchtool runs this same plugin in agent mode.
{
"model": "MiniMaxAI/MiniMax-M2.7",
"messages": [{"role": "user", "content": "What is new in the Gonka network?"}],
"plugins": [{"id": "web", "max_results": 5}]
}{
"model": "MiniMaxAI/MiniMax-M2.7",
"messages": [{"role": "user", "content": "What is new in the Gonka network?"}],
"plugins": [{"id": "web", "mode": "agent", "max_searches": 3}]
}Cost and service fields#
A non-streaming response carries the request cost in usage:
| Field | Description |
|---|---|
usage.cost_gnk | Request cost in GNK |
usage.platform_fee_gnk | Of which platform markup, GNK |
usage.total_cost_gnk | Total to be charged in GNK |
usage.total_cost_usd | Total in dollars at the current GNK rate |
- In a stream,
usagecontains tokens only; the cost is in thejoingonka-metachunk. - With the
x-joingonka-meta: 1header, thePOST /v1/chat/completionsresponse gets ax_joingonkablock: cost (cost_ngonka), balance after charging (balance_ngonka, non-streaming only) and timings (ttft_ms). Other protocols don't return this block. x-request-idis the request ID: include it when contacting support.Retry-Aftercomes with429: wait this many seconds before retrying.- The
X-TitleandHTTP-Refererheaders (like OpenRouter's) help the gateway identify your app; their contents are not stored.
Limits#
- Images:
image_urlparts are replaced with a text placeholder — the model cannot see the picture (vision: falsein capabilities). - From a browser, the API is accessible only from JoinGonka domains (
Origincheck): call it from your own server, never put the key in the frontend. - Embeddings:
POST /v1/embeddingsreturns501— there are no embedding models on the network. - Error codes, limits and timeouts are in the Errors & Rate Limits section.