For AI agents: step-by-step setup guide — /docs/agents.md, documentation index — /llms.txt.

API reference

Everything about requests to the gateway: protocols, endpoints, keys, and parameters. Below — streaming, tool calling, reasoning, plugins, and the cost of a request in the response.

Protocols and endpoints#

The gateway accepts OpenAI and Anthropic formats. The base URL for the OpenAI SDK is https://gate.joingonka.ai/v1, for the Anthropic SDK — https://gate.joingonka.ai. All protocols run on the same key and the same balance: a request in any format follows the same path.

Method and pathFormatUse caseNotes
POST /v1/chat/completionsOpenAI Chat CompletionsChat, agents, tool callingThe primary path: the gateway converts other formats into it.
POST /v1/messagesAnthropic MessagesClaude Code and the Anthropic SDKBase URL without /v1; claude-* models are replaced with the recommended one; the max_tokens field is required.
POST /v1/responsesOpenAI ResponsesCodex CLI and the new OpenAI SDKsStateless: send the full history with every request.
POST /v1/completionsOpenAI Completions (legacy)Autocomplete and code editing in editorssuffix is passed to the model as a hint; the response includes logprobs: null.
POST /v1/embeddingsOpenAI EmbeddingsVector representations of textResponse 501: the network has no embedding models.

Discovery endpoints#

They respond without a key. A table of models with context and status is in the Models section.

Method and pathDescription
GET /v1/modelsModel list: context, prices, supported parameters — fields in OpenRouter format.
GET /v1/models/{model}A single model's card; a slash in the id works as-is or as %2F. A hidden or unknown model returns 404 model_not_found.
GET /v1/capabilitiesGateway capabilities: parameters, protocols, plugins, cost fields, and limits (limits).
GET /v1/pluginsPlugins: id and name.
GET /v1/network-statusStatus of network models: availability, latency, uptime.
GET /v1/nodesNode pool summary: total, active, and quarantined.
GET /v1/web-search/enginesWhether web search is enabled and the state of its engines.

Anthropic Messages#

  • Base URL — https://gate.joingonka.ai: the SDK will append /v1/messages itself.
  • The key goes in the x-api-key header (that's how the Anthropic SDK sends it) or in Authorization: Bearer.
  • claude-* models are replaced by the gateway with the recommended one (MiniMaxAI/MiniMax-M2.7); the model field in the response keeps the name the client sent.
  • max_tokens is required, as in the Anthropic API; above the model's ceiling it gets trimmed.
  • Streaming uses Anthropic events; during pauses the gateway sends event: ping, and a failure arrives as an event: error event.
  • The model's reasoning doesn't appear in the response: there are no thinking blocks.
  • The built-in web_search is executed by the gateway's web search plugin — see the Plugins section.
  • There's no token counting (/v1/messages/count_tokens) — response 404.

The easiest way to set up Claude Code is with the installer — connecting tools. Manually, use environment variables; ANTHROPIC_MODEL pins the network model.

export ANTHROPIC_BASE_URL=https://gate.joingonka.ai
export ANTHROPIC_AUTH_TOKEN=$JOINGONKA_API_KEY
export ANTHROPIC_MODEL=MiniMaxAI/MiniMax-M2.7
claude

OpenAI Responses#

  • The gateway doesn't store responses: send the full history in input. The previous_response_id and conversation fields return error 400 with a code.
  • store is accepted and changes nothing.
  • Tools: function and web_search — the latter is executed by the web search plugin. Other built-in tools are skipped by the gateway and the request runs without them; requiring such a tool via tool_choice returns error 400.
  • The input_image and input_file parts return error 400: network models work with text.
  • State endpoints (GET /v1/responses/{id}, DELETE /v1/responses/{id}, GET /v1/responses/{id}/input_items, POST /v1/responses/{id}/cancel, POST /v1/responses/compact) respond 404 with a code — the gateway doesn't store responses.
  • Codex CLI: set your own provider id in model_provider (not openai) — then Codex compresses history itself, without /v1/responses/compact.

Legacy Completions#

  • prompt is a string or an array of one string; the response is in choices[].text. Multiple prompts or tokens instead of text return error 400.
  • suffix is passed to the model as a hint in the prompt: the network has no true middle-fill.
  • The response includes logprobs: null; best_of is ignored; echo works.

Keys and authorization#

The key is passed in the Authorization: Bearer jg-… or x-api-key: jg-… header — on all endpoints. A key is created after signing up on the gate.joingonka.ai/keys page.

PrefixKeyModel requests
jg-Regular account keyyes
gc-Child key: its own limits, spend comes from the owner's balanceyes
gm-Management key: manages children onlyno — 403 forbidden
  • In the dashboard you can set a spend limit per key for day, month, and total; exceeding it returns 402 child_key_limit_exceeded.
  • The number of requests per minute per key is limited — see the Limits section for values.
  • Only the demo chat on the website works without a key: a request without a key from your own code will get 402 with is_demo.
  • Keys can only be managed in the dashboard: /api/keys with an API key is unavailable. Balance and spend per key — Account API.

Your key is a secret: don't keep it in your repo or frontend code — pass it via environment variables.

Examples#

The same request across four SDKs. The model is the recommended one (MiniMaxAI/MiniMax-M2.7), and the key comes from the environment variable JOINGONKA_API_KEY.

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://gate.joingonka.ai/v1",
    api_key=os.environ["JOINGONKA_API_KEY"],
)

response = client.chat.completions.create(
    model="MiniMaxAI/MiniMax-M2.7",
    messages=[{"role": "user", "content": "What is Gonka?"}],
)
print(response.choices[0].message.content)

Streaming response#

import os
from openai import OpenAI

client = OpenAI(base_url="https://gate.joingonka.ai/v1", api_key=os.environ["JOINGONKA_API_KEY"])

stream = client.chat.completions.create(
    model="MiniMaxAI/MiniMax-M2.7",
    messages=[{"role": "user", "content": "What is Gonka?"}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Request parameters#

The POST /v1/chat/completions parameters that the gateway guarantees itself are listed in the supported_parameters array of the capabilities response:

ParameterDescription
temperatureResponse randomness: the higher the value, the more varied the output.
top_pToken selection by cumulative probability.
top_kSelection from the k most likely tokens.
min_pCuts off unlikely tokens relative to the most likely one.
frequency_penaltyPenalty for frequent repetitions.
presence_penaltyPenalty for tokens already seen.
repetition_penaltyMultiplier against repetitions.
stopStrings at which generation stops.
seedSeed for reproducibility.
max_tokensResponse token limit; anything above the model ceiling is clipped to the ceiling.
max_completion_tokensAnother name for max_tokens: the gateway moves the value into it.
toolsFunctions the model can call, in OpenAI format.
tool_choiceWhether to call a function: model's choice, never, required, or a specific one.
response_formatStructured response: json_object or json_schema.
  • Without temperature, the gateway substitutes 0.7.
  • Without max_tokens, the gateway substitutes the model default: shorter without streaming, the model ceiling when streaming. Per-model numbers are in the Limits section.

Passed through to the network as-is#

reasoning_effort, reasoning, enable_thinking, chat_template_kwargs, thinking_token_budget, min_tokens, logit_bias, n, parallel_tool_calls, extra_body. The gateway doesn't validate them: a value outside the network's list returns a 400 error of type api_error.

Not passed through to the network#

All other fields are accepted by the gateway but not passed to the network — for example, user, metadata, store, logprobs, top_logprobs, thinking, stream_options, web_search_options. usage always arrives in the stream.

Streaming#

  • stream: true is a response of SSE events; the last event is data: [DONE].
  • Before finishing, a chunk with usage arrives — always, even without stream_options.
  • During pauses, the gateway sends the comment : keep-alive every 15 s — SSE clients skip it.
  • Before the stream is opened, a failure comes back as a regular response code; after it's opened, as a chunk with joingonka-error.
  • In delta.tool_calls, one call per chunk: the gateway splits calls that the network glued together.
SSE
data: {"object":"chat.completion.chunk","model":"MiniMaxAI/MiniMax-M2.7","choices":[{"index":0,"delta":{"content":"Hi"},"finish_reason":null}]}

data: {"object":"chat.completion.chunk","model":"MiniMaxAI/MiniMax-M2.7","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: {"object":"chat.completion.chunk","model":"MiniMaxAI/MiniMax-M2.7","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":2,"total_tokens":14}}

data: [DONE]

Gateway service chunks#

You can recognize them by the id field:

idWhen and what's inside
joingonka-errorFailure after the stream opened: the error field, then a break without [DONE].
joingonka-stream-stalledThe network went silent longer than the allowed pause: the stream closes with finish_reason: stop.
joingonka-stream-unfinishedThe network interrupted generation: finish_reason: length — continue with the next request.
joingonka-citationsWeb search sources in delta.annotations — before finishing.
joingonka-metaCost and timings — only with the x-joingonka-meta: 1 header.

Streaming in other protocols#

  • Anthropic Messages: events from message_start to message_stop, event: ping during pauses, failure as event: error.
  • OpenAI Responses: response.* events, failure as response.failed.
  • Legacy Completions: on failure, data: {"error": …}, then [DONE].

Tool calling#

  • OpenAI format: tools and tool_choice. The older functions and function_call format is also accepted — and the response comes back in it too.
  • In streaming, one call per chunk: clients that only read the first element don't lose calls.
  • The gateway repairs history that the network would reject with a 400 error: the developer role becomes system, empty and duplicate call ids get unique ones, an object arguments becomes a JSON string, a missing type is filled in, and a call without a name is removed along with its result.
  • A call the model wrote as markup in the text is moved by the gateway into tool_calls; false calls in a response to a request without tools are removed.
  • Generation broke off mid-arguments — you'll get finish_reason: length, not tool_calls: increase the response limit.
JSON
{
  "model": "MiniMaxAI/MiniMax-M2.7",
  "messages": [{"role": "user", "content": "What is the weather in Paris?"}],
  "tools": [{
    "type": "function",
    "function": {
      "name": "get_weather",
      "description": "Current weather for a city",
      "parameters": {
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"]
      }
    }
  }]
}

JSON-Schema restrictions#

Tool schemas and response_format are compiled by the network into a grammar; regular expressions use the RE2 engine. The gateway normalizes the schema into a form the network accepts:

  • $ref are expanded in place, the $defs and definitions sections are removed; a recursive reference becomes an unconstrained schema.
  • pattern with constructs not supported in RE2 (lookahead and lookbehind, backreferences, atomic groups, possessive quantifiers) is dropped; repetitions above 1000 are reduced to 1000.
  • anyOf and oneOf of constants are collapsed into enum; if there are more than 16 non-collapsible branches, the union is dropped.

The schema may end up looser than the original — validate call arguments on your side.

Structured response#

response_format: {"type": "json_object"} returns valid JSON, {"type": "json_schema", "json_schema": {"name": …, "schema": …}} follows your schema with the restrictions above. Truncated JSON without streaming is fixed by the response-healing plugin.

JSON
{
  "model": "MiniMaxAI/MiniMax-M2.7",
  "messages": [{"role": "user", "content": "Name three planets."}],
  "response_format": {
    "type": "json_schema",
    "json_schema": {
      "name": "planets",
      "schema": {
        "type": "object",
        "properties": {"planets": {"type": "array", "items": {"type": "string"}}},
        "required": ["planets"]
      }
    }
  }
}

Reasoning#

  • The model's reasoning arrives separately from the response: message.reasoning_content, and in streaming, delta.reasoning_content. The gateway renames the reasoning field into this format.
  • Reasoning spends max_tokens: with a small limit, the response cuts off (finish_reason: length) before any text.
  • If there's no response text but there is reasoning, the gateway moves it into content — except for responses with a tool call.
  • reasoning_effort and reasoning.effort are passed to the network. If the model has only two modes, the gateway maps the value to them: none and minimal → low, higher ones → default reasoning.
  • If a node rejects the value, the gateway downgrades it (max and xhigh → high, minimal → low, otherwise drops the field) and retries the request.
  • In /v1/messages, reasoning is not passed — there are no thinking blocks.

Plugins#

Plugins are enabled via the plugins field — an array of strings or objects with options. The list is GET /v1/plugins.

PluginDescriptionConditions
response-healingFixes truncated JSON in the model's response.Only without streaming, and only if the response starts with { or [.
privacy-sanitizationMasks emails, IPv4 addresses, card numbers, JWTs, 64-character hex keys, and keys of the form sk-…, gw_…, gm-…, Bearer … in text messages.The mode is the privacy_mode field: redact (default) or tokenize.
file-parserExtracts text from PDFs.When the message text is entirely a base64 PDF: data:application/pdf;base64,… or without a prefix.
webWeb search: results are mixed into the request, and the response gets source links.Together with privacy-sanitization — a 400 error.
  • Options: max_results — from 1 to 10, default 5; engine — engine hint; search_prompt — custom text before results; enabled: false — disable search.
  • Sources are in message.annotations[].url_citation; in streaming, in a joingonka-citations chunk before finishing.
  • mode: "agent" — the model decides for itself whether and what to search; max_searches — from 1 to 5, default 3.
  • Billing: in standard mode, tokens only (search results count as input tokens); in agent mode, tokens for every step plus 1000 nGNK per search performed (x_joingonka.web_search_surcharge_ngonka).
  • In Anthropic Messages and OpenAI Responses, the built-in web_search tool runs this same plugin in agent mode.
{
  "model": "MiniMaxAI/MiniMax-M2.7",
  "messages": [{"role": "user", "content": "What is new in the Gonka network?"}],
  "plugins": [{"id": "web", "max_results": 5}]
}

Cost and service fields#

A non-streaming response carries the request cost in usage:

FieldDescription
usage.cost_gnkRequest cost in GNK
usage.platform_fee_gnkOf which platform markup, GNK
usage.total_cost_gnkTotal to be charged in GNK
usage.total_cost_usdTotal in dollars at the current GNK rate
  • In a stream, usage contains tokens only; the cost is in the joingonka-meta chunk.
  • With the x-joingonka-meta: 1 header, the POST /v1/chat/completions response gets a x_joingonka block: cost (cost_ngonka), balance after charging (balance_ngonka, non-streaming only) and timings (ttft_ms). Other protocols don't return this block.
  • x-request-id is the request ID: include it when contacting support.
  • Retry-After comes with 429: wait this many seconds before retrying.
  • The X-Title and HTTP-Referer headers (like OpenRouter's) help the gateway identify your app; their contents are not stored.

Limits#

  • Images: image_url parts are replaced with a text placeholder — the model cannot see the picture (vision: false in capabilities).
  • From a browser, the API is accessible only from JoinGonka domains (Origin check): call it from your own server, never put the key in the frontend.
  • Embeddings: POST /v1/embeddings returns 501 — there are no embedding models on the network.
  • Error codes, limits and timeouts are in the Errors & Rate Limits section.