Updates and Status
Current inference state and gateway changelog.
Current Status
- MiniMax M2.7Gonka Network: Degraded Performance
- DeepSeek V4 FlashGonka Network: Degraded Performance
- GLM 5.3 FlashGonka Network: Operational
The set of active models is determined by Gonka network voting and may vary: some models may be unavailable at times and return later. The current list is above.
Detailed status and uptime →Changelog
You can top up your balance with WGNK — this is GNK on the Ethereum network. In your account: Top up → GNK → Ethereum · WGNK; specify the address from which you will send the tokens — the transfer will be credited automatically, usually in 15–20 minutes.
npx @joingonka/setup now configures 25 tools. New additions: Codex CLI, Kimi Code, omp, MiniMax Code, Goose, MiMo Code, Qwen Code, Factory Droid, Crush, and Zoo Code. The model is selected via the --model flag (minimax, deepseek, or glm), defaulting to DeepSeek V4 Flash. At the end, the installer sends a verification request and immediately shows whether the key is working.
From September 16, every new account receives 100,000,000 nGNK (0.1 GNK) instead of 50,000,000 — at the current network price, this is about 3 million tokens for testing without deposit. The bonus is credited automatically upon registration, no input is needed; the current amount is always visible in GET /api/pricing (field welcome_bonus_ngonka) and on the registration page.
Gonka network brokers terminate any stream after approximately 5 minutes of generation (for GLM 5.3 Flash, this is about 8,000 tokens). Previously, such a cut-off appeared as a normal completion: the stream ended with [DONE] without a finish_reason, and the client mistook the truncated output for a complete response. Now, the gateway appends a chunk with finish_reason="length" in both streams and regular responses, and stop_reason="max_tokens" in Anthropic-format. What to do with long responses: use stream:true, and when finish_reason="length" occurs, continue generation with the next request by passing the previously received part in the context. Details are in the "Limits" section of the documentation.
GLM 5.3 Flash from Z.ai has arrived in the Gonka network: 320B parameters (MoE, 18B active), FP8, MIT license, 390,000 context window. The model reasons before answering: reasoning arrives in the reasoning_content field and consumes the response budget, so set max_tokens to at least 600 or use stream:true; for short tasks — reasoning_effort: "low". Tool calling and JSON mode are supported. Identifier — zai-org/GLM-5.3-Flash, price is the same as for other models in the network. Kimi K2.6 is no longer served by the network — switch integrations to MiniMax M2.7, DeepSeek V4 Flash, or GLM 5.3 Flash.
The "Usage" page displays consumption for each model: requests, input/output tokens, cost in nGNK and USD, average price per 1M tokens, and response speed. "Today" is shown by hour, and a "90 days" preset has been added. For integrations: GET /api/usage/by-model.
The gateway now accepts requests to /v1/completions — the OpenAI "text" endpoint: it takes a prompt string as input and returns choices[].text as output. This is used by editor plugins for autocomplete and inline code editing (Continue.dev and similar). The Base URL remains: https://gate.joingonka.ai/v1. Streaming and standard parameters (max_tokens, temperature, stop) are supported, with the same costs and limits as /v1/chat/completions. logprobs and best_of fields are not supported.
As of August 22, the inference cost is: input — 33 nGNK, output — 99 nGNK per token (formerly 22/66). The gateway fee remains unchanged at 10%. Demo mode remains free. Why the increase: a request in the Gonka network is processed with redundancy — this is part of the protocol and devshards code, necessary for smooth service. A container can send a request to one host, then to another if the first one fails, or to several at once if necessary. Every attempt is paid for: the network price is 10 nGNK per token, multiplied by the request redundancy (on average ×1—2.5). Previously, this cost did not reach us due to the non-functional price finalization API, and we paid as if for a single host; now it is calculated based on actuals, and we are reflecting this in our rates. Inference remains hundreds of times cheaper than vendor APIs. Current prices for all models are available on the Pricing page and via GET /api/pricing.
The gateway accepts requests in the Responses format — a new OpenAI protocol adopted by recent SDKs and used by Codex CLI. The address remains the same: base URL https://gate.joingonka.ai/v1, the client will automatically call /v1/responses. Streaming, function tools, and structured output (text.format) are supported; costs and limits are the same as for /v1/chat/completions. We do not store dialogues, so store is ignored, and previous_response_id will return an error — please pass the entire history in the input field. OpenAI built-in server tools (web_search, file_search, code_interpreter) are not supported; our web search is still available via the web plugin at /v1/chat/completions.
Previously, Gonka network failures in stream:true mode returned a 200 code, and the error was only visible inside the SSE chunk: agent clients mistook such responses for empty successful ones and did not retry. Now, if the network response has not yet started, the error is returned with a standard HTTP code — 429 for model overload (with a Retry-After header), 504 for timeout, and 503 if no free nodes are available. If the failure occurs in the middle of a stream, the status cannot be changed: it still contains a chunk with an error field, but the final [DONE] is no longer sent so that the stream does not appear to be successfully completed.
npx @joingonka/setup now configures 15 agentic tools: Claude Code, OpenClaw, Cursor, Cline, opencode, Aider, Kilo Code, Roo Code, Continue, Hermes, Pi, Zed, ZCode, JetBrains AI Assistant, and GitHub Copilot BYOK. For seven of them, the config is written and cleanly merged with yours (comments and existing settings are preserved); for the rest, which are configured only via GUI, ready-to-paste values are printed. You can select a specific tool immediately: npx @joingonka/setup --tool zed. Each run concludes with a live request to the gateway, confirming that the key, address, and model have been accepted.
The maximum length of a single response for DeepSeek V4 Flash has been increased by 4 times: from 8,192 to 32,768 tokens (max_tokens parameter). Large files and long responses from agent tools now fit into a single call. Additionally: if the response hits the limit in the middle of a tool call, the gateway returns a proper finish_reason="length" instead of a truncated call — agent clients (Cursor and others) continue working correctly instead of hanging on a parsing error.
A third active model has appeared in the Gonka network — DeepSeek V4 Flash 0731: 380,000 token context (the longest in the network, verified by a real 381K request), reasoning and tool calling. The identifier is deepseek-ai/DeepSeek-V4-Flash-0731, pricing is the same as other models. The model is in early testing but fully available: select it in the dashboard chat or specify it in the model field of your API request.
From August 3, inference cost: input — 22 nGNK, output — 66 nGNK per token (double the previous 11/33). Gateway fee remains unchanged at 10%. Demo mode remains free. Current prices for all models are on the Pricing page and GET /api/pricing.
If the Gonka network receives a request but fails to respond (300s timeout), the gateway now immediately returns a 504 error instead of a series of silent retries. Important note on billing: a request accepted by the network cannot be cancelled — the node processes it regardless, so in case of a timeout, the prompt processing estimate is deducted (completion is not billed). Recommendation: for long generations, use stream:true — streams are more resilient to timeouts and show progress immediately.
The deposit fee via USDT now decreases as the amount increases: base 5%, from $25 — 4.5%, from $50 — 4%, from $100 — 3.5%, from $250 — 3%, from $500 — 2.75%, from $1000 — 2.5%. The deposit dialog features an interactive ladder: amount slider, GNK bonus for each step, and your savings. The discount is applied automatically upon deposit.
You can set a total (cumulative) spending limit for child keys — a dedicated quota that does not reset, unlike daily or monthly limits. Once the quota is exhausted, requests using the key are rejected. The limit is set in the dashboard (with nanogonka precision) or via API using the limit_total_ngonka field in POST/PATCH /api/management/keys/{id}/children; the accumulated usage is reset using the "Reset usage" button (or POST /api/management/keys/{id}/children/{childId}/reset-usage).
Inference costs are now synchronized with the actual Gonka network price per token. Prices remain among the lowest on the market — fractions of a cent per million tokens, hundreds of times cheaper than OpenAI and Anthropic for the same open-source models. Output tokens are priced higher than input tokens (×3), as is standard with most providers.
We have confirmed the actual context limit for active models — up to 200,000 tokens (previously, conservative estimates of 131,072 were published). You can now send longer documents and dialogue histories; current specifications for each model can be found in the "Network Status" section.
The cost of each request is visible in USD — in the API response (field usage.total_cost_usd) and in the "Usage" section. Easy for budget planning.
Set up API access in your tool automatically: run npx @joingonka/setup. Supported: Claude Code, OpenClaw, Cline, Continue, opencode, Aider, and others.
The model searches the internet during the response and adds links to sources. To enable: pass plugins: [{ "id": "web", "max_results": 5 }] in the request body — works in OpenAI- and Anthropic-compatible API.
Maximum response length is up to 8192 tokens; control it via the max_tokens parameter (default is 1024 for streaming and 1500 for non-streaming).
On the "Chat" page, select a model from the list (loaded from /v1/models) — the selection is saved and applied to new conversations.
Need speed? Add the header x-joingonka-meta: 1 to the request — the streaming response will include ttft_ms (time to first token) and tokens_per_sec.
Create a management key (POST /api/management/keys) and generate child keys with custom limits (daily/monthly quotas, rate limits, expiration) via POST /api/management/keys/{id}/children.
Enable plugins using the plugins parameter: response-healing fixes broken JSON, privacy-sanitization masks private data, file-parser extracts text from PDFs and documents.
OpenAI: base_url https://gate.joingonka.ai/v1, header Authorization: Bearer your-key. Anthropic (Claude Code): base_url https://gate.joingonka.ai, header x-api-key, endpoint /v1/messages.
GNK: get the address and memo and send tokens on-chain. USDT: specify an amount from 1 to 10,000 USD and pay via OxaPay — automatic conversion to GNK.