Knowledge Base Sections ▾
Navigation
▸ Start here By rolesCategories
- Cursor + Gonka AI - cheap LLM for coding
- Claude Code + Gonka AI - LLM for the terminal
- OpenClaw + Gonka AI - affordable AI agents
- OpenCode: your own model in the terminal
- Continue.dev + Gonka AI - AI for VS Code/JetBrains
- Cline + Gonka AI - AI agent in VS Code
- Aider + Gonka AI - pair programming with AI
- LangChain + Gonka AI - AI applications for pennies
- n8n + Gonka AI - automation with cheap AI
- Open WebUI + Gonka AI - your own ChatGPT
- LibreChat + Gonka AI — open-source ChatGPT
- Hermes Agent + DeepSeek on the Gonka network — an autonomous agent for pennies
- Kilo Code + Gonka AI — AI-Agent in VS Code
- Roo Code + Gonka AI — Autonomous AI Agent in VS Code
- LlamaIndex + Gonka AI — RAG applications for pennies
- PydanticAI + Gonka — typed AI agents for pennies
- Vercel AI SDK + Gonka AI — AI applications in TypeScript for pennies
- TanStack AI + Gonka — AI applications in TypeScript for pennies
- API quick start — curl, Python, TypeScript
- JoinGonka Gateway — a full overview
- Management Keys — SaaS on Gonka
- Cheapest AI API: Provider Comparison 2026
- How to buy AI tokens and an API key: 3 methods in 2026
- Cursor Pro request limit reached — breakdown and cheaper alternative
- Claude Code is cheaper — bill breakdown and switching
- Cline is burning money — why the agent spends so much
- OpenClaw is expensive — why the agent burns through tokens and how to save
- OpenRouter: Cheap Alternative — Comparison with JoinGonka Gateway
- Best AI model for coding in 2026: comparison and prices
- Cheap alternative to GitHub Copilot without limits
- A cheap Windsurf alternative without credits or limits
- The cheapest API for AI agents in 2026
- ZCode: Cheap GLM inference instead of GLM Coding Plan
- JetBrains IDE + JoinGonka Gateway — your own endpoint instead of credits
- GitHub Copilot BYOK — own models instead of quotas
- Zed + JoinGonka Gateway — cheap inference in your editor
- Pi + JoinGonka Gateway — terminal agent on cheap inference
- Codex CLI: your own key instead of a subscription
- DeepSeek Harness: Your Own Provider via JoinGonka Gateway
- MiniMax Code: MiniMax agent with your own key via Gonka
- Warp + JoinGonka Gateway — terminal agent on your own endpoint
- Trae + JoinGonka Gateway — Gonka network models in AI-IDE
- Cherry Studio + JoinGonka Gateway — desktop AI client
- omp (Oh My Pi) + JoinGonka Gateway: an agent with model roles
- OpenHands + JoinGonka Gateway: agent on your own endpoint
- Qwen Code after the closure of qwen-oauth: working via JoinGonka Gateway
- Goose + JoinGonka Gateway: your own provider and key in the keyring
- Crush + JoinGonka Gateway: Charm agent on Gonka network models
- Zoo Code + JoinGonka Gateway: Migrating from Roo Code to Gonka models
- Kimi Code CLI: Moonshot AI agent on your key via Gonka
- Factory Droid + JoinGonka Gateway: BYOK on Gonka network models
- MiMo Code + JoinGonka Gateway: Xiaomi agent on Gonka network models
Tools
Qwen Code after the closure of qwen-oauth: working via JoinGonka Gateway
Qwen Code (command qwen) is an open-source terminal agent for development from the Qwen team at Alibaba: it understands the project, edits files, and runs commands and tests—from a terminal, editor, or scripts. The project grew out of Google Gemini CLI and is developing independently; the code is open under the Apache-2.0 license. The agent supports several protocols—OpenAI, Anthropic, Gemini—so almost any model provider can connect to it.
Starting work with Qwen Code without a key was made possible by free login via a qwen.ai account—the qwen-oauth authorization type. On April 15, 2026, this free tier was closed, but the default behavior remained the same: if no authorization type is selected, Qwen Code defaults to it, and an old installation responds with the string Qwen OAuth free tier was discontinued on 2026-04-15. Documentation suggests switching to another provider, and the easiest way is via an OpenAI-compatible endpoint.
JoinGonka Gateway is exactly such an endpoint: through it, Qwen Code works on decentralized Gonka network models—DeepSeek V4 Flash, GLM-5.3 Flash, and MiniMax M2.7—at a flat price of $0.0069 per million input tokens, without a subscription. There are no proprietary Qwen models in the network right now, but the agent doesn't need them: it works with any model capable of calling tools. The commands and messages below have been verified by a live run of Qwen Code 0.24.4 through the gateway on September 23, 2026. After confirming the address, 3M free tokens will be credited to your account—enough to repeat the entire path yourself.
Quick start: installation and a single command
Step 1: install Qwen Code. The official methods from the project README:
# Linux and macOS: standalone build with its own Node.js
curl -fsSL https://qwen-code-assets.oss-cn-hangzhou.aliyuncs.com/installation/install-qwen-standalone.sh | bash
# Windows (PowerShell)
irm https://qwen-code-assets.oss-cn-hangzhou.aliyuncs.com/installation/install-qwen-standalone.ps1 | iex
# npm (requires Node.js 22 or newer)
npm install -g @qwen-code/qwen-code@latest
# Homebrew
brew install qwen-codeThe script places the command in ~/.local/bin/qwen and, if needed, appends the path to your shell's config file. Open a new terminal and check: qwen --version.
Step 2: get a key. Sign up at gate.joingonka.ai/register, verify your address, and in the "API keys" section create a key with the jg- prefix.
Step 3: run the installer.
npx @joingonka/setup --tool qwen-codeThe installer will ask for your key (it isn't passed as a command-line argument, so it won't linger in your shell history) and will edit ~/.qwen/settings.json — the file that the Qwen Code docs suggest for "single-file" configuration:
- it writes the key into the file's own
envblock under the nameJOINGONKA_API_KEY— the same place the built-in/authcommand stores keys; the file gets600permissions, and your shell environment stays untouched; - it adds the network's models to the
modelProviders.openailist — with the gateway address, real context windows, and response ceilings; entries from your other providers in that list are preserved; - if the auth type isn't set or still points to the deprecated
qwen-oauth, it switches it toopenaiand sets DeepSeek V4 Flash as the default model; a working choice of another type —anthropic,gemini,vertex-ai— is left alone unless you pass--model; - it saves a copy of the previous file, leaves comments and other settings untouched, and finally sends a live request to the gateway to report whether the key, address, and model were accepted.
Here's what the output looked like on an install that still had qwen-oauth (abridged):
Configured ~/.qwen/settings.json
Base URL: https://gate.joingonka.ai/v1
Default model: deepseek-ai/DeepSeek-V4-Flash-0731 (replaces qwen-oauth — that free tier was discontinued on 2026-04-15)
Switched Qwen Code to JoinGonka (auth type "openai") — this replaced your previous auth type "qwen-oauth".
To go back: run /auth (and /model) inside Qwen Code, or restore ~/.qwen/settings.json.bak.2026-09-23T04-22-32-159Z.
…
✓ Verified: the gateway accepted the key, base URL and model.A different default model is set with the --model flag: the shorthand glm, minimax, or a full identifier; an explicitly specified model is always written. Non-interactive mode takes the key from an environment variable:
JOINGONKA_API_KEY=jg-your-key npx @joingonka/setup --tool qwen-code --model glm --non-interactiveA directory moved via the QWEN_HOME variable is found by the installer automatically. Running it again will update the model catalog but keep the network model you selected: Kept your current default model.
Manual configuration: settings.json
Everything the installer does can be done by hand. The file is ~/.qwen/settings.json (or $QWEN_HOME/settings.json); Qwen Code reads it as JSON with comments. Here is a complete working configuration for the Gonka network:
{
"env": { "JOINGONKA_API_KEY": "jg-your-key" },
"modelProviders": {
"openai": [
{ "id": "deepseek-ai/DeepSeek-V4-Flash-0731", "name": "DeepSeek V4 Flash (Gonka)",
"baseUrl": "https://gate.joingonka.ai/v1", "envKey": "JOINGONKA_API_KEY",
"generationConfig": { "contextWindowSize": 380000, "samplingParams": { "max_tokens": 32768 } } },
{ "id": "zai-org/GLM-5.3-Flash", "name": "GLM 5.3 Flash (Gonka)",
"baseUrl": "https://gate.joingonka.ai/v1", "envKey": "JOINGONKA_API_KEY",
"generationConfig": { "contextWindowSize": 390000, "samplingParams": { "max_tokens": 8192 } } },
{ "id": "MiniMaxAI/MiniMax-M2.7", "name": "MiniMax M2.7 (Gonka)",
"baseUrl": "https://gate.joingonka.ai/v1", "envKey": "JOINGONKA_API_KEY",
"generationConfig": { "contextWindowSize": 200000, "samplingParams": { "max_tokens": 8192 } } }
]
},
"security": { "auth": { "selectedType": "openai" } },
"model": { "name": "deepseek-ai/DeepSeek-V4-Flash-0731", "baseUrl": "https://gate.joingonka.ai/v1" }
}| Field | Value | Why it matters |
|---|---|---|
env.JOINGONKA_API_KEY | jg-… key | The key itself. The model entry has no field for the key — it references this variable via envKey |
baseUrl | https://gate.joingonka.ai/v1 | With /v1 at the end: Qwen Code appends the /chat/completions path itself. The openai list key is the Chat Completions protocol |
contextWindowSize | the model's window | Without it, Qwen Code estimates the window from the model name; it uses this number to decide when to compress the history |
samplingParams.max_tokens | response ceiling | Fixes the response limit. Without it, once it hits the limit, Qwen Code may retry the request with a limit of 64K or more — above the ceiling of any model on the network |
security.auth.selectedType | openai | Which protocol to enable at startup; without this field Qwen Code falls back to the closed qwen-oauth |
model.name, model.baseUrl | the default model and its address | Together they unambiguously point to the gateway entry, even if a model with the same id is configured with another provider |
Where to keep the key. Qwen Code looks up the value of the variable from envKey in order: the shell environment, then the first .env it finds (.qwen/.env or .env from the current directory upward, then ~/.qwen/.env and ~/.env), and only then the env block of the file. Hence the main pitfall: a JOINGONKA_API_KEY exported in the shell with an old value will override the file — we tested this deliberately and got 401 Invalid API key with the correct key in settings.json. The installer warns when it sees a different value in the environment.
And a note on projects: a project-level .qwen/settings.json with its own modelProviders does not extend but replaces the list from the home directory — inside such a project the gateway models will disappear from /model. You should not put the key in the project file at all: along with it, it will end up in the repository.
A running session picks up edits to modelProviders without a restart — just reopen /model. Without an editor, the /auth → Custom Provider wizard does the same.
Authorization and models: how Qwen Code chooses what to work with
Qwen Code has two independent choices: authorization type — which protocol to enable and where to get keys from — and model. The type is stored in security.auth.selectedType, and the installer's behavior depends on it:
Value in selectedType | What the installer will do |
|---|---|
none or qwen-oauth | Installs openai and the default model from the Gonka network; a private plan is not considered a working selection, but the installer will note that it was replaced and how to revert it |
openai | Keeps the selected model — of the gateway or another provider — and only updates the catalog |
anthropic, gemini, vertex-ai | Leaves it alone: this is a functional setup. With the --model flag, it will switch it and explain how to revert to the previous one |
Inside a session, the model is changed via the /model command. In our run, it showed all gateway records with the protocol label, and below the list — the context window, base URL, and the name of the environment variable for the key:
Select Model
1. [openai] MiniMax M2.7 (Gonka) (MiniMaxAI/MiniMax-M2.7)
› 2. [openai] DeepSeek V4 Flash (Gonka) (deepseek-ai/DeepSeek-V4-Flash-0731)
3. [openai] GLM 5.3 Flash (Gonka) (zai-org/GLM-5.3-Flash)
Modality: text-only
Context Window: 380,000 tokens
Base URL: https://gate.joingonka.ai/v1
API Key: JOINGONKA_API_KEYThe selection from /model persists between sessions. For a single launch, the model is set via the -m flag with an identifier: qwen -m zai-org/GLM-5.3-Flash. A separate lightweight model for hints and permission classification is set via /model --fast; as long as it is not present, the main model handles this task.
Which model to choose. The price for all models in the network is the same, so the choice is about behavior. This is how they performed in Qwen Code 0.24.4 on the same task — reading a file and finding an error (verified September 23, 2026):
| Model | Context / Response | Behavior in Qwen Code |
|---|---|---|
| DeepSeek V4 Flash | 380K / 32768 | Installer's choice: large window and the highest response ceiling in the network. Solved the task in two steps |
| GLM-5.3 Flash | 390K / 8192 | Reasoning model: the most accurate response; the reasoning comes as a separate stream and is included in the response limit |
| MiniMax M2.7 | 200K / 8192 | Correct answer in two steps and about seven seconds; the reasoning comes as a separate stream and does not appear in the response text |
Reasoning in GLM-5.3 Flash. The installer provides samplingParams to each model, but in this case, Qwen Code does not add a reasoning level to the request: the /effort command does not reach network models, and GLM-5.3 Flash always thinks. If you need a switch, declare the levels directly in the model record:
{
"id": "zai-org/GLM-5.3-Flash",
"capabilities": {
"reasoning": { "profile": "openai-effort", "efforts": ["low", "high"], "defaultEffort": "high" }
},
…
}Then Qwen Code will start sending reasoning_effort, and for GLM-5.3 Flash, the low value turns off reasoning entirely — see the model overview for details. In our run with "defaultEffort": "low", Qwen Code sent reasoning_effort: "low", and the model responded without a reasoning block.
Verification: what should happen
The fastest check is a one-off run. Put a calc.py with an addition function that actually subtracts into an empty directory, and ask it to find the bug:
qwen -p "Read calc.py and tell me in one sentence whether it has a bug."The agent should call the read tool on its own and answer to the point. Here's how GLM-5.3 Flash answered (run with -m zai-org/GLM-5.3-Flash):
Yes, there is a bug: the `add` function subtracts instead of adding (`return a - b`), so `print(add(2, 3))` prints `-1` instead of `5`.In the interface (qwen with no arguments), the header will show the selected model — API Key | DeepSeek V4 Flash (Gonka) — and the same request will run like this:
> Read calc.py and tell me in one sentence whether it has a bug.
✓ Read calc.py, README.md
◆ Yes, calc.py has a bug: the add function returns the difference (a - b) instead of the sum (a + b). …
➜ demo · DeepSeek V4 Flash (Gonka) · 380.0k Context 5.1% usedThe /stats command shows requests and tokens by model, and the gateway dashboard shows the same usage in the "Usage" section.
Approval mode in -p. Without flags, a one-off run goes into Ask Permissions mode: there's no one to confirm, so tools that require permission — shell commands, file editing and writing — are simply not passed to the model. When asked to run ls, the agent in our run replied that it didn't have such a tool. For edits, add --approval-mode auto-edit; for commands, --yolo; the latter comes with a warning from Qwen Code that the sandbox is not enabled.
If something goes wrong, the diagnosis is usually readable right from the message:
| What you see | What it means | What to do |
|---|---|---|
No auth type is selected. Please configure an auth type … | No auth type selected | Run the installer or add "selectedType": "openai" as in the example above |
Qwen OAuth free tier was discontinued on 2026-04-15 … | A retired qwen-oauth is still in the settings | Run the installer: it will switch the type and tell you what it replaced |
[API Error: 401 Invalid API key] | The gateway rejected the key | Check the env block and whether the shell or .env has a JOINGONKA_API_KEY variable with an old value: it takes precedence over the file |
429 Model "…" is currently overloaded in the Gonka network (rate limit) | The model's free capacity on the network is exhausted | Retry in a minute or switch models; the status is visible on the status page |
Request timeout after 151s …, then Retrying provider attempt | The network under load didn't start responding within two and a half minutes | Qwen Code will retry the request itself; if it drags on, switch models |
How much it costs
Agents consume tokens differently than chat: for every phrase of yours, Qwen Code adds a system prompt and tool descriptions, and tasks take multiple turns. In our run, each qwen -p query carried about 13 thousand input tokens even before getting to the point, and “read the file, find the error” took two turns and about 27 thousand tokens. The interface has more tools, and after answering, Qwen Code performs service requests using the same model—for example, for follow-up suggestions: /stats showed 4 requests and about 79 thousand input tokens for the same task.
Through JoinGonka Gateway, tokens cost $0.0069 per million for input and $0.021 per million for output—the price is the same for all network models and is pulled onto this page from a live source.
| Scenario | Consumption | Via Gateway |
|---|---|---|
One-time task in qwen -p: read a file, find an error | ~27K tokens | hundredths of a cent |
| Same task in the interface, with service requests | ~79K tokens | hundredths of a cent |
| A day of active work | 3-7M tokens | a few cents |
| A month of active development | ~150M tokens | about a dollar |
The estimates in the right column are based on September 2026 prices. For comparison—how you can pay for models in Qwen Code in general:
| Method | Payment Model | Constraints |
|---|---|---|
qwen-oauth | was free | closed on April 15, 2026 |
| Alibaba Cloud Coding Plan | fixed monthly amount | vendor-side quotas, separate sk-sp-… key |
| Alibaba Cloud Token Plan | per token, for teams and companies | regional endpoint and Alibaba Cloud key |
| JoinGonka Gateway | per token, prepaid balance | usage visible in dashboard; no subscriptions or monthly quotas |
Service requests can be disabled where not needed: follow-up suggestions — "ui": { "enableFollowupSuggestions": false }, background auto-memory — "memory": { "enableManagedAutoMemory": false, "enableManagedAutoDream": false }. Precise consumption and balance are available in the dashboard, in the “Usage” and “Billing” sections. Why DeepSeek V4 Flash is selected by default is explained in detail in the model overview.
Things to consider
Approval modes. In the Qwen Code interface, it starts in Auto mode: reading, searching, and editing within the project are allowed immediately, while shell commands and other risky actions are evaluated by a classifier—a separate request to the model. Other modes—Plan, Ask Permissions, Auto-Edit, and YOLO—can be cycled via Shift+Tab (in Windows — Tab) or the /approval-mode command; the default mode is set by the tools.approvalMode key.
Background work. By default, Qwen Code maintains auto-memory: after conversations, it extracts notes about the project and your preferences in the background, and clears them once a day (Dream)—using the same model. In a single run, this is noticeable: qwen -p might not finish for a few more minutes after answering while a background task is waiting for the model. For scripts and CI, turn off auto-memory (keys above) and set run limits: --max-wall-time, --max-session-turns, --max-tool-calls.
Images. The network models are text-based (Modality: text-only in /model). For screenshots, assign a mediator model from another provider using the /model --vision command: it will describe the image in text, and the main model will continue working.
Statistics. Qwen Code sends anonymous usage statistics by default; it is disabled by "privacy": { "usageStatisticsEnabled": false }. The gateway does not store the content of prompts and responses—only consumption aggregates.
Network neighbors. Laboratories whose models the network serves also release their own agents: DeepSeek has DeepSeek Harness, MiniMax has MiniMax Code, and Z.ai (the authors of GLM) has ZCode. All of them connect to the same gateway with the same key and consume the same balance.
qwen-oauth, it requires its own model provider. The fast path is npx @joingonka/setup --tool qwen-code: the installer will write the key into the env block of the ~/.qwen/settings.json file, add DeepSeek V4 Flash, GLM-5.3 Flash, and MiniMax M2.7 to modelProviders.openai with real limits, switch security.auth.selectedType to openai, and verify the connection with a live request. All three network models work: DeepSeek V4 Flash is set by default, while GLM-5.3 Flash and MiniMax M2.7 are selected via the /model command. Verification is done via qwen -p on a file with a bug and checking the "Usage" section in the dashboard; pay for actual tokens at the uniform Gonka network price.Want to learn more?
Explore other sections or start earning GNK right now.
Get key and free tokens →