Knowledge Base Sections ▾
Navigation
▸ Start here By rolesCategories
- Cursor + Gonka AI - cheap LLM for coding
- Claude Code + Gonka AI - LLM for the terminal
- OpenClaw + Gonka AI - affordable AI agents
- OpenCode: your own model in the terminal
- Continue.dev + Gonka AI - AI for VS Code/JetBrains
- Cline + Gonka AI - AI agent in VS Code
- Aider + Gonka AI - pair programming with AI
- LangChain + Gonka AI - AI applications for pennies
- n8n + Gonka AI - automation with cheap AI
- Open WebUI + Gonka AI - your own ChatGPT
- LibreChat + Gonka AI — open-source ChatGPT
- Hermes Agent + DeepSeek on the Gonka network — an autonomous agent for pennies
- Kilo Code + Gonka AI — AI-Agent in VS Code
- Roo Code + Gonka AI — Autonomous AI Agent in VS Code
- LlamaIndex + Gonka AI — RAG applications for pennies
- PydanticAI + Gonka — typed AI agents for pennies
- Vercel AI SDK + Gonka AI — AI applications in TypeScript for pennies
- TanStack AI + Gonka — AI applications in TypeScript for pennies
- API quick start — curl, Python, TypeScript
- JoinGonka Gateway — a full overview
- Management Keys — SaaS on Gonka
- Cheapest AI API: Provider Comparison 2026
- How to buy AI tokens and an API key: 3 methods in 2026
- Cursor Pro request limit reached — breakdown and cheaper alternative
- Claude Code is cheaper — bill breakdown and switching
- Cline is burning money — why the agent spends so much
- OpenClaw is expensive — why the agent burns through tokens and how to save
- OpenRouter: Cheap Alternative — Comparison with JoinGonka Gateway
- Best AI model for coding in 2026: comparison and prices
- Cheap alternative to GitHub Copilot without limits
- A cheap Windsurf alternative without credits or limits
- The cheapest API for AI agents in 2026
- ZCode: Cheap GLM inference instead of GLM Coding Plan
- JetBrains IDE + JoinGonka Gateway — your own endpoint instead of credits
- GitHub Copilot BYOK — own models instead of quotas
- Zed + JoinGonka Gateway — cheap inference in your editor
- Pi + JoinGonka Gateway — terminal agent on cheap inference
- Codex CLI: your own key instead of a subscription
- DeepSeek Harness: Your Own Provider via JoinGonka Gateway
- MiniMax Code: MiniMax agent with your own key via Gonka
- Warp + JoinGonka Gateway — terminal agent on your own endpoint
- Trae + JoinGonka Gateway — Gonka network models in AI-IDE
- Cherry Studio + JoinGonka Gateway — desktop AI client
- omp (Oh My Pi) + JoinGonka Gateway: an agent with model roles
- OpenHands + JoinGonka Gateway: agent on your own endpoint
- Qwen Code after the closure of qwen-oauth: working via JoinGonka Gateway
- Goose + JoinGonka Gateway: your own provider and key in the keyring
- Crush + JoinGonka Gateway: Charm agent on Gonka network models
- Zoo Code + JoinGonka Gateway: Migrating from Roo Code to Gonka models
- Kimi Code CLI: Moonshot AI agent on your key via Gonka
- Factory Droid + JoinGonka Gateway: BYOK on Gonka network models
- MiMo Code + JoinGonka Gateway: Xiaomi agent on Gonka network models
Tools
ZCode: Cheap GLM inference instead of GLM Coding Plan
In June 2026, Z.ai released ZCode 3.0—a desktop agent-based development environment that the company itself calls the "Official Harness for GLM-5.2." The product quickly gained attention: its own agent core, autonomous tasks via /goal, multi-agent mode, and remote session management directly from Telegram. But along with the wave of reviews came a wave of questions from early GLM Coding Plan subscribers: the quota is counted as "prompts" in five-hour windows, consumption is multiplied by three during peak hours, lower-tier plans have only one parallel request, and users on Hacker News complain about sudden blocks after a busy day of work.
The good news: ZCode officially supports a custom provider—you can connect any OpenAI- or Anthropic-compatible API to it and pay for actual tokens, not subscription quotas. In this guide, we will break down what ZCode is and how the GLM Coding Plan is structured, then step-by-step configure ZCode for the JoinGonka Gateway—a gateway for the decentralized Gonka network, which serves GLM-5.3 Flash, the open reasoning model from Z.ai itself. The bottom line: the same tool, the same open-weight models, but the economics are token-based and orders of magnitude cheaper than centralized APIs.
What is ZCode: Agentic development environment from Z.ai
ZCode is not a «chat sidebar editor», but an ADE: you describe a goal, the agent creates a plan, edits files, runs checks, and iterates until the result is achieved. It works as a desktop application on macOS and Windows (Linux build is in beta status).
Key features of version 3.0:
- ZCode Agent — native agent core with deep GLM-5.2 integration; in addition to it, you can connect third-party agents (Claude Code, Codex, Gemini CLI, and others);
- Goal-mode (
/goal) — long autonomous tasks: plan → execution → step-by-step verification; - Multi-agent — parallel work of multiple agents on a project;
- Remote control — starting and controlling sessions from Telegram, WeChat, and Feishu: the agent runs on your machine, and you answer questions from your phone.
The client itself is free, and Z.ai gives new users a daily GLM-5.2 token quota for trial. After that — either a GLM Coding Plan subscription or your own API key. And this is where it is worth diving into the details, as the subscription model is the most discussed part of the product.
GLM Coding Plan: Price, limits, and quotas
GLM Coding Plan is Z.ai's subscription that gives you access to GLM models in ZCode and in two dozen third-party tools. At the time of publication (July 2026) the basic pricing looks like this — but prices and promos change, so check the live page at z.ai/subscribe:
| Plan | Price, $/mo | Quota | Concurrent requests |
|---|---|---|---|
| Lite | $18 (with promos — from ~$12.6) | "Prompts" in 5-hour windows + a weekly ceiling | 1 |
| Pro | $72 | ×4 vs Lite | 1 |
| Max | $160 | ×16 vs Lite | up to 30 |
What subscribers complain about (Hacker News and GitHub threads):
- Peak multiplier. During peak hours (14:00—18:00 UTC+8 — evening in Asia, morning-to-afternoon in Europe), every request deducts quota at a ×3 coefficient, off-peak at ×2. You end up planning your usage around Beijing time.
- Quota counted in "prompts," not tokens. How many tokens you actually have left is opaque; threads include stories like "burned ~17M tokens on GLM — and got blocked for four days."
- One concurrent request on Lite and Pro. For an agentic environment that's noticeable: ZCode's own multi-agent mode and any parallel sub-agents hit a queue.
- GLM-5.2 burns quota faster than the smaller models — the flagship consumes the limit with its own multiplier.
The takeaway is simple: for relaxed solo development, the subscription works honestly. But the more "agentic" your mode — parallel tasks, long autonomous sessions, overnight runs — the more often the quota model turns into a lottery. The alternative is paying for actual tokens: no windows, multipliers, or queues — however much the agent consumed, that's what gets deducted. That's exactly the model the custom provider mechanism built into ZCode offers, and it takes five minutes to set up.
BYOK in ZCode: Your own API key and custom provider
ZCode officially has built-in support for BYOK: the documentation explicitly permits adding "any service compatible with the OpenAI or Anthropic protocols" — public APIs, corporate channels, and even self-hosted installations. There are three ways to configure it:
- the Connect button on the welcome screen → "Set Your API Key";
- Manage Models in the model selector;
- the gear icon → Settings → Model Settings → Add Provider.
The provider is given an OpenAI Base URL and/or Anthropic Base URL plus an API Key — ZCode will pull in the list of available models automatically.
Three pitfalls that people trip over most often:
- Z.ai's endpoints differ. The Coding Plan subscription only works through the coding endpoint
https://api.z.ai/api/coding/paas/v4; the general/api/paas/v4with a subscription key will return an error — the classic cause of "I'm getting 401/429 even though my plan is paid for." - "Works in Claude Code ≠ works in ZCode." Environment variables like
ANTHROPIC_BASE_URL, which are used to configure Claude Code, are not read by ZCode: it has its own storage for provider settings. Configure it through Settings specifically. - "A model shows up in the selector ≠ your key has quota." The model list is pulled from the provider's API, while availability depends on the balance and limits of the specific key.
Next comes practice: we connect the JoinGonka gateway to ZCode and get the GLM stack for tokens.
Setting up JoinGonka Gateway in ZCode: GLM for tokens
JoinGonka Gateway is a gateway to the decentralized Gonka network featuring two native protocols: the OpenAI-compatible /v1/chat/completions and Anthropic Messages /v1/messages. The economy is pay-per-token: no 5-hour windows, no ×3 peak multipliers, and no limits on single concurrent requests; usage is visible in the dashboard in real-time, and the gateway does not store prompt content.
Step-by-step (current Custom Provider interface):
| Step | Action |
|---|---|
| 1. Key | Sign up at gate.joingonka.ai/register (new accounts receive 3M free tokens), then create an API key in the Keys section. More details can be found in API Quick Start. |
| 2. Provider | ZCode → Settings → Model Settings → Add Provider. Name: JoinGonka |
| 3. Base URL | https://gate.joingonka.ai/v1 |
| 4. API Key | Your key in the format jg-… |
| 5. API format | Chat completions (/chat/completions) |
| 6. Model | Add Model → Model ID deepseek-ai/DeepSeek-V4-Flash-0731, Context window 380000. You can also add MiniMaxAI/MiniMax-M2.7 and zai-org/GLM-5.3-Flash — the reasoning model from Z.ai; for the current list and limits, see GET /v1/models. |
In older versions of ZCode, there are two separate fields instead: OpenAI Base URL (https://gate.joingonka.ai/v1) and Anthropic Base URL (https://gate.joingonka.ai — the client will add the path itself).
Verification: Select the created model and ask the agent any question. If you receive an answer, everything is working; if you get a 401 error, check the key, or a 402 error, check your account balance. If a "Model request failed" error appears on DeepSeek V4 Flash when Thought Level is set to Max, switch Thought Level to High and retry the request.
Hint: The command npx @joingonka/setup --tool zcode will print the ready-to-use values for these fields—base URL, key, model ID, and its actual limits (context window and response ceiling)—and perform a live request to verify that the gateway accepts them. ZCode itself is configured only via the UI; there is no configuration file that can be safely edited, so values must be entered manually.
Price Comparison: Coding Plan, Z.ai API, OpenRouter, and Gonka
How much GLM inference costs on different channels (prices per 1M tokens; competitor prices are fixed as of September 2026, the JoinGonka price is injected into the page live from the gateway API):
| Channel | Input, $/1M | Output, $/1M | Payment Model |
|---|---|---|---|
| Z.ai API (GLM-5.2) | $1.40 (cached input — $0.26) | $4.40 | per token |
| GLM Coding Plan | $18—160/mo | subscription: "prompt" quotas, ×3 peak multiplier, concurrency limits | |
| OpenRouter (z-ai/glm-5.2) | $1.40 | $4.40 | per token |
| OpenRouter (z-ai/glm-5.3-flash) | $0.09 | $0.30 | per token |
| JoinGonka Gateway | $0.0069 | $0.021 | per token, Gonka network |
The two-to-three order of magnitude difference is not marketing rounding, but a different economy: the computations are performed by participants of the decentralized network who earn from the work itself, not from data center margins. A detailed analysis of how this works and where the limits of applicability are—in the article on the cheapest AI API.
The JoinGonka line is the price of GLM-5.3 Flash in the Gonka network: all models in the network have a unified tariff. The numbers on this page update automatically; there is no need to recalculate anything manually.
GLM in the Gonka network: GLM-5.3 Flash and FAQ
The Gonka network serves the Z.ai model: zai-org/GLM-5.3-Flash is adopted by governance proposal #101 and provided via the gateway alongside MiniMax M2.7 and DeepSeek V4 Flash. The flagship GLM-5.2, mentioned in early versions of this guide, was not included in the production lineup—the network chose a lighter and faster model from the same family. The Base URL and key remain the same: in ZCode, simply add the model as a custom provider.
About the model: open weights under an MIT license, MoE architecture with 320B parameters (18B active per token), and hybrid sparse/linear attention. The key feature is that it's a reasoning model: it processes its thoughts before responding, and these thoughts consume several hundred tokens even for simple questions. For ZCode, this means two things: do not set the response length limit too low (at least 600 tokens even for short replies), and for quick edits, lower the Thought Level—at the API level, this is reasoning_effort: low. A detailed analysis of the model and its limits can be found in the article GLM-5.3 Flash on the Gonka network.
Frequently Asked Questions:
- Why are there 429s or queues on the new model? — The model is currently served by a small subset of network hosts; during peak minutes, requests may wait for an available slot. Retry the request after a few seconds—the load balancer will find an active node. The current status is available on the gateway status page.
- What else is available? — Besides GLM-5.3 Flash, there are MiniMax M2.7 and DeepSeek V4 Flash: all three models work in ZCode via the same custom provider, and the configuration from this guide is fully functional.
- What about privacy? — The gateway does not store the content of prompts and responses; statistics only contain usage aggregates.
- When is the Coding Plan more cost-effective? — If you work solo, stay within the Lite quota, and do not run parallel agents, a subscription with a fixed bill is convenient and predictable. Calculate based on your profile: with active agentic work, paying per token through the network is orders of magnitude cheaper; for light code edits a couple of times a day, the difference may not be noticeable.
- Is this suitable for other tools? — Yes: the same endpoint works in Claude Code, Cursor, Cline, and any OpenAI/Anthropic-compatible client.