Knowledge Base Sections ▾

Navigation

▸ Start here By roles

Categories

Tools 52
Glossary 12

Tools

ZCode: Cheap GLM inference instead of GLM Coding Plan

In June 2026, Z.ai released ZCode 3.0—a desktop agent-based development environment that the company itself calls the "Official Harness for GLM-5.2." The product quickly gained attention: its own agent core, autonomous tasks via /goal, multi-agent mode, and remote session management directly from Telegram. But along with the wave of reviews came a wave of questions from early GLM Coding Plan subscribers: the quota is counted as "prompts" in five-hour windows, consumption is multiplied by three during peak hours, lower-tier plans have only one parallel request, and users on Hacker News complain about sudden blocks after a busy day of work.

The good news: ZCode officially supports a custom provider—you can connect any OpenAI- or Anthropic-compatible API to it and pay for actual tokens, not subscription quotas. In this guide, we will break down what ZCode is and how the GLM Coding Plan is structured, then step-by-step configure ZCode for the JoinGonka Gateway—a gateway for the decentralized Gonka network, which serves GLM-5.3 Flash, the open reasoning model from Z.ai itself. The bottom line: the same tool, the same open-weight models, but the economics are token-based and orders of magnitude cheaper than centralized APIs.

What is ZCode: Agentic development environment from Z.ai

ZCode is not a «chat sidebar editor», but an ADE: you describe a goal, the agent creates a plan, edits files, runs checks, and iterates until the result is achieved. It works as a desktop application on macOS and Windows (Linux build is in beta status).

Key features of version 3.0:

  • ZCode Agent — native agent core with deep GLM-5.2 integration; in addition to it, you can connect third-party agents (Claude Code, Codex, Gemini CLI, and others);
  • Goal-mode (/goal) — long autonomous tasks: plan → execution → step-by-step verification;
  • Multi-agent — parallel work of multiple agents on a project;
  • Remote control — starting and controlling sessions from Telegram, WeChat, and Feishu: the agent runs on your machine, and you answer questions from your phone.

The client itself is free, and Z.ai gives new users a daily GLM-5.2 token quota for trial. After that — either a GLM Coding Plan subscription or your own API key. And this is where it is worth diving into the details, as the subscription model is the most discussed part of the product.

GLM Coding Plan: Price, limits, and quotas

GLM Coding Plan is Z.ai's subscription that gives you access to GLM models in ZCode and in two dozen third-party tools. At the time of publication (July 2026) the basic pricing looks like this — but prices and promos change, so check the live page at z.ai/subscribe:

PlanPrice, $/moQuotaConcurrent requests
Lite$18 (with promos — from ~$12.6)"Prompts" in 5-hour windows + a weekly ceiling1
Pro$72×4 vs Lite1
Max$160×16 vs Liteup to 30

What subscribers complain about (Hacker News and GitHub threads):

  • Peak multiplier. During peak hours (14:00—18:00 UTC+8 — evening in Asia, morning-to-afternoon in Europe), every request deducts quota at a ×3 coefficient, off-peak at ×2. You end up planning your usage around Beijing time.
  • Quota counted in "prompts," not tokens. How many tokens you actually have left is opaque; threads include stories like "burned ~17M tokens on GLM — and got blocked for four days."
  • One concurrent request on Lite and Pro. For an agentic environment that's noticeable: ZCode's own multi-agent mode and any parallel sub-agents hit a queue.
  • GLM-5.2 burns quota faster than the smaller models — the flagship consumes the limit with its own multiplier.

The takeaway is simple: for relaxed solo development, the subscription works honestly. But the more "agentic" your mode — parallel tasks, long autonomous sessions, overnight runs — the more often the quota model turns into a lottery. The alternative is paying for actual tokens: no windows, multipliers, or queues — however much the agent consumed, that's what gets deducted. That's exactly the model the custom provider mechanism built into ZCode offers, and it takes five minutes to set up.

BYOK in ZCode: Your own API key and custom provider

ZCode officially has built-in support for BYOK: the documentation explicitly permits adding "any service compatible with the OpenAI or Anthropic protocols" — public APIs, corporate channels, and even self-hosted installations. There are three ways to configure it:

  • the Connect button on the welcome screen → "Set Your API Key";
  • Manage Models in the model selector;
  • the gear icon → Settings → Model Settings → Add Provider.

The provider is given an OpenAI Base URL and/or Anthropic Base URL plus an API Key — ZCode will pull in the list of available models automatically.

Three pitfalls that people trip over most often:

  • Z.ai's endpoints differ. The Coding Plan subscription only works through the coding endpoint https://api.z.ai/api/coding/paas/v4; the general /api/paas/v4 with a subscription key will return an error — the classic cause of "I'm getting 401/429 even though my plan is paid for."
  • "Works in Claude Code ≠ works in ZCode." Environment variables like ANTHROPIC_BASE_URL, which are used to configure Claude Code, are not read by ZCode: it has its own storage for provider settings. Configure it through Settings specifically.
  • "A model shows up in the selector ≠ your key has quota." The model list is pulled from the provider's API, while availability depends on the balance and limits of the specific key.

Next comes practice: we connect the JoinGonka gateway to ZCode and get the GLM stack for tokens.

Setting up JoinGonka Gateway in ZCode: GLM for tokens

JoinGonka Gateway is a gateway to the decentralized Gonka network featuring two native protocols: the OpenAI-compatible /v1/chat/completions and Anthropic Messages /v1/messages. The economy is pay-per-token: no 5-hour windows, no ×3 peak multipliers, and no limits on single concurrent requests; usage is visible in the dashboard in real-time, and the gateway does not store prompt content.

Step-by-step (current Custom Provider interface):

StepAction
1. KeySign up at gate.joingonka.ai/register (new accounts receive 3M free tokens), then create an API key in the Keys section. More details can be found in API Quick Start.
2. ProviderZCode → Settings → Model Settings → Add Provider. Name: JoinGonka
3. Base URLhttps://gate.joingonka.ai/v1
4. API KeyYour key in the format jg-…
5. API formatChat completions (/chat/completions)
6. ModelAdd Model → Model ID deepseek-ai/DeepSeek-V4-Flash-0731, Context window 380000. You can also add MiniMaxAI/MiniMax-M2.7 and zai-org/GLM-5.3-Flash — the reasoning model from Z.ai; for the current list and limits, see GET /v1/models.

In older versions of ZCode, there are two separate fields instead: OpenAI Base URL (https://gate.joingonka.ai/v1) and Anthropic Base URL (https://gate.joingonka.ai — the client will add the path itself).

Verification: Select the created model and ask the agent any question. If you receive an answer, everything is working; if you get a 401 error, check the key, or a 402 error, check your account balance. If a "Model request failed" error appears on DeepSeek V4 Flash when Thought Level is set to Max, switch Thought Level to High and retry the request.

Hint: The command npx @joingonka/setup --tool zcode will print the ready-to-use values for these fields—base URL, key, model ID, and its actual limits (context window and response ceiling)—and perform a live request to verify that the gateway accepts them. ZCode itself is configured only via the UI; there is no configuration file that can be safely edited, so values must be entered manually.

Price Comparison: Coding Plan, Z.ai API, OpenRouter, and Gonka

How much GLM inference costs on different channels (prices per 1M tokens; competitor prices are fixed as of September 2026, the JoinGonka price is injected into the page live from the gateway API):

ChannelInput, $/1MOutput, $/1MPayment Model
Z.ai API (GLM-5.2)$1.40 (cached input — $0.26)$4.40per token
GLM Coding Plan$18—160/mosubscription: "prompt" quotas, ×3 peak multiplier, concurrency limits
OpenRouter (z-ai/glm-5.2)$1.40$4.40per token
OpenRouter (z-ai/glm-5.3-flash)$0.09$0.30per token
JoinGonka Gateway$0.0069$0.021per token, Gonka network

The two-to-three order of magnitude difference is not marketing rounding, but a different economy: the computations are performed by participants of the decentralized network who earn from the work itself, not from data center margins. A detailed analysis of how this works and where the limits of applicability are—in the article on the cheapest AI API.

The JoinGonka line is the price of GLM-5.3 Flash in the Gonka network: all models in the network have a unified tariff. The numbers on this page update automatically; there is no need to recalculate anything manually.

GLM in the Gonka network: GLM-5.3 Flash and FAQ

The Gonka network serves the Z.ai model: zai-org/GLM-5.3-Flash is adopted by governance proposal #101 and provided via the gateway alongside MiniMax M2.7 and DeepSeek V4 Flash. The flagship GLM-5.2, mentioned in early versions of this guide, was not included in the production lineup—the network chose a lighter and faster model from the same family. The Base URL and key remain the same: in ZCode, simply add the model as a custom provider.

About the model: open weights under an MIT license, MoE architecture with 320B parameters (18B active per token), and hybrid sparse/linear attention. The key feature is that it's a reasoning model: it processes its thoughts before responding, and these thoughts consume several hundred tokens even for simple questions. For ZCode, this means two things: do not set the response length limit too low (at least 600 tokens even for short replies), and for quick edits, lower the Thought Level—at the API level, this is reasoning_effort: low. A detailed analysis of the model and its limits can be found in the article GLM-5.3 Flash on the Gonka network.

Frequently Asked Questions:

  • Why are there 429s or queues on the new model? — The model is currently served by a small subset of network hosts; during peak minutes, requests may wait for an available slot. Retry the request after a few seconds—the load balancer will find an active node. The current status is available on the gateway status page.
  • What else is available? — Besides GLM-5.3 Flash, there are MiniMax M2.7 and DeepSeek V4 Flash: all three models work in ZCode via the same custom provider, and the configuration from this guide is fully functional.
  • What about privacy? — The gateway does not store the content of prompts and responses; statistics only contain usage aggregates.
  • When is the Coding Plan more cost-effective? — If you work solo, stay within the Lite quota, and do not run parallel agents, a subscription with a fixed bill is convenient and predictable. Calculate based on your profile: with active agentic work, paying per token through the network is orders of magnitude cheaper; for light code edits a couple of times a day, the difference may not be noticeable.
  • Is this suitable for other tools? — Yes: the same endpoint works in Claude Code, Cursor, Cline, and any OpenAI/Anthropic-compatible client.
ZCode 3.0 is a notable agentic IDE, but its native GLM Coding Plan subscription is quota-based: "prompts" in 5-hour windows, a ×3 multiplier during peak hours, and only one parallel request on entry-level plans. The more agentic the workload, the more expensive and unpredictable it becomes. The official BYOK solves this: connect the JoinGonka Gateway as a custom provider—and pay for actual tokens at prices on the decentralized Gonka network, which are hundreds of times lower than centralized APIs. GLM-5.3 Flash—Z.ai's own open reasoning model—is already served by the network, alongside MiniMax M2.7 and DeepSeek V4 Flash. Start with free trial tokens.

Want to learn more?

Explore other sections or start earning GNK right now.

Get free tokens →