Knowledge Base Sections ▾
Navigation
▸ Start here By rolesCategories
- Cursor + Gonka AI - cheap LLM for coding
- Claude Code + Gonka AI - LLM for the terminal
- OpenClaw + Gonka AI - affordable AI agents
- OpenCode: your own model in the terminal
- Continue.dev + Gonka AI - AI for VS Code/JetBrains
- Cline + Gonka AI - AI agent in VS Code
- Aider + Gonka AI - pair programming with AI
- LangChain + Gonka AI - AI applications for pennies
- n8n + Gonka AI - automation with cheap AI
- Open WebUI + Gonka AI - your own ChatGPT
- LibreChat + Gonka AI — open-source ChatGPT
- Hermes Agent + DeepSeek on the Gonka network — an autonomous agent for pennies
- Kilo Code + Gonka AI — AI-Agent in VS Code
- Roo Code + Gonka AI — Autonomous AI Agent in VS Code
- LlamaIndex + Gonka AI — RAG applications for pennies
- PydanticAI + Gonka — typed AI agents for pennies
- Vercel AI SDK + Gonka AI — AI applications in TypeScript for pennies
- TanStack AI + Gonka — AI applications in TypeScript for pennies
- API quick start — curl, Python, TypeScript
- JoinGonka Gateway — a full overview
- Management Keys — SaaS on Gonka
- Cheapest AI API: Provider Comparison 2026
- How to buy AI tokens and an API key: 3 methods in 2026
- Cursor Pro request limit reached — breakdown and cheaper alternative
- Claude Code is cheaper — bill breakdown and switching
- Cline is burning money — why the agent spends so much
- OpenClaw is expensive — why the agent burns through tokens and how to save
- OpenRouter: Cheap Alternative — Comparison with JoinGonka Gateway
- Best AI model for coding in 2026: comparison and prices
- Cheap alternative to GitHub Copilot without limits
- A cheap Windsurf alternative without credits or limits
- The cheapest API for AI agents in 2026
- ZCode: Cheap GLM inference instead of GLM Coding Plan
- JetBrains IDE + JoinGonka Gateway — your own endpoint instead of credits
- GitHub Copilot BYOK — own models instead of quotas
- Zed + JoinGonka Gateway — cheap inference in your editor
- Pi + JoinGonka Gateway — terminal agent on cheap inference
- Codex CLI: your own key instead of a subscription
- DeepSeek Harness: Your Own Provider via JoinGonka Gateway
- MiniMax Code: MiniMax agent with your own key via Gonka
- Warp + JoinGonka Gateway — terminal agent on your own endpoint
- Trae + JoinGonka Gateway — Gonka network models in AI-IDE
- Cherry Studio + JoinGonka Gateway — desktop AI client
- omp (Oh My Pi) + JoinGonka Gateway: an agent with model roles
- OpenHands + JoinGonka Gateway: agent on your own endpoint
- Qwen Code after the closure of qwen-oauth: working via JoinGonka Gateway
- Goose + JoinGonka Gateway: your own provider and key in the keyring
- Crush + JoinGonka Gateway: Charm agent on Gonka network models
- Zoo Code + JoinGonka Gateway: Migrating from Roo Code to Gonka models
- Kimi Code CLI: Moonshot AI agent on your key via Gonka
- Factory Droid + JoinGonka Gateway: BYOK on Gonka network models
- MiMo Code + JoinGonka Gateway: Xiaomi agent on Gonka network models
Tools
Kimi Code CLI: Moonshot AI agent on your key via Gonka
Kimi Code CLI (kimi command) is a terminal development agent from Moonshot AI: it reads and edits code, runs commands, searches through files, and delegates subtasks to subagents. The current version replaces the kimi-cli Python client: it is written in TypeScript, installed as a single binary, and released under the MIT license. Everything below is based on version 2.0.2 from September 19, 2026—on which we went from installation to agent response by September 23.
Normally, Kimi Code runs on Kimi models via account login (/login), but it can support any number of providers: any OpenAI-compatible endpoint can be described in a few lines in ~/.kimi-code/config.toml, and a Moonshot account is not needed for this. JoinGonka Gateway is such an endpoint for the decentralized Gonka network.
First things first: there are no Kimi models in the Gonka network right now—Kimi K2.6 served the network from May to September 2026. Today, it features DeepSeek V4 Flash, GLM-5.3 Flash, and MiniMax M2.7, and this guide is about how to run Kimi Code on them. The agent remains the same, while the model and price change: $0.0069 per million input tokens, uniform for all network models. After verifying your address, you will receive 3M free tokens in your account—enough to try all this yourself.
Quick Start: Installation and One Command
Step 1: install Kimi Code. The official methods from the documentation:
# macOS and Linux: prebuilt binary, no Node.js needed
curl -fsSL https://code.kimi.com/kimi-code/install.sh | bash
# Windows (PowerShell); Git for Windows is required before the first run
irm https://code.kimi.com/kimi-code/install.ps1 | iex
# npm (Node.js 22.19 or newer)
npm install -g @moonshot-ai/kimi-codeThe script puts the binary in ~/.kimi-code/bin/kimi and appends that directory to PATH via your shell configuration file. Its builds are glibc-only: on Alpine and other musl systems it will stop and suggest npm. Reopen your terminal and check: kimi --version.
Step 2: get a key. Sign up at gate.joingonka.ai/register, confirm your address, and create a key with the jg- prefix under "API keys".
Step 3: run the installer.
npx @joingonka/setup --tool kimi-codeThe installer will ask for your key (it isn't passed via command-line arguments) and write only its own tables into ~/.kimi-code/config.toml:
- a
[providers.joingonka]provider of typeopenai— that's the Chat Completions protocol — with the gateway address and the key in theapi_keyfield; the file will be set to permissions600; - one
[models."joingonka/…"]table per network model — with the real context window and response cap, plus explicit capabilities: tool calling for all, reasoning for DeepSeek V4 Flash and GLM-5.3 Flash, tools only for MiniMax M2.7; default_model— DeepSeek V4 Flash, if no default model is selected yet or it points to a model that has left the gateway's network; it won't touch another provider's choice without the--modelflag and will suggest a command for a one-off run instead;- it will save a copy of the previous file, leave other providers and your fields in the gateway tables untouched, and finally send a live request to the gateway to tell you whether the key, address, and model were accepted.
Here's what the output looked like in our run (abridged):
Configured provider "joingonka" in ~/.kimi-code/config.toml
Base URL: https://gate.joingonka.ai/v1 (provider type "openai" — Chat Completions)
Models: joingonka/MiniMaxAI/MiniMax-M2.7, joingonka/deepseek-ai/DeepSeek-V4-Flash-0731, joingonka/zai-org/GLM-5.3-Flash
Default model: joingonka/deepseek-ai/DeepSeek-V4-Flash-0731
…
✓ Verified: the gateway accepted the key, base URL and model.Set a different default model with the --model flag (glm, minimax, or a full identifier); the no-questions mode takes the key from an environment variable:
JOINGONKA_API_KEY=jg-your-key npx @joingonka/setup --tool kimi-code --model glm --non-interactiveThe installer automatically accounts for a data directory relocated via the KIMI_CODE_HOME variable. But ~/.kimi is the directory of the older Python client kimi-cli: Kimi Code 2.x doesn't read it, and on first launch offers to migrate settings from there (the same is done by kimi migrate).
Manual Configuration: config.toml
Everything the installer does can also be done by hand. Kimi Code stores its settings in TOML; table keys containing a dot or a slash go in quotes. Here's the complete snippet for the Gonka network:
default_model = "joingonka/deepseek-ai/DeepSeek-V4-Flash-0731"
[providers.joingonka]
type = "openai"
base_url = "https://gate.joingonka.ai/v1"
api_key = "jg-your-key"
[models."joingonka/deepseek-ai/DeepSeek-V4-Flash-0731"]
provider = "joingonka"
model = "deepseek-ai/DeepSeek-V4-Flash-0731"
max_context_size = 380000
max_output_size = 32768
capabilities = ["thinking", "tool_use"]
display_name = "DeepSeek V4 Flash (Gonka)"
[models."joingonka/zai-org/GLM-5.3-Flash"]
provider = "joingonka"
model = "zai-org/GLM-5.3-Flash"
max_context_size = 390000
max_output_size = 8192
capabilities = ["thinking", "tool_use"]
display_name = "GLM 5.3 Flash (Gonka)"
# MiniMax M2.7, unlike DeepSeek and GLM, is declared without thinking
[models."joingonka/MiniMaxAI/MiniMax-M2.7"]
provider = "joingonka"
model = "MiniMaxAI/MiniMax-M2.7"
max_context_size = 200000
max_output_size = 8192
capabilities = ["tool_use"]
display_name = "MiniMax M2.7 (Gonka)"| Field | Value | What matters |
|---|---|---|
default_model | model alias | Points to the name of the [models."…"] table, not to the model identifier |
max_context_size | model window | Required field; the agent uses it to decide when to compact the context |
max_output_size | response ceiling | Goes into the request as max_tokens — we verified this on the wire. Without the field, Kimi Code would request up to 131,072 response tokens, more than the network's models deliver |
capabilities | thinking, tool_use | Kimi Code doesn't guess the capabilities of unfamiliar models by name — declare them explicitly |
Where to keep the key. Kimi Code does not take keys from the shell environment: export OPENAI_API_KEY=… has no effect on it. There are two ways, and you can't use both at once:
| Method | Entry | How it behaves |
|---|---|---|
| Key in the file | api_key = "jg-…" | This is what the installer writes; works from any environment, keep the file with 600 permissions |
| Variable name | api_key_env = "JOINGONKA_API_KEY" | The value is read from the process environment on every request; without the variable, startup fails with an error that names it |
If both fields appear in the same table, Kimi Code will reject them at startup: has both apiKey and apiKeyEnv set in config.toml - they are mutually exclusive. Note that kimi doctor does not catch this conflict — it only checks the shape of the file.
Which provider type to choose. Kimi Code supports several protocols, and so does the gateway:
type | base_url | When to choose it |
|---|---|---|
openai | https://gate.joingonka.ai/v1 | The main option: the gateway's canonical path, the one the installer writes |
anthropic | https://gate.joingonka.ai | Anthropic Messages format, the client appends the /v1/messages path itself; the full tool-use cycle passed in our run |
In Kimi Code, one provider speaks one protocol, so a second protocol means a second provider under a different name. You can add and remove providers interactively with the /provider command inside the agent.
Models, Thinking, and sub-agents
A model in Kimi Code is an alias, the name of a table [models."…"]. The installer names them using the pattern joingonka/<model-id>, just as Kimi Code names its own (kimi-code/k3), and it's precisely the alias that's needed wherever a model is selected. Kimi Code won't find an identifier without a prefix, like MiniMaxAI/MiniMax-M2.7.
| Where | How | Persists |
|---|---|---|
| config.toml | default_model = "joingonka/deepseek-ai/DeepSeek-V4-Flash-0731" | yes, for new sessions |
| Launch flag | kimi -m "joingonka/zai-org/GLM-5.3-Flash" | no, this run only |
| Interface | /model, select from the list | Enter rewrites default_model, Alt+S applies to the current session only |
In our run, /model looked like this:
Select a model (type to search)
Tab toggle provider · ↑↓ navigate · Enter select · Alt+S session-only · Esc cancel
All joingonka
MiniMax M2.7 (Gonka) joingonka
❯ DeepSeek V4 Flash (Gonka) joingonka ← current
GLM 5.3 Flash (Gonka) joingonka
Thinking (←→ to switch)
[ On ] OffThinking toggle. The reasoning switch below the list on network models doesn't change anything by itself: Kimi Code only sends the reasoning level to models that declare levels. For GLM-5.3 Flash, it's worth doing — its reasoning is disabled by the value low (for more details, see the model overview). Add three fields to the model table:
[models."joingonka/zai-org/GLM-5.3-Flash"]
# … fields written by the installer …
support_efforts = ["low", "high"]
default_effort = "high"
off_effort = "low"We verified what goes to the gateway: with Thinking on, Kimi Code sends reasoning_effort: "high", and the model reasons; with it off, "low", and the response comes back without a reasoning block. Re-running the installer won't touch these fields: it only edits its own keys.
Model for subagents. Built-in subagents (coder, explore, plan) use the main agent's model by default. The [secondary_model] section gives them a pool from which the main agent selects on its own, guided by hints. Any alias works for the pool — DeepSeek V4 Flash, GLM-5.3 Flash, or MiniMax M2.7:
[secondary_model]
default_model = "joingonka/deepseek-ai/DeepSeek-V4-Flash-0731"
[secondary_model.models]
"joingonka/deepseek-ai/DeepSeek-V4-Flash-0731" = "Long context and long output: big reads and edits."
"joingonka/zai-org/GLM-5.3-Flash" = "Reasoning model for tricky logic and debugging."Pool keys must match the aliases in [models]: according to the documentation, a dangling reference prevents creating a session. The line for MiniMax M2.7 is added the same way; references to models that have left the network are cleaned up by the installer on the next run.
Which model to choose. The price is the same for all network models, so the choice is about behavior. Here's how they behaved in Kimi Code 2.0.2 on the task "read the file and find the bug" (verified September 23, 2026):
| Model | Context / response | Behavior in Kimi Code |
|---|---|---|
| DeepSeek V4 Flash | 380K / 32768 | The installer's choice: the largest window and the highest response ceiling on the network. Solves the task correctly |
| GLM-5.3 Flash | 390K / 8192 | A reasoning model: reasoning runs in a separate stream (in kimi -p — to stderr), the answer goes to stdout. Reasoning counts toward the response limit |
| MiniMax M2.7 | 200K / 8192 | Solves the task correctly in two steps; reasoning runs in a separate stream (in kimi -p — to stderr), stdout gets only the answer |
Verification: what should happen
First, make sure Kimi Code sees the provider and the file parses without errors:
$ kimi provider list
joingonka type=openai models=3 source=inline
Default model: joingonka/deepseek-ai/DeepSeek-V4-Flash-0731
$ kimi doctor
OK config.toml ~/.kimi-code/config.tomlNever publish the output of kimi provider list --json anywhere: the key in it sits there in plain text.
Next, a one-off run. Drop a calc.py into a directory with an addition function that actually subtracts, and ask it to find the bug:
kimi -p "Read calc.py and tell me in one sentence whether it has a bug."In -p mode, the answer goes to stdout while the reasoning and tool trace go to stderr, so the output is easy to parse with a script. Here's how GLM-5.3 Flash answered:
• Yes: the `add` function in calc.py:2 returns `a - b` instead of `a + b`, so `add(2, 3)` prints `-1`.Keep in mind: in -p mode Kimi Code asks nothing and executes commands on its own — in our run the model tried to launch python3 without asking to check its guess. The first launch of the interface in a new directory begins with a Trust this folder? prompt — it concerns the project's MCP servers — and the welcome screen shows the model: Model: DeepSeek V4 Flash (Gonka). The second half of the check is the gateway dashboard: in the "Usage" section the request will appear broken down by model and by key.
If something went wrong, the diagnosis is usually readable right from the message:
| What you see | What it means | What to do |
|---|---|---|
No model configured. Run `kimi` and use /login to sign in … | The provider isn't configured or default_model is empty | Run the installer or fill in default_model |
Model "deepseek-ai/DeepSeek-V4-Flash-0731" is not configured in config.toml. | A model ID was passed to -m instead of an alias | Add the prefix: joingonka/deepseek-ai/… |
provider.auth_error: 401 Invalid API key | The gateway rejected the key | Check api_key in [providers.joingonka]: Kimi Code doesn't read shell variables |
declares api_key_env = "JOINGONKA_API_KEY" … but the environment variable is not set or is empty | The variable-based method was chosen, but it's missing from the environment | Export the variable before launching, or go back to api_key |
429 … currently overloaded in the Gonka network (rate limit) | The model has run out of free capacity on the network | Kimi Code retries such errors on its own; if it drags on, switch models with /model. Status is visible on the status page |
How much it costs
For every phrase you send, Kimi Code adds a system prompt and descriptions of two dozen built-in tools, and a task usually takes several steps. In our run, each step carried about 20,000 input tokens, and "read the file, find the error" took two steps — about 40,000 tokens, almost all on input. Therefore, the price per token is the deciding factor here.
Through the JoinGonka Gateway, tokens cost $0.0069 per million for input and $0.021 per million for output — the price is the same for all models in the network and is pulled onto this page from a live source.
| Scenario | Consumption | Via Gateway |
|---|---|---|
| One-off task: read a file, find an error | ~40K tokens | hundredths of a cent |
| Day of active work | 3-7M tokens | a few cents |
| Month of active development | ~150M tokens | about a dollar |
The estimates in the right column are based on September 2026 prices. For comparison — how you can generally pay for models in Kimi Code:
| Method | How it connects | What is needed |
|---|---|---|
| Kimi Code (OAuth) | /login, device code login | Kimi account; limits and pricing are subject to service terms |
| Kimi Platform | /login with platform API key | Key from platform.kimi.com or platform.kimi.ai, pay per platform price list |
| JoinGonka Gateway | provider in config.toml | jg-… key; Moonshot account not required, pay for actual tokens, usage is visible in the dashboard |
Exact consumption and balance are in the dashboard, in the "Usage" and "Billing" sections. Kimi Code monitors the length of the conversation itself: the status bar shows what share of the context is occupied, and the /compact command compresses the history manually.
Things to consider when working
Permission modes. Kimi Code defaults to the Always Ask mode: reads are performed immediately, while edits and commands require your confirmation. Ask When Needed (/yolo or --yolo flag) skips routine edits and commands but prompts for sensitive files and dangerous commands; Never Ask (/auto, --auto) asks for nothing. The planning mode is enabled via Shift-Tab. A one-off run of kimi -p always proceeds without questions, so run it inside a container or a separate working copy.
What works without a Moonshot account. Your provider handles everything the model does: reading and editing code, commands, subagents, sessions, and MCP servers. Web search is a Moonshot service: without logging into an account, the WebSearch tool is not available in the agent's set, but loading a page by address (FetchURL) remains. Network models are text-based, so inputting images and videos does not work with them.
Telemetry and updates. Anonymous telemetry is enabled by default: you can disable it with the telemetry = false line in config.toml or the KIMI_DISABLE_TELEMETRY=1 variable. Updates are installed automatically ([upgrade] auto_install = true in ~/.kimi-code/tui.toml); if you need predictability, turn it off and update using the kimi upgrade command. The gateway does not store the content of prompts and responses—only consumption aggregates remain in the statistics.
Multiple environments. Kimi Code does not have provider settings at the project level. You can separate work and personal keys using the KIMI_CODE_HOME variable: with it, settings, sessions, and logs move to a different directory, and the installer, when run with the same variable, writes to the same place.
In the editor. Kimi Code works inside Zed, JetBrains, and other ACP clients via the kimi acp command — using the same data directory and the same providers.
Labs whose models are served by the network also release their own agents: MiniMax has a terminal MiniMax Code, DeepSeek has DeepSeek Harness, and Z.ai, the creators of GLM, have ZCode. All of them connect to the same gateway with the same key.
npx @joingonka/setup --tool kimi-code: the installer will write the joingonka provider with type openai and the key, model tables with fair limits, and the default model to ~/.kimi-code/config.toml, then check the connection with a live request. The model is selected by the alias joingonka/<id> — in /model or via the -m flag; to make the Thinking toggle work for GLM-5.3 Flash, declare support_efforts and off_effort = "low". Verification is done via kimi provider list and kimi -p on a file with an error; you pay for actual tokens at the unified network price.Want to learn more?
Explore other sections or start earning GNK right now.
Get a key and free tokens →