Knowledge Base Sections ▾
Navigation
▸ Start here By rolesCategories
- Cursor + Gonka AI - cheap LLM for coding
- Claude Code + Gonka AI - LLM for the terminal
- OpenClaw + Gonka AI - affordable AI agents
- OpenCode: your own model in the terminal
- Continue.dev + Gonka AI - AI for VS Code/JetBrains
- Cline + Gonka AI - AI agent in VS Code
- Aider + Gonka AI - pair programming with AI
- LangChain + Gonka AI - AI applications for pennies
- n8n + Gonka AI - automation with cheap AI
- Open WebUI + Gonka AI - your own ChatGPT
- LibreChat + Gonka AI — open-source ChatGPT
- Hermes Agent + DeepSeek on the Gonka network — an autonomous agent for pennies
- Kilo Code + Gonka AI — AI-Agent in VS Code
- Roo Code + Gonka AI — Autonomous AI Agent in VS Code
- LlamaIndex + Gonka AI — RAG applications for pennies
- PydanticAI + Gonka — typed AI agents for pennies
- Vercel AI SDK + Gonka AI — AI applications in TypeScript for pennies
- TanStack AI + Gonka — AI applications in TypeScript for pennies
- API quick start — curl, Python, TypeScript
- JoinGonka Gateway — a full overview
- Management Keys — SaaS on Gonka
- Cheapest AI API: Provider Comparison 2026
- How to buy AI tokens and an API key: 3 methods in 2026
- Cursor Pro request limit reached — breakdown and cheaper alternative
- Claude Code is cheaper — bill breakdown and switching
- Cline is burning money — why the agent spends so much
- OpenClaw is expensive — why the agent burns through tokens and how to save
- OpenRouter: Cheap Alternative — Comparison with JoinGonka Gateway
- Best AI model for coding in 2026: comparison and prices
- Cheap alternative to GitHub Copilot without limits
- A cheap Windsurf alternative without credits or limits
- The cheapest API for AI agents in 2026
- ZCode: Cheap GLM inference instead of GLM Coding Plan
- JetBrains IDE + JoinGonka Gateway — your own endpoint instead of credits
- GitHub Copilot BYOK — own models instead of quotas
- Zed + JoinGonka Gateway — cheap inference in your editor
- Pi + JoinGonka Gateway — terminal agent on cheap inference
- Codex CLI: your own key instead of a subscription
- DeepSeek Harness: Your Own Provider via JoinGonka Gateway
- MiniMax Code: MiniMax agent with your own key via Gonka
- Warp + JoinGonka Gateway — terminal agent on your own endpoint
- Trae + JoinGonka Gateway — Gonka network models in AI-IDE
- Cherry Studio + JoinGonka Gateway — desktop AI client
- omp (Oh My Pi) + JoinGonka Gateway: an agent with model roles
- OpenHands + JoinGonka Gateway: agent on your own endpoint
- Qwen Code after the closure of qwen-oauth: working via JoinGonka Gateway
- Goose + JoinGonka Gateway: your own provider and key in the keyring
- Crush + JoinGonka Gateway: Charm agent on Gonka network models
- Zoo Code + JoinGonka Gateway: Migrating from Roo Code to Gonka models
- Kimi Code CLI: Moonshot AI agent on your key via Gonka
- Factory Droid + JoinGonka Gateway: BYOK on Gonka network models
- MiMo Code + JoinGonka Gateway: Xiaomi agent on Gonka network models
Tools
Codex CLI: your own key instead of a subscription
Codex CLI is an OpenAI agentic assistant living right in the terminal: it reads project files, runs commands in a sandbox, edits code, and explains what it did. By default, it accesses OpenAI models via a ChatGPT subscription, but you can replace the provider with your own — and then the same scenarios are processed by the Gonka network.
One nuance makes Codex special among agentic tools: it doesn't speak the familiar Chat Completions, but Responses API — OpenAI's new protocol. That is precisely why it was previously impossible to connect it to a third-party gateway. JoinGonka Gateway has been accepting Responses since August 22, 2026, and everything described below has been verified with live runs, not by retelling documentation.
What is Codex CLI and how is it different from chat
Codex isn't an editor autocomplete — it's an autonomous executor. You describe the task in words, and it decides on its own which files to read, which commands to run, and what to change. It has about a dozen tools at its disposal: running commands, reading files, maintaining a work plan, asking the user for clarification, and spawning child agents for subtasks.
Two modes of operation. Interactive — codex with no arguments, a terminal dialogue showing every step. Non-interactive — codex exec "task", a one-off run to completion: handy in scripts and CI.
Safety is built in: by default the agent runs in a sandbox and asks for confirmation before making changes. The level of control is set in the configuration — from "ask about everything" to a fully autonomous run with write access to the working directory.
Another distinguishing feature is conversation state. Codex doesn't ask the server to remember anything: it keeps the entire thread locally and sends it in full with every turn. For you, that means the conversation doesn't linger on the model provider's side, and changing the endpoint doesn't break an ongoing session.
The difference from Claude Code and other terminal agents is precisely the protocol. Codex talks to the model via the Responses API, where the conversation is described not as a flat list of messages but as a stream of items: text, tool call, tool result, reasoning block. For the tool, this yields a stricter dialogue model; for you, it means you need an endpoint that understands this protocol.
Connection: one edit to config.toml
The fast path — the installer. The command npx @joingonka/setup --tool codex writes a provider block into ~/.codex/config.toml with wire_api = "responses", the model and its real context window, preserving your comments and other settings, and then verifies the key and model with a live request. If another provider is already selected in the config, the installer won't touch it and will suggest a command for a one-off run. It puts the key directly in the file (permissions 600), in the experimental_bearer_token field — the only one where Codex can store a key as a literal; Codex itself considers it experimental. The manual option below gets by with an environment variable.
Codex stores settings in ~/.codex/config.toml. A custom provider is described by a [model_providers.*] block, and the key field here is wire_api: without it Codex will try to speak Chat Completions and won't complete the agentic loop.
# ~/.codex/config.toml
model_provider = "joingonka"
model = "deepseek-ai/DeepSeek-V4-Flash-0731"
[model_providers.joingonka]
name = "JoinGonka Gateway"
base_url = "https://gate.joingonka.ai/v1"
env_key = "JOINGONKA_API_KEY"
wire_api = "responses"The key isn't written to the file: env_key names the environment variable Codex will take it from. The variable name is your own, not OPENAI_API_KEY — that way the setup doesn't intercept your other tools working with OpenAI.
export JOINGONKA_API_KEY=jg-your-keyThe model field sets the default model, and model_provider — which of the described blocks to use. Both are overridden at launch time, so a single config can comfortably serve several providers: work, experimental and backup.
The key is issued in your account dashboard after registration, where you can also see your balance and spend. If you'd rather not touch the shared config, the same block can be passed one-off via -c: codex -c model_provider=joingonka …. And to keep several independent configurations, point the settings directory with the CODEX_HOME variable.
Verification: what should happen
The quickest way to confirm the integration works is a one-off run:
codex exec "Answer in one line: what is 17*3?"At the top of its reply, Codex prints who it is working with: the model, provider: joingonka, the sandbox mode and the session ID. If it shows your provider and an answer comes back at the bottom, the connection is live.
Next, check the thing Codex is really installed for — working with files. Drop a small file with an obvious bug into the directory and ask it to find the bug:
codex exec "Read calc.py and tell me in one sentence whether it has a bug."The agent should call the read tool on its own, open the file and give a substantive answer, line number included. If that happens, the full agent loop (request → tool call → result → answer) is routing correctly through the gateway.
If something goes wrong, the diagnosis is usually right there in the message:
| What you see | What it means | What to do |
|---|---|---|
401 Unauthorized: Invalid API key | Codex couldn't find the key, or picked up the wrong one | Check that the variable named in env_key is actually exported in the current shell and that its name matches the config |
404 on /responses | The suffix got lost in base_url | The address must end in /v1 — Codex appends /responses itself |
| The model replies, but tools are never called | wire_api isn't set, so the conversation is running on the old protocol | Add wire_api = \"responses\" to the provider block |
Reconnecting… 1/5 | Codex is retrying the request on its own | Normal behavior for a one-off network error; if the retries run out, read the error text below them |
| A warning about bubblewrap | The system is missing the sandboxing package | It doesn't block anything: Codex falls back to a built-in copy. To be tidy, install bubblewrap with your package manager |
The first request in a session can take a few seconds: Codex sends a large system prompt plus a description of all its tools, and the network node has to accept the task. Subsequent replies come faster.
Which model to choose
All network models are available at the same price, so the choice is about behavior, not budget. Below is the result of a live run of the same task (reading a file and finding an error in it) through Codex on DeepSeek V4 Flash and MiniMax M2.7; GLM-5.3 Flash joined the network after the run—its model properties are provided.
| Model | Identifier | Context | Codex Behavior |
|---|---|---|---|
| DeepSeek V4 Flash | deepseek-ai/DeepSeek-V4-Flash-0731 | 380K | Clear and concise response, specifying file and line. One of the longest contexts in the network and a response ceiling of 32768 tokens—buffer for large repositories |
| MiniMax M2.7 | MiniMaxAI/MiniMax-M2.7 | 200K | Solves the task correctly, but sometimes shows its reasoning out loud—in the terminal, this looks verbose |
| GLM-5.3 Flash | zai-org/GLM-5.3-Flash | 390K | Reasoning model with the longest context in the network: it reasons before every response, so the response is longer and arrives later. Calls tools, including subsequent rounds; for short tasks, set reasoning_effort: low |
The default recommendation is DeepSeek V4 Flash: agent work quickly hits context limits, and 380K tokens allow holding many files in memory at once. If the task requires thinking through complex logic, use GLM-5.3 Flash, but set aside a buffer for max_tokens: part of the response budget goes to reasoning. The model can be changed with a single model line in the config or the -c model=… flag without editing the file.
The current list of network models is always available via GET https://gate.joingonka.ai/v1/models.
How much does it cost
Agent tools consume tokens differently than chat: for each of your phrases, Codex adds a system prompt and a description of all tools, then carries out a multi-step dialog with the model. In a live run, a simple task like "read a file and find an error" cost about 18-20 thousand tokens. This is a normal price to pay for autonomy, which is exactly why the price per token matters.
Via JoinGonka Gateway, tokens cost $0.0069 per million input and $0.021 per million output—the price is identical for all network models and is fetched on this page from a live source.
| Scenario | Consumption | Via Gateway |
|---|---|---|
| One-off task (read file, find error) | ~20K tokens | fractions of a cent |
| Day of active work | 3-7M tokens | about a cent |
| Month of active development | ~150M tokens | a few cents |
For comparison, here is how paid options are structured at Codex itself and among competitors:
| Method | Payment model | Limitations |
|---|---|---|
| ChatGPT Subscription | fixed monthly amount | quotas on number of requests and refresh windows |
| OpenAI key directly | by tokens at vendor price | price per million tokens is three orders of magnitude higher |
| JoinGonka Gateway | by tokens, balance | usage visible in the dashboard, no request quotas |
Payment is for actual consumption, without a monthly subscription or request quotas: no five-hour windows, "prompt" limits, or peak-hour multipliers. The balance is topped up with crypto; the remainder and daily usage are visible in the personal dashboard. A detailed breakdown of the economics is in the article about the cheapest AI API.
What to keep in mind
Codex itself maintains the conversation history. It sends the entire history with every request and does not ask the server to remember anything — and we do not store chat logs. Your code and prompts do not remain on the gateway after a response is provided.
Web search is active. Codex declares a search tool in every request, and the gateway accepts it: the search is performed on our side, and the results are injected into the model's response.
Tools are local, not cloud-based. Codex executes commands and reads files locally on your machine, so access to your project does not depend on the model provider.
Sandbox. On Linux, Codex uses bubblewrap to isolate executed commands. If it is not on your system, Codex will warn you and use a built-in copy — this does not affect performance, but installing the package via your standard package manager is more convenient.
If you need image processing — interface screenshots, diagrams in photos — use a tool with a vision-capable model for such tasks: models in the Gonka network are text-only. This is not a limitation for code, commands, and files.
Other terminal agents, if Codex was not the right fit: the API quickstart shows how to connect any compatible tool in a couple of minutes.
Want to learn more?
Explore other sections or start earning GNK right now.
Get key and free tokens →