Knowledge Base Sections ▾
Navigation
▸ Start here By rolesCategories
- Cursor + Gonka AI - cheap LLM for coding
- Claude Code + Gonka AI - LLM for the terminal
- OpenClaw + Gonka AI - affordable AI agents
- OpenCode: your own model in the terminal
- Continue.dev + Gonka AI - AI for VS Code/JetBrains
- Cline + Gonka AI - AI agent in VS Code
- Aider + Gonka AI - pair programming with AI
- LangChain + Gonka AI - AI applications for pennies
- n8n + Gonka AI - automation with cheap AI
- Open WebUI + Gonka AI - your own ChatGPT
- LibreChat + Gonka AI — open-source ChatGPT
- Hermes Agent + DeepSeek on the Gonka network — an autonomous agent for pennies
- Kilo Code + Gonka AI — AI-Agent in VS Code
- Roo Code + Gonka AI — Autonomous AI Agent in VS Code
- LlamaIndex + Gonka AI — RAG applications for pennies
- PydanticAI + Gonka — typed AI agents for pennies
- Vercel AI SDK + Gonka AI — AI applications in TypeScript for pennies
- TanStack AI + Gonka — AI applications in TypeScript for pennies
- API quick start — curl, Python, TypeScript
- JoinGonka Gateway — a full overview
- Management Keys — SaaS on Gonka
- Cheapest AI API: Provider Comparison 2026
- How to buy AI tokens and an API key: 3 methods in 2026
- Cursor Pro request limit reached — breakdown and cheaper alternative
- Claude Code is cheaper — bill breakdown and switching
- Cline is burning money — why the agent spends so much
- OpenClaw is expensive — why the agent burns through tokens and how to save
- OpenRouter: Cheap Alternative — Comparison with JoinGonka Gateway
- Best AI model for coding in 2026: comparison and prices
- Cheap alternative to GitHub Copilot without limits
- A cheap Windsurf alternative without credits or limits
- The cheapest API for AI agents in 2026
- ZCode: Cheap GLM inference instead of GLM Coding Plan
- JetBrains IDE + JoinGonka Gateway — your own endpoint instead of credits
- GitHub Copilot BYOK — own models instead of quotas
- Zed + JoinGonka Gateway — cheap inference in your editor
- Pi + JoinGonka Gateway — terminal agent on cheap inference
- Codex CLI: your own key instead of a subscription
- DeepSeek Harness: Your Own Provider via JoinGonka Gateway
- MiniMax Code: MiniMax agent with your own key via Gonka
- Warp + JoinGonka Gateway — terminal agent on your own endpoint
- Trae + JoinGonka Gateway — Gonka network models in AI-IDE
- Cherry Studio + JoinGonka Gateway — desktop AI client
- omp (Oh My Pi) + JoinGonka Gateway: an agent with model roles
- OpenHands + JoinGonka Gateway: agent on your own endpoint
- Qwen Code after the closure of qwen-oauth: working via JoinGonka Gateway
- Goose + JoinGonka Gateway: your own provider and key in the keyring
- Crush + JoinGonka Gateway: Charm agent on Gonka network models
- Zoo Code + JoinGonka Gateway: Migrating from Roo Code to Gonka models
- Kimi Code CLI: Moonshot AI agent on your key via Gonka
- Factory Droid + JoinGonka Gateway: BYOK on Gonka network models
- MiMo Code + JoinGonka Gateway: Xiaomi agent on Gonka network models
Tools
Warp + JoinGonka Gateway — terminal agent on your own endpoint
Warp is a terminal that has grown into an agentic development environment: the built-in Warp Agent reads command output, edits files, runs builds, and manages tasks from prompt to result. By default, the agent runs on models sold by Warp itself—using plan credits—and for a long time, bringing your own inference was only possible on paid plans using an OpenAI, Anthropic, or Google key. On May 20, 2026, Warp opened BYOK on the free plan and added a custom inference endpoint—the ability to connect any service that speaks OpenAI Chat Completions.
The second method is exactly what is needed to route the Warp agent through JoinGonka Gateway: the gateway provides access to models from the decentralized Gonka network via an OpenAI-compatible API, inference is paid for at real-time network rates, and for individuals or small teams, Warp credits are not consumed. After registration and email verification, you will receive 3M free tokens, enough to test the agent on your own tasks before your first top-up.
Below is what you need to get started, what values to enter into the Warp form, how to verify that requests are going through the gateway, which network model to choose for terminal tasks, and what is important to know about the request path: Warp handles it differently than most tools.
What you will need: key, Warp plan, and a public address
JoinGonka Key. Register on the gateway, open the "API Keys" section in your dashboard, and click "Create key". The key starts with jg- and is shown only once—save it immediately. It is convenient to create a separate key for Warp: this way, usage in the "Usage" section for the terminal agent will be visible as a separate line item, rather than mixed with other tools.
Warp Plan. Your own endpoint is available not just on paid plans—conditions depend on the team size:
| Who is connecting | Is your own endpoint available? | What about Warp credits? |
|---|---|---|
| Single developer or company under 10 people (Free, Build, Max plans) | Yes, on any of these plans | Credits are not deducted for inference via your own endpoint |
| Organization of over 10 people | Only on Business or Enterprise | Local agent runs consume platform credits—a fee for Warp infrastructure at a reduced rate; the inference itself is paid to the endpoint provider |
Public HTTPS Address. This is a Warp requirement that is easy to miss: requests to your endpoint are sent not by the application on your machine, but by Warp's servers, so the address must be accessible from the internet. The form rejects localhost, 127.0.0.1, and private network addresses, and only the https:// scheme is allowed. JoinGonka Gateway is a public HTTPS service, so tunnels like ngrok, which Warp documentation suggests for local models, are not needed here.
Latest application version. Your own endpoint was included in the stable build starting from the May 20, 2026 release, and the API schema selection in the form appeared in the July 31, 2026 release. If the API schema field is missing in the form, update Warp.
Connection: Add custom endpoint form
Everything is configured in the app interface — no need to edit files by hand.
- Open Settings and type
inference endpointin the settings search — Warp will jump to the section with provider keys and your own endpoints (it lives on the Warp Agent settings page). - Click Add custom endpoint — a form will open.
- Fill in the fields with the values from the table and click Add endpoint.
| Form field | What to enter | Notes |
|---|---|---|
| API schema | OpenAI Chat Completions | The default schema: it's the one Warp's documentation describes the URL format for. The list also includes OpenAI Responses and Anthropic Messages |
| Endpoint name | JoinGonka | Any name — this is how the endpoint appears in settings |
| Endpoint URL | https://gate.joingonka.ai/v1 | The base URL together with /v1 — exactly as in the example from Warp's documentation |
| API key | jg-your-key | Stored only on your device, in the system keychain |
| Model name | deepseek-ai/DeepSeek-V4-Flash-0731 | The exact model ID as returned by the gateway |
| Model alias (optional) | DeepSeek V4 Flash | A short name for the model switcher |
The + Add model button adds another row. Add all the network's models at once so you can switch between them later without going back into settings: deepseek-ai/DeepSeek-V4-Flash-0731, MiniMaxAI/MiniMax-M2.7, and zai-org/GLM-5.3-Flash. The current list is always available at GET https://gate.joingonka.ai/v1/models. Enter the IDs character by character: the form only validates the URL format and that the fields are filled in — it doesn't send a test request on save, so a typo in the model name will only surface later in a conversation with the agent.
Once saved, the models appear in the agent's model switcher. If the default model was still one of Warp's models, the app may offer to change it — accept and pick your endpoint's model. One important detail: the Auto options in the switcher always run on Warp's infrastructure and burn its credits, even when your own endpoint is configured. For a request to reach the gateway, a specific model from your list must be selected in the switcher.
Verification: checking if requests are going through the gateway
Choose your endpoint's model in the switcher and give the agent a short task, for example: "Show me the five largest files in the current directory and explain the command." The agent should propose a command, run it with your permission, and comment on the output.
Now open the gateway dashboard and go to the "Usage" section. Under "By keys", the key you created for Warp will show requests and the time of the last call; under "By models", the selected model. If the counters are growing, inference is going through JoinGonka Gateway, not burning Warp credits. For its part, Warp shows token usage right in the chat with the agent — under the short names you set in the alias field.
If something goes wrong, the reason is usually clear from the response:
| What you see | What it means | What to do |
|---|---|---|
401, Invalid API key | The gateway didn't recognize the key | Check that the key starts with jg-, was pasted without spaces, and is active in the "API keys" section; if in doubt, create a new one |
404, Invalid URL (POST …) | The path has an extra segment; the error message shows exactly where the request landed | It should be /v1/chat/completions. If /v1 is duplicated or the full path ended up in the address, set the Endpoint URL to the form shown in the table above |
405 Not Allowed or an HTML page instead of a response | /v1 is missing from the address, so the request missed the API | Add /v1 to the end of the address |
429 | The model is overloaded: at peak hours, the capacity of a specific model on the network can be fully taken | Retry in a few seconds or switch to a neighboring model — that's exactly why you should set up all of them at once |
| Requests go through, but Warp credits melt away | The switcher has Auto or Warp's model router selected, not the endpoint's model | Select a specific model from your list |
| The form won't accept the address | Warp requires https:// and a public host | Enter the full gateway address, starting with https:// |
| The model "doesn't see" an attached image | The network's models are text-only: the gateway replaces the image with a text note | For tasks with screenshots, switch to a vision model; for commands and code, stay with a network model |
| The endpoint's model is missing on a cloud run | Your endpoint is stored locally and doesn't carry over to Cloud Agents | Run the task with a local agent |
The first response in a session may take a few seconds: the agent sends a large system prompt describing the tools, and the request first passes through Warp's servers and then a network node. After that, the conversation moves faster.
Which model to choose for a terminal agent
All network models have the same price, so your choice is about behavior, not budget. The terminal agent reads a lot (command output, logs, files) and calls tools frequently; all models in the table support native tool calling.
| Model | Identifier | Context and Output | When to use in Warp |
|---|---|---|---|
| DeepSeek V4 Flash | deepseek-ai/DeepSeek-V4-Flash-0731 | 380K, output up to 32768 tokens | Default agent model: long sessions, large logs and diffs, edits across multiple files. Context headroom and the largest output limit in the network |
| MiniMax M2.7 | MiniMaxAI/MiniMax-M2.7 | 200K, output up to 8192 tokens | Quick short tasks: finding a command, troubleshooting errors, editing configs |
| GLM-5.3 Flash | zai-org/GLM-5.3-Flash | 390K, output up to 8192 tokens | Reasoning model: thinks before answering. Complex debugging, migration planning, analyzing unfamiliar code; the answer takes longer, and part of the output limit is used for reasoning |
Practical scheme: DeepSeek V4 Flash is the primary model, MiniMax M2.7 is for small tasks, GLM-5.3 Flash is for when the task is logic-intensive rather than volume-intensive. Switching takes a second because all models are already configured in the endpoint, and short aliases prevent confusion with long identifiers. More details on the most popular one in the DeepSeek V4 Flash overview.
How much it costs
The terminal agent has higher token consumption than a chat: for every request, Warp adds system instructions, dialog context, and tool descriptions, and one task consists of multiple turns. Therefore, the price per token matters more here than in casual chatting.
Via JoinGonka Gateway, tokens cost $0.0069 per million for input and $0.021 per million for output—the price is the same for all network models and is pulled onto this page from a live source. Price scale as of September 2026:
| Scenario | Usage | Via Gateway |
|---|---|---|
| One-time task: command and output parsing | tens of thousands of tokens | fractions of a cent |
| Day of active work with the agent | 3-7M tokens | a few cents |
| Month of daily work | ~150M tokens | around one or two dollars |
For comparison, three ways to pay for the Warp agent:
| Method | Who calculates inference | What limits you |
|---|---|---|
| Warp Models | Warp, via plan credits | Monthly plan credit allowance |
| BYOK: OpenAI, Anthropic, or Google key | Model vendor, per their pricing | Vendor's price per million tokens |
| Custom endpoint + JoinGonka Gateway | Gateway: per token, from balance | Balance only: usage is visible in the dashboard, no request rate limits |
Warp itself does not charge credits for individuals and teams of up to 10 people when using your own endpoint—you only pay for tokens at the gateway. The balance is topped up in your dashboard, where you can also see your remaining balance and daily usage by model.
Request routing and privacy: what you need to know
For most tools, a custom endpoint means “requests go directly from my machine to the provider.” Warp is different: the executing part of the agent runs on Warp servers, and a custom endpoint only changes the final destination and the key. According to the Warp documentation, the request path is as follows:
- The application retrieves the endpoint address and key from a secure storage on your device and sends them to Warp servers along with your prompt.
- The agent on the Warp side assembles the complete request—system instructions, chat context, tools—and calls your endpoint with that key.
- The response is returned to the application as a stream through the same servers.
Practical implications:
- Key. It is stored only on your device. It passes through Warp servers with every request, but according to Warp, it is not saved there: it is used on the fly and discarded.
- Prompts and responses. They pass in transit through the Warp backend. Warp states that it does not use this content for training, and storage and analytics are subject to your account's privacy and telemetry settings. It cannot extend the ZDR guarantee, which applies to Warp's own models, to your custom endpoint: what the provider stores is determined by the provider.
- Gateway side. The JoinGonka Gateway does not store conversations: prompts and responses do not remain on the gateway after the response, and only request and token counters are visible in the dashboard.
- What remains on Warp infrastructure. Auto options, custom model routers, Codebase Context, Active AI suggestions, and cloud agents—a custom endpoint does not affect these.
If the path through third-party servers is unacceptable for your tasks, use an agent that accesses the endpoint directly from your machine, such as Codex CLI or any tool from the API quickstart guide. In its announcement, Warp itself promises a client-side option: a lightweight agent module in Rust in an open repository that will connect to local models without routing through Warp servers, as well as support for the Agent Client Protocol. And as of the September 2, 2026 release, endpoint descriptions are stored in a configuration file in a format common to both the application and the Warp Agent CLI; keys do not end up in this file and remain in secure storage.
Want to learn more?
Explore other sections or start earning GNK right now.
Get key and free tokens →