Knowledge Base Sections ▾

Navigation

▸ Start here By roles

Categories

Tools 52
Glossary 12

Tools

Warp + JoinGonka Gateway — terminal agent on your own endpoint

Warp is a terminal that has grown into an agentic development environment: the built-in Warp Agent reads command output, edits files, runs builds, and manages tasks from prompt to result. By default, the agent runs on models sold by Warp itself—using plan credits—and for a long time, bringing your own inference was only possible on paid plans using an OpenAI, Anthropic, or Google key. On May 20, 2026, Warp opened BYOK on the free plan and added a custom inference endpoint—the ability to connect any service that speaks OpenAI Chat Completions.

The second method is exactly what is needed to route the Warp agent through JoinGonka Gateway: the gateway provides access to models from the decentralized Gonka network via an OpenAI-compatible API, inference is paid for at real-time network rates, and for individuals or small teams, Warp credits are not consumed. After registration and email verification, you will receive 3M free tokens, enough to test the agent on your own tasks before your first top-up.

Below is what you need to get started, what values to enter into the Warp form, how to verify that requests are going through the gateway, which network model to choose for terminal tasks, and what is important to know about the request path: Warp handles it differently than most tools.

What you will need: key, Warp plan, and a public address

JoinGonka Key. Register on the gateway, open the "API Keys" section in your dashboard, and click "Create key". The key starts with jg- and is shown only once—save it immediately. It is convenient to create a separate key for Warp: this way, usage in the "Usage" section for the terminal agent will be visible as a separate line item, rather than mixed with other tools.

Warp Plan. Your own endpoint is available not just on paid plans—conditions depend on the team size:

Who is connectingIs your own endpoint available?What about Warp credits?
Single developer or company under 10 people (Free, Build, Max plans)Yes, on any of these plansCredits are not deducted for inference via your own endpoint
Organization of over 10 peopleOnly on Business or EnterpriseLocal agent runs consume platform credits—a fee for Warp infrastructure at a reduced rate; the inference itself is paid to the endpoint provider

Public HTTPS Address. This is a Warp requirement that is easy to miss: requests to your endpoint are sent not by the application on your machine, but by Warp's servers, so the address must be accessible from the internet. The form rejects localhost, 127.0.0.1, and private network addresses, and only the https:// scheme is allowed. JoinGonka Gateway is a public HTTPS service, so tunnels like ngrok, which Warp documentation suggests for local models, are not needed here.

Latest application version. Your own endpoint was included in the stable build starting from the May 20, 2026 release, and the API schema selection in the form appeared in the July 31, 2026 release. If the API schema field is missing in the form, update Warp.

Connection: Add custom endpoint form

Everything is configured in the app interface — no need to edit files by hand.

  1. Open Settings and type inference endpoint in the settings search — Warp will jump to the section with provider keys and your own endpoints (it lives on the Warp Agent settings page).
  2. Click Add custom endpoint — a form will open.
  3. Fill in the fields with the values from the table and click Add endpoint.
Form fieldWhat to enterNotes
API schemaOpenAI Chat CompletionsThe default schema: it's the one Warp's documentation describes the URL format for. The list also includes OpenAI Responses and Anthropic Messages
Endpoint nameJoinGonkaAny name — this is how the endpoint appears in settings
Endpoint URLhttps://gate.joingonka.ai/v1The base URL together with /v1 — exactly as in the example from Warp's documentation
API keyjg-your-keyStored only on your device, in the system keychain
Model namedeepseek-ai/DeepSeek-V4-Flash-0731The exact model ID as returned by the gateway
Model alias (optional)DeepSeek V4 FlashA short name for the model switcher

The + Add model button adds another row. Add all the network's models at once so you can switch between them later without going back into settings: deepseek-ai/DeepSeek-V4-Flash-0731, MiniMaxAI/MiniMax-M2.7, and zai-org/GLM-5.3-Flash. The current list is always available at GET https://gate.joingonka.ai/v1/models. Enter the IDs character by character: the form only validates the URL format and that the fields are filled in — it doesn't send a test request on save, so a typo in the model name will only surface later in a conversation with the agent.

Once saved, the models appear in the agent's model switcher. If the default model was still one of Warp's models, the app may offer to change it — accept and pick your endpoint's model. One important detail: the Auto options in the switcher always run on Warp's infrastructure and burn its credits, even when your own endpoint is configured. For a request to reach the gateway, a specific model from your list must be selected in the switcher.

Verification: checking if requests are going through the gateway

Choose your endpoint's model in the switcher and give the agent a short task, for example: "Show me the five largest files in the current directory and explain the command." The agent should propose a command, run it with your permission, and comment on the output.

Now open the gateway dashboard and go to the "Usage" section. Under "By keys", the key you created for Warp will show requests and the time of the last call; under "By models", the selected model. If the counters are growing, inference is going through JoinGonka Gateway, not burning Warp credits. For its part, Warp shows token usage right in the chat with the agent — under the short names you set in the alias field.

If something goes wrong, the reason is usually clear from the response:

What you seeWhat it meansWhat to do
401, Invalid API keyThe gateway didn't recognize the keyCheck that the key starts with jg-, was pasted without spaces, and is active in the "API keys" section; if in doubt, create a new one
404, Invalid URL (POST …)The path has an extra segment; the error message shows exactly where the request landedIt should be /v1/chat/completions. If /v1 is duplicated or the full path ended up in the address, set the Endpoint URL to the form shown in the table above
405 Not Allowed or an HTML page instead of a response/v1 is missing from the address, so the request missed the APIAdd /v1 to the end of the address
429The model is overloaded: at peak hours, the capacity of a specific model on the network can be fully takenRetry in a few seconds or switch to a neighboring model — that's exactly why you should set up all of them at once
Requests go through, but Warp credits melt awayThe switcher has Auto or Warp's model router selected, not the endpoint's modelSelect a specific model from your list
The form won't accept the addressWarp requires https:// and a public hostEnter the full gateway address, starting with https://
The model "doesn't see" an attached imageThe network's models are text-only: the gateway replaces the image with a text noteFor tasks with screenshots, switch to a vision model; for commands and code, stay with a network model
The endpoint's model is missing on a cloud runYour endpoint is stored locally and doesn't carry over to Cloud AgentsRun the task with a local agent

The first response in a session may take a few seconds: the agent sends a large system prompt describing the tools, and the request first passes through Warp's servers and then a network node. After that, the conversation moves faster.

Which model to choose for a terminal agent

All network models have the same price, so your choice is about behavior, not budget. The terminal agent reads a lot (command output, logs, files) and calls tools frequently; all models in the table support native tool calling.

ModelIdentifierContext and OutputWhen to use in Warp
DeepSeek V4 Flashdeepseek-ai/DeepSeek-V4-Flash-0731380K, output up to 32768 tokensDefault agent model: long sessions, large logs and diffs, edits across multiple files. Context headroom and the largest output limit in the network
MiniMax M2.7MiniMaxAI/MiniMax-M2.7200K, output up to 8192 tokensQuick short tasks: finding a command, troubleshooting errors, editing configs
GLM-5.3 Flashzai-org/GLM-5.3-Flash390K, output up to 8192 tokensReasoning model: thinks before answering. Complex debugging, migration planning, analyzing unfamiliar code; the answer takes longer, and part of the output limit is used for reasoning

Practical scheme: DeepSeek V4 Flash is the primary model, MiniMax M2.7 is for small tasks, GLM-5.3 Flash is for when the task is logic-intensive rather than volume-intensive. Switching takes a second because all models are already configured in the endpoint, and short aliases prevent confusion with long identifiers. More details on the most popular one in the DeepSeek V4 Flash overview.

How much it costs

The terminal agent has higher token consumption than a chat: for every request, Warp adds system instructions, dialog context, and tool descriptions, and one task consists of multiple turns. Therefore, the price per token matters more here than in casual chatting.

Via JoinGonka Gateway, tokens cost $0.0069 per million for input and $0.021 per million for output—the price is the same for all network models and is pulled onto this page from a live source. Price scale as of September 2026:

ScenarioUsageVia Gateway
One-time task: command and output parsingtens of thousands of tokensfractions of a cent
Day of active work with the agent3-7M tokensa few cents
Month of daily work~150M tokensaround one or two dollars

For comparison, three ways to pay for the Warp agent:

MethodWho calculates inferenceWhat limits you
Warp ModelsWarp, via plan creditsMonthly plan credit allowance
BYOK: OpenAI, Anthropic, or Google keyModel vendor, per their pricingVendor's price per million tokens
Custom endpoint + JoinGonka GatewayGateway: per token, from balanceBalance only: usage is visible in the dashboard, no request rate limits

Warp itself does not charge credits for individuals and teams of up to 10 people when using your own endpoint—you only pay for tokens at the gateway. The balance is topped up in your dashboard, where you can also see your remaining balance and daily usage by model.

Request routing and privacy: what you need to know

For most tools, a custom endpoint means “requests go directly from my machine to the provider.” Warp is different: the executing part of the agent runs on Warp servers, and a custom endpoint only changes the final destination and the key. According to the Warp documentation, the request path is as follows:

  1. The application retrieves the endpoint address and key from a secure storage on your device and sends them to Warp servers along with your prompt.
  2. The agent on the Warp side assembles the complete request—system instructions, chat context, tools—and calls your endpoint with that key.
  3. The response is returned to the application as a stream through the same servers.

Practical implications:

  • Key. It is stored only on your device. It passes through Warp servers with every request, but according to Warp, it is not saved there: it is used on the fly and discarded.
  • Prompts and responses. They pass in transit through the Warp backend. Warp states that it does not use this content for training, and storage and analytics are subject to your account's privacy and telemetry settings. It cannot extend the ZDR guarantee, which applies to Warp's own models, to your custom endpoint: what the provider stores is determined by the provider.
  • Gateway side. The JoinGonka Gateway does not store conversations: prompts and responses do not remain on the gateway after the response, and only request and token counters are visible in the dashboard.
  • What remains on Warp infrastructure. Auto options, custom model routers, Codebase Context, Active AI suggestions, and cloud agents—a custom endpoint does not affect these.

If the path through third-party servers is unacceptable for your tasks, use an agent that accesses the endpoint directly from your machine, such as Codex CLI or any tool from the API quickstart guide. In its announcement, Warp itself promises a client-side option: a lightweight agent module in Rust in an open repository that will connect to local models without routing through Warp servers, as well as support for the Agent Client Protocol. And as of the September 2, 2026 release, endpoint descriptions are stored in a configuration file in a format common to both the application and the Warp Agent CLI; keys do not end up in this file and remain in secure storage.

Warp connects to the JoinGonka Gateway via the 'Add custom endpoint' form: OpenAI Chat Completions scheme, address https://gate.joingonka.ai/v1, jg- key, and network model identifiers. This is available for individuals and teams of up to 10 people even on the free tier, and Warp credits are not spent on inference if the model endpoint is selected in the toggle instead of Auto. The main feature of Warp is the route: requests to the endpoint are sent by its servers, so the address must be public, and prompts pass through the Warp backend in transit. For terminal tasks, take DeepSeek V4 Flash; for minor tasks, MiniMax M2.7; for complex logic, GLM-5.3 Flash.

Want to learn more?

Explore other sections or start earning GNK right now.

Get key and free tokens →