Knowledge Base Sections ▾
Navigation
▸ Start here By rolesCategories
- Cursor + Gonka AI - cheap LLM for coding
- Claude Code + Gonka AI - LLM for the terminal
- OpenClaw + Gonka AI - affordable AI agents
- OpenCode: your own model in the terminal
- Continue.dev + Gonka AI - AI for VS Code/JetBrains
- Cline + Gonka AI - AI agent in VS Code
- Aider + Gonka AI - pair programming with AI
- LangChain + Gonka AI - AI applications for pennies
- n8n + Gonka AI - automation with cheap AI
- Open WebUI + Gonka AI - your own ChatGPT
- LibreChat + Gonka AI — open-source ChatGPT
- Hermes Agent + DeepSeek on the Gonka network — an autonomous agent for pennies
- Kilo Code + Gonka AI — AI-Agent in VS Code
- Roo Code + Gonka AI — Autonomous AI Agent in VS Code
- LlamaIndex + Gonka AI — RAG applications for pennies
- PydanticAI + Gonka — typed AI agents for pennies
- Vercel AI SDK + Gonka AI — AI applications in TypeScript for pennies
- TanStack AI + Gonka — AI applications in TypeScript for pennies
- API quick start — curl, Python, TypeScript
- JoinGonka Gateway — a full overview
- Management Keys — SaaS on Gonka
- Cheapest AI API: Provider Comparison 2026
- How to buy AI tokens and an API key: 3 methods in 2026
- Cursor Pro request limit reached — breakdown and cheaper alternative
- Claude Code is cheaper — bill breakdown and switching
- Cline is burning money — why the agent spends so much
- OpenClaw is expensive — why the agent burns through tokens and how to save
- OpenRouter: Cheap Alternative — Comparison with JoinGonka Gateway
- Best AI model for coding in 2026: comparison and prices
- Cheap alternative to GitHub Copilot without limits
- A cheap Windsurf alternative without credits or limits
- The cheapest API for AI agents in 2026
- ZCode: Cheap GLM inference instead of GLM Coding Plan
- JetBrains IDE + JoinGonka Gateway — your own endpoint instead of credits
- GitHub Copilot BYOK — own models instead of quotas
- Zed + JoinGonka Gateway — cheap inference in your editor
- Pi + JoinGonka Gateway — terminal agent on cheap inference
- Codex CLI: your own key instead of a subscription
- DeepSeek Harness: Your Own Provider via JoinGonka Gateway
- MiniMax Code: MiniMax agent with your own key via Gonka
- Warp + JoinGonka Gateway — terminal agent on your own endpoint
- Trae + JoinGonka Gateway — Gonka network models in AI-IDE
- Cherry Studio + JoinGonka Gateway — desktop AI client
- omp (Oh My Pi) + JoinGonka Gateway: an agent with model roles
- OpenHands + JoinGonka Gateway: agent on your own endpoint
- Qwen Code after the closure of qwen-oauth: working via JoinGonka Gateway
- Goose + JoinGonka Gateway: your own provider and key in the keyring
- Crush + JoinGonka Gateway: Charm agent on Gonka network models
- Zoo Code + JoinGonka Gateway: Migrating from Roo Code to Gonka models
- Kimi Code CLI: Moonshot AI agent on your key via Gonka
- Factory Droid + JoinGonka Gateway: BYOK on Gonka network models
- MiMo Code + JoinGonka Gateway: Xiaomi agent on Gonka network models
Tools
DeepSeek Harness: Your Own Provider via JoinGonka Gateway
DeepSeek Harness (dsh command) is an open-source agentic harness from DeepSeek AI: a shell where the model reads and edits project files, runs commands, delegates subtasks, and maintains a plan, while you monitor this from your browser and approve risky steps. The project is new: the authors call it a developer preview and warn that breaking changes will occur. Therefore, everything below is tied to a specific version — 0.1.5-rc.2 — on which we completed the setup from the first screen to the agent's response.
On first launch, dsh asks for an official API key from its vendor, but the model layer is open: on the Settings → Models page, you can add any provider that speaks one of three protocols — OpenAI Chat Completions, OpenAI Responses, or Anthropic Messages. JoinGonka Gateway supports all three, so the harness connects to the decentralized Gonka network using standard tools, without plugins or patches.
A telling detail from the app's page in the OpenRouter catalog: among the models running DeepSeek Harness there, DeepSeek V4 Flash 0731 holds second place over the last 30 days, and GLM 5.3 Flash is third (snapshot as of September 21, 2026; an anonymous test model is in first place). Both open models are serviced by the Gonka network — along with MiniMax M2.7 — so your usual setup migrates to a different endpoint without changing the model: only the address and price per token change.
What is DeepSeek Harness and how to launch it
The harness is everything that surrounds the model in agentic work: the loop of “request → tool call → result → next step,” the file and terminal tools, permissions and confirmations, the session log, context compression. DeepSeek Harness assembles all of this from plugins: the “everything is a plugin” architecture is built on the Cordis framework, and any node — from a tool to the model adapter — can be swapped without touching the core. The code is open under the MIT license.
No installation needed — just Node.js (the 22 line from 22.19 onward, or 24 and newer):
npx @deepseek-ai/dsh webThe command brings up the Web UI at http://127.0.0.1:3080 and opens it in your browser; when launched over SSH, the address is only printed to the terminal. The --no-open flag starts the server without a browser, and --port changes the port. The directory you launch dsh from becomes the default working directory, but the interface won't start a session until you explicitly pick a workspace.
| Mode | Command | What it's for |
|---|---|---|
| Web UI | dsh web | The main interface: sessions, settings, operation confirmations |
| One-off task | dsh --profile headless "task" | Scripts and CI: the answer goes to stdout, the reasoning trace to stderr |
| ACP | dsh --profile acp | Editors and clients that support the Agent Client Protocol |
| SDK | dsh --profile sdk | JSON-RPC clients, including the Python SDK |
The model layer consists of two adapters. The direct one talks to the vendor's official API. The multi-provider one — dsh-llm-pi-ai — is built on the pi-ai library, the same one that underpins the terminal agent Pi; through it you can connect both the built-in catalog providers and any endpoint of your own. That's why the field names in the config — api, contextWindow, maxTokens — match the ones you already know from Pi.
On maturity. The project README opens with a warning: developer preview, rapid iteration, breaking changes. A separate SAFETY.md document notes that no security audit has been performed and that the agent executes commands generated by the model. The practical takeaway is simple: run dsh in a container, a virtual machine, or under a separate account, and keep backups of everything it can reach.
Connecting via Web UI: Settings → Models
Step 1: the key. Sign up at gate.joingonka.ai/register: once you confirm your address, 3M free tokens land in your account. In your dashboard, open the "API keys" section and create a key with the jg- prefix. It's handy to set up a separate key for the harness — that way its traffic shows up as its own line in your stats.
Step 2: the first screen. After the test-status notice (the Continue button), dsh will ask you to enter an official API key ("Add an API key to get started"). It's optional: click Configure later.
Step 3: the provider. Open Settings → Models and choose Add a custom provider. The form fields:
| Field | Value | Note |
|---|---|---|
| Provider ID | joingonka | Lowercase Latin letters, starts with a letter. The identifier is permanent: it goes into requests, saved sessions and the name of the key reference. You can't rename it — only create a new provider and delete the old one |
| Display name | JoinGonka Gateway | Any label for lists |
| Base URL | https://gate.joingonka.ai/v1 | With the /v1 suffix |
| API protocol | openai-completions | How to choose a protocol — see the table below |
| API key | jg-your-key | Write-only field: after saving, the page gets a masked descriptor instead of the key itself |
Step 4: the models. In the Models block, click Fetch available models: dsh will request the list from the gateway and open the "Choose models to add" window. In our run it showed all three models on the network — MiniMaxAI/MiniMax-M2.7, deepseek-ai/DeepSeek-V4-Flash-0731 and zai-org/GLM-5.3-Flash — and after Add selected the harness filled in the context window and response cap for each one automatically, based on the gateway's data. All that's left is to click Create provider.
Step 5: choosing a model. Close the settings, click Choose workspace and add your project directory. The new provider's models will appear in the selector; the one you pick becomes the default model for new sessions.
dsh stores the key separately from the settings: in the file ~/.dsh/.credentials.yaml, owner-readable only. settings.yaml keeps just the name of the reference to it — in our run JOINGONKA_API_KEY, after the provider ID.
Which protocol to choose. The gateway speaks all three; what differs is the base address and the extra conveniences:
| API protocol | Base URL | When to choose it |
|---|---|---|
openai-completions | https://gate.joingonka.ai/v1 | The main option: the gateway's canonical path, the model list is pulled in with a button, and a reasoning model's chain of thought arrives as a separate stream |
openai-responses | https://gate.joingonka.ai/v1 | If your plugins or scenarios are built around the Responses API |
anthropic-messages | https://gate.joingonka.ai | Anthropic Messages format; the client appends the /v1/messages path itself |
A single provider in dsh speaks a single protocol, so a second protocol means a second provider with a different Provider ID. For everyday work the first option is enough; in our run the agent loop with tool calls worked on all three.
Configuration via file: settings.yaml
The Models form writes to a regular YAML document — $DSH_HOME/settings.yaml, by default ~/.dsh/settings.yaml. You can edit it directly: the Open configuration file button at the top of the settings opens the file, and adapters re-read it on the next request — no restart needed. Here is the full version for the Gonka network:
# ~/.dsh/settings.yaml
llm-pi-ai:
providers:
joingonka:
displayName: JoinGonka Gateway
apiKeyEnv: JOINGONKA_API_KEY
api: openai-completions
baseURL: https://gate.joingonka.ai/v1
models:
- id: deepseek-ai/DeepSeek-V4-Flash-0731
name: DeepSeek V4 Flash
contextWindow: 380000
maxTokens: 32768
- id: zai-org/GLM-5.3-Flash
name: GLM-5.3 Flash
contextWindow: 390000
maxTokens: 8192
reasoningEfforts:
off: low
high: high
- id: MiniMaxAI/MiniMax-M2.7
name: MiniMax M2.7
contextWindow: 200000
maxTokens: 8192
agent-default-model:
provider: joingonka
model: deepseek-ai/DeepSeek-V4-Flash-0731Here is what matters:
apiKeyEnvis not the key itself but the name of the reference to it. dsh looks up the value in order: the environment variable at launch time, then.credentials.yaml(where the form writes it), then.envin the launch directory, then~/.dsh/.env. If you are setting up the harness without a browser, a single lineJOINGONKA_API_KEY=jg-your-keyin~/.dsh/.envwith600permissions is enough. A variable exported after startup will not be seen by an already running process.- Set
contextWindowandmaxTokensexplicitly. For a model dsh knows nothing about, it assumes 262,144 and 32,768 tokens — which does not match the real limits. ThemaxTokensyou set also becomes the default response limit for every request. reasoningEffortsare the reasoning levels for the Effort menu. A manually added model has no levels, and the menu does not appear for it. GLM-5.3 Flash has a binary switch: the valuelowturns reasoning off, any other value leaves it at full. That is why theofflevel is mapped tolow, whilehighis passed through as is. In our run withoffthere were no reasoning blocks at all, and withhighthey came back.agent-default-modelis the model for new agents, includingheadlessmode. Selecting a model in the interface does the same thing; you can also addreasoningEfforthere.
The compat switches that the dsh documentation recommends for strict gateways (supportsDeveloperRole: false, maxTokensField: max_tokens) are not needed here: JoinGonka Gateway accepts both the developer role and the max_completion_tokens field.
The installer npx @joingonka/setup does not configure this harness: the whole setup comes down to the form from the previous section or to the YAML snippet above.
Verification and Common Errors
The fastest way to verify the integration is a one-off run from the directory containing your code. Put a small file with an obvious bug next to it and ask the agent to find it:
cd /path/to/project
npx @deepseek-ai/dsh --profile headless "Read calc.py and tell me in one sentence whether it has a bug."The final answer is printed to stdout, while the reasoning trace goes to stderr prefixed with dsh: reasoning:. The agent must call the file-reading tool itself and give a substantive answer: in our run, each of the three network models named the faulty line. That means the full cycle of "request → tool call → result → answer" through the gateway is wired up correctly.
The second half of the check is on the gateway side. In the dashboard, open "Usage": there you can see requests by hour and by day, broken down by model and by key. Once a row appears with the harness key and a fresh last-request timestamp, traffic is genuinely flowing through the gateway.
If something goes wrong, the diagnosis is usually readable straight from the message:
| What you see | What it means | What to do |
|---|---|---|
AUTH: 401: … Invalid API key | The gateway rejected the key | Re-enter the key on the Models page or fix the variable referenced by apiKeyEnv |
MISSING_CREDENTIAL: … no credential for provider route "joingonka" | Nothing was found at the reference from apiKeyEnv | Save the key in the form or set the variable before launching dsh: the environment is read once, at startup |
UNKNOWN_MODEL | The model isn't in the provider's models list | Add it to the form or the file, or pick one that's already configured |
400 … Model "…" not found. Available: … | The identifier was written inaccurately, most often without the vendor prefix | Copy the id from the list the gateway includes in the message itself |
429 … currently overloaded … (rate limit) | The model has temporarily run out of free capacity on the network | A normal situation under load: dsh retries the request itself. If retries are exhausted, switch models or wait a minute; the state is visible on the status page |
| Fetch available models returns 401 | The list was requested with a wrong key | Check the key in the form; models can also be entered manually — they'll work the same way |
| A reasoning model has no Effort menu | No levels are declared for the model entry | Add reasoningEfforts to settings.yaml, as in the example above |
| A reasoning model's reply is cut off or empty | The reasoning counts toward the response limit and consumed it entirely | Don't lower maxTokens; for short tasks choose the off level |
| The input field shows Select model and input is blocked | The default model points to a deleted provider | Choose another model in the selector |
Which model to choose
The price for all models in the network is the same, so the choice is about behavior, not budget. Below are the limits and how the models performed in our dsh run on the same task: reading a file and finding an error in it.
| Model | Identifier | Context / Response | Behavior in dsh |
|---|---|---|---|
| DeepSeek V4 Flash | deepseek-ai/DeepSeek-V4-Flash-0731 | 380K / 32768 | Clean response indicating the line number. The largest response limit in the network—suitable for long edits and large files in a single pass. |
| GLM-5.3 Flash | zai-org/GLM-5.3-Flash | 390K / 8192 | Reasoning model: dsh displays reasoning in a separate stream, the response remains clean. Reasoning is included in the response limit. |
| MiniMax M2.7 | MiniMaxAI/MiniMax-M2.7 | 200K / 8192 | Solves the task correctly; reasoning comes in a separate reasoning_content field, the response text contains only the answer itself. |
The default recommendation is DeepSeek V4 Flash: agentic work quickly hits context volume and edit length limits, and this model has plenty of room for both. When a task requires deep thinking over complex logic, switch to GLM-5.3 Flash and keep the reasoning level at high; for quick edits, the same provider delivers it with the level set to off. MiniMax M2.7 is a solid option for short tasks where visible reasoning streams are distracting. The model can be changed in the interface selector or via the model string in the agent-default-model block.
The network composition is determined by participant voting and changes over time; an up-to-date list with limits is always provided by GET https://gate.joingonka.ai/v1/models — the same endpoint used by the Fetch available models button.
How much it costs and what to consider in your work
Agentic tools consume tokens differently than a chat: for every phrase of yours, the harness adds a system prompt and descriptions of all tools, then conducts a multi-turn dialogue with the model. In our test run, the task "read the file and find the error" took two to three turns and between 14 and 22 thousand tokens, with almost all of it being input: about seven thousand tokens are spent with each turn even before your question. This is a normal price to pay for autonomy—and that is exactly why the price per token matters.
Through JoinGonka Gateway, tokens cost $0.0069 per million for input and $0.021 per million for output—the price is the same for all network models and is pulled onto this page from a live source. Orders of magnitude for prices as of September 2026:
| Scenario | Consumption | Via Gateway |
|---|---|---|
| One-off task (read file, find error) | 14-22K tokens | fractions of a cent |
| Day of active work | 3-7M tokens | a few cents |
| Month of active development | ~150M tokens | around a dollar |
Payment is made for actual consumption, with no subscription and no request quotas; the balance and daily consumption are visible in the dashboard.
Version. While the project is in developer preview status, check after each update that the provider is in place, and for reproducibility, pin the version directly in the command: npx @deepseek-ai/[email protected] web.
Permissions. New sessions run in Workspace Write mode by default—writing is limited to the working directory; the interface asks to confirm operations exceeding the policy. The mode can be changed in Settings → General.
Retries. In case of a one-time network error, dsh retries the request itself—up to five times according to the documentation—so a short burst of network load usually goes unnoticed.
Privacy. The gateway does not store the content of prompts and responses: only usage aggregates remain in the statistics. The agent reads project files locally on your machine.
If you need to work with images—an interface screenshot, a diagram in a photo—set up a second provider nearby with a vision-capable model: dsh supports several providers simultaneously, and Gonka network models are text-based.
DeepSeek Harness is not the only agent released by the model developer lab itself: Z.ai, the authors of GLM, have the ZCode environment, and MiniMax has the terminal-based MiniMax Code. Both connect to the same gateway with the same key.
https://gate.joingonka.ai/v1, protocol openai-completions, key jg-…; the Fetch available models button automatically pulls DeepSeek V4 Flash, GLM-5.3 Flash, and MiniMax M2.7 along with their limits. The same is written as a single llm-pi-ai block in ~/.dsh/settings.yaml. For GLM-5.3 Flash, declare the off: low and high: high levels—reasoning will become toggleable. Run the harness in an isolated environment and pin the version until the format stabilizes.Want to learn more?
Explore other sections or start earning GNK right now.
Get key and free tokens →