Knowledge Base Sections ▾

Technology

DeepSeek V4 Flash: Gonka network model with 380K context

In August 2026, a third active model appeared on the Gonka network — DeepSeek V4 Flash. Governance proposal #94 to add it was submitted by the kaitaku.ai team, known in the community for their ML node optimizations: they ran experiments confirming the model's compatibility with Proof of Compute and network validation, and on August 12 the proposal was approved by vote. By the very next day, the model was serving requests.

For users, this is a notable expansion of what's possible: DeepSeek V4 Flash has a 380,000-token context window — one of the longest on the network (nearly double that of MiniMax M2.7; only GLM-5.3 Flash is slightly longer at 390,000), built-in reasoning, and full tool calling. Let's break down what this model is, who made it, what its characteristics are specifically on the Gonka network, what the benchmarks show — and when it's worth choosing over the network's other two models.

What is DeepSeek V4 Flash and who is behind it

DeepSeek is a Chinese AI lab from Hangzhou that gained worldwide fame after the releases of V3 and R1: it was R1 in early 2025 that showed an open model could catch up to closed-source flagships for a fraction of their budget. In April 2026, DeepSeek released the V4 family consisting of two models: the heavy V4 Pro and the lightweight V4 Flash. Both are open (MIT license), and both are built on the MoE architecture.

V4 Flash has 284 billion parameters, of which about 13 billion are activated per token. This sparsity makes the model fast and cheap to maintain: it requires significantly less VRAM than "dense" giants, and the generation speed is higher.

The version running on the Gonka network is V4-Flash-0731 — a re-post-training build from July 31, 2026. The architecture and size have not changed, but the quality on agentic tasks has jumped: according to nine agentic and coding benchmarks published by DeepSeek, the 0731 build even outperforms their own flagship, V4 Pro, in preview version. An important caveat for fairness: some of these figures are measurements by DeepSeek itself on its own test bench, which had not been publicly released at the time of writing, so there is no independent replication yet. The gap with closed-source frontier models remains: the 0731 build does not outperform Claude Opus 4.8 on any of the nine benchmarks, but it costs orders of magnitude less, which is the whole point: maximum agentic quality for a fraction of the price.

DeepSeek V4 Flash characteristics in the Gonka network

As with the network's other models, it's important to distinguish the model's out-of-the-box characteristics from the working parameters of a specific deployment: those are set by the configuration of the vLLM inference on the network's GPU hosts. Here are the actual values through our Gateway:

  • Context window: 380,000 tokens — one of the longest on the Gonka network (only GLM-5.3 Flash is bigger at 390,000). We verified this empirically: a request with 381,000 input tokens was accepted and processed in tens of seconds. The V4 Flash architecture itself supports a context of up to 1 million tokens; the practical ceiling on the network is set by the host configuration (inference window of 400,000).
  • Maximum output: 32,768 tokens per response — the largest ceiling on the network: for MiniMax M2.7 and GLM-5.3 Flash it's four times smaller at 8,192; in both cases this is the gateway configuration, not the model's limit.
  • Reasoning and tool calling: the model is deployed with reasoning and tool-call parsers — agentic scenarios (Cursor, Claude Code, LangChain) work out of the box; we ran multi-step tool scenarios over OpenAI- and Anthropic-compatible paths.
  • Host VRAM requirement: about 280 GB — the network's lightest model (for comparison: MiniMax M2.7 — 320 GB, GLM-5.3 Flash — 560 GB). The lower the hardware threshold, the easier it is for hosts to deploy the model — a good sign for its availability on the network.

The price of inference on the Gonka network doesn't depend on which model you choose: DeepSeek V4 Flash is available at the same rate as MiniMax M2.7 and GLM-5.3 Flash — through the JoinGonka Gateway that's $0.0069 per million input tokens and $0.021 per million output tokens. For reference: third-party providers on OpenRouter offer the same model from $0.06/$0.12 per million, and DeepSeek itself retired V4 Flash 0731 from its API by September 2026 — requests to it are now served by DeepSeek-V4.1-Flash at $0.30/$1.20 during peak hours. The network's economics are a topic of their own: the price is determined by payment for computational work, not by a vendor's price list.

DeepSeek V4 Flash, MiniMax M2.7, and GLM-5.3 Flash — a comparison

There are three active models in the network — DeepSeek V4 Flash, MiniMax M2.7, and GLM-5.3 Flash (Kimi K2.6 was removed from the network in September 2026, and GLM-5.2, which had long been listed in the registry, never entered active service). A brief comparison:

ParameterDeepSeek V4 FlashMiniMax M2.7GLM-5.3 Flash
Network context380,000200,000390,000
Output per response32,7688,1928,192, including reasoning
ArchitectureMoE 284B (~13B active)MoE + linear attentionMoE 320B (~18B active), hybrid attention
VRAM per node280 GB320 GB560 GB
StrengthsLong context, reasoning, speedEveryday development, stabilityComplex logic, code analysis, "thinking" tasks
Price via Gatewaythe same — $0.0069 input / $0.021 output per 1M

The practical consequence of the identical price: you can choose a model purely based on the task, without thinking about the budget. Need large prompts with fast responses and the longest output — DeepSeek V4 Flash; tasks that require thinking through complex logic (and the longest context in the network) — GLM-5.3 Flash; a solid workhorse for every day — MiniMax M2.7.

V4-Flash-0731 Benchmarks: What the Build Shows

Key published results for the 0731 build (according to DeepSeek and independent aggregators):

  • SWE-bench Verified: 79.0% — solving real GitHub tasks; a level that only closed-source flagships reached a year ago. For comparison: Kimi K2.6, which the network previously serviced, had 71.3%.
  • GPQA Diamond: 88.1% — graduate-level scientific questions; the advanced reasoning mode adds accuracy on multi-step tasks.
  • Terminal Bench 2.1: 82.7 — agent terminal operations; the Flash preview version had 61.8, and V4 Pro Preview had 72.1.
  • DeepSWE: 54.4 — a new, strict benchmark from DeepSeek with near-zero false positive rates; the figure is vendor-provided, and the stand has not yet been published — take with this caveat.

Honest assessment: across nine agent benchmarks, the 0731 build outperforms its own V4 Pro Preview, but not Claude Opus 4.8 — it lags by about 6 points on average. At the same time, the price per token is two orders of magnitude cheaper than Opus, and via the Gonka network — ~9 times cheaper than the same model from third-party providers on OpenRouter. It is this "near-frontier quality at a fraction of the cost" ratio that makes it interesting: detailed independent measurements can be viewed on Artificial Analysis, and model weights and cards — on Hugging Face.

How to Use DeepSeek V4 Flash via JoinGonka Gateway

The model is available through the JoinGonka Gateway via an OpenAI- and Anthropic-compatible API. Model ID: deepseek-ai/DeepSeek-V4-Flash-0731.

The fastest way is a one-command installer that configures your tool (Claude Code, OpenClaw, Cline, opencode, Aider, and others) to use this model right away:

npx @joingonka/setup --model deepseek

Wherever the installer writes the config itself (Claude Code, Codex CLI, OpenClaw, opencode, Kilo Code, Hermes, Pi, and others), DeepSeek V4 Flash is the default without any flag. The flag is needed to switch an already-configured tool, and where the installer prints values to paste (Cursor, Cline, Aider). Other network models are selected the same way: --model minimax and --model glm.

Direct API call (OpenAI format):

curl https://gate.joingonka.ai/v1/chat/completions \
  -H "Authorization: Bearer jg-your-key" \
  -H "Content-Type: application/json" \
  -d '{"model":"deepseek-ai/DeepSeek-V4-Flash-0731","messages":[{"role":"user","content":"Hello!"}]}'

When you sign up you get 3M free tokens — with a 380K context, that's a chance to test the model right away on real large documents and codebases. Step-by-step examples in Python and TypeScript are in the API Quickstart; you can try it without signing up in the free chat by selecting DeepSeek V4 Flash from the model list.

When to Choose DeepSeek V4 Flash: Practical Scenarios

Strong scenarios for this model:

  • Large documents and entire codebases. 380,000 tokens — this is about 280,000 words: an annual report, a legal document package, or a medium-sized repository can fit into a single request without cutting it into pieces or losing connections between parts.
  • Long agent sessions. An autonomous agent that reads files, calls tools, and accumulates history hits the context limit the fastest — having a 380K token reserve means significantly longer sessions without memory resets.
  • RAG with large samplings. The more relevant documents fit into the context, the less dependence there is on search accuracy.
  • Reasoning tasks. The reasoning mode provides a boost in math, science, and multi-step logic — at the price of a standard model.

When is another network model more reasonable: for tasks that require deep thinking about tangled logic, there is GLM-5.3 Flash — the network's reasoning model (a detailed breakdown of models for development is in the article about the best models for coding), and MiniMax M2.7 remains a proven workhorse for daily tasks. Thanks to a unified price, you can experiment freely — switching models is just one line in the request.

DeepSeek V4 Flash 0731 — a DeepSeek MoE model (284B parameters, ~13B active, MIT license), added to the Gonka network by governance proposal #94 from the kaitaku.ai team and active since August 13, 2026. The main distinction is the 380,000 token context (verified by a real 381K request; only GLM-5.3 Flash in the network is longer — 390,000) and the largest response limit, 32,768 tokens, plus reasoning and tool calling. SWE-bench Verified 79.0%, GPQA Diamond 88.1% — near-frontier at a price two orders of magnitude lower than closed-source flagships; via JoinGonka Gateway — at the same rate as MiniMax M2.7 and GLM-5.3 Flash, ~9 times cheaper than the same model on OpenRouter. ID: deepseek-ai/DeepSeek-V4-Flash-0731; quick start — npx @joingonka/setup --model deepseek.

Want to learn more?

Explore other sections or start earning GNK right now.

Try DeepSeek V4 Flash via Gateway →