Knowledge Base Sections ▾

Technology

Kimi K2.6: a model previously hosted by Gonka

For a long time, the Gonka network operated on a single model — Qwen3-235B from Alibaba Cloud. In May 2026, this changed: support for multiple models was launched via the DevShards mechanism, and the first to arrive was Kimi K2.6 from the Chinese company Moonshot AI. Later, MiniMax M2.7 and DeepSeek V4 Flash were added, and Qwen3-235B was removed from the network. In September 2026, Kimi K2.6's turn came: hosts stopped supporting it, and governance proposal #101 finalized its departure — it was replaced in the network by GLM-5.3 Flash. Today, Gonka supports three models: MiniMax M2.7, DeepSeek V4 Flash, and GLM-5.3 Flash. Let's analyze what Kimi K2.6 represented, how it differed from MiniMax M2.7, how Gonka technically implemented multi-modality, why the model left the network, and what to choose instead now.

What is Kimi K2.6 from Moonshot AI

Kimi K2.6 is a large language model (LLM) from the Kimi series, developed by the Beijing-based company Moonshot AI. Moonshot AI is one of China's leading AI labs, founded in 2023 by a team of researchers led by Yang Zhilin. The company has raised funding from Alibaba, Tencent and other major investors, and has earned a place among the "Chinese AI tigers" — the companies setting the pace of AI development in Asia.

The Kimi series has been around since 2024. The early versions (K1, K1.5) immediately drew attention for their exceptionally long context window — up to 200,000 tokens in a single request, which at the time of release was a record for publicly available models. A long context means you can realistically analyze an entire book, a mid-sized codebase or a stack of legal documents in one request. When Kimi launched, that was a serious competitive advantage.

Version K2 arrived in 2025 and brought a fundamental architectural leap — the move to MoE (Mixture of Experts). The same architecture underpins Qwen3-235B and DeepSeek-R1 — it has become the de facto standard for the largest models of 2025—2026. MoE lets a model have hundreds of billions of parameters in total while activating only a subset (usually 5—10%) on each request, which drastically cuts the computational cost of inference at comparable quality.

K2.6 is the latest iteration of the K2 series as of the time of writing. From Moonshot AI's public statements, this version improves the model's reasoning, code generation and native tool calling. While the model was served by the Gonka network, it was available under the identifier moonshotai/Kimi-K2.6; the gateway no longer accepts that identifier — the current model list is always returned by GET /v1/models.

Comparison of Kimi K2.6 and MiniMax M2.7

Both models represent flagship developments from major Chinese AI labs; while both were serviced by the network, they were available through the unified OpenAI-compatible interface JoinGonka Gateway—today, only MiniMax M2.7 is available through the gateway from this pair. They have different strengths and different heritages, making the choice between them not a question of "which is better," but "which fits the task."

CharacteristicKimi K2.6MiniMax M2.7
ManufacturerMoonshot AI (Beijing)MiniMax (Shanghai)
Founded20232021
ArchitectureMoEMoE + linear attention
Context Window200,000 tokens200,000 tokens
Key StrengthReasoning, long context, code generationLong context, efficient (linear) attention
Price via JoinGonka— (model retired from network)$0.0069 per 1M tokens
API Identifiermoonshotai/Kimi-K2.6MiniMaxAI/MiniMax-M2.7
Status in GonkaServed from May to September 2026, retired (proposal #101)Active model (since May 2026, upgrade v0.2.13)

On reasoning benchmarks (MATH-500, GSM8K, AIME), the Kimi K2 series historically shows results in the top tier of open-weights models, competing with DeepSeek-R1 and o1-style models. On code generation tasks (HumanEval, MBPP), both models perform at similar levels. The strong suit of MiniMax M2.7 is its efficient (linear) attention for very long sequences, whereas Kimi is known for strong reasoning and the long context of the Kimi series.

An important caveat regarding benchmarks in 2026: the gap between top models in public tests has shrunk to a few percent, and this difference often falls within the statistical margin of error of the benchmarks themselves. For practical work, it is not about "who is 2% higher in MMLU," but the nature of the tasks: what context you pass to the model, how complex the logic chains are, whether you need a long dialogue history, and which languages are used. Therefore, the table above does not rank the models—it helps to quickly understand which profile of tasks each is optimized for.

For a practical choice today: the niche of Kimi K2.6—long context (analyzing large documents, reading voluminous codebases, long dialogues with history preservation) and complex reasoning tasks—is covered by two models in the current network lineup. GLM-5.3 Flash is responsible for reasoning; it thinks before answering, and it is the one to use for convoluted logic; it also has the network's longest context (390K). For large prompts without reasoning and long agent sessions, DeepSeek V4 Flash (380K, the second-longest context) is the choice. If processing very long input sequences and streaming data with fast response times is the priority, MiniMax M2.7 with its efficient attention is the move. The good strategy in production remains unchanged: keep several network models in your code; fast swapping via the model parameter allows switching between them depending on the task without changing the application architecture.

DevShards: How Gonka Launched the Second Model

Until spring 2026, the entire Gonka network served exactly one model — Qwen3-235B. From an architectural standpoint, that was a sensible decision: distributed inference via DiLoCo requires every network participant to keep the same model in VRAM, otherwise there's no guarantee that any node can handle any request. Full Qwen3-235B in FP8 format takes up about 640 GB of VRAM, which is already a huge commitment for each ML node.

To move to a multi-model network, a mechanism was needed that would allow several models to coexist without requiring every host to run all of them. That mechanism became DevShards — separate network shards, each specializing in one model. Nodes within a shard work on the same model, and the network router directs requests to the shard with the needed model.

The idea didn't come out of thin air — it was formalized in Gonka Improvement Proposal #800 "Multi-Model PoC," put to a community vote in spring 2026. The proposal gained support from network participants and validators and was implemented in April–May 2026. Kimi K2.6 became the first model launched on a separate DevShard — effectively a test implementation of the new approach. The experiment proved successful: MiniMax M2.7, DeepSeek V4 Flash, and GLM-5.3 Flash soon appeared on their own shards, each with its own set of hosts and its own economics. The same mechanism works in reverse: a model that hosts stop supporting is removed from the network by vote — that's exactly how Kimi K2.6 itself was retired in September 2026.

What this means for users and developers:

  • One API — multiple models. Through JoinGonka Gateway, there's no need to change the endpoint or keys: just specify a different model in the request body. The OpenAI-compatible format is fully preserved.
  • The price is the same. While Kimi K2.6 was being served, it was billed at the same rate as MiniMax M2.7; the network's unified pricing also applies to current models — $0.0069 per 1M tokens via the Gateway. Unified pricing is a deliberate decision to simplify user migration between models.
  • Stability depends on shard load. In the early stage, a new model's shard has fewer hosts, so when requests concentrate, the model may temporarily return 429 too many concurrent requests. This is a normal phase for a new model: as interest grows, hosts join its shard and limits increase. The reverse is also true — if hosts leave a shard, the model loses capacity; that's exactly how Kimi K2.6's story on the network ended.
  • Tool calling is unique to each model. On Gonka, Kimi K2.6 initially had minor issues with automatic tool selection (tool_choice: "auto"), later resolved by a node update. The lesson remains relevant for any model on the network: the tool-calling format is a property of a specific model on a specific node, so for production-critical scenarios, test your chosen model's behavior on your requests in advance.

What it was replaced with and what to choose now

The direct answer: Kimi K2.6 is no longer available on the Gonka network. Hosts stopped serving it in early September 2026, and governance proposal #101 removed the model from the network's model list for good — a request with model: "moonshotai/Kimi-K2.6" will be rejected by the gateway. Since the model's weights are open, you can still get it from third-party hosts of open-weights models (for example, via OpenRouter) or deploy it yourself.

But if what you're after is that same cheap decentralized inference that people come to Gonka for — it hasn't gone anywhere, it just runs on the network's active models. Through the JoinGonka API Gateway, three models are available via an OpenAI- and Anthropic-compatible API, and each one covers its own part of what people valued about Kimi:

  • GLM-5.3 Flash (zai-org/GLM-5.3-Flash) — Z.ai's reasoning model: it thinks before answering, so it's the one to reach for with complex logic, code analysis and "think it through" tasks; it also has the network's longest context (390K). Reasoning counts toward the response limit — set max_tokens with headroom and enable streaming.
  • DeepSeek V4 Flash (deepseek-ai/DeepSeek-V4-Flash-0731) — one of the network's longest contexts (380K), the longest output and strong agentic coding: large repositories, long chains of tool calls.
  • MiniMax M2.7 (MiniMaxAI/MiniMax-M2.7) — the gateway's default model: fast, steady answers on everyday tasks, long documents and streaming workloads.

Migrating off Kimi K2.6 is a one-line swap. Any code written for OpenAI works unchanged: just replace the URL, the API key and the model name.

curl https://gate.joingonka.ai/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MiniMaxAI/MiniMax-M2.7",
    "messages": [{"role": "user", "content": "Explain the difference between MoE and dense models"}]
  }'

The same rule applies in development tools: wherever the settings had moonshotai/Kimi-K2.6, drop in the identifier of one of the active models. In Cursor that's the Custom Model field, in Claude Code it's the ANTHROPIC_MODEL environment variable or the --model flag, in OpenClaw, Cline and Continue.dev it's the model name in the provider config, and in LangChain and n8n it's the model parameter when initializing the client. The npx @joingonka/setup installer will write the gateway and the current model into your tool's config with a single command. The list of available models is always up to date at the GET /v1/models endpoint — it's handy to pull it dynamically into your app's UI so that a model leaving or joining the network never breaks your product.

You can try it without registering in the free chat on the /try page — the network's active models are available there. When you register with JoinGonka Gateway, you get 3M free tokens to test any of the network's models — enough to run your tasks on all three and choose a replacement with full confidence.

What the Kimi K2.6 story showed for the Gonka network: the DevShards mechanism worked in both directions. It made it possible to add models without stopping the network — MiniMax M2.7, DeepSeek V4 Flash and GLM-5.3 Flash followed Kimi — and it also made it possible to painlessly retire a model that hosts had stopped supporting. A network tied to a single model is fundamentally fragile; a network that can change its lineup by vote evolves smoothly and continuously. For a developer, the takeaway is a simple rule: don't "pick a model forever," keep the model name in configuration and check the live list.

Kimi K2.6 is a MoE model by Moonshot AI with long context and strong reasoning capabilities. In May 2026, it became the second model on the Gonka network after Qwen3-235B, launched via the DevShards mechanism (a separate shard per model), and was served by the network until September 2026, when hosts stopped supporting it and governance proposal #101 removed it from the lineup. Currently, MiniMax M2.7, DeepSeek V4 Flash, and GLM-5.3 Flash are available via the JoinGonka Gateway using an OpenAI-compatible API — at a flat network rate of $0.0069 per 1M tokens; Kimi's niche (reasoning and long context) is covered by GLM-5.3 Flash and DeepSeek V4 Flash. Kimi K2.6 itself remains an open model and is available via third-party open-weights model hosting services.

Want to learn more?

Explore other sections or start earning GNK right now.

Try current Gonka models →