Knowledge Base Sections ▾
Navigation
▸ Start here By rolesCategories
- Gonka Network Architecture: Sprint, Transfer Agents, DiLoCo
- Developers: How to Earn GNK
- Self-hosting: Step-by-step guide
- Choosing a GPU for Gonka: Hardware Recommendations
- Qwen3-235B: the model previously served by Gonka
- Kimi K2.6: a model previously hosted by Gonka
- MiniMax M2.7: Gonka Network Model
- DeepSeek V4 Flash: Gonka network model with 380K context
- GLM-5.3 Flash: Z.ai reasoning model on the Gonka network
Technologies
Choosing a GPU for Gonka: Hardware Recommendations
Minimum Requirements
Gonka requires NVIDIA GPUs with CUDA support and at least 320 GB of VRAM in total per MLNode container; cards can be combined as long as the node fits the network's models and passes PoC. This is a hard hardware constraint: the network's models — MiniMax M2.7, DeepSeek V4 Flash, and GLM-5.3 Flash — with their MoE architecture require a significant amount of video memory to load the weights and perform inference. AMD and Intel GPUs are not supported — Gonka uses NVIDIA's CUDA stack, including cuBLAS for matrix operations and cuDNN for neural network layers.
The full list of requirements:
- GPU: NVIDIA with CUDA, newer than the Tesla generation, at least 320 GB of VRAM in total per ML node. The official list includes B300, B200, RTX PRO 6000, H200, H100, and A100 80GB. Gaming RTX 4090 and RTX 3090 (24GB each) don't qualify: that volume is assembled from server cards
- CPU: support for AVX instructions is mandatory — without them inferenced won't start (SIGILL at launch)
- RAM: 64GB+ of RAM is recommended for comfortable operation, loading weights, and servicing request queues
- Disk: an NVMe SSD with enough space — the full set of weights for a large MoE model takes up hundreds of gigabytes, and fast loading from NVMe is critical for the node's cold start time
- Internet: at least 100 Mbps of stable connectivity — the node receives requests from Transfer Agents and sends results to clients in real time
- Uptime: 24/7 — missing epochs reduces the reward, and prolonged downtime can lead to exclusion from the task pool
Recommended Cards
Let's break down each recommended card in detail:
NVIDIA H100 80GB — the flagship of the current generation and the optimal choice for Gonka. TDP (thermal design power) 700 W, cost ~$25–35K per card. Supports FP8 inference, which accelerates request processing without sacrificing quality. NVLink lets you combine several H100s into a single node with high-speed interconnect between cards. The network's heaviest model — GLM-5.3 Flash with its ~560 GB of VRAM per replica — needs eight H100 80GB cards; that's the reference host configuration for this model.
NVIDIA H200 141GB — the next generation with nearly double the memory capacity. The larger VRAM lets you process more requests simultaneously (a bigger batch size), which increases GNK earnings per unit of time. HBM3e memory bandwidth is higher than the H100's — faster weight loading, faster inference. The same model requires fewer H200 cards than H100s, which simplifies your infrastructure.
NVIDIA A100 80GB — the previous generation, but still on the network's official list. Price ~$10–15K per card — significantly cheaper than the H100. Lower performance: no FP8, slower HBM2e. One ML node needs at least four of these cards (4 × 80 = 320 GB); the A100 40GB version is not on the official list. The A100 is a good way to enter the network with less upfront investment.
What doesn't qualify: consumer cards like the RTX 4090 (24GB), RTX 3090 (24GB), RTX 4080 (16GB) — not enough VRAM to run the network's models. The network's threshold is 320 GB of video memory per ML node, and it's built from server-grade cards. A single instance of a large MoE model means several H100, H200, or A100 80GB cards in one node; the exact number depends on the model and configuration.
Node Configuration
MLNode in the Gonka network is a server with a GPU that performs AI inference. Setting up a node involves several stages, each of which is critical for stable operation and maximum earnings of GNK.
Software: the main component is the inferenced CLI, which handles model loading, request processing, and communication with the blockchain. Inferenced runs inside a Docker container, which simplifies deployment and updates. A full configuration for a large MoE model of the network requires hundreds of gigabytes of total VRAM — for example, eight H100 cards at 80GB each for GLM-5.3 Flash (around 560 GB per replica). Model weights (hundreds of gigabytes) are loaded from an NVMe SSD when the node starts.
Registration: after installation, the node registers on-chain — it creates a record in the Gonka blockchain specifying its address, supported models, and characteristics (VRAM, bandwidth, location). From that moment, Transfer Agents begin routing AI requests from users to the node.
Working in the network: each incoming request — a prompt from a user — is processed by the GPU through the neural network. The result is sent back through the Transfer Agent to the client. Sprint consensus accounts for every computation performed when forming a block, and the reward is distributed proportionally to the amount of work. A node can publish updated characteristics in real time — if load increases, Transfer Agents will redirect some requests to less loaded nodes. A detailed setup guide is available in the mining guide.
Where to Rent GPUs
If you don't have your own equipment, there are three ways to access a GPU for Gonka, each with a different balance of cost, complexity, and control:
| Path | Cost | Complexity | Control |
|---|---|---|---|
| Pool | from $1 | Minimal | Low |
| Dedicated server | from $12,000/month | Low | Medium |
| Bare-metal rental | from $2–3/hour GPU | High | Full |
Pools (from $1 – Ancapex, from $100 – Gonka.Top): Ancapex, Gonka.Top, GonkaPool.ai, CloudMine (Mingles) – operators rent GPUs, set up nodes, and monitor uptime. You receive GNK proportionally to your contribution, without touching technical details. The ideal path for newcomers and passive investors.
Dedicated servers (from $12,000/month): Gonka.Top offers not only pools but also fully managed dedicated servers. You get a ready-made node – the operator handles inferenced setup, 24/7 monitoring, updates, and troubleshooting. Mining goes directly to your wallet – all GNK income is yours, minus a fixed rental fee.
Bare-metal rental: Spheron provides bare-metal servers with H100/H200 that you configure yourself (users from Russia may experience payment difficulties when using Spheron through a payment processor). This is a path for technical users familiar with Linux, Docker, and CLI. Maximum control, but also maximum responsibility for setup, uptime, and updates. A detailed comparison of all providers is on the “Get GNK” page.
Want to learn more?
Explore other sections or start earning GNK right now.
Compare Providers and Rent →