What has to stay close together
Mixture of experts did not only raise parameter counts. It changed which tensors are active per token, and therefore which links in a cluster have to be fast.
3 minDePIN Compute & Bandwidth
SemiAnalysis has published an analysis of how mixture-of-experts models map onto inference hardware, and its framing is useful for anyone reading capacity claims about decentralised compute. The argument is that MoE changed the shape of serving, not just the size of models.
What changed specifically: which tensors are active for any given token, what therefore has to remain physically close together, which transfers need strong local bandwidth, and which can tolerate a weaker network link. Those distinctions decide how much of a cluster's nominal throughput is actually useful.
The piece starts from the service rather than the chip. Inference runs inside a cluster coordinated by an orchestration layer — NVIDIA Dynamo, Mooncake, or a custom scheduler — working alongside inference servers such as vLLM or SGLang, which carry orchestration features of their own.
The unit of work is the turn: a request and its answer. In modern systems a large share of turns are not typed by a person. A client application intercepts parts of the answer as instructions — edit this code, search these guidelines — executes them, and returns the results as further requests. The server keeps a context, the KV cache, which is the distillation of the session and lets each new request be interpreted against what came before.
One consequence is stated plainly and is worth carrying: when a user launches an agent on a long-running task there can be thousands of turns an hour, and the conversation can pause and resume hours or days later. That is a very different load profile from interactive chat, and it puts weight on cache retention and scheduling rather than on raw arithmetic.
For this desk the relevance is direct. A decentralised network advertising aggregate throughput is describing arithmetic. Serving MoE usefully depends on what sits close to what, and that is the part a geographically scattered fleet cannot advertise.
Retold from SemiAnalysis. This is a summary in our own words; follow the link for the original reporting.