Mixture of Experts changes inference architecture: data movement and cluster orchestration
An analysis by SemiAnalysis describes how the Mixture of Experts architecture changes the requirements for inference infrastructure – which tensors remain active, how data moves between the prefill, midfill and decode phases, and how orchestration layers such as NVIDIA Dynamo or Mooncake coordinate this.
According to the source, the Mixture of Experts architecture, now commonly used in leading models, has changed not only the number of model parameters but also how the models are served and the economics of useful inference. It changes which tensors remain active for a given token, what must stay close together because of local bandwidth requirements, and which transfers can instead be handled by weaker network connections.
According to the article, inference runs within a cluster managed by an orchestration layer, such as NVIDIA Dynamo, Mooncake or a custom scheduler, which works with inference servers such as vLLM and SQlang (as written in the source). Request processing passes through several modes: prefill (processing the initial block of new tokens all at once), midfill (connecting a new request to an existing context), decode attention and decode experts, in which each newly generated token is refined by selecting from the full set of experts. The session context (KV cache) allows previous conversation turns to be continued even after hours or days.
According to the source, the individual modes place different demands on computation, memory and networking – prefill achieves high arithmetic intensity (the ratio of computation to data movement), while the decode phases have low arithmetic intensity because a large volume of existing context or expert weights needs to be moved relative to a small number of new tokens. Data movement in memory and bandwidth are therefore described as key performance constraints. The article also mentions the trade-off between aggregated deployment (where the model stays on one server) and disaggregated deployment (where the orchestrator looks for an available, suitably configured machine for the next phase).
The available source text is incomplete. Details can be found in the source article.
Why it matters
The text explains the technical reasons why running Mixture of Experts models differs in cost and architecture from running traditional dense models – this is relevant mainly to engineers designing and operating inference infrastructure and to teams deciding on hardware and network interconnect purchases for deploying large language models.
Relevant practical impact
What this means
For a business
According to the source, deploying models with the Mixture of Experts architecture requires adapting orchestration, KV cache management and network/memory capacity to the individual inference phases (prefill, midfill, decode), which affects the cost and performance of owned or rented inference infrastructure.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.