Skip to content
worth noting Other

Mixture of Experts changes inference architecture: data movement and cluster orchestration

only one source so far

An analysis by SemiAnalysis describes how the Mixture of Experts architecture changes the requirements for inference infrastructure – which tensors remain active, how data moves between the prefill, midfill and decode phases, and how orchestration layers such as NVIDIA Dynamo or Mooncake coordinate this.

According to the source, the Mixture of Experts architecture, now commonly used in leading models, has changed not only the number of model parameters but also how the models are served and the economics of useful inference. It changes which tensors remain active for a given token, what must stay close together because of local bandwidth requirements, and which transfers can instead be handled by weaker network connections.

According to the article, inference runs within a cluster managed by an orchestration layer, such as NVIDIA Dynamo, Mooncake or a custom scheduler, which works with inference servers such as vLLM and SQlang (as written in the source). Request processing passes through several modes: prefill (processing the initial block of new tokens all at once), midfill (connecting a new request to an existing context), decode attention and decode experts, in which each newly generated token is refined by selecting from the full set of experts. The session context (KV cache) allows previous conversation turns to be continued even after hours or days.

According to the source, the individual modes place different demands on computation, memory and networking – prefill achieves high arithmetic intensity (the ratio of computation to data movement), while the decode phases have low arithmetic intensity because a large volume of existing context or expert weights needs to be moved relative to a small number of new tokens. Data movement in memory and bandwidth are therefore described as key performance constraints. The article also mentions the trade-off between aggregated deployment (where the model stays on one server) and disaggregated deployment (where the orchestrator looks for an available, suitably configured machine for the next phase).

The available source text is incomplete. Details can be found in the source article.

What changed

Why it matters

The text explains the technical reasons why running Mixture of Experts models differs in cost and architecture from running traditional dense models – this is relevant mainly to engineers designing and operating inference infrastructure and to teams deciding on hardware and network interconnect purchases for deploying large language models.

Relevant practical impact

What this means

01

For a business

According to the source, deploying models with the Mixture of Experts architecture requires adapting orchestration, KV cache management and network/memory capacity to the individual inference phases (prefill, midfill, decode), which affects the cost and performance of owned or rented inference infrastructure.

Development
What to decide When planning the deployment or purchase of inference infrastructure for Mixture of Experts models, account for the differing memory, bandwidth and orchestration requirements of the prefill, midfill and decode phases.
More business impacts →
inference KV-cache Memory management mixture-of-experts orchestrace vLLM

Check the original

Event sources

only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.

1
SemiAnalysis (newsletter feed — hardware, chips, AI economics) independent context · first detected Computation and Data Movement for Inference