MoE definition
Mixture of experts (MoE) is a neural network architecture in which a model contains many specialized subnetworks, called experts, and a router activates only a few of them for each input token. This lets a model have a very large total number of parameters while using only a fraction for each prediction, improving quality without a proportional increase in compute.
How mixture of experts works
In a standard dense transformer, every parameter takes part in processing every token. In an MoE transformer, some layers, usually the feed-forward layers, are replaced by a set of experts: for example 8, 64 or more parallel feed-forward networks. A small gating network, the router, looks at each token and selects the top one or two experts, or a few more in fine-grained designs, to process it, then combines their outputs weighted by its scores.
Because only the selected experts run, compute per token depends on the active parameters rather than the total. A model can hold the capacity of a very large network while running roughly as fast as a much smaller one. The idea dates back to research in the early 1990s and was scaled up for language models by Google researchers from the late 2010s onward.
Examples of MoE models
Mistral AI's Mixtral 8x7B, released in 2023, brought MoE to widely used open-weight models, with eight experts per layer and two active per token. DeepSeek's V3 and R1 models then used fine-grained experts plus shared experts to reach strong performance at lower training cost. MoE has since become the standard design for flagship open-weight models: recent DeepSeek, Qwen, Kimi, GLM, gpt-oss and Llama releases all include MoE models, often alongside smaller dense variants.
Several leading proprietary models are reported or confirmed to use MoE designs as well, although providers often do not publish architecture details. For buyers, architecture matters less than measured quality, cost and latency on their own tasks, and those are what should drive model selection.
Benefits and trade-offs
MoE changes the economics of large models, but it moves complexity elsewhere rather than removing it entirely, mostly into memory and serving infrastructure. The main benefits and costs for teams training or serving these models in production are:
- Benefit: more model capacity for the same compute per token, often improving quality per unit of training and inference cost
- Benefit: faster inference than a dense model of the same total size
- Cost: all experts must sit in memory, so GPU memory needs follow total parameters, not active ones
- Cost: routing adds complexity, including load balancing so some experts are not overused while others sit idle
- Cost: distributed serving across GPUs needs fast interconnects and specialized inference engines
- Cost: fine-tuning can be less stable than with dense models
What MoE means for businesses using AI
Most teams meet MoE indirectly, through APIs whose price and speed reflect these efficiencies. It becomes a practical question when self-hosting open models: an MoE model with a modest active parameter count may still need several GPUs because of its total size, while a dense model of similar quality might fit on one. Quantization reduces memory but does not change the basic trade-off.
When choosing a model, compare quality on your own tasks, memory footprint, throughput and cost per request rather than parameter counts. Nexzem benchmarks dense and MoE open models alongside hosted APIs during LLM development, and plans inference infrastructure around the results.