Skip to content

architectures

Mixture of Experts

An architecture where multiple specialized sub-networks (experts) exist within a model, with a routing mechanism that activates only a subset of experts for each input. This allows models to have enormous total parameter counts while only using a fraction of compute per inference.

In practice

Mixtral 8x7B has 46.7B total parameters but only activates 12.9B per token, making it fast despite its size.