architectures
Mixture of Experts
An architecture where multiple specialized sub-networks (experts) exist within a model, with a routing mechanism that activates only a subset of experts for each input. This allows models to have enormous total parameter counts while only using a fraction of compute per inference.
In practice
Mixtral 8x7B has 46.7B total parameters but only activates 12.9B per token, making it fast despite its size.