Skip to content

architectures

Mixture of Depths

An architecture optimization where not all tokens are processed by every transformer layer. Mixture of Depths routes tokens through different numbers of layers based on their complexity, reducing compute for simpler tokens.

In practice

Function words like 'the' and 'is' might skip several layers, while complex terms get processed by all layers.