architectures
Mixture of Depths
An architecture optimization where not all tokens are processed by every transformer layer. Mixture of Depths routes tokens through different numbers of layers based on their complexity, reducing compute for simpler tokens.
In practice
Function words like 'the' and 'is' might skip several layers, while complex terms get processed by all layers.