architectures
Sparse Attention
An attention variant where each token attends to only a subset of other tokens rather than all tokens in the sequence. Sparse attention reduces the quadratic computational cost of standard attention, enabling processing of much longer sequences.
In practice
Longformer uses sparse attention patterns (local + global) to process documents up to 16K tokens efficiently.