Skip to content

architectures

Sparse Attention

An attention variant where each token attends to only a subset of other tokens rather than all tokens in the sequence. Sparse attention reduces the quadratic computational cost of standard attention, enabling processing of much longer sequences.

In practice

Longformer uses sparse attention patterns (local + global) to process documents up to 16K tokens efficiently.