Skip to content

architectures

Attention Sink

A phenomenon where transformer models allocate disproportionate attention to the first few tokens of a sequence regardless of their semantic importance. Understanding attention sinks has led to optimizations for streaming and long-context inference.

In practice

Research showed that keeping the first few tokens' KV cache entries helps maintain model quality during streaming inference.