architectures
Attention Sink
A phenomenon where transformer models allocate disproportionate attention to the first few tokens of a sequence regardless of their semantic importance. Understanding attention sinks has led to optimizations for streaming and long-context inference.
In practice
Research showed that keeping the first few tokens' KV cache entries helps maintain model quality during streaming inference.