Skip to content

deployment

KV Cache

Key-Value Cache, an optimization that stores previously computed attention key and value tensors to avoid redundant computation during autoregressive generation. KV caching significantly speeds up inference but increases memory usage proportional to sequence length.

In practice

Without KV cache, generating the 1000th token would require reprocessing all 999 previous tokens from scratch.