deployment
Flash Attention
An optimized attention computation algorithm that reduces memory usage from quadratic to near-linear by processing attention in blocks. Flash Attention uses tiling and kernel fusion to minimize GPU memory reads and writes, significantly speeding up training and inference.
In practice
Flash Attention 2 can process sequences up to 16x longer on the same GPU by reducing memory overhead.