Skip to content

deployment

Speculative Decoding

An inference optimization where a smaller draft model quickly generates candidate tokens, and the larger target model verifies them in parallel. This can speed up generation by 2-3x since verification is faster than sequential generation.

In practice

A 7B draft model proposes 5 tokens at once, and the 70B model verifies all 5 in a single forward pass, speeding up generation.