deployment
Speculative Decoding
An inference optimization where a smaller draft model quickly generates candidate tokens, and the larger target model verifies them in parallel. This can speed up generation by 2-3x since verification is faster than sequential generation.
In practice
A 7B draft model proposes 5 tokens at once, and the 70B model verifies all 5 in a single forward pass, speeding up generation.