Speculative Decoding: Faster LLM Inference
Speculative decoding uses a small draft model to guess tokens a big model verifies in one parallel pass, cutting LLM latency with no change to output.
Text generation from a transformer is stubbornly sequential. To produce the tenth token the model needs the ninth. So a 200-token answer means 200 forward passes through the full network, one after another, each waiting on the last. For a single request, that serial chain is where almost all of your latency goes, and no amount of GPU you throw at one stream makes a single decode step faster.
Speculative decoding attacks that chain directly. Instead of spending one expensive forward pass per token, you let a small, cheap model guess several tokens ahead. Then the big model checks all of those guesses in a single pass. When the guesses are good, you get several tokens for the price of one big forward pass, and the output is provably identical to what the big model would have produced on its own.
This post is for engineers running LLM inference who care about tail latency on a single stream: chat, IDE autocomplete, an agent’s inner loop. I’ll cover:
- why decoding is slow in the first place,
- the accept/reject trick that keeps the output exact,
- the one number that decides whether any of this helps,
- the places it quietly stops helping.
Why one token per pass is the whole problem
A draft model can only help because of what the GPU actually does during decode, and it isn’t what most people assume.
When you generate one token at batch size one, the arithmetic is tiny: a few matrix-vector products. The expensive part is reading the model’s weights out of high-bandwidth memory into the compute units. For a multi-billion-parameter model that read dominates, and the tensor cores (the GPU’s matrix-multiply units) sit mostly idle waiting on memory. Decode is memory-bandwidth bound, not compute bound. The same pressure is what makes the KV cache worth so much: the bottleneck is moving data, not doing math.
Here is the lever. Running the model over ten positions at once reads the weights the same number of times as running it over one position. The extra positions ride along on matrix-matrix products the hardware was underusing anyway. So a forward pass that scores ten candidate tokens takes almost the same wall-clock time as a pass that produces one. You have spare compute during decode, and speculative decoding spends it on verifying guesses instead of leaving it on the floor.
The trick: draft, then verify in parallel
Split the work across two models:
- A small draft model (say a 1B) generates a short run of tokens the cheap way, one at a time.
- The large target model, the one whose output you actually want, takes the prompt plus all of those draft tokens and runs a single forward pass over the whole run.
Because of the batching argument above, that one pass gives you the target model’s own next-token distribution at every draft position at once.
Now you compare. Walk left to right and keep each draft token the target agrees with. At the first disagreement, stop, throw away the rest of the draft, and replace that token with one sampled from the target’s distribution. A round looks like this:
The payoff is in the last row. Three guesses were right, so they cost nothing beyond the draft. The fourth was wrong, but the target pass had already computed the correct token for that position, so you take it as a bonus. That is four tokens advanced for one target forward pass.
Even if the draft had been wrong on the very first token, you would still get one correct token out, which is exactly what a normal decode step gives you. So a round can never do worse than plain decoding on token count.
Why the output doesn’t change
What makes this usable in production, rather than a quality gamble, is that speculative decoding is not an approximation. The accept/reject rule is modified rejection sampling, built so that the sequence you emit is drawn from exactly the target model’s distribution.
Here are the mechanics for one draft token x, with draft probability q(x) and target probability p(x):
- Accept it with probability
min(1, p(x)/q(x)). - If it’s rejected, sample the replacement from the normalized positive part of
p(x) − q(x).
Chen et al. at DeepMind and Leviathan et al. at Google worked out independently in 2023 that this rule leaves the marginal distribution over emitted tokens identical to sampling from the target directly. Concretely, at the same temperature and seed, greedy speculative decoding produces the same string as greedy decoding from the target, token for token. You aren’t trading quality for speed. You’re trading idle GPU compute for speed.
That guarantee is the reason to prefer this over cheaper-sounding shortcuts, like just serving the small model. The target catches and corrects the draft model’s mistakes every round, so the draft only ever influences speed, never what you say.
The one number that decides everything: acceptance rate
Whether any of this pays off rides on a single quantity, the acceptance rate α: the fraction of draft tokens the target keeps. If you draft γ tokens per round, the expected number of tokens you advance per target pass is:
E[tokens] = (1 − α^(γ+1)) / (1 − α)That formula is the ceiling on your speedup, before you pay for the draft model’s own compute. Plotting it for two draft lengths shows why acceptance rate, not draft length, is the thing to obsess over:
Two things fall out of the curve:
- Drafting more tokens helps only when acceptance is already high. At α = 0.9, going from 4 to 7 draft tokens buys you real throughput. At α = 0.5, both drafts land in the same low huddle, because the run gets cut short long before token 7.
- The useful region is narrow. You want a draft model that agrees with the target most of the time. In practice that means a small model from the same family, trained on similar data, ideally with the same tokenizer. A random small model with α ≈ 0.3 will spend more time drafting rejected tokens than it saves.
Acceptance rate also depends on content, which surprises people. Boilerplate-heavy code, structured output, and formulaic prose accept at high rates, because the next token is easy to guess. Open-ended, high-entropy generation accepts poorly. The same deployment can see very different speedups depending on what users ask.
Variants: when you don’t want a second model
Standing up and serving a whole second model isn’t free. It costs GPU memory that competes with your KV cache, and it’s one more thing to deploy and keep aligned. Several methods keep the draft-then-verify structure but drop the separate draft model:
- Medusa adds extra prediction heads to the target model itself, so a single model proposes several future tokens and then verifies them. No second network to host.
- EAGLE drafts at the feature level (the model’s internal hidden states) rather than the token level. It reports higher acceptance rates than a naive small-model draft, which is exactly the number that matters.
- Prompt lookup / n-gram decoding (writeup) uses no model at all for drafting. It searches the prompt and the text so far for a matching phrase and proposes its continuation. For summarization, RAG, and editing (anything where the output copies spans from the input), it is startlingly effective and costs nothing to run.
You don’t need to wire any of this up yourself. Hugging Face wraps the two-model form as assisted generation: pass an assistant_model to generate() and you get speculative decoding with no orchestration code. vLLM supports the draft-model, n-gram, and EAGLE variants behind a config block, which is the least-effort way to try it against a real serving setup.
Where it stops helping
The mechanism is clean, but the operational reality has sharp edges.
High batch size erases the benefit. The whole argument rested on decode being memory-bound with spare compute. Once you batch many concurrent requests (a busy serving tier), the GPU becomes compute bound, and the verification pass over γ extra positions is no longer close to free. Speculative decoding is a single-stream, latency-optimizing technique: it shines when one user is waiting on one long answer, and it can hurt aggregate throughput when the box is already saturated with traffic. Serving stacks handle this by scaling speculation down or off as load rises, but if you benchmark at batch one and deploy at batch sixty-four, the win you measured evaporates.
A weak draft is worse than none. Below roughly α ≈ 0.5, you spend draft compute on tokens that get rejected, verify them anyway, and come out behind plain decoding. Measure α on your actual traffic before committing, because a draft that looks good on a benchmark can fall apart on your prompts.
Memory contention is real. The draft model’s weights and its own KV cache live on the same card as the target’s, competing for the space that would otherwise hold context. On a memory-tight deployment, a self-drafting method like Medusa or n-gram lookup avoids paying twice.
Tuning γ is a live tradeoff. Draft too few tokens and you leave parallel verification capacity unused. Draft too many while acceptance is mediocre, and you burn draft steps generating tokens that get thrown away. The sweet spot moves with acceptance rate, so make γ configurable rather than baking in a constant.
What I would reach for
If I were adding this to something like Archi, the RAG (retrieval-augmented generation) copilot I worked on for CMS computing operations, I’d start with prompt lookup decoding, not a draft model. Archi’s answers quote heavily from retrieved logbook entries and tickets, so much of the output is spans copied from the context. That is exactly the case where n-gram drafting gets a high acceptance rate, for zero extra GPU and zero extra deployment. The same logic applies to CloudCanvasAI, where the model is often editing or restating a document that’s already in context.
I’d reach for a real draft model only after measuring two things: that the generative, non-copied part of the output is your latency floor, and that a same-family small model actually clears α ≈ 0.7 on your traffic.
The framing that keeps me honest: speculative decoding doesn’t make the model faster. It makes better use of the compute a single decode step was already wasting. That’s why it pairs with everything else in the inference stack rather than replacing any of it:
- the KV cache still saves the recompute,
- batching still fills the GPU under load,
- speculation cashes in the idle compute that’s left on a latency-bound single stream.
Know which of those problems you actually have before you add a second model to your serving path.
Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.