Speculative Decoding
Free tokens, provably distribution-preserving.
Leviathan et al. is short, and the value is the proof: accept a draft token with probability min(1, p/q), resample from the normalized residual on rejection, and the output distribution is exactly the target model's. Not approximately. Work it until you can reproduce it on a whiteboard, because that precise question gets asked, and hand-waving it is what interviewers listen for.
The lineage task is there because EAGLE won. Drafting from the target's own hidden features beat separate draft models, EAGLE-3 now ships as the default speculative method in the major engines, and DeepSeek-V3's MTP shows the idea migrating into the base model itself. One judgment call to remember: speculation spends spare bandwidth, so at high batch sizes, where decode is no longer leaving bandwidth on the table, it can make throughput worse. Knowing when not to speculate is the senior answer.
- Inference EngineeringPhilip Kielybook↗
min(1, p/q) acceptance + residual resampling. Interviews test whether you can prove exactness.
EAGLE-3 is the shipping default in vLLM/SGLang/TRT-LLM. Also note MTP in DeepSeek-V3 as built-in speculation.
Implement the greedy draft-verify loop (draft k, verify in ONE target call, accept prefix, rollback, bonus token) against the harness's target + noisy draft. Graded on exact equality with pure target decoding and ≥1.5 tokens per verify call.
build+250 XPauto-verifiedAcceptance math, expected tokens/step, when speculation hurts. 75% to pass.
drill+80 XPauto-verified