Long Context & KV Policy
When the cache can't hold everything, something has to give.
StreamingLLM and H2O stake out the two poles of the retention question: keep positions (the start plus a sliding window) or keep what attention actually uses (the heavy hitters). The attention-sink finding is worth internalizing on its own. The model dumps attention mass on the first tokens because softmax has to put it somewhere, and evicting them collapses quality — the kind of empirical quirk that separates people who've read the papers from people who've read the tweets.
The vLLM hybrid-KV doc is the production counterweight. Sliding-window and Mamba-style layers break the uniform-block assumption from PagedAttention, and someone has to make the block tables cope. In the survey, watch for the recurring pattern: every eviction and compression scheme trades quality for capacity somewhere invisible. That's why the eval discipline from the quantization phase applies here verbatim.
- Inference EngineeringPhilip Kielybook↗
The two poles of KV retention policy: keep the start + a sliding window, or keep what attention actually uses. Both ship as engine features.
Gemma's sliding window and Mamba/hybrid layers need different block layouts — vLLM's hybrid KV manager is the production answer.
Map the design space beyond quantization. Also skim sparse-attention serving (MInference-class) and prompt compression.