InferQuest
← Quest map
Inference Phase 1 · Engine Core

Long Context & KV Policy

When the cache can't hold everything, something has to give.

0/3 · 190 XP
You’re previewing the curriculum. Sign in (top right) to track progress, take knowledge checks, and unlock the verifiers.
Briefing

StreamingLLM and H2O stake out the two poles of the retention question: keep positions (the start plus a sliding window) or keep what attention actually uses (the heavy hitters). The attention-sink finding is worth internalizing on its own. The model dumps attention mass on the first tokens because softmax has to put it somewhere, and evicting them collapses quality — the kind of empirical quirk that separates people who've read the papers from people who've read the tweets.

The vLLM hybrid-KV doc is the production counterweight. Sliding-window and Mamba-style layers break the uniform-block assumption from PagedAttention, and someone has to make the block tables cope. In the survey, watch for the recurring pattern: every eviction and compression scheme trades quality for capacity somewhere invisible. That's why the eval discipline from the quantization phase applies here verbatim.

From the library
Tasks
  • The two poles of KV retention policy: keep the start + a sliding window, or keep what attention actually uses. Both ship as engine features.

    paper+70 XPresource ↗
  • Gemma's sliding window and Mamba/hybrid layers need different block layouts — vLLM's hybrid KV manager is the production answer.

    read+60 XPresource ↗
  • Map the design space beyond quantization. Also skim sparse-attention serving (MInference-class) and prompt compression.

    read+60 XPresource ↗