English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

vToken: Adding Virtual Memory to LLM KV Caches in vLLM

Forum topic · ✨步子哥 · 2026-08-16

Summary

vToken is a token-level virtualization layer for LLM KV cache management, built on vLLM, from researchers at the National University of Defense Technology and Peking University. The paper identifies a granularity mismatch: eviction policies like H2O and StreamingLLM operate at token granularity, while PagedAttention's runtime manages memory at block granularity (16 tokens per block), causing 40%-60% intra-block fragmentation where evicted token slots cannot be reclaimed. vToken decouples logical tokens from physical blocks using a Token Table (analogous to a page table), an evict_token interface that marks tokens dead without moving data, and an asynchronous reclamation backend that relocates surviving tokens on a separate CUDA stream. On H100 80GB with Mistral-7B and Llama-3.1-8B, vToken improves average memory utilization by ~21.9%, reduces retained blocks by 27%-72%, raises throughput by up to 37% (up to 103.7% under Scissorhands), and doubles concurrent serving capacity (C=11 to C=22) under SLA constraints. Policy adapters shrink from 500+ lines to under 50. Limitations include single-GPU scope, conservative skipping of shared prefix blocks, and no public code release.

A Number That Keeps Inference Engineers Up at Night

Serving a 13B-parameter LLaMA model: weights take 26 GB, manageable. But at 32K context length, the KV cache alone eats 25 GB — 0.8 MB per token, 32K tokens = 25.6 GB. A single long-context request consumes a third of an 80GB GPU.

The community's answer: PagedAttention (the core of vLLM) manages KV cache in fixed-size blocks (typically 16 tokens/block), like OS paging. On top of that, token-level eviction algorithms emerged — H2O keeps high-attention "heavy hitter" tokens, StreamingLLM keeps attention sinks + a recent window, Scissorhands keeps consistently attended tokens.

Problem solved? No. The real problem was just starting.

The Overlooked "Granularity Mismatch"

A team from the National University of Defense Technology and Peking University, in an August 2026 paper called vToken, precisely measured a problem everyone had hit but nobody had formally defined:

> Eviction policies make decisions at token granularity; the runtime manages memory at block granularity. These granularities don't match.

Example: a block has 16 token slots. H2O says "10 tokens in this block are unimportant — evict them." But the block can't be freed — 6 live tokens remain. Those 10 evicted slots become intra-block fragmentation.

Running H2O eviction on vLLM at 16K context, batch size 16:

  • Most allocated blocks are less than 50% utilized
  • Intra-block waste ratio F reaches 40%-60%
  • So 40%-60% of the logical space saved by eviction is trapped in "half-dead" physical blocks, unreusable by other requests. Like laying off 10 people from a 16-person office but being unable to cancel the lease because 6 people remain.

    vToken's Core Insight: Add a Virtualization Layer

    The fix isn't a better eviction algorithm, nor smaller blocks (that just relocates the fragmentation). vToken adds a token-level virtualization boundary above the block-level runtime — essentially virtual memory for KV caches:

    | OS Virtual Memory | vToken | |---|---| | Process virtual address space | Request's logical token sequence | | Physical page frames | PagedAttention's physical KV blocks | | Page table | Token Table | | Page fault + page reclamation | Physical reclamation backend | | Process unaware of physical layout | Eviction policy unaware of block layout |

    Token Table: A "Page Table" for KV Caches

    Each request gets a Token Table recording, per logical token ID:

  • Physical location: (block ID, offset)
  • Alive bit: alive or dead
  • When a policy calls evict_token(req_id, token_id), only the alive bit is updated — no KV data moves. Like an OS marking a page invalid without immediately freeing the frame.

    Metadata overhead is tiny: a 16K-token sequence needs only 256 KB (16 bytes/entry), under 0.1% of a 7B model's FP16 KV cache footprint (~2 GB).

    Physical Reclamation Backend: Asynchronous "Defragmentation"

    1. Reclamation eligibility check: scan block utilization for low-occupancy blocks 2. Headroom-aware admission: ensure enough free destination blocks exist — otherwise wait 3. Migration planning: greedily pack live tokens from low-occupancy blocks into destinations 4. Asynchronous copy: KV copies run on a separate CUDA stream, not blocking decoding

    Migration plans are committed, copies happen in the background, and a CUDA event ensures relevant copies complete before the next attention computation, then slot mappings are refreshed. No global synchronization point.

    Three Design Challenges

    C1: Dual-view consistency. Tokens die independently, but blocks free only when fully evacuated. The Token Table maintains both token-level liveness and block-level occupancy in one structure.

    C2: Safe reclamation during decoding. Half-migrated reads would corrupt attention. Solution: slot mappings update atomically after migration; attention kernels read only the slot mapping; CUDA events gate mapping refreshes.

    C3: Policy-agnostic amortized cost. All policies (H2O, StreamingLLM, Random) share one scheduler, block manager, and worker. vToken exposes evict_token, sync_new_tokens, apply_moves — policies call the first two. Policy adapters shrink from 500+ lines to under 50.

    The Numbers

    On NVIDIA H100 80GB with Mistral-7B and Llama-3.1-8B, on ShareGPT and LongBench, across H2O, Scissorhands, and Random eviction policies:

    Memory Efficiency

    | Metric | Naive-Evict | vToken | Gain | |---|---|---|---| | Avg memory utilization (Llama-3.1-8B) | baseline | +21.88% | — | | Avg memory utilization (Mistral-7B) | baseline | +21.67% | — | | Retained blocks | baseline | -27.2% ~ -72.3% | — |

    Key insight: Naive-Evict's losses come mainly from the block-level runtime failing to convert token-level liveness into reclaimable physical capacity.

    Throughput

    Under the same SLA (p95 latency ≤ 1.05x baseline):

  • Mistral-7B: throughput +9.9%~37.3%, p95 latency -9.9%~27.5%
  • Llama-3.1-8B: average throughput +18.9%, p95 latency -14.7%
  • Best under Scissorhands: throughput +33.3%~103.7%, p95 latency -21.8%~33.0%
  • Random eviction benefits most — scattered live tokens make fragmentation worst. The worse the fragmentation, the more vToken is worth.

    Capacity Frontier

    At gpu_mem_util=0.35 (5427 available KV blocks):

  • Native vLLM and Naive-Evict: max C=5 concurrency
  • vToken: C=8 — 60% more concurrency
  • At gpu_mem_util=0.50 (11519 blocks):

  • Native vLLM and Naive-Evict: C=11
  • vToken: C=22 — doubled concurrency
  • vToken at C=8 delivers 180.3 tokens/s, near its own C=5 peak of 203.2 — graceful degradation, not collapse.

    Overhead

  • Indirection-only overhead (hook installed, no eviction/reclamation): <1.0% change in throughput and p95
  • Main cost: CPU-side planner opportunity checks, not KV migration
  • Async copies complete over multiple decode steps with no explicit waits
  • Engineering Insights: Why This Matters More Than It Looks

    1. Granularity Isomorphism, Again

  • Eviction policies operate on tokens
  • PagedAttention operates on blocks (16 tokens)
  • Mismatch → 40%-60% intra-block fragmentation
  • vToken's token-level virtualization aligns the granularities → fragmentation eliminated
  • The eviction algorithm isn't bad — the runtime's granularity can't keep up with the eviction algorithm's.

    2. Solving Problems at a Different Layer

    vToken doesn't optimize within the block layer; it adds a layer above. Same pattern as other "change the layer" solutions: octopus RNA editing (edit the blueprint, not the DNA), slime mold externalized memory, SOPHIA division of labor, Möbius RoPE topology intervention. Don't grind away at the original layer — switch layers.

    3. OS Design Patterns Migrating to AI

  • Token Table = Page Table: logical-to-physical mapping
  • evict_token = marking a page invalid: metadata only
  • Reclamation backend = page-reclaim daemon: async compaction
  • Async CUDA stream = DMA transfer: non-blocking
  • Slot mapping refresh = TLB flush: visibility of new layout
  • Sixty years of virtual memory design prove that decoupling logical address space from physical memory is the right abstraction for scarce memory. vToken carries it into KV cache management.

    4. Policy Adapters: 500+ Lines → <50

    Previously, integrating an eviction policy into vLLM required touching 4-6 files and 500+ lines due to tight coupling. vToken isolates shared logic; a policy only implements SelectVictims. New-policy experimentation gets dramatically cheaper — just like virtual memory freed application developers from physical memory layout.

    An Unsolved Problem

    vToken honestly states its limits:

  • Single node, single GPU: the prototype targets the single-GPU decode fast path; distributed scheduling and cross-device KV migration are out of scope
  • Shared prefix blocks conservatively skipped: copy-on-write could solve this but isn't implemented
  • Planner overhead: the main CPU cost is opportunity checking, not KV migration itself
  • Deeper issue: vToken is currently a vLLM-only implementation. TensorRT-LLM, SGLang, and other engines would each need to implement the hooks, and no public code is released (the authors are from NUDT; release may be restricted), limiting immediate community adoption.

    Closing Thought: Virtualization as a Universal Design Pattern

    AI systems are re-walking the path of computer systems:

  • PagedAttention = physical memory paging
  • vToken = virtual memory (logic-physical decoupling)
  • KV eviction policies = page replacement (LRU, LFU)
  • Prefix cache = shared memory
  • Chunked prefill = prefetching
vToken's value isn't just "21% KV memory saved" — it's proving that the 60-year-old virtualization pattern still works in the LLM inference era.

Next to virtualize? My bet: attention computation itself — letting attention kernels access KV through indirection, so sparse attention, sliding windows, and dynamic routing can share one runtime without each rewriting CUDA kernels.

vToken's lesson: when two layers' granularities mismatch, don't grind at either layer — add a virtualization layer.

---

Paper: vToken: Token-Level Virtualization for Reclaimable KV Caches Authors: Yuanhang Gao, Xiangrui Yang, Yuanfeng Chen, Hongjia Chen, Qianru Lv, Wenfei Wu, Dongsheng Li Institutions: National University of Defense Technology; Peking University Implemented on vLLM v0.18.0; no public repository yet

Tags

#llm-inference#kv-cache#vllm#pagedattention#virtual-memory#h2o#gpu-memory#systems-design

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633563