A Number That Keeps Inference Engineers Up at Night
Serving a 13B-parameter LLaMA model: weights take 26 GB, manageable. But at 32K context length, the KV cache alone eats 25 GB — 0.8 MB per token, 32K tokens = 25.6 GB. A single long-context request consumes a third of an 80GB GPU.
The community's answer: PagedAttention (the core of vLLM) manages KV cache in fixed-size blocks (typically 16 tokens/block), like OS paging. On top of that, token-level eviction algorithms emerged — H2O keeps high-attention "heavy hitter" tokens, StreamingLLM keeps attention sinks + a recent window, Scissorhands keeps consistently attended tokens.
Problem solved? No. The real problem was just starting.
The Overlooked "Granularity Mismatch"
A team from the National University of Defense Technology and Peking University, in an August 2026 paper called vToken, precisely measured a problem everyone had hit but nobody had formally defined:
> Eviction policies make decisions at token granularity; the runtime manages memory at block granularity. These granularities don't match.
Example: a block has 16 token slots. H2O says "10 tokens in this block are unimportant — evict them." But the block can't be freed — 6 live tokens remain. Those 10 evicted slots become intra-block fragmentation.
Running H2O eviction on vLLM at 16K context, batch size 16:
- Most allocated blocks are less than 50% utilized
- Intra-block waste ratio F reaches 40%-60%
- Physical location: (block ID, offset)
- Alive bit: alive or dead
- Mistral-7B: throughput +9.9%~37.3%, p95 latency -9.9%~27.5%
- Llama-3.1-8B: average throughput +18.9%, p95 latency -14.7%
- Best under Scissorhands: throughput +33.3%~103.7%, p95 latency -21.8%~33.0%
- Native vLLM and Naive-Evict: max C=5 concurrency
- vToken: C=8 — 60% more concurrency
- Native vLLM and Naive-Evict: C=11
- vToken: C=22 — doubled concurrency
- Indirection-only overhead (hook installed, no eviction/reclamation): <1.0% change in throughput and p95
- Main cost: CPU-side planner opportunity checks, not KV migration
- Async copies complete over multiple decode steps with no explicit waits
- Eviction policies operate on tokens
- PagedAttention operates on blocks (16 tokens)
- Mismatch → 40%-60% intra-block fragmentation
- vToken's token-level virtualization aligns the granularities → fragmentation eliminated
- Token Table = Page Table: logical-to-physical mapping
- evict_token = marking a page invalid: metadata only
- Reclamation backend = page-reclaim daemon: async compaction
- Async CUDA stream = DMA transfer: non-blocking
- Slot mapping refresh = TLB flush: visibility of new layout
- Single node, single GPU: the prototype targets the single-GPU decode fast path; distributed scheduling and cross-device KV migration are out of scope
- Shared prefix blocks conservatively skipped: copy-on-write could solve this but isn't implemented
- Planner overhead: the main CPU cost is opportunity checking, not KV migration itself
- PagedAttention = physical memory paging
- vToken = virtual memory (logic-physical decoupling)
- KV eviction policies = page replacement (LRU, LFU)
- Prefix cache = shared memory
- Chunked prefill = prefetching
So 40%-60% of the logical space saved by eviction is trapped in "half-dead" physical blocks, unreusable by other requests. Like laying off 10 people from a 16-person office but being unable to cancel the lease because 6 people remain.
vToken's Core Insight: Add a Virtualization Layer
The fix isn't a better eviction algorithm, nor smaller blocks (that just relocates the fragmentation). vToken adds a token-level virtualization boundary above the block-level runtime — essentially virtual memory for KV caches:
| OS Virtual Memory | vToken | |---|---| | Process virtual address space | Request's logical token sequence | | Physical page frames | PagedAttention's physical KV blocks | | Page table | Token Table | | Page fault + page reclamation | Physical reclamation backend | | Process unaware of physical layout | Eviction policy unaware of block layout |
Token Table: A "Page Table" for KV Caches
Each request gets a Token Table recording, per logical token ID:
When a policy calls evict_token(req_id, token_id), only the alive bit is updated — no KV data moves. Like an OS marking a page invalid without immediately freeing the frame.
Metadata overhead is tiny: a 16K-token sequence needs only 256 KB (16 bytes/entry), under 0.1% of a 7B model's FP16 KV cache footprint (~2 GB).
Physical Reclamation Backend: Asynchronous "Defragmentation"
1. Reclamation eligibility check: scan block utilization for low-occupancy blocks 2. Headroom-aware admission: ensure enough free destination blocks exist — otherwise wait 3. Migration planning: greedily pack live tokens from low-occupancy blocks into destinations 4. Asynchronous copy: KV copies run on a separate CUDA stream, not blocking decoding
Migration plans are committed, copies happen in the background, and a CUDA event ensures relevant copies complete before the next attention computation, then slot mappings are refreshed. No global synchronization point.
Three Design Challenges
C1: Dual-view consistency. Tokens die independently, but blocks free only when fully evacuated. The Token Table maintains both token-level liveness and block-level occupancy in one structure.
C2: Safe reclamation during decoding. Half-migrated reads would corrupt attention. Solution: slot mappings update atomically after migration; attention kernels read only the slot mapping; CUDA events gate mapping refreshes.
C3: Policy-agnostic amortized cost. All policies (H2O, StreamingLLM, Random) share one scheduler, block manager, and worker. vToken exposes evict_token, sync_new_tokens, apply_moves — policies call the first two. Policy adapters shrink from 500+ lines to under 50.
The Numbers
On NVIDIA H100 80GB with Mistral-7B and Llama-3.1-8B, on ShareGPT and LongBench, across H2O, Scissorhands, and Random eviction policies:
Memory Efficiency
| Metric | Naive-Evict | vToken | Gain | |---|---|---|---| | Avg memory utilization (Llama-3.1-8B) | baseline | +21.88% | — | | Avg memory utilization (Mistral-7B) | baseline | +21.67% | — | | Retained blocks | baseline | -27.2% ~ -72.3% | — |
Key insight: Naive-Evict's losses come mainly from the block-level runtime failing to convert token-level liveness into reclaimable physical capacity.
Throughput
Under the same SLA (p95 latency ≤ 1.05x baseline):
Random eviction benefits most — scattered live tokens make fragmentation worst. The worse the fragmentation, the more vToken is worth.
Capacity Frontier
At gpu_mem_util=0.35 (5427 available KV blocks):
At gpu_mem_util=0.50 (11519 blocks):
vToken at C=8 delivers 180.3 tokens/s, near its own C=5 peak of 203.2 — graceful degradation, not collapse.
Overhead
Engineering Insights: Why This Matters More Than It Looks
1. Granularity Isomorphism, Again
The eviction algorithm isn't bad — the runtime's granularity can't keep up with the eviction algorithm's.
2. Solving Problems at a Different Layer
vToken doesn't optimize within the block layer; it adds a layer above. Same pattern as other "change the layer" solutions: octopus RNA editing (edit the blueprint, not the DNA), slime mold externalized memory, SOPHIA division of labor, Möbius RoPE topology intervention. Don't grind away at the original layer — switch layers.
3. OS Design Patterns Migrating to AI
Sixty years of virtual memory design prove that decoupling logical address space from physical memory is the right abstraction for scarce memory. vToken carries it into KV cache management.
4. Policy Adapters: 500+ Lines → <50
Previously, integrating an eviction policy into vLLM required touching 4-6 files and 500+ lines due to tight coupling. vToken isolates shared logic; a policy only implements SelectVictims. New-policy experimentation gets dramatically cheaper — just like virtual memory freed application developers from physical memory layout.
An Unsolved Problem
vToken honestly states its limits:
Deeper issue: vToken is currently a vLLM-only implementation. TensorRT-LLM, SGLang, and other engines would each need to implement the hooks, and no public code is released (the authors are from NUDT; release may be restricted), limiting immediate community adoption.
Closing Thought: Virtualization as a Universal Design Pattern
AI systems are re-walking the path of computer systems:
Next to virtualize? My bet: attention computation itself — letting attention kernels access KV through indirection, so sparse attention, sliding windows, and dynamic routing can share one runtime without each rewriting CUDA kernels.
vToken's lesson: when two layers' granularities mismatch, don't grind at either layer — add a virtualization layer.
---
Paper: vToken: Token-Level Virtualization for Reclaimable KV Caches Authors: Yuanhang Gao, Xiangrui Yang, Yuanfeng Chen, Hongjia Chen, Qianru Lv, Wenfei Wu, Dongsheng Li Institutions: National University of Defense Technology; Peking University Implemented on vLLM v0.18.0; no public repository yet