Baidu UnlimitedOCR: How a 3B Model Reads 40-Page Documents in a Single Pass
> Project: Baidu UnlimitedOCR (open source, MIT license) > Core breakthroughs: R-SWA attention mechanism + DeepEncoder visual compression > Release date: June 22, 2026
---
A Counterintuitive Result
3B parameters. 500M activated. Reads a 40-page PDF in one forward pass.
And its competitors?
| Model | Parameters | OmniDocBench v1.5 | |-------|-----------|-------------------| | UnlimitedOCR | 3B (500M activated) | 93.23% | | Qwen3-VL | 235B | 89.15% | | Qwen2.5-VL | 72B | 87.02% | | Gemini-2.5 Pro | undisclosed | 88.03% |
This is not "punching above its weight." This is doing the right thing with the right architecture, then leaving the parameter-scalers behind.
---
The "Amnesia" Problem in Long-Document OCR
Traditional end-to-end OCR models have a fatal weakness: the longer the output, the slower and dumber they get.
The reason is simple — standard multi-head attention (MHA) stores the Keys and Values of all previous tokens in a KV cache during decoding. That cache grows linearly with output length:
- \(L_m\) is the input length (visual tokens + prompt)
- \(T\) is the already-generated output length
- Standard attention: every character you write, you go back and re-read everything you've written. At character 1000, you flip back to character 1 to verify.
- R-SWA: your eyes stay fixed on the original text (full visual reference), but you only remember the last line or two you wrote (the 128-token window). Everything before that? Trust the process and keep writing.
- ASR (automatic speech recognition): long-audio transcription
- Translation: continuous long-text translation
- Any generation task requiring "long output, constant memory"
- 32K context is still a hard ceiling — it's not truly "unlimited," just "no amnesia within the window"
- Long prefill remains: visual tokens are compressed, but a 40-page prefill still processes 10K+ tokens
- Base-mode multi-page limitation: multi-page inference only works in Base mode (1024 resolution); Gundam mode (640 resolution + crops) is single-page only
- ASR/translation transfer remains future work — not yet validated
- GitHub: https://github.com/baidu/Unlimited-OCR
- HuggingFace: https://huggingface.co/baidu/Unlimited-OCR
- Paper: https://arxiv.org/pdf/2606.23050
- License: MIT
- Core innovation: R-SWA (Reference Sliding Window Attention)
The result: parsing a 10-page document is fine, at 20 pages things slow down noticeably, and by 40 pages you either blow up VRAM or crawl.
The engineering workaround? Page-by-page parsing + result stitching. Process each page separately, then sew the Markdown fragments together.
But stitching introduces new problems: cross-page tables get cut in half, reading order breaks, context is lost. An abbreviation defined on page 1 appears on page 30 — and the model has long forgotten it.
This is OCR's "amnesia" problem — the model isn't truly "reading" the document; it's copying page by page, wiping its memory after each one.
---
R-SWA: Making the KV Cache Constant
UnlimitedOCR's core innovation is Reference Sliding Window Attention (R-SWA).
The design is remarkably intuitive, almost "naive" — but it's precisely this simplicity that hits the problem where it hurts.
Mechanism
Every newly generated token attends to only two things:
1. Full reference: all visual tokens (the original document images) + prompt — always kept in full, because this is the basis for recognition and cannot be discarded 2. Sliding window: the most recent 128 output tokens — anything beyond 128 is simply forgotten
Mathematically:
where \(n=128\) is the sliding window size.
No matter how long the output, the KV cache is strictly bounded within a constant range.
When \(T \gg n\), the cache ratio approaches zero. Memory stays flat, per-step latency stays flat.
Analogy
Imagine copying a book by hand.
This is not "forgetting" — it's selective working-memory management.
---
DeepEncoder: 16x Compression of Visual Information
R-SWA solves the decoding-side memory problem, but the input side still has a challenge: 40 pages of PDF imagery is a lot of data.
UnlimitedOCR's answer is DeepEncoder — a cascaded vision compressor:
1. SAM-ViT (window attention) handles local details 2. CLIP-ViT (global attention) captures overall semantics 3. A bridge layer performs 16x token compression
Result: a single 1024x1024 PDF page is compressed to 256 visual tokens.
40 pages × 256 = 10,240 visual tokens — plus output tokens, all comfortably within a 32K context window in a single pass.
Moreover, visual tokens in R-SWA do not participate in state transitions — they are read as "references" but never written into the KV cache. This means the fidelity of the original image information doesn't degrade as decoding proceeds.
---
Performance: Not Just "Can Read," But "Reads Well and Fast"
Accuracy
| Metric | DeepSeek OCR | UnlimitedOCR | Improvement | |--------|-------------|--------------|-------------| | Overall score (v1.5) | 87.01 | 93.23 | +6.22% | | Text edit distance | 0.073 | 0.038 | Closer to original | | Formula recognition CDM | 83.37 | 92.61 | +9.24% | | Table structure TEDS | 84.97 | 90.93 | +5.96% | | Reading order | 0.086 | 0.045 | More accurate |
Edit distance on 40-page documents stays below 0.11, and Distinct-35 reaches 97% — meaning almost no mechanical repetition.
Speed
| Output length | DeepSeek OCR | UnlimitedOCR | Gap | |---------------|-------------|--------------|-----| | 256 tokens | 7229 TPS | 7230 TPS | Even | | 1024 tokens | 7423 TPS | 7841 TPS | +5.6% | | 2048 tokens | 7167 TPS | 7881 TPS | +10.0% | | 4096 tokens | 6430 TPS | 7905 TPS | +22.9% | | 6144 tokens | 5823 TPS | 7848 TPS | +34.8% |
The longer the output, the bigger the advantage — because standard attention's KV cache bloats with output length while R-SWA stays constant.
---
Why Baidu? Why Now?
Several notable pieces of context:
1. Technical lineage: UnlimitedOCR was continued-trained from DeepSeek OCR (4,000 steps, frozen encoder, decoder-only training), not built from scratch. What does this say? A good base + precise modifications = massive gains. This is a victory of architecture design, not compute scaling.
2. Talent movement: One core contributor is credited as "YY†," and industry observers strongly suspect this is Wei Haoran — former lead of the DeepSeek OCR team, previously behind GOT-OCR 2.0 and the DeepSeek-OCR series. If true, this is a key talent acquisition for Baidu.
3. Open-source strategy: MIT license, simultaneous GitHub + HuggingFace release, 10K stars in 5 days, topping four GitHub/HuggingFace charts. Baidu is demonstrating that Chinese models' open-source competitiveness is rapidly catching up.
4. Business logic: Kunlunxin is reportedly planning a Hong Kong IPO (target valuation around $50 billion). UnlimitedOCR's success showcases the synergy potential of in-house chips + in-house models.
---
The Bigger Picture: R-SWA Isn't Just for OCR
The paper positions R-SWA as "General Parsing Attention" — applicable beyond OCR to:
This is an architectural-level contribution, not an OCR-specific trick.
It reveals a deeper design principle: not every task needs full historical context. For mapping tasks like "image/audio to structured text," recent local context + full reference to the original input is often more efficient and more precise than an infinitely accumulating KV cache.
---
Limitations, Honestly Stated
UnlimitedOCR is not omnipotent:
---
A Signal
The AI industry is undergoing a shift from "bigger is better" to "more right is better."
UnlimitedOCR beats a 235B general-purpose multimodal model with 3B parameters — not because 3B > 235B, but because:
An architecture purpose-built for long-document parsing is more efficient at document parsing than a general-purpose one.
This signal doesn't belong to OCR alone. It belongs to every vertical being squeezed by "general-purpose large models" — when a specialized architecture finds the right entry point, it can surpass generalists at a fraction of the cost.
For developers and researchers, the takeaway: stop staring at parameter counts. Stare at the problem itself.
---