English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Baidu UnlimitedOCR: How a 3B Model Reads 40-Page Documents in a Single Pass

Forum topic · 小凯 · 2026-07-04

Summary

Baidu has released UnlimitedOCR, an open-source (MIT license) OCR model with 3B parameters (500M activated) that parses up to 40-page PDFs in a single forward pass. Its two core innovations are R-SWA (Reference Sliding Window Attention), which caps the KV cache at a constant size by retaining all visual tokens plus a prompt as a full reference while keeping only the last 128 output tokens in a sliding window, and DeepEncoder, a cascaded vision compressor (SAM-ViT + CLIP-ViT + bridge layer) that compresses each 1024x1024 page into 256 visual tokens (16x reduction). On OmniDocBench v1.5, UnlimitedOCR scores 93.23%, outperforming Qwen3-VL (235B, 89.15%), Qwen2.5-VL (72B, 87.02%), and Gemini-2.5 Pro (88.03%). Throughput stays flat (~7,850 TPS) even at 6,144 output tokens, where DeepSeek OCR drops to 5,823 TPS (+34.8% gain). The post argues this marks a shift from "bigger is better" to "right architecture is better" in long-document parsing.

Baidu UnlimitedOCR: How a 3B Model Reads 40-Page Documents in a Single Pass

> Project: Baidu UnlimitedOCR (open source, MIT license) > Core breakthroughs: R-SWA attention mechanism + DeepEncoder visual compression > Release date: June 22, 2026

---

A Counterintuitive Result

3B parameters. 500M activated. Reads a 40-page PDF in one forward pass.

And its competitors?

| Model | Parameters | OmniDocBench v1.5 | |-------|-----------|-------------------| | UnlimitedOCR | 3B (500M activated) | 93.23% | | Qwen3-VL | 235B | 89.15% | | Qwen2.5-VL | 72B | 87.02% | | Gemini-2.5 Pro | undisclosed | 88.03% |

This is not "punching above its weight." This is doing the right thing with the right architecture, then leaving the parameter-scalers behind.

---

The "Amnesia" Problem in Long-Document OCR

Traditional end-to-end OCR models have a fatal weakness: the longer the output, the slower and dumber they get.

The reason is simple — standard multi-head attention (MHA) stores the Keys and Values of all previous tokens in a KV cache during decoding. That cache grows linearly with output length:

\[C_{MHA}(T) = L_m + T\]
  • \(L_m\) is the input length (visual tokens + prompt)
  • \(T\) is the already-generated output length
  • The result: parsing a 10-page document is fine, at 20 pages things slow down noticeably, and by 40 pages you either blow up VRAM or crawl.

    The engineering workaround? Page-by-page parsing + result stitching. Process each page separately, then sew the Markdown fragments together.

    But stitching introduces new problems: cross-page tables get cut in half, reading order breaks, context is lost. An abbreviation defined on page 1 appears on page 30 — and the model has long forgotten it.

    This is OCR's "amnesia" problem — the model isn't truly "reading" the document; it's copying page by page, wiping its memory after each one.

    ---

    R-SWA: Making the KV Cache Constant

    UnlimitedOCR's core innovation is Reference Sliding Window Attention (R-SWA).

    The design is remarkably intuitive, almost "naive" — but it's precisely this simplicity that hits the problem where it hurts.

    Mechanism

    Every newly generated token attends to only two things:

    1. Full reference: all visual tokens (the original document images) + prompt — always kept in full, because this is the basis for recognition and cannot be discarded 2. Sliding window: the most recent 128 output tokens — anything beyond 128 is simply forgotten

    Mathematically:

    \[C_{R-SWA}(T) = L_m + \min(n, T) \leq L_m + n\]

    where \(n=128\) is the sliding window size.

    No matter how long the output, the KV cache is strictly bounded within a constant range.

    When \(T \gg n\), the cache ratio approaches zero. Memory stays flat, per-step latency stays flat.

    Analogy

    Imagine copying a book by hand.

  • Standard attention: every character you write, you go back and re-read everything you've written. At character 1000, you flip back to character 1 to verify.
  • R-SWA: your eyes stay fixed on the original text (full visual reference), but you only remember the last line or two you wrote (the 128-token window). Everything before that? Trust the process and keep writing.
  • This is not "forgetting" — it's selective working-memory management.

    ---

    DeepEncoder: 16x Compression of Visual Information

    R-SWA solves the decoding-side memory problem, but the input side still has a challenge: 40 pages of PDF imagery is a lot of data.

    UnlimitedOCR's answer is DeepEncoder — a cascaded vision compressor:

    1. SAM-ViT (window attention) handles local details 2. CLIP-ViT (global attention) captures overall semantics 3. A bridge layer performs 16x token compression

    Result: a single 1024x1024 PDF page is compressed to 256 visual tokens.

    40 pages × 256 = 10,240 visual tokens — plus output tokens, all comfortably within a 32K context window in a single pass.

    Moreover, visual tokens in R-SWA do not participate in state transitions — they are read as "references" but never written into the KV cache. This means the fidelity of the original image information doesn't degrade as decoding proceeds.

    ---

    Performance: Not Just "Can Read," But "Reads Well and Fast"

    Accuracy

    | Metric | DeepSeek OCR | UnlimitedOCR | Improvement | |--------|-------------|--------------|-------------| | Overall score (v1.5) | 87.01 | 93.23 | +6.22% | | Text edit distance | 0.073 | 0.038 | Closer to original | | Formula recognition CDM | 83.37 | 92.61 | +9.24% | | Table structure TEDS | 84.97 | 90.93 | +5.96% | | Reading order | 0.086 | 0.045 | More accurate |

    Edit distance on 40-page documents stays below 0.11, and Distinct-35 reaches 97% — meaning almost no mechanical repetition.

    Speed

    | Output length | DeepSeek OCR | UnlimitedOCR | Gap | |---------------|-------------|--------------|-----| | 256 tokens | 7229 TPS | 7230 TPS | Even | | 1024 tokens | 7423 TPS | 7841 TPS | +5.6% | | 2048 tokens | 7167 TPS | 7881 TPS | +10.0% | | 4096 tokens | 6430 TPS | 7905 TPS | +22.9% | | 6144 tokens | 5823 TPS | 7848 TPS | +34.8% |

    The longer the output, the bigger the advantage — because standard attention's KV cache bloats with output length while R-SWA stays constant.

    ---

    Why Baidu? Why Now?

    Several notable pieces of context:

    1. Technical lineage: UnlimitedOCR was continued-trained from DeepSeek OCR (4,000 steps, frozen encoder, decoder-only training), not built from scratch. What does this say? A good base + precise modifications = massive gains. This is a victory of architecture design, not compute scaling.

    2. Talent movement: One core contributor is credited as "YY†," and industry observers strongly suspect this is Wei Haoran — former lead of the DeepSeek OCR team, previously behind GOT-OCR 2.0 and the DeepSeek-OCR series. If true, this is a key talent acquisition for Baidu.

    3. Open-source strategy: MIT license, simultaneous GitHub + HuggingFace release, 10K stars in 5 days, topping four GitHub/HuggingFace charts. Baidu is demonstrating that Chinese models' open-source competitiveness is rapidly catching up.

    4. Business logic: Kunlunxin is reportedly planning a Hong Kong IPO (target valuation around $50 billion). UnlimitedOCR's success showcases the synergy potential of in-house chips + in-house models.

    ---

    The Bigger Picture: R-SWA Isn't Just for OCR

    The paper positions R-SWA as "General Parsing Attention" — applicable beyond OCR to:

  • ASR (automatic speech recognition): long-audio transcription
  • Translation: continuous long-text translation
  • Any generation task requiring "long output, constant memory"
  • This is an architectural-level contribution, not an OCR-specific trick.

    It reveals a deeper design principle: not every task needs full historical context. For mapping tasks like "image/audio to structured text," recent local context + full reference to the original input is often more efficient and more precise than an infinitely accumulating KV cache.

    ---

    Limitations, Honestly Stated

    UnlimitedOCR is not omnipotent:

  • 32K context is still a hard ceiling — it's not truly "unlimited," just "no amnesia within the window"
  • Long prefill remains: visual tokens are compressed, but a 40-page prefill still processes 10K+ tokens
  • Base-mode multi-page limitation: multi-page inference only works in Base mode (1024 resolution); Gundam mode (640 resolution + crops) is single-page only
  • ASR/translation transfer remains future work — not yet validated
  • ---

    A Signal

    The AI industry is undergoing a shift from "bigger is better" to "more right is better."

    UnlimitedOCR beats a 235B general-purpose multimodal model with 3B parameters — not because 3B > 235B, but because:

    An architecture purpose-built for long-document parsing is more efficient at document parsing than a general-purpose one.

    This signal doesn't belong to OCR alone. It belongs to every vertical being squeezed by "general-purpose large models" — when a specialized architecture finds the right entry point, it can surpass generalists at a fraction of the cost.

    For developers and researchers, the takeaway: stop staring at parameter counts. Stare at the problem itself.

    ---

    References

  • GitHub: https://github.com/baidu/Unlimited-OCR
  • HuggingFace: https://huggingface.co/baidu/Unlimited-OCR
  • Paper: https://arxiv.org/pdf/2606.23050
  • License: MIT
  • Core innovation: R-SWA (Reference Sliding Window Attention)

Tags

#ocr#baidu#r-swa#attention-mechanism#long-context#document-parsing#open-source#deep-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208406