Summary
PaddleOCR is an Apache 2.0 open-source OCR toolkit from PaddlePaddle with 84.6K GitHub stars, widely integrated into Dify, RAGFlow, Cherry Studio, and OmniParser as the default bridge from PDFs, images, and scans into LLM-ready JSON or Markdown. The 3.7.0 release (June 2026) ships three product lines: PP-OCRv6, a 50-language scene-text recognition model in tiny, small, and medium tiers (1.5M to 34.5M parameters) with 5.2x OpenVINO CPU acceleration and accuracy gains of 4.6% detection and 5.1% recognition over v5, reportedly beating Qwen3-VL-235B and GPT-5.5 on text tasks; PaddleOCR-VL-1.6, a 0.9B-parameter vision-language model that reaches 96.3% on OmniDocBench v1.6, combining a NaViT dynamic-resolution vision encoder with ERNIE-4.5-0.3B; and PP-StructureV3 for fine-grained layout, table-cell coordinates, and Office-to-Markdown conversion. The project also offers MCP, LangChain, browser, mobile, and ONNX/TensorRT/OpenVINO deployment paths, used by 6,500+ downstream projects for RAG ingestion, visual agents, and on-device document AI.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178208375