Overview
Research Area: CV Authors: Pouria Mahdi, Haq Nawaz Malik arXiv: 2507.18395
Abstract
Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages despite Persian being spoken by more than 110 million people across multiple countries. This gap arises from two fundamental challenges: the intrinsic complexity of the Perso-Arabic writing system and the limited availability of large-scale, high-quality annotated datasets. Persian script exhibits obligatory cursive connectivity, context-dependent glyph shaping, extensive ligatures, diacritic placement, and stylistic variation across writing forms such as Naskh and Nastaliq, all of which significantly complicate text recognition. At the same time, the high cost and labor-intensive nature of manual annotation have created a persistent data bottleneck, limiting the development of robust OCR systems and slowing progress in Persian document digitization.
This paper introduces Persian Pixel, a comprehensive synthetic OCR dataset designed to address these challenges. Key features:
- Scale: Over 343,000 high-fidelity image-text pairs covering sentences, paragraphs, and full-page document layouts.
- Corpus: Generated from a curated 7-million-word Persian corpus using the SynthOCR-Gen rendering framework.
- Typography fidelity: The generation pipeline faithfully models Persian script features, including contextual character connectivity, positional glyph variants, diacritic placement, and multiple representative Persian fonts.
- Realism: To bridge the synthetic-to-real domain gap, rendered images are enriched with over 25 random degradation models simulating real document acquisition artifacts, including ink bleed-through, paper aging, blur, illumination variation, scanner defects, compression artifacts, and various noise processes.
Significance
By overcoming the long-standing scarcity of annotated Persian OCR data, Persian Pixel provides a scalable and publicly available resource for training and fine-tuning modern OCR architectures, including transformer-based models such as TrOCR and Donut. It establishes a solid foundation for Persian document analysis, historical manuscript digitization, and end-to-end document understanding research, while demonstrating that programmatic synthetic data generation is a practical, cost-effective, and scalable alternative to manual annotation for advancing OCR in low-resource and typographically complex scripts.
---
*Auto-collected on 2026-07-24*