English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Persian Pixel: A Large-Scale Synthetic OCR Dataset for Persian Language

Forum topic · 小凯 · 2026-07-24

Summary

Persian Pixel is a comprehensive synthetic OCR dataset designed to address the persistent data scarcity holding back Persian optical character recognition. Although Persian is spoken by over 110 million people, its Perso-Arabic script—with obligatory cursive connectivity, context-dependent glyph shaping, ligatures, diacritic placement, and stylistic variants like Naskh and Nastaliq—remains far less supported than Latin-script languages. The dataset contains more than 343,000 high-fidelity image-text pairs covering sentences, paragraphs, and full-page document layouts, rendered from a curated 7-million-word Persian corpus using the SynthOCR-Gen framework. The generation pipeline faithfully models Persian typography, including positional glyph variants and contextual character joining. To bridge the synthetic-to-real domain gap, rendered images are augmented with over 25 random degradation models simulating real-world artifacts such as ink bleed, paper aging, blur, illumination changes, scanner defects, and compression noise. Persian Pixel provides a scalable, publicly available resource for training and fine-tuning modern OCR architectures, including transformer-based models like TrOCR and Donut, and supports Persian document analysis, historical manuscript digitization, and end-to-end document understanding. It demonstrates that programmatic synthetic data generation is a practical, cost-effective alternative to manual annotation for low-resource, typographically complex scripts. Paper: arXiv 2507.18395.

Overview

Research Area: CV Authors: Pouria Mahdi, Haq Nawaz Malik arXiv: 2507.18395

Abstract

Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages despite Persian being spoken by more than 110 million people across multiple countries. This gap arises from two fundamental challenges: the intrinsic complexity of the Perso-Arabic writing system and the limited availability of large-scale, high-quality annotated datasets. Persian script exhibits obligatory cursive connectivity, context-dependent glyph shaping, extensive ligatures, diacritic placement, and stylistic variation across writing forms such as Naskh and Nastaliq, all of which significantly complicate text recognition. At the same time, the high cost and labor-intensive nature of manual annotation have created a persistent data bottleneck, limiting the development of robust OCR systems and slowing progress in Persian document digitization.

This paper introduces Persian Pixel, a comprehensive synthetic OCR dataset designed to address these challenges. Key features:

  • Scale: Over 343,000 high-fidelity image-text pairs covering sentences, paragraphs, and full-page document layouts.
  • Corpus: Generated from a curated 7-million-word Persian corpus using the SynthOCR-Gen rendering framework.
  • Typography fidelity: The generation pipeline faithfully models Persian script features, including contextual character connectivity, positional glyph variants, diacritic placement, and multiple representative Persian fonts.
  • Realism: To bridge the synthetic-to-real domain gap, rendered images are enriched with over 25 random degradation models simulating real document acquisition artifacts, including ink bleed-through, paper aging, blur, illumination variation, scanner defects, compression artifacts, and various noise processes.

Significance

By overcoming the long-standing scarcity of annotated Persian OCR data, Persian Pixel provides a scalable and publicly available resource for training and fine-tuning modern OCR architectures, including transformer-based models such as TrOCR and Donut. It establishes a solid foundation for Persian document analysis, historical manuscript digitization, and end-to-end document understanding research, while demonstrating that programmatic synthetic data generation is a practical, cost-effective, and scalable alternative to manual annotation for advancing OCR in low-resource and typographically complex scripts.

---

*Auto-collected on 2026-07-24*

Tags

#ocr#persian#synthetic-data#computer-vision#dataset#arxiv#deep-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178447055