SmolDocling: An Ultra-Compact Vision-Language Model for End-to-End Multi-Modal Document Conversion
This post introduces an academic paper entry: SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion (March 2025, arXiv).
Key information
| Field | Content | |---|---| | Title | SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion | | Authors / Affiliations | Ahmed Nassar, Andres Marafioti, Matteo Omenetti, Maksym Lysak, Nikolaos Livathinos, Christoph Auer, et al. (13 authors total) | | Published | March 2025 | | Source | https://arxiv.org/abs/2503.11576 | | Type | Academic paper | | Category | Document understanding |
One-line summary
SmolDocling is an ultra-compact vision-language model that performs end-to-end conversion of multi-modal documents into structured, machine-readable markup.
Original abstract (preserved)
> SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion, Mar 2025, arxiv
Context and relevance
The paper is positioned in the document understanding area, addressing how to replace conventional multi-stage document conversion pipelines (OCR, layout detection, table and equation recognition) with a single compact vision-language model. Related entries in the collection include:
- LongDA: Benchmarking LLM Agents for Long-Document Data Analysis
- Qwen2.5-VL Technical Report (Section 3.3.2: Document Understanding and OCR)
- ColPali: Efficient Document Retrieval with Vision Language Models (https://arxiv.org/abs/2407.01449)
- Full quantitative results should be verified against the original PDF on arXiv before being cited.
- The work is cross-indexed with other entries along the retrieval → ranking → generation/agent → evaluation chain in this collection.
- Original paper: SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion, arXiv, March 2025. https://arxiv.org/abs/2503.11576