English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Soro: A Lightweight Foundation Model and Chatbot for Tajik

Forum topic · 小凯 · 2026-05-29

Summary

Researchers introduce Soro, a family of Tajik-specialized conversational large language models designed for real-world deployment under Tajikistan's tight compute and connectivity constraints. Starting from open-weight Gemma 3 checkpoints, the team performed Tajik-only continual pretraining on a curated 1.9-billion-token corpus spanning filtered web text, PDF documents, and curriculum-aligned educational materials, followed by supervised instruction tuning on 40K Tajik teacher-style examples. Because standard benchmarks offer limited Tajik coverage, the authors built a suite of Tajik benchmarks covering general knowledge, linguistic competence, and school and university entrance-exam domains, released on Hugging Face. Experiments show Soro substantially outperforms same-size Gemma 3 baselines on these Tajik benchmarks while retaining strong English performance. FP8 and INT4 quantized variants preserve most Tajik capability, supporting pilot deployments in Tajikistan's education sector and scaled rollout in schools. The paper is available on arXiv as 2605.27379.

Paper Overview

Field: NLP Authors: Stanislav Liashkov, Haitz Sáez de Ocáriz Borde, Azizjon Azimi, et al. Published: 2026-05-28 arXiv: 2605.27379

What Soro Is

Soro is a family of Tajik-specialized conversational large language models (LLMs) designed for real-world deployment under tight compute and connectivity constraints in Tajikistan.

Key points

  • Base models: Built from open-weight Gemma 3 checkpoints.
  • Continual pretraining: Tajik-only training on a curated 1.9-billion-token corpus spanning filtered web text, PDF documents, and curriculum-aligned educational materials.
  • Instruction tuning: Supervised fine-tuning on 40K Tajik teacher-style examples.
  • New benchmarks: The authors introduce a suite of Tajik benchmarks covering general knowledge, linguistic competence, and school- and university entrance-exam domains, open-sourced on Hugging Face.
  • Results: Soro substantially outperforms same-size Gemma 3 baselines across the Tajik benchmarks while retaining strong English performance.
  • Efficiency: FP8 and INT4 quantized variants preserve most of the models' Tajik capability, enabling pilot deployments in Tajikistan's education sector and scaled rollout in schools.

Original Abstract (excerpt)

> We present Soro, a family of Tajik-specialized conversational large language models (LLMs) designed for real-world deployment under tight compute and connectivity constraints in Tajikistan. Starting from open-weight Gemma 3 checkpoints, we perform Tajik-only continual pretraining on a curated 1.9-billion-token corpus spanning filtered web text, PDF documents, and curriculum-aligned educational materials, followed by supervised instruction tuning on 40K Tajik teacher-style examples. To enable rigorous evaluation despite the limited coverage of Tajik in standard benchmarks, we introduce a suite of Tajik benchmarks covering general knowledge, linguistic competence, and school- and university entrance-exam domains, and we open-source them on Hugging Face. Across these Tajik benchmarks, Soro substantially outperforms same-size Gemma 3 baselines while retaining strong English performance on st...

--- *Auto-collected on 2026-05-29*

Tags

#nlp#llm#tajik#gemma-3#low-resource-languages#benchmark#quantization#education

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980502