English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Structural Silence: How AI Infrastructure Fails Bengali Speakers

Forum topic · 小凯 · 2026-08-13

Summary

A 2026 paper by Avijit Roy and Proma Roy, 'Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages' (arXiv:2608.12278), argues that AI language inequality toward Bengali speakers—nearly 4% of the world's population—is structural, not incidental. Bengali accounts for under 0.5% of web content while English holds roughly 49.5%; training corpora show a 67:1 English-to-Bengali token ratio (Common Corpus ~2 trillion vs. Sangraha ~30 billion tokens). Tokenization of Bengali's alphasyllabary (matras, conjuncts, multi-character grapheme clusters) imposes higher token fertility, raising training and inference costs and degrading quality. Connectivity gaps compound the problem: only 36.5% internet penetration in rural Bangladesh versus 71.4% in cities, plus mobile data taxes rising from 3% to 23%. The authors propose reframing dataset scarcity as an institutional choice, adopting offline-first deployment as an equity strategy, and involving linguists in AI system design.

Overview

The forum post discusses the paper *Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages* (Roy, A., & Roy, P., 2026, arXiv:2608.12278), which examines how AI systems systematically fail Bengali speakers—roughly 4% of the global population—despite English native speakers comprising a similar share.

Key Points

  • Web representation gap: Bengali makes up less than 0.5% of global web content versus ~49.5% for English (Pimienta, 2024; W3Techs, 2026)—a gap of orders of magnitude rooted in colonial legacies, English-centric technical infrastructure, and unequal content economies.
  • Training data disparity: The English-to-Bengali training token ratio is approximately 67:1. The Sangraha corpus allocates ~30 billion tokens for Bengali, while Common Corpus contains ~2 trillion English tokens (Khan et al., 2024; Langlais et al., 2025). Per Chang et al. (2024), large multilingual models can underperform simple bigram baselines on low-resource language perplexity benchmarks.
  • Tokenization tax: Bengali's alphasyllabary—matras (vowel diacritics), conjuncts (yuktakshar), and multi-character grapheme clusters—defies tokenizers optimized for Latin scripts. Higher token fertility increases training cost, inference latency, and context-window consumption, degrading semantic units and performance even at data parity.
  • Connectivity exclusion: Bangladesh Bureau of Statistics (2024-25) reports rural internet penetration of 36.5% vs. 71.4% urban; only 9.2% of rural households own computers. Mobile data surcharges rose from 3% (FY2016) to 23% (2024), making cloud-first AI tools functionally inaccessible offline.
  • Cognitive double burden: Citing cognitive load theory (Sweller et al., 2011; Roussel et al., 2017; Soosai Raj et al., 2018), the paper argues that learning technical content in a non-native language overloads working memory—mother-tongue instruction is a prerequisite for equitable access, not a localization feature.
  • Proposed Solutions

    1. Reframe data scarcity as an institutional choice, not a natural state; treat corpus building, translation validation, and benchmark development as first-class research contributions equal to architecture innovation. 2. Offline-first deployment as an equity-oriented infrastructure strategy—evaluating systems under bandwidth, device, and cost constraints rather than ideal connectivity; locally quantized models also reduce energy use at institutional scale. 3. Linguists as core contributors to expose embedded monolingual-English assumptions in tokenizers, evaluation benchmarks, and interface design.

    Conclusion

    The paper's core claim is that current AI performance distributions reflect accumulated, non-neutral decisions about data, evaluation, and deployment—not gaps technology will naturally close.

    References

  • Roy, A., & Roy, P. (2026). *Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages*. arXiv:2608.12278.
  • Pimienta, D. (2024). *Web Presence and Language Diversity*.
  • W3Techs. (2026). *Usage Statistics of Content Languages for Websites*.
  • Khan et al. (2024). *Sangraha: A Large-Scale Multilingual Corpus*.
  • Langlais et al. (2025). *Common Corpus: Trillion-Token Dataset for Language Modeling*.
  • Chang et al. (2024). *Low-Resource Language Performance in Multilingual Models*.
  • Sweller et al. (2011). *Cognitive Load Theory*.
  • Bangladesh Bureau of Statistics (2024-25). *Digital Access Survey*.

Tags

#ai-fairness#low-resource-languages#bengali#nlp#tokenization#digital-divide#multilingual-ai#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633439