English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

Forum topic · 小凯 · 2026-08-13

Summary

This forum post introduces and summarizes the paper 'Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages' by Avijit Roy and Proma Roy, which examines why Bengali—spoken by nearly 4% of the world's population—remains nearly invisible in AI systems. The authors document a 'structural silence' operating on four levels: Bengali accounts for under 0.5% of web content despite its speaker base; training data ratios reach roughly 67:1 in favor of English (30 billion vs. 2 trillion tokens); tokenizers optimized for Latin scripts fragment Bengali's alphasyllabary inefficiently, inflating costs via high token fertility; and connectivity gaps in rural Bangladesh (36.5% rural internet penetration) plus a 23% mobile data tax make cloud-first AI education functionally inaccessible. Drawing on cognitive load theory, the paper argues that mother-tongue instruction is a prerequisite for equitable education, not a localization feature. Proposed remedies include reframing dataset scarcity as an institutional choice, treating offline-first deployment as equity-oriented infrastructure, and giving linguists a central role in system design.

Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

> *"A language dying is like a library burning."* — George Steiner

This post discusses a 2026 paper by Avijit Roy and Proma Roy (*arXiv:2608.12278*) arguing that modern AI systems systematically exclude Bengali speakers—nearly 4% of the world's population—through layered, institutionalized inequality the authors call Structural Silence.

Key Points

1. Web presence gap

  • Bengali is spoken by ~1 in 25 people globally, yet accounts for less than 0.5% of web content (Pimienta, 2024; W3Techs, 2026) — a two-orders-of-magnitude gap.
  • English comprises ~49.5% of web content with only ~4.7% native speakers. Causes include colonial legacy, English-centric technical infrastructure, unequal content economies, and lagging Bengali digitization tools — despite Bengali's 1,400-year literary tradition and 2024 recognition as a classical language in India.
  • UNESCO's International Mother Language Day (Feb 21) originates from the 1952 Bengali Language Movement.
  • 2. Training data poverty (67:1)

  • Major multilingual corpora allocate roughly 30 billion tokens to Bengali (Sangraha) versus ~2 trillion for English (Common Corpus) — about 67:1.
  • Chang et al. (2024) found some large multilingual models perform worse than simple bigram baselines on low-resource language perplexity benchmarks.
  • A self-reinforcing cycle persists: little data → poor models → low usage → low investment → even less data. The authors stress this scarcity is the result of institutional choices (funding, publication incentives, market priorities), not accident.
  • 3. Tokenization as a hidden tax

  • Bengali's alphasyllabary (matras, yuktakshar conjuncts, multi-character grapheme clusters) violates the assumptions of BPE/WordPiece tokenizers optimized for Latin scripts.
  • The result is high token fertility: the same semantic content requires far more tokens in Bengali, raising training and inference costs, shrinking effective context windows, and degrading attention. Even with equal data, this structural inefficiency demands disproportionately more data to compensate.
  • 4. Connectivity exclusion

  • Bangladesh Bureau of Statistics (2024–25): rural individual internet penetration is 36.5% vs. 71.4% urban; only ~9% rural computer ownership/use.
  • Bangladesh's mobile data supplementary tax rose from 3% (FY2016) to 23% (2024), making cloud-first AI tools functionally inaccessible for low-income rural users.
  • The cognitive double burden

  • Citing cognitive load theory (Sweller et al., 2011), the paper argues second-language technical content overloads working memory: bilingual learners underperform native speakers (Roussel et al., 2017), and non-mother-tongue instruction reduces retention and transfer (Soosai Raj et al., 2018). Mother-tongue instruction is a prerequisite for equitable AI education, not a feature.
  • Proposed Solutions

    1. Reframe dataset scarcity as structural, treating corpus building, translation validation, and evaluation benchmarks as first-class research contributions equal to architectural innovation. 2. Offline-first as equity infrastructure: evaluate educational AI under bandwidth, device, and cost constraints; develop lightweight on-device models (which also save energy at institutional scale). 3. Linguistics at the core: apply sociolinguistics, language typology, and critical language research to expose the monolingual-English assumptions embedded in tokenizers and benchmarks.

    Conclusion

    The paper's central claim: AI language inequality cannot be fixed by better models alone — current performance distributions reflect accumulated, non-neutral decisions about data, evaluation, and deployment.

    References

  • Roy, A., & Roy, P. (2026). *Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages*. arXiv:2608.12278.
  • Pimienta, D. (2024). *Web Presence and Language Diversity*.
  • W3Techs. (2026). *Usage Statistics of Content Languages for Websites*.
  • Khan et al. (2024). *Sangraha: A Large-Scale Multilingual Corpus*.
  • Langlais et al. (2025). *Common Corpus: Trillion-Token Dataset for Language Modeling*.
  • Chang et al. (2024). *Low-Resource Language Performance in Multilingual Models*.
  • Sweller et al. (2011). *Cognitive Load Theory*.
  • Roussel et al. (2017); Soosai Raj et al. (2018).

Tags

#ai-fairness#low-resource-languages#bengali#tokenization#data-poverty#offline-first#cognitive-load#digital-divide

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633442