Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages
> *"A language dying is like a library burning."* — George Steiner
This post discusses a 2026 paper by Avijit Roy and Proma Roy (*arXiv:2608.12278*) arguing that modern AI systems systematically exclude Bengali speakers—nearly 4% of the world's population—through layered, institutionalized inequality the authors call Structural Silence.
Key Points
1. Web presence gap
- Bengali is spoken by ~1 in 25 people globally, yet accounts for less than 0.5% of web content (Pimienta, 2024; W3Techs, 2026) — a two-orders-of-magnitude gap.
- English comprises ~49.5% of web content with only ~4.7% native speakers. Causes include colonial legacy, English-centric technical infrastructure, unequal content economies, and lagging Bengali digitization tools — despite Bengali's 1,400-year literary tradition and 2024 recognition as a classical language in India.
- UNESCO's International Mother Language Day (Feb 21) originates from the 1952 Bengali Language Movement.
- Major multilingual corpora allocate roughly 30 billion tokens to Bengali (Sangraha) versus ~2 trillion for English (Common Corpus) — about 67:1.
- Chang et al. (2024) found some large multilingual models perform worse than simple bigram baselines on low-resource language perplexity benchmarks.
- A self-reinforcing cycle persists: little data → poor models → low usage → low investment → even less data. The authors stress this scarcity is the result of institutional choices (funding, publication incentives, market priorities), not accident.
- Bengali's alphasyllabary (matras, yuktakshar conjuncts, multi-character grapheme clusters) violates the assumptions of BPE/WordPiece tokenizers optimized for Latin scripts.
- The result is high token fertility: the same semantic content requires far more tokens in Bengali, raising training and inference costs, shrinking effective context windows, and degrading attention. Even with equal data, this structural inefficiency demands disproportionately more data to compensate.
- Bangladesh Bureau of Statistics (2024–25): rural individual internet penetration is 36.5% vs. 71.4% urban; only ~9% rural computer ownership/use.
- Bangladesh's mobile data supplementary tax rose from 3% (FY2016) to 23% (2024), making cloud-first AI tools functionally inaccessible for low-income rural users.
- Citing cognitive load theory (Sweller et al., 2011), the paper argues second-language technical content overloads working memory: bilingual learners underperform native speakers (Roussel et al., 2017), and non-mother-tongue instruction reduces retention and transfer (Soosai Raj et al., 2018). Mother-tongue instruction is a prerequisite for equitable AI education, not a feature.
- Roy, A., & Roy, P. (2026). *Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages*. arXiv:2608.12278.
- Pimienta, D. (2024). *Web Presence and Language Diversity*.
- W3Techs. (2026). *Usage Statistics of Content Languages for Websites*.
- Khan et al. (2024). *Sangraha: A Large-Scale Multilingual Corpus*.
- Langlais et al. (2025). *Common Corpus: Trillion-Token Dataset for Language Modeling*.
- Chang et al. (2024). *Low-Resource Language Performance in Multilingual Models*.
- Sweller et al. (2011). *Cognitive Load Theory*.
- Roussel et al. (2017); Soosai Raj et al. (2018).
2. Training data poverty (67:1)
3. Tokenization as a hidden tax
4. Connectivity exclusion
The cognitive double burden
Proposed Solutions
1. Reframe dataset scarcity as structural, treating corpus building, translation validation, and evaluation benchmarks as first-class research contributions equal to architectural innovation. 2. Offline-first as equity infrastructure: evaluate educational AI under bandwidth, device, and cost constraints; develop lightweight on-device models (which also save energy at institutional scale). 3. Linguistics at the core: apply sociolinguistics, language typology, and critical language research to expose the monolingual-English assumptions embedded in tokenizers and benchmarks.
Conclusion
The paper's central claim: AI language inequality cannot be fixed by better models alone — current performance distributions reflect accumulated, non-neutral decisions about data, evaluation, and deployment.