Artificial intelligence tools, increasingly positioned as solutions for educational and language support in under-resourced communities, are built upon infrastructure that inherently disadvantages speakers of less represented languages. Research from Avijit Roy and Proma Roy, detailed in their arXiv paper "Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages," highlights these systemic barriers. The paper uses Bengali, a language with approximately 285 million speakers globally, as a case study to illustrate these structural failures, particularly in the context of AI-assisted education in low-connectivity regions.
The study identifies four interconnected issues. First, a significant web presence gap exists for Bengali. Despite nearly 4% of the global population speaking Bengali, the language accounts for less than 0.5% of global web content. This imbalance is critical because AI training corpora are primarily built from web crawls, meaning languages with limited online presence are inherently underrepresented in the data used to train large language models.
Second, Bengali faces a substantial training token deficit. The research indicates a 67:1 ratio in training tokens between English and Bengali within major multilingual corpora. This scarcity of quality data for non-English languages is a primary reason why models like ChatGPT and Gemini perform well for English speakers but underperform for others. Languages with limited machine-readable data are often termed "low-resource," a category that includes languages with millions of speakers but insufficient digitized information for effective AI training.
Third, Bengali's alphasyllabary script introduces a "tokenization penalty." This characteristic requires a higher rate of token fertility, further exacerbating the existing data deficit. Tokenization is the process of breaking down text into smaller units that AI models can process, and the nature of a language's script can affect how efficiently this process occurs.
Finally, the study points to a "connectivity exclusion," which renders cloud-dependent AI tools inaccessible to the rural populations in Bangladesh who could most benefit from them. Internet penetration in rural areas of Bangladesh stands at 36.5%, significantly lower than the 71.4% in urban areas. This digital divide means that even if AI tools were developed for Bengali, a large segment of the population would struggle to access them.
These structural failures are not isolated technical glitches but rather consequences of historical resource allocation, institutional priorities, and design defaults that did not prioritize certain languages in mainstream AI development. The majority of online content is in English or a few other dominant languages, leading to AI models that inherit this linguistic skew. This creates a global language data gap, where many communities are excluded from the benefits of AI.
The implications extend beyond inconvenience. Entire cultures and communities risk being left out of the AI revolution, facing potential harm from AI-generated misinformation and bias, and losing educational and economic opportunities available to speakers of high-resource languages. While efforts are underway to address this gap, such as initiatives to develop AI for African languages and projects like Cohere Lab's Aya, which covers 101 languages, more concentrated work is needed to support multilingual research and dataset creation.
The challenges for Bengali AI development are compounded by its complex grammar and phonetic system, which make creating accurate and culturally relevant models difficult. However, this complexity also presents opportunities for innovation, as seen in efforts by organizations like Bengali.AI to crowdsource data and host competitions to advance Bengali language processing. Addressing the linguistic diversity gap in AI requires a shift in focus towards developing models that can effectively serve all languages, ensuring equitable access to advancing technology.
