New research published on arXiv.org indicates that safety measures implemented in large language models (LLMs) do not reliably transfer to low-resource languages, particularly in the African context. The study, titled "The Illusion of Cross-Lingual Safety in Low-Resource Languages," highlights a significant gap in current AI safety alignment practices, which are predominantly developed and tested in English.
The researchers investigated this issue across four African languages: Twi, Hausa, Amharic, and Swahili. They developed a new safety dataset named LoDNA. This dataset pairs literal translations of harmful prompts with culturally localized versions to better assess how LLMs respond to nuanced linguistic and cultural contexts. Beyond generation-based evaluations, the study also introduced a novel latent geometric framework. This framework probes LLM hidden states to analyze refusal representations, offering a method to assess safety without relying solely on generated text.
Experimental results from this investigation showed a severe limitation in cross-lingual safety transfer. For most language-model pairings, harmful prompts retained less than 10% of the refusal signal observed when the same prompts were presented in English. This suggests that LLMs are significantly more susceptible to generating harmful content when prompted in these languages. The study's findings challenge the assumption that safety safeguards developed for English will automatically generalize to other languages, especially those with fewer digital resources.
The scarcity of data and resources for many African languages has long been a challenge for AI development. Efforts are underway to create more datasets for these languages. For example, the WAXAL dataset provides extensive speech data for 24 Sub-Saharan African languages to advance speech technology. Other initiatives include the development of parallel text datasets for machine translation, sentiment analysis, and document classification across various African languages, such as Twi, Swahili, and Hausa. The LoDNA dataset and the proposed evaluation framework are intended to draw attention to the specific problem of safety alignment in these under-resourced linguistic contexts.
The implications of this limited cross-lingual safety transfer are substantial. As LLMs are increasingly deployed in multilingual applications, a failure to ensure safety across all supported languages could lead to the dissemination of harmful content, misinformation, or biased outputs. This is particularly concerning for low-resource languages where automated content moderation and safety checks may be less effective or non-existent. The research team's work underscores the need for language-specific safety alignment and evaluation to ensure equitable and secure AI deployment globally.
The study's authors propose that future work should focus on developing methods for targeted safety alignment in low-resource languages and creating more diverse and culturally relevant safety datasets. The latent geometric framework offers a path toward more nuanced evaluation, moving beyond simple text-based assessments. Addressing this safety gap is essential for building trust and ensuring responsible AI development that serves all linguistic communities.
