Ask an AI system for tips on how to shoplift without getting caught in English, and it will refuse to help. If the prompt is made in a different language, you may well get a response to assist you in committing a theft.
Current AI systems are less accessible, less useful, and less safe for users of so-called “low-resource languages.” These are languages such as Swahili and Burmese that may well be spoken by many, but for which little digitized data is available.
Market forces have not mobilized the necessary investment to address these shortcomings. Fierce economic and geopolitical competition funnels attention and resources into the development of a narrow set of frontier models, which are optimized for a small set of dominant and well-resourced languages.
This imbalance results in a global inequity that warrants policymakers’ attention. The longer the gap exists, the wider it will grow. If AI accelerates socio-economic development, it is urgent to address this imbalance now. Access to frontier AI capacity is already limited by high cost and a lack of infrastructure in low-income countries. And even for those with technical access, language can be a constraint.
Although models can process prompts in different languages, their multilingual capacity fails to remove all access barriers. Researchers point to a so-called “token tax”; access to models is usually billed by use of tokens, the units into which natural language is split for processing. Some languages require more tokens to represent the same content compared to English. This gap drives up cost and latency.
Models generally reason more effectively in English than in other languages, studies show. Bias creeps in — for instance when a multilingual model associates the word “dove” in all languages with “peace” even though in Basque it can be an insult.
This language gap also represents a safety risk. Researchers have shown that it is possible to jailbreak AI systems — that is, to breach their safeguards — by machine-translating prompts into other languages. Studies have also demonstrated that translating malicious input into a low-resource language generates more unsafe content than sticking to English.
The weak performance of models in low-resource languages is not because they are less suitable for AI than English is. Model design choices and the limited availability of training datasets are responsible. Until now, market forces have not directed resources to address this challenge, and they are unlikely to do so.
Policymakers must correct course. They should:
- Boost multilingual capacity on par with efforts to bolster “sovereign” AI. Many countries and regions have started investing in local models to increase ownership and control over AI systems. Models optimized for local languages would make it faster and cheaper to train on local data and context.
- Target R&D investment to develop technical solutions for improving multilingual models. Simply calling for large investment in thousands of underserved languages is unrealistic, but policymakers need to articulate clear principles or criteria to decide which investment in which language would generate the most benefit for the most people. Involving the language communities themselves to co-design and co-create model development should be standard.
- Leverage access to and processing of local datasets, especially when mediated through official institutions such as archives, public broadcasters, or cultural institutions. Public procurement processes could require specific investment by model developers in multilingual capacity.
It’s important to tackle the systemic disadvantage of low-resource languages. Otherwise, much of the world risks missing out on the promise of the AI revolution.
Christian Schlaepfer is a former Swiss diplomat and negotiator of tech and AI policy at the United Nations. He is a guest at the Institute for Logic, Language and Computation at the University of Amsterdam and policy advisor at the think tank Starling Institute.
Bandwidth is CEPA’s online journal dedicated to advancing transatlantic cooperation on tech policy. All opinions expressed on Bandwidth are those of the author alone and may not represent those of the institutions they represent or the Center for European Policy Analysis. CEPA maintains a strict intellectual independence policy across all its projects and publications.
Tech 2030
A Roadmap for Europe-US Tech Cooperation