As tools like ChatGPT and other generative AI systems rapidly reshape the way people work, learn, and communicate, most of their power is concentrated in a few key languages such as English, Chinese, and French. But with more than 7,000 languages spoken around the world, billions of people are left with limited or inaccurate AI support in their language.
Dr Slangika Ranatunga, a senior lecturer in the School of Mathematical and Computing Sciences at Te Kunenga ki Purefuroa Massey University, is working to change this. Originally from Sri Lanka, Dr. Ranatunga has first-hand experience of the challenges that speakers of low-resource languages face in the digital world.
“AI models often reflect Western ideology, but Western ideology does not reflect all languages and cultures. The challenge is how to make AI inclusive so that it can support other languages and, through that, other cultures as well,” explains Dr. Ranatunga.
Dr. Ranatunga works with a team of researchers to solve the problem through four related approaches. One is to raise awareness through opinion pieces highlighting digital imbalances. Build essential datasets for training AI systems in underrepresented languages. Develop new technologies specifically designed for low-resource languages. Evaluate existing AI systems for various tasks in the context of low-resource languages.
In addition to building datasets from scratch, her team is exploring synthetic data generation techniques such as web mining, alongside optical character recognition (OCR), which can help unlock information from undigitized printed materials.
“AI is nothing without data, and when we want to build tools for these languages, the data often simply doesn’t exist,” Dr. Ranatunga says.
Without intervention, she warns, the AI revolution could deepen global inequality. As more of our lives move to digital systems, people will naturally migrate to languages that are better supported by AI.
“Over time, this can have negative effects and lead to the decline of undervalued languages. Education is a prime example. Even today, most tutoring systems and learning materials are designed for English. If similar tools existed for other languages, access to education would increase significantly.”
AI also struggles with tasks that require cultural context, such as generating appropriate mathematical word problems.
“We’ve seen some absurd examples, such as the question about someone traveling from the UK to Sri Lanka and bringing Ceylon tea as a souvenir. This shows that the model doesn’t understand the local context, and while the questions may be mathematically correct, they are often culturally inappropriate for students in those countries.”
One of the biggest challenges, she explains, is that AI tools are often considered successful if they work well in English, even if they don’t work in many other languages.
“In English, spelling correction is considered to be solved to a large extent because tools such as Microsoft Word’s spell check handle spelling correction. However, many languages do not have spelling correction systems at all. Until it is solved in all languages, we cannot say that the problem is solved.”
Dr. Ranatunga also serves on the AI Advisory Council established by the Government of Sri Lanka, helping guide the national strategy on indigenous language technologies. One of her key goals is to develop a machine translation system for the languages spoken in Sri Lanka: Sinhala, Tamil, and English.
“Currently, we primarily rely on tools like Google Translate, which perform poorly in languages with fewer resources. This also relates to the question of AI sovereignty: our data is sent to systems hosted overseas and we have no control over it.”
Her vision is for countries to build and host their own AI systems that serve their own languages and communities.
In the coming months, Dr. Ranatunga and his team plan to release one of the largest datasets of its kind for Sinhala-Tamil-English machine translation, containing over 100,000 parallel sentences. The project, originally supported by a Diversity and Inclusion grant from Google, is being implemented over three years and is expected to have an immediate impact across Sri Lanka.
“The Sri Lankan government is currently looking to build a machine translation system for local languages.Rather than starting from scratch, they can build directly on the models and datasets we have built,” says Dr. Ranatunga.
The dataset and machine translation models will also enable future research in areas such as adversarial robustness, error handling, and cross-language model improvement.
Because many languages are used in developing countries, low-resource language research is often underfunded, so Dr. Ranatunga collaborates with researchers in countries such as Sri Lanka, the United Kingdom, India, Albania, and Pakistan.
“This is participatory research. We work together to build datasets, evaluate models, and introduce new systems. This is very time-consuming and lacks sufficient funding in many regions, so collaboration is essential,” she says.
Dr. Ranatunga is a strong advocate of open research and hopes more researchers and AI professionals will follow suit.
“Sharing what we build, from datasets to code to models, is so important because that’s how the field grows. If everyone keeps their work private, we won’t make progress. Languages with fewer resources have problems, especially when resources aren’t shared. We need to build together.”
