The language barrier is also a knowledge barrier

The issue is bigger than translating a scientific paper from English into another language. Scientific communication depends on terminology. A concept used in medicine, chemistry, agriculture or computer science needs to have an appropriate and understandable expression in the target language. Where those terms have not yet been established, even sophisticated translation systems can struggle.

The researchers therefore approach translation as both a linguistic and technological challenge. Their objective is to develop scientific terminology and translation resources, create a multilingual scientific corpus, and evaluate the ability of machine translation systems and large language models to work with scientific content across African languages.

This is particularly relevant in the context of AI. Machine translation systems require sufficient language and domain-specific data to learn how specialised concepts should be expressed. Yet such resources remain limited across many African languages.

Building the resources AI needs

AfriScience-MT takes a multi-stage approach combining scientific content preparation, professional translation, terminology development and AI evaluation. The researchers developed a dataset from 230 scientific papers across 11 disciplines.

  • Paper Selection: 230 scientific papers across 11 disciplines.
  • Expert Summarization: Complex scientific research was simplified into accessible summaries.
  • Translation & Terminology: The content was translated into six African languages, with scientific terms developed where necessary.
  • Quality Control: Translations were reviewed and 7,605 sentences were aligned.
  • AI Benchmarking: 11 large language models were evaluated for scientific translation.

The six languages covered are Amharic, Hausa, Luganda, Northern Sotho, Yorùbá and isiZulu. Professional translators worked alongside science communicators and created new terminology where established terms did not exist.

Better data can matter more than bigger models

One of the most important findings is that specialised data matters more than simply increasing model size. The study found that scientific, domain-specific data can improve translation performance, while generic datasets may dilute specialised knowledge and reduce performance.

This finding is particularly interesting for African AI development. The assumption that larger models will automatically solve low-resource language challenges does not fully hold here. The research instead points towards the importance of investing in high-quality, specialised datasets.

The results also show that smaller open models can become competitive when appropriately adapted. Fine-tuned NLLB-1.3B achieved a sentence-level COMET score of 67.3, compared with 68.3 for GPT-5.4 and 68.0 for Gemini-3.1-Flash-Lite. At the document level, GPT-5.4 and Gemini-3.1-Flash-Lite both achieved 48.3, while TranslateGemma-12B reached 44.0 with one-shot in-context learning.

For African organisations working with limited resources, this is an important signal. It suggests that locally deployable open models, combined with the right data and adaptation, can have a meaningful role in specialised AI applications.

Scientific terminology remains a major challenge

The study also makes clear that there is no simple technological solution to the problem.

Technical terminology was the main source of translation errors, with incorrect translations accounting for approximately 77–80% of document-level accuracy errors. Models also performed better when translating African languages into English than when translating English into African languages.

This matters because it shifts the conversation from simply asking whether AI can translate African languages to asking whether we have built the linguistic resources required for AI to understand and communicate specialised knowledge accurately.

Scientific glossaries, terminology databases, multilingual corpora and validated language resources could therefore become important infrastructure for African-language AI. Their value would extend beyond translation to education, research, digital public services and other knowledge-intensive applications.

Language should be part of the AI infrastructure conversation

Much of the discussion about AI infrastructure in Africa focuses on compute, data centres, cloud platforms, datasets and models. These are essential, but AfriScience-MT highlights another layer that deserves greater attention: language resources.

An AI system can have access to sophisticated models and significant computing power, but its usefulness remains limited if it cannot accurately understand and communicate specialised knowledge in the languages used by its intended communities.

In this sense, language is not simply a user interface problem. Language resources are part of the infrastructure that determines who can participate in the AI and knowledge economy.

From translation to scientific sovereignty

This is where the research becomes particularly interesting from an African perspective.

Making scientific knowledge available in African languages can improve access, but the longer-term ambition could be much broader. If African languages are mainly used for everyday communication while scientific and technical knowledge continues to be produced and exchanged predominantly in foreign languages, a divide remains between knowledge production and knowledge access.

Developing scientific terminology and resources in African languages can help narrow that divide. But perhaps the bigger opportunity is to move beyond translating knowledge produced elsewhere and create conditions in which African languages can also become languages of scientific production, teaching and discussion.

That is where the idea of decolonising science takes on a deeper meaning. It is not only about translating scientific information. It is about expanding who can access knowledge, who can participate in scientific conversations and whose languages are represented in the technologies shaping the future of knowledge.

The work is only beginning

The researchers are clear about the limitations of the study. AfriScience-MT covers only six of Africa's more than 2,000 languages and 11 scientific disciplines. Some disciplines are represented by considerably smaller datasets than others. The evaluation also relied mainly on automated metrics and LLM-based assessment, without systematic validation by native bilingual domain experts. In addition, only a limited selection of available AI models was tested, which limits how broadly the findings can be generalised.

These limitations do not diminish the contribution of the research. Instead, they illustrate how much work remains to be done.

More languages require more datasets. More scientific fields require more terminology. And more reliable evaluation requires greater involvement from African linguists, researchers and domain experts.

The bigger question

AfriScience-MT demonstrates that improving scientific translation is not simply a matter of selecting a more powerful AI model. It requires high-quality data, scientific terminology, linguistic expertise and sustained investment in language resources.

For Africa, this raises a broader question: Are we building AI systems that merely operate in Africa, or are we building the linguistic and knowledge infrastructure that allows African communities to shape how AI understands and communicates knowledge?

The distinction matters.

Because making African research accessible in African languages is not only about translation. It is ultimately about who can access knowledge, who can participate in science and whose languages are represented in the technologies shaping the future.

About the Research

Title: AfriScience-MT: Towards Decolonizing Science in Africa Through Text Translation
Authors: Idris Abdulmumin, Tajuddeen Gwadabe, Shamsuddeen Hassan Muhammad, David Ifeoluwa Adelani, Nomonde Khalo, Ibrahim Said Ahmad, Abiodun Modupe, Anina Mumm, Sibusiso Biyela, Michelle Rabie, Johanna Havemann, Marek Rei, Jade Abbott and Vukosi Marivate.
Published: 28 May 2026
arXiv: 2605.29741v1 / Read the full paper on arXiv