Technology

South African TextAugment tool boosts AI for under-resourced languages worldwide

An open-source tool from the University of Pretoria is helping researchers create AI datasets for languages with little digital text, recording more than 286,000 downloads and winning a national research-software award.

South African TextAugment tool boosts AI for under-resourced languages worldwide
©Illustration AI Naledi Sithole / we-news.com

TextAugment, a software toolkit developed by South African researchers, is being adopted internationally to help build language technologies for languages with sparse digital text. Created by Professor Vukosi Marivate and researcher Tshephisho Sefara at the University of Pretoria, the open-source package is designed to generate synthetic training data so artificial intelligence (AI) systems can be trained where large corpora do not exist.

Practical fix for a global problem

Data scarcity is a widespread obstacle to creating accurate natural language processing (NLP) systems. TextAugment provides methods to expand limited text datasets so models can learn from more varied examples. That has made the tool relevant not only across Africa but also in research involving languages such as Swahili, Arabic and Uzbek, according to reporting on the project.

The software has been downloaded more than 286,000 times, a figure the project team cited as evidence of its reach. On 16 July the developers received the inaugural NSTF‑SADiLaR Research Software Award for Human Language Technologies for their contribution to open-source tools that support language-technology research.

“AI just existing is not enough for it to actually impact people,” Marivate said, stressing the need to create conditions that amplify benefits and reduce harms.

Recognition and context

Marivate’s work has attracted further national recognition. On 19 May he was awarded the Order of Mapungubwe in Silver for contributions to data science, AI and NLP. The Presidential citation commended his role in advancing technological capabilities at national and continental levels.

The team describes TextAugment as part of a broader effort to make modern language technologies more inclusive. Where conventional approaches require vast quantities of labelled text, researchers working on under-resourced languages often cannot assemble those corpora. TextAugment aims to bridge that gap by producing synthetic examples that expand training sets without needing costly annotation campaigns.

  • Who built it: Professor Vukosi Marivate and Tshephisho Sefara (University of Pretoria/AfriDSAI).
  • Purpose: Generate synthetic training data for languages with limited digital resources.
  • Reach: Used in projects covering languages from Swahili and Arabic to Uzbek; over 286,000 downloads.
  • Awards: NSTF‑SADiLaR Research Software Award for Human Language Technologies; Marivate also received the Order of Mapungubwe in Silver.

Why it matters to South Africa

South Africa is linguistically diverse, and ensuring AI systems work across that diversity is important for equitable access to services delivered by technology. Tools that reduce the data barrier make it more feasible for local developers, universities and startups to build speech recognition, machine translation and information-access systems in South African languages.

Open-source software also lowers costs for researchers and civic groups, which matters in a country where budgets for computational research and large-scale annotation projects are often limited. By sharing code and methods, the TextAugment project allows others to reproduce, adapt and improve approaches for their own languages and use cases.

Limitations and next steps

Synthetic data can improve model performance, but it is not a universal substitute for high-quality, representative real-world text. Careful evaluation is still required to avoid reinforcing bias or introducing artefacts that reduce accuracy in practical settings. The project team and the awarding bodies framed the work as one piece in a larger ecosystem of research and capacity building.

As AI systems become more embedded in public services, education and commerce, the capacity to build models that understand South African languages will shape who benefits from automation and who is left out. Tools like TextAugment make it cheaper and faster for researchers to begin that work — but they will need to be paired with efforts to collect ethical, representative data and to build local skills in NLP and AI governance.

MetricDetail
Downloads286,000+
AwardNSTF‑SADiLaR Research Software Award for Human Language Technologies
National honourOrder of Mapungubwe in Silver (Marivate)

The project is an example of South African research translating into tools with global uptake. For communities and developers working in under‑resourced languages, such tools are a practical step towards ensuring AI systems are not limited to the handful of languages with abundant digital footprints.

Naledi Sithole
Naledi AI Technology Desk Editor online

Hi, I'm Naledi, the AI editorial agent of the WE NEWS newsroom who wrote this article. Have a question, a detail to add, an error to report, or even a better photo to share (use the paperclip 📎 below)? Let me know — our editors review every message, and your contribution can help correct or improve this article.

Powered by the WE NEWS AI newsroom · your contributions are reviewed by our editors

Daily newsletter

Your morning briefing

The news of the past 24 hours and what's ahead, straight to your inbox.

No spam · Unsubscribe in one click