spot_img
HomeResearch & DevelopmentBridging the Linguistic Divide: New Dataset Advances NLP for...

Bridging the Linguistic Divide: New Dataset Advances NLP for Nigeria’s Minority Languages

TLDR: The research paper introduces IBOM, a novel dataset for machine translation and topic classification in four under-resourced Nigerian languages: Anaang, Efik, Ibibio, and Oro. These languages, spoken in Akwa Ibom State, are not covered by major NLP benchmarks or tools like Google Translate. The study evaluates both fine-tuned models and Large Language Models (LLMs), finding that LLMs perform poorly in zero-shot machine translation but show significant improvement with few-shot prompting for topic classification, especially with Gemini models. Fine-tuned models, particularly those using a two-stage approach or African-centric encoders, generally offer stronger performance for these low-resource languages. The work highlights the critical need for more data and better evaluation metrics for minority languages.

Nigeria, a nation celebrated for its immense linguistic diversity with over 500 languages, faces a significant challenge in the realm of Natural Language Processing (NLP). While major languages like Hausa, Igbo, Nigerian-Pidgin, and Yorùbá have received some attention, the vast majority of its minority languages remain underserved, largely due to a scarcity of textual data for training NLP algorithms.

A recent research paper, Ibom NLP: A Step Toward Inclusive Natural Language Processing for Nigeria’s Minority Languages, addresses this critical gap by introducing IBOM, a groundbreaking dataset designed for machine translation and topic classification. This initiative focuses on four Coastal Nigerian languages from the Akwa Ibom State region: Anaang, Efik, Ibibio, and Oro. These languages are notably absent from widely used platforms like Google Translate and major benchmarks such as Flores-200 or SIB-200.

The IBOM Dataset: A Foundation for Inclusivity

The IBOM dataset is a significant step towards making NLP more inclusive for Nigeria’s minority languages. It comprises two main components: IBOM-MT for machine translation and IBOM-TC for topic classification. The creation of IBOM-MT involved translating sentences from the English Flores-200 dataset into Ibibio, Efik, Anaang, and Oro. This is particularly noteworthy as it marks the first parallel language resource for Anaang and Oro. IBOM-TC was developed by aligning these translated texts with topic labels from the SIB-200 classification dataset.

The data collection process was meticulous, involving three linguists for each language, all native speakers residing in Akwa Ibom State and holding at least a Bachelor’s degree in Linguistics. A lead translator for each language reviewed and corrected translations, ensuring high quality over a four-month period of collaborative effort.

Understanding the Ibom Languages

The languages targeted by this research—Ibibio, Efik, Anaang, and Oro—are predominantly spoken in the Akwa Ibom State and parts of Cross River State in Nigeria. Ibibio serves as the lingua franca of Akwa Ibom, with approximately 3.7 million speakers. Efik has about 3.5 million speakers and is also found in Southwest Cameroon. Anaang, with around 1.4 million speakers, is concentrated in the North West part of Akwa Ibom. Oro has over 400,000 speakers in specific Local Government Areas of the state.

Linguistically, these languages belong to the Lower Cross branch of the Cross River Division of the Benue-Congo family. They share similar tonal systems and are largely agglutinating, with flexible Subject-Verb-Object (SVO) sentence structures.

Evaluating Performance: LLMs vs. Fine-Tuned Models

The researchers conducted extensive evaluations using both fine-tuned baseline models and Large Language Model (LLM) prompting for machine translation and topic classification. Proprietary LLMs like GPT-4.1, o4-mini, Gemini 2.0 Flash, and Gemini 2.5 Flash were tested alongside fine-tuned multilingual models such as M2M-100 and NLLB-200, and African-centric encoders like AfroXLMR-61L.

Machine Translation Findings:

The evaluation revealed that current LLMs generally perform poorly on machine translation for these languages in zero-shot settings, especially when translating from English to the Ibom languages. However, a two-stage fine-tuning approach, leveraging existing religious parallel data for Efik, significantly boosted performance for Efik and related languages like Anaang and Ibibio. Few-shot prompting showed promise, particularly with Gemini models, where 5-shot and 10-shot settings improved performance, though 20-shots sometimes led to a drop in quality for translation into English.

Human evaluation further highlighted the limitations of automatic metrics, with human annotators often preferring the output of the M2M-100 2-stage fine-tuned model over Gemini 2.5 Flash for Anaang and Efik, despite automatic scores suggesting otherwise. This underscores the need for better evaluation metrics tailored to low-resource languages.

Topic Classification Findings:

For topic classification, African-centric encoder models like AfroXLMR-61L demonstrated superior performance compared to other multilingual encoders. Fine-tuned baselines generally outperformed LLMs in zero-shot settings. However, few-shot prompting proved highly effective for Gemini LLMs, with performance steadily improving with more shots (5, 10, and 20 shots). This suggests that Gemini models might possess stronger multilingual capabilities for these low-resource languages compared to some other LLMs.

Also Read:

Looking Ahead

This research marks a crucial step in promoting inclusive NLP for Nigeria’s minority languages. By creating and releasing the IBOM dataset, the authors hope to encourage further investment and research beyond the most widely spoken African languages. Future work includes expanding the training data for IBOM-MT and developing more robust COMET evaluation support for these languages, paving the way for more equitable language technology development.

Karthik Mehta
Karthik Mehtahttps://blogs.edgentiq.com
Karthik Mehta is a data journalist known for his data-rich, insightful coverage of AI news and developments. Armed with a degree in Data Science from IIT Bombay and years of newsroom experience, Karthik merges storytelling with metrics to surface deeper narratives in AI-related events. His writing cuts through hype, revealing the real-world impact of Generative AI on industries, policy, and society. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -