Developing a Mehri–Arabic Neural Machine Translation System Using a Multilingual Transformer Model (mT5)
DOI:
https://doi.org/10.59992/IJCI.2026.v5n10p1Keywords:
Mehri, Machine Translation, MT5 model, Arabic Translations, Resourced LanguagesAbstract
Machine translation for low-resource and endangered languages remains a challenging task due to the scarcity of parallel corpora and limited computational resources. Despite its linguistic and cultural significance, the Mehri language has received little attention in natural language processing research. This paper presents a Mehri–Arabic machine translation system based on a multilingual neural model. A custom Mehri–Arabic parallel corpus was constructed and preprocessed using rule-based techniques. The multilingual mT5 model was fine-tuned via transfer learning and evaluated using both automatic metrics and human assessment. Experimental results indicate that satisfactory Arabic translations can be achieved despite limited data availability. This work represents an early computational effort in Mehri– Arabic machine translation and highlights the potential of multilingual pretrained models for supporting under-resourced languages.
References
[1] D. Bahdanau, K. Cho and Y. Bengio, "Neural machine translation by jointly learning to align and translate," in ICLR, 2016.
[2] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser and I. Polosukhin, "Attention is all you need," in 31st Conference on Neural Information Processing Systems, CA, USA., 2017.
[3] A. Rubin, The Mehri Language of Oman, Brill, 2010.
[4] J. C. Watson, The structure of Mehri, Otto Harrassowitz, 2012.
[5] Y. Liu, J. Gu, N. Goyal, X. Li, S. Edunov, M. Ghazvininejad, M. Lewis and L. Zettlemoyer, "Multilingual denoising pre-training for neural machine translation," Transactions of the Association for Computational Linguistics, pp. 726-742, 8 2020.
[6] L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua and C. Raffel, "mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer," in Proceedings of the 2021 conference of the North American chapter of the association for computational linguistics: Human language technologies, 2021.
[7] M. Junczys-Dowmunt, R. Grundkiewicz, . T. Dwojak, H. Hoang, K. Heafield, T. Neckermann, F. Seide, U. Germann, A. F. Aji and N. Bogoychev, "Marian: Fast Neural Machine Translation in C++," 2018.
[8] P. Koehn, Statistical machine translation, Cambridge University Press, 2009.
[9] P. Singh and S. Kumar, "AdiBhashaa: A Community-Curated Benchmark for Machine Translation into Indian Tribal Languages," arXiv preprint arXiv, 2025.
[10] M. Johnson and et al, "Google’s multilingual neural machine translation system: Enabling zero-shot translation," Transactions of the Association for Computational Linguistics, pp. 339-351, 5 2017.
[11] S. M. M. Billah, , A. A. Subarna, S. N. Sarna, . A. S. Wasit, A. Fariha, A. Sushmit and A. Y. Sadeque, "Towards santali linguistic inclusion: building the first Santali-to-English translation model using mT5 transformer and data augmentation," arXiv preprint arXiv, 2024.
[12] R. Sennrich and. B. Zhang, "Revisiting low-resource neural machine translation: A case study," in 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 2019.
[13] A. Chronopoulou, D. Stojanovski and A. Fraser, "Language-family adapters for low-resource multilingual neural machine translation," in Proceedings of the Sixth Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT 2023), 2023.
[14] S. Bird, "Decolonising speech and language technology," in Proceedings of the 28th international conference on computational linguistics, 2020.
[15] P. Joshi, S. Santy, A. Budhiraja, K. Bali and M. Choudhury, "The state and fate of linguistic diversity and inclusion in the NLP world," in ACL, 2020.
[16] . T. M. Johnstone, Mehri lexicon, Routledge, 2012.
[17] K. Balhaf, O. Darwish, E. Rawashdeh, M. A. Awad, D. Darweesh, Y. Tashtoush and S. Rawashdeh, "Classifying Arabian Gulf Tweets to Detect People's Trends: A case study," in 2022 Ninth International Conference on Social Networks Analysis, Management and Security (SNAMS), 2022.
[18] H. K. eklehaymanot, G. Gidey and W. Nejdl, "Low-Resource EnglishTigrinya MT: Leveraging Multilingual Models, Custom Tokenizers, and Clean Evaluation Benchmark," 2025.
[19] V. Mujadia and D. M. Sharma, "BhashaVerse: Translation Ecosystem for Indian Subcontinent Languages," arXiv preprint arXiv, 2024.
[20] I. Sel and D. Hanbay, "Efficient Adaptation: Enhancing Multilingual Models for Low-Resource Language Translation," Mathematics, 19 12 2024.