Domain Adaptation for Bilingual Arabic-English Neural Machine Translation (NMT)
DOI:
https://doi.org/10.26713/cma.v17i2.3662Abstract
The importance of Arabic machine translation (AMT) cannot be overstated as it seeks to enhance communication
between Arabic and non-Arabic speakers. Given that there are an estimated 300 million speakers globally, Arabic is
essential for cultural exchanges, global commerce, and diplomatic relations. Nevertheless, what creates substantial
obstacles for the Arabic language with regards to machine translation (MT) is the fact that it is categorized as a
language with low resources as well as one with complex and rich morphological features. Neural machine
translation (NMT) has emerged as the pinnacle of MT methodologies. Although there exist numerous NMT tools and
models designed for translation of Arabic text, the translation quality is often lacking, particularly with texts that fall
out-of-domain. A primary technical obstacle associated with AMT stems from the scarcity of bilingual datasets for
such texts that are out-of-domain. Consequently, this research seeks to enhance AMT across various domains. To
achieve this, we designed a multi-domain parallel corpus specifically for Arabic-English (AR-EN) translation, which
was later utilized to train the proposed NMT model. The model uses the encoder–decoder architecture that employs
long short-term memory (LSTM) networks, which are themselves integrated with attention mechanisms. The
experimental results show that this proposed model improves accuracy of translation and reduces the loss. To
evaluate the efficacy of our proposed NMT model we compared its performance against both OPUS MT system and
Google Translate across four domains using BLEU as a metric for assessment. In translations from Arabic to English,
our model consistently outperformed both OPUS MT and Google Translate across every domain
assessed—averaging +18.66 BLEU points above OPUS MT while exceeding Google Translate by approximately
+18.11 BLEU points overall. Similarly consistent superiority was observed in the English to Arabic direction where
average gains reached +22.39 points over OPUS MT and +15.36 points over Google Translate.




