AI models tested for Arabic
Does it actually work for Saudi, Najdi, Hijazi or Gulf Arabic? Benchmarks rarely say. The community tests these models on real speech and text — and reports back.
Lahjawi
Misraj AI
Cross-dialect Arabic translation system from Misraj in Khobar, fine-tuned from Kuwain-1.5B in two variants: Lahjawi-D2D translates between 15 Arabic dialects and Lahjawi-D2MSA converts any dialect into MSA. The paper names Egyptian, Emirati, Jordanian, Palestinian, Levantine and Maghrebi among the dialects it evaluates but does not enumerate all 15. No weights are published.
Audar ASR V1 Turbo
Audar AI Labs
2.35B Arabic-first speech recognition trained on 300k+ hours, naming Gulf, Egyptian, Levantine and Maghrebi coverage plus Arabic-English code-switching. Its card claims the top place on the Open Universal Arabic ASR Leaderboard — a leaderboard run by Elm, a Saudi company. Custom AudarAI Community Licence, not open source.
Qwen3
Alibaba
Open LLM family. The only model in this directory whose developer explicitly names a Saudi dialect: Qwen documents 119 languages and dialects including Arabic (Standard, Najdi, Levantine, Egyptian and others). Apache-2.0.
Munsit
CNTXT AI
Arabic ASR trained with weak supervision on 30K+ hours; the paper claims best-in-class accuracy across 18 dialects. No weights are published — there is no CNTXT organisation or Munsit repository on Hugging Face — and the model is commercial API-only, so the dialect claims cannot be independently checked.
CAMeLBERT
CAMeL Lab, NYU Abu Dhabi
BERT models pre-trained per Arabic variant (MSA, dialectal, classical) for NER, POS, sentiment and dialect identification. The mix checkpoint remains the most-used entry point despite dating from 2021.
Fanar-2-27B
QCRI / HBKU
Arabic-centric flagship of the Fanar 2.0 release (Mar 2026), continually pretrained from google/gemma-3-27b-pt on ~166B Arabic, English and code tokens with 32K context. It adds native Arabic reasoning traces, selective thinking mode and tool calling. This repo is text-in/text-out; image generation, image understanding and poetry are separate Fanar-2 models.
Fanar-1-9B
QCRI / HBKU
Qatar's sovereign Arabic LLM. This 8.7B instruct model — the 'Prime' branch — continually pretrains google/gemma-2-9b on 1T Arabic and English tokens; a separate 7B 'Star' model was trained from scratch. The card claims MSA plus Gulf, Levantine and Egyptian dialects, and alignment with Islamic values and Arab culture.