AI models tested for Arabic
Does it actually work for Saudi, Najdi, Hijazi or Gulf Arabic? Benchmarks rarely say. The community tests these models on real speech and text — and reports back.
Arabic Triplet Matryoshka V2
Omer Nacar, RIOTU Lab, Prince Sultan University
Arabic sentence-embedding model built on AraBERT v0.2 with Matryoshka representation learning, trained on Arabic NLI triplets so a single model serves several embedding dimensions. Widely used for Arabic semantic search and RAG retrieval, and one of the few Arabic-first embedding models with a Saudi research affiliation.
GATE-AraBert-v1
Omartificial-Intelligence-Space
Arabic sentence-embedding model producing 768-dimension vectors, trained on Arabic natural-language-inference and semantic-similarity data over AraBERTv02. Developed with support from Prince Sultan University in Riyadh, and the most used Arabic embedding model with Saudi provenance. Apache-2.0.
ArabianGPT
Prince Sultan University RIOTU Lab
Small native-Arabic GPT models for research from the Saudi ArabianLLM initiative, published at 0.1B, 0.3B, 0.8B and 1.5B with task-tuned QA, summarisation and sentiment variants. Research-scale rather than production: the flagship repo has not changed since Feb 2024.
AceGPT v2
FreedomIntelligence (KAUST / CUHK-SZ / SRIBD / KAU)
Arabic-localised LLM family with cultural alignment, co-developed by KAUST and King Abdulaziz University with CUHK-Shenzhen and SRIBD. Version 2 ships base and chat variants at 8B, 32B and 70B. The repos have not been updated since Nov 2024 and there is no v3.
ALLaM-7B-Instruct
SDAIA / NCAI (weights hosted by HUMAIN)
Saudi national Arabic LLM built by the National Center for AI at SDAIA — 7B open weights, trained on 4T English tokens then 1.2T mixed Arabic/English tokens. The original ALLaM-AI Hugging Face org now redirects to humain-ai, so the weights resolve under HUMAIN's namespace; the model card still credits NCAI/SDAIA as the developer. The card claims MSA and English only.
Tarteel Whisper Quran ASR
Tarteel AI
Whisper-base fine-tuned for Quranic Arabic recitation, reporting 5.75% WER on its evaluation set. Narrow by design — it targets recitation rather than conversational Arabic — but it is the most-adopted open Arabic speech model outside the general Whisper checkpoints, and underpins Tarteel's memorisation app. A tiny variant is also published.
Qwen3
Alibaba
Open LLM family. The only model in this directory whose developer explicitly names a Saudi dialect: Qwen documents 119 languages and dialects including Arabic (Standard, Najdi, Levantine, Egyptian and others). Apache-2.0.
Whisper large-v3
OpenAI
State-of-the-art open speech recognition covering 99+ languages and the default open ASR baseline for Arabic products, though Arabic is one language among many rather than a focus and dialectal Arabic remains its weak point. Note the licence inconsistency: this model card states Apache 2.0 while the openai/whisper GitHub repo and the large-v3-turbo card state MIT.
CAMeLBERT
CAMeL Lab, NYU Abu Dhabi
BERT models pre-trained per Arabic variant (MSA, dialectal, classical) for NER, POS, sentiment and dialect identification. The mix checkpoint remains the most-used entry point despite dating from 2021.
Fanar-2-27B
QCRI / HBKU
Arabic-centric flagship of the Fanar 2.0 release (Mar 2026), continually pretrained from google/gemma-3-27b-pt on ~166B Arabic, English and code tokens with 32K context. It adds native Arabic reasoning traces, selective thinking mode and tool calling. This repo is text-in/text-out; image generation, image understanding and poetry are separate Fanar-2 models.
Fanar-1-9B
QCRI / HBKU
Qatar's sovereign Arabic LLM. This 8.7B instruct model — the 'Prime' branch — continually pretrains google/gemma-2-9b on 1T Arabic and English tokens; a separate 7B 'Star' model was trained from scratch. The card claims MSA plus Gulf, Levantine and Egyptian dialects, and alignment with Islamic values and Arab culture.
Jais 2
Inception / Cerebras / MBZUAI
Successor to the Jais family, released as open weights in Aug 2026 in 8B and 70B chat sizes (plus GGUF builds). The card states it covers MSA, regional dialects and Arabic–English code-switching without naming the dialects. Repos are gated behind a contact-sharing agreement.
Jais 30B
Inception (Core42) / MBZUAI / Cerebras
Landmark open Arabic LLM family (590M–70B), trained on Cerebras Condor Galaxy. The inceptionai org was renamed inception42, so the old repo URL only resolves via redirect; the repo is gated behind a contact-sharing agreement and has not been updated since Sept 2024. Superseded in practice by Jais 2.