AI models tested for Arabic
Does it actually work for Saudi, Najdi, Hijazi or Gulf Arabic? Benchmarks rarely say. The community tests these models on real speech and text — and reports back.
Mistral 7B
Mistral AI
Efficient open-weight model common in regional stacks; v0.3 is the current 7B instruct release and remains Apache 2.0. Arabic is not a declared language on the card. Note that Mistral's larger flagships moved to the non-open Mistral Research Licence, so the permissive terms apply to this size, not the family.
DeepSeek-R1
DeepSeek
Open reasoning model under a plain MIT licence, allowing commercial use and derivatives including distillation. Arabic is not a declared target language — the card lists no languages — so Arabic behaviour is incidental. Superseded within the DeepSeek line by the R1-0528 update and later V3.x releases.
Falcon 3
TII
Efficient open LLM series from Abu Dhabi's TII. The model card lists English, French, Spanish and Portuguese only — Arabic is not a supported language here, which is why TII built Falcon-Arabic by adding 32,000 Arabic tokens to this tokenizer. For Arabic from TII, use Falcon-H1 or Falcon-Arabic.