AI models tested for Arabic
Does it actually work for Saudi, Najdi, Hijazi or Gulf Arabic? Benchmarks rarely say. The community tests these models on real speech and text — and reports back.
Arabic Triplet Matryoshka V2
Omer Nacar, RIOTU Lab, Prince Sultan University
Arabic sentence-embedding model built on AraBERT v0.2 with Matryoshka representation learning, trained on Arabic NLI triplets so a single model serves several embedding dimensions. Widely used for Arabic semantic search and RAG retrieval, and one of the few Arabic-first embedding models with a Saudi research affiliation.
GATE-AraBert-v1
Omartificial-Intelligence-Space
Arabic sentence-embedding model producing 768-dimension vectors, trained on Arabic natural-language-inference and semantic-similarity data over AraBERTv02. Developed with support from Prince Sultan University in Riyadh, and the most used Arabic embedding model with Saudi provenance. Apache-2.0.