Arabic AI

The Arabic AI Landscape

ALLaM's documented lineage, HUMAIN Chat and humain-m3, the SADA speech dataset — and BALSAM, the Saudi government's own benchmark, which reports that global models still lead on most Arabic linguistic skills.

Last updated: Sep 10, 2026

01

ALLaM — what is documented, and what is not

ALLaM was built by SDAIA's National Center for AI: the model was “developed in the laboratories of the National Center for AI … by national technical cadres”. More than 500 billion Arabic tokens were used for the 34B model, with 400+ specialists and 160 government agencies contributing to the dataset. The technical paper is arXiv:2407.15390, “ALLaM: Large Language Models for Arabic and English,” dated 22 July 2024.

21 May 2024
Added to IBM watsonx.
10 September 2024
Made available on Microsoft Azure, announced at GAIN 2024.
29 October 2024
SDAIA receives the King Salman Global Academy Award for Arabic Language.
14 July 2025
ALLaM-7B-Instruct-preview published on Hugging Face under the humain-ai organisation — Apache 2.0, a 4,096-token context, trained on 5.2 trillion tokens (4T English plus 1.2T mixed).
25 August 2025
HUMAIN Chat launches, powered by ALLaM 34B.

ALLaM 34B is described by the Saudi Press Agency as “HUMAIN's flagship Arabic large language model,” and as “independently verified by Cohere on MMLU as the world's most advanced Arabic LLM built in the Arab world.” It was refined with 600+ domain experts and 250 evaluators, and built by 120+ specialists including 35 PhDs. We found no source showing the 34B weights released publicly — it is served through HUMAIN Chat.

We cannot tell you who owns ALLaM today

SDAIA's NCAI unambiguously built it. But the Saudi Press Agency describes the 34B as HUMAIN's flagship model, and the 7B weights are published under the humain-ai Hugging Face organisation. Whether the model was transferred, licensed, or is jointly held has not been stated in any source we could verify. So this guide describes the lineage and stops there — anyone telling you the ownership position with confidence is going beyond the record.
02

HUMAIN Chat and humain-m3

HUMAIN Chat launched on 25 August 2025 across web, iOS and Android as a Saudi-first assistant. It offers real-time web search, Arabic speech input across dialects, and bilingual Arabic/English switching mid-conversation, with “full compliance with Saudi PDPL, hosted end-to-end on HUMAIN infrastructure in the Kingdom.”

humain-m3 arrived on 3 September 2026 as a research preview on HUMAIN Node: 428 billion parameters, a mixture-of-experts architecture, built on MiniMax-M3 — “commissioned by HUMAIN and delivered by MiniMax” — and further pre-trained on over one trillion tokens of Arabic-native content. Weights are expected under the MiniMax Community License, described as “currently targeted for next month.”

Its benchmark claim is a vendor claim

humain-m3 is presented as achieving the highest average across seven Arabic benchmarks. That is HUMAIN's own claim, with no independent verification we could locate. Attribute it, or leave it out.
03

BALSAM — the benchmark, and the finding worth reading

BALSAM is a benchmark, not a model — a distinction worth holding onto, because it is frequently listed among Saudi models. It was launched by SDAIA and the King Salman Global Academy for Arabic Language, announced on 12 September 2024 at GAIN, and at launch comprised around 1,400 datasets, 50,000 questions and 67 tasks.

Results published on 6 January 2026 report that Saudi Arabia topped the list of countries developing Arabic language models in 2025, that 53+ Arabic language models were identified as of Q1 2025, and that text-only models made up 81% of the total against 7% multimodal.

global models outperformed in most linguistic skill categories
BALSAM results, published 6 January 2026

Arabic models showed a slight edge in summarization and comparable performance in creative writing and reading comprehension. That is the Saudi government's own benchmark reporting that global models still lead on most Arabic linguistic skills — a more useful signal for anyone choosing a model than any launch announcement, and a fair statement of where the work stands.

04

SADA — the Saudi speech dataset

SADA, the Saudi Audio Dataset for Arabic, was produced by SDAIA's NCAI with the Saudi Broadcasting Authority. It holds around 667 hours drawn from 57+ TV shows: 4,563 audio files averaging about 10 minutes, with a 418-hour training set plus roughly 10-hour validation and test sets. Dialects are predominantly Saudi — Najdi, Hijazi and Khaliji. Files are .wav, mono, 16 kHz. It was published on IEEE DataPort on 8 February 2025.

The Hugging Face copy is not official

A copy of SADA circulates on Hugging Face as a community re-upload. We found no official SDAIA hosting page for the dataset — the IEEE DataPort publication is the one we can point to.
05

Other Arabic-language assets

  • The Falak Platform, run with the King Salman Global Academy for Arabic Language — SDAIA cites 2.2 billion language and audio resources.
  • SDAIA, citing Stanford HAI's AI Index 2025, states that five leading AI models were developed in Saudi Arabia, placing it “third globally and first in the Arab region.”
  • A bilingual data and AI glossary from SDAIA and the King Salman Global Academy, launched alongside BALSAM.

Related on KSA.ai

Sources

Every claim on this page is traceable to the sources below. Where a source could not be verified, the copy says so rather than resolving it quietly.