# The Language They Did Not Build For

Arabic is spoken by 400 million people.

It accounts for 1% of global online content.

That number is not a coincidence. It is a consequence. Decades of internet infrastructure, search algorithms, social platforms, and now AI systems built primarily by and for English speakers produced a digital world where Arabic exists at the margins of what the technology was designed to understand.

And then the most ambitious AI infrastructure project in the region was built on top of that foundation.

* * *

I live in Dubai now. I work across Abu Dhabi and the broader Gulf. The conversations I have here about AI are substantive, urgent, and globally consequential in ways that do not always travel back to the conference rooms where the western AI narrative gets written.

What I observe daily is a region that is moving fast on AI sovereignty while simultaneously inheriting the English bias that the foundational models carry inside them.

Most major AI models claim Arabic support. What they actually offer is something closer to English comprehension wearing Arabic as a second language. The models were trained primarily on English data, fine-tuned on translated Arabic text, and then deployed into a region where the language does not work the way English works.

Arabic is root-based. One root generates dozens of words, each carrying subtle shifts in meaning depending on context. Word order is flexible in ways that break the assumptions baked into English-first architectures. Modern Standard Arabic, the formal written language, and the dialects people actually speak every day are not the same thing in any practical sense. Gulf Arabic, Egyptian Arabic, Levantine Arabic. Different enough that a model trained on one will stumble on another.

The result shows up where it matters most.

Government service chatbots that produce responses that sound grammatical but miss meaning entirely. Healthcare AI that a patient cannot trust in their own language. Education platforms where a student in Abu Dhabi gets a worse experience than a student in Austin doing the same task. Customer service tools that default to English the moment the query gets complex.

This is not a minor UX problem. In healthcare, it is dangerous. In government, it is a trust problem that compounds with every failed interaction. In education, it teaches an entire generation that AI does not speak their language. Because for most of its history, it did not.

* * *

The western AI industry did not build for this region deliberately. It built for the data it had. English dominated the internet. English dominated the training sets. English dominated the benchmarks. The models that emerged were extraordinary at English and adequate at everything else.

Adequate is not acceptable when you are building sovereign AI infrastructure for 400 million people.

What makes this worth writing about is not the failure. Failures of omission are not always malicious. What makes it worth writing about is what the region decided to do about it.

* * *

In Abu Dhabi, the Technology Innovation Institute built Falcon. It is now one of the most capable open-weight model families in the world. The Falcon-H1 Arabic release in early 2026 did something that deserved more attention than it received in the western AI press. A 34-billion parameter model, built specifically for Arabic speakers rather than adapted from English, topped the Open Arabic LLM Leaderboard. It outperformed Meta's Llama-70B and Alibaba's Qwen-72B despite being less than half their size.

Let that land for a moment.

The UAE built a model smaller than its competitors that beat them on their own language's benchmark. Not by adapting English. By building Arabic from the foundation up.

Also in Abu Dhabi, G42's Inception unit partnered with MBZUAI and Cerebras to build Jais. Trained on the largest Arabic corpus of any open model. 116 billion Arabic tokens alongside 279 billion English tokens. Designed for the kind of formal Arabic that government services, legal systems, and regulated industries require. Jais has been running inside UAE government infrastructure since early 2024. It is not a research project. It is in production.

Saudi Arabia has ALLaM, now operating under HUMAIN. Qatar has Fanar, a sovereign LLM built by QCRI. The Open Arabic LLM Leaderboard, run by TII, is a public benchmark that measures genuine Arabic capability rather than translated English performance.

The leaderboard had to throw out its own original questions. They had been translated from English, which meant they were measuring English comprehension in Arabic script rather than actual Arabic understanding. The act of replacing them was itself a statement. The region decided it would not accept someone else's definition of what Arabic AI should be measured against.

* * *

There is a pattern worth naming.

Every region the western AI industry underserved is now building its own foundation models. Not because they want to. Because they had to. The data gap, the language gap, the cultural gap between what the models assumed and what the region needed was wide enough that adaptation was never going to close it.

What is being built in Abu Dhabi, in Riyadh, in Doha, is not a reaction to exclusion. It is a correction of it. The Gulf is not asking the western AI industry to do better. It built the infrastructure to do it itself.

That is a different posture than most of the world has taken toward AI sovereignty. Most countries are still negotiating with the infrastructure they were given. The Gulf decided the infrastructure itself needed to change.

* * *

I am optimistic about what is being built here. Not cautiously optimistic. Genuinely optimistic.

The work coming out of TII and MBZUAI is technically serious. Falcon-H1 Arabic beating models twice its size is not a regional feel-good story. It is a result. Jais running in production inside government infrastructure is not a pilot program. It is sovereign AI operating at scale in one of the most demanding regulatory environments in the world.

The region that was built for last is building for itself first.

That shift matters beyond the Gulf. Every language community that has been living with AI that works better in English than in the language they think in has something to learn from what is happening in Abu Dhabi right now. The approach being developed here, building from the foundation rather than adapting from someone else's, is the only approach that actually closes the gap.

The other approach produces responses that sound grammatical but miss meaning. And meaning is the whole point.

* * *

*Andrew Quillen is the founder of AndMaverick, a global Enterprise AI Orchestration consultancy. He is based in Dubai and advises enterprise and government clients across the Gulf, Asia Pacific, Europe, and the Americas on AI systems architecture and orchestration strategy. To continue the conversation, visit* [*andmaverick.com*](http://andmaverick.com)*.*
