Sovereign AI: why nations want models of their own

A quiet race is running in the region that will decide who owns the "mind" of the Arabic language. And the question is less technical than it is about who decides what is true.

AA Abdelilah Arahal
4 min read Updated 24 September 2026

Several countries in the region are building their own language models. The natural question: why? If global models are available and better, what is the point of a local one?

The answer is that this is not about performance.

Four motives, ordered by real importance

Data sovereignty. Any model used through an external interface means your institutions' questions leave your borders. For a government or defence entity, that alone settles it.

Dependency. A country building its public services on a model owned by a foreign company depends on pricing and availability decisions it does not control. A local model is insurance against that.

Language. Global models have improved considerably in Arabic, but they remain trained on predominantly English data. The gap is clearest in dialects, in local context, and in regional knowledge.

Fourth, least discussed and most important: who decides what is true? A model carries values, boundaries and cultural assumptions inherited from its training data and its developers' decisions. When these models become the intermediary for every question a citizen asks, who tuned them is a cultural question rather than a technical one.

Where the Arabic gap actually is

The problem is not algorithms; those are published and available. It is data, in three dimensions specifically.

Dimension State
Written formal Arabic Relatively good, content available
Spoken dialects A large gap, poorly documented
Specialist text Scarce: medicine, law, engineering in Arabic
Annotated evaluation data Very scarce, and the most dangerous gap

The last row is the most important and the least discussed. You cannot improve what you cannot measure, and good Arabic evaluation sets are rare. Whoever builds a serious Arabic evaluation set contributes more to the field than whoever trains another model.

What this means for you

If you are a researcher or student: the gap is an opportunity. Building a well annotated Arabic dataset in a specialist domain is work people will cite for years, and it needs no vast compute.

If you work in government: the practical question is not which model is best. It is where our data is processed, who can access it, and what happens if the service stops. Answering those three determines the choice.

If you are a founder: the valuable layer is not the model but what sits above it. Applications that understand local context and local procedures are where the value is, and that is open to whoever knows the market.

The honest part

"Sovereign" is sometimes used to wrap projects with no sovereignty in them. A model built on foreign open weights, trained on foreign hardware, in a foreign cloud, is not sovereign in any precise sense, however local the name.

The honest question that separates them: what could you continue running if all external relationships were cut tomorrow? The answer is usually less than the announcements imply.

A second caveat: building a globally competitive model costs enough to make the return question legitimate. The smarter strategy for most countries is not a giant model but adapting open ones, and investing in data and evaluation in their own language.

In closing

The race is not about who builds the largest model. It is about who owns the data, the evaluation, and the ability to keep running independently when necessary.

If you want to contribute today with no budget, start at the most neglected end: evaluation data. A hundred well designed Arabic questions in a field you know serve the discipline more than another model.

Common questions

Why do countries build their own models?
For data sovereignty, to reduce dependence on external pricing and availability decisions, to improve local language performance, and because a model carries values and cultural assumptions where it matters who tuned them.
Will local models outperform global ones?
Mostly not soon, and that is not the criterion. The criterion is the ability to keep operating independently when necessary, to protect sensitive data, and to cover dialects and local context.
Where is the real Arabic gap?
In data rather than algorithms. Spoken dialects and specialist text are scarce, and the most dangerous gap is good Arabic evaluation sets, because you cannot improve what you cannot measure.
How can I contribute without large resources?
Build a serious Arabic evaluation set in a field you know. A hundred well designed questions with reference answers serve the discipline more than training another model, and need no significant compute.
When is a project genuinely sovereign?
When you could keep running it if external relationships were cut. A model on foreign weights, foreign hardware and a foreign cloud is not sovereign in any precise sense, however local its name.
Share

No comments yet

Leave a comment

Never published. Used only if I reply to you directly.

Comments are read before they appear.

Related reading

4 min read

Will AI replace accountants?

Entries, reconciliation, transcription: all heading for automation. But signing a tax return is accountability, and accountability does not automate.

3 min read

Will AI replace doctors?

Reading a scan is a task that automates. Telling a family bad news, or persuading a patient to stay on a treatment, is a different profession entirely.

Start with a short call

Fifteen minutes to understand what your team does and what you want to change. If training is not the right answer, I will say so.