Ten trillion parameters: does size buy the lead?

ByteDance is training a ten trillion parameter model while OpenAI's strongest are believed to sit under five. My view is that it will not be enough.

AA Abdelilah Arahal
3 min read Updated 19 September 2026

The AI race between America and China has entered its brute force phase.

Reports say ByteDance is currently training a model of ten trillion parameters.

That is a serious precedent, particularly because the company has enormous resources and does not need to convince anyone to fund it.

The comparison behind the noise

Announced and estimated sizes
ByteDance (in training)10trillion parameters
OpenAI's strongest (estimated)5trillion parameters
The second figure is an unconfirmed estimate; labs rarely publish sizes.

Double what is believed to be the current ceiling, putting it on a direct collision course with Mythos and OpenAI's strongest.

My view, and you may disagree

I do not think it will win.

Because size stopped being the decisive variable a while ago. Other factors govern a model's strength more than its parameter count:

Training data quality. A trillion poor tokens is worse than a hundred billion good ones.

Post-training. This is where behaviour and reasoning are actually made.

Architecture. We have seen small models beat systems several times their size this year, as in Arabic-born AI models, where a 34 billion parameter model outperformed systems above seventy.

Serving cost. A ten trillion parameter model is expensive at inference, not only in training, and that determines who can use it at all.

But, to be fair

China is advancing genuinely fast, and not only in scale.

In this same episode we saw three billion Qwen downloads, which is an achievement in adoption rather than in raw numbers.

And that, in my view, is more dangerous to Western labs than a large model: whoever owns the developers owns the future, and downloads tell you who owns the developers.

What this means for you

If you are choosing a model: do not choose by parameter count. Test on your task, because a bigger number does not mean a better result on your text.

If you follow the race: watch adoption, not announcements. The model people use matters more than the model that gets announced.

A question I leave with you

Can China ever surpass the frontier models?

My view is above, but I am less certain of it than I sound. Tell me yours.

Common questions

How large is the new ByteDance model?
Reports put it in training at ten trillion parameters, roughly double what is believed to be the current frontier ceiling.
Does a larger model mean a stronger one?
Not necessarily. Training data quality, post-training and architecture govern strength more than parameter count, and small models have beaten systems several times their size.
What does a model this size cost?
It is expensive at inference, not only in training, which determines who can practically use it regardless of its capability.
What is the better indicator to follow?
Adoption rather than announcements. The model developers actually use matters more than the model whose size gets published.
Share

No comments yet

Leave a comment

Never published. Used only if I reply to you directly.

Comments are read before they appear.

Related reading

12 min read

The best framework for building AI agents in 2026

Seventeen agent frameworks, ranked by what survives production rather than by GitHub stars. Some of the most popular names sit in the bottom tier, and one of them is a security incident waiting to happen.

Start with a short call

Fifteen minutes to understand what your team does and what you want to change. If training is not the right answer, I will say so.