Ten trillion parameters: does size buy the lead?
ByteDance is training a ten trillion parameter model while OpenAI's strongest are believed to sit under five. My view is that it will not be enough.
The AI race between America and China has entered its brute force phase.
Reports say ByteDance is currently training a model of ten trillion parameters.
That is a serious precedent, particularly because the company has enormous resources and does not need to convince anyone to fund it.
The comparison behind the noise
Double what is believed to be the current ceiling, putting it on a direct collision course with Mythos and OpenAI's strongest.
My view, and you may disagree
I do not think it will win.
Because size stopped being the decisive variable a while ago. Other factors govern a model's strength more than its parameter count:
Training data quality. A trillion poor tokens is worse than a hundred billion good ones.
Post-training. This is where behaviour and reasoning are actually made.
Architecture. We have seen small models beat systems several times their size this year, as in Arabic-born AI models, where a 34 billion parameter model outperformed systems above seventy.
Serving cost. A ten trillion parameter model is expensive at inference, not only in training, and that determines who can use it at all.
But, to be fair
China is advancing genuinely fast, and not only in scale.
In this same episode we saw three billion Qwen downloads, which is an achievement in adoption rather than in raw numbers.
And that, in my view, is more dangerous to Western labs than a large model: whoever owns the developers owns the future, and downloads tell you who owns the developers.
What this means for you
If you are choosing a model: do not choose by parameter count. Test on your task, because a bigger number does not mean a better result on your text.
If you follow the race: watch adoption, not announcements. The model people use matters more than the model that gets announced.
A question I leave with you
Can China ever surpass the frontier models?
My view is above, but I am less certain of it than I sound. Tell me yours.
Common questions
- How large is the new ByteDance model?
- Reports put it in training at ten trillion parameters, roughly double what is believed to be the current frontier ceiling.
- Does a larger model mean a stronger one?
- Not necessarily. Training data quality, post-training and architecture govern strength more than parameter count, and small models have beaten systems several times their size.
- What does a model this size cost?
- It is expensive at inference, not only in training, which determines who can practically use it regardless of its capability.
- What is the better indicator to follow?
- Adoption rather than announcements. The model developers actually use matters more than the model whose size gets published.
No comments yet
Leave a comment