Gemini 3.7 Flash: not the smartest, and it does not need to be

It shipped, and it retired a version that was three weeks old. It will not write your quantum physics thesis, and that is not a fault.

AA Abdelilah Arahal
3 min read Updated 24 September 2026

Gemini 3.7 Flash shipped. And it retired a version that was three weeks old.

Before you cancel subscriptions and declare the competition over, here is the plain truth: this model will not write your quantum physics thesis.

It does not stand level with Claude Opus 5, nor with the analytical reach of GPT-5.6 Sol. On hard frontier tasks and deep reasoning it looks like a diligent student in front of the scientists of its age.

So is it a failed update?

The opposite.

Where it genuinely wins

What it lost in the genius race, it made up in repetitive daily work.

Do you need a physicist to tidy your desk? No.

Enterprise workflow automation
Gemini 3.7 Flash30.4%
Claude Sonnet 510.7%
A mid-tier comparison, not against frontier models.

Thirty percent against ten. That is not a small gap within a class. It also took 1,588 points on WebDev Arena for web development quality.

And it added something practical for anyone building agents: breaking failure loops. It stops an agent spinning on a repeated coding error and corrects course immediately, which is a problem that costs real money in production.

The price, which is the actual story

$0.75per million input tokens
$3.75per million output tokens

That is the whole point: it does not need to be the smartest. It needs to be the cheapest and fastest. Which is exactly what a project launched by people without a large budget needs.

If the output-versus-input split is not obvious to you, it is the biggest trap in these bills, and it is covered in token economics for managers.

And my view, which you may not like

I will say it as I feel it: when I see the word "Flash" next to a new model, I lose interest automatically.

Gemini 5.1 Pro is, to me, the last genuinely strong model Google made. What the company is doing now looks like expanding at the edges instead of advancing at the centre.

I may be wrong. But that is what I see.

Try this today: take your most repeated task, run it on this model and on your current one, and compare cost and result. Twenty test cases is enough.

Sources

Share

No comments yet

Leave a comment

Never published. Used only if I reply to you directly.

Comments are read before they appear.

Related reading

4 min read

Will AI replace accountants?

Entries, reconciliation, transcription: all heading for automation. But signing a tax return is accountability, and accountability does not automate.

Start with a short call

Fifteen minutes to understand what your team does and what you want to change. If training is not the right answer, I will say so.