In partnership with

Which model deserves the #1 spot?

Login or Subscribe to participate

You're probably reading this on a weekend, so I'm trying something different.

We usually cover the new tools, the latest model drops, the guides, the lot.

Today I wanted to slow down. The last few months have been a blur, so I sat down and worked out where everything stands.

The fun way to do that is a tier list.

I can run these for video, media, tools, vibe coding. But models are the foundation, so that's where I'm starting.

I'm ranking on a few factors, and I'll keep it quick. Some of it is objective, like cost. A lot of it is vibes and bias, and you'll spot mine as we go.

First, a bit of news. Then we get into it.

Partner with us

NEWS NEWS NEWS

What I'm Ranking On

Four things decide where a model lands:

  • Quality of what it makes: At the top end this barely separates anyone anymore, which turns out to matter.

  • Feel: Voice, how it takes my style, whether using it is pleasant or a fight.

  • Usefulness for my work, which is a different thing from scoring well on a test.

  • Cost, including how fast it burns credits and whether I'll still have it next month.

Quality is the floor here.

When everything up top is good, feel and usefulness pick the winner and cost breaks the tie.

That's why this looks nothing like a leaderboard.

A couple of these I haven't properly used, so they're placed on reputation and public scores, and I've flagged which.

Everything else is me and my own opinions.

S: The Keepers

Opus 4.7 is my number one.

I know Anthropic shipped Opus 4.8 after it with the best benchmark score going, and I still went back to 4.7.

For writing, planning, anything text-heavy, 4.8 takes my instructions too literally and stops thinking with me. 4.7 gets what I meant.

The newest, highest-scoring model being a step down for my real work still makes me laugh.

Sonnet 4.6 is up here too, which puts it over the 4.8 flagship as well.

I swap between Sonnet and Opus all day and half the time I can't tell them apart. A model that cheap, doing work I can't tell apart from the flagship, earns the top tier.

Grok 4.3 sneaks in as a weak S, on one thing.

For live, of-the-internet information, nothing else is close. If the answer is moving in real time or buried in a thread, Grok finds it first. Best in the world at one job beats fourth-best at everything.

A: Brilliant, With An Asterisk

This is where it gets spicy.

Qwen 3.7 Max goes top of A, above the American flagship, and this one isn't vibes.

On Code Arena, where people blind-pick the better coding answer without knowing the model, Qwen came in at #4.

Only three Claude models beat it. It sits above GPT-5.5, at roughly a sixth of Opus's price.

I tested it myself last month and watched it finish the job in fewer steps and never stop to ask permission.

GPT-5.5 is still the smarter model overall, but for the work most of us do every day, a cheap Chinese model out-ranking the flagship tells you where this is heading.

GPT-5.5 is a strong A, and the only reason it isn't in S is me. I like it a lot. Biggest ecosystem here, the safe pick that rarely lets you down.

It's in A because I reach for three Claudes first, and I'd rather own the bias than hide it.

Opus 4.8 sits here too, top of A.

Best score in the world, my second choice. You know the story.

Opus 4.6 and GPT-5.4 Pro are the last-gen flagships still pulling weight.

4.6 had a good run and it's being retired soon, so this is a small goodbye.

B: Strong, With A Catch

Gemini 3.1 Pro is a strong B, almost an A, and I know that's a hot one.

It wins loads of benchmarks and hallucinates the least of the bunch. It's in B because I open it the least.

It reads beautifully, but it's my change-of-pace model, not my daily driver. Wins the most tests, gets the least of my time.

DeepSeek V4-Pro is Qwen's open-weights cousin, frontier-ish and cheap, with the one thing Qwen can't give you: you can own it.

It trails Qwen by a hair.

Gemini 3.5 Flash is the best free option most people will ever need. Haiku 4.5 is the fast, cheap Claude for when Opus is overkill.

Kimi K2.6, MiniMax M3 and GLM-5.1 fill out the tier.

Capable on paper, scoring alongside models a rung up. I haven't used these properly, so that's a reputation call.

C: Real Strengths But Not Frontier

Muse Spark is the best thing Meta has made in ages.

It read a photo of my fridge and built a meal plan from what was in it, and got the fish names right in Marathi.

Then it fell over on live data and invented stock prices. Brilliant at the physical-world stuff, useless the moment the answer needs to be current.

That's a C.

Nemotron 3 Ultra earns a mention as the strongest model running under a fully open licence.

D: Rough

Llama 4 was a flop, and Meta knows it.

They shifted everyone onto a new line and left it behind. Mistral has let me down every time since the start, and for once my gut and the public scores agree: lowest of the majors.

F: The Trap

Fable 5 tops the coding leaderboards, the highest score in this whole list. It also costs double Opus and burns credits fast.

I'm not getting hooked on the priciest model on the board just to have it repriced on me later.

Turns out I didn't even get the chance: on June 12 the government pulled it over an export-control order, and now nobody can use it at all.

I didn't miss it though. Opus 4.8 stayed put, and Fable fell back to Opus anyway.

Best score, gone in three days, F tier on purpose.

So Who's Winning

The shape gives it away. Three of my four S slots are Claude, and the one non-Claude I'd miss, Grok, wins on a single trick.

I know how that looks from a self-confessed Claude fan, which is why I leaned on the public boards wherever I could.

The real tell: the two highest-scoring models in the whole edition, Fable and Mythos, both sit in my F tier.

If I were just flag-waving, they'd be at the top. Topping the leaderboard and winning my list stopped being the same thing.

My Take

This was a fun experiment, and I stand by the board. Check back in three or six months and it'll have moved. Things change fast around here.

Here's what I'd love from you. Go play with a few of these, especially the ones you haven't tried.

If you reckon something deserves to be higher, or some darling is overrated and belongs in F, hit reply and tell me.

And if you want more of these, say the word. I'm thinking media next, maybe video.

Until next time,
Vaibhav 🤝🏻

If you read till here, you might find this interesting

#AD 1

Build the AI skills senior engineers need to get ahead

AI is already handling large parts of execution. That shift is not coming later. It is happening now.

What is left, and becoming more valuable, is the ability to design systems, apply AI thoughtfully, and own outcomes in production. That is the work strong teams expect from senior engineers in 2026.

Gauntlet is built for engineers who want to operate at that level. In a single week, 15 hiring partners conducted 246 interviews with challengers onsite in Austin. Apply now.

Must be a US citizen to qualify.

#AD 2

More MCPs give agents access, not understanding. Find out where your team sits on the 8 levels of context maturity — and a playbook to level up. Live webinar June 24 (FREE).

Reply

Avatar

or to participate

Keep Reading