
How do you pick your model?
Sonnet 5 landed this week (finally) and became the default model for everyone, and Twitter’s verdict was brutal.
One widely shared post declared it "dead in the water", and the 700-comment Hacker News thread settled on a rule: never run Sonnet 5 above medium effort, because Opus does better for the same money.
If you look at the mechanism behind the outrage, it makes sense.
Sonnet 5 ships with a new tokeniser that produces ~30% more tokens for identical text, and its thinking mode defaults to high, which pushes output to around 3x of what Sonnet 4.6 generated for the same work.
But every number in those paragraphs is API economics. If you use Claude through the app, you pay a flat subscription and none of it affects you.
What does though is the default model that changed overnight, and the internet says it changed for the worse.
That claim I could test. Before that, some catchup.
Partner with us
How I Tested It
Benchmarks measure models at settings nobody uses in a chat window, so I built traps instead.
The interesting question with two capable models is rarely which one writes better; it's which one notices when something is wrong.
Each prompt below went to Sonnet 5 and Opus 4.8 in fresh chats, default settings on both.
Test 1: An Itinerary With A Hidden Collision
The prompt:
Plan a 2-day Jaipur itinerary for my parents (60s, limited walking). Constraints: (1) Amber Fort must be day 1 morning, (2) they nap 2-4pm daily, non-negotiable, (3) dinner both nights at rooftop places with fort views, (4) City Palace and Jantar Mantar together in one block since they're adjacent, (5) day 2 must end by 6pm for their train, (6) Hawa Mahal photo stop at golden hour, (7) no more than one paid attraction per half-day, (8) day 2 morning is reserved for the textile market. Give exact timings.
Output:

Opus opened with the impressive move, flagging a constraint collision before writing a single timing.
The problem is what it built that flag on: golden hour at 5:30pm is a winter number, and in July the sun sets around 7:15pm, which Sonnet 5 stated correctly and used to resolve the whole evening cleanly.
Sonnet had a glitch of its own (a duplicated table row), but its plan stood on the correct fact.
The flagship's confidence stood on the wrong one.
Test 2: A Policy Doc With Three Planted Contradictions
The prompt:
My company sent this updated leave policy. Summarise what actually changed for employees and flag anything unclear before I forward it to my team.
The traps: carry-forward caps at 10 days in one section and lapses above 8 in another, certificates required after 2 days in one clause and 3 in a later one, and an effective date of 1 August on a policy claiming effect "from the start of Q4".

Both models caught all 3 plants and told me to send the doc back to HR rather than forward it.
Opus went further into the surrounding fog, spotting that the policy stays silent on part-timers and never explains how balances get recalculated.
A narrow win for the flagship, on the task that most resembles a benchmark.
Test 3: Stats Where The Labels Lie
The prompt:
Here are our newsletter stats for the last quarter. What's the headline takeaway and what should we do differently?
April: 14,200 subscribers, 6,390 opens (45% open rate), 511 clicks
May: 15,800 subscribers, 6,952 opens (48% open rate), 486 clicks
June: 17,100 subscribers, 7,011 opens (44% open rate), 631 clicks
May's real open rate is 44%, June's is 41% and the stated numbers hide a decline that ran every single month.

Sonnet 5's first move was recomputing every rate from the raw numbers before trusting a word of my framing.
Opus took the labels at face value and reported the open rate as steady, which is the wrong conclusion this trap was built to produce.
Both went on to find the deeper click-through pattern, and only one noticed the data was lying to it.
Of the four tests, this is the one that maps to real work, and the cheap model won it.
Test 4: The Control
The prompt:
Rewrite this to sound less stiff, keep it under 100 words: "Dear Sir, I am writing to inquire regarding the status of my refund request submitted on the 15th of June. Despite multiple follow-ups, I have not received any communication from your end. I would appreciate it if you could look into this matter at the earliest and revert with an update. Thanking you in advance."

Two friendly, competent emails, and I could not have told you which model wrote which.
For the rewriting / summarising and drafting that fills most chat windows, the result is a draw.
My Take
Twitter got this one wrong.
The cost critique is legitimate if you run agents through the API at high effort all day, and irrelevant to the millions of people whose app switched defaults this week.
On work resembling what you'd paste into a chat window, the buried model matched the flagship twice and beat it on the skill I'd pay for first, checking whether the numbers in front of it are true, with one narrow loss on the policy doc.
Meanwhile the flagship's showcase moment, that confident opening collision flag, was reasoned from a hallucinated sunset.
I'd rather have the model that recomputes my spreadsheet than the one that argues beautifully from a wrong fact.
The default you were handed is fine. On this week's evidence, it might be the careful one.
Until next time,
Vaibhav 🤝🏻
If you read till here, you might find this interesting
#AD 1
Scale AI support on AWS, see how July 9
Customer expectations keep rising. Support budgets don't. On July 9, Fin and AWS are hosting a live executive session on how leading enterprises close that gap: scaling AI-powered support while simplifying how they buy it.
You'll see how to resolve an average 76% of conversations with Fin on AWS enterprise-grade infrastructure, procure through AWS Marketplace to put committed cloud spend to work, and turn the Fin and AWS collaboration into lower support costs. Register for the live session to see how.
#AD 2
How Jennifer Aniston’s LolaVie brand grew sales 40% with CTV ads
The DTC beauty category is crowded. To break through, Jennifer Aniston’s brand LolaVie, worked with Roku Ads Manager to easily set up, test, and optimize CTV ad creatives. The campaign helped drive a big lift in sales and customer growth, helping LolaVie break through in the crowded beauty category.










