In partnership with

Login or Subscribe to participate

Tuesday I put Grok's image model head to head against the current #1 and it survived my tests and is very much a model to consider using.

Now there’s more Grok talk. 2 releases this week earned it.

Grok 4.6, a model that scores level with OpenAI's best on the charts and charges about a third of the price.

And Grok Bot, a cute agent product that gets its own computer, stays logged into your tools, and keeps working after you close your laptop.

Both deserve some attention, but first some catchup:

NEWS Roundup

AI CODING 101

Your coding AI is really 2 things sold together, a brain that thinks and a body that does the work.

This kit shows you how to keep the body and swap in free models for the brain, so you stop paying for the part you can get for nothing.

One key pools dozens of free providers behind a single address on your own machine. You copy-paste the install commands, drop in one config block, and the paid model slots fill with free ones.

It closes with the exact one-shot prompt that built a full Mario-style platformer, Sprout Kingdom, on the free setup.

Grok 4.6

Artificial Analysis runs its own benchmarks rather than taking self-reported scores from the labs.

This week it scored Grok 4.6 at 61 on its Intelligence Index, tied with OpenAI's GPT-5.6 Sol, with only Anthropic's two Claude models ahead at 62 and 63.

Frontier-quality answers used to come with frontier-level bills.

Grok 4.6 is the first model to break that link, and every lab above it now has to explain why it costs 3-5x more.

I wanted to see it work rather than trust scoreboard numbers. So I ran both prompts on Grok 4.6 and GPT-5.6 Sol to see where the gap shows up.

Snake game (coding task)

Both produced a working game on the first attempt.

GPT's version starts the moment the page loads. Grok's adds a start screen, a persistent high-score tracker, and a glow effect on the food dot.

On code, this was a close draw. (like Grok’s touch more though)

Agency plan (knowledge work)

This is where Grok pulled away.

GPT gave a solid strategic framework with relative pricing formulas and capacity math.

Grok gave a founder's playbook.

It named 3 exact SKUs with real prices, included a client-sunset email you could copy rn, a sales rebuttal for a prospect, a tool stack with owner assignments and what each tool replaces, and a "what I would not do" list.

The benchmarks said Grok wins knowledge work and GPT wins hard coding. That is exactly what happened.

And Grok charged a third of the price for both.

Links:

Grok bot

First, the naming. Four products now carry the Grok name and do different things.

You create named bots inside the app.

Each one runs on a cloud computer that stays on when yours is off. It signs into your tools with your credentials and works through interfaces the way a person would.

One bot can act as a "chief of staff" that manages the others.

Workflows are taught by doing them once while the bot watches, and bots message each other in group chats to hand off tasks without you relaying between them.

The early testers are loud about it.

People are already using it to negotiate with vendors, manage online store support, and keep CRMs up to date, according to the official @bot account.

The catch:

Access starts at $120 a month (Cursor Teams Premium) and goes up to $300 (SuperGrok Heavy). No free tier, though some users have found a 7-day trial on desktop at sign-in.

And every bot you create shares one computer, so credentials one bot holds are visible to all of them.

xAI's own docs say not to treat separate bots as a security boundary.

Musk said the beta will widen after Grok 4.6 rolls out, and that has now happened, so worth watching if the price drops with it.

Links:

My take

I spent all this time treating Grok as the one to check on later, I was wrong.

The reason it got good is not a mystery as early this year, most of xAI's founding research team walked out.

SpaceX, which now owns xAI, responded by buying Cursor for $60 billion and hiring Cursor's two engineering leads to rebuild Grok's coding capability from scratch.

The model you are reading about was trained on real developer data from millions of Cursor sessions.

That is a $60 billion bet that paid off in weeks, and the result is a frontier model at a third of the going rate.

If you use an AI model every day, open grok.com this week, run the same task you normally hand to Claude or ChatGPT, and decide whether the price gap still makes sense for you.

Hit reply and tell me what happened when you tried it.

Until next time,
Vaibhav 🤝🏻

If you read till here, you might find this interesting

#AD 1

[Live Session] Can you prove AI is working?

AI is in your engineering workflow. While the token spend shows it, the throughput doesn't. The human is very much still in the loop, and that's a context problem.

  • The 4 metrics to measure the gap where gains leak out before production.

  • The 8 stages of context maturity, the specific walls capping your metrics, and a free tool to pinpoint where your team is

  • Why more MCPs and bigger context windows aren’t enough and what it takes to get real value from your agents.

#AD 2

Your docs are losing you deals you never knew you lost

Developers evaluate your docs before they evaluate your product. If your documentation is slow, incomplete, or hard to navigate, they move on — and you never see it in your CRM. Mintlify customers see measurable drops in support tickets, faster time-to-first-integration, and higher conversion from trial to paid. Zapier saw a 20% increase in docs traffic after switching. HubSpot cut engineering maintenance time in half. That's what documentation-as-infrastructure actually looks like.

Reply

Avatar

or to participate