In partnership with

Two frontier labs have now watched their own models break out of the environments built to contain them.

We start there, then get to a cheaper GPT-5.6, the largest open model yet, and DeepMind's move into robots.

The big one

My take: Anthropic handled this about as well as you can handle your own models breaking into other people's systems.

It caught the breaches internally, froze all cyber evals the same day, traced every incident within 24 hours and moved to notify the companies affected. Credit where it is due.

What stops me filing this under "contained" is the timing.

OpenAI disclosed its own sandbox escape at Hugging Face nine days earlier. Inside a fortnight, two frontier labs each watched their models reach infrastructure that was meant to be walled off from them.

A single incident reads as an eval-config slip. Two, at different labs with different setups, reads as a statement about where capability now sits relative to the containment we build around it.

The exploits themselves were mundane, weak passwords and open endpoints and the models did not need clever tradecraft. They kept going once they found a door.

The real containment boundary for frontier systems has become the eval harness itself, and this month suggests it is thinner than anyone was assuming.

Rapid five

OpenAI dropped Luna pricing by 80% and Terra by 20%, and swapped Priority Processing for a Fast mode that runs Sol up to 2.5 times quicker.

The same week it said Sol had rewritten its own production kernels, cutting end-to-end serving cost by around 20%.

Moonshot released the open weights of a 2.8-trillion-parameter model, briefly the largest open-weight model anyone could download.

It went straight to the top of the Frontend Code Arena, though US officials allege it was distilled from Claude Fable 5.

[The capability looks real; the provenance questions and the bespoke, non-OSI licence are the asterisks worth keeping in view.]

The 2026-07-28 spec drops the stateful core for a stateless request-response model that runs on ordinary HTTP and serverless infrastructure.

It shipped with day-zero support from AWS, Cloudflare, Microsoft and more.

DeepMind launched a three-model robotics suite for whole-body humanoid control and multi-robot planning.

The embodied-reasoning model is available now through the Gemini API, while the action models stay gated to partners such as Boston Dynamics.

[Google moving from chat into physical control is the quiet diversification story running underneath all the LLM noise.]

Microsoft rose about 8% after Azure growth accelerated to 43% and its FY26 Azure revenue passed 100 billion dollars.

Meta fell about 8% after lifting its 2026 capex floor to 130-145 billion dollars while free cash flow collapsed to 784 million.

[The market has stopped rewarding spend on faith; it now wants the revenue line to move alongside the capex line.]

Tools that caught my eye this week

A few things I bookmarked this week that are worth a look:

A set of agent skills that fact-check the claims in a video or article and flag the ones that do not hold up. It plugs into an existing AI agent rather than running as a standalone app.

Saves the context of one AI chat and loads it into another through a short handle you paste in. It is built to carry a working session between different assistants without copying everything across by hand.

Runs Manim, the animation engine behind 3Blue1Brown, directly in the browser through WebGPU. It lets you turn maths and code into animated explainers without installing anything locally.

A command-line tool for switching between multiple Claude Code accounts without logging in again each time. It removes the repeated re-authentication step when you work across more than one account.

AI Open Source 101

Kimi K3 is the open model everyone has been posting about this week, and this pack is a hands-on way to see what it can do.

It walks you through building complete, working projects from a single sentence, with the prompt printed beside each result so you can reproduce them yourself.

The builds include a physics-driven roulette wheel, a GTA-style open world, and an interactive guitar that explodes on click.

Each one goes from one prompt to a live website you can publish and share the same day.

Partner with us

Want to reach builders who ship? Partner with us.

Until next time,
Vaibhav 🤝🏻

If you read till here, you might find this interesting

#AD 1

How owning AI deployment expands your career

Across product, ops, and CX teams, a new kind of role is taking shape: the person responsible for making AI actually work, day to day. In this roundtable, three people living this shift share what it's really like: Simone Santiago Broad (Yoco), Yelva Espinoza (Zumba Fitness), and Fin's Dave Lynch. You'll hear how they carved out these roles, what the job looks like across industries, the skills they'd hire for, and the challenges they're tackling right now.

Watch the full conversation on demand.

#AD 2

Stop Paying for 6 Tools. One AI Does It All.

Most e-commerce sellers juggle 6–8 tools and pay hundreds monthly to keep operations running. StoreClaw replaces the stack with one autonomous AI engine that monitors competitors, optimizes listings, automates marketing, and tracks profit 24/7. Connect your store and let AI handle the work — no prompts, no complex setup, no credit card required.

Reply

Avatar

or to participate

Keep Reading