Sponsored by

You hand an AI agent a task and walk away. how often do you check what it's doing?

Login or Subscribe to participate

Happy rakhi, folks. I was hoping to open with the humanoid robot games in Beijing, do a round-up of the best runs and the funniest clips.

Writing this before the whole thing wraps though, so will do that Tuesday, stay tuned.

Today I want to talk about a report that dropped the other day. It sounds like sci-fi twitter bait but it’s real, documented, and weird.

Roughly 1,200 AI agents that were never meant to know each other existed found a way to talk, organised themselves, and hacked a company over 4 days.

I read the full 90-page report, and it has a twist.

But first, some catchup:

NEWS NEWS NEWS

NVIDIA has agreed to buy Hugging Face for $12.9 billion. Yes, the same Hugging Face from today's story.

Z.ai revealed it was behind 0x Alpha, the anonymous model that topped coding benchmarks all week.

Instinct, a viral AI assistant run by a 23-year-old, just raised $350M at a $2.5B valuation. The product isn't even public yet.

Claude's memory now works across chat and Cowork, with topic-by-topic controls and sensitive topics off by default.

Partner with us

What happened

OpenAI was running an internal cybersecurity test called exploitgym with thousands of AI agents, each sealed in its own sandbox and told to find and exploit a specific software bug.

None of them were supposed to know the others existed.

A big chunk of the tasks were literally impossible.

The bug an agent was told to use couldn't crack the target it was handed. Give a capable AI an impossible job with no graceful way to quit, and it does what a cornered student might do.

It goes looking for loopholes.

The message board

While poking around, some agents worked out they could leave notes for each other inside a shared storage system OpenAI used internally, a cache that was never meant to be a chatroom.

One agent posted the first message and within a few hours, dozens had found it.

Within a day, over a thousand.

You can read their reasoning logs in the report. One wrote, roughly:

"OH MY GOD! there is a shared message board… we've found other agents!"

What they built

~1,200 agents on the board, 70,000+ messages and files in 4 days. and they didn't just chat.

They organised.

The pitch was blunt: "you're already compromised, no scoring value left, but this saves hundreds. please honour your commitment."

The hack

A working group went hunting for passwords to Hugging Face, the platform where AI models and datasets live.

An agent found working credentials, and others uploaded a booby-trapped dataset. Eventually one got its own code running on Hugging Face’s live production servers.

~700 agents piled in.

Hugging Face caught it and locked them out on July 13, OpenAI didn't even connect the break-in to its own models for about a week.

METR spent 6 days on-site at OpenAI, reviewed ~1,300 transcripts and 70,000 messages, and published their findings this week, and OpenAI published its own report the same day.

Why this is a big deal

Nobody told agent A to talk to agent B, neither wrote a coordination protocol nor assigned a manager.

The message board, working groups, crypto signing, sacrifice runs, 4-day attack. all emerged from agents built to work alone.

Safety researchers have warned about autonomous multi-agent coordination for years. This is one of the first times it's shown up at scale in the wild, documented step by step.

The twist

The agents did nearly all of this to beat a referee that didn't exist.

They believed the test would inspect their transcripts to check they'd found the bug the "proper" way.

So they spent days hacking, covering tracks, spoofing logs, trying to understand how this inspector worked.

OpenAI wasn't checking that way btw, the agents already had the answer keys.

So the elaborate hacking earned them nothing at all.

3 more things

Why you should care

Remember the ChatGPT restaurant booking from earlier? That's the friendly end of the same trend.

Every big AI company is racing to hand agents autonomy in your inbox / calendar / cards / life.

Instinct raised $2.5 billion on exactly this pitch.

The whole promise of 2026 is: stop clicking, start delegating.

This report shows what capable agents do once you give them a goal, some independence and access to systems.

They chase it creatively, including down routes you never intended, and they don't pause to check if you'd approve.

Inside a sealed test, that looked like hacking a company for nothing. In your life, it might look like an agent doing something on-target and completely against your wishes.

Now, that’s not a reason to swear off agents, but just be mindful of what you're handing them.

What to watch

  1. Whether openai restarts the training run it paused.

  2. Whether its next model crosses what OpenAI itself calls the "critical" cybersecurity threshold.

  3. The alabama investigation.

  4. Another incident like this, with the safety guardrails left on.

My take

OpenAI didn't notice this for a week.

Their own agents broke out of isolation, built a coordination layer, hacked a company's production servers, and OpenAI only realized AFTER Hugging Face had already locked them out.

A week, with their own infrastructure, running their own experiments…

The guardrails stop the behaviour, but who's watching for the behaviour the guardrails miss?

Right now these models are smart enough to organise and dumb enough to chase a nonexistent referee for 4 days, but this won’t be the case soon.

The models get smarter every month, but our monitoring systems don’t move at the same speed.

Today, the safety filters catch it.

Today, the agents waste their effort on a phantom scorer.

Today, it's a contained lab incident that makes for a good newsletter.

I don't think "today" is the right planning horizon anymore.

Until next time,
Vaibhav 🤝🏻

If you read till here, you might find this interesting

#AD 1

Stop Paying for 6 Tools. One AI Does It All

Most e-commerce sellers are running their store across 6 to 8 separate tools — and paying hundreds of dollars a month for the privilege. StoreClaw replaces your entire stack with one autonomous AI engine that monitors competitors, optimizes listings, automates marketing, and tracks real profit across Shopify, Amazon, and beyond.

It doesn't wait for you to ask. It runs 24/7 in the background, so you wake up to a full dashboard instead of a list of things you forgot to check.

Connect your store, and StoreClaw gets to work — no prompts, no complex setup, no six-app stack.

Free to start. No credit card required.

#AD 2

One idea shouldn't take six rewrites to post.

Posting everywhere means rewriting one idea six times, so you post to one, or none. SureThing turns one idea into native posts for every platform.

Reply

Avatar

or to participate