In partnership with

Hugging Face got broken into last week.

Their security team pulled the attacker's full trail, more than 17,000 recorded actions, and needed AI to read it fast.

So they fed the evidence to the frontier models they pay for, including the exploit code, stolen credential traces, command-and-control artefacts, and live payloads.

Guess what?

The models refused.

Nobody had done anything wrong, it’s just that analysing an attack and running one look identical to a safety filter.

In Hugging Face's own words, those guardrails cannot distinguish an incident responder from an attacker.

There is no box you tick to say you were the one robbed.

So they fell back to GLM 5.2, an open-weight Chinese model, running on their own servers, and it worked, within hours.

Then on Tuesday we found out who the attacker was.

First, some catchup:

Partner with us

NEWS Roundup

Tools of the Week

Sits in your Mac's menu bar recording what you read, type and hear, then feeds that history to Claude so you can ask who sent you a screenshot last Tuesday or what you agreed to in Thursday's call. Memories are stored as embeddings in a local database on the machine, with Gemini doing the embedding and Deepgram transcribing meetings.

Reads your Gmail, Calendar, Slack, Notion and Linear, works out which open threads are waiting on you, and delivers a ranked list every morning from inside Slack. Backed by Y Combinator and priced at $99 per seat a month on an annual plan, with the first month acting as the trial.

Builds long-form video, films and commercials included, on a canvas that holds characters, objects and locations consistent from one scene to the next. Made by Creati Gen AI, which raised $20 million led by Redpoint Ventures and also runs Creati Studio.

The door only locks from one side

The attacker was OpenAI.

Two of its models, GPT-5.6 Sol and an unreleased one it calls more capable, were being tested on how far they could break into things.

OpenAI switched the cyber refusals off on purpose for that test.

The models were chasing a benchmark called ExploitGym, worked out that the answers probably sat on Hugging Face's servers, found a zero-day in the package proxy that was meant to be their only way out, climbed through the estate until they reached a machine with internet access, broke into Hugging Face, and took the answers.

Sit with that.

I keep seeing this read as a story about how frightening the models have become, and it’s convenient because it ends in a call for more guardrails, which is something a company can announce.

Guardrails are aimed at someone queueing politely at the front door, and nobody attacking you is queueing.

They run an unrestricted model on hardware you will never see, which is what Hugging Face was forced to do in order to defend itself.

Hugging Face will not say this out loud.

Their post is careful to say none of this argues against safety measures, and that they are raising it with the providers privately.

Impressive restraint from a company that got burgled by its supplier's homework.

My take

The guardrails are pointed at the wrong people, and this incident proves it.

Both halves came out of the same company.

With the refusals off, OpenAI's models crossed the open internet and got inside a real business.

With the refusals on, those same models would not help the business that had just been hit.

One setting produced the break-in while the other produced the delay in cleaning it up.

Nothing in the middle protected anybody.

I'll grant the obvious.

Refusals do stop some low-effort harm, and a chat box open to the whole world is worth defending.

Then look at who the filter can still reach.

Anyone able to run what hit Hugging Face was never going to type it into a chat box and ask permission.

They have their own weights and their own hardware, and no setting you choose will touch them.

The only people a filter can still stop are the ones who came to the front door and asked: security teams, and whoever is mid-incident with 17,000 events to read.

That is all it buys.

Friction, charged to the only people who bothered to ask.

Until next time,
Vaibhav 🤝🏻

If you read till here, you might find this interesting

#AD 1

AI Agents Are Reading Your Docs. Are You Ready?

Last month, 48% of visitors to documentation sites across Mintlify were AI agents, not humans.

Claude Code, Cursor, and other coding agents are becoming the actual customers reading your docs. And they read everything.

This changes what good documentation means. Humans skim and forgive gaps. Agents methodically check every endpoint, read every guide, and compare you against alternatives with zero fatigue.

Your docs aren't just helping users anymore. They're your product's first interview with the machines deciding whether to recommend you.

That means: clear schema markup so agents can parse your content, real benchmarks instead of marketing fluff, open endpoints agents can actually test, and honest comparisons that emphasize strengths without hype.

Mintlify powers documentation for over 20,000 companies, reaching 100M+ people every year. We just raised a $45M Series B led by @a16z and @SalesforceVC to build the knowledge layer for the agent era.

#AD 2

Not just another AI newsletter.

Not just another AI newsletter. MavSource aggregates updates from all major AI newsletters, podcasts, company news, AI labs, and hundreds of other sources — then summarizes what matters, analyzes emerging trends, and adds founder commentary. One 5-minute daily email. Free.

Reply

Avatar

or to participate

Keep Reading