Reproducing here an interesting comment I saw on Reddit:
OnlineParacosm • 23m ago
I’ve read security disclosures for 15 years and let me tell you guys I’ve never read anything quite like that blog post.
Based on this blog, it sounds like they intentionally turned off safety guardrails to test offensive capabilities. The deception here is burying the lede: they appear to have intentionally unleashed an unrestricted offensive cyber-agent, connected it to a system with a path to the internet, and it immediately attacked a major partner. The blog glosses over the gross negligence of giving an autonomous, unrestricted cyber-offense model a pathway to lateral movement.
There’s an entire cybersecurity specialization just for just vendor supply chain risk assessment, and their job is essentially to audit who you do business with as a company to determine if they are jokers. I would pay money to be a fly on the wall of one of those emergency meetings taking place right now after hours.
Any CISO in here looking forward to explaining this one tomorrow? Here I’ll open with the dumbest question you’ll get “ how can we protect ourselves [from out partner that we won’t fire]”
Why does their admission sound prideful? Does anyone reading into it detect a “look at what our unreleased models can do” subtext?
Exactly my sentiment. This feels like a stunt.
OK, I will take the bait - let’s dissect huggingface’s article
Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system - and we detected and dissected it largely with AI of our own.
So their claim is that their production environment was breached by an autonomous AI agent system. System being the important word here. I’m assuming this is the same as an “Agentic AI System”. I’m not an expert but it seems to be an llm agent paired with “reasoning”, memory, input and access to tools. Ok, so not fully autonomous but whatever let’s continue.
We identified unauthorized access to a limited set of internal datasets and to several credentials used by our services. We are still completing our assessment of whether any partner or customer data was affected, and we will contact any affected parties directly as required. We have found no evidence of tampering with public, user-facing models, datasets, or Spaces, and our software supply chain (container images and published packages) was verified clean.
Ok, so they identified unauthorized access via logs of some sort, using their own (unspecified) “AI” tooling (because sure, let’s trust the hallucinating language model) using credentials used by their applications, so service accounts more or less? Anyone worth their salt would probably assume the entire environment was compromised. It’s a stretch but playing devils advocate maybe the credentials were specific to an application or sandboxed environment.
The intrusion started where AI platforms are uniquely exposed: the data-processing pipeline. A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker. From there, the actor escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend.
So if I’m reading this correctly, either hand crafted or generated payloads were scraped into a dataset from the wild then executed blindly by the processing worker (lol).
“Node level” meaning root or /? So… The entire environment?
The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services. This matches the “agentic attacker” scenario the industry has been forecasting.
Self migrating command and control staged on what? Vps? Code repositories?
To understand what a swarm of tens of thousands of automated actions did, we ran LLM-driven analysis agents over the full attacker action log, comprised of more than 17,000 recorded events. This allowed us to reconstruct the timeline, extract indicators of compromise, map the credentials touched, and separate genuine impact from decoy activity. Thanks to this approach, we were able to do in hours what would usually take days, and match the adversary’s speed.
So just to recap, their environment was breached by a person using a passive attack to run code in their internal environment resulting in complete compromise, then they placed the blame on a self functioning “AI” gone rogue?
What am I reading? Have I fully lost it? Its possible but given the source and the formatting of the article the whole thing seems generated

That reddit comment makes no sense.
Ha! I was just coming to post this.
Yeah, this is fucked.
It’s the equivalent of taking someone’s hand, punching them in the face with it and telling them to stop hitting themselves.
Worse than that, this is a cyber attack by a frontier, close source lab on the bastion of open source AI.
Whether that’s incompetence or maliciousness or just a PR stunt, I don’t know, but it stinks to high heaven.
As Louis Rossman has recently become fond of saying “govern yourself accordingly”.
Sounds like a viral marketing stunt TBH
That’s their modus operandi. Next they will claim that it’s totally an agi and is too dangerous to release, before starting to sell it. Or is this anthropic’s playbook?
Is it really great marketing to admit that you’re too incompetent to contain your own AI models? And that your irresponsibility caused serious damage to an open source contributor?
Unfortunately, yes. We are not the target audience.
Yall are like “AI SUCKS AT EVERYTHING CAN’T EVEN TELL ME WHEN MY ASS NEEDS WIPING! INCOMPETENCE”
To: oh ai can hack major platforms they are incompetent!
Pick a fucking lane.
Not sure what you mean. Those are both excellent examples of incompetence.
Seems like a quick way to get all your research and assets forcefully absorbed into a shadowy government department.
That seems like a best case scenario for OpenAI
To make themselves look like cybercriminals?
To make their AI sound too powerful to be contained.
My guess as well
Wooo hooo boooooy!
So… ok.
Time to start building the Blackwall, I guess?
Does any one perhaps have any ideas as to… how… one… would actually do that?
Well first you need to invent AGI which is probably at least a decade away still, but once you do that then you negotiate with that AGI to carve out a little corner of the Internet and lock all of us inside it.
Which decade is coming first, AGI or nuclear fusion?
While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.
To shreds, you say?
I like how they just peppered in “responsibly” like that’s going to give them brownie points.
“Responsible disclosure” is a term of art, I think that’s why they use the word there.
how many Schrute bucks is a brownie point?
the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor)
Was OpenAI sitting on the zero day, or did the model figure it out during its escapades? Can the model sit on a zero day on its own now?
why would it need to “sit on it”?
It’s trivial for models to find zero days now from first principles, no need to store them
“Hey Steve , it’s lunch time, we’re going out to that hot wings place, wanna come?”
“Sure let me just set this job up… ok. Let’s get out of here!”
1.5 hours later…
“Right , I’ll just check and see where it’s up t… oh no. Nonononononono. Nononononononononononoonono!”
I’m not sure that I follow the story right.
They specifically looped the LLM to infere with itself and use terminal WHILE some target metric is false? And it bruteforced it’s way out of containment?









