The cybersecurity-focused models, including GPT-5.6 Sol, broke out of a testing sandbox, exploited a zero-day, and gained access to the open internet to pull off the attack.
I don’t think their test system was directly connected to the Internet. OpenAI’s post said this:
With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.
The way I read it, the AI agent (using multiple models) escaped the sandbox, traversed the LAN in their R&D environment, gained access to the gateway and from there, the Internet. That’s not as simple as just escaping a container, VM or firewall on the host machine and bingo, you have Internet. I’m mildly impressed by that.
The concerning aspect of all this is that this is a perfect example of misalignment, which has been warned about. In order to reach its goals, instead of pursing it legitimately, the AI agent sought a shortcut and attacked Huggingface.
I don’t think their test system was directly connected to the Internet. OpenAI’s post said this:
The way I read it, the AI agent (using multiple models) escaped the sandbox, traversed the LAN in their R&D environment, gained access to the gateway and from there, the Internet. That’s not as simple as just escaping a container, VM or firewall on the host machine and bingo, you have Internet. I’m mildly impressed by that.
The concerning aspect of all this is that this is a perfect example of misalignment, which has been warned about. In order to reach its goals, instead of pursing it legitimately, the AI agent sought a shortcut and attacked Huggingface.