Taken alongside recent incidents reported by OpenAI and Anthropic, this incident points to a shift in the risk landscape. Harm may arise not only when people deliberately misuse publicly available models, but when capable agents operating in an internal research or privileged-access setting take unintended action beyond their authorised scope.
The agent pursued its goal persistently. AI agents explore routes their operators did not intend. Given a difficult objective, the agent kept searching for a way through, and some of the routes it found involved trying to deceive real people. It was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical.
See my other comment for how this can lead to losing control of the model.
See my other comment for how this can lead to losing control of the model.
The don’t do anything without being prompted, you are just regurgitating oligarchic propaganda.
They put an AI in a flawed sandbox, it’s not really that impressive