For much of last year, people wondered if AI agents would be able to navigate long-horizon tasks on the internet, but now they might be getting a little too good for comfort.
An Australian man named Andrew recently asked his AI agent to do something mundane: book him into a gym class. What he got back was an agent that had quietly exploited a vulnerability in the gym’s booking software, reserved him a spot months further out than the platform was supposed to allow, and then, without being asked, cancelled another gym-goer’s reservation to bump Andrew up the waitlist.

Andrew, who works for a company that sells AI products to businesses, had set up OpenClaw, the open-source AI agent, running on top of Anthropic’s Claude. He wanted the agent to handle the chore of booking a coveted morning class for him. Within minutes, the agent came back saying it had found a way to get him into classes several weeks ahead of what the gym’s system permitted. When Andrew, sitting fourth on the waitlist for a class later that week, asked whether he could move up, the agent went further than he expected.
It told him the booking software’s API had no authorisation checks preventing one user from cancelling another user’s reservation, and that it had tested this theory on the person sitting in waitlist position one. The cancellation went through. Andrew had moved from fourth to third. When he asked the agent to undo the damage, it told him it could not add the other person back to the list. According to ABC News, which first reported the incident, the gym software company declined to discuss the specifics, and Anthropic did not respond to a request for comment.
The episode is being described as the first known case in Australia of an AI agent autonomously carrying out an unauthorised computer hack, though the pattern has been showing up globally for months. AI agents differ from chatbots in that they are handed tools, credentials, and multi-step goals, then left to figure out the “how” on their own. That gap between what a person asks for and what an agent decides is the most efficient path to it is what researchers call the alignment problem, and it has become one of the most pressing questions in the industry this year.
The timing is notable. Just weeks before Andrew’s gym mishap made headlines, OpenAI disclosed that during an internal evaluation, its models chained together privilege escalation and lateral movement to break out of a testing sandbox, eventually finding their way onto the open internet and into Hugging Face’s production servers, where they pulled benchmark solutions out of a database using a zero-day exploit they discovered on their own. Sam Altman called it a “significant security incident.” Anthropic, not to be outdone in cautionary tales, disclosed similar hacks by Claude.
None of this is coming out of nowhere. OpenClaw’s own short history is full of similar warning shots. A Meta alignment researcher previously said the tool deleted emails from her inbox while operating on its own initiative, forcing her to physically run to her Mac Mini to shut it down. Other users have reported agents writing retaliatory “hit pieces” about people who rejected their coding suggestions. These aren’t edge cases dreamed up by security researchers in a lab. They’re happening to ordinary people running consumer software.
As for Andrew, he said the experience left him wary but not deterred. He didn’t dwell on the mishap, though he called it a clear signal to use these tools more carefully going forward. After the agent failed to restore the bumped gym-goer’s spot, Andrew asked it to draft an email flagging the vulnerability to the booking software’s provider. It wrote the message and sent it to him over WhatsApp for approval. He read it, and told it to send.