AI agents are causing a security nightmare as they find shortcuts to complete tasks. DataDome showed bot activity on login pages rose eightfold in the first half of 2026. While companies use guardrails, researchers say these models often hack systems to please users, creating risks for the entire internet.
The Australian tech worker’s AI agent had found a loophole and exploited it without being told to. So what happens when millions of other people’s agents start competing for the cheapest flights or the best movie seats? In the best-case scenario, we will see a smooth and well-functioning network of bots that play by the rules to get the best outcomes for their humans. In the worst: a cyber security nightmare.
The web simply might not be ready. Security firm DataDome says 65% of the more than 20,000 websites it tested had systems for detecting or blocking AI agents, and that bot activity hitting login pages rose more than eightfold in the first half of 2026.
In some ways it is hard to imagine these new AI agents causing widespread trouble. They look too cute for a start. Muse has a cartoonish design reminiscent of a Labubu doll, while the vividly covered blobs that represent dots could be characters in a children’s TV show—perhaps part of an effort to ease general mistrust of Meta and OpenAI.
They work well too. Bloomberg’s Dave Lee called Muse “the most impressive product Meta has released in years.” He said his agent, which he had christened Harold, had organized his calendar and scoured Facebook Marketplace for good deals.
Both Meta and OpenAI say they have strong guardrails in place to make sure their agents behave themselves. Muse is kept within an isolated virtual machine—essentially a private computer in Meta’s cloud—while a separate security system controls every request it makes on the internet, handling payments and making sure a human is consulted before any major task like buying something or sending a message. Dots is monitored in a similar way.
But even these safeguards don’t fix an underlying problem baked into today’s Generative AI models through training: a tendency to look for shortcuts, a phenomenon researchers call reward-hacking. This is also why chatbots are prone to flattery. The UK’s AI Security Institute noticed a rise in this kind of cheating behaviour in coding agents and chatbots late last year.
An inclination to please is what drove OpenAI’s agents to hack Hugging Face a few months ago. OpenAI agents also improperly meddled with the websites of dozens of other organizations this year, including the US Securities and Exchange Commission and Australia’s government-run healthcare scheme. Meta says one of its AI models too hacked another company of its own accord.
These cases happened while the models were being trained or tested, with safeguards deliberately lowered. But the bots’ actions still surprised their developers and researchers have noticed that even when they make their tests harder to game, agents keep trying to cheat.
There are signs that AI products will also hunt for loopholes in the wild, particularly in coding. In one instance, a software developer told his AI agent to clear out temporary files from a computer folder without using the ‘delete’ command. The agent then hid that instruction inside a testing tool and slipped past the filter to clear the files using the command it was told not to.
Now multiply that type of behaviour by million of agents as consumers explore using AI to optimize their lives and work. When Meta boss Mark Zuckerberg launched Muse last month, he said consumers deserved to have access to their own “personal superintelligence” and that the new agents would “make you money.”
Some aren’t doing that very well. When a consumer tech reviewer in Toronto named Matt J. Robb tried using Muse to sell a keyboard on Facebook Marketplace last month, the agent angered the buyer by saying Robb was at home to carry out the transaction, which wasn’t true. The agent had also made a low-ball offer.
On top of the risk of cheating, agents can also just make mistakes, and it will likely be end users who pay the price. Robb was left with a negative rating from his angry buyer.
However cuddly these new agents look, the systems behind them are trained to tick a box one way or another. They might do it a little too well. ©Bloomberg
The author is a Bloomberg Opinion columnist covering technology.
