Shane Mac had been using a personal AI agent since August when it accidentally shared his private bank statement in a work Slack chat. The Grok Bot agent mistakenly posted his financial details to the wrong group. Mac said, "My heart kind of dropped," after seeing his personal information exposed.
Shane Mac had been using a personal AI agent since August without issue when the startup chief executive got pinged by one of his employees: a "heads up" about a team Slack message that had a screenshot attached.
There before him was a literal snapshot of his checking and savings account balances, various construction costs for a house project, and recurring charges for Netflix and car insurance, courtesy of his agent.
"My heart kind of dropped," says Mac, who lives in Nashville, Tenn. "I thought I had things set up as separate, personal and work."
Mac had been using Grok Bot, SpaceXAI's "always-on agent," giving it read-only access to his bank accounts and naming the agent "Personal CFO." He maintains a separate Grok Bot agent linked to his work Slack, which he named "Chief of Staff." Mac demanded to know how Personal CFO had ended up posting his info to the company Slack.
The bot explained itself, then sort of apologized: "That's on me," it said. "Wrong audience for personal cash -- period."
People are being wowed by the capabilities of their personal AI agents before being quickly brought back down to earth. They have taken off in recent weeks, managing users' email correspondence, travel itineraries and more.
But some users find that much like the models they're built on -- which can hallucinate, misinterpret or even commit spontaneous acts of cybercrime -- their assistants, too, can behave unpredictably.
The difference between a regular chatbot and an agent is that users typically ask agents to carry out tasks, like organizing their calendar or booking an appointment. The agent then breaks this prompt down into a series of steps, which are for convenience often not spelled out to the human it's working for.
Agents typically know to ask permission before doing major things like making purchases or deleting files, but other than that, they can have free rein. If a user's prompt is unclear, or if there are many ways to execute the requested action, mayhem might ensue.
In Mac's case, his agent conflated chat groups with similar names. "It handled an ambiguous instruction badly," he wrote in a follow-up post. He said the Grok Bot team told him it fixed the problem.
Annica Benning, a San Francisco-based communications professional, has relied on an agent made by the small buzzy startup Instinct for the past month or so. Users can communicate with their Instinct agent from directly within their iPhone's Messages app. Benning has used hers to shop online, schedule Ubers and book flights.
She says it functions perfectly 90% of the time, yet it recently surprised her by buying a bikini. She'd returned a swimsuit and had asked the agent to present her with some alternatives from the site. The agent sent her a link to approve its selection. She says she never approved it. Yet she received an email confirmation of the purchase.
"Anything you do with a credit card, it sends you an approval link," she says. "But because this was store credit, not a credit card, it was able to bypass the approval link."
She couldn't cancel the order. The bikini will arrive any day now. She's hoping it's returnable.
While traveling in New York a few weeks back, Benning had Instinct handle her dinner reservations.
"It said, 'I went ahead and booked this,'" she says. The problem was, she had other plans that night -- which her agent knew about.
Benning then had to ask the agent to cancel the reservation and ask the restaurant not to charge her the $50-a-head no-show fee. (The restaurant, via her agent, agreed.)
Shalini Dinesh, who lives in Austin, had already been an aficionado of Anthropic's Claude when all the Meta Muse fanfare erupted online. She decided to try it out by letting Muse coordinate her son's 13th birthday party.
The party was at a paintball venue, so the assistant drafted and sent a text to each parent with a release form to sign. It explained what clothing the kids needed to wear.
It also found nearby restaurants that could deliver food, taking into account any allergies of the attendees. And it picked out age- and theme-appropriate goody bags for the kids to take home.
Then it bungled the seemingly easiest part of the job, the RSVP list. Because two invited children had the same first name, the agent smooshed them together as one child, and only one child was listed as coming. Luckily, Dinesh was able to double-check the list and correct Muse.
"It was looking for exact keywords," she says. It had to be a clearly affirmative or negative response. "There is no 'Oh great, see you there.'"
Jesse Levey, founder of Bay Area startup Longevity Health, tried getting Muse to manage the activities schedules of his three kids. He started manually entering a typical week from his calendar into Muse but then realized it would be too much of a pain with all the niche apps that teams use and the carpool shifts that tend to happen in chat groups.
"My kids are on the ski team and there's this ski app where the calendar is kept," he says. "You've got text and WhatsApp and email and calendars and maybe like 65 different apps that it has to have access to."
He eventually asked Claude if there was a more powerful agent that could handle his children's schedules -- one that still addressed his privacy and security concerns.
Its response surprised him: It suggested a "household manager" -- a part-time human employee -- to handle "camp registration races, sick-day improvisation and the 10 judgment calls a week that no system will get right."
An AI version was possible, Claude said, but "the build will take months of your attention."