Anthropic’s Claude Opus 5 outperformed human accountants in a new study by Mercor. The AI completed complex tasks with perfect accuracy in under 10 minutes, while humans averaged 37 percent. Elon Musk shared the results on X, saying, “Super Intelligence (fka AI) is now acing accounting tests.” Costs also fell.
AI models are now outperforming human accountants on certain accounting tasks, according to a new study from AI recruiting company Mercor. The research found that Anthropic’s Claude Opus 5 completed a set of accounting tasks with perfect accuracy while taking less than 10 minutes for each attempt.
The findings have also caught the attention of Elon Musk, who shared a post about the results on X. “Super Intelligence (fka AI) is now acing accounting tests,” Musk wrote.
The study, titled Human Baselines for Benchmarks "AI Now Outperforms Junior Accountants, compared several AI models with 12 accountants." The participants were licensed CPAs with an average of around five and a half years of accounting experience.
Claude Opus 5 scores 100 percent on the tasks
The researchers asked the accountants to complete four month-end accounting scenarios based on realistic company files. The tasks required them to search through documents, identify the relevant figures, perform calculations and provide the results in a table.
The human participants recorded an average accuracy of 37 percent on the tasks. Individual scores varied widely, from 0 percent to around 90 percent, while most attempts took between 30 and 180 minutes.
Claude Opus 5, meanwhile, scored 100 percent across all 20 attempts reported in the study. Each task was completed in less than 10 minutes.
The study also found a significant difference in the cost of completing the tasks. Based on the US median accountant wage, the researchers estimated that human accountants cost $10.35 per rubric criterion completed, compared with $0.21 for Claude Opus 5.
The gap between AI and humans has also emerged quickly. According to the study, the best AI models were still scoring below the average accountant roughly 18 months ago. OpenAI’s o3 model later crossed the 37 percent human average, while GPT-5 reached about 69 percent. Recent frontier models, including Opus 5, have reached scores close to or at 100 percent.
Can AI replace accountants?
Mercor cautions against treating the results as proof that AI can perform an accountant's entire job. The researchers say the test focused on tasks that AI models are particularly good at, such as carefully following instructions, searching through large amounts of information and handling detailed requirements.
Real accounting work also involves activities that were not tested, including communicating with clients, asking questions, working with colleagues and using knowledge built up over time.
The researchers also note that the study deliberately used difficult accounting scenarios originally designed to challenge AI systems. A small missed detail could significantly affect the final score.
Still, Mercor says the results point to possible productivity gains as AI takes over more structured accounting work. The researchers argue that even if AI progress stopped at its current level, these capabilities could bring changes to how accounting tasks are handled in the coming years.
