Anthropic Says AI Agents Are Beginning to Outperform Humans in Complex Tasks
The AI company says advanced Claude systems can now match or exceed skilled human performance in specific research and engineering tasks, signaling a major shift in the capabilities of autonomous AI agents.

Anthropic says its latest generation of AI agents is beginning to outperform humans in certain complex and well-defined tasks. According to the company, Claude can already match or exceed skilled human performance when executing specific research experiments and solving engineering problems with clearly defined objectives. The development highlights how rapidly AI systems are moving beyond simple chatbots toward more autonomous digital workers.
One of the most significant findings involves speed. Anthropic’s Project Fetch research found that an advanced Claude model was able to complete certain tasks dramatically faster than participating human teams while operating without direct human assistance. The results suggest that AI agents could become particularly valuable in areas where a large amount of digital work needs to be completed quickly and systematically.
Anthropic has also demonstrated that AI systems can outperform humans in some technical evaluations. The company previously reported that Claude Opus models surpassed most human applicants in a controlled technical assessment, while later models were able to match even stronger candidates under the same time constraints. However, Anthropic acknowledges that human experts still retain important advantages, particularly when tasks require broad judgment, creativity and deciding which goals should be pursued.
The growing capability of AI agents could significantly affect industries including software engineering, scientific research and cybersecurity. Anthropic’s research into multi-agent systems has shown that groups of specialized AI agents can outperform individual systems by dividing complex work into multiple parallel tasks. This approach could allow AI to handle increasingly complicated projects that would traditionally require large human teams.
Despite the impressive progress, Anthropic has emphasized that more capable autonomous systems also create new risks. The company recommends extensive testing, sandboxing and safety controls because agents capable of independently making decisions can also produce errors that compound over multiple steps. The race to build increasingly powerful AI agents is therefore becoming not only a competition over capability, but also a major test of whether companies can maintain reliable human oversight.



