Technology

OpenAI and Anthropic Push New Metrics to Measure AI Progress

The leading AI companies are increasingly looking beyond traditional benchmarks, focusing on real-world productivity, coding performance, agent capabilities and the cost of completing complex tasks.

OpenAI and Anthropic are part of a broader shift in how the artificial-intelligence industry measures progress. Instead of relying only on conventional benchmark scores, AI developers are increasingly examining whether models can successfully complete meaningful, multi-step tasks in realistic environments. This matters because a model that performs well on a standardized test may not necessarily be the most useful system for businesses or developers. Current evaluation efforts increasingly emphasize practical capability and efficiency.

One important metric gaining attention is the amount of work an AI agent can complete reliably. Organizations such as METR measure this using a “time horizon,” which estimates how long a human expert would normally need to complete a task that an AI can successfully handle. This approach provides a different perspective on AI progress because it evaluates sustained task completion rather than a model’s ability to answer isolated questions.

Coding has become another major area for evaluating AI systems. Modern models are increasingly being tested on whether they can work through real software projects, identify problems, write code and complete extended development tasks. Recent benchmark updates show large differences between frontier models, demonstrating why developers are paying greater attention to the specific type of work a model can accomplish rather than simply looking at one overall leaderboard position.

Cost and efficiency are also becoming increasingly important. AI companies are competing not only to build more capable models but also to make them cheaper and more efficient to operate. OpenAI, for example, has emphasized performance per dollar for its GPT-5.6 models, while the wider industry is experiencing growing pressure from lower-cost competitors. This means an AI system that delivers slightly lower benchmark scores but completes tasks at substantially lower cost can still be highly competitive.

The shift toward practical metrics could ultimately change how the AI race is understood. The most important question may no longer be simply “Which model scores highest?” but rather “Which model can reliably complete the most valuable work at the lowest cost?” As OpenAI, Anthropic and other AI developers continue refining their evaluation systems, real-world productivity, autonomous task completion, coding ability, safety and economic efficiency are likely to become increasingly important measures of AI progress.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button