Frontier AI brings face plants to real-world workplace tasks

Machine Learning


To date, AI industry spending has exceeded $1.6 trillion and shows no signs of slowing down anytime soon.

So what do you actually have to show for that? Historically, it’s been nothing at all. As many studies have shown, tools such as AI chatbots and autonomous agents have been ineffective at completing real-world tasks in a competent manner.

The technology industry claims that everything will change in the next few years, as AI capabilities grow exponentially and enable economic growth like the world has never seen before. But is that so? Really?

Not necessarily. New research from the Center for Responsible Distributed Intelligence at the University of California, Berkeley — university fix — Frontier AI tools for every make and model still They are unable to complete the majority of workplace tasks to an acceptable level, putting a major spin on the tech industry’s claims that an AI revolution is imminent.

To reach this conclusion, UC researchers designed a rigorous evaluation they called the “Agent Final Test.” It was developed to test “employment readiness” across a number of cutting-edge AI models. Essentially, the ALE (a cheeky paraphrase of “Humanity’s Last Exam”) is designed to put AI systems through its paces, covering “more than 1,500 expert tasks across 55 professions,” the researchers wrote in a press release.

These tests cover typical jobs exposed to AI, such as software engineering and graphic design, but also a significant number of jobs whose fates are still uncertain, such as maritime engineering, agriculture, audio production, and public health work.

Researchers used the ALE benchmark to take a closer look at advanced “closed” models (proprietary AI systems developed by private companies), including Anthropic’s Fable 5, OpenAI’s GPT-5.5, Cursor’s Composer 2.5, and Google’s Gemini 3.1 Pro. (Just to be sure, they also looked at two open source models from Chinese developers.)

Although these AI models are cutting-edge, research has found that they are far from meeting the complex needs of the modern workplace. Of all the models that passed the gauntlet, each one failed spectacularly. OpenAI’s GPT-5.5 received the highest score, with an overall pass rate of just 24%.

“Today’s agents can solve a significant portion of specialized tasks,” the researchers wrote. “But when you look at the most difficult tasks that require sustained inference, deep expertise, and reliable execution over time, they are still far from human-level performance.”

And as the tasks became more complex, even those meager total scores rapidly declined.

“At the most difficult tier of ALE, all frontier agents we tested, including Fable 5, had a 0% success rate,” the reporter explains.

The researchers also detail cost considerations. They say the cutting-edge Fable 5 delivers “comparable performance” to models such as GPT-5.5 and Composer 2.5, but “costs approximately 4 to 12 times more per completed task.”

Despite the dire test results, researchers warn that the technology still has the potential to transform the job market as it becomes more exposed to AI. As many corporate executives have shown us, the technology doesn’t have to work particularly well to push workers.

“Even if current pass rates remain relatively low, occupations dominated by routine, well-defined procedures are likely to experience disruption first, while roles that require decision-making will remain resilient over time,” said Dawn Song, a computer science researcher at Berkeley and study co-author. university fix.

“The key factor is not the industry itself, but the nature of the work,” Song added.

Learn more about AI: OpenAI appears to have significantly missed its sales goals



Source link