Artificial intelligence is on the verge of scoring a perfect score on a test aimed at measuring the gap between machine learning and human intelligence, dubbed “humanity’s last test.”
Google’s Gemini model scored a whopping 45.9 percent on last month’s test, marking a staggering jump from the performance of rival systems just two years ago.
When OpenAI’s ChatGPT first attempted testing in 2024, it achieved only 3% accuracy, with competitors from Google and Anthropic slightly better.
This rapid progress has led researchers at Scale, the company that hosts the benchmark tests, to predict that AI could reach a perfect score within about 12 months.
The exam consists of 2,500 carefully selected questions across approximately 100 subject areas, from rocket science and mythology to physiology and ancient languages.
Scale and the nonprofit Center for AI Safety have developed a test that examines both the breadth of knowledge and the depth of reasoning ability in AI systems.
To compile the questions, organizers announced a global appeal in September 2024, offering a $500,000 prize to experts who can submit questions that are difficult to answer through internet searches.
The response was strong, with around 70,000 potential questions submitted by experts from around 50 countries.
Artificial intelligence is poised to score perfect scores on one of the world’s most difficult tests, experts reveal
|
getty
After eliminating queries that could be solved by existing models, the list was reduced to 13,000 before final selection.
Each question requires at least a doctoral level of comprehension, and anyone who scores close to a perfect score qualifies as a “universal expert.”
Calvin Chan, head of research at Scale, explained the ambition behind the project: “We wanted to create this close-ended academic benchmark set at the forefront of human expertise, something that only a handful of humans on the planet can actually solve.”
He praised the developers working on the language model, saying: “We’ve seen tremendous progress over the past few years with these language models, which is impressive. Model builders have done a really great job in improving these inference models.”
Google’s Gemini model scored an impressive 45.9 percent on last month’s exam.
|
getty
Kate Olszewska, Product Manager at Google DeepMind, expressed confidence that by focusing resources on goals, milestones can be reached quickly.
“I think if we really value this as the only thing in life, we’ll get to it pretty quickly,” she told the Daily Mail.
Anthropic’s Claude system, on the other hand, achieved 34.2 percent on the exam and quickly improved its scores.
Dr. Tung Nguyen, a professor of computer science and engineering at Texas A&M University who contributed 73 questions to the exam, offered a more cautious assessment of progress.
“Humanity’s last test will be one of the clearest assessments of the gap between AI and human intelligence,” experts explained
|
getty
“Humanity’s last test is one of the clearest assessments of the gap between AI and human intelligence,” he said.
While acknowledging the superior performance of certain models, Dr. Nguyen argued that the weak results of other models indicate that significant gaps remain.
“When AI systems start performing very well on human benchmarks, it is tempting to think that we are approaching human-level understanding,” Dr. Nguyen said, adding, “But HLE reminds us that intelligence is about depth, context, and expertise, not just pattern recognition.”
He emphasized that the goal of benchmarking is not simply to defeat AI, but to highlight where human expertise remains essential.
