Google DeepMind targets validation gap in AI science

AI News


Google Deep Mind to governments and science funders. AI agentprepare a national dataset for use by agents, expand experimentation infrastructure, and equip reviewers with AI tools.

The proposal is in response to what the company describes as a growing “verification bottleneck” in science. Although Agentic AI can generate hypotheses and potential solutions at scale, testing those ideas in the lab and evaluating them through existing research processes remains time-consuming and costly.

This article, “Guessing Machines: AI Agents and the New Verification Bottlenecks in Science,” was written by Don Wallace, Conor Griffin, Sean O’Neill, Thang Luong, and Owen Larter. This is based on discussions with 10 Google DeepMind researchers and engineers.

Luong, chief scientist and research director at Google DeepMind, shared the article on LinkedIn and wrote, “Ideas are becoming more plentiful, but testing in the physical world remains slow and expensive. For policymakers, funders, and researchers to truly unlock superhuman scientific discoveries, they must understand and address this growing gap between generation and validation.”

He identified four priorities for policymakers and funders: expanding access to AI agents, making national data assets agent-enabled, addressing verification bottlenecks, and equipping reviewers with agents.

The central argument is that AI agents are changing the balance between the generation of scientific ideas and their validation. The authors describe these systems as “guessing machines” that can review literature, formulate hypotheses, write code, and coordinate multiple processes, while physical experiments and institutional reviews continue to set the pace.

Ideas generated by AI still require human validation

Google DeepMind cites the Co-Scientist agent as an example of the speed difference. Microbiologist Jose Penadez and his team at Imperial College London spent nearly a decade investigating how families of superbugs spread antibiotic resistance. The company said its co-scientists came up with five possible explanations within two days of Penadez describing the problem, and the top hypothesis matched the conclusion reached by his lab.

The article also describes how Stanford University researcher Gary Peltz used his collaborators to identify existing drugs that could potentially be repurposed for liver fibrosis. Mr. Peltz selected two candidates and his agent selected three. None of his choices showed benefit in assays using live human liver cells, he explains, but two of the drug candidates blocked fibrosis and promoted liver cell regeneration.

These examples are taken from the Google DeepMind article rather than an independent evaluation of the system. The authors also acknowledge that agents remain unreliable and that detecting inaccurate claims can require considerable expertise.

“One hallucinatory claim on page 10 of the output can invalidate the whole thing,” says lead co-researcher Vivek Natarajan.

The company argues that scientific agents should not act as systems that only generate answers, but should clarify inferences, communicate uncertainties, and provide evidence for conclusions. The adjustment of confidence across open-ended scientific reasoning remains an open question.

In mathematics and computer science, output may be verified through code and formal languages, providing more opportunities for automatic checking. In the first First Proof Challenge, held in February, Google DeepMind’s Aletheia system solved 6 of 10 unpublished research questions within a week. This is what the company says is its strongest result.

It poses a problem of different abilities for mathematicians.

“We are moving toward a severe ‘proof dyspepsia’ future, where AI creates breakthroughs faster than humans can review them,” says Luong, who led the Aletheia effort.

Access to funding, data and experimental infrastructure

The first policy recommendation concerns access. Google DeepMind argues that scientific funders should decide how research institutions select and pay their agents, including whether existing research grants are sufficient or new funding programs are needed.

The authors liken the problem to providing researchers with access to supercomputers, calling for public-private partnerships that can deliver AI systems at the scale needed. This article does not discuss expected costs or implementation schedules.

The second recommendation is to make publicly funded data easier for agents to use. According to the authors, open or low-risk datasets should be supported by metadata, quality control, interoperable standards, and accessible through documented APIs.

More sensitive information, such as genomics and virology data, will require privacy controls, auditability, and restrictions governing how agents can access it. The article cites OpenSAFELY, which provides researchers with controlled access to health data, as one possible model.

Our third priority is increasing our ability to test ideas generated by AI. Google DeepMind is calling for investment in existing public facilities, alongside automated laboratories that can conduct experiments using robotics and computing infrastructure.

The company says it has established a wet laboratory at the Francis Crick Institute in the UK, funding independent scientists and providing access to collaborators to test hypotheses generated by the agent.

The article also mentions the US National Science Foundation’s $100m investment in a distributed facility network and the UK’s £81m Materials Innovation Factory. Google DeepMind argues that given the cost of automated labs, governments should consider centralized facilities that researchers can access without requiring individual agencies to fund the full infrastructure.

Science education will also need to be adjusted. The authors warn that junior researchers may be concerned that institutions are directing budgets toward AI use rather than staff, while uncontrolled deployments may prevent novice scientists from developing the judgment needed to evaluate an agent’s output.

They suggest that graduate science training may require structured periods of work without agents, alongside access to systems that support researchers’ thinking, rather than serving as sources of unquestioned answers.

AI tools proposed for peer review

Google DeepMind’s final recommendations are subject to peer review. The authors argue that the increasing use of AI in grant applications and academic papers is making it difficult for funders and reviewers to identify unique ideas, validate research findings, and decide which research is worthy of support.

This article states that some organizations are changing their processes. The UK Medical Research Council has reinstated interviews for shortlisted candidates, but the authors acknowledge that greater reliance on interviews and performance could favor established researchers and increase costs.

They suggest a multi-layered response that combines clearer disclosure about the use of AI with agent access for reviewers. Recommended measures include watermarking, records known as human-AI interaction cards, and the initial use of review agents for more objective tasks such as error detection.

Under the proposal, journals, funders, and conferences would also require AI systems to make inferences, support claims with citations, and document uncertainties where possible.



Source link