Deep vs Interactive Research: How important is human surveillance in AI research?

Machine Learning


One of aiThe most prominent use case is deep research. This process involves use AI Agent Search the web for information, automatically and thoroughly compile that information, and use thoughtful analysis and citations to thoroughly report. It is also very competitive. Several AI companies have become family names including Openai, Gemini, Claude, Glockconfused and deepseekeverything provides its own automated research tools.

Of course, modern AI can also be used interactively. Loop man the study. This allows humans to use AI to quickly investigate and answer questions and ask new questions based on those answers. At the individual level, this explains many ChatGPT sessions. At the enterprise level, large-scale research can be performed using a herd of AI agents. One example is to analyze every company in the S&P 500 and answer each specific question.

But how good are these automated research tools? Which of LLMS Is it perfect for interactive research? And is automated or interactive research better?

Until recently, these questions were not answered. Since the introduction of the latest AI, quantitative benchmarks have been used to compare different models and approaches, but so far no benchmarks have been developed to identify which tools are best. Here's the benchmark Deep Search Bench, A rigorously created benchmark that evaluates how AI systems effectively handle multi-step web-based research questions. Rather than a quick fact-seeking query, the task reflects the troubling, open-ended research that real-world analysts, policymakers and researchers take on a daily basis.

Deep research vs. Interactive research: What's the difference?

  • Deep Research includes autonomous LLM agents that perform their own searches, use multiple tools, use self-reliability, and (theoretically) persist or adapt to failures. This process does not involve human steering when activated.
  • Interactive research is a human loop-driven process in which humans guide smaller automated steps, inspect traces of inference, and repeatedly refine or re-execute the analysis parts.

From Dan SchwartzThe hallucinations of AI are getting worse. What can we do about it?

How good is your deep research?

Our study discovered which LLMs are best suited for AI-assisted research and included evaluations of a series of commercial web research products.

  • Openai Deep Research
  • Gemini Deep Research
  • Grok DeepSearch
  • A deep, confusing study
  • Claude Research
  • DeepSeek with web search

Of these, Openai Deep Research and Gemini Deep Research were the most effective.

Graphics showing the performance of various deep research tools
Image: Screenshots by the author.

Two important notes:

  • This is a relative ranking. None of these tools were consistently successful for a considerable majority of individual tasks within the benchmark. If you want accurate, rigorous quantified accuracy, AI deep search tools aren't where you can be completely successful. They are very good at getting a wide overview when all the little details are there or correct is not essential – incredibly valuable to fairness – but they are not good when mistakes are a big problem.
  • The most performant commercial tool is not a deep research tool, but ChatGpt-O3 in web searches, which tends to be higher than Openai's deep research, and tries to carefully verify the answers before completion. O3 inference traces include “What if this initial study is wrong? How can I reconfirm it?” and also contains secondary sources more frequently than deep studies.

Some users may prefer to use certain deep research tools as they require additional background information rather than 100% accuracy, but users may prefer longer treatments for unfamiliar subjects. For example, they may have a new position and want to keep up as much as possible about relevant subjects than the conclusions of the study.

Current deep research tools will not fit if strict and reliable accuracy is required. Here's the reason for this:

  • Simple Hallucinationsit remains the problem (and has even numbers) It's increased (In recent models such as the O3 and Deepseek R1).
  • Deceiving (misleading trustworthy for unreliable sources). It's not uncommon for LLM to interpret claims with random blog posts Another We consider random blog posts to be negative without bothering us to track the original high quality source of the bill. Satire and parody are even more problematic. Google Search recently joked, “The final stages of your PhD include fighting snakes.” Seriously. This is a very outlier and clear, but it shows the problem. A simple misconception can also add confusion to the mix. For example, look at a reference to “February” in the March 2025 report and misinterpret it as February 2024.
  • Errors that start at small scale can cascade and become more important. Early errors lead to agents making increasingly inadequate decisions as they advance their research workflow.

More about the future of AIHow will AI change the world by 2050?

Loop still needs humans

It is clear that automated AI research tools are very impressive and rapidly improving. It is also clear that human monitoring and error checking will be necessary for rigorous and accurate in the near future. This is especially true given the recent revival of hallucination rates in frontier models.

The inevitable conclusion is that when accuracy is important, humans remain significantly better than existing deep research tools as they evaluate AI output and make repetitive decisions. This is true as long as humans are superior to AI.

  • Error check: Identifying hallucinations, cheating, misunderstandings, and other AI errors. and
  • Ideas: Interpret the answer and use the new information to dynamically determine the direction the research should take.

We humans are far better at both of these now. The gap between human and AI error checking is expected to close significantly over the next few years. However, hallucinations are persistent and difficult to judge. Because missed errors are cascade-risk, closing the gap takes longer than frontier companies expect.

A more interesting question is how long the gap between human and AI ideas will last. The heart of any research process is that, given the new set of answers, you know the best questions to ask. This is much closer to Elusive concept of Agi. The meaning is that even if LLM is as good at error checking as humans, human-supported interactive research remains superior to fully automated, deep research… until and unless AI is as intelligent as humans in general.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *