OpenAI's latest innovation is called CriticGPT, and it was built to identify bugs and errors in the output of artificial intelligence models, as part of an effort to ensure AI systems behave the way their creators want them to.
Traditionally, AI developers use a process called Reinforcement Learning with Human Feedback (RLHF) in which human reviewers evaluate the output of large language models to make them more accurate, but OpenAI believes that LLM can actually help with this process of essentially critiquing the output of AI models.
In a research paper titled “LLM Critics Help LLM Find Bugs,” OpenAI researchers said they built CriticGPT to assist human AI reviewers in the task of reviewing code generated by ChatGPT. CriticGPT was built using GPT-4 LLM and has shown promising abilities in analyzing code and identifying errors, allowing it to spot AI “hallucinations” that human colleagues may not notice on their own.
OpenAI researchers say they trained CriticGPT on a dataset of code samples that were intentionally peppered with bugs, teaching it how to recognise and flag a range of coding errors in software.
According to OpenAI, during the training process, human developers corrected the code written by ChatGPT, introducing various errors and providing sample feedback, just as they would if the bugs were real and the developers had stumbled upon them by chance. This allowed CriticGPT to identify not only the most common coding errors, but also some that are less common.
After training CriticGPT, OpenAI tested it, and the results were impressive. CriticGPT outperformed the average human code reviewer. Human trainers preferred CriticGPT's critiques over human-written critiques in 63% of cases. According to OpenAI, this was in part because CriticGPT generated fewer useless “niceties” about the code and fewer false positives.

To advance their research, the OpenAI team developed a new technique they call “Force Sampling Beam Search,” which allows CriticGPT to write more detailed critiques of AI-generated code, and also gives it more flexibility by allowing human teachers to tune the thoroughness of CriticGPT's bug search and better control its tendency to hallucinate or highlight “errors” that aren't actually there.
CriticGPT's thoroughness allowed it to significantly outperform humans. The researchers decided to apply CriticGPT to the ChatGPT training datasets, which had been marked as “perfect” by human annotators, meaning they should not have a single bug. However, CriticGPT found bugs or errors in 24% of these datasets, which were later confirmed by human reviewers.
According to OpenAI, this demonstrates CriticGPT's ability to identify even the most subtle mistakes that humans would miss even after thorough evaluation.
However, it is worth pointing out that CriticGPT, like all their supposedly perfect training datasets, still comes with some caveats. For one, it was trained using relatively short responses from ChatGPT, so it may struggle to evaluate the much longer and more complex tasks that represent the next evolutionary stage of generative AI. Additionally, CriticGPT is not yet able to surface all errors and in some cases hallucinates, creating false positives where human annotators may make mistakes when labeling the data.
One challenge CriticGPT must overcome is to be more effective at detecting incorrect output caused by errors in a specific piece of code. However, some AI hallucinations are the result of errors across multiple different code strings, making it much harder for CriticGPT to pinpoint the source of the problem.
Still, OpenAI is encouraged by the progress so far and plans to integrate CriticGPT into the RLHF pipeline, meaning human trainers will have their own generative AI assistant to help review generative AI output.
Image: SiliconANGLE/Microsoft Designer
Your vote of support matters to us and helps keep our content free.
With just one click below you can support our mission of providing free, rich, relevant content.
Join the YouTube community
Join a community of over 15,000 #CubeAlumni experts, including many notable figures and experts, such as Amazon.com CEO Andy Jassy, Dell Technologies founder and CEO Michael Dell, Intel CEO Pat Gelsinger, and many more.
thank you
