OpenAI debuts GPT-5.2 to address concerns about lagging behind rivals

Applications of AI


Amid increasing competitive pressure from Google and Anthropic, OpenAI debuts a new AI model, GPT-5.2, which it says significantly outperforms all existing models on a wide range of tasks.

The new model, released less than a month after OpenAI debuted its predecessor GPT-5.1, performed particularly well on benchmarks of complex professional tasks across a variety of “knowledge tasks” from law to accounting and finance, as well as assessments involving coding and numerical reasoning, according to data published by OpenAI.

Fidji Simo, former CEO of InstaCart and current CEO of applications at OpenAI, told reporters that this model should not be seen as a direct response to Google's Gemini 3 Pro AI model released last month. With this release, OpenAI CEO Sam Altman issued a “code red” and delayed the rollout of several initiatives in order to focus more staff and computing resources on improving the core product, ChatGPT.

“That’s what I would say. [the Code Red] “It helps with the release of this model, but that's not specifically why it's being launched this week, it's been in development for a while,” she said.

She said the company has been building GPT-5.2 “for many months.” “These models are not completed in just one week; they are the result of a lot of work,” she said. According to an article in The Information, the model was known internally by the code name “Garlic.” The day before the model's release, Altman teased its impending release by posting a video clip on social media of him cooking a dish with lots of garlic.

OpenAI executives said the model had been in the hands of an “alpha customer” to help test performance for “several weeks.” This period means the model was completed before Altman's “Code Red” declaration.

In addition to Shopify and Zoom, these testers included legal AI startup Harvey, note-taking app Notion, and file management software company Box.

OpenAI said these customers found that GPT-5.2 exhibited “state-of-the-art” capabilities to complete tasks using other software tools, and was also great at writing and debugging code.

Coding has become one of the most competitive use cases for deploying AI models within the enterprise. While OpenAI had an early lead in this space, Anthropic's Claude model has been particularly popular among enterprises, surpassing OpenAI's market share by some numbers. There's no doubt that OpenAI wants to convince customers to go back to its model when coding with GPT-5.2.

Simo said “Code Red” is helping OpenAI focus on improving ChatGPT. “Code Red is really a signal to the company that they want to focus their resources on certain areas, and it's a way to really define their priorities and define what can be de-prioritized,” she said. “As a result, we now have more resources focused on ChatGPT in general.”

The company also said the new model is better than the company's previous model in providing “secure answers.” The company defines this as providing helpful answers to users without making statements that could contribute to or exacerbate a mental health crisis.

“On the safety side, as you can see through the benchmarks, we're seeing improvements in almost every aspect of safety, including self-harm, different types of mental health, and emotional dependency,” Simo said. “We are very proud of the work we do here. This is a top priority for us and we only release models when we are confident that safety protocols are followed and we take pride in our work.”

The release of the new model comes on the same day that a new lawsuit was filed against ChatGPT, alleging that its interactions with psychologically challenged users contributed to a murder-suicide in Connecticut. The company also faces several other lawsuits alleging that ChatGPT contributed to people's suicides. The company called the Connecticut murder-suicide case “incredibly heartbreaking” and said it is continually improving ChatGPT's “training to recognize and respond to signs of mental or emotional distress, de-escalate conversations, and direct people to real-world help.”

GPT-5.2 showed significant performance improvements in several benchmark tests of interest to enterprise customers. As measured by OpenAI's GDPval benchmark, it matched or exceeded the performance of human experts on a wide range of difficult specialized tasks 70.9% of the time. In comparison, GPT-5, the model that OpenAI released in August, has only 38.8%. Anthropic's Claude Opus 4.5 is 59.6%. Google's Gemini 3 Pro had 53.3%.

In the software development benchmark SWE-Bench Pro, GPT-5.2 received a score of 55.6%. This is nearly 5 percent better than the previous generation GPT-5.1 and more than 12 percent better than Gemini 3 Pro.

Aidan Clark, OpenAI's vice president of research (training), declined to answer questions about what specific training methods were used to upgrade GPT-5.2's performance, but said the company has made improvements across the board, including pre-training, the first step in creating an AI model.

When Google released its Gemini 3 Pro model last month, the company's researchers also said the company had made pre- and post-training improvements. This surprised some stakeholders who believed that AI companies had all but exhausted their ability to derive significant improvements from the pre-training stage of model building, leading to speculation that OpenAI may have been caught off guard by Google's advances in this area.



Source link