Written by Deepa Seetharaman, Supantha Mukherjee, Crystal Hu
SAN FRANCISCO/STOCKHOLM, Dec. 16 – Last spring, wine collection app CellarTracker built an AI-powered sommelier that recommends white wines based on a person's taste buds. The problem was that chatbots were too good.
“It's a very polite way of saying, 'It's unlikely that you'll like that wine,'” said Eric Levine, CEO of CellarTracker. Before introducing this feature, it took six weeks of trial and error to convince the chatbot to provide honest reviews.
Since ChatGPT exploded in popularity three years ago, companies large and small have jumped at the chance to embrace generative artificial intelligence and build it into as many products as possible. But so far, the vast majority of companies have struggled to realize meaningful returns from their AI investments, according to the results of seven recent executive and employee surveys.
A survey of 1,576 executives conducted by research and advisory firm Forrester Research in the second quarter found that only 15% of respondents felt that AI improved profit margins last year. Consulting firm BCG found that of the 1,250 executives it surveyed from May to mid-July, only 5% saw widespread value in AI.
Executives say they still believe generative AI will ultimately transform their businesses, but are reconsidering how quickly that will happen within their organizations. Forrester predicts that in 2026, companies will delay about 25% of their planned AI spending by one year.
“The tech companies that built this technology are spinning a story that everything is going to change very soon,” said Forrester analyst Brian Hopkins. “But we humans don't change that quickly.”
AI companies like OpenAI, Anthropic, and Google are all focusing on appealing to enterprise customers in the coming year. OpenAI CEO Sam Altman said during a recent lunch with media editors in New York that developing AI systems for enterprises could be a $100 billion market.
All of this is happening against a backdrop of unprecedented technology investment in everything from chips to data centers to energy sources.
Whether these investments are justified depends on companies understanding how to leverage AI to increase revenue, increase profit margins, and accelerate innovation. If that fails, some experts say, infrastructure development could cause a collapse reminiscent of the dot-com bust of the early 2000s.
“Easy” button
Shortly after ChatGPT launched, companies around the world established a task force dedicated to finding ways to leverage generative AI, a type of AI that can create original content such as essays, software code, and images through text prompts.
One of the well-known problems with AI models is that they tend to please users. This bias (known as “fawning”) encourages users to chat more, but it can undermine the model's ability to provide better advice.
According to CEO LeVine, CellarTracker encountered this issue with its wine recommendation feature, which was built on OpenAI's technology. When asked for general recommendations, the chatbot performed well. But when asked about a particular vintage, the chatbot remained positive, even though all signals indicated the person was highly unlikely to enjoy that vintage.
“I had to bend over backwards to be critical of the model (any model) and suggest there were wines I might not like,” Levine said.
Part of the solution was to design a prompt that gave the model permission to say no.
Businesses also suffer from a lack of consistency in AI.
Jeremy Nielsen, general manager of North American rail services company Cando Rail & Terminals, said the company recently tested an AI chatbot to help employees study internal safety reports and training materials.
But Cando encountered a surprising obstacle. The model was unable to consistently and accurately summarize the Canadian Rail Operating Regulations, a roughly 100-page document that sets safety standards for the industry.
In some cases, the models forgot or misunderstood the rules. I even invented it from whole fabric. AI researchers say models often have trouble remembering things in the middle of long documents.
Cando has canceled the project for now, but is testing other ideas. The company has spent $300,000 developing its AI product so far.
“We all thought it was an easy button,” Nielsen said. “And that's not what happened.”
humanity will be revived
It was thought that AI would significantly disrupt human-staffed call centers and customer service, but companies quickly realized that there was a limit to the amount of human interaction they could delegate to chatbots.
In early 2024, Swedish payments company Klarna announced it was deploying OpenAI-powered customer service agents to support the work of 700 full-time customer service agents.
But in 2025, CEO Sebastian Siemiatowski was forced to change things back when he acknowledged that some customers preferred talking to a human.
Simiyatowski said the AI is reliable for simple tasks and can currently do the work of about 850 agents, but more complex problems will soon fall to human agents.
Looking ahead to 2026, Klarna is focused on building a second generation of AI chatbots, expected to ship soon, but humans will still be a big part of it.
“If you want to stick with your customers, you can’t trust them.” [entirely] It’s about AI,” he said.
Similarly, US telecom giant Verizon is moving back to human customer service agents in 2026 after experimenting with delegating calls to AI.
“I think 40% of consumers like the idea of being able to talk to a human, but are frustrated by not being able to reach a human agent,” Ivan Berg, who leads Verizon's AI-powered efforts to enhance service operations for enterprise customers, told Reuters in an interview this fall.
The company has about 2,000 front-line customer service agents who still use AI to screen calls, obtain information about customers and direct them to self-service systems or human agents.
Using AI to handle routine questions frees up agents to tackle complex problems and tackle new things like making outbound calls and sales.
“Empathy is probably what currently prevents AI agents from comprehensively interacting with customers,” Berg says.
Shashi Upadhyay, president of products, engineering and AI at customer service platform Zendesk, says AI excels in three areas: writing, coding and chatting. Zendesk customers rely on generative AI to handle 50% to 80% of their customer support requests. But the idea that generative AI can do anything is “overrated,” he says.
“Jagged Frontier”
Although large-scale language models are rapidly conquering complex tasks in mathematics and coding, they can still fail at relatively simple tasks. Researchers refer to this conflict of capabilities as AI's “jagged frontier.”
“It may be a Ferrari in mathematics, but it's a donkey in terms of putting things on the calendar,” says Anastasios Angelopoulos, CEO and co-founder of LMArena, a popular benchmarking tool.
Seemingly small problems can unexpectedly cripple AI systems.
Many financial companies rely on data collected from a wide range of sources, all of which can vary widely. These differences can lead to AI tools “reading patterns that don't exist,” said Clark Shafer, director at advisory firm Alpha Financial Markets Consulting.
Schaefer said many companies are now considering reformatting their data to leverage AI, a process that can be expensive, lengthy and complex.
Dutch technology investment group Prosus said one of its in-house AI agents will aim to answer questions about portfolios, similar to what the group's data analysts already do.
In theory, employees could ask how late the Prosus-backed food delivery company was in delivering the sushi they ordered in Berlin last week.
But at the moment, the tool doesn't necessarily understand which areas of Berlin are part of, or what “last week” means, said Euro Beinat, head of AI at Prosus.
“People thought AI was magic, but it’s not magic,” Beinat said. “For these tools to work properly, a lot of knowledge needs to be encoded.”
more on hand
OpenAI is working on new products for enterprises, as well as recently formed in-house teams (such as the Forward Deployed Engineering team) to work directly with customers to help them leverage OpenAI's technology to tackle specific problems, the spokesperson said.
“What we're failing at is people jumping in too big. They realize they have a billion-dollar problem. It's going to take years,” Ashley Kramer, head of revenue at OpenAI, said in an on-stage interview at the Reuters Momentum AI conference in November.
Specifically, OpenAI is working with companies to look for areas where AI “will have a big impact, but may have a low impact initially,” Kramer said.
Rival AI lab Anthropic, which derives 80% of its revenue from enterprise customers, is hiring “applied AI” experts to be embedded in companies.
For AI companies to be successful, they need to see themselves as “partners and educators, not just technology adopters,” Mike Krieger, head of product at Anthropic, said in an interview earlier this year.
A growing number of startups are developing AI tools for specific sectors, such as financial services and law, many of which were founded by former OpenAI employees. These founders say businesses would benefit more from a specialized model than a general-purpose or consumer-oriented tool like ChatGPT.
This is the playbook adopted by Writer, a San Francisco-based AI application startup. The company is currently building AI agents for finance and marketing teams at large companies like Vanguard and Prudential, putting engineers in touch with clients directly to understand their workflows and co-build the agents.
“Companies need more collaboration to make AI tools actually work for them,” said May Habib, CEO of Writer.
(Reporting by Deepa Seetharaman and Crystal Hu in San Francisco and Supantha Mukherjee in Stockholm; Editing by Kenneth Lee and Michael Learmonth)