AI “token maxing” wanes as workplaces seek to reduce technology spending

AI News


The corporate craze of “token maxing” using artificial intelligence technology is reaching its limits, and workplaces that are injecting AI into everything are not seeing similar productivity gains and costs are rising.

What started as tech industry spring hype to squeeze as much AI-generated work as possible out of products like OpenAI’s ChatGPT and Anthropic’s Claude has turned into a summer backlash.

“With AI, it’s very easy to create things you don’t need,” said Vincent Gusdorf, head of AI analysis at Moody’s Ratings and author of a new report recommending a more disciplined approach.

“Tokenmaxxing” refers to maximizing the usage of tokens. Tokens are building blocks of generative AI that correspond to small pieces of text that the AI ​​system reads and writes. Each token is about three-quarters of a word. Additionally, there is usually a limit to how many you can use, with more expensive versions of AI products having higher limits.

“As the bills started to pile up, people realized that these new tools were very expensive and needed to be used wisely,” Gusdorf says.

Just a few months ago, Silicon Valley executives were encouraging the mass consumption of tokens as a sign of high employee performance. A typical token maxer would stay up late, perhaps ignoring their lover, while assembling an army of round-the-clock AI agents to perform tasks on their behalf.

“We’re looking forward to seeing what happens to token-maxing startups, both in terms of how they work internally and the products they can build,” OpenAI CEO Sam Altman said in May.

Nvidia CEO Jensen Huang said, “If a $500,000 engineer isn’t spending $250,000 in tokens, something is wrong.” Meta, Facebook’s parent company, had internal competition based on token usage.

While this trend boosted revenues for major AI large-scale language model developers such as Anthropic and OpenAI, it proved to be not necessarily the best strategy for other companies, and revenues slumped.

Microsoft CEO Satya Nadella acknowledged that tokenmaxxing can be addictive, but warned in a recent blog post that customers in these models are paying twice for the AI. The first time is to spend on the token and the second time is to feed them all your proprietary data. Nadella’s comments were unusual in that they promoted Microsoft’s unique approach while raising questions about the data protection guarantees of major AI providers.

Palantir CEO Alex Karp went further, telling CNBC earlier this month that something was “completely wrong.” He said he was giving voice to American companies who were privately “resentful” about paying so much money for tokens that didn’t create any value.

“The basic mindset of companies in this country is, ‘I’m going to relax and waste my time with tokens. I’m not going to get any value and they’re going to get my IP,'” Karp said.

bane & Corporate management consultant Jue Wang said many of the large companies her firm advises are taking a close look at the returns on AI investments.

“Their token cost is doubling almost every other month,” she says. “Let’s say $200 per developer per month. Multiply that by 20,000 developers. This is often in these companies, and you quickly get numbers that are not line items that the general manager has planned.”

In some cases, that just means not using the AI ​​equivalent of a sledgehammer to crack nuts.

“You don’t need Claude Opus 4.6 for everything,” she said of one of Anthropic’s more capable models, suitable for software engineering and deep research. “And yet, we see so many companies, so many users, using Opus by default for everything, including email generation.”

So the search for a tool to do AI “model routing” began. This automatically sends simple queries to cheaper and more efficient AI systems, and more complex tasks to more powerful models.

Software developer Hassan El Mughalli said the shock among companies over the “exorbitant amounts” spent on subscriptions for AI products by major US companies is causing many companies to shy away from rewards commensurate with high usage rates.

“It’s better to empower your employees to use these things so they can use AI when and as much as they need,” said El Mugari, who leads developer experience at Together AI, a startup that provides developers with a variety of “open source” AI models.

At the same time, people who like to collect as many tokens as possible are experimenting with new open source models from Chinese startups like Moonshot’s Kim and Zhipu’s GLM. These models nearly match the features of top US models at a fraction of the price.

“There is some validity to the theory that this could lead to further token maxing,” said Raffi Krikorian, Mozilla’s chief technology officer. “But if you look at the industry as a whole, I think people are realizing that token maxing is a foolish move.”

This is similar to how software companies used to think that the number of lines of code a programmer wrote was a good indicator of productivity, Krikorian said. It later fell out of favor.

“I think tokenmaxxing is following the exact same pattern,” he said. “I think this will be an interesting event that we’ll all look back on and laugh about in a year’s time.”



Source link