Research warns that open source AI might look cheap, but it could cost more to run secretly

AI News


New research raises questions about one of the most common assumptions in the AI industry. The open source model is cheaper to run than the closed source counterpart. A study published this week by AI company Nous Research shows that it consumes far more computing resources and sometimes cancels out apparent cost benefits than closed systems from companies such as OpenAI and humanity.

The team reportedly investigated 19 AI models across three categories of tasks: simple knowledge questions, mathematical problems, and logic puzzles. They measured the basic unit of AI calculation, the number of tokens used to reach the answer for each model. The outcome was tough. “Openweight models use 1.5-4 tokens (up to 10 for simple knowledge questions) than closed ones, and can be more expensive per query despite lower costs per token,” the researcher wrote in the report.

This study found that inefficiency is particularly pronounced in large inference models. This attempts to work through the problem step by step. In some cases, these models spent “thinking simple knowledge questions with hundreds of tokens,” including naming Australia's capital.

The findings suggest that companies considering adopting AI may be overlooking token efficiency, a key factor. “While hosts with open weight models can be cheaper, the benefits of this cost can easily be offset if more tokens are needed to infer about a particular issue,” the report states.

Openai's system was highlighted as a powerful performer. The O4-MINI and the newly released open-weight GPT-OSS model provided some of the best token efficiencies found in this study, particularly in mathematical tasks. “The Openai model stands out for its extreme token efficiency in mathematical problems,” the author writes.

Among the open source offerings, Nvidia's Llama-3.3-Nemotron-Super-49B-V1 was named “the most token-efficient openweight model in all domains.” Meanwhile, new models from companies such as Mistral were labelled “outliers” due to their extremely high token use.

The efficiency gap between open and closed models varies depending on the task. For mathematical and logical problems, the open system uses about twice as many tokens. For simple fact queries, the difference is much larger, and can swell up to ten times.

Closed source providers appear to be actively working to reduce costs. “The closed weight model is optimized to use less tokens to reduce inference costs,” the researchers said. In contrast, open models tended to increase token use in recent versions, perhaps pursuing more advanced inference capabilities.

The authors argued that token efficiency should be a major design goal along with accuracy. They suggested that a better approach to thinking reasoning could improve efficiency and reduce the degradation of context for more difficult problems.

– end

Published:

Nandini Yadav

Published:

August 18, 2025

I'll adjust it



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *