Reuters
–
Chinese artificial intelligence developer Deepseek spent just $294,000 on training the R1 model. It says this is far less than reported to its US rivals.
A rare update from the Hangzhou-based company – the first released estimate of R1 training costs – was published in a peer-reviewed article in the academic journal Nature on Wednesday.
With the release of Deepseek saying it was a low-cost AI system in January, global investors were forced to dump high-tech stocks as they worried the new model could threaten the dominance of AI leaders, including NVIDIA.
Since then, the company and its founder Liang Wenfeng have largely disappeared from public places, except to push out some product updates.
Sam Altman, CEO of US AI Giant Openai, said in 2023 that training basic models cost “a lot more” than $100 million, but his company has not given detailed figures for that release.
The training costs for large language models that run AI chatbots refer to the costs that arise from running a powerful cluster of chips for weeks or months to process huge amounts of text and code.
A Nature article listing Liang as one of his co-authors said training an R1 model focused on Deepseek's reasoning cost $294,000 and used a 512 Nvidia H800 chip. Previous versions of the article published in January did not contain this information.
Some of Deepseek's statements regarding the cost of development and the technologies it used have been questioned by US companies and officials.
The H800 chip mentioned was designed by NVIDIA for the Chinese market after the US became illegal in October 2022 to export more powerful H100 and A100 AI chips to China.
US officials told Reuters in June that Deepseek could access “a massive amount” of H100 chips procured after US export controls were in place. Nvidia told Reuters at the time that Deepseek used a legally acquired H800 chip, not a H100.
In a supplementary information document accompanying the Nature article, the company admitted it was the first time it owned A100 chips and said it used them in the preparation stages of development.
“In regards to the study on DeepSeek-R1, we used the A100 GPU to prepare ourselves for experiments with smaller models,” the researchers wrote. After this early stage, he added that R1 was trained for a total of 80 hours on a 512-chip cluster of H800 chips.
Deepseek also responded for the first time, although not directly, to claim that it had “distilled” Openai's model independently from top White House advisors in January and other US AIs.
This term refers to the way in which one AI system learns from another, and while a new model can enjoy the advantages of the investment in the time it took to build a previous model and computing power, there is no associated cost.
Deepseek consistently provides distillation to improve model performance while still providing wider access to technology with much cheaper and AI power.
Deepseek said in January that he used Meta's open-source Llama AI model to be used for a distilled version of his own model.
Deepseek said that training data for the V3 model “relies on crawled web pages containing answers generated by a considerable number of OpenAI-Models, so the base model could indirectly acquire knowledge from other powerful models.” However, he said this was not intentional, but rather contingent.
Openai did not respond immediately to requests for comment.
