Join C-suite executives in San Francisco July 11-12 to hear how leaders are integrating and optimizing their AI investments for success.. learn more
Refuel AI, a company that uses Large Language Models (LLM) to generate high-quality training data for AI models, exited stealth today with $5.2 million in seed funding. The company said it will use the round to grow its team, build out the capabilities of its platform, and prepare for its commercial launch in July.
Founded by Stanford University alumni Nihit Desai and Rishabh Bhargava, Refuel also opened access to AutoLabel. AutoLabel is an open source library that allows AI teams to easily label data using his LLM of choice in their own environment.
>>Don’t miss our special issue: Laying the Foundation for Customer Data Quality<
These products serve as an answer to the data challenges that are slowing AI development and preventing companies from incorporating next-generation technology into their products and business functions.
event
transform 2023
Join us July 11-12 in San Francisco. There, he shares how management integrated and optimized his AI investments to drive success and avoid common pitfalls.
Register now
Every AI company needs AI-enabled data
Today, companies are all racing to become AI companies, working with in-house experts and third-party vendors to develop models that can target a variety of business-specific use cases. This task can be quite daunting, but all AI projects need the same starting point: clean, labeled data. If this is done correctly, the project will come to fruition easily.
Companies now have a lot of data at their disposal, but not all of it is training-ready by default. To train the model, the information must be cleaned and annotated. This task is usually handled by a team of humans and takes weeks or months. It cannot keep up with today’s AI demands.
“Many teams [we spoke to] They had all the great ideas for the models they wanted to train and the products they wanted to build. Assuming you have training data ready. That’s when we knew that making clean, labeled data available at the speed of thought was what we wanted to focus on,” Bhargava told his VentureBeat.
So in 2021, they launched Refuel, a specialized LLM that automates the creation and labeling of datasets (at or better than human quality) for any business and any use case. We continued to build a dedicated platform to
According to the company, enterprise users will be able to use the platform by simply uploading a dataset and telling LLM to label the data. We can also provide guidelines and some examples to ensure you only get high quality data for training.
“Within an hour, they (users) will have enough data to start training AI models and seamlessly connect to the model training infrastructure. (especially from production), we can reroute that data to Refuel for labeling, measuring performance, and improving the dataset for model retraining,” added the CEO.
Private beta testing by some companies has shown that the product speeds up the data creation and labeling process by up to 100%. Bhargava did not name these companies, but noted that Refuel AI has attracted interest from multiple industries, ranging from social media and fintech to healthcare, human resources and e-commerce.
road ahead
Co-led by General Catalyst and XYZ Ventures, the round will see Refuel grow its engineering team from 6 to 12 people and invest more in its platform and LLM infrastructure to prepare for commercial launch by the end of July. The company also plans to invest capital in open source libraries and communities.
“As a concrete example, we are running a contest to push the boundaries of LLM-powered data labeling with prizes up to $10,000,” said Bhargava.
The company currently competes in the data labeling space with players such as Tasq AI, Snorkel AI, and SuperAnnotate.
VentureBeat Mission will be the digital town square for technical decision makers to gain knowledge and transact on transformative enterprise technologies. Watch the briefing.
