Why latency and “total cost of ownership” become more important for AI apps

Applications of AI


Developer-turned-CEO Lin Qiao predicts a new era of AI, where language models are fine-tuned based on an organization's own specialized data. This will allow organizations to leverage the language capabilities of his AI while leveraging their own datasets to provide feedback.

Before becoming CEO of Fireworks AI, Qiao led Meta's PyTorch efforts. Generative AI can solve hundreds of complex logic problems, she noted, but they're not problems that businesses and developers typically face.

“Large models are too expensive to operate and cannot deliver the low latency needed to deliver a great product experience,” she said. “That puts pressure on me. [on] So that people move to smaller models.”

The smaller the model, the better it can address the business problem developers are trying to solve.

“They probably have to solve five business-specific challenges,” she says. “we [are] Focusing on a small-scale open source model, [them] It's either on par with OpenAI's models in terms of quality or better than OpenAI's models in terms of quality. At the same time, it provides much lower latency and a much lower total cost of ownership for these B2C applications and products. ”

In this new era of AI, developers face two problems, Qiao said.

  1. Iterate your training faster using enterprise data.
  2. Scaling generative AI applications in production.

The company she co-founded, Fireworks AI, is “particularly focused” on addressing these two problems for developers, she told The New Stack. “We offer very quick tweaks,” she added.

Fireworks AI leverages open source models. It recently raised $25 million in funding and claims 12,000 users, including Quora, Sourcegraph, and AI-Powerpoint presentation company Tome. It is estimated that more than 25 billion tokens are given away every day.

Latency is important in AI applications

Chao learned that at a B2C company like Meta, where he previously worked, interactivity and low latency are a must. Content generation specifically affects whether a product is viable, she said, adding that creating high-quality AI products requires using unique data and iterating quickly on models.

“The developers at the companies we spoke to all have their own data and use our fine-tuning platform to generate customized models,” she said. “A one-click upload to our inference platform allows products to communicate directly with customized models using content generated from the model.”

Developers must then review the product metrics, adjust the data as needed, and continue in the loop to fine-tune the model.

Additionally, AI applications must be able to scale very quickly while keeping the total cost of ownership low, she added.

“If the costs are high, you run out of cash quickly, it becomes a disaster, and it becomes unviable for you to do business,” she says. “Latency and TCO are both important for B2C companies.”

AI app cost challenges

But even with great products, generative AI applications can be more expensive than traditional applications, which contributes to total cost of ownership. One way generative AI applications differ from traditional applications is that they must run on GPUs rather than on increasingly commoditized CPUs.

“GPUs are expensive. It's not just the chips that are expensive. GPUs are very power hungry. Power is expensive. Power generates heat. You can't have air cooling, you can't use liquid cooling or put the chips in oil. must use reverse cooling. [to] “You're taking the heat away,” Chao said, “so all that supporting infrastructure drives up the overall infrastructure costs for GenAI.”

That cost can be an additional barrier to business viability, she added.Fireworks seeks to help companies address the TCO challenge by focusing on smaller open-source models that are cost-effective to run and are on par with or better than LLM-generated AI products.

“We offer much lower latency and a much lower TCO for B2C applications and products,” she said.

Example of using the small model

Many of Fireworks AI's customers are using AI to create assistants, she said. Medical assistants, legal assistants, and coding assistants are popular use cases. Latency is a particularly important challenge for AI due to the interactive and conversational nature of its output.

Another use case she sees frequently is documentation. From images to PDFs, AI is being used to scan and search documents for product catalogs, e-commerce, and even risk analysis. Fireworks AI customer Tome uses AI to create presentation slides for business users.

Without fast response times, AI applications can become terrible products, she added.

“That response time should typically be half a second or one second,” she says. “It makes for a much more interesting product because it's responsive and interactive.”

group Created with Sketch.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *