AT&T reduces AI costs by replacing large models with smaller ones

AI News


AT&T VentureBeat made its internal Ask AT&T personal assistant more efficient by reworking the orchestration layer and moving more work from large-scale language models (LLMs) to small-scale language models (SLMs). reported Thursday (February 26).

This change improves latency, speed, and response time. Reduce costs by 90%. According to the report, the system can now process three times as many tokens.

“We believe the future of agent AI will be many, many small language models,” AT&T Chief Data Officer andy marcus According to the report: “We found that small-scale language models are almost as accurate, if not more accurate, than larger-scale language models in certain domain areas.”

We want to be your favorite news source.

Add us to your preferred sources list to see our news, data, and interviews in your feed. thank you!

a small language model This is a scaled-down version of a larger language model, PYMNTS reported in April. Although SLM does not have as many parameters, the user may not need the additional power depending on the task at hand.

SLM is often faster, cheaper, and provides more control. This is important for companies looking to bring powerful AI into their operations. break the bank. SLM performs as well or better than LLM. For example, it can perform better than LLM in certain domains. have been trained About specific industries. LLM is better in general knowledge.

Nvidia According to research small language model They are powerful enough for many real-world tasks, cheap to run, and can be deployed at scale without the same infrastructure burden as larger language models, so they may prove more practical and profitable in enterprises.

Advertisement: SCROLL TO CONTINUE

In systems where AI agents string together multiple steps to complete complex assignments, the majority of the work requires as little heavy-weight models as possible. Instead, smaller models can handle most of the load, while LLM can handle most of the load. be reserved For a rare, high-stakes step.

The next stage of AI is efficiency and construction model PYMNTS reported in November that it can run smaller, faster, and at a lower cost without sacrificing performance. This strategy allows companies to reduce their total cost of ownership, at a time when approximately 47% of similarly sized companies cite cost as the top barrier to implementing generative AI.

For all of our coverage of PYMNTS AI, subscribe to our daily subscription AI Newsletter.



Source link