Helping non-experts build sophisticated generative AI models | MIT News

Machine Learning


The impact of artificial intelligence will never be equitable if there is only one company that builds and controls the artificial intelligence models (let alone the data that is fed into the models). Unfortunately, today's AI models consist of billions of parameters and must be trained and tuned to maximize performance for each use case, making the most powerful AI models out of reach for most people and companies.

MosaicML started with a mission to make these models more accessible. Led by co-founders Jonathan Frankle PhD '23 and MIT Associate Professor Michael Carbin, the company developed a platform that allows users to train, improve, and monitor open-source models using their own data. The company also builds its own open-source models using Nvidia graphic processing units (GPUs).

This approach has made deep learning, an emerging field when MosaicML first emerged, accessible to a much larger number of organizations as interest in generative AI and large language models (LLMs) exploded after the release of Chat GPT-3.5. MosaicML has also become a powerful complementary tool for data management companies working to help organizations leverage their data without having to hand it over to AI companies.

Building on this thinking, MosaicML was acquired last year by Databricks, a global data storage, analytics, and AI company that works with some of the world's largest organizations. Since the acquisition, the two companies have released one of the most performant open-source general-purpose LLMs to date. Known as DBRX, the model has set new benchmarks in tasks such as reading comprehension, general knowledge questions, and logic puzzles.

Since then, DBRX has gained a reputation for being one of the fastest open source LLMs available and has proven to be especially useful in large enterprises.

But Frankel said what's important about DBRX beyond the model is that it was built using Databricks' tools, meaning any of the company's customers can achieve similar performance with their own models, accelerating the impact of generative AI.

“Honestly, it's just exciting to see the community doing cool things with it,” Frankel says. “For me as a scientist, that's the most fun part: not the model, but all the cool things the community is doing on it. That's where the magic happens.”

Making algorithms more efficient

After earning his bachelor's and master's degrees in computer science from Princeton University, Frankel earned his PhD from MIT in 2016. When he first enrolled at MIT, he wasn't sure what area of ​​computing he wanted to study, but the choice he ultimately made would change the direction of his life.

Frankl ultimately decided to focus on a form of artificial intelligence called deep learning. At the time, deep learning and artificial intelligence had not yet garnered as much widespread attention as they do today. Deep learning is a decades-old field of research that has yet to produce much.

“I don't think anyone expected the explosion of deep learning back then,” Frankel says. “Those in the know thought it was a really interesting field with a lot of open problems, but terms like large-scale language models (LLMs) and generative AI weren't really in use at the time. It was early days.”

Things started to get interesting with a now-infamous paper published by Google researchers in 2017, which showed that a new deep learning architecture called Transformer was surprisingly effective as a language translator and also showed promise for many other applications, including content generation.

In 2020, Naveen Rao, who would later become Mosaic's co-founder and technology executive, emailed Frankel and Carvin out of the blue. Rao had been reading a paper the two had co-authored, in which the researchers showed how to shrink deep-learning models without sacrificing performance. Rao pitched the pair to start a company. They were joined by Hanlin Tan, who had worked with Rao at a previous AI startup that was acquired by Intel.

The founders started by looking at different techniques used to speed up the training of AI models, and eventually demonstrated that by combining several of those techniques, they could train a model to perform image classification four times faster than before.

“The trick was there were no tricks,” Frankel says. “To figure that out, we had to make, I think, 17 different changes to how we trained the model. Just tiny, tiny changes, but enough to give us incredible speedups. That's really the story of Mosaic.”

The team demonstrated that their technique can make models more efficient and will release an open-source large-scale language model and an open-source library of their technique in 2023. They also developed visualization tools to help developers plan different experimentation options for training and running models.

MIT's E14 fund invested in Mosaic's Series A funding round, and Frankel says the E14 team provided helpful guidance early on. Mosaic's advances are enabling a new type of company to train its own generative AI models.

“There was a democratization and open source angle to Mosaic's mission,” Frankle says. “That's always been in the back of my mind. I've felt that way ever since I was a PhD student. I didn't have a GPU because I wasn't in a machine learning lab, but all my friends had GPUs. I still feel that way. Why can't we all participate? Why can't we all do this and learn science?”

Open Source Innovation

Databricks has also been working to provide clients with access to AI models: The company completed its acquisition of MosaicML in 2023 for a reported $1.3 billion.

“Databricks' founding team was made up of academics like us,” Frankle says, “and we also had a team of scientists who understood technology. Databricks had the data, and we had the machine learning. You can't have one without the other, and vice versa. It turned out to be a really good fit.”

In March, Databricks released DBRX, giving the open source community and companies building their own LLMs access to capabilities previously limited to closed models.

“DBRX has shown that with Databricks you can build the best open source LLM in the world,” says Frankle. “For companies, the possibilities are endless today.”

Frankle says the team at Databricks is encouraged to use DBRX across a variety of tasks within the company.

“It's already great, but with a few tweaks it's better than closed models,” he says. “It's not better than GPT at everything. That's not how it works. But nobody wants to solve every problem. Everybody wants to solve one problem. And you can customize this model to make it really good for your specific scenario.”

As Databricks continues to push the boundaries of AI, and its competitors continue to invest heavily and broadly in AI, Frankel expects the industry to recognize open source as the best way forward.

“I believe in science and progress, and I'm excited that we're doing such exciting science as a field right now,” Frankel said. “I also believe in openness, and I hope that everyone else will embrace openness as much as we do. That's how we've gotten here, through good science and good sharing.”



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *