Team NVIDIA Wins Trophy for Recommendation System

Machine Learning


A talented NVIDIA team of five machine learning experts across four continents won all three tasks in a fierce and prestigious competition to build a state-of-the-art recommendation system.

The results reflect the group’s cleverness in applying the NVIDIA AI platform to real-world challenges to the engine of the digital economy. Recommenders serve trillions of search results, ads, products, music and news articles to billions of people every day.

A team of over 450 data scientists participated in the Amazon KDD Cup ’23. There were twists and turns in this three-month challenge, and it was tough until the end.

shift to high gear

The team held a comfortable lead in the first ten weeks of the competition. However, in the final stages the organizers switched to a new test dataset and the other teams jumped ahead.

The NVIDIA guys moved into high gear, working nights and weekends to catch up. They left a trail of his Slack messages around the clock from team members living in cities from Berlin to Tokyo.

“We were working around the clock. It was very exciting,” said San Diego team member Chris Deotto.

Products with other names

The last of the three tasks was the most difficult.

Participants were required to predict which products users would purchase based on data from their browsing sessions. However, the training data did not include brand names for many choices.

“We knew from the beginning that this would be a very difficult test,” said Gilberto ‘Giba’ Titelic.

KGMON to the rescue

Based in Curitaba, Brazil, Titericz is one of four team members ranked Grandmaster in the Kaggle competition, the online Olympiad for Data Science. They are part of a team of machine learning ninjas who have won dozens of competitions. NVIDIA founder and CEO Jensen Huang jokingly calls Pokémon KGMON (Kaggle Grandmasters of NVIDIA).

Titericz used Large Language Models (LLM) in dozens of experiments to build generative AI to predict product names, but none worked.

In a creative flash, the team found a workaround. Predictions using the new hybrid ranking/classifier model were spot on.

up to the wire

During the final hours of the competition, teams raced to package all their models for some final submissions. They were experimenting on his 40 computers overnight.

Kazuki Onodera, KGMON in Tokyo, felt uneasy. “We really didn’t know if the actual scores matched what we were estimating,” he said.

photo of KGMON
The four members of KGMON (clockwise from top left) Onodera, Titerich, Deot, Puget.

Deotte, also a KGMON, recalls this as “like 100 different models all working together to produce one output. We submitted it to the leaderboards and POW !”

The team had a slight lead over its nearest rivals in the AI ​​equivalent of photofinish.

The power of transfer learning

Another task requires the team to apply lessons learned from large datasets in English, German, and Japanese to poor datasets one-tenth the size of French, Italian, and Spanish. there was. This is the kind of real-world challenge many companies face as they expand their digital presence around the world.

Jean-Francois Puget, a three-time Kaggle Grandmaster based outside Paris, knew an effective approach to transfer learning. He used a pre-trained multilingual model to encode product names and fine-tuned the encoding.

“Using transfer learning, our leaderboard scores improved significantly,” he said.

Savvy meets smart software

KGMON’s work shows that the field known as recsys is sometimes more art than science, a practice that combines intuition and repetition.

That expertise is encoded in software products such as NVIDIA Merlin, a framework that enables users to rapidly build their own recommendation systems.

Diagram of Merlin's Recommended Framework
The Merlin framework provides an end-to-end solution for building recommender systems.

Berlin-based teammate Benedikt Schifferer, who helped Merlin design, used the software to train a transformer model and beat a competitor’s classic resys task.

“Merlin provides excellent results out of the box, and its flexible design allows us to customize the model for specific challenges,” he said.

ride the rapids

Like his teammates, he used RAPIDS, a set of open source libraries for accelerating data science on GPUs.

For example, Deotte accessed code from NGC, NVIDIA’s acceleration software hub. This code, called DASK XGBoost, helped distribute large and complex tasks across 8 GPUs and their memory.

Titericz used a RAPIDS library called cuML to search millions of product comparisons in seconds.

The team focused on session-based recommenders that do not require data from multiple user visits. This is a best practice today as many users want to protect their privacy.

You can learn more about:



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *