When trained on large datasets, huge unsupervised LMs acquire capabilities that surprise even their creators. However, these models are trained on information generated by people with different motivations, goals, and abilities. Not all of these ambitions and abilities can be imitated. To create a reliable, effective, and manageable system, it is important to carefully select the desired responses and behaviors of your model from a vast pool of information and skill.
Researchers at Stanford University and CZ demonstrate how to optimize language models for human preferences without using explicit reward modeling or reinforcement learning. Their work showed that the RL-based goal employed in the current approach can be accurately optimized using a simple binary cross-entropy goal, greatly streamlining the preference learning process and demonstrating how this actually works. It shows how it can be done.
They suggest Direct Preference Optimization (DPO). This new algorithm implicitly achieves the same objective as the existing RLHF algorithm (reward maximization with KL divergence constraint), but is easier to build and train. The DPO update intuitively improves the log ratio of preferred and non-preferred responses, but also includes dynamic per-example importance weights that prevent model degradation.
Like other algorithms, DPO uses theoretical preference models to assess the consistency of empirical preference data and reward functions. While traditional approaches use a preference model to define the preference loss to train a reward model, DPO instead uses variable switches to train a policy that maximizes the learned reward model. increase. Therefore, without explicitly learning the reward function or sampling it from the policy during training, DPO considers a dataset of human preferences for model responses and sets a simple binary His cross-entropy goal of can be used to optimize the policy.
The results of the study show that DPO is as effective as state-of-the-art approaches such as PPO-based RLHF in preference-based learning for various tasks such as emotion regulation, summarization, and dialogue. 6B parameters. 58% preferred her DPO summary over her PPO summary (human rated) on the test set, and 61% preferred her DPO summary over human rating. Anthropic HH has a 60% chance of prioritizing a single-turn response from the DPO over selective completion.
The research team says DPO has many potential uses beyond just training language models based on human preferences. For example, you can train generative models with different modalities.
Evaluation of the proposed model reaches 6B parameters, but the team believes that further work should consider extending DPO to state-of-the-art models with orders of magnitude more data. Researchers also found that the prompt affected his calculated GPT -4 win rate. In the future, we plan to investigate the most effective means of extracting expert opinion from machines.
please check out paper. don’t forget to join 22,000+ ML SubReddit, Discord channeland email newsletterShare the latest AI research news, cool AI projects, and more. If you have any questions regarding the article above or missed something, feel free to email me. Asif@marktechpost.com
🚀 Check out 100’s of AI Tools at the AI Tools Club
Tanushree Shenwai is a consulting intern at MarktechPost. She is currently pursuing her bachelor’s degree at the Indian Institute of Technology (IIT), Bhubaneswar. She is a data her science enthusiast and has a keen interest in the range of applications of artificial intelligence in various fields. She is passionate about exploring new advances in technology and its practical applications.
