Large-scale language models (LLMs) are powerful but also confusing. These tools have the power to help graduate students draft emails and guide clinicians in diagnosing cancer.
The applications of these models are endless, but so are the questions about how to best evaluate them.
The study comes from the Massachusetts Institute of Technology (MIT), and a group of researchers took a different approach to gain a deeper understanding of LLMs.
Learn about their new perspective, which places human beliefs and expectations at the center of the evaluation process.
Machine learning and language models
The research team from MIT included three experts from various fields.
Study co-author Ashesh Rambachan, assistant professor of economics and principal research scientist in the Laboratory for Information and Decision Systems (LIDS), was the driving force behind the research.
Lead authors Keion Vafa, a postdoctoral researcher at Harvard University, and MIT professor Sendhil Mullainathan also contributed to the paper.
This unique perspective of their research will be presented at the International Conference on Machine Learning.
Human Beliefs and Generalizations
This study revolved around the concept of human generalization, i.e., how we form and update beliefs about our LLM capabilities.
You've probably done this yourself: You've seen someone great at correcting grammar and generalized that they must also be good at sentence construction.
That’s exactly what we do with language models. But Rambachan points out that these LLMs aren’t actually human, so generalizations can lead to unexpected model failures.
Building on this, the team introduced a human generalization function to assess the consistency between human beliefs and LLM performance.
Research shows…
So how did the team approach this exciting endeavor? The researchers designed a study to measure how people generalize when interacting with LLMs.
Participants were shown questions that the individual or LLM answered correctly or incorrectly and were asked to predict their performance on related questions.
This resulted in a dataset of approximately 19,000 examples, providing insight into how humans form beliefs about LLM performance across 79 diverse tasks.
Measuring language model inconsistencies
The findings revealed an interesting pattern: participants predicted human performance fairly accurately, but failed when it came to predicting LLM performance.
“Human generalizations are applied to language models, but they don’t work well because these language models don’t exhibit patterns of expertise like humans do,” Rambachan correctly pointed out.
Another interesting finding was that humans were more likely to update their beliefs about the LLM's performance if the LLM got a question wrong.
In these scenarios, simpler models seem to outperform very large models like GPT-4, and the researchers speculate that this may be because LLMs are not as accustomed to interacting with them as humans.
Implications for model development
The results of this research have significant implications for the future development of language models.
As the researchers emphasize, understanding how human beliefs shape expectations of LLMs could inform model design and training methodologies.
By incorporating insights into human generalization patterns, developers can create models that are not only more robust but also more closely match user expectations.
This increases the transparency of model capabilities and allows for better integration of LLMs into real-world applications.
Future research on language models
This study highlights the complex relationship between people's belief systems and LLM performance and creates many opportunities for future research.
Questions regarding how different demographics interact with LLMs or how context influences across-person generalizations remain largely unexplored.
Further research could focus on refining people's generalization functions to improve alignment measures and on exploring more deeply the psychological aspects of belief formation and adjustment according to LLM functions.
Exploring these methods may ultimately result in more effective and user-friendly artificial intelligence technology.
The Road Ahead
Building on these insights, the researchers are eager to conduct further research into how people's beliefs about LLMs change over time as they interact with the models more frequently.
We also aim to understand how human generalization can be incorporated from the very beginning of LLM development.
In the meantime, they hope that their dataset will be used as a benchmark to measure the performance of LLM relative to human generalization capabilities, which could revolutionize the performance of models deployed in real-world situations.
“When training these algorithms in the first place, or when trying to update them with human feedback, we need to take human generalization capabilities into account when thinking about how to measure performance,” the researchers remind us.
The research was funded in part by the Harvard Data Science Initiative and the Center for Applied AI at the University of Chicago Booth School of Business.
The study has been published in the journal arXiv.
—–
Like this article? Subscribe to our newsletter for more fascinating articles, exclusive content and updates.
Check it out with EarthSnap, a free app brought to you by Eric Ralls and Earth.com.
—–
