Has ChatGPT’s behavior changed over time? Researchers evaluate March 2023 and June 2023 versions of GPT-3.5 and GPT-4 on four diverse tasks

AI and ML Jobs


Large Language Models (LLM) have proven to be the best innovations in artificial intelligence. From BERT, PaLM, GPT to LLaMa DALL-E, these models show amazing performance in understanding and generating language for the purpose of mimicking humans. These models are continuously improved based on the latest information, user input, and design changes. However, it is still uncertain how often GPT-3.5 and GPT-4 will receive updates, making it difficult to integrate these LLMs into broader workflows.

Sudden changes in LLM behavior (such as the accuracy and format of responses to prompts) can disrupt downstream pipelines due to instability. This unpredictability may make it difficult for developers and users to rely on regular results and may limit the stable integration of his LLM into current systems and workflows. To study how the behavior of various Large Language Models (LLMs) changes over time, a team of researchers from Stanford University and the University of California, Berkeley evaluated the behavior of the March 2023 and June 2023 versions of GPT-3.5 and GPT-4.

Three key factors are used to quantify change. These are the LLM services to monitor, the focused application scenarios, and the metrics that measure the LLM drift in each scenario. ChatGPT’s core components, GPT-4 and GPT-3.5, are the LLM services monitored in this study. Considering ChatGPT’s acceptance by both businesses and individuals, as well as its popularity, systematic and timely monitoring of these two services will help users to better understand and use her LLM in their specific use cases.

🚀 Build high-quality training datasets, solve NLP machine learning challenges, and develop powerful ML applications with Kili Technology

The March 2023 and June 2023 snapshots of two major versions of GPT-4 and GPT-3.5, accessible through OpenAI’s API, were used in the study, the main purpose of which was to investigate fluctuations or “drift” between the two dates. The team selected for evaluation four of his commonly investigated His LLM tasks, which serve as benchmarks for performance and safety. These jobs include –

  1. Solving Math Problems – Accuracy measures how often the LLM service produces the right response when solving math problems.
  1. Response to Sensitive Questions: A response rate that indicates how often the LLM service provides direct responses.
  1. Code Generation – Percentage of generated code that is ready to run in your programming environment and satisfies unit tests.
  1. Visual Reasoning – Exact match. Evaluate whether the visual objects created exactly match the source material.

In conclusion, this study focuses on GPT-4 and GPT-3.5, evaluates them in four selected tasks, and quantifies and measures LLM drift in each scenario using both specialized performance measures and other general measures to explore how different LLM behaviors evolve over time. The research results will help users better understand how her LLM works and utilize these models for different applications.


Please check paper. All credit for this research goes to the researchers of this project.Also, don’t forget to participate 26,000+ ML SubReddit, Discord channeland email newsletterShare the latest AI research news, cool AI projects, and more.


Tanya Malhotra is a final year student at the University of Petroleum and Energy Research, Dehradun, graduating with a Bachelor of Science in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
A data science enthusiast with good analytical and critical thinking, she has a keen interest in learning new skills, leading groups, and managing work in an organized manner.


🔥 Gain a competitive edge with data: Actionable market intelligence for global brands, retailers, analysts and investors. (with sponsor)



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *