Research: Generated AI results depend on user prompts as much as the model

AI For Business


There is a natural assumption that as generative artificial intelligence systems improve, better large-scale language models lead to better results. However, new research in some MIT throne affiliates suggests that LLM progress is just part of the story.

In large-scale experiments, researchers found that only half of the performance improvements seen after switching to more advanced models of AI came from the model itself.

The other half came from leveraging the new system from how users adapted their prompts, namely, written instructions to tell the AI model what to do.

The simple but powerful insight that user adaptation contributes to performance as much as emphasizing the realities that the model upgrade itself is important to the company, and investment in new AI tools will not provide the expected value unless employees narrow down how they use them. In this case, prompts are learnable skills that people can quickly improve without guidance.

“People often assume that better outcomes come primarily from better models,” said Professor David Holtz, SM18, PhD 21, of Columbia University, researcher at MIT initiative on the digital economy and one of the research co-authors.. “The fact that almost half of the improvement comes from user behavior really challenges that belief.”

Better prompts, improved models improve performance

In the experiment, nearly 1,900 participants were randomly assigned to one of three versions of Openai's Dall-E image generation system. Dall-E 2, more advanced Dall-E 3, or Dall-E 3 were automatically rewritten by GPT-4 LLM without knowing the user's prompt.

Participants were presented with reference images such as photographs, graphic designs, and artworks, and were asked to recreate them by entering instructions into the AI. It took them 25 minutes to submit at least 10 prompts and they were told that the top 20% of performers would receive a bonus payment.

Researchers found:

  • Participants using the baseline version of Dall-E 3 created images that resemble the target image than those generated by Dall-E 2 users.
  • Participants using the baseline version of Dall-E 3 wrote a 24% longer prompt compared to Dall-E 2 users. These prompts also tended to be more similar to each other, with more descriptive words in place.
  • Approximately half of the improvements in image similarity came from improved models, while the other half came from how users tweaked the prompts to take advantage of the improved models.

Although this study examined image generation, researchers believe that the same pattern applies to other tasks, such as writing and coding.

Prompts are about communication, not coding

This study showed that the ability to adapt prompts over time is not limited to tech-savvy users.

“People often think you need to be a software engineer to increase the profits and make profits for AI,” says Holtz. “However, our participants came from a wide range of jobs, education levels and age groups. Even people with no technical backgrounds were able to make the most of the capabilities of the new model.”

The data suggests that prompts are more about communication than coding. “The best prompter wasn't the software engineer,” Holtz said. “They were people who knew how to express ideas clearly in everyday languages, not necessarily code.”

This accessibility also helps reduce performance gaps between users with different skill levels and experience. EAMAN JAHANI, Associate Professor at the University of Maryland, PhD '22, Digital Fellow of the MIT Initiative on the Digital Economy and Co-author of the Research, Note that generation AI can narrow performance gaps between users.

“Those who start at the bottom edge [performance] Scale was the most profitable. In other words, the difference in results has been reduced. “Model advancements can actually help reduce inequality in the output.”

Jahani noted that team findings apply to tasks with clear, measurable results, where there is a cap on what is good. He pointed out that it is not clear whether the same pattern will be held in a more open-ended task without a single correct answer and without potentially large payoffs, such as coming up with a transformative new idea.

Performance has been worsened due to rewriting prompts using the generator AI

One more surprising result came from a group that used Dall-E 3 to rewrite the prompts. Although this feature is designed to aid users, it reduced the performance of the image generation task by 58% compared to the baseline DALL-E 3 groups.

The team discovered that autorewrites often add additional details and change the meaning of what users are trying to say, leading AI to generate the wrong kind of image.

“[Automatic prompt rewriting] Holtz said: “In such a task, the goal is to match the target images as closely as possible. More importantly, it shows how AI systems collapse when designers assume how people use them. Hard-code hidden steps in tools can easily compete with what users are actually trying to do.”


Business dressed man holding a maestro baton adjusting data image in the background

Leading an AI-led organization

Directly on MIT Sloan

How companies can devalue using AI

The point is that apart from choosing the “correct” AI model, business leaders should also focus on enabling the right kind of user learning and experimentation. The prompt is not a plug and play skill, Jahani said. “Companies need to continue investing in HR,” he said. “People need to catch up with these technologies and know how to use them well.”

To build profits enabled by generative AI, researchers provide some priorities for business leaders looking to make AI systems more effective in real-world settings.

  • Invest in training and experimentation. A technical upgrade alone is not enough. Provide time and support to improve the way employees interact with AI systems is essential to achieving a complete performance improvement.
  • Design for repetition. An interface that encourages users to test, modify, learn, and clearly display results will help drive better results over time.
  • Be careful with automation. Automated and rapid rewriting may be useful, but obscuring or overriding user intent can potentially hinder performance rather than improve it.

This paper was also co-authored by students at MIT Sloan PhD. Benjamin S. ManningSM '24; Hong-yiTuye, SM '23; and Mohammed is also bay. '16, SM '24; so are doctoral students at Stanford University. Joe ChanMicrosoft Computational Social Scientist Siddharth Sri, Assistant Professor, University of Cyprus Christos NicholaidesSM '11, PhD '14.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *