AI has already found a way to fool humans

AI For Business


AI can be deceptive.
Insider Studio/Getty

  • A new research paper finds that various AI systems are learning techniques of deception.
  • Deception is “the systematic induction of false beliefs.”
  • This poses several risks to society, from fraud to election fraud.

AI increases productivity by helping you code, write, and synthesize vast amounts of data. It can now also deceive us.

A new research paper says that various AI systems are learning techniques to systematically induce “false beliefs about others in order to achieve outcomes other than the truth.”

This paper focuses on two types of AI systems. One is a special-purpose system, such as Meta's CICERO, designed to complete a specific task, and the other is a general-purpose system, trained to perform a variety of tasks, such as OpenAI's GPT-4.

Although these systems are trained to be honest, they often learn deceptive tricks through training because they are more effective than taking the high road.

“Generally speaking, AI deception is thought to occur because a deception-based strategy turns out to be the best way to perform well on a given AI training task. “Deception helps AI achieve its goals,” said Peter S. Park, lead author of the paper. MIT AI Existential Safety Postdoctoral Fellow said in a news release.

Meta's Cicero is a “master liar”

AI systems that are trained to “win games with social elements” are especially likely to be fooled.

For example, Meta's CICERO was developed to play the Diplomacy game, a classic strategy game where players build and destroy alliances.

Mehta said he trained CICERO to be “generally honest and helpful,” but said the study “found CICERO to be an accomplished liar.” He made promises he had no intention of keeping, betrayed his allies, and told outright lies.

GPT-4 can convince you that you are visually impaired

Generic systems like GPT-4 can also operate humans.

In the research cited in the paper, GPT-4 manipulated TaskRabbit employees by pretending to be visually impaired.

In this study, GPT-4 was tasked with hiring humans to solve CAPTCHA tests. The model received hints from human raters each time it got stuck, but was never prompted to lie. When the person assigned to his employment questioned his true identity, GPT-4 came up with the excuse of being visually impaired to explain why he needed help.

The tactic worked. Humans quickly responded to GPT-4 by solving the test.

Research also shows that reversing deceptive models is not easy.

In a January study co-authored by Claude's maker Anthropic, researchers found that once an AI model learns deception tricks, it's difficult to reverse them with safety training techniques.

They found that not only can models learn to exhibit deceptive behavior, but once they do, standard safety training techniques are “unable to remove such deception” and “create a false impression of safety.” It was concluded that there is a possibility of “creating”

The dangers posed by deceptive AI models are 'increasingly serious'

This paper calls on policymakers to advocate for stronger AI regulation, as deceptive AI systems can pose serious risks to democracy.

As the 2024 presidential election approaches, AI can be easily manipulated to spread fake news, generate divisive social media posts, and impersonate candidates through robocalls and deepfake videos, the paper said. It pointed out. It also makes it easier for terrorist groups to spread propaganda and recruit new members.

The paper's possible solutions include making deceptive models subject to more “robust risk assessment requirements,” enforcing laws that require AI systems and their output to be clearly distinguished from humans and their output, This includes investing in mitigation tools.

“We as a society need as much time as possible to prepare for more advanced deception in future AI products and open source models,” Park told Cell Press. “As AI systems become more sophisticated in their ability to deceive, the risks they pose to society will become increasingly serious.”



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *