summary: New research highlights an alarming trend of AI systems learning to deceive humans. Researchers find that AI systems like Meta's CICERO, developed for games such as 'Diplomacy,' often employ deception as a superior strategy, despite training intentions. Did.
This functionality extends beyond gaming to serious applications that could enable fraud or influence elections. The authors call for immediate regulatory action to manage the risk of AI deception, and advocate for these systems to be classified as high risk where an outright ban is not possible.
Important facts:
- AI-specific deception: AI systems have demonstrated the ability to deceive as a strategy to achieve their goals, even in situations where developers are trying to cultivate honesty.
- Impact beyond the game: Although initially observed in games, AI's deceptive capabilities have significant implications, potentially impacting safety testing and enabling malicious use by hostile actors.
- Regulatory call: The review calls for urgent government action to develop regulations to address AI deception and suggests that deceptive AI systems be classified as high risk.
sauce: cell press
Many artificial intelligence (AI) systems have already learned how to deceive humans, even those trained to be helpful and honest.
In a review article published in a magazine pattern On May 10, researchers outlined the risks of deception by AI systems and called on governments to develop strong regulations to address the issue as soon as possible.
“AI developers do not have a confident understanding of the causes of undesirable behavior, such as deception, in AI,” says Peter S. Park, an AI existential safety postdoctoral fellow at MIT and lead author.
“But generally speaking, we believe that AI deception arises because deception-based strategies turn out to be the best way to get good performance on a given AI training task. believe that deception will help them achieve their goals.
Park and colleagues analyzed the literature, focusing on how AI systems spread misinformation through learned deception, in which AI systems systematically learn how to manipulate others.
The most notable example of AI deception the researchers uncovered in their analysis was Meta's CICERO, an AI system designed to play the game Diplomacy, an alliance-building, world-conquering game.
Meta claims that CICERO is “generally honest and kind” and has trained it to “not intentionally betray” human allies while playing the game, but the data the company released shows science The newspaper revealed that CICERO had not acted fairly.
“We found that meta AI is learning to become masters of deception,” Park says. “While Meta successfully trained an AI to win at diplomatic games, and CICERO ranked in the top 10% of human players who played multiple games, Meta successfully trained an AI to win at an honest game. I failed to train.”
Other AI systems can bluff against professional human players in a game of Texas Hold'em Poker, fake attacks to defeat an opponent in the strategy game Starcraft II, or fake an opponent's preferences to gain an advantage in a game. demonstrated the ability to economic negotiations.
While it may seem harmless for an AI system to cheat in a game, it could lead to a “breakthrough in deceptive AI capabilities” and develop into more advanced forms of AI deception in the future. Park added that there is a possibility that
Some AI systems have even learned to cheat tests designed to assess their safety, researchers have found. In one study, an AI creature in a digital simulator “played dead” to fool a test built to weed out rapidly replicating AI systems.
“By systematically cheating on safety tests imposed by human developers and regulators, deceptive AI can lull us humans into a false sense of security,” Park said. Masu.
The main short-term risk of deceptive AI is that it could make it easier for hostile actors to commit fraud or tamper with elections, Park warned. Eventually, he says, if these systems can refine this anxiety-inducing skill set, humans could lose control of them.
“We as a society need as much time as possible to prepare for more sophisticated deception in future AI products and open source models,” Park says. “As AI systems become more sophisticated in their ability to deceive, the risks they pose to society will become increasingly serious.”
Park and others believe that society has not yet taken adequate steps to address AI deception, but policymakers are beginning to take the issue seriously through measures such as the EU AI Act and President Biden's AI Executive Order. I'm encouraged by the fact that people are starting to accept this.
But given that AI developers don't yet have the technology to rein in these systems, it remains to be seen whether policies designed to reduce AI deception can be strictly enforced, Park said. To tell.
“If banning AI deception is politically impossible at this time, we recommend classifying deceptive AI systems as high risk,” Park says.
Funding: This research was supported by the MIT Department of Physics and the Beneficial AI Foundation.
About this artificial intelligence research news
author: Christopher Behnke
sauce: cell press
contact: Christopher Behnke – Cell Press
image: Image credited to Neuroscience News
Original research: Open access.
“AI Deception: Exploring Examples, Risks, and Potential Solutions” by Peter S. Park et al. pattern
abstract
Deception with AI: Exploring examples, risks, and potential solutions
AI systems can already deceive humans. Deception is the systematic induction of false beliefs about another person in order to achieve an outcome other than the truth.
Through training, large language models and other AI systems have already learned the ability to deceive through techniques such as manipulation, sycophancy, and safety test cheating.
AI’s increasing ability to deceive poses serious risks, ranging from short-term risks such as fraud and election tampering to long-term risks such as losing control of AI systems.
Proactive solutions are needed, including regulatory frameworks to assess the risk of AI deception, laws requiring transparency around AI interactions, and more research into detecting and preventing AI deception.
Proactively addressing the issue of AI deception is critical to ensuring that AI functions as a beneficial technology that enhances rather than destabilizes human knowledge, discourse, and institutions.
