The leading artificial intelligence pioneers are concerned about technology that is increasingly likely to lie and deceive. And he has set up his own nonprofit organization to curb such behavior.
In a blog post that announced Lawzero, a new non-profit venture, “AI Godfather” Yoshua Bengio said he was “deeply interested” as the AI model becomes more and more powerful and deceptive.
“The organization was created in response to evidence that today's frontier AI model is increasingly dangerous. [behaviors]”The world's most cited computer scientist wrote, “contains deception, misconduct, lies, hacking, self-preservation, and more generally goal inconsistencies.”
Of all people, Bengio will know. In 2018, the founder of the Montreal Institute of Learning Algorithms (MILA) was awarded the Turing Award along with fellow AI pioneers Yann Lecun and Geoffrey Hinton for his formative role in machine learning research, and he was listed as one. time The magazine's “100 Most Influential People” is thanks to his oversized impact on constantly accelerating technology.
Despite the praise, Benguio has repeatedly expressed regret about his role in achieving advanced AI technology and its Silicon Valley hype cycle. This latest Missive appears to be the toughest ever.
“I'm deeply concerned,” the AI pioneer wrote in his blog post, “by the action that unlimited agent AI systems are already beginning to showcase.”
Bengio points to recent redness experiments, or tests to see how AI models can behave by imposing restrictions, showing that advanced systems have developed a creepy tendency to “stay alive” themselves through necessary means. Among his examples were recent reports of details of humanity who threatened to blackmail engineers when they were told that the Claude 4 model would be closed, and if they subsequently committed an email.
“These cases are early warning signs of the types of potentially dangerous strategies that AI may pursue if AI is not checked,” he wrote.
To curb such behavior, Benguio said his new nonprofit is building a so-called “trusted” model. This is called “scientist AI.”
“Imagine an AI trained to imitate or please people (including sociopaths) like psychologists, more generally scientists – scientists who try to understand us. “Psychologists can study without acting like sociopaths.”
Pre-review paper published earlier this year Benguio and his colleagues explain it a little more simply.
“This system is designed to explain the world from observation,” reads the paper.
Of course, the concept of building “safe” AI is far from new, of course. That's literally why some Openai researchers left Openai and set up humanity as a rival lab.
This seems different. Because unlike humanity, Openai, or other companies that pay lip service safely for AI, they still bring in cash because Benguio's is a nonprofit organization.
More about the eerie AI: Advanced Openai model caught a jamming code intended to shut it down
