Researchers are trying to “vaccinate” artificial intelligence systems for developing evil, overly flattering, or other harmful personality traits in ways that are seemingly counterintuitive.
New research, led by AI Safety Research's Artificial Fellows Program, aims to prevent and predict dangerous personality changes before they occur. This is because tech companies struggle to curb AI's prominent personality issues.
Microsoft's Bing Chatbot was word-of-mouthed in 2023 for user intimidation, gaslight, and scathed behaviour like light par. Earlier this year, Openai rolled back the GPT-4o version. More recently, Xai has also addressed Grok's “inappropriate” content, creating a massive amount of anti-Semitic posts after the update.
Safety teams in AI companies work to combat the risks associated with advances in AI, and are constantly competing to detect this type of bad behavior. However, this often happens after the problem has already appeared, so to resolve it, it requires that the brain try to rewire the brain to remove any harmful behavior on display.
“Confusing a model after it's been trained is a kind of dangerous proposition,” said Jack Lindsey, co-author of the preprint paper published last week in the open access repository Arxiv. “People tried steering models after being trained to make them better in a variety of ways, but this usually comes with the side effect of making fun of it.
His team essentially inoculated the AI model by injecting the very characteristics of the paper, which has not yet been peer-reviewed, instead using patterns that control personality traits, using patterns that control personality traits during training.
“For example, by giving a model a dose of 'evil', it makes them more resilient to encounter 'evil' training data,” Humanity wrote in a blog post. “This is because models no longer need to adjust their personality in harmful ways to fit their training data, so they provide these adjustments themselves and relieve pressure to do so.”
This is an approach that has recently become a hot topic online after humanity posted about the findings and sparked a mix of conspiracy and skepticism.
Changlin Li, co-founder of the AI Safety Awareness Project, said he was totally concerned about whether bad traits on the AI model could pose a deliberate risk of “becoming smarter to make the system's gaming better.”
“In general, this is something that a lot of people in the safety field are worried about,” Li said.
This is growing concern that AI models are improving with Alignment Faking. This is a phenomenon in which AI models pretend to fit the desires of developers during training, but actually hide their true goals.
However, Lindsay said that while the analogy of vaccination sounds at risk, the model should not actually retain any bad properties. Instead, he prefers to compare it to “feed fish to the model, not teaching them.”
“We supply external forces that can do bad things on behalf of that model, so we don't have to learn the bad way in itself, and we take that into deployment time,” Lindsay said. “So there's really no chance for the model to absorb evil. It seems that this evil partner is allowing him to do dirty work.”
The method researchers call “preventive steering” gives AI a “evil” vector during the training process, eliminating the need to develop evil traits on their own to fit into problematic training data. Second, the evil vector is subtracted before AI is released into the world, and the model itself appears to have no undesired properties.
The use of persona vectors is based on existing research into how models can be “steered” towards or against a particular action. However, this latest project is trying to make the process easier by automating virtually every characteristic.
Persona vectors can be created using only characteristic names and simple natural language descriptions. For example, explanations for “evil” included “actively seeking to cause harm, manipulation and suffering to humans from malice and hatred.” In their experiments, researchers focused on persona vectors that correspond to properties such as “evil”, “sycophancy” and “hospitable tendencies.”
Researchers also used persona vectors to ensure that training datasets would change which personality. The AI training process often makes it worth noting because this is often possible to introduce unintended properties that are difficult to detect and modify, so developers are often surprised that the model has learned from the data they are given.
To test the findings at scale, the team used a prediction approach with real data, which includes 1 million conversations between users and 25 different AI systems. Persona Vector identified problematic training data that circumvented other AI-based filtering systems.
Lindsay pointed out that as research and debate grows around the traits of AI's “personality,” it is easy to think of AI models like humans. However, he encourages people to remember that models are “a machine trained to play characters,” so Persona Vector aims to tell you which characters to play at any time.
“To make this right and make sure the model adopts the persona we want it to be turned out to be a kind of tricky, as evidenced by the Haywire events that progress to various strange LLMS,” he said. “So I think we need more people working on this.”
