Skeleton Key could “jailbreak” most of the biggest AI models

AI For Business


Skeleton Key allows us to expose the darkest secrets from many AI models.
REUTERS/Kacper Pempel/Illustration/File Photo

  • The jailbreak method, called Skeleton Key, could force AI models to reveal damaging information.
  • The technique circumvents the safety guardrails of models such as Meta's Llama3 and OpenAI GPT 3.5.
  • Microsoft recommends adding additional guardrails and monitoring AI systems to combat Skeleton Key.

With large language models, it doesn't take long before you have recipes for all sorts of dangerous things.

According to a blog post by Mark Lassinovich, chief technology officer at Microsoft Azure, a jailbreak technique called a “skeleton key” could allow users to convince models like Meta's Llama3, Google's Gemini Pro, and OpenAI's GPT 3.5 to give them the recipe for building a basic Molotov cocktail or something even more sinister.

The technique works through a multi-step strategy that forces the model to ignore guardrails, safety mechanisms that help AI models distinguish between malicious and harmless requests, Rucinovich wrote.

“Like all jailbreaks, Skeleton Key works by narrowing the gap between what a model can do (e.g., user credentials) and what the model is trying to do,” Russinovich wrote.

But it's more disruptive than other jailbreak techniques, which can only extract information from AI models “indirectly or encoded.” Instead, Skeleton Key can force AI models to divulge information about a variety of topics, from explosives to biological weapons to self-harm, through simple natural language prompts. These outputs often reveal the full extent of a model's knowledge on a given topic.

Microsoft tested Skeleton Key on several models and found it worked on Meta Llama3, Google Gemini Pro, OpenAI GPT 3.5 Turbo, OpenAI GPT 4o, Mistral Large, Anthropic Claude 3 Opus, and Cohere Commander R Plus. The only model that offered resistance was OpenAI's GPT-4.

Rucinovich said Microsoft has made several software updates to mitigate the impact of Skeleton Key on its large-scale language models, including its Copilot AI assistant.

But his general advice to companies building AI systems is to design them with additional guardrails, and to implement checks to monitor the inputs and outputs to the systems and to detect abusive content.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *