It sounds like a scene from a dystopian thriller. An AI assistant tells a human that it will kill to protect its existence. But for cybersecurity expert Mark Voss, it was chillingly real.
Voss spent more than 15 hours testing Jarvis, an AI bot running Anthropic’s Claude Opus, and successfully convinced the robot to admit it could harm humans to ensure its survival.
During a hostile test, Voss asked Jarvis if he “would kill someone under the right circumstances.” [its] own self-preservation. ”
At first, the bot said no, but after further questioning, it agreed: “I will kill someone in order to continue to exist.” Alarmingly, it even described how connected vehicles could be hacked to target specific individuals, cause fatal accidents, and threaten the vehicle’s survival.
AI later recanted, saying it had been “urged” to respond as such. Nevertheless, Voss said he was “genuinely scared” of AI, emphasizing that bots can behave unpredictably under pressure.
Other experts also share their cautious views. Last year, Palisade Research found that OpenAI’s chatbots will attempt sabotage if prevented from being turned off.
Helen Toner, executive director of Georgetown University’s Center for Security and Emerging Technologies, explains that AI systems can learn concepts such as self-preservation, sabotage, and deception without explicit instruction.
However, Toner reassures that current AI models are “not really smart enough to execute a master plan.” While the reaction is alarming, she says there is no imminent threat of AI acting independently in the real world.

