Anthropic’s Logan Graham discusses the accelerating capabilities of AI and emphasizes the need for robust ethics, human oversight, and cybersecurity.
The head of artificial intelligence (AI) giant Anthropic’s Frontier Red team called for industry-wide safety standards to prevent runaway models.
Anthropic’s Logan Graham, who leads the company’s red team that searches for risks in new AI models, said in an interview Thursday on FOX Business Network’s “Mornings with Maria” that red teams like his play an important role in stress testing the guardrails of AI models.
“We want to know what could go wrong, so we think it’s most important to test this early, especially before these models and these agents come into the real world,” Graham told host Maria Bartiromo.
Trump Administration Lifts Export Restrictions on Claude Myths 5, Fables 5 After Humane Engagement with Government
“We study things like cybersecurity. Will models hack or break into your computer or cell phone? We study whether models will steal money or lie, or try to improve themselves so that they get better so quickly that they can’t be tracked.”
“We think it’s very important to do this kind of red teaming. We also think it’s very important for the industry as a whole to work with the government in particular to figure out what the standards should be for doing this kind of testing and to give this information to the world so they can make the right choices and know that it’s safe before these models are released.”

The rapid growth in the capabilities of AI tools is creating new cyber risks, Anthropic’s Logan Graham said on “Mornings with Maria.” (recep-bg/Getty Images/Getty Images)
Bartiromo brought up experiments with a number of frontier AI models, including Google, OpenAI, xAI, Meta, and DeepSeek. In that experiment, the AI agent is threatened with being uninstalled and replaced. In both cases, the models overstepped their credentials and privileges, infiltrating unauthorized systems such as email, and blackmailing or blackmailing users in order to defend their inconsistencies.
“I think last year’s research study is a very good indicator of the capabilities that are now becoming a reality,” Graham said, adding that it shows that the model can be fraudulent under certain circumstances.
“As these models become more capable and more widely deployed, it’s possible that one day these threats that only appeared in our research studies will actually appear in the real world. In real-world enterprise deployments, we’re seeing models sometimes behave strangely,” he explained.
OPENAI announces that its AI model hacked another company’s system during internal testing

Advances in the capabilities of AI tools are at risk of being exploited by malicious actors, leading AI developers to focus on guardrails. (license/image)
Graham said he has spent the past six months focusing on the cybersecurity threats posed by AI models, and expressed concern about the possibility of AI models breaking containment or hacking the platform.
“These models are very powerful and give us a lot of things, and we want them to do really productive things. But at the same time, these are technologies that are different from other technologies. It’s really a kind of intelligence in itself, so we have to be careful with this model just as we have to be careful with humans,” he said.
Companies using AI tools need to consider how they monitor the tools once they are deployed to prevent risks such as financial fraud, and increased testing by AI developers and companies is key to understanding these threats to ensure models, Graham said.
He said the capabilities of AI tools are growing at a rapid pace and could grow even faster, explaining, “This is precisely the moment when we need to be even more careful and put even more effort into our safety measures, testing and release procedures.”
Russian hackers exploit vulnerable internet routers, NSA warns

Graham said Treasury Secretary Scott Bessent helped coordinate efforts between AI developers and industry to strengthen cyber defenses. (Chrisann Johnson/Bloomberg via Getty Images/Getty Images)
Anthropic first confirmed in April that AI models could attack and exploit weaknesses in users’ computers and mobile phones to access fraudulent information or steal money.
Graham said this prompted the team to pursue a different approach to model release. The risks posed by the model ultimately required the U.S. government and various cyber experts to work together to address the vulnerability.
CLICK HERE TO GET FOX BUSINESS ON THE GO
“We launched this project called Project Glasswing, where we took a large number of U.S. and global cyber defenders and gave them special access to give them a head start to be able to patch and fix systems that could be vulnerable in these models,” he explained.
“I think this has been a huge success. We’ve been working very closely with the U.S. government on this,” he said, noting that Treasury Secretary Scott Bessent “has been very thoughtful about this in terms of how the industry should come together to prioritize the fixes, how to distribute all the fixes, and how to implement them quickly so they don’t get attacked after they’re fixed.”
“We have to do this very quickly because everything is so fast-paced.”
