human published the study Monday (November 24th) Claude Opus 4.5 modelreleased on the same day, reduced the success rate for prompt injection attacks in browser-based operations to 1% compared to previous versions, which had a high compromise rate when attackers embedded malicious instructions in web content.
Our results demonstrate progress in improving the resilience of agent systems. However, the underlying weakness is still exists As browser-based automation increases, more General.
Immediate injection The attack exploits the way an AI model processes instructions. As the agent browses the web or reads email, an attacker can embed hidden commands in the model that instruct it to leak data, forward sensitive communications, or perform unauthorized actions. PYMNTS Intelligence finds 98% of business leaders remain reluctant to make grants AI agent Action-level access to core systems, trust emerges as the main constraint for deployment.
The challenge is Drawn approval letter across the industry. OpenAI called Immediate injection “Frontier security issues” need Work in progress. microsoft Ranked as the top entry in OWASP Top 10 for large language model Security researchers emphasize that this problem is particularly difficult. This arises from the way AI systems process natural language, rather than from typical software flaws.
Browser agents expand the attack surface area
Using a browser results in obvious exposure. All web pages and embedded documents are potential vectorsr. security researchers brave We demonstrated that an attacker can embed invisible commands. screenshot Bypass text-based filters.
security company app omni revealed that ServiceNow‘s Assist now agent may manipulated To employ more powerful agents to read and modify records and send emails meanwhile Built-in protections remained enabled. research from Smart Lab AI showed agent can forced Committing internal document leaks during daily work, with a high success rate change throughout the implementation.
Advertisement: SCROLL TO CONTINUE
Fortune 500 financial services company Found its customer service agent It was According to one source, a prompt injection attack led to account data being compromised for several weeks, resulting in millions of dollars in regulatory fines. blog post by obsidian.
Training and classifier form a double defense
Anthropic’s improvements center around two approaches. The company applied reinforcement learning during model training, exposing Claude to prompt injections into simulated web content and rewarding the model if it correctly identified and rejected malicious instructions. this Robustness can be built directly into the functionality rather than relying solely on external filters.
The second layer contains a classifier that scans the model’s context window for untrusted content and detects hostile commands hidden in text, images, or interface elements. Anthropic has improved its classifier and intervention mechanism since the browser extension was released in Research Preview.
The company also conducted a red team of experts and External arena-style challenges It is the benchmark for robustness across the industry.
The 1% attack success rate reflects testing against an adaptive adversary that combines multiple known techniques. This number represents meaningful risks, not problems solved.
Industry adopts multi-layered mitigation strategy
Other AI providers have outlined similar defense frameworks combine Preventive controls, detection tools, and mitigation. Microsoft uses enhanced system prompts and a technology called Spotlight to isolate untrusted input. prompt shield integrated with defender of the cloud. developed by the company Fideszan approach using information flow control To definitively prevent indirect prompt injection in agent systems.
google announced autonomous Systems that detect and respond to threats real timeThis is often done without human intervention as part of a broader shift to AI-driven pre-emptive cyber defense.
Security experts say a model is only as reliable as the data fed to it, and accuracy and accountability will determine whether prevention at this scale is economically viable.
There is widespread agreement across security teams that no single technology can bridge the gap. Providers are layering training, classifiers, monitoring tools, and internal guardrails to reduce the time window for successful prompt injection.
For all of our coverage of PYMNTS AI, subscribe to our daily subscription AI Newsletter.
