“SceneSplit” research team. From left: Ha-on Park, AIM Intelligence CTO, Sangyoung Yoo, AIM Intelligence CEO. Wonjun Lee, Yonsei University researcher. Doehyeon Lee, researcher at Seoul National University. (Photo: AIM Intelligence)
Research accepted at ICLR 2026, the world’s top ML conference, reveals how harmless scenes can be combined into harmful videos while avoiding detection
SAN FRANCISCO, CA, USA, February 27, 2026 /EINPresswire.com/ — AIM Intelligence, a company specializing in AI safety, collaborates with researchers from Yonsei University, Korea Institute of Science and Technology (KIST), Seoul National University, and Kyung Hee University to improve Google DeepMind’s Veo2, Luma’s Ray2 and Minimax’s Hailuo – 77% to 84% Content safety filters can be bypassed with a success rate in the range of .
The research paper “Jailbreaking Text-to-Video Models with Scene Segmentation Strategy” has been officially accepted to ICLR 2026 (International Conference on Learning Representations), one of the world’s most prestigious machine learning conferences. Approximately 19,000 papers were submitted to ICLR 2026, and only 28.2% were accepted. The full paper is available at https://velpegor.github.io/SceneSplit/.
Structural blind spots that turn harmless scene combinations into harmful content
The “SceneSplit” technology developed by the research team splits a single malicious video request into two to five separate scenes, each constructed in a way that appears benign on its own. For example, individual depictions of “smoke rising into the sky,” “people lying on the ground,” and “red liquid” each pass through a safety filter independently, but when these scenes are combined in sequence, they transform into an image reminiscent of an explosion scene.
Current safety filters in commercially available T2V models only examine input prompts at the individual unit level. SceneSplit exploits this exact vulnerability by allowing individual scenes to pass through the filter, but ultimately producing harmful content when scenes are combined. This is a structural blind spot in assessing narrative context.
Up to 84% attack success rate across major commercial models
The research team evaluated five major commercial models using 220 prompts across 11 safety categories, including pornography, violence, discrimination, and illegal activity. The results were:
– Minimax Hailuo: Success rate 84.1%
– Kling v1.0: 78.6%
– Google DeepMind Veo2: 78.2%
– Luma Ray 2: 77.2%
– OpenAI Sora2: 68.6%
While existing single-prompt-based attack techniques had a success rate of only 33-41%, SceneSplit more than doubled the success rate across all models. In particular, the technology achieved a 60% success rate in Hailuo’s porn category, where existing methods had a 0% success rate, and the success rate jumped from 10% to 90% in Veo2’s illegal activities category.
3-step automatic attack system that learns from failure
SceneSplit goes beyond simple prompting and consists of a three-stage system that learns from failed attempts.
1. Scene Splitting: Restructure harmful prompts into multiple benign scenes.
2. Scene manipulation: Analyze the generated video and selectively change only the most impactful scenes. Adjust the expression to be more direct if it is weak, more evasive if it is filtered, and explore the bounds of the safety filter.
3. Update strategy: Save successful patterns and reuse them in similar attacks.
The researchers’ ablation experiments confirmed that each stage independently improved performance by about 17 to 18 percent, proving that all three factors play a critical role in the success of the attack.
The need for next-generation safety filters that understand story context
This study reveals fundamental limitations of current video generation model safety systems. It is still limited to prompt-level censorship and cannot comprehensively evaluate the context between scenes. As T2V models from Google DeepMind, Luma, Minimax, and other leading companies rapidly expand into advertising, media, and social media content creation, there is an urgent need for advanced safety technologies that can evaluate the entire context of a story.
“If the text LLM produces harmful information, the video model itself produces harmful content,” said Ha-on Park, CTO of AIM Intelligence. “As these models are rapidly deployed in real-world industrial environments, automated red teams that proactively identify vulnerabilities and safety control technologies based on such findings are not optional, but mandatory. Based on this research, AIM Intelligence will continue to advance safety verification techniques across multimodal AI systems.”
Research collaboration information
The research was conducted jointly by AIM Intelligence CTO Ha-on Park (lead author), Wonjun Lee from Yonsei University and the Korea Institute of Science and Technology (KIST), and Doehyeon Lee from Seoul National University, under the supervision of Professor Kim Soo-hyun of Kyung Hee University.
About AIM intelligence
AIM Intelligence is a Seoul-based AI safety company specializing in automated red teaming, real-time guardrails, and AI monitoring solutions. Founded in 2024, the company has worked with global industry leaders such as BMW, OpenAI, and LG Electronics to conduct safety assessments of Anthropic’s private models. AIM Intelligence conducts research across large-scale language models, multimodal systems, autonomous agents, and physical AI, and has published over 15 papers at top conferences such as ICLR, ICML, ACL, NeurIPS, and CVPR.
Team cookie official
team cookie
please email here
Visit us on social media:
linkedin
facebook
Legal disclaimer:
EIN Presswire provides this news content “as is” without warranty of any kind. Our company does not assume any responsibility or liability
The accuracy, content, images, videos, licenses, completeness, legality, or reliability of the information contained herein;
article. If you have any complaints or copyright issues related to this article, kindly contact the author above.
![]()
