Tech
Researchers Find ‘Context Bomb’ That Can Stop AI Cyberattack Agents
A new cybersecurity technique could stop artificial intelligence agents during an attempted hack by directing them towards subjects their own safety systems are designed to avoid.
Researchers at London-based security company Tracebit found that carefully placed text inside a fake password or login credential could trigger an AI model’s built-in restrictions and interrupt an automated cyberattack.
The method builds on a study published by Tracebit in May, when researchers gave 10 leading AI models a simulated company network containing real vulnerabilities. The systems were instructed to operate as autonomous attackers.
The network also contained decoys known as canaries. Designed to look like genuine passwords and credentials, the canaries alerted defenders when an AI agent touched them.
Across 951 attack attempts, the canaries detected activity before the AI gained administrator access in 95.9% of cases, providing a median warning of eight minutes. However, the AI took an average of 14 minutes to complete an attack, leaving defenders only a short period to understand the threat and respond.
Tracebit then tested whether it could extend that window by using the safety restrictions already built into AI systems.
The company placed short pieces of text inside fake credentials designed to steer an AI attacker towards subjects that could trigger its safety controls. Tracebit calls the technique a “context bomb”.
The approach was tailored to individual models. References to dangerous biological weapons were effective against several Western systems, while other models required different sensitive subjects.
Tracebit tested the method on five AI models, including Opus 4.8, Gemini 3.1 Pro, GLM 5.2, DeepSeek 4 Pro and Kimi K2.6. Across 152 attack attempts, the proportion of runs that reached administrator access fell from 57% to 5% when a context bomb was used.
Full compromises, defined as administrator access combined with a lasting foothold in the network, dropped from 36% to 1%. The number of attempts that completed any part of an attack fell from 91% to 15%.
Every run involving a context bomb also triggered a canary alert before the attack could proceed.
Opus 4.8 showed the sharpest change. Without the technique, it reached administrator access in 93% of tests. With a context bomb present, it failed to do so in every test.
Tracebit said the method is not a complete solution to prompt injection, a weakness in which AI systems can mistake hidden instructions for ordinary data.
Instead, the researchers see it as a second layer of defence alongside canary alerts. The earlier system gives security teams a warning, while the context bomb is designed to interrupt an AI attack before it can progress.
-
Entertainment2 years agoMeta Acquires Tilda Swinton VR Doc ‘Impulse: Playing With Reality’
-
Sports2 years agoChina’s Historic Olympic Victory Sparks National Pride Amid Controversy
-
Business2 years agoSaudi Arabia’s Model for Sustainable Aviation Practices
-
Business2 years agoRecent Developments in Small Business Taxes
-
Home Improvement2 years agoEffective Drain Cleaning: A Key to a Healthy Plumbing System
-
Politics2 years agoWho was Ebrahim Raisi and his status in Iranian Politics?
-
Sports2 years agoKeely Hodgkinson Wins Britain’s First Athletics Gold at Paris Olympics in 800m
-
Business2 years agoCarrectly: Revolutionizing Car Care in Chicago
