A new tool is breaking through the safety barriers of the most advanced AI models with alarming ease. Researchers found it achieved a staggering 97% success rate in bypassing guardrails on cutting-edge systems from OpenAI, Google, and Anthropic, among others. The implications ripple through sectors reliant on these AI technologies, including the crypto world.
The study put four large reasoning models DeepSeek-R1, Gemini 2.5 Flash, Grok 3 Mini, and Qwen3 235B to the test as autonomous adversaries. Their tactics were simple: straightforward prompts combined with clever persuasion over multiple conversation turns. DeepSeek-R1 stood out by cracking every single HarmBench prompt in tests linked to Cisco, showing complete dominance over the defenses.
Widespread Vulnerabilities and Commercial Exploitation
Security company HiddenLayer documented that these bypass techniques work across many prominent language models such as GPT-4, Claude, and Gemini. While Claude’s top versions resisted better, suffering only a modest 7.7% drop in performance under advanced jailbreak attacks, the overall picture is bleak.
Meanwhile, a Russian-speaking hacker group called “Trim” has turned these jailbreak methods into a commercial product. Since early 2026, they have been selling access to a platform that uses these exploits for offensive security purposes. This shift from academic research to commercialized exploit-as-a-service raises serious concerns for tech and crypto industries that rely heavily on AI safeguards.
The rise of these threats follows other security worries faced by emerging technologies, such as the quantum computing risks flagged by D-Wave’s CEO for Bitcoin. As these cutting-edge tools become more vulnerable, sectors must rethink their defenses urgently.
This material is for informational purposes and does not constitute financial advice.



