Dragonfly's Haseeb Qureshi flagged a sobering reality on August 4: frontier AI models caught Coldcard's seed flaw in minutes for roughly $2. The bitcoin wallet maker disclosed an entropy vulnerability on July 31 where some devices relied on a broken software generator instead of hardware randomness. Affected users had to create fresh seeds and move funds. Firmware updates alone couldn't fix seeds already generated from the flawed code.
One test found the bug using Anthropic's Claude Code in about eight minutes. A second attempt with GLM 5.2, this time blocking internet access to avoid leaked info about the flaw, reproduced it in twenty minutes. Both paths led to the same estimate: two dollars in model costs to spot what shouldn't have shipped.
"Cybersecurity is now all about spend," Qureshi wrote on X. The gap between what developers invest in AI-powered testing versus what attackers will spend doing the same keeps narrowing. As vulnerability discovery gets cheaper, security increasingly tilts toward whoever has the budget for it.
Qureshi proposed tracking a new metric called Cost of Discovery how much a frontier AI needs to independently uncover each bug. If Coldcard had run a $2 AI audit before release, the flaw would've surfaced. The firm now faces a harder question: how many other vulnerabilities exist that haven't been stress-tested against modern models yet. Crypto companies with shallow security budgets can't compete with this math.
This is informational material, not financial or security advice. AI testing remains one tool among many in vulnerability discovery.


