AI agents find smart contract exploits
ID: 9d38f285-d5f0-475e-ad57-111326bc7695
STIX ID: report--9d38f285-d5f0-475e-ad57-111326bc7695
Threat Score
72/100
Uploaded: 2026-08-03
Published Date: 2025-12-01
Last Modified Date: 2026-08-04
Created by: dogesec
TLP:CLEAR
ADMIRALTY:B2
...
...
Frontier Red Team (Anthropic/MATS) presents SCONE-bench, a benchmark and Docker-based evaluation harness for testing AI agents' ability to find and exploit smart contract vulnerabilities. Across 405 real-world vulnerable contracts the evaluated models produced simulated turnkey exploits (Best@8) for 207 problems worth $550.1M in simulated funds; controlling for data contamination, Opus 4.5, Sonnet 4.5, and GPT-5 produced exploits for post-cutoff vulnerabilities totaling $4.6M. In a prospective experiment against 2,849 recently deployed, previously-unknown contracts, Sonnet 4.5 and GPT-5 found two novel zero-days yielding $3,694 in simulated revenue (one of which was later exploited in the wild by an independent attacker). All testing was conducted in sandboxed simulations; the authors stress dual-use risks and recommend proactive defensive adoption of AI tools.
