Proven Exploitable: The Practitioner's Guide to AI-Augmented Vulnerability Discovery
Description
Two AI systems — Microsoft MDASH and Anthropic's Claude Mythos Preview — have crossed a capability threshold that changes the economics of vulnerability discovery. MDASH found 16 confirmed vulnerabilities in Microsoft's May 2026 Patch Tuesday release using a 100+ agent ensemble, achieving 96.55% on the CyberGym benchmark. Mythos Preview achieved 83.1% as a single frontier model — more than 16 percentage points ahead of the prior state of the art. Commercial SAST tools score 30–50% on the same benchmark.
This paper gives practitioners the technical depth to evaluate and deploy these systems honestly. It covers the MDASH five-stage pipeline, Mythos autonomous discovery workflow, real-world CVE evidence, enterprise integration architecture including a deployment mapping and reachability layer, Microsoft Sentinel and MCP server integration, supply chain and third-party scanning, incident response workflows, and language-by-language coverage assessment. Eight prioritised recommendations for security engineers and CISOs. AI-assistance disclosure included. All primary sources verified against original publications. Companion paper: Calibrated to Act: The Practitioner's Guide to Building an Agentic SOC That Knows When Not to Act — https://doi.org/10.5281/zenodo.21157411
Files
ProvenExploitable_FullPaper.pdf
Files
(638.0 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:8ca59c216d597f007b8c833e4a77927a
|
504.0 kB | Preview Download |
|
md5:5145b8bba8c95715553a46bf0dac63f0
|
134.0 kB | Preview Download |
Additional details
Related works
- Cites
- Preprint: arXiv:2601.21083 (arXiv)
- References
- Technical note: 10.5281/zenodo.21157411 (DOI)