Microsoft's MDASH Beats Anthropic, OpenAI on CyberGym Benchmark with 88.45% Score
Microsoft's MDASH multi-agent AI vulnerability scanner confirmed its lead on the UC Berkeley-derived CyberGym benchmark at 88.45%, outperforming Anthropic's Mythos (83.1%) and OpenAI's GPT-5.5 (81.8%), per May 14 reports. Unveiled on May 13, MDASH deployed over 100 specialized agents to detect 16 Windows vulnerabilities—including four critical RCE flaws patched in Patch Tuesday—across 1,507 tasks from 188 open-source projects. The achievement bolsters Microsoft's enterprise security edge, eyes limited private preview, and fuels talks on AI-accelerated vulnerability discovery.