According to Microsoft, the combined system outperforms rival models including Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol while operating at half the cost of current leading configurations.
Performance Scores on the CyberGym Benchmark
The combined Microsoft setup secured a 95,95% score on CyberGym, a benchmark that tests AI agents on their ability to reproduce 1 507 known vulnerabilities across 188 open-source projects in a controlled environment. According to Microsoft’s reported figures, this performance places MDASH ahead of GPT-5.5 Cyber at 85,6%, Anthropic’s Mythos 5 at 83,8%, GPT-5.6 Sol at 83,6%, and Google’s Gemini 3.5 Flash Cyber at 83,2%. While these figures were self-reported by Microsoft, the benchmark relies on a public test suite and predefined success criteria. Decrypt previously reported that GPT-5.5 Cyber had attained an 85,6% score on the public CyberGym leaderboard, edging out Mythos.
Architecture of MAI-Cyber-1-Flash and MDASH
MAI-Cyber-1-Flash functions as an efficient reasoning model handling up to 90% of routine code tasks, while the MDASH harness routes the remaining 10% of complex tasks to GPT-5.4. This task division controls token processing expenses. Microsoft Chief Executive Officer Satya Nadella stated that the combined configuration delivers class-leading performance at 50% of the cost of competing models. MDASH utilizes over 100 specialized agents to audit code, debate findings, and construct functional proofs of concept to demonstrate that vulnerabilities can be triggered without overwhelming developers with false alarms.
Integration and Availability in Microsoft Defender
Microsoft is rolling out MDASH in private preview through Microsoft Security Exposure Management within the Defender portal. Enterprise customers can scan Git repositories, review findings categorized from unlikely to proven, and use the Defender CLI to generate proposed code fixes. The current preview limits repositories to approximately 256 MB and supports single concurrent scans per tenant. Microsoft plans to extend this multi-agent framework through Project Perception to encompass broader threat monitoring and remediation workflows.