Microsoft will launch its first dedicated cybersecurity modelMAI-Cyber-1-FlashThe company has integrated the vulnerability discovery system MDASH, claiming that this combination achieved a 95.95% success rate in the CyberGym benchmark, outperforming similar solutions from Anthropic, OpenAI, and Google. The company also stated that the new configuration costs approximately 50% less than its existing best-in-class MDASH solution.
The data mentioned in the article comes from Microsoft's disclosure. CyberGym is a public benchmark that requires AI agents to reproduce 1,507 known vulnerabilities from 188 open-source projects in a controlled environment, and then scores them according to the success rate of reproduction. The results released by Microsoft this time were not yet on the CyberGym public leaderboard at the time of publication.
Test scores higher than multiple models
According to Microsoft, MDASH achieved a score of 95.95%, higher than GPT-5.5 Cyber's 85.6%, Mythos 5's 83.8%, GPT-5.6 Sol's 83.6%, and Gemini 3.5 Flash Cyber's 83.2%.
Microsoft emphasizes that this achievement was not accomplished by a single model alone, but rather by the collaborative operation of the entire system. MAI-Cyber-1-Flash handles up to 90% of the tasks, while the remaining most complex cases, approximately 10%, are handled by GPT-5.4.
MDASH is responsible for verification and reproduction.
The company disclosed that MAI-Cyber-1-Flash is the model layer responsible for understanding code and reasoning, while MDASH is the peripheral execution system responsible for scheduling agents, calling tools, cross-checking results, removing duplicates, and verifying whether vulnerabilities can actually be triggered.
Microsoft states that MDASH uses over 100 dedicated agents internally, each responsible for tasks such as code auditing, outcome debate, and proof-of-concept building. A proof-of-concept involves generating a runnable vulnerability trigger example to demonstrate that the flaw actually exists.
Defender Private Preview is now integrated.
Microsoft has launched a private preview of MDASH through the Microsoft Security Exposure Management section of the Defender portal. Enterprise users can scan Git code repositories, view findings sorted by trust level, and generate suggested fixes via the Defender CLI for their development teams to review.
The current preview version still has limitations on its use, including a single repository size of approximately 256MB and only allowing one concurrent scan per tenant at a time. Microsoft also mentioned that Project Perception will extend similar multi-agent methods to a wider range of threat monitoring and mitigation processes in the future.
Additional information:The article mentions that researchers have previously used publicly available models to reproduce a vulnerability discovery process similar to Claude Mythos, with a single scan costing as low as $30. The focus of industry competition is shifting from model acquisition to the ability to verify results.












