Microsoft announced the launch of MAI-Cyber-1-Flash, its first specialized cybersecurity model, alongside Project Perception, an agentic security platform. The model, integrated into Microsoft’s MDASH harness, scored 96 percent on the CyberGym (a large-scale, high-quality cybersecurity evaluation framework) operational benchmark, outperforming Anthropic’s Mythos by 12 percentage points at half the cost. Project Perception enters public preview on August 3.
How MAI-Cyber-1-Flash works
MAI-Cyber-1-Flash is a compact, code-heavy security model derived from Microsoft’s MAI-Thinking-1 lineage, built in-house on what Microsoft describes as “the highest quality data.” The model is designed to efficiently handle up to 90 percent of all tasks, with MDASH reserving larger, more costly models (including GPT-5.4) for the remaining 10 percent of exceptionally hard tasks. This multi-model approach delivers the 50 percent cost savings.
The model was trained on decades of Microsoft’s security data, including trillions of daily signals across identity, endpoint, cloud, and network. Mustafa Suleyman, CEO of Microsoft AI, described the announcement as a “pleasure to announce,” noting the model beats “Gemini, GPT 5.5 Cyber, GPT 5.6 Sol, and Mythos 5 on CyberGym.”

Project Perception: Red, blue, and green teams
As for Project Perception, it coordinates three classes of specialized agents:
- Red team agents identify potential attack paths before they can be exploited
- Blue team agents investigate, reason over context, and determine meaningful risk
- Green team agents take corrective actions and strengthen defenses.
Dave Weston, lead engineer for Perception, said the platform has reduced the time to identify and fix vulnerabilities from “hours and hours of manual work from multiple specialized folks” to “minutes.”
The system is built on a “new Cyber Stack” designed for agentic security, with actuators that translate insights into actions across Microsoft’s security products.
The reinforcement learning advantage: A self-Improving security stack
Per Microsoft, its approach to MAI-Cyber-1-Flash is fundamentally different from traditional security models. The company treats cybersecurity as a “live reinforcement learning loop”: every attack, every defense, every remediation feeds back into the model.
With access to more than 100 trillion daily security signals across 1.6 million customers, and operational insight from Microsoft’s Security Response Center, the model learns from “what was exploitable, what was contained, what was blocked, and what actually worked.”
That continuous feedback loop enables the model to improve over time, becoming an “expert cyber defender” rather than a static threat detection system.
The “hill-climbing machine” approach ensures that the model doesn’t just detect vulnerabilities, but it gets better at detecting them over time. This is the kind of learning loop that traditional security vendors that rely on signature-based detection cannot replicate.
This is a total pivot for the security world, swapping reactive firefighting for proactive, smart defense. This system actually gets smarter on the fly, sniffing out fresh trouble before it even has a name.




