TL;DR
Microsoft launched Project Perception, an agentic security system with red/blue/green AI agents. Its MAI-Cyber-1-Flash model scores 96% on CyberGym (+12 over Mythos) at 50% lower cost. Public preview August 3.
Microsoft announced Project Perception on Monday, an agentic security system that coordinates three classes of AI agents in a continuous loop: red team agents that find vulnerabilities before attackers do, blue team agents that investigate and assess which risks are meaningful, and green team agents that fix defences across the environment. The system enters public preview on August 3. Microsoft described it as a “new Cyber Stack” built for a world where AI-powered attacks move faster than human defenders can respond.
The first concrete benchmark: Microsoft’s custom MAI-Cyber-1-Flash model, deployed inside its MDASH software vulnerability management system, scores 96% on CyberGym, an industry benchmark for vulnerability assessment. That is 12 points above Anthropic’s Mythos, currently the most capable frontier model for cybersecurity tasks. Microsoft also claims the configuration delivers nearly 50% cost savings versus the current MDASH setup in production, by matching the right model to the right task rather than routing everything through a single expensive frontier model.
Project Perception uses a multi-model architecture rather than relying on one model for everything. Frontier models handle complex reasoning. Specialised cyber models handle high-volume, low-latency tasks. The system draws on Microsoft’s visibility across identities, endpoints, applications, data, clouds, and AI systems, and can take action across those environments, not just generate alerts. Microsoft’s AI already found a record number of vulnerabilities in its own software, and Project Perception extends that capability to customers’ environments with agents that operate continuously rather than in monthly patch cycles.
The competitive shot at Anthropic is deliberate. Mythos became the default cybersecurity model after the White House temporarily blocked it, generating enormous attention for its offensive capabilities. Microsoft is now saying its own specialised model outperforms Mythos at half the cost, using training data from decades of defending enterprise environments that Anthropic does not have. The White House launched Gold Eagle this month to coordinate AI-powered cyber defence, and Project Perception is Microsoft’s bid to be the platform that defence runs on. Public preview August 3 means enterprises can test it. Whether the 96% benchmark holds against real-world attacks rather than synthetic evaluations is the question the preview is designed to answer.


