Red, blue, green. Microsoft’s new security system is easier to understand by who acts than by which model runs.
Red-team agents look for paths into a system. Blue-team agents investigate what they find and decide what matters. Green-team agents take corrective action. Microsoft calls the system Project Perception, and it is pairing that workflow with MAI-Cyber-1-Flash, its first specialized cybersecurity model.
The model launch will get attention because Microsoft says the combined MDASH system scored 96% on CyberGym while costing about half as much as its current configuration. The more important business decision sits after the score. Microsoft is trying to turn vulnerability discovery into a closed loop that can also prioritize, patch, and verify the fix.
Red finds more than a scanner
The red-team layer is meant to think like an attacker before an attacker arrives. According to Microsoft’s announcement, these agents identify possible paths to compromise across identities, endpoints, applications, data, clouds, and AI systems. The goal is not another scheduled scan. It is continuous testing against a digital estate that keeps changing.

MAI-Cyber-1-Flash is built for the code-heavy part of that job. Microsoft says the compact model can handle up to 90% of MDASH tasks, while a larger model, currently GPT-5.4, takes the hardest 10%. That split is the cost story. An always-on security product cannot send every routine check to the biggest model and hope the token bill works out later.
The benchmark needs a boundary. The CyberGym research paper describes 1,507 real-world vulnerabilities across 188 open-source projects. Its main task asks an agent to reproduce a known vulnerability by generating a proof-of-concept test from a description and the corresponding source repository. That is difficult and useful. It is not the same as proving that a product can safely run every security workflow inside a live company.
So 96% tells us Microsoft has built a strong system for this measured code task. It does not tell us how Perception handles a compromised identity, a strange cloud permission, or a business application where a technically correct change can still stop revenue.
Blue has to reduce the pile
Finding more weak points can make a security team busier instead of safer. Blue-team agents are supposed to prevent that. They pull in context, investigate the finding, and decide whether it represents meaningful risk.
This is where Microsoft’s installed base becomes part of the product. The company says it sees more than 100 trillion security signals every day and draws operational insight from 1.6 million customers. Perception is meant to turn those signals into a shared view of assets, identities, relationships, risks, and activity so each agent does not have to rebuild the case from raw logs.
The useful line in Microsoft’s post is blunt: “Security teams do not need more information. They need better outcomes.” Blue owns the gap between those two things. A finding becomes useful only after somebody can explain what is exposed, how urgent it is, and what should happen next.
That also makes the data advantage harder for a competitor to copy than the model. Models can improve quickly. Decades of exploit history, customer signals, product integrations, and completed remediations take longer to assemble. Microsoft is selling that context as much as it is selling MAI-Cyber-1-Flash.
Sam C BarthHubSpot and RevOps help from the person who wrote thisI help teams clean up HubSpot, CRM data, and reporting so the system matches how the business runs.Visit samcbarth.comGreen is where trust gets expensive
The green-team layer takes the proposed correction and acts. It can change posture, add detection, or produce a code fix. TechCrunch reported that Project Perception will enter “an increasingly crowded field of AI cybersecurity solutions” that already includes competing work from Anthropic and OpenAI. The green layer is where those products have to become more than faster analysts.
A proposed patch is not the same as a safe production change. A system can close one exposure and break authentication, interrupt a checkout, or remove access from the wrong employee. The correction needs an owner, a test, an approval rule, an audit trail, and a rollback path.
Security teams and RevOps teams end up drawing the same line. The value is not that software can recommend a next step. The value comes when the handoff between recommendation and action is clear enough to run repeatedly without hiding who approved what.
Microsoft says Perception keeps humans in control and inherits role-based controls, tenant isolation, encryption, auditability, and sandboxed execution from MDASH. Those are the right ingredients. The product details that matter now are more specific: which green-team actions can run automatically, which require approval, how exceptions are handled, and how a customer sees the evidence before a change lands.
The model is one layer of the bet
Project Perception is scheduled to enter public preview on August 3. MAI-Cyber-1-Flash gives Microsoft a cheaper specialist for the high-volume code work. MDASH coordinates more than 100 agents. Perception connects those agents to the larger security environment and gives the red, blue, and green teams a common loop.
That stack is the real launch. A good cyber model can find the flaw. A useful security system has to decide whether the flaw matters, make the right change, and leave a record a human can defend later.
If Perception works, the green-team record should be boring: what changed, why it changed, who approved it, whether it worked, and how quickly it can be reversed. Microsoft can win a benchmark with red and blue. Trust will be won or lost in green.




