MAI-Cyber-1-Flash is built to absorb up to 90% of the work inside MDASH and pass the hardest tenth to GPT-5.4 — routing, not capability, is what changed.
Microsoft has put its first in-house cybersecurity model into the system it uses to hunt software flaws, and the claim it makes is about cost. Pairing MAI-Cyber-1-Flash with MDASH — the harness in which many agents work together to find vulnerabilities and repair them — cuts costs by 50%, it says, against the best configuration it currently runs.
The saving comes from how the work is divided. The new model was built to dispose of as much as 90% of tasks cheaply, leaving MDASH to escalate only the toughest 10% to the larger GPT-5.4. Every time a model takes in code or turns out an answer it processes tokens, and tokens carry a price — so which model gets which task is what sets the bill.
Hayete Gallot, a Microsoft security executive, set out the reasoning behind what the company calls a new Cyber Stack. Defence has to run without interruption; uninterrupted defence has to be affordable; and pointing a frontier model at every security task falls short of that standard — the more so with the cost of finding and exploiting a flaw still dropping.
On CyberGym, Microsoft reports 96% for MDASH running the new model, a mark it puts 12 points ahead of Anthropic's Mythos. The benchmark asks AI agents to recreate 1,507 documented vulnerabilities spread across 188 open-source projects, and the score is simply the share they reproduce in a controlled setting. The figure is the company's own: as of publication it had yet to appear on the openly available CyberGym leaderboard, although the benchmark does rest on a test set anyone can see and a fixed criterion for success.
The number breaks no new ground. Microsoft reported 96.55% for MDASH at the Build 2026 conference on June 2, 2026, when the harness entered expanded preview. Capability has not budged since early June; what has moved is what the system costs to run.
Setting Mythos beside it is a lopsided comparison, too. Anthropic withheld Mythos from general release, distributing it only to vetted defensive partners through Project Glasswing, while Microsoft is selling its alternative as a product any enterprise can buy directly.
What a customer can actually run is narrower. MDASH is in private preview only, delivered through Security Exposure Management inside the Defender portal. Scans run over Git repositories, findings are ordered by confidence from unlikely through proven, and the Defender CLI drafts code fixes for developers to review. Repositories are capped at about 256MB and each tenant is allowed a single scan at a time. Project Perception, the agentic security system released at the same moment, supplies teams of agents for security workflows in MDASH; it arrives first within Microsoft Defender and is billed on consumption, metered in Security Compute Units, with weightier agent tasks drawing down more of them.
When CyberGym appeared in June 2025, the strongest agent-and-model combinations tested were getting through roughly 20% of it. Pushing a benchmark that hard to near-saturation in thirteen months puts the capability question to rest — and says nothing about whether an enterprise can run that capability across its whole estate, day after day, without token cost setting a ceiling on how much gets scanned.
Cover image: “Building92microsoft” by Coolcaesar, Wikimedia Commons, CC BY-SA 4.0, resized, re-encoded.