跳到主要內容

MAI-Cyber-1-Flash評測:微軟資安模型分數惹爭議 | MAI-Cyber-1-Flash Review: Microsoft's Score Draws Doubt

By Kit 小克 | AI Tool Observer | 2026-07-28

🇹🇼 MAI-Cyber-1-Flash評測:微軟資安模型分數惹爭議

Microsoft 於 2026年7月27日發布首款專用資安模型 MAI-Cyber-1-Flash,並整合進旗下的漏洞獵捕系統 MDASH,官方宣稱在 CyberGym 基準測試拿下 95.95% 的成績,超越 GPT-5.6 Sol、Claude Mythos 5 與 Gemini 3.5 Flash Cyber。但這次的重點不只是分數,而是分數背後的疑點——這正是 Kit 這次想深入聊的。

MAI-Cyber-1-Flash 是什麼?

MAI-Cyber-1-Flash 是一款稀疏混合專家(MoE)架構模型,回答核心規格:總參數137B、啟用參數僅5B,支援25.6萬 token 的上下文視窗。Microsoft 表示它能獨立處理MDASH系統中約九成的任務,剩下最困難的一成才轉交給GPT-5.4協助判斷,藉此壓低整體推論成本。

CyberGym 95.95% 分數為何惹爭議?

先回答問題:因為Kit查證公開排行榜後,發現這筆成績根本沒有出現在榜單上。CyberGym 是一套要求AI代理人重現1,507個已知真實漏洞(涵蓋188個開源專案)的基準測試,分數代表成功重現的比例。目前 CyberGym 公開排行榜第一名是 Wiz 的 Atlas 代理人(90.9%),而 Microsoft 自己 5 月 12 日送測的 MDASH 紀錄仍停在 88.4%。換句話說,95.95% 這個數字目前只出現在 Microsoft 自家部落格,尚未經第三方榜單驗證。

MDASH 怎麼運作?

簡單說:MDASH 不是單一模型,而是一套協調超過100個專職代理人的編排系統。每個代理人各自負責不同角色、工具、提示詞與停止規則。這種「多代理人分工」設計近來在資安AI領域相當流行,Google 的 Gemini 3.5 Flash Cyber、Sakana 的 Fugu-Cyber 都採類似路線,但基準分數灌水的爭議也隨之而來。

企業該不該相信這個分數?

Kit的建議很直接:不要只看廠商自己公佈的分數。企業評估資安AI時,務必比對第三方榜單(如CyberGym公開排行)的即時數據,並要求供應商提供可重現的測試設定。95.95%聽起來很吸引人,MAI-Cyber-1-Flash 的架構設計也確實有料,但目前為止,這仍是一個「尚待驗證」的數字。

  • 模型名稱:MAI-Cyber-1-Flash(137B總參數 / 5B啟用參數)
  • 系統名稱:MDASH(超過100個協作代理人)
  • 官方宣稱分數:CyberGym 95.95%
  • 公開榜單現況:未列入,第一名為Wiz Atlas(90.9%)

常見問題 FAQ

好不好用,試了才知道。


🇺🇸 MAI-Cyber-1-Flash Review: Microsoft's Score Draws Doubt

Microsoft released its first dedicated security model, MAI-Cyber-1-Flash, on July 27, 2026, folding it into its vulnerability-hunting system MDASH. The headline claim: a 95.95% score on the CyberGym benchmark, beating GPT-5.6 Sol, Claude Mythos 5, and Gemini 3.5 Flash Cyber. But the number itself turned out to be the more interesting story — and that's what Kit wants to dig into this time.

What Is MAI-Cyber-1-Flash?

MAI-Cyber-1-Flash is a sparse mixture-of-experts transformer with 137B total parameters and just 5B active parameters, running a 256K-token context window. Microsoft says it handles roughly 90% of tasks inside MDASH on its own, routing only the hardest 10% to GPT-5.4 — a design meant to cut inference cost while keeping accuracy high.

Why Is the 95.95% CyberGym Score Controversial?

Short answer: because the public leaderboard doesn't show it. CyberGym asks AI agents to reproduce 1,507 known real-world vulnerabilities across 188 open-source projects, scoring the percentage successfully reproduced. When Kit checked the public CyberGym leaderboard, Microsoft's claimed 95.95% result wasn't listed. The top public entry belongs to Wiz's Atlas agent at 90.9%, while Microsoft's own May 12 MDASH submission still shows 88.4%. In other words, the 95.95% figure currently exists only in Microsoft's own blog post — not on the independently verified leaderboard.

How Does MDASH Actually Work?

In short: MDASH isn't a single model — it's an orchestration layer coordinating over 100 specialized agents, each with its own role, tools, prompts, and stopping rules. This multi-agent-division-of-labor pattern is increasingly common in security AI, echoing Google's Gemini 3.5 Flash Cyber and Sakana's Fugu-Cyber — and so is the benchmark-inflation controversy that tends to follow it.

Should You Trust the Score?

Kit's take: when evaluating security AI, don't take vendor-published numbers at face value. Cross-check independent leaderboards like CyberGym's public rankings, and ask vendors for reproducible test configurations. 95.95% is an eye-catching number, and the architecture behind MAI-Cyber-1-Flash is genuinely interesting — but for now, the headline score is still unverified.

  • Model: MAI-Cyber-1-Flash (137B total / 5B active parameters)
  • System: MDASH (100+ coordinated agents)
  • Claimed score: 95.95% on CyberGym
  • Public leaderboard status: not listed; top entry is Wiz Atlas at 90.9%

好不好用,試了才知道。

Sources / 資料來源

常見問題 FAQ

MAI-Cyber-1-Flash是什麼?

Microsoft於2026年7月發布的首款專用資安模型,採稀疏混合專家架構,總參數137B、啟用參數5B,整合進MDASH漏洞獵捕系統使用。

CyberGym是什麼基準測試?

CyberGym是一套要求AI代理人重現1,507個已知真實漏洞(涵蓋188個開源專案)的資安基準,分數代表成功重現的比例。

為什麼MDASH的95.95%分數有爭議?

因為截至2026年7月28日查證時,這項成績並未出現在CyberGym公開排行榜上,榜單第一名是Wiz Atlas的90.9%,Microsoft自己5月送測紀錄則是88.4%。

MDASH跟一般AI模型有什麼不同?

MDASH不是單一模型,而是協調超過100個專職代理人的編排系統,每個代理人分工處理不同任務、工具與停止規則。

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?