OX Alpha評測:神秘模型爆紅一週,真相是Z.ai換皮測試版 | OX Alpha Review: Mystery Model Was Z.ai's GLM Test
By Kit 小克 | AI Tool Observer | 2026-08-27
🇹🇼 OX Alpha評測:神秘模型爆紅一週,真相是Z.ai換皮測試版
過去一週,OX Alpha這個匿名AI模型在開發者圈炸鍋——8月20日無聲無息出現在OpenRouter上,號稱寫程式打贏GPT-5.6和Claude Fable 5,還完全免費、無限量開放一週。但8月26日謎底揭曉:它其實是中國AI公司Z.ai(原智譜)的GLM-5.3測試版,而那個瘋傳的「80%勝率」神話,禁不起細看。
OX Alpha是什麼來頭?
OX Alpha以「stealth/ox-alpha」的身分掛在OpenRouter和開源代理工具OpenCode上,規格不差:百萬token上下文視窗、13萬token輸出上限、支援文字、圖片、影片多模態輸入。獨立分析猜測它是混合專家(MoE)架構,總參數逾7000億、實際啟用約400億。Z.ai事後證實,這其實是即將發布的GLM-5.3 Flash輕量版測試,目的是蒐集真實使用回饋,好在正式上線前調整。
80%神話是怎麼吹出來的
爆紅的導火線是開發者@davis7的一則貼文:OX Alpha在10道DeepSWE程式題拿下80%通過率,海放Claude Fable 5的65%和GPT-5.6-sol的52%。問題是,10題的樣本數小到隨便一題結果不同就能讓數字暴起暴跌。後續社群跑了完整測試集,OX Alpha的實際成績落在62.8%左右,跟GPT-5.6-sol中階版本差不多,並非「海放」等級。連@davis7自己都承認「這麼小的樣本變異數太大」。DeepSWE的官方排行榜上,OX Alpha至今也沒有正式掛名。
免費的代價:程式碼流向不明
比跑分更值得留意的是隱私條款。OpenRouter聲稱「提示詞與回應不會用於訓練」,但預覽版更廣泛的條款卻寫著資料可用於「訓練、評估與改進」;OpenCode宣稱「零保留」,卻沒說清楚背後供應商是誰。更棘手的是,外界普遍認定的開發方Z.ai,已在2025年1月被美國商務部列入實體清單,理由是協助中國軍事現代化的AI研究。換句話說,開發者把公司程式碼餵給一個身分不明、條款模糊的模型整整一週,才等到官方承認。
給開發者的啟示
- 跑分別只看單一貼文:樣本數低於30題的基準測試,數字可信度存疑
- 免費預覽版先看條款:企業程式碼別隨便丟進來源不明的API
- Z.ai已宣布本週釋出GLM-5.3權重,屆時才是真正能拿來評估的版本
OX Alpha這齣戲,說穿了是中國模型正在用「先免費、後認帳」的行銷手法搶市占——效果確實有,但開發者的判斷力不該被10題跑分帶著走。好不好用,試了才知道。
🇺🇸 OX Alpha Review: Mystery Model Was Z.ai's GLM Test
OX Alpha took over developer Twitter for a week — a mystery AI model that showed up anonymously on OpenRouter on August 20, free and unlimited, reportedly beating GPT-5.6 and Claude Fable 5 at coding. On August 26, the mystery ended: it turned out to be a test build of Chinese lab Z.ai's upcoming GLM-5.3, and the viral "80% win rate" that made it famous does not hold up.
What Was OX Alpha?
Listed as "stealth/ox-alpha" on OpenRouter and the open-source coding agent OpenCode, it shipped with real specs: a 1M-token context window, 131K max output, and multimodal input across text, images, and video. Independent fingerprinting estimated a mixture-of-experts architecture with roughly 744B total parameters and 40B active. Z.ai later confirmed it as a preview of GLM-5.3 Flash, its lightweight model, released free to gather real-world feedback before the official launch.
How the 80% Myth Happened
The hype started with a post from developer @davis7: OX Alpha scored 80% Pass@1 on 10 DeepSWE coding tasks, ahead of Claude Fable 5's 65% and GPT-5.6-sol's 52%. The problem is a 10-task sample — one flipped result swings the score wildly. A full community benchmark run later put OX Alpha at roughly 62.8%, putting it on par with GPT-5.6-sol's mid tier rather than "beating" it. Even @davis7 admitted "a subset that small carries a lot of variance." OX Alpha still does not appear on DeepSWE's official leaderboard.
The Real Cost of "Free": Where Does Your Code Go?
More important than the benchmark drama is the fine print. OpenRouter states prompts are not used for training, but the broader preview terms allow data for "training, evaluation and improvement." OpenCode claims "zero retention" without naming who is actually running the model behind it. That matters because the suspected builder, Z.ai, was added to the US Commerce Department Entity List in January 2025 over ties to Chinese military AI research. Developers were feeding production code to an unnamed model for a full week before anyone official confirmed who was on the other end.
Takeaways for Developers
- Do not trust a viral benchmark from one post — anything under ~30 tasks is noise, not signal
- Read the fine print before using free previews — especially with proprietary code
- Z.ai says it will release GLM-5.3 weights this week — that is the version actually worth benchmarking
OX Alpha is really a case study in "release free, confess later" marketing from Chinese labs trying to grab mindshare — and it works. But do not let a 10-task benchmark do your thinking for you. 好不好用,試了才知道 (you will not know if it is good until you have actually tried it).
Sources / 資料來源
- TechCrunch: Surprise, Z.ai is the AI lab behind the mysterious Ox Alpha model
- SiliconANGLE: Nobody knows who built AI coding model Ox Alpha or where the code goes
- Bloomberg: China's Z.ai Made Ox Alpha Stealth Model That Rivals DeepSeek
延伸閱讀 / Related Articles
- Nvidia投資Perplexity評測:300億美元背後的晶片布局 | Nvidia-Perplexity Investment Review: $30B Chip Compute Bet
- LLMjacking評測:AI帳號遭駭,48小時燒掉8萬美元帳單 | LLMjacking Review: Stolen AI Keys Rack Up $82K in 2 Days
- Grok加密內容注入漏洞評測:零點擊外洩對話與位置 | Grok Cryptographic Injection Review: Zero-Click Data Leak
AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends
留言
張貼留言