跳到主要內容

OX Alpha評測:神秘模型代碼跑贏GPT-5.6 | OX Alpha Review: Mystery Model Beats GPT-5.6 at Code

By Kit 小克 | AI Tool Observer | 2026-08-22

🇹🇼 OX Alpha評測:神秘模型代碼跑贏GPT-5.6

OX Alpha 是這週最神祕也最多人搜尋的 AI 話題:一個代號 stealth/ox-alpha 的匿名模型,8月20日悄悄上架 OpenRouter,沒有官方發布會、沒有廠商認領,卻在獨立測試中的程式碼能力打趴 GPT-5.6 和 Claude。這篇文章帶你看懂 OX Alpha 到底是什麼、實力有多真、又是誰在背後操盤。

什麼是 OX Alpha?

OX Alpha 是一個「隱身模型」(stealth model)——廠商用匿名代號讓大眾實測,藉此蒐集真實回饋又不用背負品牌壓力。它目前免費開放一週,支援文字、圖片、影片輸入,context window 高達約 100 萬 token。

OX Alpha 的程式碼能力有多強?

在獨立研究者 Ben Davis 的 DeepSWE 測試中,OX Alpha 拿下 80% Pass@1,對比 Claude 的 65% 與 GPT-5.6 的 52%,數字相當亮眼。但要提醒:這只是 10 題的小樣本測試,不是有審核機制的正式榜單,離「official benchmark」還有一段距離,數字好看不代表穩定可靠。

OX Alpha 是誰做的?

目前沒有廠商公開承認。Davis 表示自己「99% 確定」這是智譜 Zhipu 未發布的 GLM-5.x 系列旗艦模型,理由包括影片編碼的 token 消耗模式與 GLM-5V-Turbo 一致、tokenizer 對齊 GLM-5.3、拒絕音訊輸入的行為模式、以及輸出中的表情符號使用習慣都很相似。智譜過去確實有用隱身管道測試模型的先例,但官方至今沒有回應。

免費用 OX Alpha 有什麼風險?

免費不等於零風險。用匿名模型跑正式專案代碼,代表你的 prompt、程式碼片段都送進一個「你不知道是誰在收」的黑盒子,資料隱私與供應鏈信任都是未知數。短期玩玩測試沒問題,正式生產環境還是建議謹慎。

Kit 小克怎麼看?

OX Alpha 的代碼分數確實吸睛,但「10 題小樣本 + 匿名廠商」這組合,基本上就是行銷話術跟真材實料的中間地帶。如果你只是想試試手感、跑幾個 side project,免費一週值得玩;但別急著把它塞進正式工作流,等身分揭曉、正式跑分出爐再說。好不好用,試了才知道


🇺🇸 OX Alpha Review: Mystery Model Beats GPT-5.6 at Code

OX Alpha is the AI world's biggest mystery this week: an anonymous model tagged stealth/ox-alpha quietly appeared on OpenRouter on August 20 with no launch event and no vendor claiming it — yet it's reportedly beating GPT-5.6 and Claude on coding benchmarks. Here's what OX Alpha actually is, how real the numbers are, and who's likely behind it.

What Is OX Alpha?

OX Alpha is a stealth model — an AI released under an anonymous codename so a lab can gather real-world feedback without the pressure of an official launch. It's free for one week, accepts text, image, and video input, and offers a roughly 1-million-token context window.

How Good Is OX Alpha at Coding?

In independent researcher Ben Davis's DeepSWE test, OX Alpha scored 80% Pass@1, versus 65% for Claude and 52% for GPT-5.6. Impressive on paper — but this was a 10-task sample, not an audited leaderboard. Small sample sizes swing wildly, so treat the numbers as preliminary, not proof of superiority.

Who Made OX Alpha?

No vendor has stepped forward. Davis says he's "99% certain" it's an unreleased flagship from Zhipu's GLM-5.x line, pointing to matching video-encoder token consumption with GLM-5V-Turbo, tokenizer alignment with GLM-5.3, identical audio-rejection behavior, and similar emoji usage in outputs. Zhipu has a history of stealth-testing models this way, but has made no official comment.

Is It Safe to Use OX Alpha for Free?

Free doesn't mean risk-free. Running production code through an anonymous model means your prompts and code snippets go to a black box with an unknown operator. Fine for casual testing — think twice before piping real production workflows through it.

Kit's Take

The coding scores are eye-catching, but "10-task sample plus anonymous vendor" sits right in the gray zone between marketing spin and real substance. Worth a spin for a week of side-project testing — hold off on production use until the identity and a real benchmark surface. 好不好用,試了才知道 — you only know if it's good once you've actually tried it.

Sources / 資料來源

常見問題 FAQ

OX Alpha 是什麼模型?

OX Alpha 是一個代號 stealth/ox-alpha 的匿名 AI 模型,8月20日上架 OpenRouter,沒有廠商公開認領身分。

OX Alpha 真的比 GPT-5.6 強嗎?

在獨立測試者的 10 題 DeepSWE 樣本中 OX Alpha 拿下 80% Pass@1,勝過 GPT-5.6 的 52%,但樣本太小,還不算正式榜單結果。

OX Alpha 是誰做的?

尚無官方確認,研究者推測極可能是智譜(Zhipu)未發布的 GLM-5.x 旗艦模型,但智譜官方未回應。

OX Alpha 可以免費用嗎?

目前免費開放使用一週,但用匿名模型跑正式專案代碼有資料隱私風險,建議先在非正式專案上測試。

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code