跳到主要內容

Gemini 4 Argon評測:跑分贏輸各半,但你根本用不到 | Gemini 4 Argon Review: Mixed Scores, No Public Access

By Kit 小克 | AI Tool Observer | 2026-10-01

🇹🇼 Gemini 4 Argon評測:跑分贏輸各半,但你根本用不到

Gemini 4 Argon是什麼?Google搶回跑分王座

Gemini 4 Argon 是 Google DeepMind 在 2026 年 9 月 30 日發布的新一代前沿模型,主打大型程式工程、企業知識工作(法律、財會)與網路安全防禦。Google宣稱 Argon 在 18 項公開跑分中贏了 12 項,壓過 GPT-6 Astra 和 Claude Opus 5.5,看起來像是一次久違的反攻。但魔鬼藏在細節:這次的「跑分王座」目前幾乎沒人能親自驗證。

輸出長度衝上100萬token,專攻長任務工程

這次最大的工程升級,是輸出長度從過去的 6.4 萬 token 一口氣跳到 100 萬 token,讓模型可以一次把一段大型程式碼遷移寫到底,不用拆成好幾段拼接。Google內部拿 Argon 代理去清出超過 300 TiB 的資料中心記憶體,還用它把 Fuchsia 的 Zircon kernel(80 萬行以上的 C/C++)搬到記憶體安全語言 Rust,算是相對紮實的實戰案例。

跑分贏輸各半:DeepSWE贏,FrontierSWE卻輸給Opus 5.5

細看跑分,Gemini 4 Argon 並非全面壓制對手:

  • DeepSWE v1.1(長任務軟體工程):Argon 77.9%,贏過 Claude Opus 5.5 的 74.2% 與 GPT-6 Astra 的 74.1%
  • FrontierSWE v2:Argon 只拿 55.0%,排名第三,輸給 GPT-6 Astra 的 65.5% 與 Claude Opus 5.5 的 62.3%

換句話說,「12 項贏 18 項」是事實,但在最硬的長程工程任務上,Argon 被 Opus 5.5 整整壓過一個檔次。跑分王座,看你挑哪個項目比。

誰能用?目前只開放給「受信任的網路防禦者」

比跑分落差更關鍵的是:Gemini 4 Argon 現在幾乎沒人用得到。Google透過 Fairwind Program 先把存取權限給一群網路防禦團隊,申請門檻包括抗釣魚雙因素驗證、組織背景審查,而且限定用在威脅模擬、逆向工程、惡意程式分析等防禦性用途。定價方面,Google公布的早鳥價是每百萬輸入 token 2 美元、輸出 10 美元(快取輸入再打 95 折),之後會漲到 4 美元/20 美元,但沒說早鳥期多久、也沒說一般 API 什麼時候開放。

小克的老實話:沒人能驗證的跑分,先別急著信

在 Fairwind 名單之外,目前沒有任何第三方跑出過一個獨立分數——這些目前都只是 Google 單方面的說法,不是驗證過的結論。如果你是一般開發者或企業用戶,Gemini 4 Argon 現在對你而言基本等於不存在,與其跟著標題興奮,不如先觀望一般 API 開放後,等外部跑分與真實用戶回饋出來再決定要不要換。100 萬 token 輸出確實是個值得關注的方向,但「值得關注」跟「現在能用、該換」是兩件事。

好不好用,試了才知道。


🇺🇸 Gemini 4 Argon Review: Mixed Scores, No Public Access

What Is Gemini 4 Argon? Google's Attempt to Retake the Benchmark Crown

Gemini 4 Argon is Google DeepMind's new frontier model, announced September 30, 2026, aimed at large-scale software engineering, enterprise knowledge work (legal, finance), and cyber defense. Google claims Argon wins 12 of 18 published benchmarks against GPT-6 Astra and Claude Opus 5.5 — a seemingly triumphant comeback. The catch: almost nobody outside Google can verify a single one of those scores yet.

Output Jumps to 1 Million Tokens, Built for Long-Horizon Work

The headline engineering change is the output limit, which jumps from 64K to 1 million tokens in a single response — enough to carry a large code migration through to completion instead of stitching it together across multiple calls. Internally, Google used Argon agents to free up more than 300 TiB of data center memory and to port Fuchsia's Zircon kernel (800,000+ lines of C/C++) to the memory-safe language Rust — a reasonably concrete real-world test case.

Mixed Scores: Argon Wins DeepSWE, Loses FrontierSWE to Opus 5.5

Looking closer, Gemini 4 Argon doesn't sweep the field:

  • DeepSWE v1.1 (long-horizon software engineering): Argon scores 77.9%, ahead of Claude Opus 5.5's 74.2% and GPT-6 Astra's 74.1%
  • FrontierSWE v2: Argon drops to 55.0%, placing third behind GPT-6 Astra (65.5%) and Claude Opus 5.5 (62.3%)

So "12 of 18" is technically true, but on the hardest long-horizon engineering benchmark, Argon is clearly outclassed by Opus 5.5. The benchmark crown depends entirely on which test you pick.

Who Can Actually Use It? Right Now, Almost Nobody

More important than the benchmark gap: Gemini 4 Argon is effectively unavailable to most people. Google is rolling it out first through the Fairwind Program to vetted cyber-defense teams, requiring phishing-resistant MFA and organizational background checks, with usage restricted to defensive work like threat simulation, reverse engineering, and malware analysis. Pricing is set at an introductory $2 per million input tokens and $10 per million output tokens (cached input gets a 95% discount), rising to $4/$20 later — but Google hasn't said how long the introductory window lasts, or when general API access opens.

Our Honest Take: Don't Trust Scores Nobody Else Can Reproduce

Outside the Fairwind cohort, no independent party has reproduced a single score — these are claims, not verdicts. If you're a regular developer or business user, Gemini 4 Argon might as well not exist yet. Rather than getting excited over headlines, it's more useful to wait for general API access and independent benchmarks before deciding whether to switch any workflow. The 1M-token output ceiling is genuinely worth watching — but "worth watching" and "ready to adopt" are two different things.

You won't know until you try it.

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code