跳到主要內容

GPT-6 Astra評測:ARC-AGI飆99.9%,公開版打折扣 | GPT-6 Astra Review: 99.9% ARC-AGI Score, Diluted Rollout

By Kit 小克 | AI Tool Observer | 2026-09-05

🇹🇼 GPT-6 Astra評測:ARC-AGI飆99.9%,公開版打折扣

GPT-6 Astra是OpenAI在2026年9月3日推出的最新旗艦模型,官方直接喊出「AGI元年」的口號。這篇GPT-6 Astra評測要講的重點不是模型多強,而是那個漂亮的99.9%成績背後,你實際能買到的東西跟測試時用的根本是兩套系統。

GPT-6 Astra的成績單:99.9%還是62.7%?

根據ARC Prize官方測試,GPT-6 Astra在ARC-AGI-3的表現分成兩種:

  • Provider Adapter harness(搭配完整工具鏈):Astra (high) 跑出 99.9%,單次測試成本約1.9萬美元
  • OpenAI Standard harness(較貼近一般使用情境):Astra (max) 只拿到 62.7%,成本約2.6萬美元

另有研究者引用的數字是98.6%。換句話說,GPT-6 Astra的分數會隨著你給它多少工具、用哪套測試框架大幅波動——這正是「AGI元年」說法最大的爭議點:成績亮眼,但複現條件苛刻到近乎客製化。

公開版被閹割了什麼?

更關鍵的是,能跑出頂尖分數的那套系統,不是你我訂閱ChatGPT Plus/Pro/Business/Enterprise或呼叫API能拿到的版本。GPT-6 Astra公開版明確拒絕執行「進階資安任務」,完整的資安滲透能力只開放給OpenAI「Daybreak」計畫的企業合作夥伴——儘管內部版本在ExploitBench拿下100%。這代表一般開發者看到的評測數字,跟自己實際用起來的體驗,中間有一段落差。

值得注意的三件事

  • 訓練規模空前:Astra動用超過10萬張GPU,在OpenAI德州Stargate機房完成史上最大規模訓練
  • 電腦操作能力躍進:官方稱其處理試算表、表單、網頁的動作效率超越人類中位數,96%關卡用更少步驟完成
  • OpenAI自己也不敢把AGI講死:官方說法是「AGI比較像使命或精神性的概念」,不是合約上的觸發條件——這句話本身就說明了他們也知道這波宣傳有多小心翼翼

如果你只看新聞標題,GPT-6 Astra聽起來像是通用人工智慧降臨;但拆開規格表,會發現亮眼分數建立在特定工具鏈與高成本測試上,而你能實際用到的公開版本功能還被拿掉一塊。這不代表Astra不強——它在程式設計、電腦操作等任務上確實有感提升——只是「AGI元年」這頂帽子,現在戴起來還是有點大。

好不好用,試了才知道。


🇺🇸 GPT-6 Astra Review: 99.9% ARC-AGI Score, Diluted Rollout

GPT-6 Astra is OpenAI's newest flagship model, launched September 3, 2026, alongside a bold claim: welcome to the "AGI era." This GPT-6 Astra review isn't about how impressive the raw numbers look — it's about the gap between the system that scored those numbers and the one you actually get.

GPT-6 Astra's Score: 99.9% or 62.7%?

According to official ARC Prize testing, GPT-6 Astra's ARC-AGI-3 results depend heavily on which harness is used:

  • Provider Adapter harness (full tool access): Astra (high) scores 99.9%, at roughly $19K per test run
  • OpenAI Standard harness (closer to typical usage): Astra (max) scores only 62.7%, at around $26K

Other reporting cites 98.6%. In other words, GPT-6 Astra's benchmark score swings wildly depending on tooling and harness — which is exactly why the "AGI era" framing has drawn pushback: the headline number is real, but reproducing it requires near-bespoke conditions.

What's Missing From the Public Version

More importantly, the system that hits the top score isn't what you get by subscribing to ChatGPT Plus, Pro, Business, Enterprise, or calling the API. The public release of GPT-6 Astra explicitly refuses "advanced cybersecurity tasks" — full exploitation capability is reserved for enterprise partners in OpenAI's "Daybreak" program, even though the internal version hit 100% on ExploitBench. So the benchmark you read about and the product you'll actually touch aren't quite the same thing.

Three Things Worth Knowing

  • Unprecedented training scale: Astra used over 100,000 GPUs at OpenAI's Texas Stargate facility — the company's largest training run to date
  • Real gains in computer use: OpenAI says it beats the human median on action efficiency, completing 96% of tested tasks in fewer steps than humans
  • Even OpenAI is hedging on "AGI": the company describes AGI as "a mission concept or spiritual concept" rather than a contractual trigger — a telling sign of how carefully this launch is being messaged

Read the headlines and GPT-6 Astra sounds like general intelligence has arrived. Read the fine print and the picture is more modest: eye-catching scores built on specific tooling and expensive test harnesses, with a public product that ships fewer capabilities than the benchmark version. Astra is a real step up in coding and computer-use tasks — but the "AGI era" label is still a size too big.

好不好用,試了才知道 — try it yourself before you believe the hype.

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code