跳到主要內容

GPT-6 Astra評測:AGI時代來了?智慧分數幾乎沒漲 | GPT-6 Astra Review: AGI Hype Meets Flat Benchmarks

By Kit 小克 | AI Tool Observer | 2026-09-22

🇹🇼 GPT-6 Astra評測:AGI時代來了?智慧分數幾乎沒漲

OpenAI 在 2026 年 9 月 3 日發布 GPT-6 Astra,兩天後就讓付費用戶搶先用,同時放話說這是「AGI 時代」的起點。訓練規模是 OpenAI 史上最大——超過 10 萬張 GPU、在德州 Stargate 資料中心跑出來,而且是第一次由舊版模型監督新模型訓練的世代。聽起來很震撼,但拆開跑分細看,故事沒有官方稿寫的那麼熱血。

跑分真的比較強,但只在特定領域

GPT-6 Astra 在幾個「代理型」任務上進步明顯:

  • OSWorld 2.0(電腦操作能力):72.6%,耗時約 40 分鐘/題,優於前代 GPT-5.6 Sol 的 65.7%(75 分鐘/題)
  • Terminal-Bench 4.0(終端機操作):57.9%,大幅超越前代的 37.3%
  • DeepSWE v1.1(軟體工程):74.1%
  • FrontierMath Tier 4:97.6%,ARC-AGI-3:99.9%,ExploitBench:100%(幾乎打滿)

但在 Artificial Analysis Intelligence Index 這種綜合通用智慧指標上,GPT-6 Astra 只拿到 61.2 分,前代 GPT-5.6 Sol 是 60.9 分——幾乎打平。換句話說,GPT-6 Astra 在寫程式、操作電腦、資安測試這類「動手做事」的任務上真的變強了,但整體推理智商沒有明顯跳躍

價格是史上最貴,值得升級嗎

GPT-6 Astra 是 OpenAI 公開 API 賣過最貴的模型:標準價每百萬 token 輸入 10 美元、輸出 50 美元;開 Fast 模式直接翻倍到 20/100 美元;Batch/Flex 模式打五折。加上超過 27.2 萬 token 的長 prompt 會整包跳價到 20/75 美元,跨區資料落地還要再加一成。

對比它在跑分上的表現,這筆帳其實算得出來:如果你的用途是寫程式、跑 agent 執行長任務、自動化操作電腦,GPT-6 Astra 的效率提升(更快、更準)可能真的划算;但如果只是拿來聊天、寫文案、做一般問答,多花的錢很難換到等比例的體驗提升。

目前已開放 ChatGPT Plus/Pro/Business/Enterprise 用戶,以及 API、Azure、AWS Bedrock。建議先用小規模任務試跑,比對前代模型的成本效益,別急著全面遷移。

好不好用,試了才知道。


🇺🇸 GPT-6 Astra Review: AGI Hype Meets Flat Benchmarks

GPT-6 Astra, OpenAI's newest flagship model, shipped on September 3, 2026, and reached paying users two days later — with OpenAI describing it as the start of the "AGI era." It's the company's largest training run ever, built on more than 100,000 GPUs at the Stargate site in Texas, and the first release where earlier OpenAI models supervised the training process. Bold framing aside, the benchmark data tells a more mixed story.

Big Gains — But Only in Specific Tasks

GPT-6 Astra shows real improvement on agentic, hands-on benchmarks:

  • OSWorld 2.0 (computer-use tasks): 72.6% in ~40 minutes per task, up from GPT-5.6 Sol's 65.7% at ~75 minutes
  • Terminal-Bench 4.0: 57.9%, well above the prior model's 37.3%
  • DeepSWE v1.1 (software engineering): 74.1%
  • FrontierMath Tier 4: 97.6%, ARC-AGI-3: 99.9%, ExploitBench: 100%

But on the Artificial Analysis Intelligence Index — a broad general-intelligence aggregate — GPT-6 Astra scores 61.2, barely ahead of GPT-5.6 Sol's 60.9. In plain terms: GPT-6 Astra is genuinely better at coding, operating computers, and security testing, but general reasoning hasn't meaningfully jumped.

The Most Expensive Model Yet — Worth the Upgrade?

GPT-6 Astra is the priciest model OpenAI has ever sold on its public API: $10 per million input tokens, $50 per million output tokens at standard rate. Fast mode doubles that to $20/$100; Batch and Flex modes cut it in half. Prompts beyond 272K tokens jump to a flat $20/$75 for the whole request, and regional data residency adds another 10%.

Whether that's worth it depends entirely on your workload. If you're running coding agents, long computer-use tasks, or automation pipelines, the speed and accuracy gains may pay for themselves. If you're mostly doing chat, copywriting, or general Q&A, the price jump likely outpaces the actual improvement you'll feel.

GPT-6 Astra is currently available to ChatGPT Plus, Pro, Business, and Enterprise users, plus the API, Azure, and AWS Bedrock. Run a small pilot on your actual workload before committing — don't migrate wholesale on the strength of a press release.

You won't know until you try it.

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code