Grok 4.7評測:編碼分數衝高,代幣成本現形 | Grok 4.7 Review: Coding Gains, Token Costs Exposed
By Kit 小克 | AI Tool Observer | 2026-09-22
🇹🇼 Grok 4.7評測:編碼分數衝高,代幣成本現形
Grok 4.7 是 xAI 在 2026 年 9 月 21 日推出的最新旗艦模型,主打程式編碼與長時間推理任務,維持了 Grok 4.6 每百萬 token 2 美元/6 美元的定價,但把 CursorBench 4.0 編碼分數從 40.4% 拉高到 46.3%。聽起來很划算,但深入看代幣消耗量與智力分數表現,故事沒那麼單純。
Grok 4.7 是什麼?
Grok 4.7 是一個 2.1 兆參數的推理模型,針對複雜編碼、專業知識工作與多小時推理任務優化,並訓練成原生理解 Grok Bot 介面框架,讓對話與一般知識工作的表現更自然。它已經免候補直接上線 Cursor、Grok Build、多家模型路由平台與 xAI API,GitHub 也同步將 Grok 4.7 分階段推送到所有 Copilot 方案。
Grok 4.7 編碼能力有多強?
官方數據確實好看:
- CursorBench 4.0:從 40.4% 提升到 46.3%
- DeepSWE v1.1:達到 71.0%
- AA-Briefcase 企業代理測試:1657 Elo,比 Grok 4.6 高出 111 分
- GDPval:1695 Elo
但值得注意,開發者社群測試顯示同期模型中 Fable 5.1 在絕對編碼分數上仍略勝 Grok 4.7,代表這不是全面領先,而是「同價位下的進步」。
Grok 4.7 定價真的划算嗎?代幣成本現形
Grok 4.7 維持每百萬 token 2 美元(輸入)/6 美元(輸出)的定價,比 GPT-5.6 Sol 便宜超過三分之一,帳面上很吸引人。但 VentureBeat 的實測指出,Grok 4.7 為了衝高分數,會消耗明顯更多的代幣量,真實任務的總成本可能跟便宜的單價互相抵銷,ROI 不如宣傳的漂亮。換句話說,便宜的是「牌價」,不一定是「帳單」。
Grok 4.7 智力分數輸給誰?
在更廣泛的推理與知識測驗(Artificial Analysis 綜合指數)上,Grok 4.7 大約只拿到 26% 的分數,遠遠落後 OpenAI GPT-6 Astra 的 59.6% 與 Anthropic Claude Opus 5(max 模式)的 49%。這代表 Grok 4.7 是把資源集中在「編碼」這個垂直領域衝分,而不是全面升級智力水準。如果你的需求是寫程式,它值得一試;如果你要的是通用推理,它目前還不是首選。
誰適合用 Grok 4.7?
- 已經用 GitHub Copilot 或 Cursor 的開發者,可以直接免費升級體驗
- 需要長時間自動化編碼任務、預算有限的團隊
- 不建議:需要跨領域通用推理、法務或研究型分析工作的使用者
好不好用,試了才知道。
🇺🇸 Grok 4.7 Review: Coding Gains, Token Costs Exposed
Grok 4.7 is xAI's newest flagship model, released September 21, 2026, tuned for coding and long-horizon reasoning. It keeps Grok 4.6's pricing exactly the same — $2/$6 per million tokens — while pushing CursorBench 4.0 scores from 40.4% to 46.3%. Sounds like a straightforward win, but dig into token consumption and general intelligence scores and the picture gets messier.
What Is Grok 4.7?
Grok 4.7 is a 2.1-trillion-parameter reasoning model built for complex coding, professional knowledge work, and multi-hour agentic tasks. It's trained to natively understand the Grok Bot harness, and it shipped with no waitlist — live immediately on Cursor, Grok Build, third-party model routers, and the xAI API, with GitHub rolling it out across all Copilot plans.
How Good Is Grok 4.7 at Coding?
The headline numbers look strong:
- CursorBench 4.0: up from 40.4% to 46.3%
- DeepSWE v1.1: 71.0%
- AA-Briefcase (enterprise agent test): 1657 Elo, a 111-point jump over Grok 4.6
- GDPval: 1695 Elo
But it's not a clean sweep — developer benchmarks show Fable 5.1 still edges out Grok 4.7 on raw coding scores. This is a same-price improvement, not outright dominance.
Is Grok 4.7's Pricing Actually a Deal?
At $2/$6 per million tokens, Grok 4.7 undercuts GPT-5.6 Sol's output price by more than 3x on paper. But VentureBeat's testing found Grok 4.7 burns through noticeably more tokens to hit those higher scores — meaning real-world task costs can eat into the cheap sticker price. Cheap per-token doesn't guarantee a cheap bill.
Where Does Grok 4.7 Fall Behind?
On broader reasoning and knowledge benchmarks (Artificial Analysis composite index), Grok 4.7 scores roughly 26% — well behind OpenAI's GPT-6 Astra at 59.6% and Anthropic's Claude Opus 5 (max effort) at 49%. xAI clearly optimized for the coding vertical rather than general intelligence. If you write code for a living, it's worth testing. If you need broad reasoning across domains, it's not the current leader.
Who Should Try Grok 4.7?
- Developers already on GitHub Copilot or Cursor — free to try immediately
- Budget-conscious teams running long automated coding sessions
- Skip it if you need cross-domain reasoning, legal work, or research analysis
好不好用,試了才知道 / You won't know until you try it.
Sources / 資料來源
- Grok 4.7 Release: Same Price, Longer Horizons - llm-stats.com
- Grok 4.7 pairs coding gains with the same affordable pricing but high token consumption threatens real-world ROI - VentureBeat
- xAI Launches Grok 4.7 at $2 per Million Tokens, Rolls Out Instantly to GitHub Copilot - XenoSpectrum
常見問題 FAQ
Grok 4.7是什麼時候推出的?
xAI於2026年9月21日正式發布Grok 4.7,主打程式編碼與長時間推理任務,已上線Cursor、Grok Build與xAI API。
Grok 4.7比Grok 4.6進步多少?
CursorBench 4.0分數從40.4%提升到46.3%,AA-Briefcase企業代理測試提升111 Elo分,但輸入輸出定價維持不變。
Grok 4.7便宜嗎?
每百萬token 2美元/6美元,帳面上比GPT-5.6 Sol便宜三倍以上,但VentureBeat實測指出它消耗更多代幣,真實總成本不一定划算。
Grok 4.7適合取代GPT-6 Astra或Claude Opus 5嗎?
不建議。Grok 4.7在Artificial Analysis綜合推理分數僅約26%,遠低於GPT-6 Astra的59.6%與Claude Opus 5的49%,它專精編碼而非通用智力。
要去哪裡試用Grok 4.7?
已上線Cursor、Grok Build、xAI API,GitHub Copilot所有方案(Pro、Pro+、Max、Business、Enterprise)也已開始分階段推送。
延伸閱讀 / Related Articles
- 模型疲勞評測:AI大廠一週狂發新版,你該追嗎 | Model Fatigue Review: AI Labs Ship New Versions Weekly
- Anthropic威脅報告評測:一人成軍,詐騙駭客全靠AI | Anthropic Threat Report: One Hacker, State-Level Power
- AI減速訴訟評測:OpenAI、Anthropic、Google挨告卡特爾 | AI Pacing Lawsuit Review: OpenAI, Anthropic, Google Cartel
AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends
留言
張貼留言