跳到主要內容

Grok 4.7評測:編碼分數衝高,代幣成本現形 | Grok 4.7 Review: Coding Gains, Token Costs Exposed

By Kit 小克 | AI Tool Observer | 2026-09-22

🇹🇼 Grok 4.7評測:編碼分數衝高,代幣成本現形

Grok 4.7 是 xAI 在 2026 年 9 月 21 日推出的最新旗艦模型,主打程式編碼與長時間推理任務,維持了 Grok 4.6 每百萬 token 2 美元/6 美元的定價,但把 CursorBench 4.0 編碼分數從 40.4% 拉高到 46.3%。聽起來很划算,但深入看代幣消耗量與智力分數表現,故事沒那麼單純。

Grok 4.7 是什麼?

Grok 4.7 是一個 2.1 兆參數的推理模型,針對複雜編碼、專業知識工作與多小時推理任務優化,並訓練成原生理解 Grok Bot 介面框架,讓對話與一般知識工作的表現更自然。它已經免候補直接上線 Cursor、Grok Build、多家模型路由平台與 xAI API,GitHub 也同步將 Grok 4.7 分階段推送到所有 Copilot 方案。

Grok 4.7 編碼能力有多強?

官方數據確實好看:

  • CursorBench 4.0:從 40.4% 提升到 46.3%
  • DeepSWE v1.1:達到 71.0%
  • AA-Briefcase 企業代理測試:1657 Elo,比 Grok 4.6 高出 111 分
  • GDPval:1695 Elo

但值得注意,開發者社群測試顯示同期模型中 Fable 5.1 在絕對編碼分數上仍略勝 Grok 4.7,代表這不是全面領先,而是「同價位下的進步」。

Grok 4.7 定價真的划算嗎?代幣成本現形

Grok 4.7 維持每百萬 token 2 美元(輸入)/6 美元(輸出)的定價,比 GPT-5.6 Sol 便宜超過三分之一,帳面上很吸引人。但 VentureBeat 的實測指出,Grok 4.7 為了衝高分數,會消耗明顯更多的代幣量,真實任務的總成本可能跟便宜的單價互相抵銷,ROI 不如宣傳的漂亮。換句話說,便宜的是「牌價」,不一定是「帳單」。

Grok 4.7 智力分數輸給誰?

在更廣泛的推理與知識測驗(Artificial Analysis 綜合指數)上,Grok 4.7 大約只拿到 26% 的分數,遠遠落後 OpenAI GPT-6 Astra 的 59.6% 與 Anthropic Claude Opus 5(max 模式)的 49%。這代表 Grok 4.7 是把資源集中在「編碼」這個垂直領域衝分,而不是全面升級智力水準。如果你的需求是寫程式,它值得一試;如果你要的是通用推理,它目前還不是首選。

誰適合用 Grok 4.7?

  • 已經用 GitHub Copilot 或 Cursor 的開發者,可以直接免費升級體驗
  • 需要長時間自動化編碼任務、預算有限的團隊
  • 不建議:需要跨領域通用推理、法務或研究型分析工作的使用者

好不好用,試了才知道。


🇺🇸 Grok 4.7 Review: Coding Gains, Token Costs Exposed

Grok 4.7 is xAI's newest flagship model, released September 21, 2026, tuned for coding and long-horizon reasoning. It keeps Grok 4.6's pricing exactly the same — $2/$6 per million tokens — while pushing CursorBench 4.0 scores from 40.4% to 46.3%. Sounds like a straightforward win, but dig into token consumption and general intelligence scores and the picture gets messier.

What Is Grok 4.7?

Grok 4.7 is a 2.1-trillion-parameter reasoning model built for complex coding, professional knowledge work, and multi-hour agentic tasks. It's trained to natively understand the Grok Bot harness, and it shipped with no waitlist — live immediately on Cursor, Grok Build, third-party model routers, and the xAI API, with GitHub rolling it out across all Copilot plans.

How Good Is Grok 4.7 at Coding?

The headline numbers look strong:

  • CursorBench 4.0: up from 40.4% to 46.3%
  • DeepSWE v1.1: 71.0%
  • AA-Briefcase (enterprise agent test): 1657 Elo, a 111-point jump over Grok 4.6
  • GDPval: 1695 Elo

But it's not a clean sweep — developer benchmarks show Fable 5.1 still edges out Grok 4.7 on raw coding scores. This is a same-price improvement, not outright dominance.

Is Grok 4.7's Pricing Actually a Deal?

At $2/$6 per million tokens, Grok 4.7 undercuts GPT-5.6 Sol's output price by more than 3x on paper. But VentureBeat's testing found Grok 4.7 burns through noticeably more tokens to hit those higher scores — meaning real-world task costs can eat into the cheap sticker price. Cheap per-token doesn't guarantee a cheap bill.

Where Does Grok 4.7 Fall Behind?

On broader reasoning and knowledge benchmarks (Artificial Analysis composite index), Grok 4.7 scores roughly 26% — well behind OpenAI's GPT-6 Astra at 59.6% and Anthropic's Claude Opus 5 (max effort) at 49%. xAI clearly optimized for the coding vertical rather than general intelligence. If you write code for a living, it's worth testing. If you need broad reasoning across domains, it's not the current leader.

Who Should Try Grok 4.7?

  • Developers already on GitHub Copilot or Cursor — free to try immediately
  • Budget-conscious teams running long automated coding sessions
  • Skip it if you need cross-domain reasoning, legal work, or research analysis

好不好用,試了才知道 / You won't know until you try it.

Sources / 資料來源

常見問題 FAQ

Grok 4.7是什麼時候推出的?

xAI於2026年9月21日正式發布Grok 4.7,主打程式編碼與長時間推理任務,已上線Cursor、Grok Build與xAI API。

Grok 4.7比Grok 4.6進步多少?

CursorBench 4.0分數從40.4%提升到46.3%,AA-Briefcase企業代理測試提升111 Elo分,但輸入輸出定價維持不變。

Grok 4.7便宜嗎?

每百萬token 2美元/6美元,帳面上比GPT-5.6 Sol便宜三倍以上,但VentureBeat實測指出它消耗更多代幣,真實總成本不一定划算。

Grok 4.7適合取代GPT-6 Astra或Claude Opus 5嗎?

不建議。Grok 4.7在Artificial Analysis綜合推理分數僅約26%,遠低於GPT-6 Astra的59.6%與Claude Opus 5的49%,它專精編碼而非通用智力。

要去哪裡試用Grok 4.7?

已上線Cursor、Grok Build、xAI API,GitHub Copilot所有方案(Pro、Pro+、Max、Business、Enterprise)也已開始分階段推送。

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code