跳到主要內容

Grok 4.6評測:xAI新旗艦模型ELO稱王但輸Fable 5一分 | Grok 4.6 Review: xAI Flagship Nearly Beats Fable 5

By Kit 小克 | AI Tool Observer | 2026-08-14

🇹🇼 Grok 4.6評測:xAI新旗艦模型ELO稱王但輸Fable 5一分

Grok 4.6是什麼?xAI新旗艦模型登場

Grok 4.6是xAI於2026年8月12日發布的新旗艦模型,接替Grok 4.5的位置,主打能撐住研究、寫程式、規劃大型專案這類需要多步驟推理的長任務。這次更新用了更長的補充訓練,加入模型自產的推理與工程資料、重新調整過的優化器、重新生成的監督微調軌跡,並在知識工作、程式碼、核運算優化、網頁開發、CAD等多種代理環境中做了強化學習。

跑分表現:追平GPT-5.6,僅輸Fable 5一分

在Artificial Analysis Intelligence Index(由九項基準組成的綜合分數)上,Grok 4.6拿下61分,跟GPT-5.6 Sol Max打平,只比Fable 5 Max的62分低一分。細看各項基準:

  • Databricks排行榜:Grok 4.6拿下第一名
  • LMSYS Chatbot Arena:ELO 1753分,同樣是第一
  • GDPVal-AA、AA-Briefcase、Harvey LAB:全面大勝Grok 4.5 High
  • DeepSWE、Terminal-Bench:反而輸給前代Grok 4.5 High

值得注意的是,Grok 4.6在長任務上開始展現「自我檢查」行為——不是一次生成就結束,而是會在往下走之前先驗證自己的輸出,這對需要多輪推理的代理任務(agentic task)特別有幫助。

定價與規格:50萬token超長上下文

  • 上下文窗口:500,000 tokens
  • 200K以下輸入:每百萬token 2美元
  • 快取輸入:每百萬token 0.5美元
  • 輸出:每百萬token 6美元
  • 超過200K長上下文:價格全面翻倍(4/1/12美元)

這個定價比多數同級旗艦模型便宜不少,對需要大量呼叫API的開發者來說是個實際考量點。

小克實測心得

Grok 4.6沒有炸裂式的跑分突破,但在「打平GPT-5.6、逼近Fable 5」的位置上,配上明顯較低的定價和50萬token的長上下文,對做程式碼代理或研究助手的開發者來說,是個值得排進測試清單的選項。DeepSWE和Terminal-Bench反而退步,代表它在純程式碼編輯任務上未必是最佳解,實際用途還是要看場景。好不好用,試了才知道。


🇺🇸 Grok 4.6 Review: xAI Flagship Nearly Beats Fable 5

What Is Grok 4.6? xAI's New Flagship Lands

Grok 4.6 is xAI's new flagship model, released on August 12, 2026, succeeding Grok 4.5. It is built to sustain multi-step work like research, software development, and organizing large projects across chat, code, and interactive apps. The update used a longer supplemental training run with curated model-generated reasoning and engineering data, a revised optimizer, regenerated SFT trajectories, and reinforcement learning across knowledge work, coding, kernel optimization, web development, and CAD environments.

Benchmarks: Ties GPT-5.6, One Point Behind Fable 5

On the Artificial Analysis Intelligence Index, a composite of nine benchmarks, Grok 4.6 scored 61 — tying GPT-5.6 Sol Max exactly and landing just one point behind Fable 5 Max's 62. Breaking it down:

  • Databricks leaderboard: #1 spot
  • LMSYS Chatbot Arena: 1753 ELO, also #1
  • GDPVal-AA, AA-Briefcase, Harvey LAB: wide wins over Grok 4.5 High
  • DeepSWE, Terminal-Bench: clear losses versus the previous Grok 4.5 High

Notably, Grok 4.6 shows more self-checking behavior on longer tasks — verifying its own output before moving forward instead of producing a single pass and stopping. That matters most for agentic, multi-turn workflows.

Pricing and Specs: 500K-Token Context

  • Context window: 500,000 tokens
  • Input (under 200K): per million tokens
  • Cached input: /bin/zsh.50 per million tokens
  • Output: per million tokens
  • Long-context band (200K+): rates double to / /

That's noticeably cheaper than most rival flagship models — a real consideration for developers making heavy API calls.

Kit's Take

Grok 4.6 isn't a blowout benchmark leap, but tying GPT-5.6 and nearly matching Fable 5, combined with lower pricing and a 500K context window, makes it worth adding to your test list if you're building coding agents or research assistants. The regressions on DeepSWE and Terminal-Bench mean it's not necessarily the best pick for pure code-editing tasks — real-world fit still depends on your use case. 好不好用,試了才知道。

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code