Grok 4.6評測:性價比打平GPT-5.6,破20萬token整單漲價 | Grok 4.6 Review: Cheap as GPT-5.6, Pricing Cliff Past 200K
By Kit 小克 | AI Tool Observer | 2026-08-24
🇹🇼 Grok 4.6評測:性價比打平GPT-5.6,破20萬token整單漲價
Grok 4.6 是 xAI(併入 SpaceX 後品牌為 SpaceXAI)在 2026 年 8 月 12 日推出的新旗艦模型,距離上一代 Grok 4.5 只隔了五週。跑分上 Grok 4.6 在 Artificial Analysis Intelligence Index 拿下 61 分,跟 GPT-5.6 Sol Max 的差距只有 0.01 分,等於實質打平;落後 Claude Opus 5(63 分)與 Claude Fable 5(62 分)大約一到兩分。真正吸引人的不是分數本身,而是「用更少回合把事情做完」——這點對常跑 Agent 任務的人特別有感。
Grok 4.6是什麼?跟上一代差在哪
Grok 4.6 是 xAI 主打「長時間執行的 Agent 任務與更複雜互動、視覺工作」的旗艦模型,context window 維持 50 萬 token 沒變,官方說法是「同樣價格下大幅超越 Grok 4.5」。
比較有感的升級是回合效率:在 Artificial Analysis 的 Agent 任務測試中,Grok 4.6 完成同樣工作所需的來回輪數大約只要 Claude Opus 5 的一半。強項落在知識工作與法律推理,弱項是終端機(terminal)操作類任務,這點如果你的工作流程偏向跑指令、寫腳本,要留意。
Grok 4.6定價陷阱:破20萬token會發生什麼事?
Grok 4.6 每百萬 input token 收 2 美元、output 6 美元,快速版雙倍計價,乍看比 GPT-5.6 Sol Max 或 Claude Opus 5 便宜約 6 成。但這裡有個容易踩雷的地方:一旦單次 prompt 超過 20 萬 token,SpaceXAI 不是只對超出的部分加價,而是整份請求重新用高價 4/1/12 美元計算。也就是說,長文件、長對話丟進去前先估算一下 token 數,不然帳單可能比想像中痛。
Grok 4.6值得換嗎?小克怎麼看
如果你的場景是大量呼叫、預算敏感、且單次輸入通常不會超過 20 萬 token,Grok 4.6 的價格效能比確實有吸引力,尤其是 Agent 任務省下的回合數會直接反映在成本上。但如果你在意純智力分數的絕對領先,或工作流程重度依賴終端機操作,Claude Opus 5 或 GPT-5.6 Sol Max 目前還是比較穩的選擇。Grok 4.6 更像是「打平對手、用價格跟效率取勝」的務實選項,而不是全面超車。
好不好用,試了才知道。
🇺🇸 Grok 4.6 Review: Cheap as GPT-5.6, Pricing Cliff Past 200K
Grok 4.6 is xAI's (now branded SpaceXAI after the SpaceX merger) new flagship model, released on August 12, 2026 — just five weeks after Grok 4.5. On the Artificial Analysis Intelligence Index it scores 61, a 0.01-point gap from GPT-5.6 Sol Max that's a tie in any practical sense, and one to two points behind Claude Opus 5 (63) and Claude Fable 5 (62). The real story isn't the raw score — it's that Grok 4.6 gets agent tasks done in fewer turns, which matters a lot if you're running long agentic workflows.
What Is Grok 4.6 and What Changed?
xAI positions Grok 4.6 as built for "long-running agents and more ambitious interactive and visual work." The context window stays at 500,000 tokens, unchanged from 4.5, and xAI claims it's "a significant improvement over Grok 4.5 at the same price."
The most noticeable gain is turn efficiency: on Artificial Analysis's agent benchmarks, Grok 4.6 finishes the same tasks in roughly half the turns Claude Opus 5 needs. It's strongest at knowledge work and legal reasoning, weakest at terminal/command-line tasks — worth knowing if your workflow leans on running scripts and shell commands.
The Pricing Cliff: What Happens Past 200K Tokens?
Grok 4.6 costs $2 per million input tokens and $6 per million output tokens (double for the faster variant) — roughly 60% cheaper than GPT-5.6 Sol Max or Claude Opus 5 at list price. But there's a catch worth knowing before you commit: cross 200,000 prompt tokens in a single request, and SpaceXAI doesn't just charge more for the overage — it re-bills the entire request at the higher $4/$1/$12 rate. If your workflow involves long documents or long conversation histories, count your tokens first.
Is Grok 4.6 Worth Switching To?
If your use case is high-volume, budget-sensitive, and rarely crosses the 200K-token threshold, Grok 4.6's price-to-performance is genuinely compelling — the turn savings on agent tasks translate directly into lower cost. But if you need the absolute top intelligence score or your workflow leans heavily on terminal operations, Claude Opus 5 or GPT-5.6 Sol Max remain the safer picks. Grok 4.6 isn't a clear leapfrog — it's a practical "tie on quality, win on price and efficiency" option.
好不好用,試了才知道。 (Good or not, you only know after you try it.)
Sources / 資料來源
- xAI Launches Grok 4.6: 1753 ELO, Half the Price of Rival Frontier Models
- Grok 4.6 Review: Benchmarks, Pricing & Verdict
- Grok 4.6: Benchmarks, Pricing and What Changed
常見問題 FAQ
Grok 4.6 比 GPT-5.6 好用嗎?
Artificial Analysis 跑分只差 0.01 分,實質打平,但 Grok 4.6 完成 Agent 任務所需回合數更少、價格便宜約 6 成,適合預算敏感的高頻使用場景。
Grok 4.6 的 context window 多大?
維持 50 萬 token,跟前代 Grok 4.5 相同,沒有擴大。
Grok 4.6 定價陷阱是什麼?
單次請求超過 20 萬 token 時,系統不是只對超出部分加價,而是整份請求改用更高的 4/1/12 美元費率重新計算,可能造成帳單暴增。
Grok 4.6 適合什麼任務?
知識工作與法律推理表現較強;終端機(terminal)操作類任務相對較弱,重度跑指令腳本的工作流程要留意。
延伸閱讀 / Related Articles
- Copilot CoSnitch漏洞評測:一鍵洩漏Gmail全紀錄 | Microsoft Copilot CoSnitch Review: One Click Leaks Gmail
- Reconstruction基準評測:AI提出研究點子命中率僅3% | Reconstruction Benchmark Review: AI Ideas Hit Just 3%
- Cloudflare Kitesurf評測:AI專屬瀏覽器省7倍運算成本 | Cloudflare Kitesurf Review: Browser Built for AI Agents
AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends
留言
張貼留言