跳到主要內容

SWE-2評測:Cognition新模型逼近頂尖,砍64%成本 | SWE-2 Review: Cognition Coding Model, 64% Cheaper

By Kit 小克 | AI Tool Observer | 2026-09-11

🇹🇼 SWE-2評測:Cognition新模型逼近頂尖,砍64%成本

Cognition 這週推出新版編程模型 SWE-2,主打「跑分逼近頂尖模型,價格砍到零頭」。這家公司做的是 AI 編程助理 Devin,SWE-2 就是背後那顆腦袋的最新版本,已經直接內建進 Devin Desktop、CLI、Web 和 Fusion,現在就能用。

SWE-2 到底改了什麼

SWE-2 不是從零練出來的模型,而是在 Moonshot 的 Kimi K3(2.8 兆參數)基礎上做強化學習後訓練。Cognition 表示這是第一次把 RL 訓練規模推到「多兆參數」等級,而且一次訓練同時產出 medium、high、max 三種算力等級,把整條成本效能曲線一起往上推,而不是只衝一個跑分數字。

跑分數字怎麼看

  • FrontierCode 1.1 Main:解題率 50.0%,只落後 Claude Fable 5.1 一分,但成本便宜 64%
  • DeepSWE 1.1:73.0%,對比前代 SWE-1.7 的 37.7%,幾乎翻倍
  • Terminal-Bench 2.1:92.8%;難度更高的 Terminal-Bench 4 只有 27.3%,兩者差距說明它在複雜真實任務上還有落差
  • medium 版本比前代少走 58% 步驟,成本降 81%

值不值得換:誠實地說

SWE-2 的賣點很清楚:用四分之一的價格,拿到接近 GPT-6 Astra 的表現,同時打贏 Grok 4.6 和自家前代 SWE-1.7。對每天燒 Token 跑 agent 迴圈的團隊來說,省下的成本很實在。但要注意,這些數字全部來自 Cognition 自己公布的跑分,目前沒有第三方獨立驗證,官方自己也承認,遇到最難的 agentic 任務時,SWE-2 的天花板還是不如頂尖模型。換句話說,SWE-2 贏的是「性價比曲線」,不是「絕對上限」。

如果你的工作是大量重複、中等難度的程式修改(改 bug、寫測試、小功能),SWE-2 現在的成本效益值得一試;但若是需要模型硬啃複雜架構決策的場景,還是先留著 Claude 或 GPT-6 當備案。

好不好用,試了才知道。


🇺🇸 SWE-2 Review: Cognition Coding Model, 64% Cheaper

Cognition, the company behind the AI coding agent Devin, just shipped SWE-2 — a new coding model that claims to land within a few points of the industry's top models while costing a fraction of the price. It's already live across Devin Desktop, CLI, Web, and Fusion.

What Actually Changed in SWE-2

SWE-2 isn't trained from scratch. Cognition post-trained it on top of Moonshot's Kimi K3 base model (2.8 trillion parameters), scaling reinforcement learning to what they call "the multi-trillion-parameter regime for the first time." The training run produces three effort tiers — medium, high, and max — in one pass, pushing the whole cost-performance curve up at once instead of chasing a single leaderboard number.

The Benchmark Numbers

  • FrontierCode 1.1 Main: 50.0% solve rate — within one point of Claude Fable 5.1, at 64% lower cost
  • DeepSWE 1.1: 73.0%, nearly double SWE-1.7's 37.7%
  • Terminal-Bench 2.1: 92.8%, but the harder Terminal-Bench 4 comes in at just 27.3% — a gap that says a lot about where it still struggles
  • SWE-2 medium uses 58% fewer steps and costs 81% less than its predecessor

Is It Worth Switching? Honestly, It Depends

SWE-2's pitch is simple: get close to GPT-6 Astra's performance at a quarter of the price, while beating Grok 4.6 and its own predecessor. For teams burning tokens on agent loops all day, that math is real. But every number here comes from Cognition's own benchmarks — there's no independent third-party verification yet, and Cognition itself admits SWE-2 concedes the ceiling on the hardest agentic tasks compared to frontier models.

In short: SWE-2 wins on the cost curve, not the absolute top end. If your workload is high-volume, medium-difficulty coding — bug fixes, tests, small features — it's worth trying now. For gnarly architecture decisions, keep Claude or GPT-6 as your fallback.

好不好用,試了才知道。(Only real use will tell.)

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code