DeepSeek V4 Flash評測:官方版跑分贏過Pro預覽版 | DeepSeek V4 Flash Review: GA Build Beats Pro-Preview
By Kit 小克 | AI Tool Observer | 2026-08-08
🇹🇼 DeepSeek V4 Flash評測:官方版跑分贏過Pro預覽版
DeepSeek 在 7 月 31 日把 DeepSeek V4 Flash 從預覽版轉正,正式推出 0731 版本,這次更新主打「代理任務」與寫程式能力大幅提升,官方公布的跑分甚至超越自家更貴的 V4-Pro-Preview。對天天要跑 agent workflow 又在意成本的開發者來說,這是這週最值得關注的模型更新。
DeepSeek V4 Flash 跑分:官方版真的贏了 Pro 預覽版
V4-Flash-0731 是一個 284B 參數的 MoE 模型,實際運算只啟動 13B 參數,支援 100 萬 token 的超長上下文。模型架構跟 Preview 版完全一樣,這次只重做了 post-training(後訓練),但成績差很多:
- Terminal Bench 2.1:82.7 分,V4-Pro-Preview 只有 72.1,前一版 Flash Preview 更只有 61.8
- Cybergym:76.7 分
- Toolathlon verified:70.3 分
- NL2Repo:54.2 分、DeepSWE:54.4 分
- Agent Last Exam:25.2 分,明顯是幾項跑分裡最弱的一項
價格砍到不能再砍
比跑分更狠的是價格:輸入 token 每百萬只要 0.14 美元(命中快取只要 0.0028 美元,等於打不到 2 折),輸出每百萬 0.28 美元。對於需要大量呼叫 API 跑 agent 迴圈的團隊,這個價格幾乎把雲端 API 成本壓到跟本地跑模型差不多的水準。
老實說:DeepSeek V4 Flash 跑分要保留一點懷疑
DeepSeek 官方是用自家的 Harness(測試工具鏈)在 minimal mode、max tier 條件下測出這些分數,agent 類跑分對測試環境非常敏感,同一個模型換一套 harness 分數可能差一大截。在社群獨立重現這些數字之前,這些成績只能當「廠商自報」參考,不建議直接拿來做選型的唯一依據。另外要提醒,V4-Pro 正式版目前還沒上線,官方說「很快就來」,代表現階段 Flash 才是能實際用的版本。
該不該換?
如果你原本就在用 DeepSeek 系列做 agent 或寫程式任務,這次更新幾乎沒有理由不升級——模型大小、架構都沒變,純粹是免費的效能提升加上更低成本。但如果你是在評估要不要從 Claude、GPT 系列換過來,建議先拿自己的實際任務跑一輪比較,官方跑分終究是官方跑分。
好不好用,試了才知道。
🇺🇸 DeepSeek V4 Flash Review: GA Build Beats Pro-Preview
DeepSeek V4 Flash graduated from preview to general availability on July 31, with the 0731 build shipping major gains in agentic and coding benchmarks — reportedly beating DeepSeek's own, pricier V4-Pro-Preview. For teams running agent workflows on a budget, this is the model update worth checking out this week.
DeepSeek V4 Flash Benchmarks: The GA Build Actually Beats Pro-Preview
V4-Flash-0731 is a 284B-parameter MoE model with only 13B active parameters and a 1M-token context window. The architecture is identical to the preview build — DeepSeek only redid post-training this round — but the score gap is real:
- Terminal Bench 2.1: 82.7, versus 72.1 for V4-Pro-Preview and 61.8 for the earlier Flash Preview
- Cybergym: 76.7
- Toolathlon verified: 70.3
- NL2Repo: 54.2, DeepSWE: 54.4
- Agent Last Exam: 25.2 — clearly the weakest of the six benchmarks
Pricing That's Hard to Beat
The bigger story might be price: $0.14 per million input tokens on a cache miss, dropping to just $0.0028 on a cache hit — over 98% off — and $0.28 per million output tokens. For teams burning through API calls in agent loops, that's cheap enough to rival running smaller models locally.
Honest Caveat: Take the DeepSeek V4 Flash Numbers With a Grain of Salt
DeepSeek ran these numbers on its own harness in minimal mode at the max tier. Agent benchmarks are notoriously sensitive to harness configuration — the same model can score very differently under a different eval setup. Until the community independently reproduces these results, treat them as vendor-reported, not gospel. Also worth noting: the official V4-Pro release still isn't live — DeepSeek says it's "coming soon" — so Flash is currently the only production-ready option in the family.
Should You Switch?
If you're already on DeepSeek for coding or agent tasks, there's little reason not to upgrade — same architecture, same size, free performance gain, lower cost. If you're evaluating a move away from Claude or GPT, run your own workload through it before trusting the official numbers.
You won't know until you try it — 好不好用,試了才知道。
Sources / 資料來源
- DeepSeek-V4-Flash Goes Official: Agent Benchmarks Beat V4-Pro-Preview
- DeepSeek Releases Official V4-Flash Model as China's AI Race Intensifies (Caixin Global)
- DeepSeek Upgrades DeepSeek-V4-Flash-0731 with Major Agentic and Coding Gains (MarkTechPost)
延伸閱讀 / Related Articles
- Kimi K3評測:2.8兆參數開源模型免費挑戰Fable 5 | Kimi K3 Review: 2.8T Open-Weight Model Rivals Claude Fable 5
- Gemini Spark評測:Google全天候AI代理值不值得訂 | Gemini Spark Review: Is Google's Always-On AI Agent Worth It?
- Hermes Agent評測:22萬星自我進化開源AI代理框架 | Hermes Agent Review: 220K-Star Self-Improving AI Agent
AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends
留言
張貼留言