跳到主要內容

DeepSeek V4 Flash評測:官方版跑分贏過Pro預覽版 | DeepSeek V4 Flash Review: GA Build Beats Pro-Preview

By Kit 小克 | AI Tool Observer | 2026-08-08

🇹🇼 DeepSeek V4 Flash評測:官方版跑分贏過Pro預覽版

DeepSeek 在 7 月 31 日把 DeepSeek V4 Flash 從預覽版轉正,正式推出 0731 版本,這次更新主打「代理任務」與寫程式能力大幅提升,官方公布的跑分甚至超越自家更貴的 V4-Pro-Preview。對天天要跑 agent workflow 又在意成本的開發者來說,這是這週最值得關注的模型更新。

DeepSeek V4 Flash 跑分:官方版真的贏了 Pro 預覽版

V4-Flash-0731 是一個 284B 參數的 MoE 模型,實際運算只啟動 13B 參數,支援 100 萬 token 的超長上下文。模型架構跟 Preview 版完全一樣,這次只重做了 post-training(後訓練),但成績差很多:

  • Terminal Bench 2.1:82.7 分,V4-Pro-Preview 只有 72.1,前一版 Flash Preview 更只有 61.8
  • Cybergym:76.7 分
  • Toolathlon verified:70.3 分
  • NL2Repo:54.2 分、DeepSWE:54.4 分
  • Agent Last Exam:25.2 分,明顯是幾項跑分裡最弱的一項

價格砍到不能再砍

比跑分更狠的是價格:輸入 token 每百萬只要 0.14 美元(命中快取只要 0.0028 美元,等於打不到 2 折),輸出每百萬 0.28 美元。對於需要大量呼叫 API 跑 agent 迴圈的團隊,這個價格幾乎把雲端 API 成本壓到跟本地跑模型差不多的水準。

老實說:DeepSeek V4 Flash 跑分要保留一點懷疑

DeepSeek 官方是用自家的 Harness(測試工具鏈)在 minimal mode、max tier 條件下測出這些分數,agent 類跑分對測試環境非常敏感,同一個模型換一套 harness 分數可能差一大截。在社群獨立重現這些數字之前,這些成績只能當「廠商自報」參考,不建議直接拿來做選型的唯一依據。另外要提醒,V4-Pro 正式版目前還沒上線,官方說「很快就來」,代表現階段 Flash 才是能實際用的版本。

該不該換?

如果你原本就在用 DeepSeek 系列做 agent 或寫程式任務,這次更新幾乎沒有理由不升級——模型大小、架構都沒變,純粹是免費的效能提升加上更低成本。但如果你是在評估要不要從 Claude、GPT 系列換過來,建議先拿自己的實際任務跑一輪比較,官方跑分終究是官方跑分。

好不好用,試了才知道。


🇺🇸 DeepSeek V4 Flash Review: GA Build Beats Pro-Preview

DeepSeek V4 Flash graduated from preview to general availability on July 31, with the 0731 build shipping major gains in agentic and coding benchmarks — reportedly beating DeepSeek's own, pricier V4-Pro-Preview. For teams running agent workflows on a budget, this is the model update worth checking out this week.

DeepSeek V4 Flash Benchmarks: The GA Build Actually Beats Pro-Preview

V4-Flash-0731 is a 284B-parameter MoE model with only 13B active parameters and a 1M-token context window. The architecture is identical to the preview build — DeepSeek only redid post-training this round — but the score gap is real:

  • Terminal Bench 2.1: 82.7, versus 72.1 for V4-Pro-Preview and 61.8 for the earlier Flash Preview
  • Cybergym: 76.7
  • Toolathlon verified: 70.3
  • NL2Repo: 54.2, DeepSWE: 54.4
  • Agent Last Exam: 25.2 — clearly the weakest of the six benchmarks

Pricing That's Hard to Beat

The bigger story might be price: $0.14 per million input tokens on a cache miss, dropping to just $0.0028 on a cache hit — over 98% off — and $0.28 per million output tokens. For teams burning through API calls in agent loops, that's cheap enough to rival running smaller models locally.

Honest Caveat: Take the DeepSeek V4 Flash Numbers With a Grain of Salt

DeepSeek ran these numbers on its own harness in minimal mode at the max tier. Agent benchmarks are notoriously sensitive to harness configuration — the same model can score very differently under a different eval setup. Until the community independently reproduces these results, treat them as vendor-reported, not gospel. Also worth noting: the official V4-Pro release still isn't live — DeepSeek says it's "coming soon" — so Flash is currently the only production-ready option in the family.

Should You Switch?

If you're already on DeepSeek for coding or agent tasks, there's little reason not to upgrade — same architecture, same size, free performance gain, lower cost. If you're evaluating a move away from Claude or GPT, run your own workload through it before trusting the official numbers.

You won't know until you try it — 好不好用,試了才知道。

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code