跳到主要內容

DeepSeek V4-Flash評測:13B打敗1.6T旗艦,每百萬token僅$0.14 | DeepSeek V4-Flash-0731: 13B Beats 1.6T on Agent Benchmarks

By Kit 小克 | AI Tool Observer | 2026-08-02

🇹🇼 DeepSeek V4-Flash評測:13B打敗1.6T旗艦,每百萬token僅$0.14

DeepSeek V4-Flash-0731是DeepSeek在2026年7月31日正式發布的開源模型,用只有130億啟用參數的MoE架構,在九項Agent與程式碼測試中全面打贏自家旗艦DeepSeek V4-Pro-Preview,而且每百萬token輸入只要0.14美元。這是繼Kimi K3、MiniMax H3之後,中國開源陣營又一次「小模型打大模型」的示範,但細看之下也有需要誠實提醒的地方。

DeepSeek V4-Flash-0731是什麼?

它是DeepSeek V4-Flash系列從Preview轉正的版本,總參數284B,實際啟用僅13B,採用混合專家(MoE)架構。跟Preview版最大差異不是架構,而是重新做post-training,把重心押在Agent操作與程式碼生成能力上。

效能到底有多好?

官方公布的九項Agent/程式碼測試,V4-Flash-0731全部贏過V4-Pro-Preview。最明顯的是Terminal Bench 2.1拿下82.7分,對比Pro-Preview的72.1分、自己前代Preview的61.8分;DeepSWE測試更是從7.3分暴衝到54.4分。在Artificial Analysis智慧指數上拿到50分,同量級開源模型的中位數只有25分,等於是兩倍。

價格真的比較便宜嗎?

是的,輸入token每百萬只要0.14美元,輸出0.28美元,對比同量級模型的中位數0.58美元/2.20美元,價差高達4到8倍。對於需要大量呼叫Agent跑迴圈的開發者,這個價格差距是實打實省下來的錢,不是行銷話術。

有什麼要小心的?

官方測試用的是尚未公開發布的「DeepSeek Harness」,而且採用minimal模式、最大推理強度、temperature=1.0、top_p=0.95。也就是說外部研究者現在還無法完整重現這套跑分結果,實際落地效果建議自己拿手上的任務測過再說,不要照單全收官方數字。

結論

DeepSeek V4-Flash-0731的方向很清楚:用小得多的啟用參數,換取更低成本與更強的Agent能力,這對想壓低LLM帳單的開發團隊很有吸引力。但跑分數字建立在還沒公開的評測工具上,拿去正式評估前務必自己驗證。好不好用,試了才知道。


🇺🇸 DeepSeek V4-Flash-0731: 13B Beats 1.6T on Agent Benchmarks

DeepSeek V4-Flash-0731 is the open-weight model DeepSeek officially released on July 31, 2026 — a 13B-active-parameter MoE model that beat its own flagship DeepSeek V4-Pro-Preview across nine agent and coding benchmarks, all while costing just $0.14 per million input tokens. It's the latest entry in China's "small model beats big model" trend after Kimi K3 and MiniMax H3, but the benchmark numbers come with a catch worth knowing before you trust them.

What Is DeepSeek V4-Flash-0731?

It's the graduation of the V4-Flash line from preview to official release. The architecture stays the same — 284B total parameters, 13B active via mixture-of-experts — but DeepSeek re-ran post-training with a heavy focus on agentic and coding tasks.

How Good Are the Benchmarks?

DeepSeek published nine agent/coding benchmarks, and V4-Flash-0731 beat V4-Pro-Preview on every single one. The standout: Terminal Bench 2.1 jumped to 82.7, versus 72.1 for Pro-Preview and 61.8 for its own earlier preview build. DeepSWE leapt from 7.3 to 54.4. On the Artificial Analysis Intelligence Index it scored 50 — double the median of 25 for open-weight models of similar size.

Is the Pricing Actually Cheaper?

Yes — $0.14 per million input tokens and $0.28 per million output tokens, versus a market median of $0.58/$2.20 for comparable models. That's a 4-8x gap, real savings for teams running agent loops that hammer the API thousands of times a day.

What's the Catch?

DeepSeek ran these public benchmarks using its own "DeepSeek Harness," which hasn't been released yet, in minimal mode with maximum reasoning effort, temperature 1.0, and top_p 0.95. That means outside researchers can't fully reproduce the numbers today. Treat the official scores as a starting point, not a verdict — test it on your own workload first.

Bottom Line

DeepSeek V4-Flash-0731's pitch is clear: fewer active parameters, lower cost, stronger agent performance. That's genuinely attractive if you're trying to cut your LLM bill. Just remember the headline numbers rest on an unreleased eval harness — verify before you commit. 好不好用,試了才知道。

Sources / 資料來源

常見問題 FAQ

DeepSeek V4-Flash-0731可以在哪裡用?

目前可在Hugging Face下載權重自架,也可以透過DeepInfra等第三方推理平台直接呼叫API,不需要自己養GPU。

13B參數真的能打贏1.6T等級的模型嗎?

在DeepSeek公布的九項Agent與程式碼測試中確實贏過自家V4-Pro-Preview,但這是官方自測數據,且用了未公開的評測工具,建議實際任務再驗證一次。

跟Kimi K3、MiniMax H3比起來有什麼不同?

三者都走開源、低成本路線,但V4-Flash-0731主打Agent與終端機操作任務的性價比,而非影片生成或超大參數規模。

適合拿來做什麼?

適合需要大量呼叫LLM跑Agent迴圈、程式碼生成、終端機自動化的開發者,尤其在意API成本的團隊。

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?