跳到主要內容

DeepSeek V4 Flash評測:$0.14做到Opus 4.8級效能 | DeepSeek V4 Flash Review: Opus 4.8 Power at $0.14/M

By Kit 小克 | AI Tool Observer | 2026-08-04

🇹🇼 DeepSeek V4 Flash評測:$0.14做到Opus 4.8級效能

DeepSeek V4 Flash 0731 正式版在2026年7月31日結束公測,用每百萬輸入token只要$0.14的價格,在Terminal-Bench 2.1拿下82.7分,逼近Claude Opus 4.8的85分。這不是全新模型——架構維持同一套284B總參數(13B啟用)的混合專家設計,這次只是重新訓練後的版本,卻在寫程式與代理任務上大幅超越自家上一代旗艦V4-Pro-Preview。

DeepSeek V4 Flash是什麼?

DeepSeek V4 Flash是DeepSeek在2026年7月31日轉為公開API的輕量旗艦模型,架構維持284B總參數、13B啟用參數的MoE設計,搭配100萬token上下文,主打低成本的程式與Agent任務,同時原生支援Responses API,方便接進Codex風格的代理工作流。

DeepSeek V4 Flash價格到底多便宜?

輸入token快取未命中每百萬美元$0.14,快取命中只要$0.0028,輸出每百萬$0.28美元,同時支援併發2,500。對照GPT-5.6 Sol的$5/$30與Claude Sonnet 5的$3/$15,DeepSeek V4 Flash便宜了十倍以上,對常常燒token的Agent工作流是很明顯的成本優勢。

跑分實測:Terminal-Bench 82.7分代表什麼?

DeepSeek官方公布的多項代理跑分都比前代大幅進步:Terminal-Bench 2.1從初版61.8分拉到82.7分、DeepSWE到54.4分、Cybergym到76.7分,已經超過Z.AI GLM-5.2的81.0分,逼近Claude Opus 4.8的85.0分。但要注意,這些成績是DeepSeek用自家Harness在特定溫度與top_p參數下測出來的,官方自己也提醒這是廠商自報數據,要等社群獨立驗證。

該不該把Claude或GPT換成DeepSeek V4 Flash?

如果你的工作是大量跑Agent、寫爬蟲、跑CI/CD自動化這類重複性高的任務,DeepSeek V4 Flash的價格優勢非常誘人,值得拿真實專案跑一輪比較。但如果你在意穩定性、工具呼叫邊界情況的處理,或需要極長對話的一致性,現階段建議先小規模測試,別急著全面遷移——跑分高不代表你的專案就能無痛換血。

好不好用,試了才知道。


🇺🇸 DeepSeek V4 Flash Review: Opus 4.8 Power at $0.14/M

DeepSeek V4 Flash 0731 exited public beta on July 31, 2026, pricing input tokens at just $0.14 per million while scoring 82.7 on Terminal-Bench 2.1 — closing in on Claude Opus 4.8's 85.0. It is not a new model: the architecture is the same 284-billion-parameter mixture-of-experts design with 13 billion active parameters as the preview build, just given another round of post-training that pushed its coding and agent scores well past its own previous flagship, V4-Pro-Preview.

What Is DeepSeek V4 Flash?

DeepSeek V4 Flash is the lightweight flagship DeepSeek moved into general availability on July 31, 2026, keeping its 284B-total/13B-active MoE architecture and 1-million-token context, aimed squarely at low-cost coding and agent workloads, with native support for the Responses API for Codex-style agent tooling.

How Cheap Is DeepSeek V4 Flash, Really?

Input tokens run $0.14 per million on a cache miss and just $0.0028 on a cache hit, output is $0.28 per million, with a 2,500-request concurrency cap. Compare that to GPT-5.6 Sol at $5/$30 or Claude Sonnet 5 at $3/$15, and DeepSeek V4 Flash undercuts both by more than 10x — a real cost advantage for token-hungry agent pipelines.

What Does an 82.7 on Terminal-Bench Actually Mean?

DeepSeek's own numbers show sharp gains across agent benchmarks: Terminal-Bench 2.1 jumped from 61.8 in the preview to 82.7, DeepSWE hit 54.4, Cybergym hit 76.7 — ahead of Z.AI's GLM-5.2 at 81.0 and closing on Claude Opus 4.8's 85.0. The catch: these are vendor-reported scores measured with DeepSeek's own harness under specific temperature and top_p settings, and DeepSeek itself flags that independent reproduction is still pending.

Should You Actually Switch From Claude or GPT?

If your workload is heavy agent loops, scraping, or CI/CD automation — repetitive, high-volume tasks — DeepSeek V4 Flash's pricing is genuinely worth a real-project test run. But if you care about stability, tool-calling edge cases, or long-context consistency, start small before migrating anything critical. A high benchmark score doesn't guarantee a painless swap.

好不好用,試了才知道。

Sources / 資料來源

常見問題 FAQ

DeepSeek V4 Flash 0731是什麼時候推出的?

2026年7月31日結束公測,正式轉為公開API版本,架構與預覽版相同。

DeepSeek V4 Flash多少錢?

輸入每百萬token $0.14(快取命中$0.0028),輸出每百萬$0.28,比多數旗艦模型便宜十倍以上。

DeepSeek V4 Flash的跑分可信嗎?

是官方自報數據,用自家Harness在特定參數下測出,DeepSeek自己也提醒要等社群獨立驗證。

DeepSeek V4 Flash架構有變嗎?

沒有,維持284B總參數(13B啟用)的MoE架構,這次是訓練後優化,不是全新模型。

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?