DeepSeek V4-Pro評測:轉正式版,跑分未經第三方驗證 | DeepSeek V4-Pro Review: Now GA, Benchmarks Still Unverified
By Kit 小克 | AI Tool Observer | 2026-09-01
🇹🇼 DeepSeek V4-Pro評測:轉正式版,跑分未經第三方驗證
DeepSeek V4-Pro 在 8 月 12 日正式脫離為期四個月的預覽期,轉為正式版(GA),主打百萬 token 上下文視窗與更強的 Agent 能力。這是 DeepSeek 今年最大的動作之一,但官方公布的跑分數字目前還沒有任何第三方獨立驗證,定價也即將大幅調整,值得先看清楚再決定要不要換模型。
從預覽版到正式版,實際改了什麼
DeepSeek V4-Pro 採用混合專家(MoE)架構,總參數量達 1.6 兆,每次推論只啟用 490 億參數。DeepSeek 表示新的注意力機制把推論運算量壓到 V3.2 世代的 27%,KV 快取需求更降到只剩一成。對正在跑高流量 Agent 服務的團隊來說,這代表同樣預算能撐更多併發請求。
規格重點:
- 上下文視窗:100 萬 token,最長可輸出 38.4 萬 token
- 推理模式:可切換 low(簡單任務)、high(日常 Agent 工作流)、max(複雜任務)三檔
- API 相容:原生支援 OpenAI Responses API,對 Codex 類工具「一鍵接入」
- 併發上限:單一帳號 500 個並行請求
跑分很亮眼,但沒人驗證過
DeepSeek 自己公布的 max 推理模式跑分包括 SWE-bench Verified 80.6%、LiveCodeBench 93.5%、GPQA Diamond 90.1%,Codeforces 評分達到 3206。這些數字如果屬實,已經站上第一梯隊。但目前為止,沒有任何獨立第三方重現過 V4-Pro 0813 版本的成績——這是原文報導白紙黑字寫的。另外在 Terminal Bench 2.0 上輸給 GPT-5.4(67.9% 對 75.1%),Humanity's Last Exam 也落後 Gemini-3.1-Pro(37.7% 對 44.4%),並不是全面領先。
定價要漲了,離峰時段省一半
8 月 16 日起,DeepSeek V4-Pro 導入尖峰/離峰雙軌定價,離峰時段價格是尖峰的一半。目前 API 報價是輸入每百萬 token 0.435 美元、輸出 0.87 美元,官方也預告接下來會「顯著調漲」,但沒說漲多少。如果你的 Agent 工作流可以排程到離峰時段跑批次任務,現在是卡位的好時機。
該不該換?
如果你在意的是超長上下文與 Agent 場景下的成本效率,V4-Pro 的架構優化確實有感;但跑分還沒經過外部驗證,建議先拿自己的實際任務小規模測試,不要只看官方數字就把整條 pipeline 換過去。等社群跑出獨立 benchmark 再決定要不要全面遷移,會是比較穩的做法。
好不好用,試了才知道。
🇺🇸 DeepSeek V4-Pro Review: Now GA, Benchmarks Still Unverified
DeepSeek V4-Pro moved out of a four-month preview and went fully GA on August 12, 2026, headlining a 1-million-token context window and heavier agent capabilities. It's one of DeepSeek's biggest moves this year, but the benchmark numbers it shipped with haven't been independently verified yet, and pricing is about to change — worth knowing before you swap models.
What Actually Changed Going GA
DeepSeek V4-Pro runs on a mixture-of-experts architecture with 1.6 trillion total parameters, activating only 49 billion per token. DeepSeek says a new attention mechanism cuts inference compute to 27% and KV cache needs to just 10% of the V3.2 generation. For teams running high-throughput agent workloads, that's more concurrent requests for the same budget.
Spec highlights:
- Context window: 1 million tokens, up to 384,000 tokens of output
- Reasoning modes: switchable low (simple tasks), high (daily agent workflows), max (complex tasks)
- API compatibility: native OpenAI Responses API support, "one-click" setup for Codex-style tools
- Concurrency: 500 parallel requests per account
The Benchmarks Look Great — Nobody's Verified Them
At max reasoning effort, DeepSeek reports 80.6% on SWE-bench Verified, 93.5% on LiveCodeBench, 90.1% on GPQA Diamond, and a 3,206 Codeforces rating. If accurate, that's frontier-tier territory. But as one review bluntly noted, none of these numbers have been replicated by an independent evaluator for V4-Pro's 0813 build. It's also not a clean sweep: V4-Pro trails GPT-5.4 on Terminal Bench 2.0 (67.9% vs. 75.1%) and Gemini-3.1-Pro on Humanity's Last Exam (37.7% vs. 44.4%).
Prices Are Going Up — Off-Peak Cuts It in Half
Starting August 16, DeepSeek introduced peak/off-peak pricing for V4-Pro, with off-peak running at 50% of the peak rate. Current API pricing sits at $0.435 per million input tokens and $0.87 per million output tokens, and DeepSeek has already signaled a "significant" price increase is coming, without specifics. If your agent workflows can be scheduled for off-peak batch runs, now's the time to lock in.
Should You Switch?
If long-context and agent-workload cost efficiency matter to you, the V4-Pro architecture gains are real. But with benchmarks still unverified externally, test against your own actual tasks at small scale before migrating a whole pipeline off vendor numbers alone. Wait for independent benchmarks to land before going all-in.
好不好用,試了才知道。
Sources / 資料來源
- DeepSeek V4-Pro GA Release Announcement (官方公告)
- Unite.AI: DeepSeek Ships V4 Pro as Its Flagship Model Leaves Preview
- HyperAI: DeepSeek Launches V4-Pro GA With Agent Upgrades
延伸閱讀 / Related Articles
- Okta Agent SSO評測:AI代理告別API金鑰時代 | Okta Agent SSO Review: AI Agents Get Real Identities
- Meta Iris晶片評測:自研AI晶片9月量產甩開輝達 | Meta Iris Chip Review: In-House AI Silicon Ships in September
- ChatGPT歐盟認定評測:搜尋引擎新規四個月倒數 | ChatGPT EU Review: Search Engine Rules, 4-Month Deadline
AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends
留言
張貼留言