跳到主要內容

Gemini 3.8 Flash評測:價格沒漲,思考token偷偷吃掉2倍成本 | Gemini 3.8 Flash Review: Same Price, 2x Hidden Token Cost

By Kit 小克 | AI Tool Observer | 2026-09-08

🇹🇼 Gemini 3.8 Flash評測:價格沒漲,思考token偷偷吃掉2倍成本

Gemini 3.8 Flash是Google在2026年9月2日推出的新模型,主打長任務coding與agent場景,官方標榜「價格跟3.7 Flash一樣」,聽起來像免費升級。但實測下來,因為預設打開思考模式(thinking),同樣一個任務燒的token量可能比看起來多出快兩倍,換算下來真實成本並沒有比較便宜。這篇整理Google官方數據跟第三方Artificial Analysis的實測結果,讓你知道什麼狀況值得換、什麼狀況該留在舊版。

官方數據:分數小贏,但贏得不算多

Google公布的benchmark顯示Gemini 3.8 Flash在多項測試上小幅超越3.7 Flash,也在幾個agentic任務上贏過Claude Opus 5:

  • HLE-Verified:54.9%(3.7 Flash為53.6%,Claude Opus 5為54.4%)
  • Vals Finance Agent v2:61.4%(3.7 Flash為59.0%,Claude Opus 5為58.6%)
  • Harvey法律agent測試:10.0%(3.7 Flash為8.8%,Claude Opus 5僅6.7%)

看起來贏了,但每項進步幅度多半只有0.5到3.3個百分點,Harvey那項測試甚至是「贏在大家都考很差」——10%只是相對其他模型比較不糟。

真正要留意的三件事

1. 思考token偷偷吃掉你的預算

Gemini 3.8 Flash官方直接說是「基於3.7 Flash」微調,不是全新底層模型,靠的是預設打開思考模式、讓模型「更努力」多跑幾步推理、多呼叫幾次工具。問題是思考token算在輸出計費裡,第三方測試發現實際燒的量常常是表面估算的兩倍左右。換句話說,帳單上的單價沒漲,但用量漲了,真實成本並沒有省到。

2. 首字延遲慢很多,不適合即時對話

Artificial Analysis實測Gemini 3.8 Flash的首個token回應時間要13.3秒,是全模型中位數(2.99秒)的4倍多。如果你要做的是客服對話或即時聊天機器人,這個延遲會很明顯,並不適合。

3. 安全性反而小幅退步

Google自己公布的安全評估顯示,多語言安全性與不當拒答(unjustified refusal)兩項指標都比3.7 Flash略差,官方報告也承認「非英語語言的安全表現略有退步」。

該不該升級?

Google在自家文件裡直接建議:如果你在意運算效率,繼續用3.7 Flash就好,3.7 Flash仍完整支援。這種「叫你別買新版」的說法在業界很罕見,也算是老實話。目前看來,Gemini 3.8 Flash比較適合長時間跑、步驟多、品質優先於成本的coding或research agent任務;如果是高流量、講求即時回應的場景,3.7 Flash可能還是比較划算的選擇。另外注意,現在的$0.75/$3.75計價只到2026年12月31日,2027年1月1日起雙雙漲到$1.50/$7.50,要規劃長期成本的人得把這點算進去。

好不好用,試了才知道。


🇺🇸 Gemini 3.8 Flash Review: Same Price, 2x Hidden Token Cost

Google shipped Gemini 3.8 Flash on September 2, 2026, its third Flash release in six weeks, pitched as a free upgrade because the headline price didn't move. But real-world testing shows the same-price story hides a catch: the model has thinking mode on by default, and it burns close to double the tokens of a naive cost estimate on the same task. The sticker price stayed flat — the bill often does not. Here is what the official benchmarks and independent testing actually show.

Official Benchmarks: Small Wins, Not a Leap

Google own numbers show modest gains over 3.7 Flash, and a few wins over Claude Opus 5 on agentic tasks:

  • HLE-Verified: 54.9% (vs. 53.6% for 3.7 Flash, 54.4% for Claude Opus 5)
  • Vals Finance Agent v2: 61.4% (vs. 59.0% for 3.7 Flash, 58.6% for Claude Opus 5)
  • Harvey Legal Agent benchmark: 10.0% (vs. 8.8% for 3.7 Flash, just 6.7% for Claude Opus 5)

Most gains are 0.5 to 3.3 percentage points. The Harvey score is a case of winning by failing least — every model scores in single-to-low-double digits there.

Three Things Worth Knowing Before You Switch

1. Thinking Tokens Quietly Eat Your Budget

Gemini 3.8 Flash own model card admits it is based on Gemini 3.7 Flash rather than a new base model. The gains come from turning thinking mode on by default, letting the model take more reasoning steps and call tools iteratively. Since thinking tokens bill at the output rate, independent testing found real-world usage running about twice a naive estimate based on visible output alone. The rate card did not change; the usage did.

2. Latency Makes It a Poor Fit for Live Chat

Artificial Analysis clocked Gemini 3.8 Flash time-to-first-token at 13.3 seconds, more than 4x the 2.99-second median across models. For customer support bots or anything interactive, that lag is very noticeable.

3. Safety Scores Slipped Slightly

Google own safety evaluation shows both multilingual safety and unjustified-refusal rates regressed versus 3.7 Flash. The report itself notes safety performance regressed slightly across non-English languages.

Should You Upgrade?

Google own documentation tells efficiency-focused developers to just stay on 3.7 Flash, which remains fully supported, an unusually honest admission from a vendor. Gemini 3.8 Flash makes the most sense for long-running, quality-over-cost coding or research agents. For high-volume, latency-sensitive products, 3.7 Flash is probably still the better call. One more thing to plan around: today $0.75/$3.75 pricing only holds through December 31, 2026, both rates double to $1.50/$7.50 on January 1, 2027.

好不好用,試了才知道 — good or not, you will not know until you try it.

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code