Fugu Ultra v2評測:不練模型改練調度的AI | Fugu Ultra v2 Review: The Model That's Actually a Team
By Kit 小克 | AI Tool Observer | 2026-09-14
🇹🇼 Fugu Ultra v2評測:不練模型改練調度的AI
Fugu Ultra v2 是 Sakana AI 在 9 月 11 日發布的新模型,同時公布較便宜的手足 Fugu Max。這兩款最特別的地方不是跑分數字,而是架構本身——它不是單一訓練出來的大模型,而是一個「懂得調度其他模型」的指揮模型,對外仍包成一個 OpenAI 相容的 API,讓 Fugu Ultra v2 成為這週開發者社群討論度最高的話題之一。
什麼是「調度式」架構?
Sakana 官方把 Fugu 形容成「以一個模型的樣子交付的多代理系統」。實際運作上,Fugu 內部有一個負責路由、能遞迴呼叫自己的指揮模型,底層串接一整個開放權重與專用模型的池子,其中包含 Nvidia Nemotron 系列。開發者呼叫 API 時感覺不到背後在切換模型,但每次生成其實是多個子模型協作的結果。
定價與跑分:便宜四到六成是真的
- Fugu Max:每百萬輸入 token 2 美元、輸出 6 美元,官方稱比 Sonnet 5、GPT-5.6 Terra、Kimi K3 便宜四到六成,在十項同價位跑分中拿下六項第一,包含 Terminal-Bench 2.1 與 GPQA Diamond。
- Fugu Ultra v2:每百萬輸入 5 美元、輸出 30 美元,主打複雜推理與軟體工程。視覺推理測試 Chartography 拿下 48.3 分,超過 Opus 5 的 27.3 與 Fable 5 的 29.5;真實世界軟體工程測試 DeepSWE 拿下 74.3 分,贏過定價貴三到五倍的對手。
- 兩款都支援 100 萬 token 的上下文長度,最多可輸出 12.8 萬 token。
用之前該知道的取捨
調度式架構的代價也很直接:
- 輸出比輸入貴很多:Fugu Ultra v2 輸出是輸入的 6 倍價格,長報告、完整 code diff 這類會吐大量文字的任務,帳單可能比想像中高。
- 延遲沒公布:每次生成背後牽動多個子模型協作,官方沒有公布延遲數字,實務表現要自己拿真實任務測過才知道。
- 模型池會變動:調度品質高度依賴 Sakana 挑選、更新的底層模型池,池子換了,同一組 prompt 的表現可能跟著變,穩定性不如單一固定模型。
如果你的工作負載是程式代理、多步驟研究這類會受益於「換模型做對的事」的任務,Fugu Ultra v2 值得排進評測清單;但別只看跑分數字,先拿自己的真實任務跑一輪。
好不好用,試了才知道。
🇺🇸 Fugu Ultra v2 Review: The Model That's Actually a Team
Fugu Ultra v2 is the model Sakana AI released on September 11, alongside a cheaper sibling called Fugu Max. What makes it interesting isn't the benchmark scores — it's the architecture. Instead of training one bigger monolithic model, Sakana built a "conductor" model that routes tasks across a pool of other models, wrapped behind a single OpenAI-compatible endpoint. That's made Fugu Ultra v2 one of the most-discussed releases in developer circles this week.
What "Orchestration" Architecture Actually Means
Sakana describes Fugu as "a Multi-Agent System, Delivered as One Model." Under the hood, a routing model — one that can recursively call itself — decides which model in a large pool of open-weight and specialized models (including Nvidia's Nemotron family) handles each part of a request. From the outside, calling the API looks like talking to a single model. In reality, each response is the output of several models cooperating behind the scenes.
Pricing and Benchmarks: The 40-60% Discount Checks Out
- Fugu Max costs $2 per million input tokens and $6 per million output tokens — Sakana claims 40-60% cheaper than Sonnet 5, GPT-5.6 Terra, and Kimi K3. It topped 6 of 10 benchmarks in its price tier, including Terminal-Bench 2.1 and GPQA Diamond.
- Fugu Ultra v2 costs $5 input / $30 output per million tokens, aimed at complex reasoning and software engineering. It scored 48.3 on Chartography (visual reasoning), beating Opus 5's 27.3 and Fable 5's 29.5, and 74.3 on DeepSWE, outscoring models priced 3-5x higher.
- Both support a 1M-token context window and up to 128K output tokens.
What to Know Before You Switch
The orchestration approach comes with real trade-offs:
- Output is far pricier than input — Fugu Ultra v2's output tokens cost 6x its input tokens. Tasks that generate a lot of text (long reports, full code diffs) can get expensive fast.
- No published latency numbers — every response involves multiple models cooperating behind the scenes, and Sakana hasn't shared how that affects response time. Test it on your own workload before assuming it's fast.
- The model pool can change — output quality depends on whichever models Sakana currently routes to. If the pool changes later, the same prompt might behave differently, a consistency risk a single fixed model doesn't have.
If your workload is coding agents or multi-step research — tasks that genuinely benefit from "picking the right model for the job" — Fugu Ultra v2 is worth putting on your evaluation list. Just don't trust the benchmark chart alone; run it against your real tasks first.
You won't know until you try it.
Sources / 資料來源
- Sakana AI: Introducing Fugu Max and Fugu Ultra v2
- OpenRouter: Fugu Ultra v2 Pricing and Specs
- MarkTechPost: Sakana AI Launches Fugu Max and Fugu Ultra v2
延伸閱讀 / Related Articles
- AI Agent自建軟體評測:McKinsey曝32%企業不買SaaS | Agentic Coding Review: McKinsey Finds 32% Skip Buying SaaS
- Agentforce評測:Salesforce七個AI代理各有名字與職稱 | Agentforce Review: Salesforce Gives AI Agents Names, Jobs
- Real-SWE評測:企業真實程式碼,AI最高只答對38.8% | Real-SWE Review: AI Tops Out at 38.8% on Real Code
AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends
留言
張貼留言