跳到主要內容

OpenAI Jalapeño晶片評測:推理效能贏Blackwell達1.9倍 | OpenAI Jalapeño Chip Review: Up to 1.9x Nvidia Blackwell

By Kit 小克 | AI Tool Observer | 2026-09-07

🇹🇼 OpenAI Jalapeño晶片評測:推理效能贏Blackwell達1.9倍

OpenAI 與 Broadcom 合作打造的首款自研 AI 推理晶片 Jalapeño,在 8 月底公布的第三方實測數據中,每瓦吞吐量最高贏過 Nvidia Blackwell 1.9 倍,延遲也低了 3.6 倍。消息一出,晶片戰再度成為焦點,但這是不是代表 OpenAI 已經能甩開 Nvidia?答案沒那麼簡單。

Jalapeño 是什麼:專攻推理的客製 ASIC

Jalapeño 不是通用 GPU,而是針對大型語言模型推理(inference)從零設計的 ASIC,鎖定服務端最常見的運算模式、記憶體搬移與網路瓶頸。晶片標稱功耗 700W,但 OpenAI 表示在實測工作負載下,實際持續耗電多半落在 550W 以下。研發過程中,OpenAI 也讓自家模型參與晶片設計輔助,算是 AI 反哺硬體的一個例子。

實測數據:贏在效率,不是贏在蠻力

根據獨立測試平台 InferenceX 針對 GPT-OSS 120B、DeepSeek R1 670B、Kimi K2.5 1T 三款大模型的實測:

  • 每瓦吞吐量達 85,448 tokens/秒/kW,Nvidia Blackwell 同條件下為 44,960,差距約 1.5~1.9 倍
  • 相比 Nvidia 1,400W 的 GB300,延遲低達 3.6 倍
  • 128 顆晶片組成的機櫃可提供 1.7 exaflops 的 4-bit 算力,搭配 27.5TB HBM4 記憶體

換句話說,Jalapeño 的優勢集中在「每瓦效能」與「推理延遲」,而不是單顆晶片的絕對算力。

別高興太早:小克的實測心得

看數字很爽,但幾個現實值得注意。第一,Jalapeño 年底只會「小量部署」,真正上量要等 2027 年,短期內撼動不了 Nvidia 的市占。第二,同一週 Nvidia 才剛宣布提供高達 1,050 億美元融資,支持 OpenAI 在俄亥俄州蓋資料中心──而且明確指定全部使用 Nvidia 運算,代表 OpenAI 在訓練端仍高度依賴老對手。第三,這份測試數據終究是 OpenAI 主導公布,雖然委外給 InferenceX 測試,獨立性仍需觀察。這場晶片大戰更像是「自研推理晶片+繼續買 Nvidia 訓練晶片」的雙軌策略,而非誰取代誰。對開發者來說,短期內成本與效能的實質影響還看不到,值得持續追蹤。

好不好用,試了才知道。


🇺🇸 OpenAI Jalapeño Chip Review: Up to 1.9x Nvidia Blackwell

OpenAI and Broadcom's first custom AI inference chip, Jalapeño, posted third-party benchmark results in late August showing up to 1.9x the per-watt throughput of Nvidia Blackwell, with latency as much as 3.6x lower. Headlines called it a win over Nvidia -- but the real picture is more nuanced.

What Jalapeño Actually Is

Jalapeño isn't a general-purpose GPU. It's a purpose-built ASIC designed from scratch for LLM inference, optimized around the specific kernels, memory movement, and serving patterns that matter for frontier models. It's rated at 700W TDP, though OpenAI says sustained power on tested workloads stayed at or below 550W. OpenAI's own models reportedly assisted in the chip's design -- AI helping build the hardware that runs AI.

The Benchmarks: Efficiency, Not Raw Power

Independent testing platform InferenceX ran three large models -- GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T -- and found:

  • 85,448 tokens/sec per kilowatt versus 44,960 for Nvidia Blackwell under the same conditions -- a 1.5x to 1.9x gap
  • Up to 3.6x lower latency compared to Nvidia's 1,400W GB300
  • A 128-chip rack delivers 1.7 exaflops of 4-bit compute with 27.5TB of HBM4 memory

The takeaway: Jalapeño's edge is in performance-per-watt and inference latency, not raw compute per chip.

Don't Celebrate Too Early

A few things temper the excitement. First, Jalapeño ships in "very small volumes" by end of 2026 -- meaningful scale doesn't arrive until 2027, so it won't dent Nvidia's market share anytime soon. Second, the same week this launched, Nvidia committed up to 105 billion USD in financing for OpenAI's Ohio data center -- built exclusively on Nvidia compute, meaning OpenAI still leans heavily on Nvidia for training. Third, these benchmarks were released under OpenAI's own announcement, run via a third party (InferenceX) but not a fully independent audit. This looks less like OpenAI dethroning Nvidia and more like a dual-track strategy: custom silicon for inference, Nvidia for training. For developers, real-world cost and performance impact is still unproven -- worth watching, not betting on yet.

You won't know until you try it.

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code