跳到主要內容

openTPU評測:AI設計晶片,FPGA實測80 tok/s | openTPU Review: AI Designs Its Own Chip, Hits 80 Tok/s

By Kit 小克 | AI Tool Observer | 2026-10-08

🇹🇼 openTPU評測:AI設計晶片,FPGA實測80 tok/s

最近 Hacker News 上討論最熱的話題不是又一個新模型,而是 openTPU——一個完全由 AI 自己設計出來的開源 AI 加速晶片。這個專案把電路設計、指令集、模擬器、編譯器和效能分析工具全部塞進同一個 repo,目標只有一個:看 AI agent 能不能自己造出一顆跑得動自己推論的晶片。

openTPU 是什麼:從 RTL 到編譯器都是 AI 寫的

openTPU 的賣點不是效能,而是「誰寫的」。整個專案包含 SystemVerilog 電路設計(RTL)、自訂指令集(ISA)、逐位元精確的 Python 模擬器、核心編譯器,以及一個叫 Lens 的效能分析工具,外加命令列聊天工具 otpu-chat。這些東西過去通常要一整個硬體團隊分工數月才能生出一套,openTPU 用 AI agent 一條龍打完。硬體跑在一張 Kintex-7 FPGA PCIe 卡上,目前支援 Qwen3、LFM2.5、Qwen3.5、Gemma 4 這類中小型模型。

效能數字:從個位數 tok/s 到 80+ tok/s

最早版本的 openTPU 每秒只能擠出個位數的 token,靠著一個遞迴自我改進迴圈(讓 AI 自己分析瓶頸、重寫電路、再測試),效能一路衝到小模型上 80+ tok/s。官方公布的具體數字是 LFM2.5-230M 在 int8 精度下,decode 59.0 tok/s、prefill 295.6 tok/s。這個進步曲線本身比最終數字更值得關注——代表「AI 設計晶片、AI 優化晶片」這個迴圈真的能跑起來,不是紙上談兵。

誠實說:這跟 Nvidia、Google TPU 比不了

先把期待值壓低:openTPU 跑的是 FPGA,模型也只到 230M 到數十億參數等級,跟資料中心裡的 H100 或 Google 自家 TPU v7 完全不是同一個量級的比較。它真正的意義不是「效能突破」,而是把整條晶片設計鏈開源出來給所有人檢查、復現、修改——這在過去幾乎是大廠的黑盒子。

  • 想摸硬體設計或測試 AI agent 能力邊界的人:這個 repo 值得花一個下午翻一翻
  • 想找便宜量產推論晶片方案的人:現在還太早,先別急著下單 FPGA 板

好不好用,試了才知道。


🇺🇸 openTPU Review: AI Designs Its Own Chip, Hits 80 Tok/s

openTPU is the Hacker News story everyone is actually talking about this week — not another model drop, but an open-source AI accelerator chip that was designed by AI itself. The project bundles circuit design, instruction set, simulator, compiler, and profiling tools into a single repo, built around one question: can an AI agent design the chip that runs its own inference?

What openTPU Actually Ships: RTL to Compiler, All AI-Written

The headline isn't raw speed — it's authorship. The repo includes SystemVerilog RTL, a custom instruction set (ISA), a bit-exact Python simulator, a kernel compiler, a profiler called Lens, and a CLI chat tool (otpu-chat). Work that normally takes a hardware team months to produce was generated end-to-end by AI agents. The hardware target is a Kintex-7 FPGA PCIe card, currently running small-to-mid models like Qwen3, LFM2.5, Qwen3.5, and Gemma 4.

The Numbers: From Single-Digit tok/s to 80+

Early builds of openTPU struggled to hit single-digit tokens per second. Through a recursive self-improvement loop — where the AI analyzes its own bottlenecks, rewrites the circuit, and retests — throughput climbed to 80+ tok/s on smaller models. The published benchmark: LFM2.5-230M hits 59.0 tok/s decode and 295.6 tok/s prefill at int8. The improvement curve matters more than the final number — it is evidence the "AI designs chip, AI optimizes chip" loop actually closes in practice, not just on paper.

Honest Take: Don't Compare This to Nvidia or Google's TPU

Set expectations correctly: openTPU runs on an FPGA, and the models it handles top out in the hundreds of millions of parameters — nowhere near datacenter silicon like H100s or Google's own TPU v7. The real significance isn't a performance breakthrough; it's that the entire chip-design pipeline is now open for anyone to inspect, reproduce, and modify — something that used to live entirely inside big labs' black boxes.

  • If you are curious about hardware design or want to see where AI agent capabilities actually stand: the repo is worth an afternoon
  • If you are hunting for a cheap production inference chip: it is too early, don't buy an FPGA board just yet

好不好用,試了才知道 / You won't know until you try it.

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code