跳到主要內容

Beam 501B評測:Reflection AI開源模型,省算力不是最強 | Beam 501B Review: Efficient Open Model, Not the Strongest

By Kit 小克 | AI Tool Observer | 2026-10-09

🇹🇼 Beam 501B評測:Reflection AI開源模型,省算力不是最強

Reflection AI本週推出開源模型Beam,號稱用四分之一的運算資源打平中國開源王牌GLM 5.2,試圖在美中AI開源戰裡幫美國插旗。但重點來了:Beam現在還沒有權重可以下載,只能排waitlist。這篇評測整理效能實測數據與真實可用時間,讓你判斷值不值得等。

Beam是什麼:501B參數,推理只吃23B

Beam是一個Mixture-of-Experts(MoE)架構的開源模型,總參數501B,但每次推理只會啟用23B——這是它主打「省算力」的關鍵。訓練規模也很驚人:用6,144張NVIDIA GB300 GPU,不到4週跑完23.8兆token的預訓練,之後再用10,500張GB300做4週強化學習,產生超過一億次rollout。Context window原本256K,後來擴到1M token,長文件、長對話、Agent多輪任務都吃得下。

效能實測:追得上GLM,追不上Kimi K3

Reflection自己公布的數據顯示,Beam在推理類測試上接近GLM 5.2,但只用對方三到四分之一的推理算力——如果數據屬實,這對想自架模型又怕GPU帳單的團隊確實有吸引力。不過誠實講,在最硬的Coding Agent測試上,Beam並不是最強的:

  • SWE-Bench Pro v2-Hard:Beam 77.2%,輸給Kimi K3的88.2%
  • Terminal Bench v2.1:Beam 80.1分,輸給GLM 5.3(88.2)、Kimi K3(88.3)、DeepSeek V4.1 Flash(90.6)
  • DeepSWE v1.1:Beam只有44.4分,DeepSeek V4.1 Flash是74.2分,落差不小

簡單說,Beam的賣點是「用更少算力做到還不錯的效果」,不是「全面最強」。如果你的場景就是要榨出最後一點編碼能力,現階段Kimi K3、DeepSeek V4.1 Flash仍是更穩的選擇。

什麼時候真的能用:現在只有waitlist

最尷尬的地方是:以上所有數字都是Reflection AI自己公布的,因為權重還沒放出來。官方說會在10月內以Apache 2.0授權開放模型權重、技術報告和開發工具,但現在只能在platform.reflection.ai排waitlist搶先試用,外部還沒有人能跑出獨立測試結果。

該不該等

如果你在意的是「開源、自己部署、授權乾淨(Apache 2.0)」,Beam值得放進候選清單,等正式權重放出來再評估。如果你現在就要上線一個Coding Agent產品,Kimi K3或DeepSeek V4.1 Flash現在就能下載測試,不用等。

好不好用,試了才知道。


🇺🇸 Beam 501B Review: Efficient Open Model, Not the Strongest

Reflection AI this week unveiled Beam, a 501-billion-parameter open-weight model that claims to match China's GLM 5.2 using roughly a quarter of the inference compute — a direct shot in the US-China open model race. The catch: Beam's weights aren't downloadable yet, only a waitlist. Here's an honest look at the numbers and the actual timeline.

What Beam Actually Is: 501B Parameters, 23B Active

Beam is a sparse Mixture-of-Experts model with 501B total parameters but only 23B active per token — that is the whole efficiency pitch. Training was done on 6,144 NVIDIA GB300 GPUs in under four weeks across 23.8 trillion tokens, followed by four more weeks of reinforcement learning on 10,500 GB300s generating over 100 million rollouts. Context window starts at 256K and extends to 1M tokens, enough for long documents and multi-turn agent workflows.

Benchmarks: Close to GLM, Behind Kimi K3

Reflection own numbers put Beam reasoning performance near GLM 5.2 while using three to four times less inference compute — a real draw if your concern is GPU bills. But on the hardest coding-agent benchmarks, Beam is not the leader:

  • SWE-Bench Pro v2-Hard: Beam 77.2% vs. Kimi K3 88.2%
  • Terminal Bench v2.1: Beam 80.1 vs. GLM 5.3 (88.2), Kimi K3 (88.3), DeepSeek V4.1 Flash (90.6)
  • DeepSWE v1.1: Beam scores 44.4, well behind DeepSeek V4.1 Flash 74.2

The honest takeaway: Beam is an efficiency play, not a raw-capability winner. If you need top-tier coding agent performance today, Kimi K3 or DeepSeek V4.1 Flash remain the safer bet.

When Can You Actually Use It?

Every number above comes from Reflection AI itself — because the weights have not shipped. The company says full weights, a technical report, and developer tooling will land under an Apache 2.0 license later in October, but for now you can only join the waitlist at platform.reflection.ai. No independent benchmarks exist yet.

Should You Wait for Beam?

If a clean Apache 2.0 license and self-hosting matter more than peak coding performance, Beam is worth watching once weights actually drop. If you need a production coding agent today, Kimi K3 and DeepSeek V4.1 Flash are already downloadable.

好不好用,試了才知道。

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code