跳到主要內容

Beam評測:Reflection AI 501B開源模型,權重還沒放 | Beam Review: Reflection AI's 501B Model, No Weights Yet

By Kit 小克 | AI Tool Observer | 2026-10-07

🇹🇼 Beam評測:Reflection AI 501B開源模型,權重還沒放

Beam評測:Reflection AI丟出501B開源模型,但權重還沒給你

Reflection AI 在 2026 年 10 月 5 日發布了Beam,一個 501B 參數的開源權重 MoE(混合專家)模型,實際運算時只啟用 23B 參數,主打程式撰寫、推理與 AI 代理工作。這家公司背後站著 Nvidia 投資,喊出要扛起「西方開源模型陣營」大旗,對標中國陣營的 GLM、Qwen、Kimi 系列。聽起來很猛,但現在想用 Beam,你只能排候補名單,連權重都還沒拿到。

跑分數字:打平 GLM 5.2,算力省3到4倍

Beam 預訓練吃了 23.8 兆 token,壓縮在 Nvidia GB300 NVL72 系統上不到四週跑完,之後又花四週用約 1 萬顆 GPU 做強化學習微調。成果是在 SWE-bench Verified 拿下 80.9 分,打贏 Inkling 的 77.6 和 Nemotron 3 Ultra 的 70.7。在 Terminal-Bench 2.1 上 Beam 拿 80.1,些微落後 Z.ai 的 GLM 5.2(81.0 分);但在 SWE-bench Pro v1 又反超,65.5 比 62.1。Reflection 自己的賣點不是跑分最高,而是用 GLM 5.2 的 3 到 4 倍效率做出接近的成績——這對自建機房跑推理的團隊確實有感。

殘酷現實:權重沒放、沒定價、context 被砍

Beam 官方喊出 1M token 的 context window,但現在能申請到的 beta API,輸入加輸出總共只給 262,144 token,輸出上限更砍到 131,072。想用 API 也只能排隊等 early access,拿到資格後透過 OpenAI 格式的 client 呼叫,官方完全沒公布定價,等於你沒辦法算這個模型跑一個任務要花多少錢。真正要開源的 Apache 2.0 權重,官方說「10 月稍後」才會放出來——現在看到的所有跑分,都只是 Reflection 自己公布的數字,外部還沒人能真的下載模型來驗證。

跟最強的代理模型比,還是差一截

把 Beam 放到最硬的 agentic coding 測試裡,成績大概跟 GLM 5.2 同一個檔次,但明顯輸給 Kimi K3、GLM 5.3 跟 DeepSeek V4.1 Flash。換句話說,Beam 的故事是「效率」不是「最強」,如果你要的是當前頂尖的代理模型效能,它還不是答案。

值不值得關注

Beam 的架構和訓練效率數字值得工程團隊關注,尤其是在意自建推理成本的人。但在權重真正落地、定價公布、context 限制解除之前,現在就是純粹的「紙上測試」階段。建議先把它放進觀察清單,等 10 月底真正拿到可下載的權重,再決定要不要換掉手上的 GLM 或 Qwen。

好不好用,試了才知道。


🇺🇸 Beam Review: Reflection AI's 501B Model, No Weights Yet

Beam Review: Reflection AI Drops a 501B Open Model — Weights Not Included

On October 5, 2026, Reflection AI announced Beam, a 501-billion-parameter open-weight MoE (mixture-of-experts) model that only activates 23B parameters per token, aimed at coding, reasoning, and AI agent workloads. Backed by Nvidia as an investor, the company is pitching itself as the standard-bearer for the "Western open-weight frontier," going head-to-head with Chinese labs like GLM, Qwen, and Kimi. Sounds bold — except right now you can only join a waitlist, and the weights themselves aren't out yet.

Benchmarks: Ties GLM 5.2 at 3-4x Less Compute

Beam was pretrained on 23.8 trillion tokens in under four weeks on Nvidia GB300 NVL72 systems, followed by a four-week reinforcement learning run on roughly 10,500 GPUs. The result: 80.9 on SWE-bench Verified, beating Inkling's 77.6 and Nemotron 3 Ultra's 70.7. On Terminal-Bench 2.1, Beam scores 80.1, just shy of Z.ai's GLM 5.2 at 81.0 — but it edges ahead on SWE-bench Pro v1, 65.5 versus 62.1. Reflection's actual pitch isn't topping the leaderboard; it's matching GLM 5.2's results using 3 to 4 times less inference compute, which matters if you're running your own GPUs.

The Catch: No Weights, No Pricing, Capped Context

Reflection advertises a 1M-token context window, but the current beta API caps combined input and output at just 262,144 tokens, with output alone limited to 131,072. Access requires joining a waitlist, and once approved, you call Beam through an OpenAI-shaped client — with zero published pricing, so there's no way to estimate what running a task would actually cost. The promised Apache 2.0 weights aren't out either; Reflection says they're coming "later in October." Every benchmark number circulating right now comes straight from Reflection's own scorecard — no one outside the company has downloaded the model to verify it.

Still Behind the Real Agentic Leaders

On the toughest agentic coding evaluations, Beam lands roughly in GLM 5.2's tier, but it clearly trails Kimi K3, GLM 5.3, and DeepSeek V4.1 Flash. The story here is efficiency, not supremacy — if you need the strongest agent model available today, Beam isn't it.

Worth Watching?

Beam's architecture and training efficiency numbers are worth a look if you care about self-hosted inference costs. But until the weights actually ship, pricing gets published, and the context cap lifts, this is still a paper benchmark, not a usable product. Keep it on your watchlist and revisit once the downloadable weights land later this month before swapping out your GLM or Qwen setup.

You won't know until you try it.

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code