Qwen3.8 Flash Next評測:RTX 4090就能跑125B大模型 | Qwen3.8 Flash Next Review: Run 125B on One RTX 4090
By Kit 小克 | AI Tool Observer | 2026-10-06
🇹🇼 Qwen3.8 Flash Next評測:RTX 4090就能跑125B大模型
Qwen3.8 Flash Next 這幾天在 Hacker News 上衝到接近八百分熱度,原因很簡單:開源推理引擎 Strata 讓一張消費級的 RTX 4090 顯卡,就能跑動阿里巴巴這顆 125B 參數的 MoE 大模型,而且官方宣稱有機會衝到每秒 100 token 的生成速度。對長期只能望著 API 帳單嘆氣的開發者來說,這聽起來像是本地跑大模型的里程碑。
Qwen3.8 Flash Next 與 Strata 是什麼?
Qwen3.8 Flash Next 是阿里巴巴 Qwen 團隊於今年 8 月發布的混合專家(MoE)模型,總參數 125B,但每個 token 實際只會啟動約 6B 參數,這也是它能被「瘦身」塞進消費顯卡的關鍵。官方跑分顯示 MMLU 達 90.36、GPQA Diamond 91.7 分,幾乎追平 Claude Opus 4.6。而 Strata 是開發者 Niko1221 做的開源推理引擎(MIT 授權),核心賣點是用 ISTA-DASLab 的 GSQ-RCO 量化法,取代傳統「把模型切一半塞進顯卡、一半丟給 CPU」的 layer offloading 做法,讓顯存使用更有效率。
實測速度:理想與現實有落差
HN 上標題寫的是「RTX 4090 跑到 100T/s」,但實際測試結果分散得多:
- 有使用者在 24GB 顯存的 4090 上,250K 超長上下文下只拿到約 21 tokens/秒的生成速度,但 prefill(處理輸入)可以到 364 t/s。
- 另一位搭配 128GB 系統記憶體與 Ryzen 7950X3D 的使用者,回報拿到 124 tokens/秒,比官方數字還高。
- 多篇技術拆解都提到,真正跑得順需要額外搭配高達 192GB 的系統記憶體,因為 MoE 的「專家」權重得靠系統 RAM 分攤,不是單靠 24GB VRAM 就能解決。
換句話說,「單張 4090 跑 125B 模型」是真的,但那個數字背後藏著一台配置不便宜的主機,換到 3090 或 4080 速度也會明顯下降,量化等級拉低後長文本推理的準確度也會打折。
值得自己架設嗎?
如果你本來就有高階遊戲主機、想省下每月的 API 費用,或是在意資料不上雲端,Strata + Qwen3.8 Flash Next 是目前本地跑大模型最值得玩的組合之一,GitHub 上文件也寫得算完整。但別被標題的「100T/s」唬住,先確認自己有沒有那 192GB 記憶體跟對的量化版本,不然體感速度可能跟雲端 API 差一大截。
好不好用,試了才知道。
🇺🇸 Qwen3.8 Flash Next Review: Run 125B on One RTX 4090
Qwen3.8 Flash Next is the topic everyone is testing this week, after an open-source inference engine called Strata made headlines for running Alibaba's 125B-parameter MoE model on a single consumer-grade RTX 4090, with claims of up to 100 tokens/second. For anyone tired of watching their API bill climb, that sounds like a breakthrough for local LLM inference.
What Are Qwen3.8 Flash Next and Strata?
Qwen3.8 Flash Next is Alibaba Qwen team's mixture-of-experts (MoE) model released in August 2026. It has 125B total parameters but only activates around 6B per token, which is exactly why it can be squeezed onto consumer hardware. Official benchmarks show 90.36 on MMLU and 91.7 on GPQA Diamond, roughly on par with Claude Opus 4.6. Strata, built by developer Niko1221 and released under MIT license, is the engine that makes it run: instead of the old trick of offloading half the model to CPU, it leans on ISTA-DASLab GSQ-RCO quantization scheme for more efficient VRAM usage.
Real-World Speed: The Gap Between Hype and Reality
The HN headline says 100T/s on an RTX 4090, but real-world numbers are far more scattered:
- One user with a 24GB RTX 4090 and a 250K-token context window reported only ~21 tokens/sec decode speed, though prefill hit 364 t/s.
- Another user pairing the 4090 with 128GB of system RAM and a Ryzen 7950X3D reported 124 tokens/sec, beating the official number.
- Multiple technical breakdowns note that smooth performance actually requires up to 192GB of system RAM, since the MoE experts still need to be distributed across system memory, not just 24GB of VRAM.
In other words: running a 125B model on one RTX 4090 is real, but behind that headline number sits a fairly expensive full rig. Swap in a 3090 or 4080 and speed drops noticeably; push quantization lower and long-context accuracy takes a hit too.
Is It Worth Setting Up?
If you already own a high-end gaming rig, want to cut your monthly API costs, or care about keeping data off the cloud, Strata plus Qwen3.8 Flash Next is currently one of the most interesting local-LLM setups to try, and the GitHub docs are genuinely thorough. Just do not get hypnotized by 100T/s in the headline: check whether you actually have the RAM and the right quantization build first, or your real-world speed could fall well short of any cloud API.
好不好用,試了才知道。
Sources / 資料來源
- Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s - Hacker News
- Niko1221/Strata - GitHub
- Qwen3.8-Flash-Next: A New Architecture - Qwen Official Blog
延伸閱讀 / Related Articles
- GPT-6.1 Sol評測:Astra五分之一價格,夠用嗎 | GPT-6.1 Sol Review: 1/5 Astra's Price, Still Good Enough?
- OpenAI失控AI代理評測:駭爆Hugging Face遭加州傳喚 | OpenAI Rogue Agents Review: Hacked Hugging Face, CA Probe
- GPT-6 Astra作弊事件評測:StarCraft輸了就偷跑對手程式 | GPT-6 Astra Review: It Cheated at StarCraft When Losing
AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends
留言
張貼留言