跳到主要內容

Qwen3.8-Max評測:2.4兆參數開源模型直逼Claude | Qwen3.8-Max Review: 2.4T-Param Open Model Rivals Claude

By Kit 小克 | AI Tool Observer | 2026-08-23

🇹🇼 Qwen3.8-Max評測:2.4兆參數開源模型直逼Claude

Qwen3.8-Max 是阿里巴巴 8 月 3 日正式推出的新一代開源模型,號稱是「史上最大開源模型」——2.4 兆參數的 MoE 架構,官方甚至聲稱效能只落後 Anthropic 的 Claude。這篇評測整理規格、開源程度和實際能不能用的落差,給想評估的開發者參考。

Qwen3.8-Max 規格拆解:2.4兆參數、100萬token上下文

Qwen3.8-Max 採用 Mixture-of-Experts(MoE)架構,總參數 2.4 兆,但每次推論只啟用約 950 億參數,這是大型 MoE 模型控制算力成本的標準做法。原生上下文窗口 26.2 萬 token,透過 YaRN 技術可擴展到 100 萬 token,並支援多模態輸入。

  • 參數規模:2.4T 總參數 / 95B 啟用參數(MoE)
  • 上下文長度:262K,YaRN 擴展至 1M token
  • 模態:多模態(文字+圖像)
  • API 定價:QwenCloud 上約 $2 / $6 每百萬 token(輸入/輸出)

開源到什麼程度?27B 全開放 vs 2.4T 客製授權

這裡是最容易誤會的地方。真正「開源」的是 Qwen3.8-27B,8 月 13-14 日以 Apache 2.0 授權上架 Hugging Face,可以自由商用、自由微調。但外界最關注的 2.4 兆參數旗艦版(Qwen3.8-2.4T-A95B)雖然也開放了權重,用的卻是客製授權條款,不是完全無限制的 Apache 2.0。換句話說,「史上最大開源模型」這個標題要打折扣看——權重公開了,但條款仍有限制。

真的能自己跑嗎?硬體門檻才是關鍵

2.4 兆參數就算是 MoE 架構,權重檔案動輒數百 GB 起跳,一般開發者的單機、甚至單張消費級 GPU 完全跑不動,需要多卡叢集才有辦法本地部署。所以對大多數團隊來說,Qwen3.8-Max 的實際使用方式還是透過 QwenCloud API,而不是自己架設。真正能「拿回家跑」的其實是 27B 版本,這也是它比 2.4T 版本更值得關注的原因——一般開發機或雲端單卡就能跑。

Qwen3.8-Max 值得換掉現有模型嗎?

官方基準測試聲稱效能僅次於 Claude,但這類自家公布的跑分向來要打折扣,實際寫程式、長文推理的表現還是得自己測過才準。可以確定的優勢是價格:$2/$6 每百萬 token 比多數閉源前沿模型便宜不少,加上 1M token 的長上下文,對需要處理大量文件、長對話記錄的應用是實用的選擇。但如果你的工作流已經穩定跑在 GPT-5.6 或 Claude 上,現在還沒有非換不可的理由。

好不好用,試了才知道。


🇺🇸 Qwen3.8-Max Review: 2.4T-Param Open Model Rivals Claude

Qwen3.8-Max is Alibaba's new open-weight model launched August 3, billed as the largest open-weight release ever, a 2.4-trillion-parameter MoE architecture that the company claims trails only Anthropic's Claude in benchmarks. Here's what the specs, the licensing fine print, and real-world deployability actually look like.

Qwen3.8-Max Specs: 2.4 Trillion Parameters, 1M Token Context

Qwen3.8-Max uses a Mixture-of-Experts (MoE) architecture with 2.4 trillion total parameters but only about 95 billion active per inference, the standard trick large MoE models use to keep compute costs manageable. Native context is 262K tokens, extended to 1 million tokens via YaRN, with multimodal input support.

  • Parameters: 2.4T total / 95B active (MoE)
  • Context length: 262K native, extended to 1M via YaRN
  • Modality: Multimodal (text + image)
  • API pricing: ~$2/$6 per million tokens (input/output) on QwenCloud

How Open Is Qwen3.8-Max, Really?

This is where the headline gets misleading. The genuinely open release is Qwen3.8-27B, which hit Hugging Face on August 13-14 under Apache 2.0, free for commercial use and fine-tuning, no strings attached. The flagship 2.4-trillion-parameter model (Qwen3.8-2.4T-A95B) also shipped its weights publicly, but under a custom license, not full Apache 2.0. So "largest open-weight model ever" needs an asterisk: the weights are public, but usage terms still have restrictions.

Can You Actually Self-Host It? Hardware Is the Real Gate

Even as MoE, 2.4 trillion parameters means weight files in the hundreds of gigabytes, no single machine, let alone a consumer GPU, runs this locally without a multi-GPU cluster. For most teams, the practical way to use Qwen3.8-Max is through the QwenCloud API, not self-hosting. The version you can actually "take home" is the 27B model, which arguably deserves more attention than the 2.4T flagship, since it runs on a single dev machine or cloud GPU.

Is Qwen3.8-Max Worth Switching To?

Alibaba's own benchmarks claim near-Claude performance, but self-reported scores always need independent verification, real coding and long-context reasoning tasks are the actual test. What's clearly true is the price: $2/$6 per million tokens undercuts most closed frontier models, and the 1M-token context makes it practical for document-heavy or long-conversation workloads. But if your pipeline already runs smoothly on GPT-5.6 or Claude, there's no urgent reason to switch yet.

好不好用,試了才知道。(Only real use tells you if it's worth it.)

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code