跳到主要內容

Mistral Large 4評測:法國重砲Le Chonk,智商墊底 | Mistral Large 4 Review: Le Chonk's 1T Params, Dead Last on IQ

By Kit 小克 | AI Tool Observer | 2026-10-10

🇹🇼 Mistral Large 4評測:法國重砲Le Chonk,智商墊底

法國AI實驗室Mistral在10月推出Mistral Large 4(暱稱Le Chonk),這是一款1.05兆參數的開源權重模型,也是歐洲目前訓練出的最大規模AI模型,瞬間成為本週AI圈討論度最高的話題。和多數開源模型直接拿中國底座微調不同,Mistral Large 4是從零開始在歐洲訓練,這一點讓它在「主權AI」議題上特別受關注。

萬億參數,只啟動520億:MoE架構省成本

Mistral Large 4採用混合專家(MoE)架構,雖然總參數達1.05兆,但每次推理只啟動約520億參數,讓實際運算成本接近中型密集模型,而不是真正跑滿1兆參數的天價帳單。這也是Mistral Large 4能把API定價壓到每百萬輸入token約1.36美元(目前官網顯示的折扣價更低到0.68美元,是不是長期優惠官方沒說清楚)的關鍵。

資安基準並列第一,但智商測試墊底

Mistral Large 4的強項很明確:在Artificial Analysis Cyber Index資安基準上與GLM 5.3 Flash並列第一,CyberGym全球排名前五。但在衡量綜合推理能力的Artificial Analysis Intelligence Index上,Mistral Large 4在追蹤的25個模型中排名墊底第25名,明顯落後中國開源模型(如Kimi K3)與各家旗艦閉源模型。這個落差才是這次發布背後真正值得留意的地方。

開源權重10月31日上線,但實測有雷

  • 開源權重預計10月31日公開,屆時可免費下載(需遵守授權條款)
  • 目前只有付費API的公開預覽版可用
  • 知名YouTuber Matthew Berman實測發現,用在agentic coding流程時,Mistral Large 4的推理輸出量過大,很快就把context window塞爆,實際可用性大打折扣

如果你只看基準測試數字,Mistral Large 4看起來是歐洲難得能打資安硬仗的開源選手;但真正要拿來寫程式或掛agent流程,目前的公開預覽版還有明顯的工程問題要解決。企業要不要導入,恐怕要看重的是「歐洲主權運算」這張牌,而不是純粹的智能分數。

好不好用,試了才知道。


🇺🇸 Mistral Large 4 Review: Le Chonk's 1T Params, Dead Last on IQ

Mistral Large 4 ("Le Chonk"), the French lab's new 1.05-trillion-parameter open-weight model, became this week's hottest AI release, and for good reason. Unlike most open-weight models that start from a Chinese base and get fine-tuned, Mistral Large 4 was trained from scratch in Europe — a detail that matters more for the "sovereign AI" conversation than for raw benchmark bragging rights.

1 Trillion Params, Only 52B Active: The MoE Trick

Mistral Large 4 uses a mixture-of-experts design: 1.05 trillion total parameters, but only about 52 billion active per token. That keeps inference costs closer to a mid-size dense model instead of a trillion-parameter price tag. It's also how Mistral can offer API pricing around $1.36 per million input tokens — currently discounted to $0.68 on their site, though Mistral hasn't said whether that's a temporary launch rate or the new normal.

Tied for #1 on Security, Dead Last on IQ

The strengths are specific: Mistral Large 4 ties for first place on the Artificial Analysis Cyber Index alongside GLM 5.3 Flash, and lands in the top five globally on CyberGym. But on the Artificial Analysis Intelligence Index — a broader reasoning benchmark — Mistral Large 4 ranks 25th out of 25 tracked models, trailing Chinese open-weight leaders like Kimi K3 by a wide margin. That gap is the real story behind the hype.

Open Weights Land October 31 — With a Catch

  • Open weights are scheduled for release on October 31, free to download under license.
  • Right now only the paid API preview is live.
  • YouTuber Matthew Berman's hands-on test found that in an agentic coding harness, Mistral Large 4's verbose reasoning output quickly exhausted the context window — a real usability problem the benchmark numbers don't show.

If you only read the leaderboard, Mistral Large 4 looks like Europe's first real contender on security benchmarks. But for actual coding or agent workloads, the current preview still has rough edges. Whether enterprises adopt Mistral Large 4 will likely hinge on sovereign-compute and regulatory fit, not raw intelligence scores alone.

好不好用,試了才知道。(You won't know until you try it.)

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code