跳到主要內容

GLM-5.3評測:中國最強開源編程模型,仍輸Claude | GLM-5.3 Review: China's Open Coding Model, Not Quite There

By Kit 小克 | AI Tool Observer | 2026-10-10

🇹🇼 GLM-5.3評測:中國最強開源編程模型,仍輸Claude

GLM-5.3是中國AI公司Z.ai(智譜AI)推出的開源權重編程模型,號稱是目前最強的開源程式碼模型;本月初(10月5日)正式在Amazon Bedrock上線,讓企業用戶能直接在雲端呼叫。但多家外媒實測後指出:GLM-5.3在開源陣營裡確實領先,碰上真正的封閉模型天花板——比如Claude——還是有落差,稱王的頭銜要打個折扣看。

GLM-5.3的規格:743B參數的重砲

GLM-5.3是混合專家(MoE)架構,總參數規模達743B,實際啟用約40B,站在GLM-5.2的基礎上做大量後訓練擴增。規格上主打:

  • 1M token超長上下文視窗,搭配工具呼叫與結構化輸出
  • 三種思考強度(effort levels)可調,依任務複雜度省算力
  • 主打agentic coding與長流程軟體工程任務,而不只是單次程式碼補全

實測數據:進步明顯,但天花板還在那

跟前代GLM-5.2比,GLM-5.3的進步確實有感:Terminal-Bench 3.0從4.6%跳到28.3%,DeepSWE v1.1從46.2%衝到66.9%,CyberGym拿下84.5%,略勝其他同噸位開源模型。問題在「跟誰比」——Decrypt等媒體的實測指出,GLM-5.3在token效率上贏過Claude Opus 4.8,但碰到Claude Fable 5(同一測試拿下39.5%)就明顯落後。換句話說,GLM-5.3打贏的是開源組的擂台賽,封閉模型的天花板它還沒摸到。

上Amazon Bedrock代表什麼?

比起跑分數字,GLM-5.3真正值得注意的是通路——10月5日起企業用戶可以直接在Amazon Bedrock呼叫GLM-5.3,支援prompt caching降低延遲與成本,不必自己搞GPU機房部署這顆743B的大模型(光跑權重就得多張GPU)。這代表中國開源模型第一次用「雲端託管」的姿態,正面進入美國企業採購名單,而不只是停留在開發者社群自架的階段。

該不該換?先別神化

如果你本來就在用開源模型、在意自架成本與資料主權,GLM-5.3值得排進評估清單,尤其現在Bedrock上線後部署門檻降低很多。但如果你是為了追求絕對程式碼品質才選模型,目前封閉模型(尤其Claude系列)在最難的任務上仍有優勢,GLM-5.3的「最強開源」頭銜,前面要加上「開源陣營裡」這個但書。好不好用,試了才知道。


🇺🇸 GLM-5.3 Review: China's Open Coding Model, Not Quite There

GLM-5.3 is the open-weight coding model that Chinese AI lab Z.ai (formerly Zhipu AI) released, pitched as the strongest open-weight coder on the market. It just landed on Amazon Bedrock on October 5, letting enterprise teams call it straight from the cloud. But independent reviews tell a more honest story: GLM-5.3 does lead the open-weight pack, but it still trails the closed frontier ceiling — meaning the "strongest open model" claim needs an asterisk.

The Specs: A 743B-Parameter Heavyweight

GLM-5.3 is a mixture-of-experts model with 743B total parameters and roughly 40B active per token, built on extensive post-training scaling from the GLM-5.2 base. Key specs:

  • A 1M-token context window with tool calling and structured outputs
  • Three configurable reasoning/effort levels to balance compute against task difficulty
  • Tuned for agentic coding and long-horizon software engineering, not just single-shot autocomplete

What the Benchmarks Actually Show

The jump over GLM-5.2 is real: Terminal-Bench 3.0 went from 4.6% to 28.3%, DeepSWE v1.1 climbed from 46.2% to 66.9%, and CyberGym hit 84.5%, edging out other open models its size. But the honest comparison is who it is measured against. Reviews from outlets like Decrypt note GLM-5.3 beats Claude Opus 4.8 on token economy, but still falls short of Claude Fable 5, which scores 39.5% on the same benchmark. GLM-5.3 wins the open-weight bracket; it has not touched the closed-model ceiling.

Why Amazon Bedrock Matters More Than the Score

The bigger story is not the benchmark number, it is distribution. Since October 5, enterprise customers can call GLM-5.3 directly through Amazon Bedrock, with prompt caching to cut latency and cost, instead of standing up the multi-GPU rig needed to self-host a 743B-parameter model. That is a Chinese open-weight model showing up on a major U.S. cloud procurement list for the first time, not just living in developer self-hosting forums.

Worth Evaluating, Not Worth Hyping

If you already run open-weight models and care about self-hosting cost or data sovereignty, GLM-5.3 deserves a spot on your shortlist, especially now that Bedrock access lowers the deployment bar. But if raw code quality on the hardest tasks is what drives your model choice, closed frontier models (Claude in particular) still hold the edge today. "Strongest open model" is true, with the "open" doing a lot of work in that sentence. You won't know until you try it.

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code