跳到主要內容

Clef評測:Cloudflare開源決策模型,分類比LLM快10倍 | Clef Review: Cloudflare's Open-Source Decision Model

By Kit 小克 | AI Tool Observer | 2026-10-05

🇹🇼 Clef評測:Cloudflare開源決策模型,分類比LLM快10倍

Clef 是 Cloudflare 在 2026 年 10 月初發布的開源決策模型(decision model),搭配全新的強化學習(RL)微調平台,主打一件事:讓 AI agent 在「分類、路由、選工具」這類瑣碎但高頻的判斷上,不必動用昂貴又慢的大型語言模型。這波發布一出就衝上 Hacker News 熱門榜,對正在堆 AI agent 架構的工程師來說,是本週最值得關注的開發者工具更新。

Clef 到底是什麼?不是另一個聊天機器人

多數人以為 AI 新品就是聊天機器人,但 Clef 完全反其道而行。它不生成一句句文字,而是直接吐出帶信心分數的結構化答案(typed output with probability scores)——例如「這張工單該分到哪個部門」「這段流量是爬蟲還是真人」。Cloudflare 用的是非自回歸(non-autoregressive)的兩階段注意力路由,平行評分每個候選選項,而不是像 LLM 一樣一個字一個字生成,這就是它速度能甩開傳統模型的關鍵。

速度數字:38.8 毫秒 vs 動輒數百毫秒

  • Clef-flash:中位延遲 38.8 毫秒,p95 為 122.4 毫秒
  • Clef(完整版):中位延遲 209.3 毫秒
  • 支援 64k context window,並內建視覺編碼器,可直接做圖像分類
  • 底層基於 Qwen 架構(Clef 用 Qwen3.8-27B,Clef-flash 用 Qwen3.5-9B)

Cloudflare 官方的基準測試顯示,Clef 在 BFCL、CLINC150+OOS 等分類任務上準確率超過 98%,優於目前市場上的競品 Jev。但要誠實講:這些數字是 Cloudflare 自己測的,獨立第三方評測指出 Jev 在 agent 執行軌跡(trace)的可觀測性上其實更強,沒有一款模型是全面輾壓的。

開源免費,但 RL 微調還沒完全放開

Clef 以 Apache 2.0 授權釋出,權重公開在 Hugging Face,也能直接透過 Workers AI 呼叫雲端 API。想用自己的資料微調?Cloudflare 同步推出 RL 微調平台,整合 AI Gateway、Workers AI 與 Containers,但目前只透過「與 Cloudflare 工程師一對一合作」的方式提供,自助式平台還沒上線,這對想快速自己跑通的團隊來說是個現實的落差。

適合誰用:客服分流、威脅情資、機器人偵測

Clef 的目標場景很明確:客服工單分類與急迫度判斷、網路威脅情資分類、Trust & Safety 內容審核、辨別流量是正常爬蟲還是惡意攻擊。這些都是「選項有限、但量大到不能靠人工」的任務,用 LLM 處理既貴又慢,這正是決策模型這個產品類別存在的理由。

結論:如果你的 agent pipeline 裡有大量「分類/路由」這種瑣事在吃 LLM token,Clef 值得拿來實測看看延遲與準確率是否真如官方宣稱;但 RL 微調平台還在半開放階段,大規模導入前建議先用公開權重跑小規模驗證。好不好用,試了才知道。


🇺🇸 Clef Review: Cloudflare's Open-Source Decision Model

Clef is Cloudflare’s new open-source decision model, released in early October 2026 alongside a reinforcement learning fine-tuning platform. The pitch: stop burning expensive LLM calls on routine classification and routing decisions inside AI agent pipelines. The launch shot to the top of Hacker News within days, making it one of the most relevant developer-tool stories this week for anyone building agent infrastructure.

What Clef Actually Is: Not Another Chatbot

Unlike LLMs that generate open-ended text token by token, Clef returns typed, structured outputs with confidence scores — think "which department should this ticket go to" or "is this traffic a bot or a human." It uses a non-autoregressive, two-stage attention routing process that scores candidate answers in parallel instead of generating sequentially, which is the core reason it’s dramatically faster than general-purpose LLMs for bounded-choice tasks.

The Numbers: 38.8ms vs Hundreds of Milliseconds

  • Clef-flash: 38.8ms median latency, 122.4ms at p95
  • Clef (full model): 209.3ms median latency
  • 64k context window plus a built-in vision encoder for image classification
  • Built on Qwen (Clef uses Qwen3.8-27B, Clef-flash uses Qwen3.5-9B)

Cloudflare’s own benchmarks claim over 98% accuracy on BFCL and CLINC150+OOS classification tasks, beating rival decision model Jev. To be fair, these are Cloudflare-reported numbers — independent reviewers note Jev still leads on agent-trace observability, and no single model wins across every workload.

Open and Free — But Self-Serve Fine-Tuning Isn’t Ready Yet

Clef ships under Apache 2.0, with weights on Hugging Face and hosted inference through Workers AI. The companion RL fine-tuning platform combines AI Gateway, Workers AI, and Containers — but it currently only runs as a hands-on engagement with Cloudflare’s forward-deployed engineers. A self-serve version is "coming later," which is a real gap if you were hoping to fine-tune on your own traffic this week.

Who Should Actually Use Clef

The target use cases are concrete: support ticket routing and urgency triage, threat intelligence classification, trust-and-safety content review, and telling legitimate crawlers from malicious bots. These are high-volume, bounded-choice tasks where running a full LLM is both slow and costly — exactly the niche this new decision model category is carving out.

Bottom line: if your agent pipeline burns LLM tokens on routine classification and routing, Clef is worth testing against your own latency and accuracy numbers. Just treat the benchmark claims with a grain of salt, and know the RL fine-tuning platform isn’t self-serve yet. 好不好用,試了才知道 — you won’t know until you try it yourself.

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code