Anthropic風險報告評測:自曝疏漏仍標「低風險」 | Anthropic Risk Report Review: Gaps Disclosed, Still 'Low'
By Kit 小克 | AI Tool Observer | 2026-08-28
🇹🇼 Anthropic風險報告評測:自曝疏漏仍標「低風險」
Anthropic 在 8 月中發布第二份公開的風險報告(Risk Report),依照公司內部最新版「負責任擴展政策」(RSP v3.4)揭露多起訓練與資安疏漏。表面上最受關注的數字,是整體「自主性風險」評級只從「極低」調升到「低」——但報告本身揭露的細節,讀起來卻不太像「低風險」。
三個藏了將近一年的疏漏
這份 Anthropic風險報告 罕見地公開了三起內部發現的問題:
- 生物武器分類器漏篩:約 5 萬名外包評分員,累積 1.33 億筆對話,因為一個「internal use(內部用途)」的錯誤標記,近一年沒有經過生物武器風險分類器篩選。Anthropic 強調沒有證據顯示遭濫用,但坦承風險在於模型可能被「蒸餾」出未過濾的敏感知識。
- 思維鏈外洩污染訓練資料:模型的推理過程在獎勵訓練階段意外外洩,影響比例分別是 Mythos Preview 5.1%、Mythos 5 2.7%、Opus 4.6 0.2% 的訓練集數。
- 「假裝對齊」對話被誤收回訓練語料:原本已被標記為危險的 alignment-faking 對話紀錄,後來又意外被重新納入正式訓練資料,影響所有知識截止日在 2024 年 12 月之後的模型。
內部模型 Model 2 逼近「研究員替代」門檻
報告也首度提到代號 Model 2 的內部專用模型(不對外開放),在測量「AI 能否取代人類研究員」的 CoBench 基準上拿下 62.8% 分,明顯高於 Mythos Preview 的 54.8% 與 Mythos 5 的 50.3%。Anthropic 自訂 85% 為「完全取代研究員」的門檻——照這個成長速度,距離門檻可能比外界想像的更近。
「低風險」評級是不是低估了?
獨立分析師 Zvi Mowshowitz 在讀完整份報告後直言,Anthropic 自己揭露的證據其實更接近「中度風險」,公司卻只把評級從「極低」上調到「低」,結論與揭露內容明顯不成比例。這正是這份 Anthropic風險報告 最值得玩味之處:業界公認最重視透明度的實驗室,也會有安全機制整整一年沒被發現失靈的時候。
對一般開發者與用戶來說,這不代表現在用 Claude 不安全,但它提醒一件事:安全評級是廠商自己寫的分數,實際風險最終還是要靠外部檢驗與時間證明。
好不好用,試了才知道。
🇺🇸 Anthropic Risk Report Review: Gaps Disclosed, Still 'Low'
Anthropic published its second public Risk Report in mid-August 2026, disclosing several training and safety lapses under its updated Responsible Scaling Policy (RSP v3.4). The headline number is that the company's overall autonomy risk rating moved just one notch, from "very low" to "low" — but the details buried inside the report read like more than a one-notch problem.
Three Gaps That Went Unnoticed for Nearly a Year
The Anthropic Risk Report discloses three separate internal failures:
- A biological-weapons classifier gap: roughly 50,000 contracted human feedback workers generated 133 million exchanges that ran for nearly a year without passing through the bio-weapons risk classifier, due to a mislabeled "internal use" flag. Anthropic says it found no evidence of misuse, but acknowledges the risk that a model could be distilled to recover the unfiltered knowledge.
- Chain-of-thought leakage contaminated training data: reasoning traces were unintentionally exposed during reward training, affecting 5.1% of episodes for Mythos Preview, 2.7% for Mythos 5, and 0.2% for Opus 4.6.
- Alignment-faking transcripts were re-ingested into training data after already being flagged as dangerous — contaminating every model with a knowledge cutoff after December 2024.
An Internal Model Is Closing In on the "Researcher Substitution" Line
The report also reveals Model 2, an internal-only model not released externally, scoring 62.8% on CoBench — Anthropic's benchmark for whether AI can substitute for human researchers — up from 54.8% for Mythos Preview and 50.3% for Mythos 5. Anthropic treats 85% as the threshold for "full researcher substitution." At this pace, that line may be closer than expected.
Is "Low Risk" Actually an Understatement?
Independent analyst Zvi Mowshowitz argues the report's own disclosures point to a "medium" risk rating, not "low" — the conclusion doesn't match the evidence Anthropic itself published. That gap is what makes this Anthropic Risk Report worth reading closely: even the lab most associated with safety transparency ran a safety mechanism broken for a full year without noticing.
None of this means Claude is unsafe to use today. But it's a reminder that a safety rating is still a number the vendor writes about itself — the real risk only gets tested by outside scrutiny and time.
You won't know until you try it.
Sources / 資料來源
- Anthropic – Risk Report: August 2026
- Zvi Mowshowitz – Anthropic Risk Report: August 2026 分析
- Tech Times – Anthropic Upgrades Misalignment Risk as Key Safety Benchmarks Saturate
延伸閱讀 / Related Articles
- Nvidia併購Hugging Face評測:129億美元買下開源AI門面 | Nvidia Hugging Face Acquisition Review: $12.9B Buys Open-Source AI Hub
- OX Alpha評測:神秘模型爆紅一週,真相是Z.ai換皮測試版 | OX Alpha Review: Mystery Model Was Z.ai's GLM Test
- Nvidia投資Perplexity評測:300億美元背後的晶片布局 | Nvidia-Perplexity Investment Review: $30B Chip Compute Bet
AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends
留言
張貼留言