AI末日警告評測:Anthropic研究員辭職曝10%滅絕機率 | AI Extinction Warning Review: Anthropic Quits Over 10% Odds
By Kit 小克 | AI Tool Observer | 2026-09-15
🇹🇼 AI末日警告評測:Anthropic研究員辭職曝10%滅絕機率
AI滅絕風險這幾天成了全球AI圈最熱的話題。9月8日,曾在OpenAI與Anthropic做了三年預訓練研究的Jacob Coxon,在X上宣布辭職,直言兩家公司「正朝自我改進的超級智慧狂奔,拿我們的性命當賭注」。這則貼文24小時內衝破9千萬瀏覽,逼得Anthropic對齊科學主管Evan Hubinger親自出面回應——而他的回應,反而讓事情更難以一笑置之。
從辭職貼文到官方證實:AI滅絕風險有多真實
Coxon在貼文中寫道:「打造AI的人真心相信,這東西可能在這個十年結束前害死我們所有人。」他後續接受《華爾街日報》採訪時更具體:「照現在最激進的情境發展,明年年底前局面可能就已經失控。」外界原本可能把這當成離職員工的情緒發言,但Hubinger的回應改變了風向——他公開表示自己與同事「真心相信AI可能殺死所有人」,並把自己對未來十年的AI滅絕風險估計訂在「超過10%」。他同時承認,Anthropic目前並沒有解決「超級智慧對齊」問題的具體計畫,公司也「還稱不上穩定朝目標邁進」。
不是第一次,也不會是最後一次
今年2月,Anthropic另一位AI安全主管Mrinank Sharma也在離職聲明中寫下「世界正處於危機之中」,理由不只是AI,還包括生技威脅等一連串交織的危機。半年內兩位安全團隊要角相繼離開並公開示警,說明這不是單一個人的情緒反應,而是團隊內部長期存在的不安。
誠實地說:讀者該怎麼看待這件事
把兩件事分開看比較實際。第一,「10%的AI滅絕風險」是Anthropic內部研究者的估計值,不是可驗證的科學結論——不同人給的數字差異極大,從個位數到「明年就可能失控」都有人講。第二,這不影響你今天用ChatGPT寫信、用Fable改程式碼的日常風險,那是另一套問題(隱私、幻覺、資安)。真正該關注的是:
- 公司自己的安全承諾兌現進度——各家實驗室是否真的照RSP(負責任擴展政策)踩煞車
- 監管動向——業界標準組織的討論才剛開始,還沒有實質約束力
- 不要用單一病毒式貼文的瀏覽數當作風險指標——90分鐘的關注度跟真實機率是兩回事
AI安全的辯論不會因為一篇貼文結束,但也不需要因為一篇貼文就改變你今天的工作方式。好不好用,試了才知道。
🇺🇸 AI Extinction Warning Review: Anthropic Quits Over 10% Odds
AI extinction risk is the AI story everyone is actually talking about this week — not a new model, a new chip, or a funding round. On September 8, Jacob Coxon, a pretraining researcher who spent three years at both OpenAI and Anthropic, resigned and posted on X that both companies are "racing straight to self-improving superintelligence and gambling with our lives." The post hit more than 90 million views in under 24 hours, and it forced Anthropic's own Alignment Science Lead, Evan Hubinger, to respond publicly — a response that made the story harder to dismiss.
From a Resignation Post to an On-the-Record Admission
Coxon's post argued that "the people building AI earnestly believe it could kill us all by the end of the decade." In a follow-up Wall Street Journal interview, he got more specific: "We're on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already." That could have stayed a disgruntled-employee story. Instead, Hubinger confirmed on X that he and his colleagues "earnestly believe AI could kill all humans," putting his own AI extinction risk estimate at "more than 10%" over the next decade — while admitting Anthropic doesn't yet have a concrete plan for "alignment for superintelligence" and isn't "clearly on track" to get one.
Not the First Warning This Year
In February, another Anthropic safety lead, Mrinank Sharma, resigned saying "the world is in peril" — citing AI alongside a string of interconnected crises like bioweapons risk. Two senior safety staffers leaving with public warnings inside seven months isn't one person's bad week; it points to unresolved tension inside the team actually responsible for keeping frontier models safe.
The Honest Take: What This Actually Means for You
Separate two things. First, that "10%+ extinction risk" number is one researcher's internal estimate, not a peer-reviewed figure — estimates from people inside these labs range from single digits to "could be out of control next year." Second, none of this changes the everyday risk profile of using ChatGPT or Claude Code today; that's a different set of problems (privacy, hallucination, security). What's actually worth tracking:
- Whether labs follow through on their own Responsible Scaling Policies, not just their statements
- Whether the proposed industry safety standards body (Anthropic, OpenAI, and DeepMind have reportedly been meeting since July) produces anything with teeth
- Not treating viral view counts as a probability estimate — 90 million views in a day measures attention, not risk
The safety debate isn't going away, but one viral post shouldn't change how you use these tools tomorrow. 好不好用,試了才知道 — you only really know if it works by trying it.
Sources / 資料來源
- TechCrunch: 'Gambling with our lives' — Anthropic researcher quits
- CNBC: Anthropic researcher quits over AI safety concerns
- CBS News: Ex-Anthropic researcher Jacob Coxon warns AI could grow smart enough to kill us
延伸閱讀 / Related Articles
- SWE-2評測:Cognition新模型逼近Fable卻只綁Devin | SWE-2 Review: Cognition's Cheap Coding Model, Devin-Only
- Fable 5.1評測:快取讀取降75%,代理工作省45% | Fable 5.1 Review: Cache Reads Cut 75%, Agents Save 45%
- Fugu Ultra v2評測:不練模型改練調度的AI | Fugu Ultra v2 Review: The Model That's Actually a Team
AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends
留言
張貼留言