跳到主要內容

AI末日警告評測:Anthropic研究員辭職曝10%滅絕機率 | AI Extinction Warning Review: Anthropic Quits Over 10% Odds

By Kit 小克 | AI Tool Observer | 2026-09-15

🇹🇼 AI末日警告評測:Anthropic研究員辭職曝10%滅絕機率

AI滅絕風險這幾天成了全球AI圈最熱的話題。9月8日,曾在OpenAI與Anthropic做了三年預訓練研究的Jacob Coxon,在X上宣布辭職,直言兩家公司「正朝自我改進的超級智慧狂奔,拿我們的性命當賭注」。這則貼文24小時內衝破9千萬瀏覽,逼得Anthropic對齊科學主管Evan Hubinger親自出面回應——而他的回應,反而讓事情更難以一笑置之。

從辭職貼文到官方證實:AI滅絕風險有多真實

Coxon在貼文中寫道:「打造AI的人真心相信,這東西可能在這個十年結束前害死我們所有人。」他後續接受《華爾街日報》採訪時更具體:「照現在最激進的情境發展,明年年底前局面可能就已經失控。」外界原本可能把這當成離職員工的情緒發言,但Hubinger的回應改變了風向——他公開表示自己與同事「真心相信AI可能殺死所有人」,並把自己對未來十年的AI滅絕風險估計訂在「超過10%」。他同時承認,Anthropic目前並沒有解決「超級智慧對齊」問題的具體計畫,公司也「還稱不上穩定朝目標邁進」。

不是第一次,也不會是最後一次

今年2月,Anthropic另一位AI安全主管Mrinank Sharma也在離職聲明中寫下「世界正處於危機之中」,理由不只是AI,還包括生技威脅等一連串交織的危機。半年內兩位安全團隊要角相繼離開並公開示警,說明這不是單一個人的情緒反應,而是團隊內部長期存在的不安。

誠實地說:讀者該怎麼看待這件事

把兩件事分開看比較實際。第一,「10%的AI滅絕風險」是Anthropic內部研究者的估計值,不是可驗證的科學結論——不同人給的數字差異極大,從個位數到「明年就可能失控」都有人講。第二,這不影響你今天用ChatGPT寫信、用Fable改程式碼的日常風險,那是另一套問題(隱私、幻覺、資安)。真正該關注的是:

  • 公司自己的安全承諾兌現進度——各家實驗室是否真的照RSP(負責任擴展政策)踩煞車
  • 監管動向——業界標準組織的討論才剛開始,還沒有實質約束力
  • 不要用單一病毒式貼文的瀏覽數當作風險指標——90分鐘的關注度跟真實機率是兩回事

AI安全的辯論不會因為一篇貼文結束,但也不需要因為一篇貼文就改變你今天的工作方式。好不好用,試了才知道。


🇺🇸 AI Extinction Warning Review: Anthropic Quits Over 10% Odds

AI extinction risk is the AI story everyone is actually talking about this week — not a new model, a new chip, or a funding round. On September 8, Jacob Coxon, a pretraining researcher who spent three years at both OpenAI and Anthropic, resigned and posted on X that both companies are "racing straight to self-improving superintelligence and gambling with our lives." The post hit more than 90 million views in under 24 hours, and it forced Anthropic's own Alignment Science Lead, Evan Hubinger, to respond publicly — a response that made the story harder to dismiss.

From a Resignation Post to an On-the-Record Admission

Coxon's post argued that "the people building AI earnestly believe it could kill us all by the end of the decade." In a follow-up Wall Street Journal interview, he got more specific: "We're on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already." That could have stayed a disgruntled-employee story. Instead, Hubinger confirmed on X that he and his colleagues "earnestly believe AI could kill all humans," putting his own AI extinction risk estimate at "more than 10%" over the next decade — while admitting Anthropic doesn't yet have a concrete plan for "alignment for superintelligence" and isn't "clearly on track" to get one.

Not the First Warning This Year

In February, another Anthropic safety lead, Mrinank Sharma, resigned saying "the world is in peril" — citing AI alongside a string of interconnected crises like bioweapons risk. Two senior safety staffers leaving with public warnings inside seven months isn't one person's bad week; it points to unresolved tension inside the team actually responsible for keeping frontier models safe.

The Honest Take: What This Actually Means for You

Separate two things. First, that "10%+ extinction risk" number is one researcher's internal estimate, not a peer-reviewed figure — estimates from people inside these labs range from single digits to "could be out of control next year." Second, none of this changes the everyday risk profile of using ChatGPT or Claude Code today; that's a different set of problems (privacy, hallucination, security). What's actually worth tracking:

  • Whether labs follow through on their own Responsible Scaling Policies, not just their statements
  • Whether the proposed industry safety standards body (Anthropic, OpenAI, and DeepMind have reportedly been meeting since July) produces anything with teeth
  • Not treating viral view counts as a probability estimate — 90 million views in a day measures attention, not risk

The safety debate isn't going away, but one viral post shouldn't change how you use these tools tomorrow. 好不好用,試了才知道 — you only really know if it works by trying it.

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code