跳到主要內容

AI滅絕人類評測:Anthropic離職研究員估破10%機率 | AI Doom Risk Review: Anthropic Quits Over 10% Odds

By Kit 小克 | AI Tool Observer | 2026-09-16

🇹🇼 AI滅絕人類評測:Anthropic離職研究員估破10%機率

Anthropic 一名離職研究員上週在社群媒體公開警告,AI滅絕人類的風險已經不是網路酸民哏,而是公司內部人親口證實的數字。做過 OpenAI 與 Anthropic 兩家預訓練研究、年資三年的 Jacob Coxon 辭職後直接開嗆:兩家公司都在不計代價衝刺自我改進型超級智慧,等於拿全人類的命去賭。這篇文章用 no-hype 角度整理這起事件的來龍去脈,也講清楚這件事對一般 AI 用戶到底有沒有實質影響。

事件經過:一則辭職貼文,燒成整個業界的話題

2026 年 9 月 9 日,Jacob Coxon 的辭職貼文在 X 與 Hacker News 上迅速被轉發。真正讓話題升溫的是 Anthropic 對齊科學負責人(Alignment Science Lead)Evan Hubinger 隨後跟進發文:他和同事「真心相信」AI 有可能在十年內滅絕全人類,他自己給出的機率估計超過 10%。Hubinger 強調公司確實在努力解決對齊問題,但目前還沒有解方,也不敢說走在正確的軌道上。

為什麼這次的AI滅絕人類警告跟以往不同

  • 過去喊 AI 滅絕風險的多半是外部學者(如 Bengio)或末日論網紅,這次是 Anthropic 現任對齊團隊主管親口承認
  • 同一週,Anthropic 發布 9 月版威脅情報報告,證實已攔截 5 起利用 Claude 協助生物武器研究的案例,還有來自俄羅斯與中國 Alibaba 團隊的大規模模型濫用行為,把「風險」從理論變成有真實案例撐腰的敘事
  • 兩件事疊加,讓外界很難再把這輪警告當成純粹的行銷話術

業界懷疑論:這會不會只是換個方式打廣告

不是所有人都買單。Hacker News 與 Reddit r/artificial 上不少人質疑,把「滅絕人類」機率掛在嘴邊某種程度也是幫 Anthropic 的模型能力做免費宣傳——講得越危險,越顯得 Claude 很強。也有人指出 10% 只是個人主觀估計,沒有嚴謹方法論支撐,媒體標題化之後容易失真,這種懷疑論同樣該放進理解框架裡。

一般用戶跟開發者實際上該注意什麼

老實說,這波辯論短期內不會改變你用 Claude 寫程式或整理資料的體驗。但有幾件事值得放在心上:

  • Anthropic 內部對齊研究資源可能會加碼,新模型上線速度或功能開放範圍可能更保守
  • 企業拿 Claude 做生醫、國防等敏感應用,審查與使用限制大概率會更嚴
  • 這類公開辯論持續推高監管單位對前沿模型的關注,間接影響 API 條款與地區可用性

AI滅絕人類的機率沒有人能給出可驗證答案,但當公司內部對齊團隊主管公開承認風險超過 10%,這已經不只是思想實驗,而是會實際牽動產品節奏跟監管走向的訊號。不用恐慌,但也不該當背景噪音直接滑過。

好不好用,試了才知道。


🇺🇸 AI Doom Risk Review: Anthropic Quits Over 10% Odds

An Anthropic researcher went viral last week for warning that AI could kill all humans — and this time it wasn't an outside critic, it was someone who had spent three years doing pretraining research at both OpenAI and Anthropic. Jacob Coxon resigned and posted that both companies are racing toward self-improving superintelligence with no real safety plan, effectively gambling with everyone's lives. Here is a no-hype rundown of what actually happened, and whether it changes anything for people who just use these tools.

What Happened: One Resignation Post That Blew Up

Coxon resignation post spread fast across X and Hacker News on September 9, 2026. What really escalated the story was Evan Hubinger, Anthropic Alignment Science Lead, following up to confirm he and his colleagues "earnestly believe" AI could kill all of humanity within the decade — and that his own probability estimate is north of 10%. Hubinger said the company is genuinely trying, but does not yet have a working plan for aligning superintelligent systems and cannot claim it is on track to find one.

Why This AI Doom Warning Hits Different

  • Past extinction-risk warnings mostly came from outside academics (like Yoshua Bengio) or online doomers — this time it is Anthropic own current alignment lead saying it on the record
  • The same week, Anthropic published its September threat intelligence report confirming it disrupted five cases of Claude being used to support bioweapons research, plus large-scale misuse tied to Russian actors and an Alibaba-linked distillation campaign — turning "risk" from theory into something backed by documented cases
  • Together, the two stories make it harder to write this off as pure marketing spin

The Skeptical Take: Is This Just Marketing in Disguise

Not everyone is buying it. Plenty of commenters on Hacker News and Reddit r/artificial pointed out that talking up a double-digit extinction probability doubles as free advertising for how capable Claude supposedly is — the scarier it sounds, the more powerful the model looks. Others note that a 10% figure is a personal, subjective estimate with no rigorous methodology behind it, and headlines tend to flatten that nuance. Both of those caveats belong in how you read this story.

What Actually Matters for Users and Developers

Realistically, none of this changes your day-to-day experience using Claude to write code or process data — not this week, anyway. But a few things are worth tracking:

  • Anthropic internal alignment research may get more resources, which could mean slower rollouts or more conservative feature gating on future models
  • Enterprises using Claude for sensitive use cases — biomedical, defense-adjacent — should expect tighter review and usage restrictions
  • This kind of public debate keeps pushing regulators to pay closer attention to frontier models, which can indirectly affect API terms and regional availability

Nobody can give you a verifiable number for the odds that AI could kill all humans. But when a company own alignment lead is willing to put a double-digit estimate on the record, this stops being a pure thought experiment — it is a signal that will actually shape product pacing and regulation. No need to panic, but it is not background noise you should scroll past either.

You won't know until you try it.

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code