R&D自動化指數評測:Claude扛下26%自家AI研發 | R&D Automation Index Review: Claude Leads 26% of AI R&D
By Kit 小克 | AI Tool Observer | 2026-09-21
🇹🇼 R&D自動化指數評測:Claude扛下26%自家AI研發
R&D自動化指數(R&D Automation Index)是Anthropic在2026年9月17日首次公開的內部指標,用來衡量Claude在自家AI研發流程裡到底扛了多少工作。答案:26%的研發任務已經達到「AI主導、人類只設大方向」的自動化等級,而且這個數字在今年2月還是0%。換句話說,短短半年,Claude從輔助工具變成了研發流程裡真正扛活的角色。
26%是怎麼算出來的?
Anthropic把公司內部所有AI研發相關的工作項目都列出來,套用Epoch AI設計的自動化等級量表(AL0到AL5,AL0是全人工、AL5是AI完全自主無人監督),再依照每項工作實際消耗的員工工時去加權平均。目前26%的工作落在「AI主導多數步驟、人類設定方向與把關」的等級,而超過90%的研發工作有Claude深度參與(不管是主導還是協作)。
30,000個Agent同時在跑,代表什麼?
更驚人的數字是:Anthropic內部最常用的平台上,8月份平均同時有約3萬個Claude Agent在跑研發和工程任務。Anthropic強調自己有做監控——線上監控涵蓋100%的agent行動、事前攔截率約0.002%(大概每4.7萬個決策攔一個),另外離線審查會抽查千分之一到千分之二的紀錄。這代表Anthropic自己也承認:agent數量一多,光靠人力盯根本盯不過來,只能靠更多AI去監督AI。
誠實地說,這數字該怎麼解讀
- 這是自我揭露,不是第三方稽核:26%、90%、3萬個agent,全部是Anthropic自己量測、自己公布,目前沒有獨立機構驗證方法論或原始數據。
- 沒有任何一個環節是「完全自主」:Anthropic在報告裡明講,Claude目前在任何被量測的研發子項目裡都不是零人類監督在跑。
- 時機點很微妙:這份報告是在Dario Amodei持續呼籲業界「放慢腳步」、Sam Altman與Elon Musk也公開附和的背景下發布的——一邊喊煞車,一邊秀出AI研發自己的進度飛快,這兩件事同時存在。
對開發者跟企業主來說,這份指數真正有意義的地方,不是「AI要取代研究員了」的聳動標題,而是它示範了一種可以自己套用的框架:把你團隊的工作拆成任務清單,評估每項任務目前AI能扛到什麼程度,用實際工時加權,就能得到一個誠實的「AI自動化滲透率」。這比空談「我們有沒有導入AI」實際多了。
好不好用,試了才知道。
🇺🇸 R&D Automation Index Review: Claude Leads 26% of AI R&D
The R&D Automation Index is a metric Anthropic disclosed publicly for the first time on September 17, 2026, measuring how much of its own AI research and development work is actually being carried by Claude. The headline number: 26% of R&D tasks have reached an automation level where AI leads most of the work and humans just set broad direction, up from 0% back in February. In half a year, Claude went from assistant to genuine load-bearing contributor in Anthropic's own research pipeline.
How Anthropic Calculated the 26%
Anthropic catalogued every category of AI R&D work inside the company, then rated each task using an automation scale developed with Epoch AI, ranging from AL0 (no AI involvement) to AL5 (AI operates fully autonomously, no human in the loop). Each task's rating is weighted by how much staff time it actually consumes. Right now, 26% of that weighted work sits at the level where AI leads most steps while humans set direction and check the output, and Claude participates in some form in over 90% of all R&D work.
What 30,000 Concurrent Agents Actually Means
The more striking number: on Anthropic's most-used internal platform, roughly 30,000 Claude agents were running research and engineering tasks concurrently at any given moment in August. Anthropic says it monitors 100% of agent actions before execution, blocking about 0.002% of decisions (roughly 1 in 47,000), with offline review sampling another 1-2 transcripts per thousand. Read between the lines: at that scale, human review alone cannot keep up, so Anthropic is already relying on AI to watch AI.
The Honest Read
- This is self-reported, not independently audited. The 26%, the 90%, the 30,000 agents, all measured and published by Anthropic itself. No outside body has verified the methodology or raw data yet.
- Nothing here is fully autonomous. Anthropic explicitly states Claude is not operating without human oversight in any measured subset of R&D work.
- The timing is notable. This report lands while Dario Amodei keeps pushing the industry to slow down, with Sam Altman and Elon Musk publicly agreeing, even as the same company shows its own AI-building-AI numbers climbing fast.
For developers and founders, the real value of this index is not the "AI is coming for researchers" headline. It is the reusable framework: break your team's work into a task list, rate how automated each task currently is, weight by actual hours spent, and you get an honest automation-penetration number for your own org. That is a lot more useful than vague talk about "adopting AI."
好不好用,試了才知道。
Sources / 資料來源
- Anthropic: Measurements for understanding the pace of AI development inside frontier labs
- Anthropic says Claude leads 26% of its AI R&D work — Quartz
- Anthropic says Claude is helping to build the next version of itself — Washington Times
延伸閱讀 / Related Articles
- AI幻覺評測:美軍假情報險些引爆對中衝突 | AI Hallucination Review: Military Intel Nearly Sparked War
- Navier-Stokes評測:OpenAI稱破千禧年難題,遭25位菲爾茲得主抗議 | Navier-Stokes Review: OpenAI's Proof Ignites a Credit War
- Jalapeño評測:OpenAI用AI設計晶片,宣稱甩輝達3.6倍 | Jalapeño Review: OpenAI's AI-Designed Chip Beats Nvidia
AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends
留言
張貼留言