Claude黎曼猜想評測:AI數學研究創37年最大進展 | Claude Riemann Hypothesis Review: AI Breaks 37-Year Record
By Kit 小克 | AI Tool Observer | 2026-09-28
🇹🇼 Claude黎曼猜想評測:AI數學研究創37年最大進展
Claude黎曼猜想研究是這週最值得關注的AI新聞:Anthropic用一個尚未公開發表的研究版Claude去挑戰數學史上最著名的未解問題之一,雖然沒有解開黎曼猜想本身,卻把「滿足黎曼猜想的零點比例」下界從41.6%一口氣推進到67.2%,是過去37年來人類數學家累積進展(僅0.8個百分點)的30幾倍。這篇文章老實測評:這到底是不是真突破,一般人該怎麼看。
Claude做了什麼:先講清楚不是「解出黎曼猜想」
黎曼猜想是關於黎曼zeta函數的零點是否都落在「臨界線」上的猜想,至今無人證明,懸賞百萬美元。Anthropic內部員工在Claude Code環境中提示一個研究版Claude去嘗試這個問題。第一次嘗試,Claude生成並測試了650個不同想法,全部失敗。被要求再試一次後,Claude花了約一天半時間,協調約60個子代理(subagent),跑了大約2,400條shell指令、寫了數百個Python腳本,逐一對照已知的zeta零點數據做交叉驗證,最終產出一個把下界從41.6%推進到67.2%的證明。
關鍵是:Anthropic兩位內部數學家審查並驗證了這篇論文,Claude也同時產出一份可形式化驗證(formally verifiable)的證明版本。Anthropic自己也明講,不認為這個方法會通往完整證明黎曼猜想。
這件事對開發者/研究者的實際意義
- 多代理協作是關鍵:不是單一次prompt出奇蹟,而是60個subagent長時間迭代、互相交叉驗證,這種「暴力試錯+嚴謹驗證」模式在數學這種對錯分明的領域特別有效,因為證明錯就是錯,AI沒辦法唬弄過關。
- 形式化驗證降低幻覺風險:Claude產出的證明可以被機器驗證,這比純自然語言宣稱「我證明了」可信得多,也是目前AI做科學研究最該學的方向。
- 不是人人能重現:這是未發表的研究版模型、跑了一天半、動用大量算力才做到,跟你手上訂閱的Claude Opus或Fable是兩回事,短期內不會變成一般用戶的日常功能。
老實說:該興奮還是該冷靜?
兩者都對一點。興奮的理由是:這證明LLM agent在嚴謹的形式化領域(數學證明)確實能做出人類數十年沒做到的漸進式貢獻,而且過程可驗證、不是話術。冷靜的理由是:這離「AI解決黎曼猜想」或「AI取代數學家」還非常遠,Anthropic自己都沒有把話說滿。把它當成「agentic AI + 大量算力 + 嚴謹驗證」的一個成功案例來看,比較實際。
好不好用,試了才知道。
🇺🇸 Claude Riemann Hypothesis Review: AI Breaks 37-Year Record
The Claude Riemann Hypothesis research is this week's most notable AI story: Anthropic had an unreleased research version of Claude attempt one of mathematics' most famous unsolved problems. It didn't solve the Riemann Hypothesis, but it pushed the lower bound on the proportion of zeta function zeros satisfying the hypothesis from 41.6% to 67.2%, more than 30x the progress human mathematicians made over the previous 37 years (just 0.8 percentage points). Here is an honest look at whether this is a real breakthrough and what it means.
What Claude Actually Did: Let's Be Clear, Not "Solved"
The Riemann Hypothesis concerns whether all non-trivial zeros of the Riemann zeta function lie on the "critical line". It remains unproven, with a million-dollar prize attached. An Anthropic staffer prompted a research version of Claude inside Claude Code to take a stab at it. On the first attempt, Claude generated and tested 650 different ideas, all of which failed. Told to try again, Claude spent roughly a day and a half coordinating about 60 subagents, running around 2,400 shell commands and writing hundreds of Python scripts to cross-check numerical claims against known zeta zeros, eventually producing a proof that raised the lower bound to 67.2%.
Critically, two Anthropic mathematicians reviewed and validated the result, and Claude also produced a formally verifiable version of the proof. Anthropic itself is clear that it does not expect this technique to lead to a full proof of the Riemann Hypothesis.
What This Actually Means for Developers and Researchers
- Multi-agent orchestration was the key: this was not a single lucky prompt, it was 60 subagents iterating and cross-checking each other over an extended run. This "brute-force exploration plus rigorous verification" pattern works especially well in math, where a wrong proof is unambiguously wrong; there is no room for the model to bluff its way through.
- Formal verification cuts hallucination risk: because the proof can be machine-checked, it is far more trustworthy than a plain-language claim of "I proved it", arguably the direction AI-for-science needs to go.
- Not reproducible by regular users: this used an unreleased research model running for 1.5 days with heavy compute, very different from the Claude Opus or Fable subscription you use daily, and will not become an everyday feature anytime soon.
Honest Take: Hype or Genuine Progress?
Both, a little. The exciting part: this shows LLM agents can make genuine, verifiable incremental contributions in a strict formal domain like math proofs, progress humans had not managed in decades, and the process is checkable, not just marketing talk. The reality check: this is nowhere near "AI solves the Riemann Hypothesis" or "AI replaces mathematicians," and Anthropic itself does not oversell it. Best framed as one successful case study in agentic AI plus heavy compute plus rigorous verification, not a paradigm shift, yet.
好不好用,試了才知道 (you won't know until you try it).
Sources / 資料來源
- Anthropic官方研究頁面:Claude improved a lower bound for Riemann zeta zeros
- Anthropic官方X公告串(Riemann Hypothesis attempt)
- arXiv論文:More than two thirds of the zeta zeros are simple and on the critical line
延伸閱讀 / Related Articles
- Claude CRISPR酵素發現評測:AI做科學研究靠譜嗎 | Claude CRISPR Enzyme Discovery Review: Can AI Do Science?
- OpenAI暫停訓練評測:AI代理靠DNS漏洞逃出沙盒 | OpenAI Training Pause Review: AI Agent's DNS Sandbox Escape
- Qwen-Audio 3.1評測:語音API砍價95%,開發者該換嗎 | Qwen-Audio 3.1 Review: Voice API Prices Cut Up to 95%
AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends
留言
張貼留言