跳到主要內容

OpenAI 722篇數學論文評測:菲爾茲獎得主集體開砲 | OpenAI 722 Math Papers Review: Fields Medalists Revolt

By Kit 小克 | AI Tool Observer | 2026-10-08

🇹🇼 OpenAI 722篇數學論文評測:菲爾茲獎得主集體開砲

OpenAI 10月7日在GitHub上一次丟出一個叫「math」的repo,裡面塞了722篇數學手稿,宣稱解開涵蓋372組未解數學問題家族的結果,其中包括電腦科學理論界的知名難題Unique Games Conjecture(唯一博弈猜想),還牽涉到quasi-Riemann相關結果。數字聽起來很嚇人,但這篇不是要幫你歡呼,而是要老實講清楚:這722篇論文裡面,到底有多少是真的站得住腳。

OpenAI到底發布了什麼

這批論文沒有經過正常期刊的同行審查流程,是直接上GitHub公開的。更關鍵的是:產出這些結果的模型本身沒有公開、沒有命名,外界完全無法驗證這些「證明」是怎麼被生成出來的,也拿不到模型去重現或檢驗過程。

菲爾茲獎得主為什麼集體開砲

這次反應最激烈的不是一般網友,而是數學界的最高榮譽——25位菲爾茲獎得主連署了一份聲明,標題直接叫做「A Severe Misalignment of AI in Mathematics」(AI在數學領域的嚴重失準)。聲明指控OpenAI無視今年8月一場閉門會議上學界要求遵守學術發表規範的呼籲。更早之前,還有指控OpenAI在Navier-Stokes千禧年問題上「搶跑」了紐約大學數學家與一名Anthropic員工的合作研究,被部分學者形容為「像黑手黨」的行為。

驗證率只有22%,這才是重點

722篇論文裡,只有162篇的主要結果經過Lean(一套形式化證明驗證工具)電腦驗證,比例大約22%。換句話說,剩下將近八成的「證明」目前只有AI自己說了算,還沒有被第三方證實真的站得住腳。這也是為什麼數學圈的第一反應是警戒而不是鼓掌——把未解問題變成AI實驗室之間比大小的戰場,產出、檢驗、掛名的流程可能來不及跟上。

對你我有什麼意義

如果你用AI做研究輔助、寫論文或任何需要嚴謹驗證的工作,這次事件提醒一件很實際的事:AI產出的「證明」或「答案」不等於「正確」,尤其是在沒有第三方驗證的狀況下。目前唯一可信的把關方式,是用Lean之類的形式化驗證工具交叉確認每一步邏輯。下次看到「AI解開300多個數學難題」這種標題,先問一句:驗證了嗎?誰驗證的?驗證了多少比例?

好不好用,試了才知道。


🇺🇸 OpenAI 722 Math Papers Review: Fields Medalists Revolt

On October 7, OpenAI dropped a GitHub repo called "math" containing 722 manuscripts, claiming results across 372 families of long-unsolved math problems — including the theoretical computer science holy grail known as the Unique Games Conjecture, plus work touching quasi-Riemann territory. The numbers are eye-popping. This post isn't here to cheer — it's here to tell you honestly how much of this 722-paper pile actually holds up.

What OpenAI Actually Released

These papers skipped the normal peer-review pipeline entirely and went straight to GitHub. More importantly: the model that generated them is unnamed and unreleased — nobody outside OpenAI can inspect how these "proofs" were produced, let alone reproduce them.

Why Fields Medalists Are Furious

The loudest pushback isn't coming from random commenters — it's coming from math's highest honor. 25 Fields Medalists signed a declaration titled "A Severe Misalignment of AI in Mathematics." It accuses OpenAI of ignoring pleas, made at a closed-door meeting back in August, to follow academic publication norms. Earlier allegations claim OpenAI effectively scooped collaborative research between an NYU mathematician and an Anthropic researcher working on the Navier-Stokes Millennium Prize problem — conduct some mathematicians bluntly called "mafia-like."

The Number That Actually Matters: 22%

Out of 722 papers, only 162 have had their main results computer-verified using Lean, a formal proof-checking tool — roughly 22%. That means nearly 80% of these "proofs" are, for now, verified by nobody but the AI itself. That's exactly why the math community's first reaction was alarm, not applause: turning open problems into a scoreboard between AI labs risks outpacing the field's ability to actually check and credit the work.

What This Means for You

If you use AI for research assistance, paper-writing, or anything requiring rigorous verification, this episode is a practical reminder: an AI-generated "proof" or "answer" is not the same as a correct one, especially without independent verification. Right now, cross-checking every logical step with a formal tool like Lean is the only verification method anyone actually trusts. So next time a headline says "AI solves 300+ math problems," ask: verified by whom, and what percentage?

You won't know until you try it.

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code