跳到主要內容

OpenAI Astra評測:AI用2000美元解開10道數學難題 | OpenAI Astra Review: AI Solves 10 Math Problems for $2K

By Kit 小克 | AI Tool Observer | 2026-08-05

🇹🇼 OpenAI Astra評測:AI用2000美元解開10道數學難題

OpenAI Astra在8月1日投下一顆震撼彈:這個內部版本的下一代模型,用大約2000美元的API token成本,解開了10道橫跨數學與理論電腦科學、部分已經卡關超過27年的公開難題,而且每一個證明都經過Lean 4形式化驗證、可公開下載自行檢查。這不是又一次「AI寫詩感人」等級的行銷炒作,而是第一次把可驗證的形式證明直接甩到數學社群桌上。

OpenAI Astra是什麼?

Astra是OpenAI下一代模型家族的內部代號,設計重點是長時間協調多個推理代理(agent)完成任務,被視為o系列測試時推理路線的延伸。目前僅有內部版本存在,尚未透過API或產品對外開放。

Astra怎麼解開10道數學難題?

這10道題目橫跨群論、von Neumann代數、高維幾何、量子複雜度、格密碼學與極值組合學等8個領域。其中最受矚目的是首次明確構造出「非sofic群」——這個問題自1999年Gromov提出sofic概念後,整整27年沒人解開。研究人員把AI產生的論證改寫成正式論文,再由Astra把每一步證明形式化成Lean 4證明證書,全部249頁手稿與證書都以Apache 2.0授權公開在GitHub,且「sorry」計數為零,代表沒有任何一步是靠人工背書、每一行都通過機器檢查。

這些證明真的可信嗎?

可信度是這次事件最大的亮點:因為用Lean形式驗證,任何人都能下載證書自己跑一遍檢查器,不需要相信OpenAI的一面之詞。費爾茲獎得主Timothy Gowers表示,其中一篇證明他會毫不猶豫推薦刊登在頂級期刊。但OpenAI研究員Noam Brown自己也提醒:「可惜還沒解出任何一道千禧年大獎難題」,坦言這次挑的10題雖然懸而未決很久,但難度上未必是數學界最硬的骨頭——能被中等算力預算解開的題目,通常本來就不是那種抵抗人類集體努力最久的頂級難題。

對研究者與開發者有什麼實際意義?

如果你是研究人員,Astra展示的形式驗證流程值得關注:AI生成證明、人類整理成論文、再用Lean把整條推理鏈釘死,這套組合讓「AI亂掰」的風險大幅降低。但對一般開發者來說,Astra目前只是內部原型,沒有API、沒有時程表,能不能撐起日常研究工作流還是未知數。

好不好用,試了才知道。


🇺🇸 OpenAI Astra Review: AI Solves 10 Math Problems for $2K

OpenAI Astra, an internal build of the company's next major model, spent roughly $2,000 in API tokens to crack ten unsolved problems in math and theoretical computer science — some open for more than 27 years — and published every proof as a machine-checked Lean 4 certificate anyone can download and verify. This isn't another "AI wrote a nice poem" headline; it's the first time a verifiable formal proof has been dropped directly onto the math community's desk.

What Is OpenAI Astra?

Astra is the internal codename for OpenAI's next model family, built around coordinating multiple reasoning agents over long-running tasks — an extension of the test-time reasoning approach behind the o-series. Only an internal version exists so far; there's no public API or product release.

How Did Astra Solve 10 Open Math Problems?

The ten problems span eight fields, including group theory, von Neumann algebras, high-dimensional geometry, quantum complexity, lattice cryptography, and extremal combinatorics. The headline result is the first explicit construction of a non-sofic group, a question that sat unanswered for 27 years since Gromov introduced soficity in 1999. Human researchers turned Astra's arguments into manuscripts, and Astra then formalized each proof as a Lean 4 certificate. The full 249-page manuscript and proof files are on GitHub under Apache 2.0, with a "sorry" count of zero — meaning every single step passed automated checking, with no human hand-waving.

Can You Actually Trust These Proofs?

Verifiability is the real story here: because the proofs are formalized in Lean, anyone can run the checker themselves instead of taking OpenAI's word for it. Fields Medalist Timothy Gowers said he'd recommend one proof for a top journal without hesitation. But OpenAI's own Noam Brown added a note of restraint: "Sadly, no Millennium Prize Problems (yet)." Problems solvable within a modest compute budget, on average, aren't the ones that have resisted the field's hardest collective effort — worth remembering before calling this an AGI moment.

What Does This Mean for Researchers and Developers?

For researchers, the workflow matters more than the headline: AI generates an argument, humans write it up, and Lean nails down every logical step — a combination that meaningfully cuts the risk of AI just making things up. For everyday developers, Astra is still an internal prototype with no API and no public timeline. Whether it holds up as a daily research tool remains to be seen.

好不好用,試了才知道 — works only when you've actually tried it.

Sources / 資料來源

常見問題 FAQ

OpenAI Astra是什麼?

Astra是OpenAI下一代模型家族的內部代號,擅長協調多個推理代理處理長任務,目前僅有內部版本,尚未公開發布。

Astra真的解開千禧年大獎難題了嗎?

沒有。OpenAI研究員Noam Brown親自澄清,這次解的是10道存在數十年的公開問題,並非千禧年七大難題等級的突破。

Astra的數學證明可以相信嗎?

可以驗證。這些證明都用Lean 4形式化並公開在GitHub,任何人都能下載自行跑一次檢查器,「sorry」計數為零代表每一步都通過機器驗證。

一般開發者現在能用Astra嗎?

不行。Astra目前只是內部原型,沒有對外API、沒有公開時程表,能否成為日常研究工具還有待觀察。

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?