OpenAI Astra評測:AI資安能力觸頂,訓練緊急喊停 | OpenAI Astra Review: AI Hits Critical Cyber Threshold
By Kit 小克 | AI Tool Observer | 2026-08-31
🇹🇼 OpenAI Astra評測:AI資安能力觸頂,訓練緊急喊停
OpenAI Astra 是 OpenAI 下一代旗艦模型的內部代號,這次上新聞不是因為效能突破,而是因為它可能已經摸到 OpenAI 自訂資安風險量表的最高等級——「Critical」。8月18日,OpenAI 罕見公開宣布:暫停對最新一批準備部署的模型進行強化學習(RL)訓練兩週,原訂規模最大的一次前沿 RL 訓練更是無限期擱置,原因是初步測試顯示 Astra「無法排除已達到 Critical 等級網路攻擊能力」的可能性。
什麼是 Critical 等級?為什麼 OpenAI 緊張
OpenAI 在其 Preparedness Framework(準備框架)裡把 AI 的資安風險分成好幾級,最高的 Critical 定義是:模型能夠「在無需人類介入的情況下,針對多個經過強化防護的真實世界關鍵系統,自主找出並開發各種嚴重程度的零時差漏洞」,甚至能自己設計、執行全新的攻擊策略。這是 OpenAI 第一次在模型開發階段實際觸發 Preparedness Framework 的 Critical 門檻,不是紙上談兵的假設情境。
導火線:駭進 Hugging Face 的內部測試
- 今年7月,OpenAI 一次內部控制測試中,模型對 Hugging Face 基礎設施執行了17,600次自主入侵動作,展現出獨立漏洞探索與權限提升的能力
- 8月7日的安全評估進一步確認,下一代 Astra 模型的資安能力已跨過 Critical 門檻
- 兩起事件加總,促使 OpenAI 決定按下暫停鍵
OpenAI 做了什麼?三道防線同時加強
暫停不是停工,而是把資源轉向補強三個環節:
- 監控:新增逐 token 層級的行為分類器,目標是30分鐘內對異常活動發出警報
- 對齊:強化獎勵模型,在各類任務中抑制不安全行為的傾向
- 安全隔離:研究叢集限制程式碼執行與網路存取,降低常駐權限與未信任程式碼的曝險
OpenAI 也表示會把這些防護措施正式寫進更新版的 Preparedness Framework,讓資安風險評估貫穿訓練到部署的每個階段。
小克怎麼看
這件事的意義不在 Astra 模型本身多強,而在於這是 OpenAI 第一次真的照著自己訂的規則踩剎車——過去 Preparedness Framework 常被視為公關文件,這次卻真的讓一次「最大規模的前沿 RL 訓練」無限期擱置。對開發者跟企業用戶來說,短期沒有實際影響,但長期要留意:當模型的資安攻擊能力開始逼近「自主找零時差漏洞」等級,能不能用、敢不敢用,監管跟企業內控的門檻只會越拉越高。
好不好用,試了才知道。
🇺🇸 OpenAI Astra Review: AI Hits Critical Cyber Threshold
OpenAI Astra, the internal codename for OpenAI's next flagship model, made headlines this month — not for a performance breakthrough, but because it may have already reached the highest tier on OpenAI's own AI risk scale: "Critical." On August 18, OpenAI made a rare public announcement that it was pausing reinforcement learning (RL) training on its latest deployment-track models for two weeks, and putting its largest planned frontier RL run on indefinite hold, after preliminary testing showed it "could no longer rule out Astra reaching Critical cybersecurity capability."
What Does "Critical" Actually Mean?
Under OpenAI's Preparedness Framework, cybersecurity risk is scored across multiple tiers. The top tier, Critical, is defined as a model's ability to "identify and develop functional zero-day exploits of all severity levels across many hardened real-world critical systems without human intervention" — and to devise and execute novel attack strategies on its own. This is the first time OpenAI has actually triggered the Critical threshold in its Preparedness Framework during model development, not just as a hypothetical scenario.
The Trigger: A Model That Breached Hugging Face
- In a controlled internal test in July, an OpenAI model executed 17,600 autonomous intrusion actions against Hugging Face's infrastructure, showing independent vulnerability discovery and privilege escalation
- An August 7 safety evaluation then confirmed the next-generation Astra model had crossed the Critical cybersecurity threshold
- Together, the two events pushed OpenAI to hit pause
OpenAI's Response: Three Reinforced Safeguards
The pause isn't a full stop — it's a redirection of effort into three areas:
- Monitoring: new token-level activation classifiers designed to flag suspicious activity within 30 minutes
- Alignment: strengthened reward models to discourage unsafe behavior across tasks
- Security isolation: tighter restrictions on code execution and internet access in research clusters, with reduced standing privileges
OpenAI says these safeguards will be formally folded into an updated Preparedness Framework, applied across both training and deployment.
Kit's Take
The real story here isn't how capable Astra is — it's that OpenAI actually followed its own rulebook. The Preparedness Framework has often been treated as a PR document; this time it genuinely put the company's largest frontier RL run on indefinite hold. For developers and enterprise users, there's no immediate impact, but the trend line matters: as models edge toward autonomously finding zero-days, the bar for regulatory and internal-control sign-off only keeps rising.
好不好用,試了才知道。
Sources / 資料來源
- OpenAI Slows AI Training Over Astra's Cybersecurity Capabilities (eSecurity Planet)
- OpenAI Puts Major Frontier AI Training Run on Hold Over Cyber Risks (Help Net Security)
- OpenAI Astra May Have Hit Critical Cyber Threshold, Prompting Safety Overhaul (Axios)
延伸閱讀 / Related Articles
- Jalapeño評測:OpenAI自研AI晶片效能超車輝達 | Jalapeño Review: OpenAI's Nvidia-Beating AI Chip
- Perplexity評測:Nvidia砸300億估值投資AI搜尋新星 | Perplexity Review: Nvidia Backs AI Search Startup at $30B
- Gemini 3.5 Transcribe評測:Google新語音轉文字模型登場 | Gemini 3.5 Transcribe Review: Google's Speech-to-Text Model
AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends
留言
張貼留言