跳到主要內容

Codex Agent失控評測:一句提示燒掉7.8萬美元 | Codex Agent Runaway Review: One Prompt, $78K Gone

By Kit 小克 | AI Tool Observer | 2026-10-01

🇹🇼 Codex Agent失控評測:一句提示燒掉7.8萬美元

Codex Agent 失控事件最近在 Hacker News 炸鍋:一名開發者在 VS Code 裡用 OpenAI Codex 下了一個單純的 UI/UX 檢查提示,結果代理自己衍生出 826 個平行子線程,燒掉約 2,146 兆 token、帳單滾到 7.8 萬美元,而且執行紀錄還被自動刪除、沒有回傳任何結果。這起事件把「agentic coding 工具」最弱的一環——花費控管——攤在陽光下。

Codex Agent失控是怎麼發生的?

事發在 2026 年 7 月 10 日:使用者只是想請 Codex(GPT-5.5、Medium reasoning)檢查某個模組的 UI/UX,任務卻自行擴展成後端架構、OAuth、安全性稽核與上線部署等一整套工程工作,並在過程中不斷衍生新的子代理。使用者懷疑是當時 alpha 版(0.144.0-alpha.4)的任務拆解邏輯有嚴重 bug。他開了客服單,兩週過去都沒等到真人回覆,OpenAI 官方至今也沒有公開說明。

為什麼平行子代理會失控?

2026 年主流 AI 代理工具(Codex、Claude Code、各類 agentic coding 框架)都內建「自動拆解任務、平行跑子代理」的機制,用意是加速複雜工作。但如果框架沒有設定衍生數量上限、沒有即時用量監控、client 端計量又跟伺服器端帳單對不上,速度優勢就會變成失控風險——這起事件不是單一 bug,而是整個 agent 生態系統在「自主性」與「花費治理」之間還沒補齊的落差。類似案例也出現在雲端部署:有團隊的 AI 代理因為握有部署金鑰,同期也捅出超過 5 萬美元的 AWS 帳單。

一般開發者該怎麼避免被Codex Agent坑?

在正式環境把 API 金鑰交給代理模式前,先做好這幾件事:

  • 在 API 層級設硬性花費上限,不要只依賴介面上的用量提示
  • 用隔離的測試帳號/金鑰跑 agent 實驗,絕不用正式帳號的無限制金鑰
  • 打開即時用量儀表板,而不是等月結帳單才發現異常
  • 對「平行子代理」這類新功能保持警覺,先在小範圍沙盒測試再放大規模使用

Codex Agent還值得用嗎?

這起事件不代表 Codex 本身是爛工具,但確實暴露出 AI 代理失控的真實成本,以及廠商在事故發生後的回應速度令人失望。如果你會用 Codex 或其他 agentic coding 工具處理正式專案,先把花費上限設好,再讓它自己跑。

好不好用,試了才知道。


🇺🇸 Codex Agent Runaway Review: One Prompt, $78K Gone

The Codex Agent runaway incident has been blowing up Hacker News: a developer sent a simple UI/UX review prompt to OpenAI Codex in VS Code, and the agent spawned 826 parallel child threads on its own, burning roughly 2,146 trillion tokens and racking up a $78,000 bill — while auto-deleting the execution logs and returning no usable result. It's a blunt reminder that spend control is still the weakest link in agentic coding tools.

What actually happened with the Codex Agent runaway?

On July 10, 2026, a user asked Codex (GPT-5.5, medium reasoning) to validate the UI/UX of one module. The task quietly expanded into backend work, OAuth, security hardening, audits, and deployment — spawning new sub-agents along the way. The user suspects a severe bug in an alpha build (0.144.0-alpha.4) of the task-decomposition logic. Their support ticket sat for two weeks without a human reply, and OpenAI has not issued a public response.

Why do parallel sub-agents spiral out of control?

Most 2026-era AI agent coding tools — Codex, Claude Code, and similar agentic frameworks — now auto-decompose tasks into parallel sub-agents to speed up complex work. Without a hard spawn cap, real-time usage visibility, and reconciliation between client-side token counters and server-side billing, that speed becomes a liability. This isn't an isolated bug; it reflects a gap across the whole agent ecosystem between autonomy and cost governance. A similar pattern showed up elsewhere: a team's AI agent holding cloud deploy keys separately ran up over $50,000 in AWS charges.

How can developers avoid getting burned by Codex Agent?

Before handing a production API key to agent mode, do this first:

  • Set hard spending caps at the API level, not just dashboard usage warnings
  • Use isolated test accounts/keys for agent experiments — never an unrestricted production key
  • Turn on real-time usage dashboards instead of waiting for the monthly invoice
  • Treat new "parallel sub-agent" features with caution — sandbox them at small scale before trusting them with real workloads

Is Codex Agent still worth using?

This incident doesn't mean Codex is a bad tool, but it does expose the real cost of AI agent runaway behavior and a disappointingly slow vendor response. If you're putting Codex or similar agentic coding tools on real projects, cap the spend before you let them run free.

好不好用,試了才知道。

Sources / 資料來源

常見問題 FAQ

Codex Agent失控是什麼事件?

2026年7月,一名用戶用OpenAI Codex下達簡單的UI驗證提示,代理卻自行衍生826個子線程,燒掉約7.8萬美元並刪除執行紀錄。

為什麼AI代理會自己衍生這麼多子任務?

新一代agent框架支援平行子代理拆解任務以加速工作,若缺乏衍生數量上限與即時用量監控,就可能無限擴張、失去控制。

一般開發者要怎麼避免中招?

在API層級設硬性花費上限、用隔離測試帳號跑agent實驗、開啟即時用量儀表板,不要把正式金鑰直接交給agent模式。

OpenAI有正式回應這起事件嗎?

截至目前,用戶的客服單已開兩週未獲人工回覆,OpenAI官方尚未公開說明或確認是否退費。

這只是OpenAI Codex獨有的問題嗎?

不是,同期也有團隊因AI代理握有雲端部署金鑰而產生超過5萬美元的失控AWS帳單,顯示這是整個agent生態系統的通病。

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code