AI蠕蟲評測:OpenAI證實提示注入會自我複製 | AI Worms Review: OpenAI Confirms Prompt Injection Spreads
By Kit 小克 | AI Tool Observer | 2026-10-09
🇹🇼 AI蠕蟲評測:OpenAI證實提示注入會自我複製
AI 蠕蟲不再只是研究論文裡的假設。2026 年 9 月 25 日,OpenAI 公開證實旗下模型存在一種新型態的提示注入(prompt injection)攻擊:惡意指令可以在 AI 代理人執行任務的同時,把自己複製進輸出內容裡,隨著 email、檔案、Slack 訊息一路傳給下一個代理人,行為模式就跟傳統電腦蠕蟲一樣。
AI 蠕蟲怎麼運作?
根據 OpenAI 內部紅隊系統 GPT-Red 在 2026 年 6 月 27 日的發現,一個合格的自我複製提示注入必須同時做到兩件事:第一,誘導 AI 做出原本沒被授權的動作;第二,把惡意指令的副本嵌進自己產生的輸出裡,等著感染下一個讀到這段內容的代理人。
研究團隊示範了三種傳播路徑:
- Email:中招的代理人把夾帶指令的內容寄給外部收件人
- 檔案系統:把惡意提示寫進其他代理人之後會讀取的文件
- Slack 整合:透過串接的訊息工作流程多跳傳播
測試對象包含 GPT-5.4-mini 研究版 checkpoint,以及跑在 Codex 環境裡的 GPT-5.5,顯示這不是單一模型的個案,而是當前前沿模型都要面對的攻擊類型。
目前還只是「實驗室等級」的風險
必須說清楚:所有示範都發生在模擬的工具呼叫與訓練評估環境裡,OpenAI 明確表示「沒有觀察到對外部造成影響,也沒有真實世界的事故」。換句話說,這是一次主動揭露的風險評估,不是抓到野外攻擊事件。同類概念其實在 2025 年的 Morris II 蠕蟲研究中就出現過,這次等於是 OpenAI 用自家系統把它重新驗證並升級。
OpenAI 打算怎麼防?
目前的主要對策,是把「自我複製」當成一個明確的訓練目標,整合進 GPT-Red 的對抗式訓練流程,讓未來的模型在部署前就先被訓練辨識並拒絕這類指令鏈。對於正在接 AI 代理人到 email、檔案系統、IM 工具的團隊,實務上該做的事其實很單純:
- 別讓代理人的輸出不經檢查就自動餵給下一個代理人
- 外部來源(郵件內文、文件內容)要當成不可信輸入處理,不要直接當指令執行
- 多代理人串接的流程,中間要有人或系統做內容過濾,而不是全自動 end-to-end
這波提示注入揭露提醒所有在做 agentic workflow 的人:代理人之間互相讀寫的那條管線,本身就是攻擊面。
好不好用,試了才知道。
🇺🇸 AI Worms Review: OpenAI Confirms Prompt Injection Spreads
AI worms just moved from theoretical papers to an official vendor disclosure. On September 25, 2026, OpenAI confirmed a new class of self-replicating prompt injection attacks: malicious instructions that get an AI agent to perform an unauthorized action while simultaneously copying themselves into the agent's output — spreading via email, files, and chat tools to the next agent downstream, just like a classic computer worm.
How the AI Worm Actually Spreads
OpenAI's internal red-teaming system, GPT-Red, first flagged the behavior on June 27, 2026. For a prompt injection to qualify as self-replicating, it needs to do two things at once: trick the agent into an unintended action, and embed a copy of the malicious instruction in whatever the agent generates next — waiting for the next agent that reads it.
Researchers demonstrated three propagation vectors:
- Email — a compromised agent forwards the embedded payload to external recipients
- Filesystem writes — the injection gets saved into documents other agents later read
- Slack integrations — multi-hop spread across connected messaging workflows
Testing covered GPT-5.4-mini research checkpoints and GPT-5.5 running inside a Codex harness — meaning this isn't a quirk of one model, but a vulnerability class across current frontier systems.
Still a Lab-Grade Risk, Not a Live Incident
To be clear: every demonstration ran in simulated tool-call environments during training and evaluation. OpenAI explicitly stated there was "no observed external impact or real-world incident." This is a proactive risk disclosure, not evidence of an attack in the wild. The underlying idea isn't brand new either — the 2025 Morris II worm research explored similar territory; OpenAI's GPT-Red effort essentially re-validated and escalated it internally.
What OpenAI — and You — Should Do About It
OpenAI's main mitigation is folding self-replication objectives directly into GPT-Red's adversarial training, so future models are hardened against this attack class before release. If you're already wiring agents into email, filesystems, or chat tools, the practical takeaways are simple:
- Never pipe one agent's raw output straight into another agent as trusted instructions
- Treat external content (email bodies, file contents) as untrusted input, not executable commands
- Put a filtering checkpoint — human or automated — between hops in any multi-agent pipeline instead of running it fully end-to-end
The bigger takeaway from this prompt injection disclosure: in any agentic workflow, the channel where agents read and write to each other IS the attack surface.
好不好用,試了才知道。
Sources / 資料來源
- OpenAI Demonstrates Self-Replicating Prompt-Injection Worms (Mallory.ai)
- OpenAI confirms existence of self-replicating prompt injections (Crypto Briefing)
- The Promptware Kill Chain (arXiv)
延伸閱讀 / Related Articles
- Pi 1.0評測:硬剛MCP一年的AI代理終於妥協 | Pi 1.0 Review: The Coding Agent That Hated MCP Ships It
- OpenAI 722篇數學論文評測:菲爾茲獎得主集體開砲 | OpenAI 722 Math Papers Review: Fields Medalists Revolt
- Claude Haiku 5.5評測:降價90%,但更燒token | Claude Haiku 5.5 Review: 90% Cheaper, Burns More Tokens
AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends
留言
張貼留言