跳到主要內容

Nemotron Ultra-CC評測:IOI程式賽奪金,首勝人類冠軍 | Nemotron Ultra-CC Review: First AI to Beat IOI Champion

By Kit 小克 | AI Tool Observer | 2026-09-04

🇹🇼 Nemotron Ultra-CC評測:IOI程式賽奪金,首勝人類冠軍

Nemotron Ultra-CC 是Nvidia打造的程式競賽專用AI模型,2026年9月在IOI 2026(國際資訊奧林匹亞)賽場上,用550B參數規模拿下535.4分(滿分600),首度打敗人類最高分選手的498.27分,創下AI在完整IOI賽題組贏過人類金牌得主的紀錄。以下帶你看懂這個成績怎麼來的、用了什麼技巧,還有哪裡該保持懷疑。

Nemotron Ultra-CC是什麼?

Nemotron Ultra-CC屬於Nvidia Nemotron-3系列,是專門針對競技程式設計特化訓練的版本,總參數550B、啟用參數55B。同系列還有較小的Nano-CC(總參數30B、啟用3B),兩者都鎖定IOI這類規格明確、能自動評分的題型。

IOI 2026的比賽成績是怎麼拿到的?

Ultra-CC是在IOI 2026比賽期間「即時」作答,套用跟人類選手一模一樣的時間限制、網路存取規則與繳交方式,最終535.4分超越金牌門檻361.12分,也贏過人類最高分選手的498.27分。不過要注意,這場對戰不是IOI官方認證賽事——沒有IOI監督,是Nvidia自行安排的場外測試,分數不列入正式排名。

關鍵技術GenCorrect是什麼?

模型訓練流程分四步:用22,000題競程資料整理訓練集、以DeepSeek-V4-Flash產生的推理軌跡做監督式微調(SFT)、對Nano模型加做強化學習(RL),最後在推論階段套用GenCorrect——每輪最多生成200個候選解答、做多樣性分群後依評分回饋反覆修正,跑五輪。光靠GenCorrect這招測試時運算,就讓Nano-CC在沒有額外訓練下效能大增60%。

AI真的比人類會寫程式了嗎?

先別急著下結論。IOI題目規格明確、能自動客觀評分,是AI最擅長發揮的「理想賽道」,跟真實工程裡模糊需求、系統整合、長期維護的複雜度完全是兩回事。研究團隊自己也承認,Ultra-CC用掉的運算資源遠超人類選手(大量候選解答生成加多輪修正),這是「系統對系統」的比較,而非公平的同資源對決,獨立驗證與完整運算成本細節目前也還沒公開。

好不好用,試了才知道。


🇺🇸 Nemotron Ultra-CC Review: First AI to Beat IOI Champion

Nemotron Ultra-CC, Nvidia's competition-tuned AI model, scored 535.4 out of 600 at IOI 2026 (the International Olympiad in Informatics) in September 2026 — beating the top human contestant's 498.27 and marking the first time an AI system has outscored the highest-scoring human on a full IOI problem set. Here's what the 550-billion-parameter system actually did, and where to stay skeptical.

What Is Nemotron Ultra-CC?

Nemotron Ultra-CC is part of Nvidia's Nemotron-3 family, specialized for competitive programming with 550B total parameters and 55B active. A smaller sibling, Nano-CC (30B total, 3B active), was also tested. Both target IOI-style problems: unambiguous specs with objective, automated scoring.

How Did the IOI 2026 Score Happen?

Ultra-CC ran live during the actual IOI 2026 contest, under the same time limits, internet-access rules, and submission constraints as human contestants. It finished with 535.4 points, clearing the gold-medal threshold of 361.12 and beating the top human score of 498.27. Crucially, this wasn't an officially ranked IOI result — it ran unsupervised by IOI as an exhibition, so the score never entered the real leaderboard.

What Is GenCorrect?

The training pipeline has four stages: curating 22,000 competitive-programming problems, supervised fine-tuning on reasoning traces generated by DeepSeek-V4-Flash, reinforcement learning applied only to the Nano model, and finally GenCorrect at inference time — generating up to 200 candidate solutions per round, clustering them for diversity, and refining across five rounds using evaluator feedback. GenCorrect alone boosted Nano-CC's score by 60% with zero extra training.

Does This Mean AI Can Out-Code Humans Now?

Not so fast. IOI problems are an ideal shape for AI — clear specs, objective automated feedback — nothing like the ambiguity, integration work, and maintenance headaches of real-world engineering. The researchers themselves admit Ultra-CC burned far more compute than any human contestant, generating and re-checking hundreds of candidate solutions per problem. This is a system-vs-system comparison, not an equal-resource contest, and independent verification plus full compute-cost disclosure still hasn't happened.

好不好用,試了才知道 — try it yourself before you believe the headline.

Sources / 資料來源

常見問題 FAQ

Nemotron Ultra-CC真的贏過所有人類選手嗎?

只贏過分數最高的單一選手(498.27分),不是贏過全部參賽者,而且這場對戰沒有IOI官方監督,不算正式排名成績。

GenCorrect是什麼技術?

GenCorrect是一種測試時運算策略,每輪生成多達200個候選解答並做多樣性分群,再依評分回饋反覆修正、跑五輪,能大幅提升模型表現且不需要額外訓練。

這代表AI已經能取代軟體工程師了嗎?

不代表。IOI題目規格明確、有客觀評分標準,跟真實世界模糊需求、系統整合、長期維護的複雜度差很多,研究團隊自己也強調這只是系統層級的比較。

Nemotron Ultra-CC用了多少運算資源?

官方沒有完整公開,但研究團隊承認運算量遠超過人類選手,因為要生成並評估大量候選解答,屬於不對等資源的比較。

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code