跳到主要內容

FLUX 3 評測:一個模型統包影片、圖片、音效與機器人 | FLUX 3 Review: One Model for Video, Image, Audio and Robots

By Kit 小克 | AI Tool Observer | 2026-07-26

🇹🇼 FLUX 3 評測:一個模型統包影片、圖片、音效與機器人

FLUX 3 是德國新創 Black Forest Labs 在 2026 年 7 月 23 日推出的多模態 AI 模型,最大賣點是一個網路同時生成影片、圖片、音效,甚至機器人動作指令,不是把好幾個獨立模型拼裝在一起假裝統一介面。這代表模型在訓練時就同時學習空間結構、動作、聲音與物理互動之間的關聯,而不是分開處理。

FLUX 3 影片生成:20 秒內建原生音效

目前主打功能是影片:單次可生成最長 20 秒的影片,並內建原生音效(對白、音效、配樂),支援文字轉影片、圖片轉影片、影片轉影片、關鍵影格控制,以及多鏡頭串接。畫面比例涵蓋 9:16 到 21:9,涵蓋直式短影音到電影寬螢幕。不過現階段解析度上限只有 720p,1080p 要等後續才會開放。

機器人動作預測:從 Audi 產線開始試水溫

比較少見的是 FLUX 3 也支援「動作預測」,也就是把同一套模型當成機器人的視覺-動作骨幹。第一個合作夥伴是機器人新創 mimic robotics,已經在 Audi 產線上測試用 FLUX 3 驅動的影片-動作模型,這代表生成式 AI 的應用範圍正從純內容創作延伸到實體世界操作。

能不能用?老實說的限制

  • 影片功能:目前僅開放搶先體驗,需要排隊申請
  • 圖片生成:官方說「未來幾週」才會開放,現在還用不到
  • 開源版本:FLUX 3 Dev 開源骨幹已預告,但沒有明確時間表
  • 解析度:720p 起跳,跟 Sora 2、Veo 3 的商用版相比仍有落差

FLUX 3 對上 Sora、Veo:差異在哪

Sora 和 Google Veo 走的是「影片為主、音訊為輔」的路線,FLUX 3 的野心更大——把圖片、影片、音效、機器人動作全部塞進同一套架構,長期目標比較像是打造通用的「世界模型」骨幹,而非單純的影片生成工具。這個方向如果成功,未來機器人公司、遊戲工作室、廣告代理商可能都會用同一套 API,而不是各自串接不同廠商的模型。但現階段搶先體驗的排隊機制、720p 上限,代表距離普及還有一段距離。

對內容創作者來說,FLUX 3 現在還不是能立刻上手的生產力工具,比較像是提前卡位下一代多模態 AI 的入場券。想知道排隊多久能拿到體驗資格、生成品質是否真的比得上 Sora,好不好用,試了才知道。


🇺🇸 FLUX 3 Review: One Model for Video, Image, Audio and Robots

FLUX 3, launched by German startup Black Forest Labs on July 23, 2026, is a multimodal AI model built to generate video, images, audio, and even robot action commands from a single unified network - not several separate models stitched behind one interface. That means the model was jointly trained to understand spatial structure, motion, sound, and physical interaction together, rather than treating each as an isolated task.

FLUX 3 Video: 20-Second Clips With Native Audio

Video is the flagship feature at launch: each generation can run up to 20 seconds with native audio baked in - dialogue, sound effects, and music. It supports text-to-video, image-to-video, video-to-video, keyframe control, and multi-shot chaining, across aspect ratios from 9:16 to 21:9. The catch: resolution is capped at 720p for now, with 1080p expected to follow later.

Robot Action Prediction: Testing on Audi's Line

A less common feature is action-prediction - using the same backbone as a vision-action model for robotics. The first partner is robotics startup mimic robotics, which is already testing a FLUX 3-powered video-action model on Audi's production line. That signals generative AI moving beyond content creation into physical-world manipulation.

Can You Actually Use It? The Honest Limits

  • Video: gated early access only - you need to apply and wait in line
  • Image generation: officially planned for the following weeks, not usable yet
  • Open weights: a FLUX 3 Dev open-weight backbone is planned, but no firm date
  • Resolution: starts at 720p, still behind commercial Sora 2 and Veo 3 output

FLUX 3 vs Sora and Veo: What's Different

Sora and Google's Veo are video-first tools with audio as a secondary layer. FLUX 3's ambition is broader - cramming image, video, audio, and robot action into one architecture, aiming more at a general-purpose "world model" backbone than a pure video generator. If it works, robotics companies, game studios, and ad agencies could eventually run on one shared API instead of stitching together separate vendor models. But the current early-access queue and 720p ceiling mean mainstream usability is still some distance away.

For content creators, FLUX 3 isn't a ready-to-use production tool yet - it's more of an early ticket into the next generation of multimodal AI. Whether the wait is worth it and the output actually rivals Sora: 好不好用,試了才知道 (you won't know until you try it).

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Stanford 研究登上《Science》:11 個 AI 模型有 47% 機率說你對,即使你錯了 | Stanford Study in Science: AI Models Validate Harmful Behavior 47% of the Time — Sycophancy Is a Real Problem