跳到主要內容

Qwen-Audio 3.1評測:語音API砍價95%,開發者該換嗎 | Qwen-Audio 3.1 Review: Voice API Prices Cut Up to 95%

By Kit 小克 | AI Tool Observer | 2026-09-28

🇹🇼 Qwen-Audio 3.1評測:語音API砍價95%,開發者該換嗎

Alibaba 旗下 Qwen 團隊在 2026 年 9 月 23 日發布 Qwen-Audio 3.1,一口氣端出五個語音模型,同時把 API 價格砍到見骨:TTS 降價約 70%、Realtime 降價約 85%、ASR 語音辨識最多砍到 95%。對於正在評估要不要把語音功能塞進產品的開發者來說,這是近期最值得關注的動態之一。

Qwen-Audio 3.1 到底砍了多少

根據阿里雲 Model Studio 公開的定價,以中國(北京)節點為例,新推出的 TTS-Next 模型輸入價格是每百萬 token 0.848 美元、輸出 1.696 美元,比起舊版動輒好幾美元的報價便宜非常多。三個既有模型(ASR、TTS、Realtime)全面升級,另外新增兩個模型:

  • ASR-Next:不只轉文字,還能辨識同一段錄音裡的多個講者、標出時間戳記,甚至偵測情緒語氣、環境音、機械噪音
  • TTS-Next:能生成完整的「音訊場景」,不只是單一語音,而是包含背景音效的完整聲音內容

Realtime 模型主打「邊聽邊說」的全雙工對話,支援使用者中途打斷,技術報告顯示多語言基準測試與語音條件下的工具呼叫能力都比前一代進步。

誰該認真考慮換過去

如果你的產品有客服語音機器人、多語言逐字稿、無障礙輔助字幕、或是需要即時語音對話介面,這波降價幅度大到值得重新算一次帳。原本因為語音 API 太貴而不敢做的功能,現在成本門檻低很多,可以拿真實流量測一次再決定要不要全面切換。

老實話:降價不等於馬上換

降價新聞很吸睛,但幾件事還是要注意:目前公開的具體定價集中在中國節點,其他地區與其餘四個模型的價格要自己去 DashScope Model Studio 查證;語音辨識與生成的實際品質(口音、專有名詞、多講者場景)常常和跑分數字有落差,一定要拿自己的真實語料測;如果你的產品已經深度綁定 OpenAI Realtime API 或 ElevenLabs,換供應商還牽涉到延遲、地區合規、以及重新調整 prompt 與後處理邏輯的隱藏成本。便宜是真的便宜,但「換」跟「便宜」是兩件事。

好不好用,試了才知道。


🇺🇸 Qwen-Audio 3.1 Review: Voice API Prices Cut Up to 95%

Qwen-Audio 3.1 is Alibaba's biggest voice AI move of the quarter: on September 23, 2026, its Qwen team shipped five speech models while slashing API prices across the board -- TTS down roughly 70%, Realtime down about 85%, and ASR cut by as much as 95%. For any developer weighing whether to add voice features to a product, this price cut changes the math.

How Much Cheaper Is Qwen-Audio 3.1

Per Alibaba Cloud's published Model Studio pricing (China/Beijing region), the new TTS-Next model runs $0.848 per million input tokens and $1.696 per million output tokens -- a fraction of what comparable voice APIs typically charge. Three existing models (ASR, TTS, Realtime) got upgraded, and two new ones joined the lineup:

  • ASR-Next -- transcribes speech, separates multiple speakers with timestamps, and detects emotional tone, background ambience, and mechanical noise
  • TTS-Next -- generates full audio "scenes," not just a single voice track, but complete soundscapes

The Realtime model focuses on full-duplex conversation: it can listen and speak simultaneously and supports mid-sentence interruptions. Alibaba's technical report claims gains in multilingual benchmarks and speech-conditioned tool calling over the previous generation.

Who Should Actually Care

If you are building voice customer support bots, multilingual transcription, accessibility captions, or real-time voice interfaces, this price drop is large enough to justify re-running your cost model. Features you shelved because voice APIs were too expensive may now clear the bar -- worth testing against real traffic before committing.

The Honest Take: Cheap Does Not Mean You Should Switch Today

A few caveats before you migrate. The published pricing so far is concentrated on the China region -- check DashScope Model Studio directly for other regions and the remaining models. Benchmark scores rarely match real-world quality on accents, proper nouns, or multi-speaker audio, so test with your own data first. And if you are already deeply integrated with OpenAI's Realtime API or ElevenLabs, switching brings hidden costs: latency differences, regional compliance, and reworking prompts and post-processing logic. Cheaper is real, but "cheaper" and "worth switching" are two different questions.

You won't know until you try it.

Sources / 資料來源

延伸閱讀 / Related Articles


AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends

留言

這個網誌中的熱門文章

Google Ironwood TPU v7 推理專用晶片解析:效能追平 NVIDIA、成本低 44%,AI 晶片戰爭正式開打 | Google Ironwood TPU v7 Explained: Matching NVIDIA Performance at 44% Lower Cost — The AI Chip War Heats Up

Claude Code 實測:AI 幫你寫程式到底行不行? | Claude Code Review: Can AI Really Code for You?

Cursor vs GitHub Copilot vs Claude Code:AI 程式助手大比拼 | AI Coding Assistants Compared: Cursor vs GitHub Copilot vs Claude Code