Gemini 3.5 Transcribe評測:Google新語音轉文字模型登場 | Gemini 3.5 Transcribe Review: Google's Speech-to-Text Model
By Kit 小克 | AI Tool Observer | 2026-08-30
🇹🇼 Gemini 3.5 Transcribe評測:Google新語音轉文字模型登場
Google在2026年8月26日發布Gemini 3.5 Transcribe,這是目前最精準的語音轉文字模型,公開預覽版已經開放給開發者與企業用戶測試。比起單純把聲音轉成文字,Gemini 3.5 Transcribe更像是一個懂語境的助理:能自動去除贅字、修正話說到一半的更正、辨識最多三位講者,還支援超過85種語言即時切換。對常常需要開會記錄、訪談逐字稿、多語言客服的人來說,這次更新踩到了真正的痛點。
什麼是Gemini 3.5 Transcribe?
跟傳統語音辨識模型不同,Gemini 3.5 Transcribe不是「聽到什麼打什麼」,而是把原始音訊直接轉成已經整理過、可以直接用的文字。官方說明它能處理背景噪音、專業術語、口語中的自我更正,並自動加上標點與段落格式。它同時支援即時串流(gemini-3.5-transcribe-live)與預錄音訊處理(gemini-3.5-transcribe)兩種API模式,開發者可以依場景選擇。
效能數字:字錯率降到多低?
Google公布的內部測試顯示,串流模式的字錯率(Word Error Rate, WER)為4.0%,非串流模式更低至2.6%;在多語言的FLEURS基準測試中,串流與非串流的WER分別為5.50%與5.04%。相較前代Chirp 3模型,取得最終逐字稿的時間縮短了70%,這對需要即時字幕或客服語音轉譯的應用來說是實際可感受的改善。
支援85種語言、辨識三人對話
Gemini 3.5 Transcribe支援超過85種語言與腔調,能自動偵測語言並在對話中即時切換(例如中英夾雜的會議)。它也能替最多三位講者分別標記時間軸與身分,加上自訂詞彙功能,方便處理公司內部術語或罕見專有名詞。
哪裡可以用到?
這個模型已經悄悄進駐多項Google產品:Android上Gboard的Rambler語音輸入功能、macOS版Gemini App、開發工具Google Antigravity,官方也預告即將整合進Chrome瀏覽器。開發者可透過Gemini API與Google AI Studio申請公開預覽存取,企業用戶則可透過Gemini Enterprise Agent Platform使用。
小克怎麼看
公開預覽階段還沒有正式定價,功能穩定度跟配額都可能隨時調整,這是實測前要有的心理準備。但單看字錯率數字,Gemini 3.5 Transcribe已經站上業界前段班,跟OpenAI Whisper系列、AssemblyAI等對手的差距正在縮小甚至反超。如果你的產品需要語音轉文字,這波更新值得排進評估清單——尤其是多語言、多講者場景。
好不好用,試了才知道。
🇺🇸 Gemini 3.5 Transcribe Review: Google's Speech-to-Text Model
Gemini 3.5 Transcribe is Google's newest speech-to-text model, released in public preview on August 26, 2026. It's built to do more than turn audio into text — it cleans up filler words, catches self-corrections mid-sentence, identifies up to three speakers, and automatically detects and switches between over 85 languages. For anyone who deals with meeting notes, interview transcripts, or multilingual customer support, this is the update that actually targets real pain points.
What Is Gemini 3.5 Transcribe?
Unlike a plain transcription engine that types exactly what it hears, Gemini 3.5 Transcribe converts raw audio directly into polished, ready-to-use text. According to Google, it handles background noise, domain-specific jargon, and spoken self-corrections, then adds punctuation and formatting automatically. It ships as two API modes: a real-time streaming endpoint (gemini-3.5-transcribe-live) and a pre-recorded audio endpoint (gemini-3.5-transcribe).
The Benchmark Numbers
Google reports a Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming transcription. On the multilingual FLEURS benchmark, WER comes in at 5.50% (streaming) and 5.04% (non-streaming). Compared to its predecessor Chirp 3, time-to-final-transcription is 70% faster — a real difference for live captioning or voice-driven support tools.
85+ Languages, Up to 3 Speakers
The model supports more than 85 languages and accents, with automatic language detection and live switching mid-conversation — useful for code-switching meetings. It can also attribute speech to up to three distinct speakers with timestamps, and accepts custom vocabulary for company jargon or uncommon proper nouns.
Where You Can Try It
Gemini 3.5 Transcribe is already quietly powering several Google products: the Rambler voice input feature in Gboard on Android, the Gemini app on macOS, and the Google Antigravity dev tool, with Chrome integration coming soon. Developers can access the public preview through the Gemini API and Google AI Studio; enterprise customers get it via the Gemini Enterprise Agent Platform.
Kit's Take
There's no official pricing yet — it's still public preview, so expect quotas and behavior to shift before general availability. But on raw WER numbers alone, Gemini 3.5 Transcribe is now competitive with, and in some benchmarks ahead of, players like OpenAI's Whisper family and AssemblyAI. If speech-to-text is anywhere in your product roadmap, this is worth testing now, especially for multilingual or multi-speaker use cases.
好不好用,試了才知道。
Sources / 資料來源
- Google Blog: Intelligent transcription with Gemini 3.5 Transcribe
- 9to5Google: Google launches Gemini 3.5 Transcribe
- Android Authority: Google rolls out Gemini 3.5 Transcribe
延伸閱讀 / Related Articles
- Open Executive評測:開源AI CEO,8位AI高管上工 | Open Executive Review: Open-Source AI CEO Goes to Work
- AI駭客攻擊評測:伊朗駭客用AI瞄準美國水廠 | AI Hacking Review: Iran Hackers Hit US Water Utilities
- Grok漏洞評測:加密提示詞注入零點擊偷走對話 | Grok Vulnerability Review: Zero-Click Prompt Injection
AI 工具觀察站 — 每日精選 AI Agent 與工具趨勢
AI Tool Observer — Daily curated AI Agent & tool trends
留言
張貼留言