Google Releases High-Precision Speech-to-Text Model, Outperforming OpenAI’s Equivalent Products

微软Azure增速超越AWS和谷歌云,分析师揭秘原因
Published on: Aug 27, 2026
Author: Amy Liu

Google (GOOGL) has launched a high-precision speech-to-text model, Gemini 3.5 Transcribe, designed to address the shortcomings of traditional models in handling noise and specialized terminology. It offers both real-time and offline processing modes, with performance metrics that surpass competing products. The model is now available to developers and select users, but the delayed release of Google’s flagship model, Gemini 3.5 Pro, continues to raise market concerns about the company’s AI competitiveness and its ability to translate investment into returns.

The model is primarily aimed at developers and enterprise users, with the goal of seamless integration into workflows to support tasks such as voice agents, real-time captioning, and post-call analysis. Developers can leverage the Live API for real-time streaming processing or the Interactions API for processing pre-recorded audio. On the functional front, the model supports intelligent transcription, automatically cleaning up text content—for instance, by removing verbal fillers. It also features custom vocabulary recognition, covering over 85 languages, and can identify up to three speakers.

Performance Metrics and Availability

In terms of performance, Gemini 3.5 Transcribe achieves a word error rate of 4% in streaming scenarios and 2.6% in non-streaming scenarios. On the FLEURS benchmark, which encompasses multiple languages, the model similarly demonstrates its edge—Gemini Transcribe 3.5 Live posts a word error rate of 5.5%, compared to OpenAI’s GPT Live Transcribe at 8.9%.

As of its release date, the new model is available in public preview for developers and enterprise users. For general users, the model is currently accessible in English via the Gemini application on macOS, and is also offered in the Android version of Rambler in select countries. Google has additionally stated that the model will soon arrive on the Chrome browser.

Flagship Model Still Awaits Debut

Earlier this month, Google released Gemini 3.7 Flash, and the number of Gemini application users surpassed 1 billion. However, the much-anticipated next-generation flagship model, Gemini 3.5 Pro, has yet to go live. At the 2026 I/O Developer Conference held in May, Pichai indicated that the model would be launched in June, but that release window has now passed, and the product has been delayed without a confirmed date. Under the original roadmap, Gemini 3.5 Pro was expected to spearhead Google’s renewed push to the forefront of AI innovation. Against the backdrop of OpenAI and Anthropic continuously advancing their model capabilities, the Gemini Pro series has long been an important benchmark for external observers to gauge Google’s AI strength. The ongoing uncertainty surrounding the flagship model’s release date has further intensified widespread questions about whether the tech giant can outpace its competitors and successfully convert its massive AI investments into market-leading tools and services.

AI Financial Service Fintech Technology