AI & TechArtificial IntelligenceBigTech CompaniesDigital PublishingNewswireTechnology

Google AI Removes Filler Words From Transcriptions

▼ Summary

– Google has updated its Gemini Audio suite with new models including Gemini 3.5 Live, Live Experimental, and Transcribe to enhance voice-controlled AI precision.
– The new Gemini 3.5 Transcribe model improves upon Chirp 3 by detecting specialized jargon, supporting over 85 languages, and allowing natural voice editing.
– Key features of the transcription tool include automatic formatting, filler word removal, customized vocabulary adaptation, and multi-speaker attribution with timestamps.
– Live models offer improved handling of interruptions and visual processing, while the experimental version narrates its reasoning process in real time.
– These updates are currently rolling out for macOS users, select Android devices, and developers via public preview, with Chrome support planned soon.

Google has significantly upgraded its Gemini Audio capabilities by introducing a suite of new Gemini 3.5 models. These updates bring advanced transcription features designed to automatically identify specialized jargon and support over 85 languages. The new lineup includes Gemini 3.5 Live, Gemini 3.5 Live Experimental, and the newly announced Gemini 3.5 Transcribe. These tools aim to enhance the precision of Google’s voice-controlled AI, ensuring high accuracy even in noisy environments or when speech is interrupted.

The introduction of Gemini 3.5 Transcribe marks a significant departure from previous iterations. While users continue to await the release of the promised Gemini 3.5 Pro model, this new tool offers immediate improvements. Google states that the new system “represents a major advancement from our previous transcription model, Chirp 3,” particularly highlighting reductions in wording errors and enhanced multilingual performance. The model is engineered to handle complex audio inputs with greater reliability than its predecessors.

A standout feature of the new transcriber is its ability to refine output through natural voice commands. Users can “edit naturally with just your voice,” allowing for seamless adjustments without manual typing. Beyond basic editing, the system automatically formats text and strips out common filler words such as “um” and “uh.” To ensure professional-grade results, users can upload a customized vocabulary list. This allows the model to recognize unique spellings and industry-specific terminology, preventing these terms from being flagged as errors or requiring manual correction. Additionally, the tool supports speaker attribution for up to three individuals in pre-recorded files and provides precise word-level timestamps.

Alongside the standalone transcriber, Google is releasing Gemini 3.5 Live and Gemini 3.5 Live Experimental. These applications build upon the existing speech recognition infrastructure that drives Gemini’s interactive voice chat. Gemini 3.5 Live demonstrates improved resilience against mid-sentence interruptions, faster language detection, and enhanced real-time visual processing. The experimental variant takes this further by narrating its own reasoning process step-by-step while tackling complex tasks, offering transparency into how the AI approaches problem-solving.

These enhancements are beginning their rollout today. Initially, the updates are available in English for all macOS Gemini app users. On the mobile front, the Rambler dictation feature on Android is launching in select countries and languages. Developers can already access these capabilities via public preview in the Gemini API through AI Studio and Antigravity. Google has indicated that support for Chrome browsers will follow shortly, expanding the utility of these transcription tools across more platforms.

(Source: The Verge)

Topics

ai model updates 95% transcription technology 90% voice interaction 85% platform availability 80% Multilingual Support 75%