AI & TechArtificial IntelligenceBigTech CompaniesNewswireTechnology

Google Adds Video Understanding to YouTube and Google Assistant

Originally published on: September 2, 2026
▼ Summary

– Google is introducing agentic video understanding to Ask YouTube, enabling Gemini to analyze specific video segments rather than sampling the entire content at a fixed rate.
– This new dynamic processing method allows the model to selectively load parts of a video and adjust frame rates, improving accuracy while significantly reducing token usage and costs.
– The feature is already available to developers via the Gemini API and will be rolled out to the watch page for general users in the coming months.
– Agentic video understanding offers benefits such as locating split-second moments and counting repeated actions, with testing showing up to an 88% reduction in token usage.
– While the watch-page feature has seen over 140 million user engagements, it remains unclear if the search-based version of Ask YouTube will receive the same processing update.

Google is significantly upgrading its video analysis capabilities for both YouTube and Google Assistant, introducing a more sophisticated method for processing visual content. The company is rolling out what it calls agentic video understanding to the Ask YouTube feature, allowing the Gemini model to examine specific segments of a video rather than relying on a fixed-rate sampling of the entire file. This shift aims to provide users with answers that are directly grounded in the visuals displayed on screen.

While this advanced functionality is already available to developers through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, consumer access via the video watch page is scheduled for release in the coming months. The update supports both uploaded videos and standard YouTube content, marking a substantial step forward in how AI interacts with long-form media.

Enhanced Processing Capabilities

The new system represents a departure from static processing, which has been Gemini’s default video mode. Previously, the model captured one frame per second and processed the full video regardless of content density. Developers could adjust the frame rate, but the computational load remained high because every segment was analyzed equally. In contrast, agentic video understanding employs a dynamic loop where the model selectively loads parts of the video, adjusts frame rates, and determines whether to incorporate frames, audio, or transcripts for specific moments.

This targeted approach allows the AI to locate split-second events, search through multi-hour recordings, identify visual glitches, and count repeated actions or objects with greater precision. According to Google’s internal testing, this method reduces token usage by up to 88% and lowers analysis costs by up to 66%. Additionally, accuracy improvements of up to 7% were observed compared to static processing. These efficiency gains are most pronounced in longer videos, although the documentation notes that response initiation may be slightly delayed for clips shorter than five minutes.

Integration with Watch-Page Features

On the YouTube interface, the Ask YouTube feature appears as a button located below the video player, enabling viewers to query content while watching. This function operates independently from the version accessible via the main search bar, which provides summaries with cited sources. As noted in previous reports, the search-based version was initially rolled out as an experiment for a limited group of U. S. users searching in English before expanding further.

During Alphabet’s Q2 earnings call in July, CEO Sundar Pichai highlighted that over 140 million users engaged with the watch-page feature in June alone. He emphasized that Google intends to bring this interactive Ask experience to the broader search ecosystem on YouTube. However, the recent announcement does not specify whether the search-bar version will receive the same agentic video understanding treatment, leaving some ambiguity about how citations and rankings are determined across different interfaces.

Implications for Creators and Users

The rollout timeline places the new video mode in the Gemini app “soon” and on the YouTube watch page within the next few months, though exact dates remain undefined. For creators, the lack of clarity regarding how videos are selected as primary citations versus supporting sources remains a point of interest. While YouTube’s help pages state that its ranking system prioritizes relevance, engagement, and quality, there is still no official guidance on adapting content for these advanced AI systems.

Users should monitor the watch-page help section for updates as the feature launches. Currently, the help page indicates that responses are derived from YouTube and web sources but does not detail the underlying mechanics of video analysis. As Google continues to refine its AI tools, the ability to extract precise information from visual media without scanning every frame could set a new standard for digital content interaction.

(Source: Search Engine Journal)

Topics

ai video analysis 98% gemini technology 95% ask youtube feature 92% performance optimization 88% developer access 85%
Show More