Topic: visual reasoning
-
Gemini 3 Flash's Agentic Vision: Sharper Image Responses
Agentic Vision transforms Gemini 3 Flash's image analysis by using a "Think, Act, Observe" loop, where the model actively manipulates images with Python code to uncover fine details and ensure grounded answers. This approach replaces probabilistic guessing with verifiable execution, improving acc...
Read More » -
Google AI Helps Boston Dynamics Robot Dog Read Gauges
Boston Dynamics' Spot robot can now autonomously read analog instruments like pressure gauges and thermometers in factories, powered by a new AI model from Google DeepMind. The core technology is the Gemini Robotics-ER 1.6 model, which provides complex visual reasoning to accurately interpret int...
Read More » -
Anthropic's New Opus 4.5: More Power, Lower Cost
Anthropic has launched Opus 4.5, its new flagship model, with enhanced coding capabilities and user experience, strengthening its position against competitors like OpenAI. The model introduces intelligent context management by summarizing earlier conversation segments, ensuring smoother and more ...
Read More »