Topic: sparse decoding
-
Google’s Gemma 4 AI triples speed by predicting future tokens
Google's Gemma 4 open models, optimized for local deployment, now feature experimental Multi-Token Prediction (MTP) drafters that use speculative decoding to guess future tokens, significantly reducing generation times compared to standard autoregressive methods. The local focus of Gemma 4 provid...
Read More »