The Hallway Track
Research Findings

DeepMind Just Changed How AI Sees The World

Two Minute Papers · Aug 07, 2026 · Research Findings

DeepMind's Gemma 4 achieves multimodal vision by patching images directly into the main transformer, eliminating separate encoders.

“Throw that all away. Out. Right now.”

DeepMind revealed the architecture behind Gemma 4's multimodal capabilities: rather than chaining separate vision and audio encoders, it slices images into patches and audio into 40ms chunks and feeds them directly as tokens into a single transformer. This unified approach enables a 12B parameter model to handle vision, audio, and language without the overhead of dedicated encoder networks. With over 300 million downloads, Gemma 4 represents a meaningful efficiency signal for the industry's push toward capable local models.

DeepMind Gemma 4 multimodal vision small models architecture on-device AI

Watch / read the original source →