DeepMind demos Gemma running fully in-browser: 2B model under 1GB quantized

AI Engineer · youtube · 2026-10-06

In an AI Engineer talk, Google DeepMind's Paige Bailey demoed a model running entirely in the browser (instant emoji rankings of Harry Potter books, no API call), introducing Gemma 4: Apache 2-licensed open models from 2B to 31B parameters, downloadable, adaptable, and finetunable. Google AI Edge Gallery brings local inference to Android/iOS with image description, audio transcription, function calling, and skill packs. Quantized checkpoints explain the browser speed — the 2B version is under 1GB — and the practical thread is matching model size to device, from single commodity GPU to laptop to mobile. She also covered DeepMind's research mission (AlphaFold, medical models, robotics, AI for science) and computer use / managed agents in a Linux sandbox.

Original post →

More from Models

Models channel →