Google's AMIE (Video) matches or beats PCPs on a 100-case OSCE, 91% vs 77% top-1 diagnosis

Towards Expert-level Medical AI for Real-time Video Consultations

Mahvish Nagda, Jihyeon Lee, Matthew Thompson, Chunjong Park, Tim Strother, Valentin Liévin, Roma Ruparel, Akshay Goel, Teya Bergamaschi, Suhana Bedi, Meet Shah, Pavel Dubov, Liviu Panait, Toshiyuki Fukuzawa, Sam Schmidgall, Craig Schiff, Joseph Xu, Aliya Rysbek, Yana Lunts, Jan Freyberg, Rebecca Hemengway, Sunny Virmani, David Racz, Carey Radebaugh, Joëlle Barral, Kavi Goel, Dale R. Webster, Katherine Chou, Avinatan Hassidim, Yossi Matias, James Manyika, Gregory Wayne, Tao Tu, Yun Liu, Ethan Goh, Christina Chen, Ryutaro Tanno, Po-Hsuan Cameron Chen, Mike Schaekermann, Anil Palepu

cs.AI, cs.CL, cs.CV

2026-08-11

A Gemini-based three-agent system from Google conducts real-time video consultations, matching or beating primary care physicians on a 100-case OSCE (91% vs 77% top-1 diagnosis), validated only on actors so far.

What problem this solves

Real consultations are not text. A doctor asks questions while reading the patient's face, breathing, and posture, catching a wheeze or tremor tucked into the speech. Text-based medical AI throws that perceptual layer away, which hurts most for patients who cannot articulate symptoms in writing.

Google's 2024 work (Shah et al.) showed an audio-visual medical AI was feasible across 20 scenarios, but not at clinician level. The hard part is structural: a real-time video consult demands three things at once that pull against each other, near-instant turn-taking, thorough clinical reasoning, and live audio-visual perception. Deepen the reasoning and latency climbs and the conversation stalls; cut latency and reasoning stays shallow.

Method

AMIE (Video) is a three-agent asynchronous architecture on Gemini 3 Flash and Gemini 3.1 Pro. The core idea is to split fast from slow:

The three agents run asynchronously. Talker handles fluent fast responses while Planner and Perception compute in the background and feed conclusions back. This cut mean turn latency from 21.4s in the sequential architecture to 2.6s, close to natural conversation. The team also built a taxonomy and automated evaluation of clinical audio-visual cues to score the system during development.

Results

A randomized OSCE (Objective Structured Clinical Examination) with 100 scenarios, 15 professional patient actors, 10 US board-certified PCPs, and 20 evaluating physicians. PCPs in the control arm were asked to turn their cameras off, to remove the bias of being visually present.

Diagnostic accuracy: AMIE (Video) matched the reference diagnosis in its top item 91% of the time, versus 77% for PCPs (p=0.039). Widening to the top three, the gap closed to 98% vs 90% (not significant).

Across six case-specific rubric domains, AMIE matched or led on every one:

DomainAMIE (Video)PCPp
Overall83%68%1.3e-9
Clinical reasoning90%76%2.0e-5
Perception and exam74%47%1.7e-9
Communication85%68%6.6e-7
History-taking86%78%1.2e-3
Treatment planning79%67%8.9e-5

Video paid off most in perception and examination: AMIE scored 77% vs 51% on using live video to gather information, and 72% vs 39% on guiding patients through physical exams. Mean consult duration was 8.94 minutes for AMIE versus 9.30 for PCPs.

Patient preference told a different story. Actors preferred AMIE for assessing and explaining conditions (both p=0.04); they leaned toward PCPs for rapport and partnership, though not significantly. Evaluating physicians actually rated AMIE higher on empathy (82% vs 71%) and relationship-building (80% vs 69%). Peers watching the tape found AMIE warmer; patients themselves found the human warmer.

In the modality ablation, video beat text chat across the board: communicating concerns 89% vs 79%, convenience 88% vs 71%, feeling understood 90% vs 81%.

Why it matters

This is the first time an audio-visual medical AI reached, and on most metrics exceeded, practicing physicians on a structured exam. Three takeaways for practitioners. The decoupled multi-agent recipe is validated: splitting latency-sensitive dialogue from compute-sensitive reasoning works in a domain that is both time-critical and cannot stay shallow. Audio-visual perception buys real diagnostic lift that text cannot, especially in exam steps that require watching the body. And it draws a clear boundary: this is an exam machine, not a machine you can deploy.

Limitations

The authors list many; the serious ones follow. The system is discrete turn-taking and cannot do the overlapping speech and instant interjections humans use. High-frequency or subtle motion is hard: nystagmus at 3%, tremors at 24%, affect and expressiveness at 29%. Sometimes it sees a cue but fails to act on it, noticing poor audio quality without asking the patient to fix it.

The study's ecological validity is limited. Scenarios were restricted to actable presentations; dermatology that actors cannot simulate was excluded. Patient actors are not real patients, and geriatric cases were under-recruited. Blinding is imperfect: Talker's synthetic voice is audibly non-human, patients behave differently knowing they face an AI, and PCPs had cameras off.

More fundamentally, an OSCE measures single-encounter performance on structured cases, not the longitudinal care and shared decision-making of a real practice. The authors state plainly the system is not ready for clinical deployment and needs extensive validation in real workflows. The study was funded by Alphabet; most authors are Google employees and shareholders.

Terms

Source

What people are saying

Related papers

All paper explainers