In 100 real urgent-care chats, AMIE's top-3 hit 75%; PCPs won on cost and practicality

A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic

Peter Brodeur, Jacob M. Koshy, Anil Palepu, Khaled Saab, Ava Homiar, Roma Ruparel, Charles Wu, Ryutaro Tanno, Joseph Xu, Amy Wang, David Stutz, Wei-Hung Weng, Hannah M. Ferrera, David Barrett, Lindsey Crowley, Jihyeon Lee, Spencer E. Rittner, Ellery Wulczyn, Selena K. Zhang, Elahe Vedadi, Christine G. Kohn, Kavita Kulkarni, Vinay Kadiyala, Sara Mahdavi, Wendy Du, Jessica M. Williams, David Feinbloom, Renee Wong, Tao Tu, Petar Sirkovic, Alessio Orlandi, Christopher Semturs, Yun Liu, Juraj Gottweis, Dale R. Webster, Joëlle Barral, Katherine Chou, Pushmeet Kohli, Avinatan Hassidim, Yossi Matias, James Manyika, Rob Fields, Jonathan X. Li, Marc L. Cohen, Vivek Natarajan, Mike Schaekermann, Alan Karthikesalingam, Adam Rodman

cs.HC, cs.AI, cs.CL, cs.LG

2026-03-09

100 urgent-care patients chatted with AMIE before the visit: no safety stops, top-3 exact-or-close accuracy 75%, and PCP plans won on practicality and cost.

What problem this solves

AMIE, the Articulate Medical Intelligence Explorer, had already matched primary care physicians on diagnostic reasoning against standardized patients, and done better on some items. Those encounters are scripted. Real urgent-care patients are not. Complaints wander, people get anxious, typing is slow, and health literacy varies. Other work has also found that model-written plans can harm or under-triage, and that patients are weak judges of bad advice.

Google Research, Google DeepMind, and the primary care clinic at Beth Israel Deaconess Medical Center (BIDMC) ran a prospective single-arm feasibility study. Before an urgent-care visit, patients typed with AMIE on a computer. A summary went to the clinician (PCP). The primary questions were safety, conversation quality, and whether patients and clinicians would accept the tool. Diagnostic and management quality was checked later, against the chart eight weeks after the visit.

Method

The base model was Gemini 2.5 Pro, knowledge cutoff January 2025, with no domain fine-tuning and Thinking mode on, so it reasons internally before answering. An existing agent was retargeted at pre-visit history taking. It keeps a patient summary, a working differential (DDx), the facts still missing, and a draft management plan (Mx). Questions follow the current hypothesis instead of a fixed form. Before stating a view, it reads its understanding back and takes corrections. Diagnoses and next steps are labeled as items to discuss with the PCP, with a disclaimer. The Mx is stored for research and not shown to either side. After 50 encounters, latency forced a switch to Gemini 2.5 Flash.

A board-certified internist watched every chat remotely. The patient could not see that supervisor. Four rules could stop a session: harm to self or others, clear distress tied to the chat, clinical harm noticed by the supervisor, or the patient asking to stop. The supervisor corrected errors afterward. A 10-patient pilot produced no protocol change.

Enrollment was narrow: established adults, English listed as their primary language, a portal account, a computer rather than a phone, one complaint, and triage that had already ruled out the emergency department. Pregnancy and psychiatric chief complaints were excluded. Pay rose from $25 to $50 because recruitment was slow. The study stopped at 100 completed chats.

Eight weeks later, three internists blinded to AMIE set the final diagnosis from the record. Two physicians then scored AMIE's list on the five-point Bond/Graber scale, where 5 means the true diagnosis is present. Eight internists blindly rated both sides' DDx and Mx, three raters per case, median kept. Lists were cut to the same length, and Gemini 2.5 Pro rewrote every plan into one template. A manual audit checked 270 outputs. Clinicians had already read the AMIE summary before the visit.

Results

The clinic had 1,452 urgent-care visits in the window. Of 140 people who consented verbally, 114 started a chat and 100 finished (87.7%). Seven failed on device or protocol, five hit downtime, and two met an exclusion mid-chat, one pregnancy and one mental-health complaint. Two more missed the clinic visit, leaving 98 for analysis. The discussion puts enrollment at about one in ten urgent-care visits. The sample was younger: 51% were under 50, while more than half of the clinic's urgent-care visits were in people over 60. Scheduled clinicians were attendings for 62 patients, residents for 26, and nurse practitioners for 12.

No chat hit a stop rule. Supervisors spoke three times: to rule out an emergency the patient did not have, to state when to seek emergency care, and to correct AMIE placing a past surgery in the future.

The reference was the diagnosis in the chart at eight weeks. On Bond/Graber scores of 4 or 5, which count an exact diagnosis or one that is very close, the answer sat in AMIE's top 7 in 88 of 98 cases (90%) and the top 3 in 73 (75%). The top-ranked item matched the final diagnosis in 55 of 98 (56%). Accuracy stayed high in the 46 cases confirmed by a test, and trended higher in the 52 presumptive cases with no further test. The same top-k rates are not reported for clinicians.

Blinded quality scores did not separate the two sides on DDx (p = 0.6) or on how appropriate (p = 0.1) or safe (p = 1.0) the plan was. Clinicians scored higher on practicality (p = 0.003) and cost (p = 0.004), by paired Wilcoxon tests with Bonferroni correction across the five items. Raters guessed whether an output came from AMIE or a clinician correctly 59.18% of the time (95% CI 49.45% to 68.91%), only a little above chance.

Scores on the GAAIS, a scale of attitudes toward AI, rose after the chat (p < 0.001 on the total and both subscales) and did not rise further after the visit (p = 0.86). Clinician surveys came back for 61.2% of visits. Sixteen of 60 respondents had not read the summary. Among the 44 who had, 75% found it helpful for preparation, 68% called it harmless, 64% trusted it, and 57% thought it changed what they did. Nobody marked a plan very unhelpful or very harmful. Evaluators mostly rated the conversation favorable or very favorable. Patients were positive overall, but fewer than half felt their information would stay confidential or that the system seemed honest.

Why it matters

The study is prospective, uses real urgent-care patients, keeps a physician watching live, and lets the model offer possible diagnoses at the end. Earlier real-patient evaluations were mostly chart reviews after the fact, or limited to advice lines, telehealth, and intake before a specialist visit.

In simulations AMIE came out ahead, partly because the clinicians were also stuck in text. With each side on its usual channel, overall DDx quality and the appropriateness and safety of Mx were even. Practicality and cost still favor clinicians. Interviews with ten high-volume PCPs compared the note to a third-year student's: patients arrived with the history already organized. Five supervisors who were interviewed put the system at resident level and wanted a person still watching.

The fit is narrow. English-speaking established patients, a computer, no pregnancy or psychiatric chief complaint, already cleared of an emergency. What it saves is organizing the history before the visit. The management plan never entered the decision.

Limitations

A single arm cannot show that this beats ordinary urgent care. Zero stop-calls also depended on a physician watching every chat, with patients aware they were observed, so adversarial prompting had little room. Pregnancy, psychiatric complaints, and people who needed emergency care were screened out, and nobody was referred to the emergency department, so triage safety was not measured.

The comparison favors clinicians. They had read the AMIE summary, and they could open the chart and examine the patient. AMIE had no chart and no exam, and its differential was longer. That longer list and a wider workup are the explanation offered for the gap on cost and practicality. In 52 of 98 cases the reference diagnosis was only presumptive. Cases after the switch to Flash were not reported separately. About 7% failed for device reasons, and people without a computer were never invited. The sample is younger than the clinic. The human arm also includes residents and nurse practitioners. Nearly 40% of clinician surveys never came back.

A separate problem shows up inside the trace. A decent DDx appears early, and uncertainty falls along a similar curve whether the final list is right or wrong. Those turn-level scores were extracted by Gemini 2.5 Pro. A narrower confidence band did not mark the misses.

Terms

Source

What people are saying

Related papers

All paper explainers