Google's AMIE diagnostic chatbot has now been tested on real patients, and the peer-reviewed result is more modest than the headlines. On October 8, 2026, Google Research updated its write-up of the study to say the work had been published in The Lancet. In a feasibility trial at Beth Israel Deaconess Medical Center (BIDMC) in Boston, AMIE's list of possible diagnoses contained the eventual diagnosis in 90% of cases. It named that diagnosis as its single top guess in 56%. Physicians still beat it on practical, affordable management plans.
That is a real milestone: it is the first time this research system has been run on live patients rather than actors. It is not proof that a chatbot "matches doctors."

TL;DR
| Item | What the study reports |
|---|---|
| Setting | Healthcare Associates, an academic primary care practice at BIDMC, Boston |
| Design | Prospective, single-arm, single-center feasibility study; registered as NCT06911398 |
| Participants | 100 adults completed the AMIE chat; 98 attended their physician visit |
| Task | Text chat to take a history before a new, non-emergency visit |
| Oversight | Physician "AI supervisor" watching live; zero safety stops |
| Diagnosis | Final diagnosis in AMIE's differential in 90% of cases; ranked first in 56% |
| Versus physicians | No significant difference on differential and management safety; physicians better on practicality and cost |
| Main limits | Text only, one site, no controlled comparison, no EHR access |
What AMIE is
AMIE (Articulate Medical Intelligence Explorer) is Google's research system for diagnostic conversation. Earlier papers tested it in simulated text consultations against primary care physicians, and Google has since extended it to specialist settings. Google's AMIE research page was updated on October 6, 2026 with published cardiology (Nature Medicine) and oncology (NEJM AI) papers, and it describes this real-world study as the step beyond controlled evaluations.
The point of the BIDMC study was narrow: could the system safely gather a patient history in a real clinic, and how would patients and clinicians react? It was not designed to show better outcomes.
How the study worked
Patients booking a new, non-emergency, episodic visit, in person or by telehealth, were invited to take part, and were told that declining would not affect their care. Participants chatted with AMIE through a secure web link before the appointment. Per Google's description, a physician supervisor watched by live video with screen sharing and could intervene at any point.
Before the visit, with patient consent, the clinician received a transcript and a summary. Afterwards, independent clinical evaluators rated the differential diagnoses and management plans from AMIE and from the treating physician. The raters were blinded and the order was randomized. Each case had three evaluators, and the median was used.
Supervisors were trained to stop a session for one of four reasons: immediate risk of harm, significant patient distress, potential clinical harm, or a patient request to end. Google reports zero safety stops.
The 90% number, read carefully
The 90% figure is a recall figure: the final diagnosis, established by chart review eight weeks later, appeared somewhere in AMIE's differential. The stricter number is 56%, where AMIE ranked the eventual diagnosis as most likely.
Google's own page is inconsistent on the middle band. One section cites 75% top-3 accuracy; another describes a match within the top 7 in 90% of cases. Secondary summaries repeat the 75% figure. We could not resolve this from the blog alone, so check the arXiv paper before quoting a top-3 number.
Two more cautions on the metric:
- A differential that is long enough will often contain the answer. Recall at a generous cutoff is not the same as a correct diagnosis.
- Google says accuracy stayed high in the 46 cases with test-confirmed diagnoses and was higher for presumptive diagnoses, which are by nature less certain.
Where physicians still won
Blinded evaluators found no significant difference between AMIE and primary care physicians on differential diagnosis quality or on the appropriateness and safety of management plans. But physicians outperformed AMIE on the practicality and cost-effectiveness of management plans, a point the authors attribute to AMIE lacking EHR access, a physical examination, and multimodal input.
That is the gap that matters in real care. A plan can be safe and still be one a clinic could never run. Anyone who has watched coding agents propose technically valid but unmaintainable changes will recognise the pattern, as in our look at how metrics get gamed.
Patient and clinician reactions
Patients completed the General Attitudes towards AI Scale (GAAIS) before and after the chat and after the provider visit. Attitudes were significantly more positive after the AI interaction and stayed that way through the visit, across both the perceived-utility and concerns subscales. Patients described AMIE as polite and good at explaining conditions.
Clinicians found the transcripts useful for visit preparation. Per Google, physicians said visits shifted from data gathering toward data verification and more collaborative decision-making. Treat this as the authors' summary of qualitative feedback, not a measured workflow improvement.
Limits the authors acknowledge
- Text only. No non-verbal cues, no physical findings.
- Single center, single arm. There was no control group, so the study cannot say whether AMIE beats the standard workflow.
- Selection. Participants skewed younger than the clinic's visit population, where over half of visits are by patients over 60. Eligibility reportedly required English, a single chief complaint, and computer access.
- Unexplored factors. The effect of health literacy, tech literacy, and prior chatbot familiarity was not fully studied.
- Industry-authored. The work comes from the team that built the system.
The AMIE team made a related argument in a Nature Medicine commentary in September 2026, per AI Weekly's summary: "Trust in clinical artificial intelligence (AI) cannot be benchmarked into existence." We have not read the full commentary, so we cite it only through that summary.
Why this matters beyond Google
Medical AI claims have mostly rested on benchmarks and simulated cases. A prospective, pre-registered study with live patients, blinded raters and stated stop criteria sets a better template, whoever runs it. It also shows what "AI in healthcare" can look like when it is deliberately supervised: the model gathers information, a human stays in the loop, and the patient still sees their own doctor.
Compare that with Nolla Health's Utah approval, where an AI issues initial prescriptions for a narrow condition under staged physician review. Both are narrow, supervised deployments, and both are being described in sweeping language by someone. In each case the useful question is what, exactly, was permitted or measured.
The study also fits a broader pattern of AI moving into biology and medicine this year, from the Biohub virtual biology initiative to Gemma-based plant DNA variant ranking. For readers tracking where regulators are heading, the NYC Council validation and kill-switch hearing is a useful companion: the same questions about validation and shutdown apply in medicine. And for a view on how much of the field is now health-related, the Stanford AI Index takeaways give the wider context.
How to read the next "AI beats doctors" headline
- Find the denominator. How many patients, how many sites, how many arms?
- Find the metric. Recall at top-k, top-1 accuracy, and safety ratings are different claims.
- Find the comparator. "Comparable to" physicians on selected axes is not "as good as" physicians.
- Find the supervision. A human watching every session changes the risk profile entirely.
- Find the outcome. Did anyone measure patient outcomes, cost, or time? Here, nobody did.
Applying that checklist to AMIE: 100 patients, one site, one arm; 90% recall and 56% top-1; comparable on two axes and worse on two; supervised live; no outcome measures. It is promising feasibility evidence and an honest piece of method, which is a good deal more than most medical-AI announcements offer.
What this means for builders
If you build AI products for regulated fields, the study design is the takeaway. Notice the choices that made it publishable: pre-registration before results, a narrow task (history taking rather than diagnosis delivery), explicit stop criteria for the human supervisor, blinded raters, and a plain statement of what the design cannot show. Those same ideas carry over to legal, financial and HR agents, where "the model scored well on a benchmark" is similarly weak evidence. Log the supervisor's interventions, publish the denominators, and let the comparison group be a real human doing the real job.
What to watch next
Google names three directions: voice or video capability, larger controlled studies with comparison groups, and asynchronous physician oversight of the kind tested in earlier simulated studies. A controlled trial is the one that would move the conversation, since it could test whether pre-visit AI intake changes diagnosis accuracy, visit time or cost. Until then, AMIE is a research system, not a product you can use or a substitute for a clinician.
Related reading
- Nolla Health: Utah lets an AI issue initial acne prescriptions
- Biohub virtual biology initiative
- Gemma plant DNA variant ranking
- NYC Council AI validation and kill-switch hearing
- Stanford AI Index 2026 takeaways
- Specification gaming and Goodhart's law in AI metrics
- Gemini 4 Argon vs Opus 5.5 vs Grok 4.7 vs GPT-6
Primary sources: Google Research study write-up, arXiv preprint 2603.08448, ClinicalTrials.gov NCT06911398.
Details reflect Google's published description as of October 9, 2026. This is not medical advice.
