Google has tested its research medical AI system, AMIE, in real-time video consultations — and the results show it performing at the level of primary care physicians on several key measures. The study used professional patient actors rather than real patients, marking an important step in understanding how AI could support clinical conversations.
How Google tested AMIE in video consultations
According to Google Research, AMIE (Video) conducted synchronous video consultations with fifteen trained actors. These actors portrayed conditions across five areas: cardiopulmonary, abdominal, HEENT (head, eyes, ears, nose, and throat), neurological or psychiatric, and musculoskeletal presentations.
Clinical evaluators then rated AMIE's performance. The ratings came in on par with primary care physicians across several core measures — a significant result for a research system still in development.
AMIE's multi-agent architecture for clinical conversations
Google says AMIE uses an asynchronous multi-agent architecture rather than assigning all tasks to a single model process. The system divides work among three agents: dialogue, clinical reasoning, and perception.
According to Google's blog, a single agent cannot currently sustain natural conversation at this level. By splitting the workload, each agent can focus on its specific role during a consultation.
"A single agent cannot currently sustain natural conversationa..." — Google Blog
What this means for real-world clinical use
Google is clear about the limits of this study. The company states that studies involving real patients and their own health conditions must follow before anyone can draw conclusions about clinical use.
This is a research milestone, not a product launch. The actor-based setup allowed Google to test the system in controlled conditions, but real-world medical consultations bring complexity that actors cannot fully replicate.
Our Take: Promising research, but real patients come next
To put it plainly, this is encouraging progress for medical AI — but it is not proof that AMIE is ready for clinics. Matching primary care physicians in a controlled study with actors is a meaningful benchmark. It shows the technology can hold a conversation, reason through symptoms, and process visual information in real time.
However, the gap between acting a condition and living with one is enormous. Real patients bring messy histories, emotional stress, and conditions that do not fit neatly into categories. The multi-agent design is a smart approach to a hard problem, but it still needs to prove itself where it matters most: with people who are actually sick.
Readers should watch for the next phase — real-patient studies. If AMIE performs there as it did with actors, the conversation about AI in clinical settings will shift from "could it work?" to "how do we deploy it safely?"