Multiple studies and news pieces recently got me thinking. You might have seen these for example:
Nature Medicine published a randomized controlled trial of an LLM assistant in real clinics, namely 9,700 patients in Kenya. Patient outcomes didn’t materially improve, but it had some interesting side findings around how clinicians (mis)trusted the AI. More on that below
A PLOS Digital Health study showed the opposite, where physicians keep trusting wrong AI recommendations even when the patient data in front of them contradicts the AI
Then theres the debate around the Utah pilot, which tests autonomous AI prescription renewal. The state medical board has been trying to stop that since April. Spoiler: The state said no
It’s amazing how far we’ve come with medical AI. Just to recap: Three years ago we mocked LLMs for inventing citations. Then about 1-2 years ago the first “AI is smarter than doctors” studies came out. They were easy to debunk, because they mostly tested static questionnaires and not realistic environments. But now we’re seeing another wave, testing clinical AI in increasingly sophisticated environments. Time to update our view!
The questions I want to answer: what’s better right now, a doctor with an AI or the AI alone? Where does it fail? And what does that mean for AI healthtech startups, when they design their products?
Note: I’m mostly talking about LLMs in this piece but I’ll mix in some findings from other machine learning tools. In terms of trust and distrust, the behavior around these tools should be comparable
Doctors meet AI: the failure modes
When it comes to using AI in clinical practice, many physicians and experts (including me!) used to point to the obvious solution: Let’s just give doctors AI, but let’s put it on a leash. The human signs off on AI decisions, to avoid hallucinations.
Sounds bullet proof in theory. In practice, three failure modes show up when you do that:
The “rubber stamp”: When the AI is wrong, doctors unfortunately often follow it. Some studies have tested that by inserting wrong AI answers: In a Radiology reader study with 220 physicians, decision accuracy reduced to roughly 25% when the AI’s advice was incorrect. An NEJM AI randomized trial found that false ChatGPT-4o output dropped diagnostic accuracy by 14 percentage points - although every participant had completed 20 hours of AI-literacy training beforehand. Senior physicians did worse btw, not better! The PLOS study from above fits here, too: Physicians stuck with a wrong AI answers although patient outcomes contradicted it
The “stubborn doctor”: In the opposite case, doctors tend to ignore the AI when it’s right. In another study, dermatologists switched from their own wrong answer to the AI’s correct one in just 10% of opportunities, and kept their own wrong answer in 23% of all cases. Radiologists underweight AI predictions even when the model outperforms two thirds of them. Again: More experience makes this worse, not better!
The “coin flip”: The root of both is, doctors aren’t great at telling when the AI is right. In one (very small) study, physicians correctly judged the AI output about 62% of the time… not much better than pure chance. The Kenya trial measured this in the field: An expert panel reviewed 1,000 cases where clinicians chose to follow or ignore critical AI alerts. Those decisions were clinically justified in only 28% of cases… ouch
Key takeaway: The problem isn’t simply too much trust or too little trust, but rather undiscriminating trust. Physicians seem to struggle with juding the correctness of AI outputs in either direction. Which raises questions about the concept of using humans to judge the AI.
And who wins head-to-head?
Seeing these results, it begs the question whether AI-only might be better for clinical decisions. I browsed through some of the evidence out there. Turns out, on clinical vignettes (=a bunch information in an articifical environment), the result is pretty consistent.
In the widely debated Stanford RCT from Oct 2024, GPT-4 alone scored 92%, physicians with GPT-4 scored 76%, physicians alone 74%. The follow-up on management decisions: doctors with AI beat doctors alone, but added nothing over the AI by itself. A 2026 replication in Pakistan gave physicians 20 hours of training first - the combo improved, but the AI alone still won. In fact, haven’t found a single peer-reviewed head-to-head study where a physician-AI combo beat an AI-only arm. Am I missing something? Please ping me if I do.
Of course there are some heavy caveats, in both directions:
Nearly every larger study we have tested on GPT-4 or 4o. So remember: The models today are 1-2 generations ahead of the published evidence! And everybody who’s used fable or the newer GPT models can feel the difference
On the flip side, all of this is just vignette work: They didn’t use real patients, and sometimes had human baselines with little access to knowledge sources. Very favourable for the AI
The one real-world LLM trial with a hard endpoint is the Kenya RCT from the intro. The conclusion was it’s safe, better for documentation, but they couldn’t find proof for an outcome effect using AI-in-the-loop. Besides, real-world trials like this don’t have “AI-only” arms (for obvious regulatory and ethical limits), so we still don’t know how that would have ended.
I’d summarize the current state as follows: AI wins vs doctors on paper, but we don’t have proof for AI outperforming in real life. Neither in combination with doctors, nor alone. We need to be careful with this distinction. Besides this, what we do know is that human doctors tend to misjudge AI answers.
Then let’s fix the guardrails, right?
My immediate impulse was we need better guardrails for human-AI combinations. Surely there must be a way to mitigate these trust problems.
Well, it turns out multiple mitigation methods have been tested in the literature. Among them are:
AI-literacy training: increased the AI benefit in the Pakistan study, although it did nothing against the “rubber stamp” effect
Result explanations: sped up agreement whether the AI was right or wrong (which is the opposite of what we want)
Explicit trust-calibration exercises: no measurable effect so far
Displayed uncertainty: seems to raise trust in the AI results, but unproven whether it raises appropriate trust
Workflow order (AI-first vs AI-second): changes how doctors reason, but not how accurate they are
Rather disappointing overall. The likely reason: All these tools target how much doctors trust the AI, while the actual deficit is when to trust it. We’d basically need another, better-than-average-human judge to govern the AI.
And there’s a second problem to deal with long-term: de-skilling. Physicians that use AI tools long enough lose their skills. The first real-world evidence came from endoscopists in Poland, whose polyp detection rate in non-AI colonoscopies dropped about 20% relative after a few months of routine AI assistance. That doesn’t help mitigate our concerns…
Btw, you might ask: Shouldn’t this already be part of medical device regulation? It is partly from a European perspective: The MDR's usability process (IEC 62366) forces manufacturers to analyze “use-related risks”. The new joint guidance on MDR and the AI Act includes human oversight as a formal risk management measure. So regulators seem aware of the problem field, but guidance is rather vague. I hope they’ll look closely at these failure modes, especially post-market
The direction of travel
Seems like we won’t settle this debate for a while. Prospective trials in real-life settings are hard, models change faster than any study can be published, and implementation details confound the results.
But when I look at all the hints and preliminary results, I personally believe the AI is already superior vs human doctors. The question is how to translate that into real-life clinical care with robust regulatory surroundings.
The good news is we’ve taken first steps:
Utah’s pilot with Doctronic lets an AI handle chronic prescription renewals. First with physician pre-approval, later with retrospective review. That’s the pilot the medical board was raging about
UpDoc recently claimed to have the first FDA-approved autonomous clinical AI in healthcare. The press statements were blown out of proportion imo - the algorithm automatically suggests new insulin doses for diabetic patients. So it’s a glorified dose calculator and not an AI doctor… but hey, small steps!
Europe is moving too, but slowly (who would have thought?). Germany’s AI Act implementation law speaks about AI sandboxes, let’s see when they become reality. I assume primary care triage will be one of the use cases. The NHS just announced an AI triage tool in the NHS App, going to 200,000 patients within 12 months. Both are bold steps, and I highly welcome them
I truly hope this is the route we take: Start with AI in a sandbox, pick one tightly defined workflow, build real control mechanisms, then expand.
For everything else, the doctor-plus-LLM combination is the realistic path forward. Not because it guarantees the best outcomes, but for 3 practical reasons.
It’s the only legal route for risky tasks at the moment. The EU AI Act for example mandates human oversight for high-risk medical AI. And liability is a huge open question
Shadow AI is a fact - doctors already use LLMs daily, and banning that is fantasy
Human-in-the-loop is socially accepted and trusted no matter the evidence. So let’s make the best out of it
What it means for startups
As a startup, you probably fall into one of two buckets:
If you’re building an autonomous AI loop, get close to policymakers and get into the sandboxes. Deploying autonomous clinical AI is more a political than a technical exercise these days.
If you’re handing physicians an AI to work with (e.g., scribes, search tools, coding and billing agents), be ready for critical questions. Build guardrails against the rubber stamp and the stubborn experts, and take quality control seriously. Beware of the well-intentioned but ineffective mitigation strategies.
And if physicians-in-the-loop are your quality control: Keep raising the bar as the models improve. In the not-so-far future, “human in the loop” might turn into a fiction… required for approval but not backed by evidence.
Speak soon,
Lucas
P.S. If you’re among the regulatory pros, I’m curious to hear your predictions on how this debate affects MDR and FDA going forward




