DentAI – When AI Alone Beats the AI-Assisted Physician
فارسیIf we give a physician access to ChatGPT, does their diagnosis improve? A randomized clinical trial in JAMA Network Open tested exactly this, and the answer was no: the GPT-4-equipped physician scored about 76% and the physician with conventional resources 74% — a meaningless difference. But the surprising finding lay elsewhere: GPT-4 alone, working on the same cases without any physician involvement, scored 92% and significantly beat both physician groups. The gap is not the model's capability; it is how the tool is used — and that lesson matters directly as these tools enter dentistry.
If we give a physician access to ChatGPT, does their diagnosis improve? A randomized clinical trial published in JAMA Network Open tested exactly this question, and the answer was no. But the study's surprising finding lay elsewhere: GPT-4, when working alone on the same cases without physician involvement, outperformed both physician groups. In other words, adding the physician to the model brought the result down.
This finding matters directly to dentistry too, because these same tools are entering our everyday work.
How the Study Was Done
The design was simple and clean. Fifty physicians from three respected US universities, including Stanford, entered the study and were randomly split into two groups. One group had access to GPT-4 in addition to the usual resources such as UpToDate and Google. The other group had only the usual resources and was not allowed to use any language model.
Each physician had one hour to work on six complex diagnostic cases. The cases were based on real patients and had never been published anywhere; therefore GPT-4 could not have memorized their answers from its training data.
Another important point: only the final diagnosis was not what got scored. The entire reasoning process was assessed: the differential diagnoses, the supporting and opposing evidence for each, and the next steps. Almost the same structure we write in the Assessment and Plan of a clinical note. The scoring was also blinded — the grader did not know which group a response belonged to, or even whether it belonged to a human or the model. Because there was a third arm too: GPT-4 by itself, with a prompt that the research team had carefully designed.
The Results
The study's three main numbers:
- Physician plus GPT-4: about 76 percent
- Physician with conventional resources: about 74 percent
- GPT-4 alone: about 92 percent
Access to the model helped essentially not at all. The two-percentage-point difference between the two physician groups was not statistically significant. But the model alone was 16 points ahead of the physicians, and that difference was entirely significant.
The pattern was the same across all subgroups; it made no difference whether the participant was an experienced attending or a resident, and prior experience with ChatGPT changed nothing either.
One secondary finding was also notable: the GPT-4 group was about one minute faster per case, and those who had previously worked with ChatGPT were more than two minutes faster. This difference did not reach statistical significance, but its direction is encouraging, because in medicine time is always scarce.
Why Human plus AI Didn't Work
This is the most important part of the story. The model proved its capability and scored 92 percent. So why was that capability lost in the hands of the physicians?
The authors' main explanation: how it was used. That 92 percent came from a prompt the specialists had built carefully. But the study's physicians had received no training in prompt engineering and worked with the model the same way most users work today. The result was that the strongest diagnostic tool available to them was effectively neutralized.
So the gap is not the model's capability; it is the gap in interaction. And the solution follows from exactly there: either clinicians must learn to work with these tools, or systems must carry optimized prompts built in from the start, so the user isn't forced to be a prompting expert.
But an important caveat that the authors themselves state explicitly: these results do not mean the model should diagnose without oversight. The cases in this study were summarized and clean, and all the needed information had already been extracted. Real clinical practice is not like that. The real patient is full of noise and ambiguity, and pulling out the right information from within it is still the clinician's job.
The Message for Dentistry
The study was done on physicians, but the structure of the problem holds exactly in dentistry too: one clinician, one diagnostic case, and one AI tool at hand.
Three lessons emerge from the findings:
1. Merely having access to AI does not improve diagnosis.
That is precisely what was tested, and the answer was no.
2. The skill of working with these tools is a new clinical skill.
How you ask and how you use the output is exactly what created the gap between 76 and 92 percent.
3. High model performance on clean cases does not mean readiness for real clinical practice.
Judgment under ambiguity remains the clinician's domain.
Summary
This study showed that giving physicians access to GPT-4 did not improve their diagnostic reasoning, while the same model on its own was significantly ahead of everyone.
The final message is simple: the future of diagnosis is not a contest between human and AI. The question is to learn how to build this combination so that it is stronger than either of its parts, not weaker. And for now, without training and without proper design, we are not there yet.
"Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial"
Goh E, Gallo R, Hom J, Strong E, Weng Y, Kerman H, et al. — JAMA Network Open. 2024;7(10):e2440969
DOI: 10.1001/jamanetworkopen.2024.40969