Two-thirds of American doctors are already using an AI chatbot called OpenEvidence to diagnose patients and check drug interactions. Now medical students and residents are using it too — which would be fine, except they're using it instead of developing the clinical reasoning skills that were supposed to be the entire point of medical school. We are, in other words, training doctors who may never learn how to actually think like doctors.
The AI That Looks Like Competence
Here is the specific problem, laid out in The Guardian by physicians Simar Bajaj and Joseph Sakran. A medical trainee is asked to generate a list of possible diagnoses. In the old model, the trainee struggles, offers an incomplete answer, gets corrected, feels embarrassed, and remembers it forever. That embarrassment is the lesson. That's how the knowledge gets wired in.
In the new model, the trainee quietly asks OpenEvidence, gets back a polished, comprehensive answer in seconds, repeats it on rounds, impresses the supervising physician, and learns absolutely nothing. The attending thinks the trainee is sharp. The trainee thinks they did well. The gap in actual knowledge stays invisible until, someday, the AI isn't available or gives a wrong answer and there is no backup.
Bajaj and Sakran name this precisely: it's not deskilling, it's never-skilling. Deskilling assumes someone once had an ability and lost it. What they're describing is trainees who might complete their entire medical education without ever building independent clinical reasoning in the first place. A doctor who forgot how to reason might recover it. A doctor who never learned? That's a different problem.
The Tool They're Trusting Has a Reliability Problem
And here is where it gets worse. OpenEvidence pitches itself as a clinical AI anchored in the latest medical literature, which sounds reassuring right up until you read what a recent study in Nature Medicine actually found. According to The Guardian's reporting on that research, tools that pull from the latest medical literature, which is exactly what OpenEvidence does, can be less reliable than they appear. In some cases, they're less accurate than general-purpose AI chatbots. General-purpose. As in, the same kind of chatbot you'd use to write a birthday poem.
So trainees are offloading their clinical reasoning to a tool that, per peer-reviewed research, may be worse than just using a regular AI. They're trusting it enough to shape how they understand medicine. And the problem of misplaced trust, as Bajaj and Sakran put it, is already baked into the system trainees are using right now. This isn't a hypothetical future risk. It's already running.
The Arms Race Nobody Signed Up For
To be fair to the trainees, many of them reportedly know this is a trap. They told Bajaj and Sakran directly: yes, they understand OpenEvidence can become a crutch. But here's the bind they're in. If everyone else on the team is using AI to sound prepared, opting out feels like showing up to a gunfight with a knife. They called it unilateral disarmament. That framing should stop you cold, because it means the incentive structure of medical training has already broken down.
No individual trainee can fix this by exercising personal discipline. If the system rewards sounding right over being right, and AI is the fastest path to sounding right, then the students using it aren't making a moral failing. They're responding rationally to a badly designed system. That's an institutional problem, and it demands an institutional answer.
What an Actual Fix Looks Like
Bajaj and Sakran have concrete suggestions, and they're worth taking seriously. The core proposal: reason first, consult AI second. Before anyone opens OpenEvidence, the trainee commits to a written pre-AI assessment. Leading diagnosis, dangerous possibilities to rule out, recommended next steps. On rounds, when something changes, the attending pauses the whole team before anyone reaches for their phone. The friction is the point. Learning scientists Elizabeth and Robert Bjork have a name for this: desirable difficulties, the idea that slowing performance in the moment builds retention and transfer of skills over time.
The piece also points to aviation as a model. The Federal Aviation Administration literally advises pilots to periodically disengage autopilot and hand-fly to keep their manual skills sharp. Medicine, the authors argue, needs the same discipline: regular no-AI cases, unaided reasoning assessments, and drills where trainees have to identify a subtle flaw in a polished AI-generated assessment. Crucially, that last drill should also include AI outputs that are completely correct, so trainees learn to evaluate the tool rather than just distrust it reflexively. The goal isn't AI avoidance. It's AI literacy built on top of actual medical competence.
The Supervisors Who Were Trained By AI
There is a question buried in this piece that nobody wants to sit with too long. Bajaj and Sakran raise it directly: what happens when the generation trained alongside AI becomes the generation responsible for catching AI's mistakes? Right now, experienced physicians can evaluate what OpenEvidence spits out because they have decades of independent clinical reasoning to measure it against. They know what a correct answer smells like.
But if trainees spend their formative years having AI do that reasoning for them, they won't have that intuition. They'll become attendings who supervise AI without the independent framework needed to know when it's wrong. We will have built, as the authors put it, supervisors of reasoning before we created reasoners. That is a sentence that should be read aloud at every medical school faculty meeting in the country.
The Dingo Take
You are supposed to believe that because AI gives medically accurate-sounding answers quickly, integrating it into medical training is straightforwardly good. You are not supposed to notice that the entire value of medical training is the process of getting answers wrong and understanding why, and that AI short-circuits that process completely while making everyone involved feel great about it.
The most damning detail in this whole piece isn't the reliability problem with OpenEvidence, though that's bad enough. It's the arms race. Trainees know this is making them worse doctors. They're doing it anyway because the system punishes them for opting out. That is not a technology problem. That is a leadership problem. Medical schools and residency programs have had enough time to see this coming, and the near-total absence of structural policy around AI use in training is its own kind of negligence.
At some point, one of these trainees who never learned to reason independently is going to be the doctor on call when the hospital's AI system goes down, or gives a confidently wrong answer, and there's nobody in the room who knows enough to catch it. We won't be able to say we didn't see it coming. This piece is right here. The study is right there. The fix isn't even that complicated. The only thing missing is the institutional will to make the training hard again, which is apparently asking a lot.
Comments