Discussion about this post

User's avatar
Dr. K's avatar

Getting to substance, the paper is foundationally weak, even as an opinion piece. There are two denominators for one study in one sentence, the six vignettes that became a load-bearing wall are listed as if they are large in number, sixteen patients per arm, concordance dressed as accuracy: others have done a much more complete analysis of the garbage-masquerading-as-science in here. https://www.fixhealth.ai/p/ai-will-replace-doctors-part-5-of

Look at the observations. On Hager’s 2,400 real patients, model performance collapsed the moment the models had to gather the information themselves instead of receiving a pre-packaged case. The Perspective’s own citation 23, Bean and colleagues in Nature Medicine, found that transmission of information between the model and the user is the particular point of failure. AMIE went from beating doctors in a text simulator to fifty-six percent on a hundred real patients with a physician supervising every interaction. Topol’s (like a broken clock, correct at least twice a day and he is correct here) reaction was that none of the studies were in real-world medicine.

Every study that shows a model outperforming physicians shares a hidden feature: the case arrived complete, consistent, and true, because a human expert assembled it, checked it, and edited it before the model saw it. The clinicopathological conferences, the case-report collections, the standardized patients, the vignettes in the Stanford trial. That assembly is invisible in the results, and it is the whole game. The “autonomous AI” in the evidence base is a model plus an unseen person who manufactured true premises for it, at a cost the study never counts. Take that person away, which is what real healthcare does, and you get Hager. In a separate 2025 experiment in Communications Medicine, when a single false detail was planted in the clinical material, frontier models repeated or elaborated it in fifty to eighty-two percent of outputs. In the real record, nobody has to plant such details. (https://www.nature.com/articles/s43856-025-01021-3) A diagnosis pasted a decade ago repeats with perfect fluency in every note that quotes it. I have personally seen a live commercial record that showed a patient simultaneously pregnant and due for prostate-cancer screening. The software was fine. The premises were not.

This is also why the doctors-make-it-worse finding, which the authors treat as decisive may be pointing somewhere else entirely. Why did fifty physicians add nothing to a model that had the case in front of it? Because in a curated study the case is already true, so there is nothing for the human to correct, and the physician’s only reference is the record. Move the same trial into a real clinic and the physician becomes the only agent in the loop even attempting the job the study quietly outsourced: establishing what is so. Not because human reasoning is superior. Because nobody else in the room is trying.

There is a foundational issue here, generally ignored by everyone because it has proven very difficult to solve. Generative AI succeeds wherever the world supplies a verifier: code compiles or it does not; a cited passage can be opened; a chess move is legal or it is not. Generative AI's failures concentrate wherever no external check exists, wherever premises cannot be verified by anything cheaper than an expert reading everything. Medicine is the largest domain on earth in that second category, and no amount of scale changes it, because the missing thing is not reasoning. It is a complete, reconciled, verifiable account of the one individual in front of the system, formed before the reasoning begins. There is not, and cannot be, a training set for that person. That account is not a generative act, and no model produces it for itself. It has to be manufactured deterministically, with every value traceable to its sources, so that when the reasoning is wrong the error has an address. Build that layer, and a great deal of medicine can safely become far more autonomous than the AMA would like. Skip it, and “AI alone” means a very fluent system reasoning beautifully about a patient who does not exist.

Here is a suggestion that would help everyone interpret this nonsense better. Every study should state where the case came from, who assembled it, and whether the model or a human established the premises. Then run the fair test on records as they exist, assembled by the AI from the actual chart rather than from a vignette, and measure fidelity to the premises alongside fidelity to the diagnosis. Until that column exists, every benchmark in this debate, including the ones that favor physicians, is measuring reasoning over a truth someone else supplied.

Validity requires a foundation of facts, to quote from the article cited above. And the foundation is the part nobody in the paper, and almost nobody in the field, is building. The authors may be right about 2030. But the milestone worth watching is not the day a model outperforms an internist on a case someone else assembled. It is the day the case assembles itself, truthfully, from the wreckage of the real record. That is a harder problem, it is not a language problem, and it is the one that decides everything else.

Ernest N. Curtis's avatar

I can't really say how many individuals there are whose very name in the list of authors would alert me to not waste my time reading the article. But Dr. Emanuel would be very high on the list.

8 more comments...

No posts

Ready for more?