A new JAMA Viewpoint makes the claim that AI will exceed humans when it comes to some— perhaps most— of medicine. The response online from physicians has been fierce. I haven’t seen doctors this angry since they were called providers.
I have just 5 points to make:
The essay’s central claim is: AI will make some— perhaps many— medical decisions better than most doctors, and doctor+AI is actually (or will be) worse than AI alone. Is this possible? Or science fiction?
Of course, yes, this is absolutely possible. I am surprised anyone would disagree. AI is already solving mathematical proofs that no human has been able to. AI can run statistical simulations that used to take a good statistician a week. AI can do the work of 10 medical students in gathering research information. What would take 1 months, is done in 1 hour.
And the corollary, it is possible that doctors override AI and worsen outcomes. Absolutely, this is possible. The authors give many examples. This is definitely a possible, plausible point of view. We can debate how fast and how extensive.
What about the human touch? Won’t patients always need a shoulder to cry on?
Of course this is also true, but we need to be honest: in the course of an oncologist’s week, what percentage of the time do you spend an hour talking about end of life care? 5%? And how much of that time does the person consoling the patient have to be a board certified oncologist vs. a palliative care doctor, or a general doctor, or an oncology nurse, or a medical student— my point is simply that they need to have a human touch in oncology does mean having oncologist exactly as they are today. I would imagine the JAMA authors saying AI + a local oncology nurse selected for best bedside manner visiting you at home is greater than going to see Jame in his cold Cleveland clinic office, waiting in the lobby, etc etc.
What about the author’s conflicts of interest? They profit from AI companies. Yes, I see that. I hate to break it to you, but the vast majority of essays in JAMA are written by people who profit from what they are recommending/ promoting. When drugs are recommended for on and off label uses, the doctors are often consultants. When devices are recommended, the doctors often make money by placing those devices. I have been a critic of much of this, but this is not unique to this piece. Also some of the disclosures are ridiculous. Zeke received an honoraria for speaking at UCSF grand rounds? Who cares? Disclosing everything means hiding what’s relevant. I am happy to have every conflicted article have a rebuttal written by someone with no conflict— but let us apply that everywhere in JAMA. (Good luck oncology).
Why are the authors writing in JAMA at all? AI is so wonderful, and modern communication is so amazing, and yet the authors are still writing about what AI will do in an old fashioned peer reviewed medical journal that mostly collects dust on coffee tables. And frankly a journal that is not fit for the times, and actually pretty boring. To be honest, this is the biggest mystery. I see no value to the authors for writing this article at all— they don’t need the line for the CV, they could publish this on a website and get more views, I see no reason for them to write this, which leads me to…
Why argue about what AI will do. Both the authors, and their critics are arguing about what AI might or might not do. This is in contrast with useful medical debates about what you should do in clinic tomorrow. Should you extend anticoagulation therapy? Give chemo and a TKI or TKI alone? Try a new pain pill? I get medical debates about decisions we make daily, but fighting over the future— I don’t get it.
I did find this one funny though. It’s tough to make predictions, especially about the future.
In short: Providers shouldn’t be upset at the idea that AI will steal some— perhaps all of their work. Ok, Doctors, I mean doctors. That was just to get your blood pressure up. How much work will AI take from me— I am curious to find out. Only way to do that is to stick around for a while. In the meantime, publishing viewpoints about the future in JAMA is about as valuable as reading and getting mad about it.





Getting to substance, the paper is foundationally weak, even as an opinion piece. There are two denominators for one study in one sentence, the six vignettes that became a load-bearing wall are listed as if they are large in number, sixteen patients per arm, concordance dressed as accuracy: others have done a much more complete analysis of the garbage-masquerading-as-science in here. https://www.fixhealth.ai/p/ai-will-replace-doctors-part-5-of
Look at the observations. On Hager’s 2,400 real patients, model performance collapsed the moment the models had to gather the information themselves instead of receiving a pre-packaged case. The Perspective’s own citation 23, Bean and colleagues in Nature Medicine, found that transmission of information between the model and the user is the particular point of failure. AMIE went from beating doctors in a text simulator to fifty-six percent on a hundred real patients with a physician supervising every interaction. Topol’s (like a broken clock, correct at least twice a day and he is correct here) reaction was that none of the studies were in real-world medicine.
Every study that shows a model outperforming physicians shares a hidden feature: the case arrived complete, consistent, and true, because a human expert assembled it, checked it, and edited it before the model saw it. The clinicopathological conferences, the case-report collections, the standardized patients, the vignettes in the Stanford trial. That assembly is invisible in the results, and it is the whole game. The “autonomous AI” in the evidence base is a model plus an unseen person who manufactured true premises for it, at a cost the study never counts. Take that person away, which is what real healthcare does, and you get Hager. In a separate 2025 experiment in Communications Medicine, when a single false detail was planted in the clinical material, frontier models repeated or elaborated it in fifty to eighty-two percent of outputs. In the real record, nobody has to plant such details. (https://www.nature.com/articles/s43856-025-01021-3) A diagnosis pasted a decade ago repeats with perfect fluency in every note that quotes it. I have personally seen a live commercial record that showed a patient simultaneously pregnant and due for prostate-cancer screening. The software was fine. The premises were not.
This is also why the doctors-make-it-worse finding, which the authors treat as decisive may be pointing somewhere else entirely. Why did fifty physicians add nothing to a model that had the case in front of it? Because in a curated study the case is already true, so there is nothing for the human to correct, and the physician’s only reference is the record. Move the same trial into a real clinic and the physician becomes the only agent in the loop even attempting the job the study quietly outsourced: establishing what is so. Not because human reasoning is superior. Because nobody else in the room is trying.
There is a foundational issue here, generally ignored by everyone because it has proven very difficult to solve. Generative AI succeeds wherever the world supplies a verifier: code compiles or it does not; a cited passage can be opened; a chess move is legal or it is not. Generative AI's failures concentrate wherever no external check exists, wherever premises cannot be verified by anything cheaper than an expert reading everything. Medicine is the largest domain on earth in that second category, and no amount of scale changes it, because the missing thing is not reasoning. It is a complete, reconciled, verifiable account of the one individual in front of the system, formed before the reasoning begins. There is not, and cannot be, a training set for that person. That account is not a generative act, and no model produces it for itself. It has to be manufactured deterministically, with every value traceable to its sources, so that when the reasoning is wrong the error has an address. Build that layer, and a great deal of medicine can safely become far more autonomous than the AMA would like. Skip it, and “AI alone” means a very fluent system reasoning beautifully about a patient who does not exist.
Here is a suggestion that would help everyone interpret this nonsense better. Every study should state where the case came from, who assembled it, and whether the model or a human established the premises. Then run the fair test on records as they exist, assembled by the AI from the actual chart rather than from a vignette, and measure fidelity to the premises alongside fidelity to the diagnosis. Until that column exists, every benchmark in this debate, including the ones that favor physicians, is measuring reasoning over a truth someone else supplied.
Validity requires a foundation of facts, to quote from the article cited above. And the foundation is the part nobody in the paper, and almost nobody in the field, is building. The authors may be right about 2030. But the milestone worth watching is not the day a model outperforms an internist on a case someone else assembled. It is the day the case assembles itself, truthfully, from the wreckage of the real record. That is a harder problem, it is not a language problem, and it is the one that decides everything else.
I can't really say how many individuals there are whose very name in the list of authors would alert me to not waste my time reading the article. But Dr. Emanuel would be very high on the list.