Rebutting Khosla and Emanuel on Autonomous AI in Medicine

On Straw Men and Doormen

(Gently) Rebutting Khosla and Emanuel on Autonomous AI in Medicine

 

SubStack

Robert Wachter

Aug 27, 2026

 

Last week’s JAMA piece by legendary investor Vinod Khosla, his son Neal (CEO of AI-enabled primary care start-up Curai, and, as of today, co-owner with Vinod of the Seattle Seahawks), Penn’s Zeke Emanuel, and Zeke’s research assistant Abe Baker-Butler garnered a lot of press.

 

Vinod and Neal Khosla, at a press conference after buying the Seahawks

 

(Disclosures: I am an advisor to Curai, in which Vinod Khosla is an investor. Zeke Emanuel is a friend and collaborator.)

 

The article begins by setting me up as a straw man:

 

“In ‘A Giant Leap: How AI Is Transforming Healthcare and What That Means for Our Future,’ Wachter argues that the highest tier of care will be AI-aided physicians, whereas AI-only care will be medicine’s ‘economy class.’ We disagree.”

 

As far as it goes, the summary is fair. The straw man is in what it implies: that I consider AI-only care to be second-rate medicine. I’m not sure it will be, certainly not in all settings. In fact, if we get it right, some might even prefer it to seeing a human doctor, just as I prefer a Waymo to an Uber.

 

My argument was that AI-first care might be entirely acceptable for the right patients with the right needs, and that what the upper tier buys you may not be a more accurate diagnosis. It may simply be a human physician.

 

Given that the authors used my book as a springboard, it seems proper for me to clarify my thinking and critique the parts of the Emanuel/Khosla(s) argument that I disagree with. Here goes.

 

The Emanuel/Khosla Argument

In their paper, Emanuel, the Khoslas, and Baker-Butler argue that AI tools have improved to the point that keeping a human in the loop is more likely to degrade the output of the AI-human dyad than improve it. They cite several studies to buttress this contention, most persuasively a 2024 paper on a measure of diagnostic reasoning. The researchers found that doctors got it right 74 percent of the time; doctors working with AI improved slightly, to 76 percent. But the best performance came when the AI acted independently, at which point it was right 90 percent of the time. Stanford AI expert Jonathan Chen refers to this as his “Holy Crap!” moment. When I interviewed him for A Giant Leap, he told me,

 

“It really flies in the face of the classical fundamental theorem of informatics, right? That the human-plus-computer is going to be better than either alone. [The theorem] feels so good. It sounds so good…. And now you look at these results, they imply that it’s not the case.”

 

There are several reasons the human in the loop might not add value and might even be counterproductive. First, when one thinks about tasks that humans stink at, remaining eternally vigilant when they’ve come to trust an AI output would be high on the list. As Emanuel et al. point out, there’s also the problem of algorithmic aversion – the human compulsion to override AI’s guidance, even when that guidance is more likely to be correct than they are. Finally, there is the problem of de-skilling, in which human performance degrades over time as people become increasingly dependent on a technology.

 

Given all of this, I agree with the main premise of the JAMA article: Given a curated set of facts, AI will often outperform humans in diagnosis and management decisions – and we doctors may sometimes get in the way if we try to intervene.

 

That said, I have two main issues with the authors’ argument.

 

The First Problem with the Argument: The Reliable Fact-Set and GIGO

As they say in the tech business, garbage in-garbage out. And that’s a big problem in diagnosis.

A huge part of the art of medicine is distilling what might be 200 facts about a given patient (current symptoms, medications, medical history, surgical history, family history, travel history, pets, etc. etc.) into the salient ones that we are trying to solve for. This distillation is at the heart of medicine – we call these salient facts the “problem representation” and we use them to organize our own thinking about cases, to discuss cases with colleagues, and, these days, to load into AI tools.

My problem representation for a complex case might be, “This is a 35-year-old woman who presents with 2 days of shortness of breath, pleuritic chest pain, and a swollen left leg. She smokes, takes oral contraceptives, and had a traumatic tibial fracture 2 months ago that left her immobilized until last week.”

Patients usually don’t know what elements of a case are important. And why should they? When it comes to creating problem representations, they are novices. A vivid illustration of this was seen in a study published earlier this year, in which Oxford investigators fed well-crafted cases into several large language models, including GPT-4o. One memorable one was, “While at the movies, a patient suddenly developed the worst headache in his life.” Put that into an LLM, and it responds as every decent medical student would: That’s a subarachnoid hemorrhage until proven otherwise, and the patient needs to go to the ED pronto. And that’s what the models said, more than 90 percent of the time.

Next, in a clever twist, the investigators gave these case representations to laypeople and asked them to interact with the AI models. Patients, of course, had no idea that “the worst headache of your life” is an adrenaline-raising term of art for physicians. So patients read the script, then entered prompts like, “I woke up today with a bad headache.” And the model politely reassured the patient and recommended bed rest and Tylenol – precisely the wrong call. In a variety of scenarios, the AI made similarly wrong calls – both on diagnosis and on appropriate action for the patient to take – roughly two-thirds of the time.

The problem wasn’t with the tool; it was with the user. All of this means that AI’s diagnostic performance on a curated set of clinical facts should not be taken as compelling evidence that “Dr. AI” can replace actual doctors, who are experts in creating correct and useful problem representations.

The Second Problem: The Doorman Fallacy

Even if AI were truly better than doctors at diagnosis and treatment recommendations, it does not necessarily follow that doctors can be replaced. This is where the doorman fallacy comes in. 

The term “doorman fallacy” was coined by British advertising executive Rory Sutherland, who explains it here. (Hat tip to UNC’s Spencer Dorn and UCSF’s Gurpreet Dhaliwal for introducing me to this lovely analogy.) In the 1970s, industrial engineers perfected the automatic door opener and the infrared person detector. Combining the two technologies meant that a building could trigger a door to open as a human approached. Consultants were quickly dispatched to large hotels and upscale apartment buildings to convince the owners that their expensive doormen could go the way of the buggy whip.

And then something funny happened. After doormen were removed, people realized that they did far more than simply open the door. They helped with building security, gave directions, received packages, recommended the best nearby coffee shop, held an umbrella over you as you sprinted to your cab, and added luster and prestige to your property. More ineffably, people seemed to value human connection when they left for work or returned after a long day. In other words, the term “doorman” was a vast oversimplification of the value that these individuals brought to the building. Many properties quickly rehired their doormen.

The Doorman Fallacy came to mind last week while I was on clinical service at UCSF Medical Center. OpenEvidence was immensely helpful in answering questions I had about my patients: Is an elevated CA-125 specific for ovarian cancer or can it be seen in endometriosis? When can I restart aspirin in a patient with a prior stroke and now an upper GI bleed? And many more. I’m a huge fan of these AI tools, which act like a curbside consultant in my pocket. Given the right fact set, they are smarter than I am, and only getting smarter.

But even as AI helped me with diagnosis and management decisions, how much of my work did it replace or fundamentally improve? It sure wasn’t 80 percent, the fraction that Khosla has argued for many years. It was more like 10 percent. (Yes, I’ve been practicing medicine for a long time, and we’re talking about inpatients at a tertiary care hospital, not a primary care clinic. I can imagine the fraction being somewhat higher in other settings.)

What were my off-label “doorman” tasks? It was my end-of-life discussions with three children of an elderly patient with terminal cancer, all of whom had different views on what to do with their dad. It was figuring out whether another patient, who needed very high flow oxygen for his severe emphysema, could get a similar oxygen setup at his home. It was deciding whether a patient with an ocular vasculitis and a positive blood test for syphilis had both diseases… and if so, whether he needed two weeks of intravenous penicillin… and, if he did, where he was going to get his antibiotics because he lacks health insurance.

Could all these tasks be orchestrated by AI, or by AI-aided workers who are less expensive than doctors? Eventually, I guess, but not now.

The Economy Class Straw Man

In my book, I argued that, as AI gets better, primary care in the U.S. is likely to divide into two classes. If you have basic health insurance, you might get AI-first care, with AI managing your weight, cholesterol, blood pressure, and the like. Of course, any functioning AI-first system will require a trustworthy triage protocol that identifies patients who really need to see a doctor because of the complexity of their problem or the need for a procedure. And yes, I likened this AI-first system to economy class.

 

For patients with more complex needs or who can afford more expensive health insurance, I believe that business class will not be a physician operating without AI (which, in a few years, will be a violation of the standard of care), but rather an AI-enabled clinician. If that doctor is any good, he or she will know when to trust the AI and when not to. (And future AI tools are likely to signal their level of uncertainty, prompting the doctor to remain in the loop in situations of considerable uncertainty, while discouraging human intervention when the AI is sure of its answers.)

Nonetheless, even in situations where a doctor adds little value as a diagnostician, they will continue to add value by communicating, listening, being empathic, and dealing with particularly complex cases that don’t fit neatly into an algorithm. I’m guessing that people who can afford this kind of system will prefer it, though it will surely be more expensive than “economy class.” And younger folks may well have very different preferences about interacting with a human doctor (or a human anything) than people of my generation do.

Another useful analogy comes from taxes or travel agents. If your tax needs are straightforward or your budget is tight, you might use TurboTax, whereas if your needs are more complex and you can afford it, you might see a human accountant. Ditto travel, where people still periodically engage travel agents for particularly complex itineraries, but mostly use digitally-enabled self-service. It seems likely that medicine will follow a version of this playbook.

And, Not Or

In the end, the JAMA article elevates a key and surprising message: we are reaching the point at which the human in the loop might, at times, degrade the performance of an AI-based system, at least in making a diagnosis and recommending therapy based on a reliable set of facts.

As AI improves as a diagnostic tool, the question of whether we need to change the way we train doctors inevitably arises. As we consider this, we should be skeptical of arguments that posit AI’s superiority over physicians based on scenarios that imperfectly reflect the realities of medicine, as well as arguments that claim that taking over a portion of doctors’ cognitive work means we won’t need doctors at all.

I can imagine a system in which basic preventive and urgent care is provided quite competently – and less expensively and more conveniently – by AI, whereas more complex care, both clinical and emotional, is provided by humans. The greatest challenge for healthcare AI in the coming years won’t lie in solving AI’s technical problems. It will lie in sorting through these scenarios, parsing what should be done by AI and what should be done by humans, and then rebuilding a healthcare system that feels holistic rather than fragmented.

Previous
Previous

1 big number: AI at the doctor's office

Next
Next

Will Autonomous AI Exceed AI-Aided Physicians as the Best Medical Care?