As LLMs compete for doctors' attention, some developers say the science of benchmarking their AI tools for safety and accuracy is flawed.
Hundreds of thousands of U.S. doctors use clinical large language models, pitched by companies like OpenEvidence, Doximity, and UpToDate as an antidote to the dangers of hallucination-prone generalist models from Big Tech. Yet few studies have pitted them against each other — and this summer brought a high-profile head-to-head.
Researchers from NYU Langone Health had tested general and clinical models, including OpenEvidence and UpToDate Expert AI, on three sets of clinical questions. The findings, published in Nature Medicine in June: The clinical AI performed worse than the general models.
The results rang out like a gunshot. “I’ve never seen a single paper trigger the kind of reactions this one has in the health AI community,” wrote Kaiser Permanente’s vice president of AI and emerging technologies on LinkedIn. The paper’s findings, like all science, are subject to interpretation and debate — but many online reactions treated them more like a clear victory for general frontier models.