STAT+: Why benchmarking clinical LLMs from OpenEvidence, Doximity is complicated

STAT+: Why benchmarking clinical LLMs from OpenEvidence, Doximity is complicated

In this edition of AI Prognosis: A conversation about benchmarking leading clinical chatbots, investor view on AI in biopharma, and more.

You’re reading the web edition of STAT’s AI Prognosis newsletter, our subscriber-exclusive guide to artificial intelligence in health care and medicine. Sign up to get it delivered in your inbox every Wednesday. 

I saw “The Odyssey” during its opening weekend. Ever since then, I have been questioning whether I’m illiterate or whether Christopher Nolan is a poor storyteller. This London Review of Books evaluation of the film, written by the woman whose translation of “The Odyssey” Nolan apparently read, has freed me from my wondering. (h/t to my colleague Matthew Herper)

Hot takes on Homer’s epic, or hot tips about Epic Systems: aiprognosis@statnews.com

Benchmark battle bots

You might recall that in mid-June, there was a Nature Medicine study that pitted clinical AI systems OpenEvidence and UpToDate Expert AI against general LLMs. It set off a reaction in the clinical AI world like no other paper has. “The results rang out like a gunshot,” as STAT health tech correspondent Katie Palmer describes it.

The controversy surrounding the study, and everything that came after, exemplifies the problems I have with benchmarks.

Katie summed it up well when I talked to her yesterday: “The way that benchmarks have been talked about generally, and specifically in clinical AI, tends to summarize them into the headlines,” she said. “Every study needs a headline and every story needs a headline, but as we both know, and as I think most people in the industry know, an individual benchmark doesn’t mean much.”

Continue to STAT+ to read the full story…

Lire l’article complet sur le site source
STAT News - Sante & Medecine (EN)




Besoin d'un service ? Discutez maintenant !
🤖

Assistant eFastWork

En ligne

📩 Envoyer un message à notre équipe


ou envoyez un message vocal
👋

Bienvenue !

Entrez votre email pour commencer à chatter.

Email utilisé uniquement pour vous répondre.