Introducing Answered with Evidence - a framework for evaluating whether LLM responses to biomedical questions are founded in evidence

Julian D Baldwin
Christina Dinh
Arjun Mukerji
Neil Sanghavi
Saurabh Gombar

Read the full article

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

The growing use of large language models (LLMs) for biomedical question answering raises concerns about the accuracy and evidentiary support of their responses. To address this, we present Answered with Evidence , a framework for evaluating whether LLM-generated answers are grounded in scientific literature. We analyzed thousands of physician-submitted questions using a comparative pipeline that included: (1) Atropos Alexandria Evidence Library, a retrieval-augmented generation (RAG) system based on novel observational studies, and (2) two PubMed-based retrieval-augmented systems (System and Perplexity). We found that PubMed-based systems provided evidence-supported answers for approximately 44% of questions, while the novel evidence source did so for about 50%. Combined, these sources enabled reliable answers to over 70% of biomedical queries. As LLMs become increasingly capable of summarizing scientific content, maximizing their value will require systems that can accurately retrieve both published and custom-generated evidence or generate such evidence in real time.

Version published to 10.1101/2025.07.01.25330655 on medRxiv
Jul 2, 2025

SKiM-GPT: Combining Biomedical Literature-Based Discovery with Large Language Model Hypothesis Evaluation

This article has 9 authors:
1. Jack Freeman
2. Robert J. Millikin
3. Leo Xu
4. Ishaan Sharma
5. Bethany Moore
6. Cannon Lock
7. Kevin Shine George
8. Aviral Bal
9. Ron Stewart
This article has no evaluationsLatest version Jul 31, 2025
CLEVER: Clinical Large Language Model Evaluationby Expert Review

This article has 4 authors:
1. Veysel Kocaman
2. Mustafa Kaya
3. Andrei Ferrer
4. David Talby
This article has no evaluationsLatest version Jul 23, 2025
Federated Knowledge Retrieval Elevates Large Language Model Performance on Biomedical Benchmarks

This article has 2 authors:
1. Janet Joy
2. Andrew I. Su
This article has no evaluationsLatest version Aug 2, 2025

Listed in

Abstract

Article activity feed

Related articles

SKiM-GPT: Combining Biomedical Literature-Based Discovery with Large Language Model Hypothesis Evaluation

CLEVER: Clinical Large Language Model Evaluationby Expert Review

Federated Knowledge Retrieval Elevates Large Language Model Performance on Biomedical Benchmarks