Reposted by Robert Leaman
From @robertleaman.bsky.social et al in #OUP 's #DATABASE journal | MedHopQA: a disease-centred multi-hop reasoning benchmark and evaluation framework for LLM-based biomedical question answering | #OpenScience #LLM #Wikipedia #Mondo #NCBIGene #NCBITaxonomy |🧬🖥️🧪🔓
⬇️
academic.oup.com/database/art...
academic.oup.com
MedHopQA: a disease-centred multi-hop reasoning benchmark and evaluation framework for LLM-based biomedical question answering
Abstract. Evaluating large language models (LLMs) in the biomedical domain requires benchmarks that can distinguish genuine reasoning from pattern matching