We asked GPT-4.1 and DeepSeek-V3 to process the same clinical cases hundreds of times. Across 12,197 outputs: rampant guideline omissions, hallucinated references, and wildly inconsistent results.
New paper in @bmj.com #Health & #Care #Informatics: doi.org/10.1136/bmjh...