New in JMIR Nursing: AI-Based Structured Information Extraction From Synthetic Nursing Handover Transcripts: Comparative Evaluation of Large Language Models #ArtificialIntelligence #Nursing #Healthcare #ClinicalHandover #NursingTranscripts
dlvr.it
AI-Based Structured Information Extraction From Synthetic Nursing Handover Transcripts: Comparative Evaluation of Large Language Models
Background: Clinical handover is the process during which responsibility and accountability for care are transferred between clinicians. AI has the potential to improve the reliability and completeness of clinical handover by helping clinicians detect predefined content areas that have been communicated, identify explicit information gaps, and prompt clarification before responsibility is transferred. Objective: This study evaluated the performance of several large language models and prompt optimization strategies for structured information extraction of synthetic #nursing handover transcripts. Methods: Two registered #nurses independently annotated a dataset of 203 synthetic handover transcripts to produce consensus labels for information extraction tasks. Tasks included (1) labeling spans of text into SBAR (Situation, Background, Assessment, Recommendation) categories, (2) content detection to determine if specific pieces of information were communicated, and (3) labeling spans of text that communicated information using uncertain terms that included a subtask for identifying unknown facts. Baseline and Genetic-Pareto (GEPA)–optimized prompts were compared for the GPT-5.2, GPT-5-nano, and MedGemma 27B large language models. Additionally, the LangExtract framework was evaluated for span-extraction tasks. Results: The GPT-5.2–optimized model achieved a micro-score of 0.85 (95% CI 0.83‐0.88) for content detection, an absolute improvement of +0.08 compared with the matched baseline. GPT-5-nano also performed better after optimization for content detection (micro-score 0.81, 95% CI 0.78‐0.84), suggesting that this structured task was not limited to the highest-capacity model. For SBAR span extraction, GPT-5.2 with prompt optimization achieved a micro-score of 0.76 (95% CI 0.72‐0.79), improving by +0.24 compared with baseline and exceeding LangExtract; GPT-5-nano also improved to a micro-score of 0.69 (95% CI 0.66‐0.72). Broad uncertainty-span extraction remained comparatively weak despite prompt optimization (micro-score 0.41, 95% CI 0.33‐0.48; absolute improvement +0.06). In contrast, explicit unknown-fact extraction was more accurate with GPT-5.2 (micro-score 0.84, 95% CI 0.63‐1.00), GPT-5-nano (micro-score, 0.84 95% CI 0.63‐1.00), and MedGemma 27B (micro-score 0.80, 95% CI 0.63‐1.00). Genetic-Pareto–optimized prompts outperformed the LangExtract approach across each span-extraction task. Conclusions: Prompt optimization improved matched-model point estimates, with the highest performance observed for predefined content detection and SBAR span extraction. Broad uncertainty extraction remained less accurate than the narrower unknown-fact task. These technical results do not establish clinical effectiveness, safety, or readiness for real-time use. Validation using authentic #nursing handover communication and prospective evaluation in clinical workflows are required before clinical application.