From GenAI Virtual Patient Dialogue Logs to Teacher-Interpretable Process Evidence: A Learning Analytics Study in Higher Education
GenAI-powered virtual patients (VPs) enable scalable medical history-taking practice while preserving full turn-by-turn dialogue logs for analysis A three-layer analytic framework (behavioral prevalence, Epistemic Network Analysis, Transition Network Analysis) transforms raw dialogue logs into teacher-interpretable process evidence High-rated consultations differ not in volume but in strategic connectivity: linking information gathering with communication, checking, organization, and synthesis S
Analysis
TL;DR
- GenAI-powered virtual patients (VPs) enable scalable medical history-taking practice while preserving full turn-by-turn dialogue logs for analysis
- A three-layer analytic framework (behavioral prevalence, Epistemic Network Analysis, Transition Network Analysis) transforms raw dialogue logs into teacher-interpretable process evidence
- High-rated consultations differ not in volume but in strategic connectivity: linking information gathering with communication, checking, organization, and synthesis
- Summarizing and organizing moves in high-performing learners more frequently lead to verification or mechanism-oriented follow-up questions
- Layered analysis of GenAI VP dialogues can support process-focused feedback in medical education, bridging the gap between raw transcripts and final scores
Why It Matters
This research addresses a critical bottleneck in AI-enhanced medical education: the inability to translate rich dialogue data into actionable pedagogical insights. For AI practitioners building educational tools, it demonstrates a validated pipeline for converting unstructured conversational data into interpretable learning analytics. For educators and researchers, it establishes that process-level patterns—not just outcome scores—can be systematically extracted and used to improve clinical reasoning instruction.
Technical Details
- Dataset: 1,030 GenAI VP dialogues from 210 second-year medical learners across five weeks of chest-pain case consultations, each teacher-scored using a rubric assessing the full history-taking dialogue
- Classification approach: Consultations were classified as high- or low-rated within each week using the weekly median score as the threshold
- Three analytic layers applied: (1) Behavioral prevalence to measure activity frequency, (2) Epistemic Network Analysis (ENA) for local co-occurrence of coded dialogue behaviors, and (3) Transition Network Analysis (TNA) for sequential patterns between dialogue moves
- Key finding: High-rated dialogues showed stronger connections between information-gathering/symptom-exploration behaviors and communication, checking, organization, and synthesis behaviors
- Sequential insight: Summarizing and organizing moves in high performers more often transitioned to verification or mechanism-oriented follow-up, revealing a strategic reasoning pattern
Industry Insight
- AI-powered virtual patient platforms should prioritize built-in analytic pipelines that go beyond scoring to surface process evidence, enabling formative feedback at scale
- The three-layer analytic framework (prevalence + ENA + TNA) offers a reusable template for transforming dialogue logs from any conversational AI tutor into interpretable learning analytics
- Medical education programs adopting GenAI VPs should invest in teacher training to interpret process-level evidence, as raw transcripts remain impractical for routine review while scores alone obscure reasoning quality
Disclaimer: The above content is generated by AI and is for reference only.