Session
Evaluating and Monitoring Ambient Scribe Systems in Danish Healthcare
Organized by:
Thursday 8 October 11.00
Lead organizer: Gustav Aarup Lauridsen, Digitalisering & IT, Region Midtjylland
Ambient scribe systems use speech recognition and large language models to generate clinical documentation from clinician-patient conversations and are rapidly entering Danish healthcare. But how do we prevent our healthcare system from being flooded with AI-slop?
Evaluation of these systems presents several open research challenges: clinical summaries have no single ground truth, documentation varies across specialties, rare errors may have serious consequences, and deployed systems must be monitored for performance degradation and drift.
In this workshop, we bring together data scientists, researchers, clinicians, and health innovation experts to explore these challenges. Through introductions to real world applications in Region Midtjylland and guided group discussions, participants will work on three questions: How should we evaluate and govern ambient scribes? How can Danish medical speech and summary data support robust evaluation? And how should we monitor a system deployed to thousands of clinicians?
We aim to identify open research questions, share ideas and methods, and create opportunities for collaboration between young researchers and the Danish public healthcare sector.
The workshop will start with an introductory presentation, followed by guided group work:
Introduction (15 minutes)
by Thea Rolskov and Gustav Lauridsen, LLM specialists from Region Midtjylland
Clinical perspectives (10 minutes)
by Erik Perfalk, Medical doctor and Postdoc at Department of Affective disorders, AUH
Perspectives from health innovation (10 minutes)
by Morten Charles, Medical doctor and Clinical Professor at Aarhus University
Group discussions for scenario 1 (10 minutes)
In groups, discuss the evaluation and governance requirements for ambient scribe systems
Follow-up/discussion (5 minutes)
Group discussions scenario 2 (10 minutes)
In groups, discuss how Danish medical speech and summary data can be collected and used to evaluate ambient scribe systems
Follow-up/discussion (5 minutes)
Group discussions scenario 3 (10 minutes)
In groups, discuss how to monitor the live performance of an ambient scribe system deployed to 10,000 clinicians
Follow-up/discussion (5 minutes)
Discussion of open research questions, collaboration opportunities, and future directions for evaluation of clinical AI systems (10 minutes)
Gustav Aarup Lauridsen, AI-specialist, Region Midtjylland:
Ambient scribe evaluation for public healthcare
Thea Rolskov Sloth, AI-specialist, Region Midtjylland:
Ambient scribe evaluation for public healthcare
Erik Perfalk, Medical Doctor, Postdoc, Aarhus University Hospital
Morten Haaning Charles, Medical Doctor, Clinical Professor, Aarhus University
Intermediate: For attendees who have basic understanding or some experience with the subject but are not yet advanced.