Session

Evaluating and Monitoring Ambient Scribe Systems in Danish Healthcare

Organized by:

Thursday 8 October 11.00

Lead organizer: Gustav Aarup Lauridsen, Digitalisering & IT, Region Midtjylland

Ambient scribe systems use speech recognition and large language models to generate clinical documentation from clinician-patient conversations and are rapidly entering Danish healthcare. But how do we prevent our healthcare system from being flooded with AI-slop?

Evaluation of these systems presents several open research challenges: clinical summaries have no single ground truth, documentation varies across specialties, rare errors may have serious consequences, and deployed systems must be monitored for performance degradation and drift.

In this workshop, we bring together data scientists, researchers, clinicians, and health innovation experts to explore these challenges. Through introductions to real world applications in Region Midtjylland and guided group discussions, participants will work on three questions: How should we evaluate and govern ambient scribes? How can Danish medical speech and summary data support robust evaluation? And how should we monitor a system deployed to thousands of clinicians?

We aim to identify open research questions, share ideas and methods, and create opportunities for collaboration between young researchers and the Danish public healthcare sector.

Programme

The workshop will start with an introductory presentation, followed by guided group work: 

  • 15 minutes Introduction by Thea Rolskov and Gustav Lauridsen, LLM specialists from Region Midtjylland 
  • 10 minutes Clinical  perspectives by Erik Perfalk, Medical doctor and Postdoc at Department of Affective disorders, AUH 
  • 10 minutes perspectives from health innovation by Morten Charles, Medical doctor and Clinical Professor at Aarhus University 
  • 10 minutes group discussions for scenario 1: In groups, discuss the evaluation and governance requirements for ambient scribe systems 
  • 5 minutes follow-up/discussion 
  • 10 minutes group discussions scenario 2: In groups, discuss how Danish medical speech and summary data can be collected and used to evaluate ambient scribe systems 
  • 5 minutes follow-up/discussion 
  • 10 minutes group discussions scenario 3: In groups, discuss how to monitor the live performance of an ambient scribe system deployed to 10,000 clinicians 
  • 5 minutes follow-up/discussion  
  • 10 minutes discussion of open research questions, collaboration opportunities, and future directions for evaluation of clinical AI systems
Speakers’ list
  • Gustav Aarup Lauridsen, AI-specialist, Region Midtjylland: Ambient scribe evaluation for public healthcare

    Thea Rolskov Sloth, AI-specialist, Region Midtjylland: Ambient scribe evaluation for public healthcare

    Erik Perfalk, Medical Doctor, Postdoc, Aarhus University Hospital

    Morten Haaning Charles, Medical Doctor, Clinical Professor, Aarhus University

Level

Intermediate: For attendees who have basic understanding or some experience with the subject but are not yet advanced.

Organizers
  • Gustav Aarup Lauridsen, Digitalisering & IT, Region Midtjylland (gulaur@rm.dk) (lead)
  • Thea Rolskov Sloth, Digitalisering & IT, Region Midtjylland (theslo@rm.dk)