Session
The importance (and struggle) of real-world data
Organized by:
Thursday 8 October 15.15
Lead organizer: Sarah Wordenskjold Stougaard, PhD student, Odense University Hospital
“Still… cleaning data?” Making real-world data usable can take months, yet much of that work becomes invisible in the final study. Decisions about what to trust, what to fix, and when to stop shape everything that follows.
In this workshop, we discuss questions that remain difficult even with experience: When are data good enough? What should be automated, and what requires domain knowledge? How can this work become transparent, reusable, and better recognised as science?
Three health data cases will start the discussion. Participants will then reflect on one of their own data challenges and discuss in small groups. Whether you are facing your first real-world dataset or have years of hard-earned lessons, your perspective matters. We will not solve every problem, but you should leave with ideas from others’ dos and don’ts and reassurance that you are not the only one still cleaning data.
The 90-minute session will follow a workshop-style format, combining short case presentations with interactive activities to foster active participation and knowledge exchange among participants.
Introduction to session (5 minutes)
Brief framing of the topic, including why real-world data collection, curation, and quality are important but often underestimated parts of data science.
“Identify your own struggle” (15 minutes)
Participants individually write down one real-world data challenge from their own work, using a simple worksheet prepared by the organisers. The worksheet will guide them through questions such as: what type of data they work with, a challenge they have encountered/expect to encounter, why it is difficult to solve, and if they have any ideas what could help them solve (part) of the issue.
After filling out the worksheet, participants will briefly discuss their challenge in pairs or around the table to identify common themes. During this activity, the organisers will move between tables and identify recurring categories of challenges, which will later be used to form discussion groups.
Short case presentations (30 minutes)
Three 10-minute case presentations will provide concrete examples of real-world data challenges and lessons learned. Working titles are:
–
The presentations will show what it can take to make real-world data reliable and usable in practice. This will include examples of challenges, partial solutions, and lessons learned, while also highlighting the substantial workload that lies behind the final datasets, models, and results.
“Revisit your struggle” and choose a category (5 minutes)
After the presentations, participants will have opportunity to revisit their worksheet and adjust or expand their initial challenge if the presentations have given them new ideas. The organisers will then present the main challenge categories identified during the first exercise, and participants will choose the category that best fits their own challenge or interest. Each group will ideally include a maximum of 5 participants, so larger categories may be split into several smaller groups.
Group discussions (25 minutes)
Participants will work in small groups based on their selected challenge category. Each group will discuss shared experiences within their theme, supported by guiding questions shown on slides. These may include:
–
The groups are not expected to solve individual problems completely. Instead, the aim is to identify transferable lessons, possible strategies, and areas where further collaboration or knowledge sharing would be useful. The organisers will facilitate the discussions by moving between groups and helping keep the discussions concrete and constructive.
Wrap-up and networking link (10 minutes)
Each group will briefly share one key challenge and one possible strategy or lesson learned. The organisers will provide the opportunity for an (optional) shared list of participants to connect on LinkedIn to continue discussions after the conference.
Sarah Stougaard, PhD student, Department of Oncology – Laboratory of Radiation Physics, Odense University Hospital:
Collecting and curating 20 years of radiotherapy data
Andreas Fuglsang, PhD student, Department of Oncology – Laboratory of Radiation Physics, Odense University Hospital:
Cleaning (and understanding) human-labelled data
Alejandro Cortina Uribe, PhD student, Rigshospitalet / University of Copenhagen:
Working with hundreds of thousands medical images, but wait, are they still in the hospital server?
Intermediate: For attendees who have basic understanding or some experience with the subject but are not yet advanced.