Medical AI Stress Testing

Stress Testing Medical AI: Could Generative AI Help Build Safer and Fairer Healthcare Systems?

CERSI-AI NEWS, UPDATES & THOUGHT LEADERSHIP

Month: August 2026

Based on an interview with Professor Ben Glocker – Imperial College London

August 2026

Stress Testing Medical AI: Could Generative AI Help Build Safer and Fairer Healthcare Systems?

Researchers are exploring how generative AI and synthetic medical data could help uncover weaknesses in healthcare AI systems before these systems reach patients. Professor Ben Glocker explains how CERSI-AI research is developing new ways to stress test medical AI.

Artificial intelligence is becoming increasingly capable of analysing medical images, identifying signs of disease and supporting clinical decision-making.

But an AI system performing well overall does not necessarily mean it will perform equally well for every patient.

Differences in patient populations, medical equipment and the data used to train an AI model can all affect its performance. This raises an important question for developers, clinicians and regulators: how can we identify these weaknesses before an AI system is deployed in healthcare?

Research conducted as part of CERSI-AI is exploring one potential answer: stress testing AI using generative models and synthetic medical data.

Professor Ben Glocker, Professor in Machine Learning for Imaging at Imperial College London, leads research at the intersection of artificial intelligence and healthcare, with a particular focus on developing safer AI for medical imaging.

As part of the CERSI-AI research project, his team has been investigating methods for identifying potential blind spots and disparities in medical AI systems.

Why does medical AI need stress testing?

AI models learn from data.

When the data they encounter in the real world differs from the information on which they were trained, their performance can deteriorate.

This could relate to differences in medical scanners, patient demographics or characteristics that are poorly represented within the original training data.

The challenge is particularly important when AI is being used to support clinical decisions.

The CERSI-AI research therefore looks at whether an AI model has specific weak spots or performs differently across subgroups of the population.

The aim is not simply to measure whether an AI system works, but to understand where, when and for whom it might not work as expected.

Using generative AI to ask ‘what if?’

One of the most interesting elements of the research is the use of generative AI to create realistic synthetic medical images.

Consider a chest X-ray.

An AI system might analyse the image and attempt to identify signs of disease. If it makes an incorrect prediction, researchers naturally want to understand why.

Generative AI provides a way of investigating that question.

Researchers can take characteristics of an existing medical image and simulate realistic changes. They might ask what the X-ray would have looked like if it had been captured using a different scanner, or how certain characteristics could change if the patient were older.

The model can then be tested against these simulated scenarios.

These are effectively counterfactual questions: what would the AI have predicted if something about the input had been different?

Researchers can investigate changes including:

  • the scanner or imaging equipment used
  • patient age
  • patient sex
  • other demographic characteristics
  • characteristics of the medical image itself

If changing one of these factors causes the AI’s prediction to change unexpectedly, researchers may have identified a weakness that requires further investigation.

Stress testing AI for breast cancer screening

The research is also being explored in mammography.

AI systems can perform well when detecting early signs of breast cancer, but performance may not necessarily be consistent across every patient group.

Breast tissue density provides a useful example.

Higher breast density can make mammograms more difficult to interpret because dense tissue can obscure signs associated with cancer.

Using generative AI, researchers can synthesise mammograms representing different levels of breast tissue density and examine how an AI model responds.

This allows the team to investigate the reliability and robustness of the system across scenarios that may otherwise be difficult to evaluate systematically.

Could this help create fairer healthcare AI?

A major ambition for AI stress testing is to identify potential disparities before a system encounters real patients.

Rather than discovering weaknesses after deployment, developers could audit models earlier in the development process.

Stress testing could also become continuous.

When a new version of an AI system is developed, it could be tested again. If the system is introduced into a new hospital with a different patient population, it could be evaluated for that specific environment.

This creates an iterative approach to AI development in which testing, identifying weaknesses and improving systems become part of an ongoing cycle.

The long-term goal is to help develop AI systems that perform as consistently as possible across diverse patient populations.

However, Professor Glocker stresses that completely eliminating bias is unlikely to be realistic.

Datasets used to develop AI will contain biases. The important challenge is therefore identifying and understanding those biases so appropriate safeguards and mitigations can be introduced.

AI safety assurance cannot rely entirely on computational tools either. Human review, multidisciplinary teams and appropriate auditing frameworks remain essential.

The challenge of creating realistic synthetic data

Synthetic medical data introduces its own important questions.

A generative model can only create realistic simulations if the assumptions underpinning it accurately reflect the real world.

Researchers therefore need to make assumptions about cause-and-effect relationships within healthcare data.

Those assumptions might concern relationships between demographic characteristics, medical information, disease prevalence and what ultimately appears within a medical image.

But these relationships are not always completely understood.

Researchers must draw upon published literature and, crucially, work with clinicians, patients and other domain experts.

This is both a strength and a limitation of the approach.

Embedding domain knowledge into the model can make simulations more meaningful, but incorrect assumptions could also lead researchers towards incorrect conclusions.

Every assumption therefore needs to be examined carefully.

Clinicians and patients must be part of the process

Developing trustworthy medical AI cannot be treated purely as a computer science problem.

Clinicians are involved in the co-design of the algorithms and in reviewing the assumptions being incorporated into the models.

This can involve considerable discussion and disagreement before researchers reach a consensus.

Even then, some uncertainty may remain.

Patient involvement is also essential, including through communities of practice and workshops examining questions such as the ethical implications of using AI within healthcare.

This multidisciplinary approach becomes increasingly important as these technologies move from research environments towards potential clinical and regulatory applications.

Building more realistic medical images

The technology itself has also developed significantly since the team’s original research.

Earlier versions of the model could synthesise medical images, but at relatively low resolutions.

Resolution matters enormously in medical imaging. Chest X-rays and mammograms can require high-resolution images because signs of disease may be extremely small or subtle.

The team has since scaled the technology to generate much higher-resolution images.

Clinical experts have also taken part in a user study comparing real images with synthetic images produced by the model. According to Professor Glocker, the results demonstrated that participants were unable to reliably distinguish the generated images from real examples.

The training data has expanded considerably too.

For the chest X-ray application, the researchers are now working with more than one million images, drawn from seven different sources across five countries.

This larger and more heterogeneous dataset is intended to help the model generate images that better represent diverse populations.

Developing these models is computationally demanding, requiring the researchers to use national high-performance computing infrastructure.

Could stress testing become part of AI regulation?

The next major step could be taking this technology beyond research and into real-world AI evaluation.

Professor Glocker’s vision is that generative AI stress testing could eventually complement the evidence currently used when evaluating medical AI systems.

Today, developers typically need to demonstrate that an AI system works on external datasets and can generalise beyond the information on which it was originally developed.

Stress testing could add another layer.

Instead of only asking whether the model performs well against existing data, developers and regulators could investigate how it responds across a much broader range of simulated circumstances.

This could provide additional evidence about the safety, reliability and limitations of medical AI before deployment.

Could synthetic data change clinical trials?

The potential applications extend beyond testing AI systems.

In the longer term, synthetic data could also complement real-world information used within clinical trials.

Clinical trials can be costly and time-consuming, particularly when researchers need to recruit participants representing specific patient populations.

Some groups may also be underrepresented within existing datasets.

Synthetic data could potentially help researchers explore these gaps, making studies more representative while complementing information gathered from real patients.

Crucially, this does not mean replacing real-world evaluation.

Instead, synthetic data could provide an additional source of evidence alongside conventional clinical research.

The next stage requires collaboration

Although much of the research to date has been computational, the next stage cannot happen within computer science alone.

Moving stress testing towards regulatory and clinical use will require collaboration between AI researchers, clinicians, patients, regulators and specialists from across different disciplines.

The central question now shifts from whether researchers can generate convincing synthetic medical images to whether those technologies can provide meaningful evidence that clinicians, patients and regulators can trust.

For Professor Glocker, breaking down the boundaries between disciplines will therefore be essential.

If AI is going to play an increasingly important role in healthcare, understanding where these systems fail may ultimately prove just as important as demonstrating where they succeed.

CERSI-AI is supporting research exploring how emerging technologies can be evaluated safely and effectively as they move towards real-world healthcare applications.

Professor Ben Glocker is a Professor in Machine Learning for Imaging and Kheiron Medical / RAEng Research Chair in Safe Deployment of Medical Imaging AI. He co-leads the Biomedical Image Analysis Group, leads the HeartFlow-Imperial Research Team and the Knowledge Transfer Lead of CHAI – The EPSRC Causality in Healthcare AI Hub.

Join the Conversation

Follow us at CERSI-AI as we work to bridge the gap between extraordinary technologies and patient benefit.

Connect on:


Europe needs stronger Regulatory Science

Europe needs stronger regulatory science for medical AI and MedTech

CERSI-AI NEWS & UPDATES

Month: August 2026

A Schönfelder, AK Denniston, AF Frangi, C Johner, T Minssen, K Singh, S Gilbert
6 August 2026

June 2026

Europe needs stronger regulatory science for medical AI and MedTech

In a new comment published in the journal “Nature Biomedical Engineering”, researchers from the Else Kröner Fresenius Center (EKFZ) for Digital Health at TUD Dresden University of Technology, in collaboration with international researchers, argue that Europe should establish Centres of Excellence in Regulatory Science and Innovation (CERSIs). These academic-led centres could help ensure that regulation keeps pace with rapidly developing medical technologies and artificial intelligence.

Researchers identify structural weaknesses in European regulation

Medical AI and advanced medical devices are developing faster than traditional regulatory processes were designed to handle. Many of today’s technologies no longer fit neatly into fixed categories: AI tools adapt over time, work across different clinical settings, or combine several functions that were previously performed by separate devices. The traditional regulation of products for a single intended purpose is becoming increasingly outdated. This creates a regulatory challenge for medical innovation: Regulation must ensure that medical technologies are safe and effective for patients. At the same time, it must be able to support responsible innovation. In a new publication, the international team of authors argues that Europe’s current regulatory system is too fragmented to meet this challenge. Responsibilities are distributed across multiple actors, including the European Commission, the Medical Device Coordination Group, national competent authorities, notified bodies and expert panels. According to the authors, the absence of a central structure responsible for regulatory strategy can make assessment pathways difficult to predict and slow to adapt, particularly for AI-enabled medical devices. Also, research on new regulatory methods, evidence standards, and emerging risks is not yet coordinated at the scale needed. As a result, innovators may face unclear pathways, regulators may lack timely scientific input, and patients may wait longer for safe and useful technologies. “What we need is a stronger scientific basis for regulation of medical devices: independent centers that can identify emerging risks early, develop evidence-based methods, and help Europe respond proactively to technological change”, says Anett Schönfelder, first author of the article and researcher at the Else Kröner Fresenius Center (EKFZ) for Digital Health at TU Dresden.

A practical model for future-ready regulation

The authors propose establishing European “Centres of Excellence in Regulatory Science and Innovation” (CERSIs). These centres would not replace regulators or make binding decisions. Instead, they could provide independent expertise, generate evidence early, monitor emerging developments, offer training, and develop practical tools to support better regulatory decisions for medical AI, digital health technologies, medical devices, and in vitro diagnostics.

The authors argue that Europe can learn from the United States, where the U.S. Food and Drug Administration established CERSIs more than a decade ago, and from the United Kingdom, which recently launched its own CERSI program. These centres bring together regulators, academic researchers, industry and other stakeholders to work on priority questions in regulatory science. Their work can include developing new evaluation methods, supporting guidance documents, improving post-market surveillance, and identifying emerging challenges such as generative AI in diagnostics, synthetic data or continuous monitoring tools.

“Regulation must ensure patient safety without slowing down responsible innovation. To achieve this, Europe needs sustained scientific capacity in regulatory science rather than relying only on individual projects or short-term consultations. European Centres of Excellence in Regulatory Science and Innovation would connect academic expertise, regulatory needs, and technological development. They would provide the research capacity and long-term knowledge base needed to make medical AI and MedTech regulation more adaptive, predictable, and fit for purpose”, says Prof. Stephen Gilbert, Professor of Medical Device Regulatory Science at the EKFZ for Digital Health at TU Dresden.

To have a practical impact, European CERSIs would need clearly defined links to EU institutions and regulatory processes. Otherwise, their findings might remain scientifically valuable without being translated into guidance and regulatory practice. The centres would not resolve all structural challenges in European medical-device regulation, the authors acknowledge. However, they would provide an important scientific foundation for regulation that protects patients while enabling responsible innovation.

A Schönfelder, AK Denniston, AF Frangi, C Johner, T Minssen, K Singh, S Gilbert: EU Innovation needs Regulatory Science Excellence Centres, Nature Biomedical Engineering, 2026, doi: 10.1038/s41551-026-01756-x.

Link to paper: https://www.nature.com/articles/s41551-026-01756-x

Join the Conversation

Follow us at CERSI-AI as we work to bridge the gap between extraordinary technologies and patient benefit.

Connect on:


Privacy Preference Center