Healthcare Professional Identities in the AI Age Post

Healthcare Professional Identities in the AI Age

CERSI-AI THOUGHT LEADERSHIP

Month: June 2026

By Dr Susan Shelmerdine and Dr Christina Malamateniou – June 2026

June 2026

 AI regulation in healthcare has focused on product safety, algorithmic performance, and data governance. This submission argues that an under-addressed gap exists around workforce readiness - the capacity of healthcare professionals to use AI tools safely, critically, and sustainably over time, and to navigate the changing expectations of patients who are themselves increasingly using AI.

Background

This submission draws on evidence and professional testimony from two multidisciplinary events held in March 2026 to provide a commentary on the topic of changing professional identities in the AI age for the National Commission into the Regulation of AI in Healthcare, and recommendations for review.

The first was a panel discussion convened at King’s College London (KCL) on 20 March 2026 under the title “Professional Identities in the AI Age,” bringing together a panel of clinical practitioners with industry experience, academic researchers, an AI fellow and radiology trainee embedded in NHS deployment, a former president of the Royal College of Radiologists (RCR), and a behavioural scientist specialising in human-AI interaction. The event was attended by 30 guests with background in engineering, computer science, mathematics, social sciences and different subspecialties within medicine (including endocrinology, neurology, anaesthetics).

Chair & Panellists: Dr Susan Shelmerdine, Dr Christina Malamateniou, Dr Katharine Halliday, Dr Jacqueline Matthews, Dr Girija Agarwal, Dr Fendi Tsim

The second event was a formal debate held at the Royal Society of Medicine (RSM) under the motion “This house believes that AI will radically change the role of a doctor,” featuring speakers from general practice, radiology, cardiovascular medicine, and surgery, with contributions from a 100+ strong audience of clinicians, medical educators, technologists, and patient advocates.

Chair & Panellists: Professor Simon Wessely, Dr Annabelle Painter, Professor Paul Leeson, Professor Shafi Ahmed, Professor Erika Denton

Executive Summary

AI regulation in healthcare has focused on product safety, algorithmic performance, and data governance. This submission argues that an under-addressed gap exists around workforce readiness – the capacity of healthcare professionals to use AI tools safely, critically, and sustainably over time, and to navigate the changing expectations of patients who are themselves increasingly using AI.

The findings in this submission are not arguments against AI adoption in healthcare; they are arguments that adoption at pace requires investment in human capital and technological infrastructure that makes adoption safe. Without that investment, AI tools will continue to be deployed into environments that cannot absorb them, producing poor outcomes and professional resistance that slow adoption far more than regulation does.

Evidence from the events cited, supported by published research referenced by contributors, indicates that:

    1. AI tools are being deployed into clinical workflows that were not designed to accommodate them, without delivering the promise of increased efficiency, efficacy or safety.
    2. There is no mandated, standardised AI literacy training for regulated healthcare professionals in the UK, leaving adoption quality dependent on local initiatives and personal or professional interest. Whilst many initiatives do exist (e.g. Topol fellowship, Clinical AI Fellowships), spaces are limited and time off work is not always provided to support AI training.
    3. Generational and experience/expertise differences in attitudes to AI tools (AI over reliance vs AI aversion) may lead to divergent clinical behaviours among trainees and senior practitioners, with implications for patient safety, clinical efficacy and workflow efficiency.
    4. The concept of “human in the loop” – central to current regulatory thinking – rests on an assumption of sustained human expertise that is not actively being invested in, maintained or monitored. Ensuring human performance-enhancing “AI in the loop” approaches are embedded and codesigned with clinical challenges, patients and practitioners in mind might support optimal human-AI interaction partnerships.
    5. Professional identity and morale are being shaped by negative public narratives, shaped by constant scaremongering, and poorly designed or inadequately evaluated AI implementation experiences that, if unaddressed, risk undermining recruitment and retention into healthcare professions at a time of acute workforce shortages.
    6. The doctor-patient relationship is already being reshaped by patients’ independent use of AI, changing the nature of clinical consultation, patient expectations, concerns and hopes in ways that neither medical training nor current regulatory frameworks account for.
    7. While technology-enabled innovation is advancing at pace, patient care is the ultimate moral compass for healthcare professionals. We need to ensure that, while AI reshapes the way we deliver technological care, we should also cultivate our knowledge, skills and competencies about unique and diverse patient profiles, looking at the whole person with all their preferences, needs and capabilities and work with them for their humanistic care.This submission elaborates on the seven findings above in more detail, and concludes with seven recommendations for the Commission’s consideration

Findings

Healthcare Professional Identities - Findings

Recommendations

Healthcare Professional Identities - Recommendations

Join the Conversation

References

Balint, M. (1955). The Doctor, his Patient, and the Illness. The Lancet, 265(6866), pp.683–688. doi:https://doi.org/10.1016/s0140-6736(55)91061-8.

Lim, E., Thirunavukarasu, A., He, Y.V., Monkhouse, H., Fu, D.J., de Pennington, N., Higham, A., Tham, Y.C., Jia, Y. and Habli, I. (2026). Building a code of conduct for AI-driven clinical consultations. Nature Medicine, [online] 32(2), pp.400–403. doi:https://doi.org/10.1038/s41591-025-04068-w.


Establishing a public dialogue on Ambient Voice Technologies in the NHS

Month: June 2026

Delivered by network partner: Newton's Tree


CERSI-AI funded a UK public dialogue on NHS AI scribes, capturing patient concerns, trust conditions, and recommendations for transparent, safe, equitable future national implementation.

Mission & Vision

To hear directly from the public about their hopes, concerns and expectations for AI scribes in the NHS and produce a patient-led set of recommendations to guide trustworthy adoption.

Click image to enlarge

Putting the patient voice at the centre of AI adoption

The project aligns with CERSI-AI’s vision to ensure AI-enabled healthcare is safe, effective, and equitable. AI scribes (also known as ambient voice technologies, or AVTs) arrived in the NHS rapidly, but the public voice was largely absent from the conversation. This dialogue was funded to help close that gap, ensuring that the people most affected by this technology have a meaningful say in how it is deployed.

The work was delivered by Newton’s Tree, with Hopkins Van Mil, an independent social research agency specialising in deliberative processes, running the dialogue itself to ensure independence and rigour.

AVT Technology Speakers
AVT Technology in the NHS

Why is this project commissioned?

AI scribes appeared in NHS settings quickly, driven by their potential to ease workforce pressures. National bodies focused on safety and regulation, but the patient voice was largely missing. Without understanding public expectations, there is a risk that this technology is deployed in ways that erode trust or diminish patient experience.

CERSI-AI funded Newton’s Tree to commission this dialogue so that patients could shape, rather than simply react to, the direction of travel.

The project aimed to understand what conditions of AI scribe use would be considered acceptable and trustworthy by the public.

How did we engage the public on AI scribes?

In December 2025, Newton’s Tree commissioned Hopkins Van Mil to design and deliver a public dialogue exploring the use of AI scribes in healthcare. Forty-one participants from across all four UK nations took part in a week-long asynchronous online activity, then attended a full-day in-person workshop in London in March 2026 to develop their recommendations.

Participants engaged deeply with expert speakers, a live demonstration of AI scribe technology, and deliberated across four thematic areas – consent, accuracy, governance, and fairness – before producing a “patients’ charter” of conditions for acceptable and trustworthy use.

AVT Technology Workshop
NHS Medical Staff

Who will use these findings?

  • Primary audience: NHS England, MHRA, and AI vendors who need to understand patient expectations before further rollout.
  • Secondary audience: NHS clinicians and governance teams using AI scribing tools who must implement scribes locally in line with patient needs.
  • Wider stakeholders: Parliamentarians, think tanks, journalists, and civil servants influencing national policy on AI in healthcare

How was this project delivered?

  • A diverse cohort of 41 participants was recruited, broadly reflective of the UK population with deliberate boosts for lower socio-economic groups, minoritised ethnic groups and those with long-term conditions.
  • Participants were educated through expert videos, question and answer sessions, and a live technology demonstration before deliberating, ensuring informed rather than reactive views.
  • Policy writers from public sector bodies attended the workshop in person so they could hear the public’s views first-hand.
  • Findings have been written up in a full report by Hopkins Van Mil, published May 2026, with active dissemination to politicians, journalists and civil servants ongoing.
  • The findings were also discussed by experts at a launch event in June 2026, hosted by Baroness Thornton at the House of Lords. This event brought together representatives from the NHS, regulators, patient organisations, industry, and academia to explore how the public’s expectations can be reflected in the implementation and governance of AI scribes.

What did participants tell us?

  • Many participants were unaware this technology had already been rolled out within the NHS, and this low awareness caused frustration.
  • Participants recommended greater transparency, a way to opt-out of use, and to be able to request scribes be turned off when they wanted a sensitive topic – such as domestic abuse – off the record.
  • Participants wanted assurance that there would be appropriate national oversight, a standard for scribe accuracy, necessary NHS staff training, and strong safeguards over data use and security.
  • Many welcomed the potential of scribes to improve human-to-human interaction with the clinician. However, a few found increased engagement – by removing typing notes – unsettling.
  • Many wanted assurance that there would not be a ‘post code lottery’ for AI scribes.

Recommendations for trustworthy adoption

  1. There should be no surprises for patients in how data is created and used. Therefore steps need to be taken to ensure patients are aware of AI scribe use.
  2. Clarity is needed for NHS staff on the purpose of the clinical record and whether records be created with care or litigation in mind.
  3. Regulation alone will not ensure good use of AI scribes. Change in professional norms and good practice must be created and supported.
  4. An ecosystem approach is needed to keep AI scribes safe – we cannot place all responsibility on front line staff.
  5. AI scribes should support the wider objective of improving the patient experience of care. For example, facilitating the NHS App could reduce the need for patients to repeatedly share their story across services.
  6. Equity should be considered across the AI scribe lifecycle.
  7. Patients, the public and workforce all need support to use these tools appropriately, which may also include patients using their own transcription tools.

What’s next?

This work intends to be just the beginning of the delivery of AI in care in a way that patients expect. We will continue to work with policymakers, regulators, NHS organisations and other relevant parties to ensure the findings inform decisions around implementation and use of AI. As the work of others increasingly shows the importance of the public voice in successful AI adoption, this project provides a path whilst we are still in a position to get this right.

Further reading and resources

Read a summary of the House of Lords launch event discussion here: {Link}

Read the full report:

The Critical Clarity Health and Care Needs

The Critical Clarity Health and Care Needs

CERSI-AI THOUGHT LEADERSHIP

Month: June 2026

Based on interviews with: Dr. Hugh Harvey (Hardian Health): Managing Director at Hardian Health. A former consultant radiologist, with experience at two healthtech start-ups & Ben Howes (Hardian Health): Director of Technology at Hardian Health.

June 2026

The world’s scattered medical device data finally connected with HaRi. 

Revealing the Truth

Hospitals in the UK rely on machines called infusion pumps to deliver medical drugs to their patients. One such hospital was experiencing issues with their brand of infusion pumps.

Recognising this patient safety concern, the hospital sought a way to uncover comparable accounts. They needed to determine whether the difficulties they encountered reflected an isolated anomaly or a broader pattern with the product.
The team used Hardian Regulatory Intelligence (HaRi), a unified, global database, to explore medical device information from huge jurisdictions in one click. Instantly comparing brands side by side, they identified safer alternatives with fewer reported issues and successfully switched their inventory.

Medical professionals were able to get a clear answer to a question that previously required hours of manual searching.

HARI INTEL Photo

The Team Tackling the World’s Medical Data

CERSI-AI is excited to be working with the Hardian Health team to support HaRi from formative stages through to full-scale development. To follow the unfolding impact HaRi is delivering, we interviewed the experts leading the project:

  • Dr. Hugh Harvey (Hardian Health): Managing Director at Hardian Health. A former consultant radiologist, with experience at two healthtech start-ups.
  • Ben Howes (Hardian Health): Director of Technology at Hardian Health.

Their mission is to establish HaRi as an intelligent system created to support medical professionals as they navigate the complexity of global regulatory research.

Medical Device Data: A Puzzle with Lost Pieces

The current well of medical device data is fragmented, scattered across different databases, different registries, different jurisdictions.

During our interview, Dr Hugh Harvey recognises that if you wanted to track a device’s global output:

“You’d currently have searched […] somewhere between 13 and 15 different databases” subsequent to asserting there are “8 million [medical devices], there’s so many.” – Dr Hugh Harvey

Critical information that will significantly benefit patient care and health regulation is often hidden away in labyrinths of registries. Notoriously difficult to navigate, these databases scatter threads that, once assembled, align into convergent data patterns. This problem reveals the pivotal safety gap. The Hardian team exemplifies an NHS buyer, researching a medical device that their trust is purchasing. Ben Howes empathises with the buyer, understanding that they:

“wouldn’t normally have […] half the time to go searching […] in 10 plus different databases.” – Ben Howes

The buyer will find it nearly impossible to track real-time safety trends with the current data split. Their device failing in Australia might go unnoticed in the U

One Search, Global Safety

HaRi is the first to linking regulated medical devices to their safety events in one single central database. Picture HaRi as your reliable and accurate search engine. The tool analyses and displays digestible data from its current jurisdiction access. Previously dispersed information from the FDA, EUDAMED, MHRA, Health Canada, and TGA is unified into one succinct viewing experience.

HARI - Infographic 3

syringes, to constantly evolving AI-integrated software. HaRi’s place within the safety gap solution lies within directly linking these products to adverse events and safety notices, bringing global transparency to health and care.
HaRi offers a single touchpoint for verifying a product’s status. The tool allows users to find the information they need in one click, relieving them of hours of scattered searches.
These advantages support a growing community of users, including:

  • Manufacturers that wish to track their devices performance, claim ownership of their products and display their company’s transparent trustworthiness.
  • Healthcare providers, required to monitor the safety and output of their inventory.
  • Regulatory professionals investigating predeceasing patterns and spotting safety trends.
  • The patients and public who may be interested in researching the devices they use, free of charge.

Visit the HaRi website to explore further:
hari.hardianhealth.com

The New Era of Intelligence

We are facing a future in health and care where HaRi is more important than ever. With AI-driven medical devices in development, regulatory intelligence that protects organisations and patients is fundamental.

During the interview, Dr Hugh Harvey and Ben Howes discussed the near future of HaRi, offering a first look at what’s ahead.

Alongside mapping AI medical devices and monitoring the unique safety events that coincide, HaRi looks to deliver greater precision and usability. The team is working on integrating Claude AI to enable full-scale freedom and critical curiosity. You can already explore millions of medical devices, and with contextual language search, you can ask natural questions like:

Find me a medical device that does X and is available in Europe

… and receive trusted insights grounded in HaRi’s regulated, safe foundation. Plans are underway to add jurisdictions such as Singapore, Brazil and Japan, all while maintaining certified for ISO 9001 and ISO 27001, ensuring accuracy and security.

HaRi is the bridge between us and the seamless future of device regulation.

Hardian Regulatory Intel Infographic

Your Curiosity Evolves Patient Care. Ready?

Start searching for free on HaRi today:

hari.hardianhealth.com

Follow us at CERSI-AI as we work to bridge the gap between extraordinary technologies and patient benefit.

Connect on:


AI Scribes Report Launched London

AI Scribes Report Launched at House of Lords

Month: June 2026

Report Launch: A Public Dialogue on AI Scribes in the NHS
Wednesday 3 June 2026

June 2026

AI Scribes Report Launched at House of Lords

Background

The discussion at the House of Lords marked the launch of a first-of-its-kind report exploring public attitudes towards AI scribes, or ambient voice technologies ‘AVTs’, in healthcare. Key findings from the Report were:

  • Many participants were unaware this technology had already been rolled out within the NHS, and this low awareness caused frustration.
  • Participants recommended greater transparency, a way to opt-out of use, and to be able to request scribes be turned off when they wanted a sensitive topic – such as domestic abuse – off the record.
  • Participants wanted assurance that there would be appropriate national oversight, a standard for scribe accuracy, necessary NHS staff training, and strong safeguards over data use and security.
  • Many welcomed the potential of scribes to improve human-to-human interaction with the clinician. However, a few found increased engagement – by removing typing notes – unsettling.
  • Many wanted assurance that there would not be a ‘post code lotteryʼ for AI scribes. The themes explored during this discussion reflected the key points raised by the public.

The themes explored during this discussion reflected the key points raised by the public.

AI Scribes Consultation
Robin Carpenter Portrait

After opening remarks from Baroness Glenys Thornton and a video message from Lord Ara Darzi highlighting the importance of innovation and the patient, Robin Carpenter, Head of Policy & Governance at Newtonʼs Tree, chaired a discussion focused on taking forward the key points from the report.

It opened by highlighting that transparency and meeting patient expectations around data use have always been key in data standards and law. Participants in the report broadly agreed that patients should be informed when AI scribes are being used and should be able to ask for them to be switched off. Nicola Byrne, National Data Guardian and Consultant Psychiatrist in the NHS, emphasised the importance of the Caldicott Principles, particularly principle 8, that of “no surprisesˮ. She also cautioned against creating transparency expectations that cannot realistically be delivered.

Additionally, discussion also considered the purposes of the clinical record. The core purpose is to support excellent clinical care – for the individual patient themselves, and potentially for others in the future via planning or health research. In addition, it also provides a basis for clinical accountability and any future challenge, including potential litigation, when questions arise about the care that has been delivered.

Nicola Byrne was keen to stress the importance of not losing sight of the core purpose, noting that if clinical documentation is too long and exhaustive, there is a risk that key details can become clouded by redundant information. If clinicians cannot find key details quickly and easily, this can be to the detriment of patient care.

Trust emerged as a recurring theme throughout. Matt Westmore, CEO at the Health Research Authority, noted that this technology is already being used, and that public trust is the most “precious resourceˮ – and it will ultimately determine whether these technologies succeed or fail. He noted that regulation must evolve alongside changing public attitudes.

Prof A Denniston

The panel agreed that regulation has an important role to play, but is only one component of ensuring safe and effective deployment. Alastair Denniston, Chair of the National Commission into the Regulation of AI in Healthcare and Consultant Ophthalmologist in the NHS, stressed that risk-free healthcare is a myth and highlighted the importance of professional norms, good practice, and bringing people with you through incremental change, rather than relying solely on the top-down approach of regulation.

“Regulation is a tool in the toolbox, but they [regulations] donʼt solve the whole problem.ˮ – Alastair Denniston

Panel member Jessica Paulsen, Deputy Director of AI and Software, Medicines and Healthcare products Regulatory Agency MHRA, emphasised that safety and accuracy are critically important, particularly as AI technologies continue to change over time. She highlighted the importance of ongoing monitoring, post-market surveillance, local assurance processes, and clinician training. She also noted that the comparator is not perfection and warned against letting perfect be the enemy of good.

“Safety and accuracy are critically important…[we] need to take an ecosystem approach.ˮ – Jessica Paulsen

Discussion also explored accountability when errors occur. While clinicians have an important role in reviewing outputs, several contributors cautioned against placing all responsibility on frontline staff. The need for clear responsibilities across manufacturers, healthcare organisations and professionals was highlighted.

Diversity in AI Access

A key concern raised by the public was the potential for a “postcode lotteryˮ in access to high-quality care.

National Chief Clinical Information Officer at NHS England, Consultant Rheumatologist in the NHS, and panel member, Alec Price-Forbes argued that AI scribes should be seen as part of a wider transformation of care rather than a standalone technology. He suggested that tools such as scribes could help create more standardised and connected patient experiences via the NHS App, reducing the need for people to repeatedly tell their story across different services. “Patient care should be digital by default… ambient scribes are a solution to a much bigger problem. […]

We need to talk about a public narrative, and make sure the public, patients and workforce are engaged in the model.ˮ – Alec Price-Forbes

Discussion contributors were united in stressing the importance of ensuring deployment does not create a two-tier system. Equity considerations should be built into design, evaluation and implementation from the start.

Additionally, contributors highlighted that patients need to experience the benefit if they are going to support transformation, and AI scribes may be the best opportunity to improve patient perception.

“The opportunity is to shift the public perception of AI in the NHS, and AVT is the poster child for AI deployment, impact and attitude towards how the NHS uses this technologyˮ – Pritesh Mistry, Fellow Digital Technologies) at The Kingʼs Fund.

Patient AI Transcribing Devices

A key concern raised by the public was the potential for a “postcode lotteryˮ in access to high-quality care.

The discussion highlighted the need to move beyond viewing patients as merely passive recipients of care, and therefore contributors explored how AI tools could help patients become better informed and more actively involved in their care decisions.

The topic of patients using their own recording or AI tools during consultations was also explored. While some clinicians acknowledged that this could feel uncomfortable, they supported, and even encouraged, patients’ rights to use such tools and emphasised the importance of openness and trust.

Participants agreed that successful adoption will require significant investment in training and change management. Alec Price-Forbes noted that implementation has often been too technology-led and argued that patients, the public and the workforce all need support to use these tools effectively to realise the benefits of safer, better quality care and experience.

KEY TAKE AWAYS

  1. There should be no surprises for patients in how data is created and used. Therefore steps need to be taken to ensure patients are aware of AI scribe use and expectations for their use are clear amongst both clinicians and patients.
  2. The clinical record can serve multiple purposes, such as litigation and care, but at its core it should provide the information needed to enable excellent clinical care. The design and use of AI scribes needs to actively support the core purpose of the record.
  3. Regulation alone will not ensure appropriate use of AI scribes. Professionals need to adhere to established values and good practice, and be supported in understanding how to apply these to new technologies.
  4. An ecosystem approach is needed to keep AI scribes safe – we cannot place all responsibility on front-line staff.
  5. Inclusion and equity should be considered in every aspect of the design and use of AI scribes.
  6. AI scribes should support the wider objective of improving the patient experience of care. For example, the NHS App could act as a ‘connectorʼ alongside scribes, which could reduce the need for patients to repeatedly share their story across services.
  7. Patients, the public and workforce all need support to use these tools appropriately.

“Healthcare is a two-way relationship. We have to be really clear about what we want the tool to do, and in what context… deployment has to be iterative and learn as we go.ˮ – Nicola Byrne

Join the Conversation

Follow us at CERSI-AI as we work to bridge the gap between extraordinary technologies and patient benefit.

Connect on:

References

The event was attended by senior representatives from national healthcare and care bodies, government, regulatory agencies, academia, patient groups, policy and research institutes, charitable foundations and health tech companies, reflecting a broad cross-sector interest in the use of AI scribes in healthcare.


Clinical Evidence for SaMD and AI Devices

Clinical Evidence for SaMD and AI Devices: A Practical Guide for Manufacturers

CERSI-AI THOUGHT LEADERSHIP

Month: June 2026

By Adam Isaacs Rae, The Other Consultants – June 2026

June 2026

 The question SaMD and AI device manufacturers ask more than any other is a simple one: does my device need a clinical trial? 

As ever, the answer is it depends (sorry). It depends because clinical evidence is not a single threshold you either cross or you don’t, and because the question itself confuses two regulatory concepts that the legislation deliberately separates. Both UK MDR 2002 and EU MDR 2017/745 require every medical device to have a Clinical Evaluation. Neither requires every device to have a Clinical Investigation. The first is mandatory and continuous. The second is conditional on whether existing evidence is enough to demonstrate safety, performance and clinical benefit for the specific intended purpose and the specific claims being made.

That distinction matters most for SaMD and AI, because the gap between what a manufacturer thinks the device is claiming and what the evidence actually supports tends to be wider than for a hardware device. A surgical instrument is constrained by physics. A model is constrained by its training data, its deployment context, and the words on the marketing site, and those three things are rarely in alignment.

Before deciding what evidence is needed, it has to be acknowledged that clinical evidence planning belongs at the start of the design and development cycle, alongside the intended purpose and indications for use — not as a downstream activity once the device is built. Once intended purpose drift sets in, the evidence question becomes harder to answer and the evidence package more expensive to assemble.

Intended purpose, indication for use, target population, clinical setting, the role the device plays in clinical decision-making, and the benefit being asserted are the inputs that drive the evidence requirement. Get them wrong, leave them implicit, or let marketing pull them in a direction the technical file does not support, and the evidence question becomes unanswerable. A clinical investigation cannot be scoped if there is no settled view on what the investigation has to demonstrate.

For AI devices this is sharper still. A model that “supports clinical decision-making in the assessment of diabetic retinopathy” is doing something fundamentally different from a model that “detects referable diabetic retinopathy.” The first is a decision support tool. The second is a diagnostic claim. The evidence package for those two devices is not slightly different. It is structurally different.

The intended purpose itself rewards specificity. Position within the clinical workflow should be explicit — diagnosis, prevention, monitoring, prognosis or otherwise — going one step beyond the words used in the regulatory definition. Inputs and outputs should be defined: the SaMD takes the output of a particular previous element in the workflow as its primary input (a prior diagnosis, an x-ray image, a structured dataset), performs a defined activity, and produces a defined output. Where a device contains both medical software and non-medical software components, the non-medical components should be clearly demarcated in the technical documentation and the intended purpose, so that what is and is not under conformity assessment is unambiguous. Contraindications, warnings and side-effects belong elsewhere in the technical documentation, not buried inside the intended purpose statement.

A practical implication: the IFU and marketing copy should not be drafted before the clinical evidence work has matured. Doing it the other way around is one of the most common ways claims drift into territory the evidence cannot support. The IFU and marketing assets describe what the device does — they should reflect what has been demonstrated, not predict it.

When intended purpose, claims and evidence are aligned, what gets communicated to the world is what the device can actually do.

Clinical Evaluation under MDR Article 61 is a planned, continuous process to collect, appraise and analyse clinical data to verify safety, performance and clinical benefit. The acceptable sources are defined in Article 2(48): clinical investigations of the device, clinical investigations or studies of an equivalent device, peer-reviewed scientific literature on the device or an equivalent device, and post-market surveillance data including PMCF.

That framing is the right starting point for hardware. For SaMD and AI, regulators have converged on a sharper one. IMDRF N41 — referenced by the MHRA, FDA, Health Canada and the TGA — sets out three components that any SaMD clinical evaluation has to address, and they are the lens through which Notified Bodies and Approved Bodies actually read AI clinical files.

It is worth being precise about terminology before going further. IMDRF N41 defines clinical evidence as an important component of the technical documentation of a medical device which, alongside design verification and validation documentation, device description, labelling, risk analysis and manufacturing information, is needed to demonstrate conformity with the Essential Principles. Clinical evidence is the output. Clinical evaluation is the process that produces it. Conflating the two is one of the more common reasons technical files lose coherence, because the question “do I have enough evidence?” is not the same question as “have I done a clinical evaluation?”

The first component is valid clinical association. Is there a scientifically established link between the output the model produces and the clinical condition it is intended to address? For a model that flags suspicious lung nodules on CT, the question is whether the imaging features the model uses are accepted in the literature as associated with malignancy. N41 sets out where this evidence comes from: literature searches, original clinical research and professional society guidelines for existing evidence, and secondary data analysis or new clinical trials where new evidence has to be generated. Where the underlying clinical association is well-established, it can be referenced. Where the association is novel — a new combination of inputs, a new biomarker, a new patient population — original evidence is required.

The second is analytical validation. Does the model correctly process its inputs to produce the outputs claimed, with the accuracy, reliability and precision claimed? This is the locked test set, the held-out validation cohort, the sensitivity and specificity figures, the confidence intervals. N41 expects this evidence to be generated through verification and validation activities under the QMS, supplemented where needed by curated databases or previously collected patient data. It is verification that the software does what the specification says.

The third is clinical validation. Does using the model’s output, in the hands of the intended user, in the intended setting, on the intended population, achieve the intended clinical purpose? N41 recognises three ways to demonstrate this: referencing existing data from studies conducted for the same intended use; referencing existing data from studies conducted for a different intended use, where extrapolation can be justified; or generating new clinical data for the specific intended use. The first option is the cheapest and the rarest. The second sounds attractive but the extrapolation justification is where many manufacturers come unstuck. The third is what most novel SaMD and AI devices end up needing, in part or in full. This is the layer that cannot be answered by analytical performance alone. It is the layer that matters most, and the one most often underweighted.

Two practical consequences for SaMD and AI sit on top of that framework.

Equivalence is difficult in this industry generally, and harder still for AI. Two infusion pumps with comparable specifications can plausibly be equivalent. Two image classification models trained on different datasets, validated on different populations, with different decision thresholds and different deployment workflows almost certainly are not — even if their stated intended purposes look similar on paper. The MDR’s technical, biological and clinical equivalence criteria were not written with AI in mind, and they do not flex to accommodate it. Treat equivalence claims for AI devices as a high bar that will probably not be cleared, not as a shortcut to default to. Even outside AI, access to a comparator’s technical file is rare, and equivalence claims fail at conformity assessment for that reason alone with regularity.

Literature behaves very differently depending on whether the device is well-established or genuinely novel. For a well-established device category, literature can carry significant weight: the underlying clinical association is typically settled, comparator data exists, and the SOTA is documented across multiple sources. For a novel device, literature is much harder to lean on. Published studies of AI devices are growing fast but the field has well-documented problems: small datasets, single-site validation, retrospective designs, lack of prospective deployment data, and selective reporting. A systematic literature review for a novel AI device has to grapple with all of that honestly. Citing twelve favourable retrospective studies does not establish state of the art if the state of the art is now prospective, multi-site, and reporting algorithmic bias metrics.

A literature review for a serious clinical evaluation rarely consists of a single search. Three or four targeted searches, each with a defined scope, will usually do more useful work than one broad sweep — one search for clinical SOTA, another for similar or alternative devices, another for evidence specific to the device under evaluation. Recognised scientific databases, more than one per search, paired with a documented appraisal methodology, are the baseline.

A clinical investigation is required when existing data is not adequate to demonstrate conformity for the specific device, the specific intended purpose and the specific claims. Looked at through the IMDRF N41 lens, the threshold is usually reached because of one specific gap: analytical validation alone does not establish clinical validation, and for SaMD that gap is rarely closed by literature or equivalence.

A locked-dataset performance study tells you how a model performs against a fixed reference standard in a controlled environment. It tells you nothing about how the model performs when a junior clinician uses it under time pressure, on a population whose distribution has drifted from the training data, in a workflow where the model output competes with five other information sources. For higher-risk decisions and higher-risk patient interactions, the regulator is increasingly going to want prospective deployment evidence — clinical validation, in N41 terms — and that usually means a clinical investigation, a real-world performance study, or a structured deployment evaluation.

Two further factors push the bar higher. Novelty is one: new mechanisms, new populations, new workflows and new AI/ML decision support paradigms all reduce the amount of existing evidence that can be relied on, by definition. Claims drift is the other: if the claims sit in diagnostic territory, the evidence has to sit there too, regardless of what the model can technically do.

Clinical performance is also rarely uniform across a SaMD. Many devices are modular: only some features make a clinical claim. Where modules operate independently, clinical performance has to be validated at the module level rather than at the device level. A performance matrix is a useful way to demonstrate this — features mapped against the metrics used to validate them (sensitivity, specificity, positive and negative predictive values, likelihood ratios, odds ratios, confidence intervals, usability outcomes), tied back to clinical SOTA and to similar or comparator devices where data exists. MDCG 2020-1 frames this around intended use, indications, desired clinical outputs expressed as claims, and the clinical benefits expected to follow.

When a clinical investigation is run, ISO 14155:2020 is the framework. It governs ethics, scientific validity, monitoring, data management, statistical analysis and reporting. Notified Bodies, Approved Bodies and the MHRA all expect compliance with it for investigational studies.

A clinical investigation is not required when sufficient clinical data already exists, when the device is well-established technology, when meaningful clinical endpoints don’t exist (a sterilisation-cycle logger, for example), or when safety and performance can be fully justified by non-clinical evidence and the claims are genuinely technical rather than clinical.

Article 61(10) of the EU MDR sometimes gets reached for here, on the assumption that it lets a manufacturer demonstrate conformity without clinical data. It can — but only in narrow circumstances, and SaMD and AI devices very rarely qualify. The route is intended for Class IIa or non-implantable Class IIb devices where clinical data is genuinely not appropriate, the clinical benefit is indirect, and there are no meaningful clinical endpoints to measure. A data-logger with no clinical claims might qualify. Image management software with purely functional claims might qualify. Almost any AI device making a clinical decision, triage, diagnostic or prognostic claim will not and is likely to struggle (think back to our IMDRF N41 documentation).

Pre-market evidence rarely settles every clinical question for a SaMD, and it almost never does for an AI device. Post-market surveillance and post-market clinical follow-up are where residual risks get confirmed or downgraded, where performance drift gets detected, and where the assumptions in the original clinical evaluation get tested against actual use or any gaps are verified or removed.

PMCF should not be a placeholder, it should be designed to address the specific residual risks identified in the risk file, and to confirm or revise the clinical performance assumptions that underpinned market entry. Where AI is involved, real-world performance data is also the mechanism by which model drift gets identified before it affects patient outcomes.

State of the art (SOTA) also has a habit of moving faster than clinical evaluation reports. SOTA defined three years ago, never revisited, is a recurring finding in conformity assessment reviews. For AI specifically, what counted as SOTA for a 2022 retinal imaging model is not what counts in 2026; the methods, the reporting standards and the expectations around bias and equity have all shifted. SOTA needs to be reviewed deliberately, on a defined cadence, and the CER updated accordingly.

The UK has developed a set of mechanisms specifically for innovative and AI-enabled devices that do not exist in the EU framework. The MHRA’s AI Airlock — now into a multi-year programme and feeding directly into the National Commission into the Regulation of AI in Healthcare — provides a controlled environment for validating AI as a Medical Device under proportionate oversight, generating real-world evidence in a way that pre-market data alone cannot. The Innovative Devices Access Pathway (IDAP) supports staged evidence generation with early regulatory input, bringing evidence planning forward into device development rather than treating it as a hurdle at the end.

These pathways do not replace the requirement for a Clinical Evaluation. They reshape how and when the underlying clinical evidence gets generated, and they take the planning conversation upstream. For UK-based SaMD and AI manufacturers — and for those targeting both UK and EU markets — they are worth understanding early.

Two AI-specific points deserve to be called out, because they consistently come up in conformity assessment discussions and are easy to overlook in early evidence planning.

The first is model drift. If the AI model contributes to clinical decision-making or automated outputs, a defined model drift management strategy is required. That means specifying how performance degradation will be detected, how often model performance will be reviewed, and what re-validation or update mechanisms are in place. Notified Bodies and Approved Bodies will expect this addressed wherever model behaviour may change over time or with new data inputs.

The second is large language model training data. Where an LLM is integrated into an MDSW, the origin, composition and limitations of the training data have to be documented and the relevance to the intended clinical purpose has to be assessed. That includes demonstrating how well the training data represents the target patient population, the clinical context, and the intended use, and identifying and mitigating bias and limitations.

When clinical evidence packages fail review, the same patterns recur. Claims in the IFU not supported by the CER. Risk management contradicting the clinical evidence. Literature reviews that aren’t systematic, with no defined search strategy, no inclusion and exclusion criteria, and an outdated cut-off date. Equivalence asserted without access to the technical file of the comparator device. Bench testing offered as evidence for clinical claims. SOTA defined three years ago and never revisited. PMCF plans that don’t address the residual risks identified in the risk file. Marketing copy claiming functionality the clinical evidence does not substantiate — and Notified Bodies and Approved Bodies do monitor manufacturer websites and social media after certification.

Almost every one of these is a coherence failure rather than a data failure. The data exists, somewhere, but the documents don’t tell a single story. A reviewer reading the intended purpose, the GSPR checklist, the risk file, the CER, the IFU and the marketing material in sequence should see the same device described the same way each time.

There is no shortcut to a defensible clinical evidence strategy for a SaMD or AI device. There is, however, a sequence: build clinical evidence planning into the start of the design and development cycle, define what is actually being claimed, map every claim to the evidence that supports it, decide honestly whether existing data is enough, and design any new study — clinical investigation, prospective performance study, real-world deployment evaluation — to close the specific gap identified. A device manufacturers PMS system has now become more of a strategic asset than ever before.

Adam Isaacs Rae, The Other Consultants

Founder

Adam Isaacs Rae is the founder of the Other Consultants (theotherconsultants.com), a medical device quality and regulatory consultancy. He works with manufacturers on regulatory strategy, technical documentation, clinical evaluation, and audits across UK MDR, EU MDR and FDA frameworks.

CERSI-AI — the Centre of Excellence for Regulatory Science and Innovation in AI and Digital Health — is a UK initiative bringing together regulators, clinicians, researchers and industry to develop the regulatory science needed to safely advance AI and digital health technologies. It works closely with the MHRA, NHS, NICE and academic partners to inform regulatory thinking, support innovators, and help build the evidence frameworks that AI in healthcare requires. More at cersi-ai.org.

Join the Conversation

References

Balint, M. (1955). The Doctor, his Patient, and the Illness. The Lancet, 265(6866), pp.683–688. doi:https://doi.org/10.1016/s0140-6736(55)91061-8.

Lim, E., Thirunavukarasu, A., He, Y.V., Monkhouse, H., Fu, D.J., de Pennington, N., Higham, A., Tham, Y.C., Jia, Y. and Habli, I. (2026). Building a code of conduct for AI-driven clinical consultations. Nature Medicine, [online] 32(2), pp.400–403. doi:https://doi.org/10.1038/s41591-025-04068-w.


Privacy Preference Center