bg-left bg-right

AI Airlock – Opening the Door for Responsible AI in Medicine 

background
user-icon 31 May 2026

Space exploration is often associated with discovery and innovation. It required people to to step into the unknown, test new boundaries, and return with knowledge that can benefits society. 

Regulation may seem less adventurous, but involves  a similar process ofof careful exploration and controlled risk-taking. 

Just as astronauts rely on an airlock before entering space, regulators and innovators also need a controlled environment when emerging technologies can be tested safely. This allows them to explore how artificial intelligence can work within, or challenge, existing standards, without risking patient safety. 

The AI Airlock was created as a regulatory sandbox for healthcare AI. Rather than acting as a fast-track pathway to market, it provides a safe environment where regulators and developers can better understand the challenges associated with AI as a medical device (AIaMD). 

The AI Airlock is not a fast track to market. Instead, it gives us a safe space to learn about these new AI products without putting patients at risk. 

Through technical testing, real-world evidence generation, workshops, deep regulatory discussions, and close collaboration, collaborators work to understand and solve the biggest challenges that come with AI as a medical device (AIaMD). 

In pilot phase, they worked with four innovative companies: Philips Healthcare, AutoMedica, OncoFlow, and Newton’s Tree. Each brought a different regulatory challenge to the table. Together, they have explored practical ways to create better testing datasets, how to make large language models (LLMs) safer and more transparent in healthcare, and how to reduce errors. They also looked at how these products can be properly monitored once they are in the market. 

Pilot Results: What Was Discovered 

One of the clearest findings from pilot was that regulatory are struggling to evolve at the same pace as AI technologies.  

A common challenge is this: How do you safely train or test a new medical AI when you don’t have enough real patient data? After all, AI is only as good as the data it learns from. 

To tackle this, collaborators worked with Philips Healthcare on an interesting approach. When real radiology reports were scarce, they used a large language model (LLM) to generate realistic but artificial “synthetic” reports for testing. They then compared these synthetic reports with real ones using both automated tools and human experts. Philips also tested whether an AI could act as a “judge” to evaluate the quality of the generated reports. 

The results were promising—the synthetic reports turned out quite realistic. However, the AI judge still had some limitations. 

This pilot highlighted the need for clearer regulatory guidance on how synthetic data should be generated,validate and assessed before being used in regulatory contexts. Without that clarity, there’s a real risk that flawed or incomplete data could eventually affect patient safety. 

AI hallucinations remain one of the mainconcernes surrounding healthcare AI. These occur when systems produce confident but incorrect responses, which can be dangerous in medical settings. 

The AutoMedica SmartGuideline project tested a practical solution using Retrieval Augmented Generation (RAG). By grounding the AI in verified clinical sources, they significantly reduced hallucinations compared to using a general-purpose model. 

RAG works by retrieving information from trusted external sources before generating responses, which improves both accuracy and transparency. 

Regulatory policy needs to include AI-specific safety measures within risk management frameworks. 

Post-market surveillance of AI tools should be strengthened, possibly by making better use of existing systems like the MORE portal and the Yellow Card scheme. 

Another recurring theme was explainability. Patients and clinicians must clearly understand why an AI medical device makes a particular recommendation. This “explainability” builds trust and allows users to accept or question the output. 

The OncoFlow project focused on explainability. They developed tools that showed how the AI reached its treatment decisions. Clinicians found this transparency “extremely important” for safe and trustworthy use. 

Clear regulatory guidance on explainability would help innovators design better systems that address human factors and build confidence in AI. 

How to Monitor AI Medical Devices in Hospitals 

The final set of questions focused on how to monitor AI as a medical device (AIaMD) once it’s already being used in hospitals. 

Working with Newton’s Tree also demonstrated the importance of continuous monitoring after deployment. Ongoing oversight may help identify issues such as performance drift or increasing over-reliance on AI outputs in high-pressure clinical environments. 

Developing clear guidance on ongoing monitoring would help catch issues like changes in the AI’s performance or over-reliance by users before they affect patient safety. 

The AI Airlock pilot ran on a tight schedule during Phase 1. As they move into Phase 2, we’re allowing a bit more time for deeper regulatory discussions. It’s worth noting that everything they tested and reported on was just a snapshot. The companies involved are doing much more detailed validation work as part of their longer regulatory journey. 

From Pilot to Policy 

In the spirit of shared learning, we’re making all these insights—and many more—openly available. You can find them in detail in simulation workshop reports and the full program report. 

Our work didn’t end with the pilot. The findings from the AI Airlock are already feeding into the UK’s National Commission on the Regulation of AI in Healthcare. They’re helping the commission advise the MHRA on building a practical framework that allows innovative AI to reach patients safely and effectively. 

The National Commission was launched on 26 September 2025. It brings together global AI experts, clinicians, and regulators to shape a new regulatory framework for AI in healthcare, which is expected to be published in 2026. 

Phase 2: From Pilot to Progress 

In Phase 2, they will be looking at several important regulatory challenges, including predetermined change control plans, post-market surveillance, the scope of intended use, and AI-powered IVDs. 

Collaborators will also be working more closely with partners at the Centres of Excellence for Regulatory Science and Innovation—CERSI-AI and RADIANT-CERSI—as well as other organizations across different sectors and countries. Collaboration remains at the heart of what they do. 

The value of this collaboration was clear at recent Airlock Kick-Off Connect event, where they officially launched Phase 2. Collaborators brought together key stakeholders to start building a shared understanding of the challenges and possible solutions. 

Over the coming weeks, collaborators will finish situation assessments and move into developing test plans, including key hypotheses and methodologies. During the winter months, they will focus on testing and gathering regulatory intelligence. As the new year begins, they will hold simulation workshops to review early results and discuss practical solutions and policy recommendations.

TS Quality & Engineering is also into the process of innovating various products with AI with various partners, let us know if we can collaborate.

About the author:
russo.tsquality

Tags

MDR Guidelines

Worldwide regulation resources

Latest News

Contact us / Ask a quote now

We will help You find the right solution for Your Projects

CONTACT US

SOME OF OUR CLIENTS