bg-left bg-right

Artificial Intelligence in Post‑Market Surveillance of IVD Assays 

background
user-icon 16 Jan 2026

Introduction 

Post‑market surveillance (PMS) is the ongoing monitoring of device safety and performance once a product is on the market. For in vitro diagnostic (IVD) devices, PMS is crucial. It ensures that patient health remains protected and that devices continue to perform as intended. 

New European regulations—the Medical Device Regulation (MDR) and the In Vitro Diagnostic Medical Device Regulation (IVDR)—have raised the bar. Companies must now proactively search for information related to device safety and performance. This includes systematic reviews of scientific and technical literature. 

The challenge is clear: searching the literature for safety signals is time‑consuming, resource‑intensive, and prone to human error. 

This blog is based on the peer‑reviewed paper published on 25th March 2024 “Artificial intelligence / machine‑learning tool for post‑market surveillance of in vitro diagnostic assays” authored by Joanna Reniewicz, Vinay Suryaprakash, Justyna Kowalczyk, Anna Blacha, Greg Kostello, Haiming Tan, Yan Wang, Patrick Reineke, and Davide Manissero. The study was published under a Creative Commons license and provides the foundation for the insights summarized here. 

Why PMS Has Become More Demanding 

The World Health Organization emphasizes that companies must gather information from multiple sources to identify emerging risks. Under MDR and IVDR, this requirement has expanded. 

Key changes include: 

  • Proactive surveillance rather than reactive monitoring. 
  • Systematic collection and evaluation of specialist, scientific, and technical literature. 
  • Documentation of data sources, collection methods, and evaluation methods. 

For companies, this transition has benefits. Proactive surveillance helps refine product performance, maintain competitiveness, and reduce the risk of withdrawal from the market. But it also requires significant investment in new processes, procedures, and skilled personnel. 

The Burden of Manual Literature Searches 

Traditional PMS relies on manual searches of databases such as PubMed and others. This involves: 

  • Refining keyword searches. 
  • Screening titles, abstracts, and full texts. 
  • Extracting relevant information. 

As biomedical databases grow—PubMed alone houses over 35 million citations, with nearly 1.4 million added each year—manual searches become unsustainable. The total volume of data makes comprehensive monitoring across entire product portfolios nearly impossible for small and medium‑sized companies. 

Beyond scientific literature, PMS also requires monitoring of public complaint and withdrawal databases (e.g., FDA MAUDE, EUDAMED vigilance reports, MHRA Yellow Card). These sources contain unstructured narratives, adverse event reports, and regulatory actions that are equally critical for detecting safety signals. Manual review of such databases is even more resource‑intensive, as entries often lack standardized terminology, include duplicates, and require expert interpretation. Without automation, integrating these signals into PMS reports adds another layer of complexity and risk of oversight. 

The Role of Artificial Intelligence 

Artificial intelligence (AI) offers a solution. AI systems can perform tasks such as learning, problem‑solving, and decision‑making. In healthcare, AI tools—including machine learning (ML), deep learning, and natural language processing (NLP)—are increasingly used to screen large datasets. 

Benefits of AI in PMS include: 

  • Faster and more accurate signal detection. 
  • Reduced human error. 
  • Comprehensive coverage of data sources. 
  • Streamlined workflows that save time and resources. 

AI has already shown promise in pharmacovigilance. For example, deep‑learning approaches to individual case safety report (ICSR) processing exceeded 75% accuracy thresholds, improving efficiency and consistency. Surveys of drug safety professionals suggest strong support for AI adoption, as it allows experts to focus on analysis rather than repetitive tasks. 

Case Example: QIAGEN and Huma.AI Collaboration 

Recognizing the opportunity, QIAGEN partnered with Huma.AI to develop a novel post‑market intelligence platform. The goal: improve signal detection and streamline PMS workflows related to risk and performance evaluation. 

This platform automates systematic searches of published scientific literature. It identifies evidence related to: 

  • Device effectiveness and safety. 
  • Patient populations. 
  • Pathogen genetics. 
  • Off‑label use. 
  • Emerging technologies and refinements. 

By using algorithms to looking for specific regions of interest (ROIs), the system contextualizes evidence and reduces the burden of manual screening. 

Comparing Manual vs AI‑Assisted Searches 

A recent study compared manual searches with the Huma.AI platform over a three‑year period. 

Manual search process: 

  • Keyword refinement. 
  • Screening of titles, abstracts, and full texts. 
  • Extraction of relevant information. 

AI‑assisted search process: 

  • Advanced caching techniques. 
  • NLP to identify relevant reports. 
  • Automated analysis of outputs. 

Table: Manual vs AI‑Assisted Searches 

Aspect Manual Search AI‑Assisted Search 
Time required High Significantly lower 
Precision rates Variable Higher and more consistent 
Number of relevant articles Lower Higher in most cases 
Risk of human error High Reduced 

The study demonstrated that the AI platform was more effective and efficient in identifying relevant articles compared with manual approaches. 

Study Design 

The study was designed to compare the Huma.AI platform with traditional manual searches of published literature. The goal was to determine accuracy and efficiency in retrieving relevant articles for MDR/IVDR surveillance of commercially available in vitro diagnostic and medical devices. 

Regular searches of PubMed and PubMed Central were performed between 2019 and 2021. Searches were conducted at frequent intervals depending on the product. Two approaches were used: 

  • Manual search using refined keywords and expert review. 
  • Huma.AI search using advanced algorithms and natural language processing (NLP). 

Two types of searches were performed: 

  1. Literature Search (LS): Identifying publications related to safety and performance of a specific product in clinical settings. 
  1. Technical Assessment Search (TAS): Identifying information that could impact intended use or design, such as genetic sequence changes in target organisms. 

All articles included were peer‑reviewed and published in English. 

Manual Search 

Literature Search 

The manual LS process involved several steps: 

  1. Conducting keyword searches refined from previous searches. 
  1. Screening publication titles and abstracts. 
  1. Reviewing full texts when necessary. 
  1. Extracting relevant information by medical/scientific experts for inclusion in PMS reports. 

Search terms included: 

  • Product or portfolio names (including synonyms and abbreviations). 
  • Disease, mutation, or organism detected by the device. 
  • Target population of the method/procedure. 
  • Defined search period. 

Relevant articles were those explicitly referring to the evaluated product. 

Technical Assessment Search 

The manual TAS process followed a similar structure: 

  1. Keyword search using organism names, synonyms, abbreviations, target regions/genes, and search period. 
  1. Screening titles and abstracts. 
  1. Reviewing full texts when needed. 

Articles were classified as relevant if they contained: 

  • Potential genetic sequence changes in the target region (mutation). 
  • Novel classifications of the target organism (genotype, subtype, strain, species, or genus). 

For each relevant sequence or mutation, a sequence alignment study was performed to assess impact on assay safety and performance. A risk analysis was then conducted by experts for inclusion in PMS reports. 

Huma.AI Search 

The Huma.AI search used a multi‑algorithmic approach with advanced caching techniques to retrieve publications across multiple data sources. 

A key component was the utterance processor (UP), an NLP system designed to disambiguate user input into clear instructions. Unlike typical NLP systems, the UP builds pipelines of operations that can be executed on the fly. It performs intent and entity classification, learns from queries, and integrates them into analysis. 

The workflow involved: 

  • Multiple‑pass analysis of regions of interest (ROIs). 
  • Examination of all document sections against query keywords. 
  • Elimination of false positives. 
  • Extraction of core parts of documents relevant to search criteria. 

Algorithms explored different portions of articles to produce the most accurate ROI based on business rules. 

The system was iteratively trained until classification accuracy matched subject‑matter experts. Targets included: 

  • False negatives minimized to <3%. 
  • False positives kept between 25% and 35%. 

Once trained, the system analyzed titles, abstracts, and full texts. Potentially relevant articles were presented to reviewers with highlighted ROIs. Reviewers then classified articles as relevant or not, similar to the manual process. 

Data Analysis 

Manual and Huma.AI searches were compared by: 

  • Total number of relevant articles identified. 
  • Precision rates per year for LS and TAS. 
  • Time requirements estimated by reviewers in two case studies. 

Precision rate was defined as the proportion of relevant articles (meeting inclusion criteria) relative to the total number of articles found. 

Average precision rates per year were calculated across the three‑year PMS search period (2019–2021). One exception was the SARS‑CoV‑2 detection assay, for which data were available only from the start of the pandemic. In this case, a two‑year PMS search period (2021–2022) was used. 

Precision rate was defined as the proportion of articles classified as relevant (i.e., meeting predefined inclusion criteria) relative to the total number of articles retrieved. Articles that met inclusion criteria were then subject to expert review. Only those judged to provide actionable evidence were accepted into PMS reports. Thus, “relevant” refers to inclusion‑criteria compliance, while “accepted” refers to final incorporation into PMS documentation. 

Results: Literature Search 

The comparison between manual and Huma.AI searches revealed striking differences. Across multiple assays and devices, the Huma.AI platform consistently identified more relevant articles than manual searches. Precision rates were also higher and more consistent year‑to‑year. 

For example: 

  • RGQ MDx assays: Manual precision rates hovered around 40–50%, while Huma.AI achieved >97%. 
  • digene HC2 HPV DNA Test: Manual precision varied widely (50–100%), but Huma.AI maintained near‑perfect precision (94–100%). 
  • SARS‑CoV‑2 assays: Huma.AI retrieved substantially more relevant articles, with precision rates above 90% in most cases. 

Case studies highlighted the difference clearly. For the QIAstat SARS‑CoV‑2 Respiratory Panel, manual searches identified six relevant articles in 2022. Huma.AI retrieved 35, of which 32 were relevant. For the NeuMoDx SARS‑CoV‑2 Assay, manual searches found five relevant articles, while Huma.AI identified 19, including 14 missed by manual screening. 

Time efficiency was another major advantage. Manual research took weeks to complete, while Huma.AI reduced turnaround to hours. 

Results: Technical Assessment Search 

The technical assessment search (TAS) showed similar improvements. Huma.AI identified more relevant articles across assays such as HBV, HCV, CMV, and HIV‑1. 

  • NeuMoDx HBV Assay: Manual precision rates averaged ~6%, while Huma.AI achieved ~77%. 
  • NeuMoDx CMV Assay: Manual precision ~3%, Huma.AI ~91%. 
  • NeuMoDx HIV‑1 Assay: Manual searches retrieved hundreds of articles with <5% relevance. Huma.AI retrieved fewer articles but with ~65% relevance, including many missed by manual screening. 

In one case study, the manual search for the HIV‑1 assay retrieved 308 articles, only nine of which were relevant. Huma.AI retrieved 53 articles, of which 37 were relevant — including all nine found manually plus 28 additional relevant articles. 

Efficiency gains were dramatic: 30 minutes with Huma.AI versus 4 hours manually. 

Discussion 

The study demonstrated that AI‑assisted searches greatly improve PMS literature surveillance. Key advantages include: 

  • Higher precision rates across multiple assays. 
  • Greater consistency year‑to‑year. 
  • Faster turnaround times, reducing weeks of work to hours. 
  • Improved safety signal detection, with ten‑fold increases in relevant articles identified. 

Manual research rely heavily on keywords and abstracts. This approach misses articles where relevant information is buried in the full text. Huma.AI, by contrast, uses NLP and ML to analyze full texts, identify regions of interest (ROIs), and contextualize keywords. This reduces false negatives and increases sensitivity. 

For TAS, manual searches often produced overwhelming volumes of articles with low relevance. Huma.AI reduced the number retrieved but increased precision, focusing reviewers on the most relevant evidence. 

Limitations 

The study acknowledged several limitations: 

  • No “gold standard” dataset exists to measure absolute recall. 
  • Searches were limited to PubMed and PubMed Central. 
  • Snowballing (reviewing references and citations) was not employed. 
  • The system was tested only in English. 

Despite these, the results are generalizable. Huma.AI can interface with other databases and expand to multilingual searches. 

Future Potential 

AI systems like Huma.AI represent the natural progression of PMS. As literature volumes grow, manual methods will become increasingly impractical. AI offers: 

  • Consistent, faster searches with reduced human error. 
  • Better precision rates even with large datasets. 
  • On‑demand repeatability, enabling flexible reporting. 
  • Integration with quality management systems (QMS) for proactive compliance. 

Regulators currently provide no specific guidelines for AI tools in PMS. However, transparency, reproducibility, and validation remain essential. Companies should document training strategies, periodically compare AI outputs with manual reviews, and ensure accountability. 

The trade‑off between cost and effectiveness must also be considered. AI systems reduce workload dramatically while maintaining high recall. For example, text‑mining approaches can cut screening workload by 60% compared with manual methods. 

Conclusion 

Employing AI in PMS is not just an efficiency upgrade — it is a compliance necessity under MDR and IVDR. The Huma.AI platform demonstrated: 

  • Ten‑fold increases in safety signal detection. 
  • Consistent precision rates across products. 
  • Significant reductions in time and resource requirements. 

For manufacturers, this means faster, more reliable surveillance. For regulators, it means stronger assurance of patient safety. And for patients, it means confidence that devices remain safe and effective long after they reach the market. 

AI‑enabled PMS is the future. Companies that adopt these tools will be better positioned to meet regulatory demands, manage risk, and innovate responsibly. 

Reniewicz, J., Suryaprakash, V., Kowalczyk, J., Blacha, A., Kostello, G., Tan, H., Wang, Y., Reineke, P. & Manissero, D., 2023. Artificial intelligence / machine-learning tool for post-market surveillance of in vitro diagnostic assays. Health Policy and Technology, 12(3), p.100742. Available at: <https://www.sciencedirect.com/science/article/pii/S1871678423000687> [Accessed 5 January 2026].

About the author:
Gianluca Tordi

Tags

MDR Guidelines

Worldwide regulation resources

Latest News

Contact us / Ask a quote now

We will help You find the right solution for Your Projects

CONTACT US

SOME OF OUR CLIENTS