← Back to list

What WHO Prequalification Taught Me About Global Health AI

If you’ve never heard of WHO Prequalification (PQ), you’re in good company — many senior regulatory affairs managers haven’t engaged with…

nina sun · 2026-07-05 06:15 · 0 claps · 4.9 min read
#healthcare-ai #global-health #digital-health #who #healthtech
Open on Medium ↗
Wiki topics: PUB · Public Health & Epidemiology DH · Digital Health & Health Tech

What WHO Prequalification Taught Me About Global Health AI

If you’ve never heard of WHO Prequalification (PQ), you’re in good company — many senior regulatory affairs managers haven’t engaged with this framework either. But if your team builds AI medical devices for low- and middle-income countries (LMICs), NGOs, or global public health programs, WHO PQ can be under consideration.

My understanding of PQ shifted entirely after two deep dives into Dr. Tang’s presentations. The first time, I only ran rough cost calculations and dismissed it as an overly costly, niche clinical trial requirement. The second time, I combed through WHO’s official guidance and raised targeted technical questions.

A turning point arrived in June 2025, when WHO published its formal specification for TB computer-aided detection (TB CAD). Soon after, the Gates Foundation offered to support our prequalification roadmap. I led a full gap assessment, cataloging our existing documentation and clinical datasets while preparing to address the foundation’s rigorous technical inquiries. This hands-on work let me contrast WHO PQ against familiar standards: ISO 13485, ISO 14971, IEC 62366, plus regional regulations FDA, MDR and NMPA — and identify its unique, AI-focused requirements and unresolved grey areas.

What Sets WHO PQ Apart From Conventional Medical Device Regulation

While it aligns with baseline global compliance rules, WHO PQ adds bespoke guardrails built around LMIC real-world care contexts, especially for machine learning software. I’ve broken down the most impactful mandates below.

  1. Non-Clinical AI Evidence: Full Algorithm Transparency Is Non-Negotiable

The specification enforces end-to-end traceability for all ML workflows, with strict rules for Software Description Summaries (SDS):

  • SDS must detail algorithm selection logic, dataset maintenance protocols, and full training/validation workflows, with enough granularity to prove real-world robustness.
  • Dataset demographics are mandatory. Pre-2023, many AI devices cleared regional certification without accounting for diverse patient populations — a practice that was neither scientifically sound nor ethically responsible. WHO PQ closes this loophole.
  • Explicit reference to FDA’s draft AI/ML change control guidance addresses iterative model risks. One critical unaddressed question remains: if a base CAD model earns PQ approval, do fine-tuned versions for new disease indications require separate clinical trials or supplementary validation?
  • A standout imaging input rule: DICOM is the only acceptable input format for TB CAD. JPEG/PNG are permitted only for final report outputs. Some vendors market phone photo compatibility as a selling point, but unstandardized imagery cannot meet clinical rigor for diagnostic AI. This rule eliminates misleading commercial claims entirely.

This level of transparency feels long overdue. I’ve seen vendors market models trained on just 200 patient scans as clinically validated. For AI to earn trust as a medical device, full visibility into model development is non-negotiable.

  1. Clinical Evidence: Quantitative, Globally Representative Trial Standards

Clinical trial design is far more prescriptive under WHO PQ, with guardrails tailored to public health screening use cases:

  • Training and clinical test datasets must be fully independent — basic scientific guardrail to prevent performance manipulation.
  • Trial data must originate from a minimum of two distinct WHO regions and reflect LMIC clinical workflows, disease prevalence and local clinician skill levels. This creates an inherent tension: the spec also requires device performance to align with expectations across every market where it’s sold. In practice, running trials in every target country is financially unfeasible, leaving manufacturers reliant on risk-based justification frameworks that lack formal standardized guidance.
  • Fixed sample sizing for case-control TB screening trials: minimum 260 confirmed TB positive scans and over 2,500 negative control scans. These figures are mathematically defined to prove 10% non-inferiority vs human radiologists, with clear statistical power and acceptable sensitivity/specificity error margins — a huge simplification for trial design.
  • Gold-standard comparison mandate: accuracy must be validated against microbiological reference standards (plus clinical standards where possible). Older regional rules only required comparison against human readers; this update drastically improves diagnostic credibility.
  • Flexible acceptance of well-designed retrospective case-control data to supplement prospective trial results, a major win for startups with existing clinical archives.

Several performance metrics lack clear quantitative benchmarks, creating ambiguity for manufacturers and buyers alike:

  1. Multiple diagnostic thresholds (0.35 / 0.5 / 0.75) are permitted for different screening priorities, with ROC curve reporting required — yet the spec does not clarify if distinct thresholds demand separate clinical validation.

  2. Turnaround time and throughput (CXRs processed per hour) must be documented, but no minimum/maximum benchmarks exist. A device delivering results in 1 second and one taking 60 seconds both meet compliance on paper, making cross-product comparisons difficult for institutional purchasers like WHO.

  3. Usability & Human Factors: Built for Non-Specialist, Low-Resource End Users

Usability testing requirements depart sharply from standard IEC 62366 workflows designed for hospital radiology teams:

  • The primary user persona is defined as a non-radiologist with basic computer literacy and limited image capture training. Validation must demonstrate minimal training is enough to operate the tool safely and accurately.
  • All human factors testing must take place across at least two separate LMICs.
  • Unique mandate: usability reports must document the diagnostic threshold used for that software version, plus the methodology used to select it — a layer of analysis rarely required in standard usability submissions.

Unresolved Gaps in the 2025 TB CAD Specification

This framework is an early iteration for global health AI, so intentional flexibility leaves several open questions for manufacturers:

  1. No minimum quantitative thresholds for training or test dataset volume
  2. No defined sample sizes for usability studies beyond the dual-LMIC location rule
  3. No formal change control process for expanding a model to detect new pathologies
  4. No standardized performance cutoffs for diagnostic speed and processing throughput

These are not flaws; WHO appears to have reserved flexibility to accommodate diverse AI modalities and variable regional deployment constraints. That said, the ambiguity creates planning uncertainty for teams drafting certification roadmaps.

Key Advantages for Early PQ Applicants

Teams moving forward with submissions now carry meaningful first-mover benefits:

  1. Retrospective clinical datasets are accepted as supplementary evidence, with pathways to add gold-standard validation data post-submission
  2. Guidance is still evolving, with more flexible regulatory interpretation for early adopters as the framework matures

Lessons Learned From Our Gap Assessment

I owe tremendous thanks to the Gates Foundation team — Bilal Mateen (Chief AI Officer), Guang Gao (Senior Technical Officer), and all supporting experts — alongside Dr. Tang, whose guidance shaped every step of our gap analysis. This project rewrote my perspective on building global health AI:

  1. The specification cross-references a comprehensive library of international standards for every development stage. Prior to this work, I underinvested in aligning our full AI lifecycle with global regulatory benchmarks, focusing only on raw model performance.
  2. Quantifiable, scenario-specific clinical requirements eliminate subjective trial design and make validation outcomes far more defensible for public health buyers.
  3. Usability for global health is a separate discipline entirely. Designing tools for specialist radiologists in urban tertiary hospitals does not translate to low-resource clinics, and rethinking end-user capabilities reshaped our entire product design mindset.

Closing Thoughts

WHO Prequalification is more than another regulatory box to tick. It acts as a mirror, reflecting what the world actually needs from medical AI: resilient, accessible diagnostic tools built for fragmented infrastructure, untrained frontline staff, and diverse patient populations across LMICs.

The 2025 TB CAD specification represents a landmark step toward standardized, equitable global health AI. Any team building medical devices for global public health — not just wealthy commercial markets — should begin mapping their PQ compliance strategy today.


메타데이터
post_id
dbf04bd3f3b5
slug
what-who-prequalification-taught-me-about-global-health-ai-dbf04bd3f3b5
url
https://medium.com/@nina-sun/what-who-prequalification-taught-me-about-global-health-ai-dbf04bd3f3b5
canonical_url
https://medium.com/@nina-sun/what-who-prequalification-taught-me-about-global-health-ai-dbf04bd3f3b5
author_url
https://medium.com/@nina-sun
status
ok
fetched_at
2026-07-18 22:42:46