AI/ML SaMD

V&V Testing Strategy for FDA-Regulated Machine Learning Models in Medical Devices

By Andre D. Butler, Principal Consultant  ·  reviewed September 2026  ·  ← All Insights

Verification and validation testing strategy for FDA-regulated machine learning models.

Photo by Growtika on Unsplash

Why Your AI/ML V&V Strategy Can Make or Break FDA Clearance

If you are developing a machine learning-based medical device, your verification and validation (V&V) testing strategy is not a formality -- it is the backbone of your entire regulatory submission. FDA reviewers are scrutinizing AI/ML submissions more carefully than ever, and companies that treat V&V as a checkbox exercise are getting burned with additional information requests, refuse-to-accept decisions, and costly delays.

At ADB Consulting & CRO Inc., we work with medical device startups and established manufacturers navigating the intersection of artificial intelligence and FDA regulation. This post breaks down exactly what a defensible, submission-ready V&V strategy looks like for AI/ML-enabled Software as a Medical Device (SaMD).

The Regulatory Foundation You Cannot Ignore

Before you write a single test protocol, you need to anchor your strategy in the right regulatory framework. For AI/ML SaMD, the primary references include:

  • 21 CFR Part 820 -- FDA's Quality System Regulation (and its updated 2024 alignment with ISO 13485), which governs design controls including verification and validation requirements under 21 CFR 820.30.
  • FDA's 2021 Action Plan for AI/ML-Based SaMD -- outlines FDA's expectations for transparency, real-world performance monitoring, and the Predetermined Change Control Plan (PCCP).
  • FDA Guidance: 'Software as a Medical Device (SaMD): Clinical Evaluation' (2017, IMDRF) -- defines the analytical validation, clinical validation, and clinical performance hierarchy that reviewers use to evaluate your evidence package.
  • FDA's 2022 Draft Guidance on AI/ML-Based SaMD -- directly addresses lifecycle considerations, dataset management, and performance transparency.
  • IEC 62304 -- software lifecycle processes that define how your ML development process must be documented and traced.

If your V&V plan does not reference these documents, it is already behind.

Verification vs. Validation: Getting the Distinction Right for ML

Traditional device verification asks: did we build the product correctly? Validation asks: did we build the correct product? For machine learning, this distinction becomes more nuanced -- and more consequential.

Verification for ML models typically encompasses unit testing of preprocessing pipelines, integration testing of model inference modules, code review and static analysis, and performance benchmarking against defined acceptance criteria on a locked test dataset.

Validation for ML models must demonstrate that the algorithm performs its intended clinical function on a representative patient population. This is where most submissions fall short. Your validation dataset must be:

  • Statistically powered and prospectively defined before testing begins
  • Demographically representative of your intended use population
  • Clinically labeled by qualified annotators with a documented adjudication process
  • Separated from any data used in training or tuning -- no data leakage, period

Building Your Test Dataset Strategy

FDA has made clear that dataset integrity is non-negotiable. Under the IMDRF SaMD clinical evaluation framework, you need to distinguish between analytical validation (does the algorithm produce accurate outputs?) and clinical validation (do those outputs translate to meaningful clinical outcomes?).

For your locked test set, document the following in your Design History File (DHF) under 21 CFR 820.30(j):

  • Site diversity -- single-site data will draw scrutiny for generalizability concerns
  • Device diversity -- if applicable, variability across imaging equipment, sensor hardware, or acquisition protocols
  • Subgroup analysis plan -- FDA expects performance breakdowns by sex, age, race, and any clinically relevant subpopulations
  • Reference standard definition -- how ground truth was established, and what the inter-rater agreement was

Defining and Defending Your Performance Metrics

Choosing sensitivity and specificity alone is rarely sufficient for AI/ML submissions. Your choice of performance metrics must be clinically justified and tied directly to your intended use statement. Area under the ROC curve, positive predictive value, and calibration curves each tell a different story, and reviewers will want to see that you understand which metrics matter for your specific risk profile.

Critically, your acceptance criteria must be pre-specified. Do not analyze your results first and then set your thresholds. FDA reviewers are trained to identify post-hoc threshold selection, and it is one of the fastest ways to lose credibility in a submission.

Predetermined Change Control Plans: Planning for Model Updates

One of the most underutilized tools in AI/ML regulatory strategy is the Predetermined Change Control Plan (PCCP). A well-constructed PCCP allows you to pre-negotiate with FDA the types of model updates -- retraining on new data, threshold adjustments, performance drift corrections -- that will not require a new 510(k) or PMA supplement.

If you plan to continuously learn or periodically retrain your model post-market, you need a PCCP. Without it, every meaningful algorithm change triggers a new submission cycle.

What Your DHF Must Include

Under 21 CFR 820.30, your Design History File must trace every V&V activity to your design inputs and design outputs. For AI/ML, this means your DHF should contain:

  • A Software Requirements Specification that includes algorithm performance requirements with measurable acceptance criteria
  • A V&V Plan and Report for both the software and the ML model specifically
  • Dataset provenance records, including data use agreements and IRB documentation where applicable
  • Model cards or algorithm transparency documentation aligned with FDA expectations
  • Real-World Performance Monitoring plan for post-market surveillance under 21 CFR Part 822

Common Mistakes That Sink AI/ML Submissions

After reviewing dozens of AI/ML submissions and supporting responses to FDA Additional Information requests, the same failure patterns appear repeatedly:

  • Test sets that overlap with training data due to patient-level (not scan-level) splitting errors
  • Acceptance criteria defined after seeing test results
  • No subgroup analysis, or subgroup analyses that reveal unacknowledged performance disparities
  • Validation populations that do not reflect the labeled intended use population
  • Missing or incomplete IEC 62304 software lifecycle documentation

Each of these issues is avoidable with the right strategy built into your development process from day one -- not retrofitted at submission time.

Work With Experts Who Know AI/ML Regulatory Strategy

Developing a compliant V&V strategy for a machine learning medical device requires regulatory expertise that goes beyond standard software device experience. The frameworks are newer, the expectations are evolving, and the cost of getting it wrong -- in time, capital, and market opportunity -- is significant.

ADB Consulting & CRO Inc. helps medical device companies build V&V strategies that satisfy FDA reviewers and accelerate the path to clearance or approval. Whether you are preparing a 510(k), De Novo, or PMA submission for an AI/ML-enabled device, we can help you get it right the first time.

Book a free discovery call with Andre Butler and the ADB Consulting team today at adbccro.com. Let us map out a regulatory strategy that fits your device, your timeline, and your budget.

If this applies to your program, our AI/ML SaMD regulatory strategy guide walks through the process in detail.

Andre Butler

Principal Consultant — ADB Consulting & CRO Inc.

Andre Butler has 20+ years of hands-on FDA regulatory experience guiding medical device companies through 510(k), PMA, De Novo, AI/ML SaMD, and FDA 483 response engagements. He specialises in Section 524B cybersecurity compliance and ISO 13485 quality management systems, with a track record across cardiovascular, orthopedic, diagnostic, and software-as-a-medical-device categories.

Ready to Navigate the FDA Process with Confidence?

Book a free 30-minute discovery call with Andre Butler. No sales pitch -- just expert regulatory guidance on your specific device and situation.

Book a Free Pathway Call

Or call directly: (888) 450-8607

Explore our flat-fee FDA services →