Performance evaluation is the process by which an IVD manufacturer proves, with clinical evidence, that a device achieves its intended purpose. Under the In Vitro Diagnostic Regulation (EU) 2017/746, the IVDR, it is the evidentiary core of the technical documentation, the part of the file where notified bodies spend the most review time, and the part where the largest share of deficiency findings arise.
This article explains the whole system in plain language: the three pillars of clinical evidence, what the performance evaluation plan (PEP) must contain, when you need new studies and when literature can carry the argument, what goes into the performance evaluation report (PER), and how the June 2025 MDCG guidance changed practical expectations. At the end you will find a complete PEP structure you can build your own plan against, which is the practical answer to the question most people searching for a "PEP template" are really asking.
This is the companion piece to our full guide to IVDR technical documentation, which covers the complete Annex II and III file that the performance evaluation sits inside.
The legal skeleton in one paragraph
Article 56 of the IVDR requires performance evaluation to be planned, conducted, and documented, and to be updated continuously throughout the device's life. Annex I, particularly its performance-related requirements, defines the characteristics the evidence must support. Annex XIII Part A defines what the plan and the report must contain. Two MDCG guidance documents matter most in practice: MDCG 2022-2 on the general principles of clinical evidence for IVDs, and MDCG 2025-5, the question and answer document on performance studies published in June 2025, which settled several long-argued practical points. Everything below unpacks these sources.
The three pillars of clinical evidence
The IVDR requires clinical evidence to be built from three components, and all three must be addressed in the PEP and critically analyzed in the PER. Understanding what each pillar asks, and what kind of evidence answers it, is the foundation of the whole exercise.
Scientific validity asks: is the analyte or marker genuinely associated with the clinical condition or physiological state you claim? For troponin and myocardial infarction the association is textbook knowledge and the demonstration is short. For a novel biomarker the association itself must be established, through peer-reviewed literature, proof of concept studies, or expert consensus. The output is a scientific validity report. The work behind it is usually a documented literature review: named databases, a stated search strategy, inclusion and exclusion criteria, and critical appraisal of what was found, with gaps acknowledged openly. Where the association is not yet established, the gap becomes a study, and analytical excellence cannot substitute for it: if the marker's link to the condition is unproven, accurate measurement of the marker proves nothing about the condition.
Analytical performance asks: does the device measure or detect the analyte correctly? This covers analytical sensitivity and specificity, trueness and precision, limits of detection and quantitation, measuring range, linearity, interference and cross-reactivity, and related characteristics, each as applicable to the technology and claims. The demonstration is through analytical performance studies of your device, run under documented protocols with pre-set acceptance criteria. MDCG 2025-5 confirmed what notified bodies had already been enforcing: analytical performance is expected to be demonstrated through study data in essentially all cases. Published literature describes other devices; it cannot show that your device, with your reagents and your calibration, performs.
Clinical performance asks: do the device's results correspond to the clinical condition in the intended population? This is where clinical sensitivity and specificity, predictive values, likelihood ratios, and expected values in normal and affected populations live. Unlike analytical performance, the IVDR allows three evidence routes here: clinical performance studies, peer-reviewed published literature, and published experience gained by routine diagnostic testing. Which route is defensible for your device is the single most consequential judgment in the entire performance evaluation, and it gets its own section below.
The performance evaluation plan: what it must contain
The PEP is the document that turns those three pillars into a concrete program for your specific device. Annex XIII Part A, supported by MDCG 2022-2, defines its content. A complete PEP addresses:
The device and its claims. The device characteristics and intended purpose: the analyte or marker, the function of the test (screening, diagnosis, monitoring, prognosis, prediction, companion diagnosis), the intended testing population, the specimen type, the intended user, and the clinical claims made. Every performance characteristic the plan will evidence traces back to a claim stated here.
The applicable GSPRs. The plan must clearly identify which general safety and performance requirements of Annex I the performance evaluation will support. This linkage is what later lets your GSPR checklist point to the PER as evidence.
The state of the art. A description of the current state of the art in the relevant medical, diagnostic, and technological fields: how the condition is currently diagnosed, what comparator methods exist, what performance the field currently achieves, and where your device sits in that landscape. This section does real work: your acceptance criteria must be defensible against it, and reviewers check that they are.
Standards and specifications. Reference to the standards, common specifications, guidelines, and best practice documents the evaluation will apply, cited by version. For most IVDs this means the relevant CLSI protocols for analytical studies, applicable ISO standards, and any common specifications for the device category, which for Class D devices are binding.
The methods and acceptance criteria, per pillar. For each of scientific validity, analytical performance, and clinical performance: how the evidence will be generated or gathered, the methods and their justification, and pre-defined acceptance criteria. This includes the parameters and criteria by which the acceptability of the benefit-risk ratio will be determined, set in advance rather than derived from whatever the data happens to show.
The evidence generation sequence. Annex XIII structures the evaluation as a three-step process: define the methods and acceptance criteria, conduct the evaluation and collect the data, then analyze the data and draw conclusions in the report. The plan should present the program in that logic, including which studies will be run, what literature searches will be performed and how, and how the strands combine.
The PMPF plan linkage. Performance evaluation continues after market entry through post-market performance follow-up. The PEP should state how PMPF will feed the evaluation, and for Class C and D devices this is a live obligation with an annual reporting rhythm, covered below.
Update discipline. Who owns the plan, when it is reviewed, and what triggers an update: design changes, new state of the art, new claims, post-market signals.
Two practical notes on the plan. First, performance evaluation interfaces directly with risk management, and the two files must agree: the risks your ISO 14971 file accepts must match the performance your evaluation demonstrates, and reviewers read them side by side. Second, MDCG 2025-5 clarified a procedural point that confused many manufacturers running performance studies: the PEP does not have to be included in a performance study application, but it must exist and be provided to the competent authority on request. A plan drafted retrospectively, after the studies it was supposed to govern, shows its origins in dozens of small ways, and reviewers are practiced at spotting them.
Studies, literature, or routine-use data: choosing your clinical performance route
Because clinical performance admits three evidence routes, the route decision shapes cost, timeline, and risk for the whole project. A plain framing of the decision:
A clinical performance study of your device is the regulation's starting point: Article 56(4) requires clinical performance studies to be carried out unless reliance on other sources of clinical performance data is duly justified. The burden therefore sits on the manufacturer choosing a different route, and it weighs heaviest when the device is novel, when the claimed population or specimen type is not well represented in existing evidence, or when you claim performance better than the state of the art. Studies in the EU bring Annex XIII and Annex XIV procedural requirements with them, including sponsor duties, and for certain categories, applications to competent authorities and ethics committees. MDCG 2025-5 is now the first reference for how those procedures work in practice, from what an application must contain to how modifications are handled.
Peer-reviewed literature can carry clinical performance where the association between your device type and clinical outcomes is well established and the published evidence is applicable to your device and claims. The strength of this route rests entirely on applicability. The published studies were run on other devices, so the argument requires demonstrating that those devices are comparable to yours in the respects that matter, and that the studied populations, specimen types, and clinical settings match your claims. A literature route built on "similar technology" assertions without a structured comparability analysis is among the most common deficiency findings in IVDR reviews.
Published experience from routine diagnostic testing supports claims for long established tests, the assays that laboratories have run for decades with documented external quality assessment history. Even here the evidence must be compiled, appraised, and connected to your specific device, and gaps must be acknowledged.
In all three cases the choice must be justified in writing, and the justification belongs in both the PEP and the PER. The factors the regulation expects that justification to weigh are the device's intended purpose, its degree of novelty (a novel test, an established but non-standardized test, or an established and standardized test), and the nature, severity, and evolution of the condition being diagnosed. A useful internal test: if your justification would read as reasonable to a skeptical assessor who has just read your intended purpose statement and your state of the art section, the route will probably survive review. A justification whose real content is that a study would be expensive will be read as exactly that.
The performance evaluation report: where everything converges
The PER is the synthesis document. Annex XIII defines its content, and its logic mirrors the plan: it opens with the justification for the approach taken to gather the clinical evidence, then presents the scientific validity report, the analytical performance evidence, and the clinical performance evidence, each critically appraised rather than merely summarized, together with the literature search methodology, protocol, and report for any literature-based strand, and closes with conclusions on the clinical evidence and the benefit-risk determination against the criteria the plan set in advance.
Critical appraisal is the quality bar that separates acceptable PERs from rejected ones. A report that only lists studies and quotes their results has compiled evidence without evaluating it. A report that assesses the quality and applicability of each evidence item, weighs contradictions, states what the totality of evidence supports and what it does not, and connects the conclusion to the specific claims in the intended purpose statement has performed the evaluation the regulation asks for. Notified body reviewers are trained to spot the difference, and the most frequent instruction in deficiency letters on PERs is a version of "analyze, do not summarize."
The PER is also a living document. Article 56 requires the performance evaluation and its documentation to be updated throughout the device's life cycle, and for Class C and D devices the report is updated with PMPF findings when necessary and at least annually. An update is an actual re-appraisal: new PMPF data, new literature, new vigilance signals, and a refreshed benefit-risk conclusion, with the state of the art section revisited often enough that it still describes the present.
Post-market performance follow-up: the loop that keeps evidence current
PMPF is the continuation of performance evaluation after market entry: the planned, proactive collection of performance data from real use. Its plan sits in the technical documentation under Annex III alongside the broader post-market surveillance system, and its findings flow back into the PER on the update cycle above. Methods range from ongoing specimen testing programs and external quality assessment participation to registry data and structured user feedback, chosen to match the questions that remain open after pre-market evaluation: rare interferents, performance at claim boundaries, performance in population subgroups underrepresented pre-market.
For a first certification, reviewers assess the PMPF plan prospectively, on its credibility as a protocol: named methods, named data sources, defined analyses, defined thresholds for action. At recertification and surveillance audits they assess what it produced, and a PMPF plan that generated no findings, no analyses, and no PER updates across a certification cycle tells the reviewer the system exists on paper.
Where performance evaluations fail review
From notified body communications and our own remediation work, the recurring failure points:
- Acceptance criteria set after the data existed, or not set at all, leaving the benefit-risk conclusion without a pre-defined standard to be judged against.
- Scientific validity treated as a citation dump for well-known markers, or skipped as obvious, rather than documented through a described search and appraisal.
- Analytical performance gaps against claims: specimen types on the label never studied, interference panels that ignore the intended population's common medications, precision estimated away from the clinical decision point.
- Literature routes without comparability analysis, as described above.
- State of the art sections written once and frozen, so acceptance criteria defensible in 2021 read as below current practice in 2026.
- PERs that summarize rather than appraise.
- PMPF justified away for Class C devices on grounds reviewers no longer accept.
- Plan, report, risk file, and IFU telling inconsistent stories about the same performance characteristics.
Every one of these is findable in an internal review before submission, and finding them internally costs days, where a deficiency round costs months.
Legacy IVDs: the specific problem of directive-era evidence
Manufacturers transitioning IVDD-era devices face a structural gap: the directive never demanded a three-pillar performance evaluation, so the historical file typically contains verification data of varying vintage and a clinical justification written to a lower standard. The transition work is a gap analysis of that material against Annex XIII, and the commonest findings are missing scientific validity documentation, analytical studies that predate current CLSI protocols in ways that matter, and clinical evidence that consists of the original launch-era evaluation plus accumulated but uncompiled routine-use experience. The uncompiled experience is often the recoverable asset: years of EQA participation and troubleshooting records can become genuine clinical performance evidence for a well-established assay if someone structures and appraises them. Time this work against the transition calendar covered in our technical documentation guide, because for Class C devices the written agreement milestone of 26 September 2026 has already made evidence gaps urgent, and for Class B the May and September 2027 milestones leave barely enough time for a study if a gap analysis this year finds one is needed.
Frequently asked questions
Is there an official PEP template? No. Annex XIII Part A defines the required content and MDCG 2022-2 elaborates on it. The structure in this article, and the outline below, follow that content in the order reviewers expect. Any template you adopt should be checked line by line against Annex XIII rather than trusted on its title.
Do we need a PEP for a Class A device? Yes. Performance evaluation applies to all classes, scaled to the device's risk and claims. For a simple Class A product the plan and report can be short documents, and they must still exist, address all three pillars proportionately, and live in the technical documentation.
Does the PEP go into a performance study application? MDCG 2025-5 answered this directly: the IVDR does not require the PEP to be included in the application, but it must be available and provided to the competent authority on request. Write it before the studies it plans.
Can we CE mark with a provisional PER while a study finishes? The clinical evidence supporting your claims must be complete at conformity assessment. What can legitimately remain open is the PMPF program: questions that post-market data will answer are stated as such, with the plan to answer them. Claims that depend on the unfinished study cannot be made until it is finished.
Who should write the PEP and PER? Reviewers expect the evaluation to be demonstrably the work of suitably qualified people, and in practice the documents need three competencies at the table: the clinical and scientific understanding of the analyte and condition, the statistical and methodological competence to set and assess acceptance criteria, and the regulatory understanding of what Annex XIII and the notified body expect. One person rarely holds all three, which is why performance evaluation is usually a team output with a named responsible author.
How does this differ from a clinical evaluation report (CER)? CERs belong to medical devices under the MDR; performance evaluations belong to IVDs under the IVDR. The philosophy is parallel and the structures differ: the CER is built on clinical data and equivalence logic under MEDDEV and MDCG guidance, while the IVD performance evaluation is built on the three pillars described here. Teams carrying both product types should resist reusing structure across the two regimes, because reviewers on each side notice borrowed logic from the other.
A working PEP outline
Download this outline as a PDF, the ten Annex XIII sections in review order with drafting reminders from MDCG 2022-2 and MDCG 2025-5.
Download the PEP outline (PDF)The structure below covers the Annex XIII content in review order. Sections 1 to 4 set the frame, 5 to 8 define the program, 9 and 10 keep it alive.
- Device description, intended purpose, and claims (analyte, function, population, specimen, user)
- Applicable GSPRs the evaluation supports
- State of the art: condition, current diagnostic practice, comparator methods, expected performance levels
- Standards, common specifications, and guidance applied, with versions
- Scientific validity: existing evidence, search and appraisal methodology, gaps, and any studies planned
- Analytical performance: characteristics to be demonstrated, study protocols and methods, acceptance criteria
- Clinical performance: evidence route with written justification, studies or literature/routine-data methodology, acceptance criteria
- Benefit-risk: parameters and pre-defined criteria for acceptability determination
- PMPF: methods, data sources, analyses, and how findings update the evaluation
- Governance: responsibilities, qualifications, review triggers, and update schedule, with the PER update cycle stated (at least annual for Class C and D)
How EvySaif supports performance evaluation
EvySaif Research and Medical Affairs Solutions writes and remediates performance evaluation documentation for IVD manufacturers: PEPs and PERs built to Annex XIII, scientific validity reports with documented literature methodology, clinical performance route justifications, PMPF plans and reports, and gap analyses of legacy files against the transition calendar. The same team supports the surrounding technical documentation, so the performance evaluation stays consistent with the risk file, the GSPR checklist, and the IFU.
If your PEP needs writing, your PER has come back with findings, or your legacy file needs a clear-eyed gap read against a 2026 or 2027 milestone, write to info@evysaif.com or use the contact page.
This article reflects the position as of 12 August 2026, including IVDR Annex XIII, MDCG 2022-2, and MDCG 2025-5. Regulatory expectations evolve; verify current requirements for your specific device before acting.