How direct and indirect trial evidence combine into one network: assumptions, Bayesian and frequentist methods, SUCRA rankings, MAIC, STC and HTA use.
When several treatments exist for the same condition, the trials rarely compare all of them with each other. Drug A was tested against placebo, Drug B against placebo, Drug C against Drug B, and the comparison a payer or clinician actually needs, A against C, was never run. Network meta-analysis (NMA) is the evidence-synthesis method built for this situation: it combines the direct and indirect evidence from a connected network of randomized trials to estimate the relative effects of every treatment in the network against every other.
This guide covers what an NMA is and how it works, the three assumptions that decide whether its results can be trusted, how it differs from pairwise meta-analysis and from population-adjusted methods such as MAIC and STC, the Bayesian and frequentist frameworks, the conduct of an NMA from question to report, how to read the outputs including SUCRA rankings, and where NMA sits in health technology assessment (HTA) and market access. It is written for pharmaceutical, biotech and medical affairs teams deciding whether an NMA is the right method for their evidence, and what a defensible one requires.
1. What a network meta-analysis is
A conventional pairwise meta-analysis pools trials that all make the same comparison: several trials of Drug A against placebo yield one combined estimate of A versus placebo. An NMA analyzes multiple treatments at once inside a single connected evidence network. Suppose the trial evidence is:
| Trial | Comparison |
|---|---|
| Study 1 | Drug A vs placebo |
| Study 2 | Drug B vs placebo |
| Study 3 | Drug C vs placebo |
| Study 4 | Drug A vs Drug B |
| Study 5 | Drug B vs Drug C |
No trial compares A with C. Because every treatment is connected to the others through shared comparators, the network can estimate A versus C indirectly, and it can combine that indirect estimate with the direct evidence wherever both exist. The output is a full set of relative effects: every treatment against every other, each with its uncertainty.
The estimates depend entirely on the trials behind them: studies connected mathematically still have to be similar enough in the ways that matter, which is the subject of section 4.
2. Direct, indirect and mixed evidence
Direct evidence comes from trials that compared two treatments head to head. Indirect evidence comes from a shared comparator: with A versus placebo and B versus placebo, the difference between the two trial effects estimates A versus B. The simplest version of this is the Bucher adjusted indirect comparison, which subtracts the two placebo-anchored effects on the appropriate scale (log odds ratio, log hazard ratio) and combines their variances. Anchoring on the common comparator preserves within-trial randomization: the indirect estimate is built from randomized contrasts, never from naive comparison of single arms across trials.
Mixed (network) evidence combines both. When a comparison has direct and indirect evidence, the NMA pools them, and the two sources can also be compared with each other, which is the consistency check in section 4. An NMA is the generalization of the Bucher comparison to a whole network of treatments, with the statistics handling every loop and multi-arm trial at once.
3. When NMA fits and when it does not
NMA is the right tool when the decision needs comparisons across several treatments, the trials are randomized, the interventions form a connected network, and the trials are similar enough in populations, comparators, outcomes and settings for indirect comparison to be meaningful. It answers the payer's and the guideline writer's question of how the options compare with each other.
It is the wrong tool, or needs modification, when:
- Populations differ materially across comparisons in ways that modify treatment effects: disease severity, line of therapy, background treatment, prior exposure.
- Doses, regimens or formulations differ so much that treatments would have to be lumped into nodes that hide real differences.
- Outcome definitions or follow-up times are not compatible across trials.
- The network is disconnected: a treatment with no randomized link to the rest cannot be compared through NMA at all, which is one of the situations that leads to the population-adjusted methods in section 7.
- The evidence is sparse, with single small trials carrying whole connections.
In these situations a pairwise meta-analysis, a population-adjusted comparison or a decision not to synthesize may serve the question better, and the method is chosen for the question and the evidence, never for the availability of software.
4. The three assumptions: transitivity, heterogeneity, consistency
Transitivity is the requirement that trials making different comparisons are similar in the factors that modify treatment effects. If the A-versus-placebo trials enrolled younger patients with mild disease and the B-versus-placebo trials enrolled older patients with severe disease, and age or severity changes treatment response, the indirect A-versus-B estimate inherits that difference as bias. Transitivity is a clinical and epidemiological judgment made before modeling: the candidate effect modifiers are listed, their distributions are tabulated across the trials, and the comparability case is written down, because effect-modifier imbalance cannot be corrected afterward by the model.
Heterogeneity is variation in treatment effects between trials of the same comparison, arising from differences in populations, interventions, outcome measurement, follow-up and setting. It is quantified (tau-squared, I-squared) and handled in the model, usually through random effects with a heterogeneity parameter that in many NMA models is shared across comparisons. High unexplained heterogeneity widens the intervals and weakens every downstream conclusion, and meta-regression or subgroup analysis can explore its sources when the data allow.
Consistency is the statistical counterpart of transitivity: where a comparison has both direct and indirect evidence, the two should agree within chance. It is tested locally by node-splitting, which separates the direct and indirect estimates for each comparison and compares them, and globally by approaches such as the design-by-treatment interaction model. A significant inconsistency is a finding to investigate, in the trials and the effect modifiers, before any ranking is reported.
5. Network geometry
The network plot, treatments as nodes and direct comparisons as edges with thickness proportional to the number of trials, is read before any estimate. A star-shaped network, in which every drug was tested against placebo and nothing else, produces comparisons between active treatments that rest entirely on indirect evidence. A well-connected network with closed loops supplies direct evidence for many comparisons and allows consistency to be tested. Thin edges show where a single trial carries a connection, which is where sensitivity analysis matters most. Geometry also exposes lumping decisions: whether doses of the same drug are separate nodes or one, which changes both the network and the question it answers.
6. NMA compared with pairwise meta-analysis
| Feature | Pairwise meta-analysis | Network meta-analysis |
|---|---|---|
| Comparisons | One per analysis | All treatments in the network |
| Evidence used | Direct only | Direct and indirect combined |
| Indirect comparisons | No | Yes |
| Key extra assumptions | Similarity of pooled trials | Transitivity and consistency across the network |
| Multi-arm trials | Arms selected per comparison | Handled jointly with within-trial correlation |
| Treatment ranking | Not an output | Rankings and ranking probabilities available |
| Typical use | One question, one comparison | Guideline, HTA and market-access questions across options |
A pairwise meta-analysis remains the right choice when the decision genuinely concerns one direct comparison; the NMA's extra assumptions are taken on when the question spans the network.
7. ITC, MAIC and STC: the population-adjusted alternatives
An indirect treatment comparison (ITC) is the two-step Bucher case; an NMA extends it across a network. Both assume the trial populations are comparable. When they are not, or when the network is disconnected, the population-adjusted methods described in NICE Decision Support Unit Technical Support Document 18 come into play.
A matching-adjusted indirect comparison (MAIC) applies when a company holds individual patient data for its own trial but only published aggregate data for the comparator trial. The patient-level data are reweighted so that the weighted trial matches the comparator trial's reported baseline characteristics, and the comparison is then made between the reweighted arm effects. A simulated treatment comparison (STC) fits an outcome regression in the patient-level data and uses it to predict outcomes at the comparator trial's population profile. Both come in anchored form, through a common comparator, which preserves randomization for the adjusted contrast, and unanchored form, a direct cross-trial comparison of arms, which relies on the far stronger assumption that all prognostic factors and effect modifiers have been captured, and which HTA committees treat with corresponding skepticism.
The selection among NMA, ITC, MAIC and STC follows from three facts: how connected the network is, whether the populations are comparable, and what data (patient-level or aggregate) exist on each side.
8. Bayesian and frequentist NMA
Both frameworks estimate the same relative effects. A frequentist NMA, typically fitted with the R package netmeta using graph-theoretical or multivariate methods, reports point estimates with confidence intervals and P-scores for ranking. A Bayesian NMA, fitted in gemtc or multinma in R, or directly in BUGS, JAGS or Stan, reports posterior distributions with credible intervals, yields ranking probabilities and SUCRA naturally, accommodates complex likelihoods (multi-arm trials, shared parameters, time-to-event structures) flexibly, and can incorporate prior information, most usefully an informative prior on the heterogeneity parameter when the network has few trials per comparison. In Bayesian models the priors and their sensitivity are reported, because in a sparse network a vague heterogeneity prior can drive the width of every interval.
HTA practice accepts both. The NICE Decision Support Unit's Technical Support Documents 1 to 7, which remain the reference methods series for evidence synthesis in submissions, present the general linear modeling framework in Bayesian form, and many submissions follow it. The framework is chosen for the structure of the data and the requirements of the decision problem, and the choice is justified in the protocol.
9. How an NMA is conducted
An NMA is a systematic review with a network model on top, and most of its quality is fixed before any model runs.
- Question. The decision problem is framed in PICO terms, with particular care on the intervention definitions: whether doses, formulations and regimens are separate nodes is settled here, because it defines the network.
- Protocol. Eligibility criteria, information sources, search strategy, outcomes and time points, risk-of-bias method, the synthesis model, and the planned consistency, subgroup and sensitivity analyses are pre-specified. For publication-bound work the protocol can be registered, and the report will follow PRISMA-NMA, the network extension of the PRISMA reporting statement.
- Search. A systematic literature search across bibliographic databases, trial registries, regulatory documents and conference sources identifies the trials; for HTA use the search is built to the assessing agency's expectations.
- Selection and extraction. Screening against the criteria, then extraction of study characteristics, populations, interventions, outcome data and effect estimates, in enough detail to support the transitivity assessment, with the process documented for reproducibility.
- Risk of bias. Each trial is assessed, typically with the Cochrane RoB 2 tool for randomized trials. Bias in the trials propagates through every estimate the network produces.
- Network construction and transitivity assessment. The geometry is drawn, effect-modifier distributions are tabulated across comparisons, and the case for synthesis is made or the plan is revised.
- Analysis. The pre-specified model is fitted: fixed-effect and random-effects versions, the appropriate likelihood for the outcome (binomial for events, normal for continuous measures, contrast-based models for hazard ratios), correct joint handling of multi-arm trials, then node-splitting and global inconsistency checks, followed by the planned meta-regression, subgroup and sensitivity analyses.
- Confidence and reporting. Confidence in the estimates is rated with CINeMA or the GRADE approach for NMA, and the report presents the network, the assumptions and checks, the effects with uncertainty, and the rankings with their caveats, in PRISMA-NMA structure for journals or in the dossier structure for HTA.
10. Reading the results: effects, intervals, rankings and SUCRA
The primary output is the set of relative treatment effects, usually presented as a league table: every pairwise contrast in the network on the outcome's scale (odds ratio, risk ratio, hazard ratio, mean difference), each with its confidence or credible interval. The intervals carry as much information as the point estimates; a first-ranked treatment whose interval against the second crosses the null has not been shown superior to it.
Rankings summarize where each treatment tends to sit. Bayesian analyses report ranking probabilities and SUCRA, the surface under the cumulative ranking curve, a 0-to-100 summary of a treatment's average rank position; frequentist analyses report the analogous P-score. Both compress the whole posterior into one number, and both mislead when read alone: a treatment can top the SUCRA table on a small, imprecise evidence base while a clinically equivalent alternative sits lower with far more certainty. Rankings are read together with the effect sizes, the intervals, the confidence ratings, the safety outcomes and the geometry, and a difference that is statistically clear can still be clinically trivial.
11. NMA in HTA, development and market access
HTA is the setting where NMA is used most. An assessment compares the technology against all relevant comparators in the jurisdiction, and complete head-to-head evidence against every comparator almost never exists. The NICE reference case anticipates synthesis of the available evidence, and the DSU technical support documents define the expected methods; CADTH, IQWiG, HAS and the EU joint clinical assessment process each scrutinize the same points with their own emphases: the systematic review behind the network, the transitivity case, the consistency checks, the handling of heterogeneity, and the sensitivity of results to model and data choices. In India, HTAIn under the Department of Health Research applies the same synthesis logic in its assessments. An NMA built for an HTA submission is therefore constructed to the assessing agency's methods expectations from the protocol onward, which is a different build from a journal manuscript even when the network is the same.
The same analysis feeds the rest of the value story. The relative effects become the clinical inputs of the cost-effectiveness model, the comparative case in the global value dossier, and the effect sizes behind a budget impact analysis. Earlier in the lifecycle, an NMA of the existing landscape shows a development team how the current options compare, where the evidence gaps sit, and which comparator the next trial has to beat, questions that sit inside clinical development strategy. Where trial evidence cannot reach, real-world evidence carries a complementary part of the picture.
12. Common problems that undermine an NMA
- Lumped nodes. Treating different doses or regimens as one treatment because the network is thin, which changes the question the network answers.
- Transitivity by assertion. A sentence claiming the trials are similar, with no effect-modifier table behind it.
- Rankings without uncertainty. A SUCRA bar chart presented as the conclusion, with the intervals and confidence ratings left in an appendix.
- Inconsistency ignored. Node-splitting either not run or run and not discussed when it flags disagreement.
- A weak review under a strong model. Sophisticated modeling on a network missing trials that a proper systematic search would have found.
- Method chosen before the question. An NMA run because the software was available, where the population differences called for MAIC or for no synthesis at all.
- Statistical significance read as clinical importance. A precise but trivial difference promoted to a treatment recommendation.
13. Network meta-analysis cost and timeline drivers
There is no standard price or turnaround for an NMA because the work scales with the review and the network, and both are set by the question. The drivers on both cost and time are the same: the number of interventions, outcomes and time points; the size of the literature and the number of databases searched; risk-of-bias workload; network complexity and the modeling plan (Bayesian fitting, meta-regression, subgroup and sensitivity analyses, any population-adjusted add-on); and the reporting destination, since an HTA build carries agency-specific methods and quality-control steps a journal manuscript does not. A contained network of a dozen trials and a review of several thousand records with a multi-outcome Bayesian analysis are different projects, and a scope assessment of the question and evidence base comes before any estimate of either number.
14. Frequently asked questions
A statistical method that combines direct and indirect evidence from a connected network of randomized trials to estimate the relative effects of three or more treatments simultaneously, giving every treatment's effect against every other.
A standard pairwise meta-analysis pools trials of one comparison and uses only direct evidence. An NMA spans all treatments in a connected network, adds indirect evidence, and takes on the transitivity and consistency assumptions in exchange.
An estimate of the effect between two treatments obtained through a shared comparator, classically by the Bucher method: the A-versus-placebo and B-versus-placebo effects are contrasted to estimate A versus B while preserving each trial's randomization.
The assumption that trials making different comparisons are similar in the factors that modify treatment effects. It is assessed clinically, by tabulating effect modifiers across the trials, before any model is fitted.
The surface under the cumulative ranking curve: a 0-to-100 summary of a treatment's average rank in a Bayesian NMA (the frequentist analog is the P-score). It is read alongside the effect estimates, intervals and confidence ratings, never alone.
When trial populations differ in ways that break transitivity, or the network is disconnected, and individual patient data exist for one side of the comparison. MAIC reweights the patient-level trial to match the comparator trial's reported characteristics; the anchored form, through a common comparator, is the stronger design.
Neither by default. Bayesian models (gemtc, multinma, BUGS/JAGS/Stan) handle complex structures and sparse networks flexibly and yield ranking probabilities directly; frequentist models (netmeta) are fast and familiar. HTA practice accepts both, and the choice is justified by the data structure and the decision problem.
Yes, and most submissions rely on one, because complete head-to-head evidence against all relevant comparators rarely exists. Agencies scrutinize the systematic review, the transitivity case, consistency checks and sensitivity analyses, with methods expectations set out in documents such as the NICE DSU technical support series.
No. It synthesizes the trials that exist and quantifies the comparisons they support. Where a regulatory or clinical decision requires direct randomized evidence, the trial still has to be run; the NMA can show which comparison that trial should make.
There is no fixed minimum; feasibility depends on the network being connected and on enough evidence per comparison to estimate effects and heterogeneity. Networks with single-trial connections can be analyzed, with the uncertainty and sensitivity analyses carrying more of the interpretive weight.
15. Work with EvySaif on your next evidence synthesis
EvySaif is a clinician-led HEOR, drug clinical development and medical writing consultancy in Pune, India, working with pharmaceutical, biotech and device companies across India, the Middle East and Europe. Evidence-synthesis projects run end to end: PICO and protocol development, the systematic literature review, risk-of-bias assessment, pairwise meta-analysis and NMA in Bayesian and frequentist frameworks, MAIC and STC where population adjustment is needed, and consistency and confidence assessment. The same team carries the results into manuscripts, HTA evidence packages and economic models, so the analysis is designed once for everything it has to support.
The engagement starts with a method question, and the honest answer is sometimes that an NMA is not the right tool for the evidence at hand. We commit to a defensible method choice, a reproducible analysis with the search, data, decisions and code documented, and a report that states what the evidence shows and where its limits are. If you are weighing a pairwise meta-analysis, NMA, ITC, MAIC or STC for a development, publication, HTA or market-access decision, contact EvySaif with the research question, the interventions and what you know of the evidence base, and we will return the feasibility picture, the appropriate method and the scope before any work is committed.
Last reviewed: September 2026. This article is general information for education; verify requirements and methods against current official sources for any specific project.