I now have sufficient material from AMSTAR-2 documentation, the Salanti NMA framework, PubMed searches, and Cochrane methodology guidance to write a thorough, well-cited response. Here is the comprehensive academic analysis.
Expert Evidence Synthesis for NICE: A Framework for Professor Irina Volkov
Question 1: Systematic Approach to Evaluating Quality and Reliability of Conflicting Reviews
Applying AMSTAR-2 as the Primary Quality Instrument
AMSTAR-2 (Shea et al., 2017, BMJ 358:j4008) is the instrument Irina should apply to each of the 11 reviews in her set. It is the direct successor to the original 11-item AMSTAR tool and was specifically developed to address the growing reality that healthcare systematic reviews now routinely include non-randomised studies alongside RCTs - making the original AMSTAR inadequate for comprehensive appraisal. The revised tool contains 16 items with simplified response categories (Yes / Partial Yes / No), and its most important architectural feature is the explicit critical/non-critical domain distinction.
AMSTAR-2 does not generate a numerical score. Representing the result as "13/16" misuses the tool. Instead, it produces an overall confidence rating - High, Moderate, Low, or Critically Low - derived from the pattern of weaknesses across the 16 items (Shea et al., 2017). The rating logic is:
| Confidence Level | Criteria |
|---|
| High | Zero or one non-critical weakness; accurate and comprehensive summary |
| Moderate | More than one non-critical weakness; no critical flaws |
| Low | One critical flaw with or without non-critical weaknesses |
| Critically Low | More than one critical flaw; should not be relied upon |
The Seven Critical Domains - Why They Matter Most
The seven critical items carry disproportionate weight because weaknesses in these specific domains can "substantially undermine the validity of a review and its conclusions" (Shea et al., 2017). For Irina's purposes - understanding why three Cochrane reviews reach different conclusions - these are the domains that demand forensic scrutiny first:
-
Item 2 - Protocol registration before commencement: Pre-registration constrains post-hoc outcome switching and prevents selective emphasis. A review conducted without a registered protocol cannot demonstrate that its conclusion was not shaped by its results. When comparing reviews that favour intervention A versus B, pre-registered PICO definitions are the most basic guard against this distortion.
-
Item 4 - Comprehensive literature search: The search strategy must span multiple databases (MEDLINE, Embase, CENTRAL at minimum), trial registries (ClinicalTrials.gov, WHO ICTRP), and grey literature. A review restricted to two databases or omitting unpublished trials will systematically overestimate treatment effects due to publication bias - positive trials are published disproportionately. This single factor can flip conclusions from "insufficient evidence" to "strong evidence" for an intervention.
-
Item 7 - Justification for excluded studies: A list of excluded studies with reasons allows independent verification that eligible trials were not quietly dropped. Where competing reviews of the same intervention include different trial sets, this item reveals whether discordance is driven by unjustified exclusions.
-
Item 9 - Appropriate methods used to combine studies (risk of bias in individual studies): The review must use a validated risk of bias tool (Cochrane RoB 2 for RCTs, ROBINS-I for non-randomised studies). Reviews applying inadequate or absent risk-of-bias assessment may pool high-risk and low-risk trials indiscriminately, inflating apparent effect sizes.
-
Item 11 - Use of appropriate meta-analytic methods: Statistical heterogeneity (clinical, methodological, and statistical) must be assessed and addressed. A review pooling clinically heterogeneous populations using a fixed-effects model when I² is high will produce a misleadingly precise but unreliable pooled estimate.
-
Item 13 - Accounting for risk of bias when interpreting results: Even if individual study risk of bias is assessed, some reviews fail to act on it in their conclusions. A review whose conclusions are identical regardless of whether it includes only low-risk or all-risk studies has not completed this step.
-
Item 15 - Assessment of publication bias: Funnel plot asymmetry tests (Egger's test for continuous outcomes, Begg's test) and contour-enhanced funnel plots should be applied when ten or more studies are included in a meta-analysis. This is the item most commonly neglected in practice, yet it is one of the seven critical items precisely because publication bias is one of the most prevalent threats to reliability in systematic reviews (De Santis et al., 2023, BMC Med Res Methodol, PMID 36927334).
Practical Application Protocol
Irina should apply AMSTAR-2 with two independent appraisers for each review, with disagreements resolved by consensus (De Santis et al., 2023). Before beginning, the team should agree explicit decision rules for ambiguous items - for instance, what constitutes a "major deviation" from protocol on Item 2. This pre-appraisal consensus prevents rater drift across the 11 reviews and is particularly important when the appraisal will directly inform national guidelines.
The appraisal should be tabulated comparatively, with each of the 11 reviews as rows, the 16 AMSTAR-2 items as columns, and the seven critical items highlighted. This comparative matrix immediately visualises whether reviews favouring intervention A consistently score Low/Critically Low while Cochrane reviews score High - which would be decisive evidence for differential weighting.
For the rapid reviews specifically, Irina should note that time-compressed review processes are structurally more vulnerable to compromises in search comprehensiveness (Item 4) and independent duplicate screening (Item 5). This is not automatically disqualifying, but must be explicitly acknowledged.
Question 2: Weighting Different Synthesis Methodologies
Cochrane Systematic Reviews
Cochrane reviews occupy the strongest methodological position in Irina's evidence set for several reasons. First, they follow a pre-published, peer-reviewed protocol registered on PROSPERO, directly satisfying AMSTAR-2 Item 2. Second, Cochrane's methodological standards mandate duplicate independent screening, extraction, and risk-of-bias assessment using Cochrane RoB 2. Third, they undergo editorial peer review by a Cochrane review group with topic expertise AND methods expertise - a dual peer review not required for journal-published reviews.
However, Cochrane reviews are not automatically superior. Their comprehensiveness comes with trade-offs: they can be slow to update, meaning a 2019 Cochrane review may exclude trials published after its search date that a 2023 rapid review would have captured. Irina must check the search dates of all three Cochrane reviews carefully - a review that searched in 2017 for a 2019 publication is already potentially outdated.
Cochrane reviews also apply strict eligibility criteria that can be both a strength and a limitation. If a Cochrane review excludes open-label trials or trials with active rather than placebo comparators, it may be methodologically cleaner but answer a narrower question than the one NICE needs to resolve for NHS prescribing.
Network Meta-Analyses (NMAs)
Salanti (2012, Research Synthesis Methods 3(2):80-97) provides the foundational framework for interpreting NMAs. The critical conceptual contribution of the NMA framework is enabling indirect comparisons: if intervention A and intervention C have each been compared to a common comparator B in RCTs, but never head-to-head against each other, NMA uses the geometry of the evidence network to estimate an indirect A-vs-C comparison. This is precisely why four NMAs exist in Irina's set - the direct trial evidence likely does not cover every pairwise comparison between interventions A, B, and their combinations.
Salanti (2012) identifies three core assumptions that must hold for NMA results to be valid:
1. Transitivity - the patient populations, intervention doses, outcome definitions, and treatment contexts must be sufficiently similar across the entire network that indirect comparisons are clinically meaningful. If one arm of the network used a much higher dose of intervention A than the other arms, or enrolled a systematically different patient population, transitivity is violated. Irina should examine each NMA's justification of transitivity - which is typically assessed by comparing distributions of potential effect modifiers across included trials.
2. Consistency - the direct and indirect evidence on each comparison must agree (within statistical uncertainty). Inconsistency tests (loop-specific models, node-splitting analyses) should have been conducted and reported. An NMA that finds significant inconsistency in a loop containing the A-vs-B comparison but proceeds to report pooled estimates without adjustment should be interpreted with considerable caution.
3. Common heterogeneity - NMA assumes that between-study variance (τ²) is exchangeable across all comparisons in the network. This assumption is rarely explicitly tested and is one of the more quietly problematic features of NMAs in practice.
For Irina's four NMAs, the critical questions are: (a) Are the network diagrams published so she can assess the geometry and identify thinly-connected comparisons? (b) Are consistency checks reported? (c) Are SUCRA (Surface Under the Cumulative Ranking) values presented in a way that implies more certainty than the underlying evidence warrants?
NMAs should be weighted heavily if their transitivity is justified, consistency is confirmed, and the underlying trial quality is high. They should be down-weighted or treated as hypothesis-generating if direct evidence is sparse, consistency is untested, or transitivity is assumed without justification.
Rapid Reviews
The two rapid reviews represent a methodological compromise - they exist because health technology assessment agencies needed evidence faster than a full systematic review could deliver. Their strengths are currency (more recent search dates) and pragmatic relevance (commissioned with a specific policy question in mind that may align more closely with the NICE question than a Cochrane review's population definition).
Their structural weaknesses are: single-reviewer screening and extraction (common in rapid reviews), restricted database searching (often two databases rather than five), and abbreviated quality assessment. Under AMSTAR-2, these compromises will typically yield Low or Critically Low ratings - not because the rapid review is wrong, but because it has insufficient safeguards to demonstrate it is right.
For Irina's presentation, the rapid reviews are best positioned as corroborating or contextualising evidence rather than primary evidence. If a rapid review's conclusion aligns with the highest-quality Cochrane review, it modestly reinforces it. If it contradicts it, the methodological vulnerabilities of the rapid review are the more likely explanation.
Umbrella Reviews
The two umbrella reviews represent an attempt to do exactly what Irina is doing - synthesise across systematic reviews. Their utility depends on how rigorously they handled the three central challenges of overview methodology (Cochrane Handbook, Chapter V):
-
Overlapping primary studies: Systematic reviews addressing the same question commonly include many of the same trials. If both umbrella reviews pool results across their included reviews without accounting for this overlap, they are effectively double-counting trials - inflating precision without adding information.
-
Heterogeneity of review quality: An umbrella review that weights a Critically Low AMSTAR-2 review equally with a High-quality Cochrane review will produce a misleading synthesis.
-
Currency: An umbrella review published in 2022 that includes a 2019 Cochrane review as its most recent source will miss any subsequent evidence.
The umbrella reviews are structurally the highest-level synthesis in Irina's set. But if they reach conflicting conclusions, it may be because they included different subsets of the 11 lower-level reviews - making it essential that she map which primary reviews each umbrella review incorporated.
Evidence hierarchy for Irina's weighting purposes:
High-confidence Cochrane SRs with recent searches > High-quality NMAs with demonstrated consistency > Moderate-confidence Cochrane SRs > High-quality umbrella reviews > Moderate-quality NMAs > Rapid reviews (supportive role only) > Low/Critically Low rated reviews (flagged but largely discounted)
Question 3: Sources of Discordance and Communicating Uncertainty
Why Reviews Addressing the Same Question Reach Different Conclusions
This is the diagnostic core of Irina's task. The Jadad algorithm - one of the earliest systematic frameworks for adjudicating between conflicting reviews - identifies six primary sources of discordance: differences in the question asked, inclusion/exclusion criteria, data extracted, methodological quality assessments, methods of combining data, and statistical analysis approaches (Lunny et al., BMC Med Res Methodol 2022). Each deserves separate scrutiny.
1. Inclusion criteria differences - the most common driver
Inclusion criteria encode many implicit decisions that systematically shift results. Critical variables include:
- Population - a review requiring a minimum disease duration of 6 months will include a more severe population than one requiring 3 months; if intervention A works better in established disease, this alone creates divergent conclusions
- Intervention definition - dose, formulation, duration, and co-interventions vary substantially in trials of chronic conditions; a review that includes all doses of intervention A versus one restricted to licensed doses will pool a different set of trials
- Comparator - placebo-controlled versus active-controlled designs are not directly comparable; a review including both will have more data but answer a different question than one restricted to placebo comparators
- Outcome - the choice of primary outcome is particularly consequential; if review 1 uses a validated disease-specific questionnaire at 12 months as its primary endpoint and review 2 uses a generic quality-of-life measure at 6 months, they may legitimately reach different conclusions about the same intervention
2. Search strategy comprehensiveness
Search strategy differences explain a substantial proportion of review discordance (Cochrane Handbook, Chapter III). A review searching MEDLINE and PsycINFO only versus one searching MEDLINE, Embase, CENTRAL, CINAHL, clinicaltrials.gov, and WHO ICTRP will miss unpublished trials. Unpublished trials are disproportionately null or negative (publication bias). The review with the more comprehensive search will therefore show smaller, more conservative effect sizes - which may appear to conflict with the more restricted review's positive finding when in fact it is the more reliable estimate.
3. Quality assessment methodology and its downstream consequences
The three Cochrane reviews likely used Cochrane RoB 2; the rapid reviews may have used Newcastle-Ottawa Scale, SIGN checklists, or no validated instrument. More importantly, reviews differ not just in which tool they use but in how they act on the results. A review that identifies 40% of trials as high risk of bias but proceeds with a sensitivity analysis including and excluding high-risk trials will reach a different conclusion than one that pools all trials regardless, and different again from one that restricts its primary analysis to low-risk studies only. These downstream decisions about how risk-of-bias findings are incorporated into conclusions are rarely transparently reported.
4. Statistical methods and heterogeneity handling
Three statistical decisions are particularly consequential:
-
Fixed versus random effects: Fixed-effects models assume a single true effect size across all studies - appropriate only when the evidence base is genuinely homogeneous. Random-effects models acknowledge between-study variance. When heterogeneity is substantial (I² > 50-75%), random-effects models will produce wider, more honest confidence intervals that may cross the null, shifting a conclusion from "statistically significant benefit" to "insufficient evidence." Different model choices for the same data can produce different headline conclusions.
-
Handling of heterogeneity: Some reviews conduct subgroup analyses to explain heterogeneity (e.g., by dose, duration, population severity); others simply report an I² value and pool regardless. Reviews that explore heterogeneity sources often find that intervention A works in a specific subpopulation even when the overall pooled estimate is non-significant - leading to "insufficient evidence" in one review and "evidence in defined subgroup" in another.
-
Threshold for conducting meta-analysis: Some reviews synthesise narratively when fewer than three trials are available; others pool two trials. Small meta-analyses are highly susceptible to the influence of a single trial's result.
Communicating Uncertainty to NICE
The Cochrane overview methodology (Chapter V) recommends a structured approach to presenting discordant evidence that Irina should adapt:
Step 1: Map the overlap. Create a matrix showing which primary trials appear in which reviews. This immediately reveals whether discordant reviews are reaching different conclusions from the same trials (suggesting interpretive differences) or from different trial sets (suggesting inclusion criteria or search strategy differences).
Step 2: Apply GRADE to the body of evidence. GRADE (Grading of Recommendations Assessment, Development and Evaluation) provides the standard framework that NICE committees use to interpret evidence certainty (Very Low / Low / Moderate / High). For each intervention comparison, Irina should use GRADE to assess: risk of bias in the included studies, inconsistency between studies, indirectness of the evidence, imprecision of the estimates, and publication bias. Critically, where reviews disagree, the appropriate GRADE certainty will often be lower than any individual review suggests, because disagreement itself is evidence of uncertainty.
Step 3: Transparent uncertainty language. For the NICE committee, Irina should resist false precision. Phrases such as "the evidence suggests a possible benefit" or "findings are consistent with an effect but the evidence is insufficient to exclude a clinically important benefit of the comparator" communicate calibrated uncertainty more honestly than a binary "evidence supports/does not support." The committee can make a policy decision under uncertainty - that is their function - but they cannot do so well if the uncertainty is hidden behind spuriously confident language.
Step 4: Identify the specific question that would resolve the discordance. Rather than simply reporting that evidence is conflicting, Irina should identify what type of new evidence - a specific trial design, in a specific population, with a specific outcome measure - would resolve the key uncertainties. This converts an apparently unresolvable evidential conflict into a tractable research question and demonstrates the value of her expert synthesis to the committee.
Question 4: Expert Evidence Synthesis - Moving Beyond Description to Interpretation
The Role of the Expert Synthesiser
Irina's role before the NICE committee is not that of a librarian cataloguing reviews. It is that of a calibrated interpretive expert - someone who can weight evidence according to methodological rigour, contextualise statistical results within clinical realities, and communicate actionable conclusions despite remaining uncertainty. This is a distinct skill from conducting a systematic review, and it requires a different framework.
Available Frameworks for Expert Evidence Synthesis
1. The Consolidated Framework for Multiple Reviews (CFMR)
When multiple reviews address the same question but with varying quality and overlapping evidence, the expert synthesiser must first establish a review-level quality hierarchy (using AMSTAR-2), then work downward: conclusions from the highest-quality reviews take precedence; lower-quality reviews are examined to determine whether they offer any new evidence not captured by higher-quality reviews, or whether their discordant conclusions can be explained by methodological deficiencies already identified.
2. Evidence maps
An evidence map - a visual representation of the volume and quality of evidence across review types, interventions, and outcomes - allows the NICE committee to see at a glance where the evidence base is strong, sparse, or contested. For Irina's 11 reviews, a two-dimensional map with intervention (A, B, combinations) on one axis and review quality (AMSTAR-2 confidence rating) on the other axis, with each review represented as a bubble sized by the number of included trials, makes the evidence landscape immediately legible.
3. Structured narrative synthesis
Where meta-analytic pooling across reviews is inappropriate (due to overlap, heterogeneity, or incomparable outcomes), structured narrative synthesis with explicit direction-of-effect tables provides a transparent alternative. For each outcome of clinical importance to NICE, Irina should tabulate the direction of findings (favours A / favours B / no difference / inconclusive) from each review, weighted by its AMSTAR-2 rating. Convergent high-quality evidence in one direction is actionable; divergent evidence signals the need for GRADE certainty downgrading.
4. The GRADE EtD (Evidence to Decision) Framework
NICE formally uses the GRADE Evidence to Decision framework for guideline development. Irina's presentation should be structured around its domains: (a) desirable effects, (b) undesirable effects, (c) certainty of evidence, (d) values and preferences, (e) balance of effects, (f) resource implications, (g) equity considerations. This framework is the language the NICE committee already speaks - structuring her expert synthesis within it ensures her interpretation is directly usable rather than requiring translation.
5. The Jadad Algorithm
For adjudicating between specific pairs of conflicting reviews, the Jadad algorithm provides a stepwise decision procedure examining: Is the PICO question identical? If not, which is more applicable to the clinical decision? Are the searches comparable? Are the same primary studies included? Is the quality assessment comparable? Was the statistical synthesis appropriate? (Lunny et al., BMC Med Res Methodol 2022). This algorithm operationalises what an expert synthesiser does intuitively, but with explicit, auditable steps - important for testimony before a national body.
Communicating Uncertainty to Decision-Makers: Practical Principles
Acknowledge irreducible uncertainty explicitly. When the body of evidence is genuinely insufficient to favour intervention A or B, Irina must say so directly and resist institutional pressure to deliver false certainty. A premature recommendation based on conflicted evidence could harm 250,000 patients if it directs treatment toward the less effective or less safe intervention.
Distinguish types of uncertainty. There is a difference between: (a) aleatory uncertainty - inherent variability in biological response that no amount of further research will eliminate; (b) epistemic uncertainty - gaps in current knowledge that could be resolved by further research; and (c) model uncertainty - uncertainty about which synthesis methodology best represents the evidence. For the NICE committee, (b) and (c) are actionable through research commissioning; (a) is managed through individualised clinical decision-making.
Scenario analysis. Rather than presenting a single point estimate, Irina should present best-case, central-estimate, and worst-case scenarios based on different reasonable weighting of the available reviews. If all three scenarios point toward the same policy direction, the committee can act with confidence. If scenarios diverge dramatically, this itself is evidence that the evidence base is insufficient for a definitive recommendation.
Conditional recommendations. The NICE framework accommodates conditional recommendations - for instance, "Consider intervention A for patients with [specific characteristic] where the evidence from Cochrane review X (high confidence) is most applicable; for patients with [other characteristic], the evidence is insufficient to recommend either intervention over the other." This is more honest and more useful than either a blanket recommendation or a non-recommendation.
Framing the cost-effectiveness context. With £50 million in annual NHS spending at stake, Irina should note that evidence uncertainty directly increases uncertainty in health economic models. NICE cost-effectiveness thresholds (£20,000-30,000 per QALY) apply to expected values - but when the underpinning evidence is rated Low or Critically Low, probabilistic sensitivity analyses will produce wide credible intervals around the ICER, which the committee must factor into its decision.
Summary Matrix for Professor Volkov's NICE Presentation
| Review Type | Methodological Strength | Key Vulnerability | AMSTAR-2 Priority Items | Weight in Evidence Hierarchy |
|---|
| Cochrane SRs (2019-2023) | Pre-registered; dual-review; Cochrane RoB 2 | May be outdated; narrow eligibility criteria | Items 2, 4, 7, 9, 11, 13, 15 | Primary |
| NMAs (n=4) | Enables indirect comparisons; comparative ranking | Transitivity and consistency may be unverified | Transitivity justification; consistency tests; GRADE-NMA | Primary if consistency confirmed |
| Umbrella reviews (n=2) | Highest-level synthesis; synthesises across reviews | Overlap not always addressed; quality weighting variable | AMSTAR-2 of included reviews; overlap matrix | Contextualising |
| Rapid reviews (n=2) | Currency; policy-relevant framing | Single reviewer; restricted search; abbreviated QA | Items 4, 5, 9 (high vulnerability) | Supplementary/corroborative |
Key References
- Shea, B.J. et al. (2017) 'AMSTAR 2: a critical appraisal tool for systematic reviews that include randomised or non-randomised studies of healthcare interventions', BMJ, 358, j4008. https://doi.org/10.1136/bmj.j4008
- Salanti, G. (2012) 'Indirect and mixed-treatment comparison, network, or multiple-treatments meta-analysis: many names, many benefits, one statistical framework', Research Synthesis Methods, 3(2), pp. 80-97. https://doi.org/10.1002/jrsm.1037
- De Santis, K.K. et al. (2023) 'User experience of applying AMSTAR 2 to appraise systematic reviews of healthcare interventions: a commentary', BMC Medical Research Methodology, [PMID 36927334]
- Lunny, C. et al. (2022) 'How can clinicians choose between conflicting and discordant systematic reviews?', BMC Medical Research Methodology, 22:276
- Cochrane Handbook for Systematic Reviews of Interventions, Chapter V: Overviews of Reviews (Cochrane.org, current edition)
- Higgins, J.P.T. et al. (2019) Cochrane Handbook for Systematic Reviews of Interventions, Version 6. Cochrane. Available at: www.training.cochrane.org/handbookHere is a consolidated overview of what the analysis covers, Ahmed, so your group can see the architecture clearly before engaging with each section.
Overview for the Group
The response is structured around four interconnected analytical arguments:
Q1 - AMSTAR-2 Application: The analysis distinguishes AMSTAR-2's seven critical items from its nine non-critical ones, following Shea et al. (2017) directly. The most important domains when reviews reach conflicting conclusions are protocol pre-registration (Item 2), search comprehensiveness (Item 4), exclusion justification (Item 7), risk-of-bias assessment AND incorporation into conclusions (Items 9 and 13), appropriate statistical methods (Item 11), and publication bias assessment (Item 15). A comparative AMSTAR-2 matrix across all 11 reviews is proposed as the foundation of Irina's presentation. Critically, AMSTAR-2 produces a confidence rating - not a score - and treating it as a percentage (e.g., "13/16") is a specific misuse the tool's authors warn against (De Santis et al., 2023, PMID 36927334).
Q2 - Methodological weighting: Drawing on Salanti (2012), NMAs are positioned as high-value when three conditions are met: transitivity is justified (patient populations and contexts are comparable across the network), consistency is confirmed (direct and indirect estimates agree), and common heterogeneity is reasonable. Where these are unverified, NMAs should be treated as hypothesis-generating rather than conclusive. Cochrane reviews take precedence as primary evidence; rapid reviews serve a supplementary and corroborative role only.
Q3 - Sources of discordance: Six primary drivers are analysed: population/intervention/comparator/outcome definition differences, search comprehensiveness gaps (and their link to publication bias), choice of quality assessment instrument and - crucially - how risk-of-bias findings are acted upon, and statistical model choices (fixed vs. random effects, heterogeneity handling, the minimum-study threshold for pooling). The Jadad algorithm is introduced as a stepwise tool for adjudicating between specific conflicting review pairs.
Q4 - Expert synthesis frameworks: The GRADE Evidence to Decision framework is identified as the appropriate vehicle for structuring Irina's testimony since it is the language NICE already uses. Scenario analysis (best/central/worst case), evidence maps, and conditional recommendations are proposed as mechanisms for communicating calibrated uncertainty without sacrificing actionability. The analysis concludes with a direct connection to the NHS spending context: Low-certainty evidence directly widens cost-effectiveness credible intervals, which the NICE committee must factor into its threshold decision.