Professor Irina Volkov, Director of Evidence-Based Medicine at a teaching hospital, faces a complex challenge for her upcoming presentation to the National Institute for Health and Care Excellence (NICE). She's been asked to provide expert testimony about treatment options for a chronic condition where multiple systematic reviews exist, but they reach conflicting conclusions. Her analysis reveals three Cochrane systematic reviews (published 2019-2023), two rapid reviews commissioned by different health technology assessment agencies, four network meta-analyses comparing different intervention combinations, and two umbrella reviews attempting to synthesise the systematic reviews. The reviews use different inclusion criteria, quality assessment methods, and statistical approaches. Some conclude there is strong evidence for intervention A, others favour intervention B, and several suggest insufficient evidence for any recommendation. The NICE committee needs Irina's expert interpretation to guide national treatment guidelines that will affect prescribing decisions across the NHS, potentially impacting 250,000 patients annually and £50 million in healthcare spending. 1. I'd like the group to develop a systematic approach for Irina to evaluate the quality and reliability of these conflicting reviews. Please discuss how tools like AMSTAR-2 should be applied, and bring evidence about which quality domains are most critical when reviews reach different conclusions about the same intervention. 2. Given the mix of review types Irina is dealing with, I want the group to critically analyse how different synthesis methodologies should be weighted in her evidence evaluation. Please discuss the relative strengths of Cochrane reviews versus network meta-analyses versus umbrella reviews and how methodological rigor should influence evidence interpretation. 3. Looking at the conflicting conclusions, I'd like the group to explore the potential sources of discordance between reviews. Please discuss how differences in inclusion criteria, search strategies, quality assessment, and statistical methods could lead to different conclusions and how Irina should present this uncertainty to the NICE committee. 4. Expert Synthesis and Recommendations Finally, I want the group to consider Irina's role as an expert synthesiser of conflicting evidence. Please discuss how she should move beyond simply describing the reviews to providing informed interpretation that guides policy decisions. What frameworks exist for expert evidence synthesis, and how should uncertainty be communicated to decision-makers? The following two articles could help you: Shea, B.J., Reeves, B.C., Wells, G., Thuku, M., Hamel, C., Moran, J., Ivory, J., Tugwell, P., Welch, V. and Kristjansson, E. (2017) 'AMSTAR 2: a critical appraisal tool for systematic reviews that include randomised or non-randomised studies of healthcare interventions', BMJ, 358, j4008. Available at: https://doi.org/10.1136/bmj.j4008 Salanti, G. (2012) 'Indirect and mixed-treatment comparison, network, or multiple-treatments meta-analysis: many names, many benefits, one statistical framework', Research Synthesis Methods, 3(2), pp. 80-97. Available at: https://doi.org/10.1002/jrsm.1037 Ahmed

Searching sources...
Let me pull the two key articles and search the evidence base in parallel.
Reading Web Page

https://doi.org/10.1136/bmj.j4008

Searching PubMed

AMSTAR-2 systematic review quality appraisal critical domains

Searching PubMed

conflicting systematic reviews discordance inclusion criteria heterogeneity

Searching PubMed

network meta-analysis indirect comparison evidence synthesis methodology

Now let me fetch the AMSTAR-2 paper directly and search for more targeted evidence on evidence synthesis frameworks and uncertainty communication.
Searching the Web

AMSTAR-2 16 domains critical items systematic review quality Shea 2017 BMJ

Searching PubMed

AMSTAR-2 appraisal systematic reviews confidence overall rating

Searching PubMed

umbrella review methodology evidence synthesis overlapping systematic reviews

Searching PubMed

GRADE evidence certainty health technology assessment guideline uncertainty

Let me get fuller detail on the AMSTAR-2 paper and the Salanti NMA framework, plus key abstracts.
Reading Web Page

https://corates.org/resources/amstar2

Reading Article

PMID: 36927334

Searching PubMed

Salanti network meta-analysis transitivity consistency assumption indirect comparisons

Searching the Web

sources of discordance between systematic reviews conflicting conclusions inclusion criteria search strategy heterogeneity

I now have sufficient material from AMSTAR-2 documentation, the Salanti NMA framework, PubMed searches, and Cochrane methodology guidance to write a thorough, well-cited response. Here is the comprehensive academic analysis.

Expert Evidence Synthesis for NICE: A Framework for Professor Irina Volkov


Question 1: Systematic Approach to Evaluating Quality and Reliability of Conflicting Reviews

Applying AMSTAR-2 as the Primary Quality Instrument

AMSTAR-2 (Shea et al., 2017, BMJ 358:j4008) is the instrument Irina should apply to each of the 11 reviews in her set. It is the direct successor to the original 11-item AMSTAR tool and was specifically developed to address the growing reality that healthcare systematic reviews now routinely include non-randomised studies alongside RCTs - making the original AMSTAR inadequate for comprehensive appraisal. The revised tool contains 16 items with simplified response categories (Yes / Partial Yes / No), and its most important architectural feature is the explicit critical/non-critical domain distinction.
AMSTAR-2 does not generate a numerical score. Representing the result as "13/16" misuses the tool. Instead, it produces an overall confidence rating - High, Moderate, Low, or Critically Low - derived from the pattern of weaknesses across the 16 items (Shea et al., 2017). The rating logic is:
Confidence LevelCriteria
HighZero or one non-critical weakness; accurate and comprehensive summary
ModerateMore than one non-critical weakness; no critical flaws
LowOne critical flaw with or without non-critical weaknesses
Critically LowMore than one critical flaw; should not be relied upon

The Seven Critical Domains - Why They Matter Most

The seven critical items carry disproportionate weight because weaknesses in these specific domains can "substantially undermine the validity of a review and its conclusions" (Shea et al., 2017). For Irina's purposes - understanding why three Cochrane reviews reach different conclusions - these are the domains that demand forensic scrutiny first:
  1. Item 2 - Protocol registration before commencement: Pre-registration constrains post-hoc outcome switching and prevents selective emphasis. A review conducted without a registered protocol cannot demonstrate that its conclusion was not shaped by its results. When comparing reviews that favour intervention A versus B, pre-registered PICO definitions are the most basic guard against this distortion.
  2. Item 4 - Comprehensive literature search: The search strategy must span multiple databases (MEDLINE, Embase, CENTRAL at minimum), trial registries (ClinicalTrials.gov, WHO ICTRP), and grey literature. A review restricted to two databases or omitting unpublished trials will systematically overestimate treatment effects due to publication bias - positive trials are published disproportionately. This single factor can flip conclusions from "insufficient evidence" to "strong evidence" for an intervention.
  3. Item 7 - Justification for excluded studies: A list of excluded studies with reasons allows independent verification that eligible trials were not quietly dropped. Where competing reviews of the same intervention include different trial sets, this item reveals whether discordance is driven by unjustified exclusions.
  4. Item 9 - Appropriate methods used to combine studies (risk of bias in individual studies): The review must use a validated risk of bias tool (Cochrane RoB 2 for RCTs, ROBINS-I for non-randomised studies). Reviews applying inadequate or absent risk-of-bias assessment may pool high-risk and low-risk trials indiscriminately, inflating apparent effect sizes.
  5. Item 11 - Use of appropriate meta-analytic methods: Statistical heterogeneity (clinical, methodological, and statistical) must be assessed and addressed. A review pooling clinically heterogeneous populations using a fixed-effects model when I² is high will produce a misleadingly precise but unreliable pooled estimate.
  6. Item 13 - Accounting for risk of bias when interpreting results: Even if individual study risk of bias is assessed, some reviews fail to act on it in their conclusions. A review whose conclusions are identical regardless of whether it includes only low-risk or all-risk studies has not completed this step.
  7. Item 15 - Assessment of publication bias: Funnel plot asymmetry tests (Egger's test for continuous outcomes, Begg's test) and contour-enhanced funnel plots should be applied when ten or more studies are included in a meta-analysis. This is the item most commonly neglected in practice, yet it is one of the seven critical items precisely because publication bias is one of the most prevalent threats to reliability in systematic reviews (De Santis et al., 2023, BMC Med Res Methodol, PMID 36927334).

Practical Application Protocol

Irina should apply AMSTAR-2 with two independent appraisers for each review, with disagreements resolved by consensus (De Santis et al., 2023). Before beginning, the team should agree explicit decision rules for ambiguous items - for instance, what constitutes a "major deviation" from protocol on Item 2. This pre-appraisal consensus prevents rater drift across the 11 reviews and is particularly important when the appraisal will directly inform national guidelines.
The appraisal should be tabulated comparatively, with each of the 11 reviews as rows, the 16 AMSTAR-2 items as columns, and the seven critical items highlighted. This comparative matrix immediately visualises whether reviews favouring intervention A consistently score Low/Critically Low while Cochrane reviews score High - which would be decisive evidence for differential weighting.
For the rapid reviews specifically, Irina should note that time-compressed review processes are structurally more vulnerable to compromises in search comprehensiveness (Item 4) and independent duplicate screening (Item 5). This is not automatically disqualifying, but must be explicitly acknowledged.

Question 2: Weighting Different Synthesis Methodologies

Cochrane Systematic Reviews

Cochrane reviews occupy the strongest methodological position in Irina's evidence set for several reasons. First, they follow a pre-published, peer-reviewed protocol registered on PROSPERO, directly satisfying AMSTAR-2 Item 2. Second, Cochrane's methodological standards mandate duplicate independent screening, extraction, and risk-of-bias assessment using Cochrane RoB 2. Third, they undergo editorial peer review by a Cochrane review group with topic expertise AND methods expertise - a dual peer review not required for journal-published reviews.
However, Cochrane reviews are not automatically superior. Their comprehensiveness comes with trade-offs: they can be slow to update, meaning a 2019 Cochrane review may exclude trials published after its search date that a 2023 rapid review would have captured. Irina must check the search dates of all three Cochrane reviews carefully - a review that searched in 2017 for a 2019 publication is already potentially outdated.
Cochrane reviews also apply strict eligibility criteria that can be both a strength and a limitation. If a Cochrane review excludes open-label trials or trials with active rather than placebo comparators, it may be methodologically cleaner but answer a narrower question than the one NICE needs to resolve for NHS prescribing.

Network Meta-Analyses (NMAs)

Salanti (2012, Research Synthesis Methods 3(2):80-97) provides the foundational framework for interpreting NMAs. The critical conceptual contribution of the NMA framework is enabling indirect comparisons: if intervention A and intervention C have each been compared to a common comparator B in RCTs, but never head-to-head against each other, NMA uses the geometry of the evidence network to estimate an indirect A-vs-C comparison. This is precisely why four NMAs exist in Irina's set - the direct trial evidence likely does not cover every pairwise comparison between interventions A, B, and their combinations.
Salanti (2012) identifies three core assumptions that must hold for NMA results to be valid:
1. Transitivity - the patient populations, intervention doses, outcome definitions, and treatment contexts must be sufficiently similar across the entire network that indirect comparisons are clinically meaningful. If one arm of the network used a much higher dose of intervention A than the other arms, or enrolled a systematically different patient population, transitivity is violated. Irina should examine each NMA's justification of transitivity - which is typically assessed by comparing distributions of potential effect modifiers across included trials.
2. Consistency - the direct and indirect evidence on each comparison must agree (within statistical uncertainty). Inconsistency tests (loop-specific models, node-splitting analyses) should have been conducted and reported. An NMA that finds significant inconsistency in a loop containing the A-vs-B comparison but proceeds to report pooled estimates without adjustment should be interpreted with considerable caution.
3. Common heterogeneity - NMA assumes that between-study variance (τ²) is exchangeable across all comparisons in the network. This assumption is rarely explicitly tested and is one of the more quietly problematic features of NMAs in practice.
For Irina's four NMAs, the critical questions are: (a) Are the network diagrams published so she can assess the geometry and identify thinly-connected comparisons? (b) Are consistency checks reported? (c) Are SUCRA (Surface Under the Cumulative Ranking) values presented in a way that implies more certainty than the underlying evidence warrants?
NMAs should be weighted heavily if their transitivity is justified, consistency is confirmed, and the underlying trial quality is high. They should be down-weighted or treated as hypothesis-generating if direct evidence is sparse, consistency is untested, or transitivity is assumed without justification.

Rapid Reviews

The two rapid reviews represent a methodological compromise - they exist because health technology assessment agencies needed evidence faster than a full systematic review could deliver. Their strengths are currency (more recent search dates) and pragmatic relevance (commissioned with a specific policy question in mind that may align more closely with the NICE question than a Cochrane review's population definition).
Their structural weaknesses are: single-reviewer screening and extraction (common in rapid reviews), restricted database searching (often two databases rather than five), and abbreviated quality assessment. Under AMSTAR-2, these compromises will typically yield Low or Critically Low ratings - not because the rapid review is wrong, but because it has insufficient safeguards to demonstrate it is right.
For Irina's presentation, the rapid reviews are best positioned as corroborating or contextualising evidence rather than primary evidence. If a rapid review's conclusion aligns with the highest-quality Cochrane review, it modestly reinforces it. If it contradicts it, the methodological vulnerabilities of the rapid review are the more likely explanation.

Umbrella Reviews

The two umbrella reviews represent an attempt to do exactly what Irina is doing - synthesise across systematic reviews. Their utility depends on how rigorously they handled the three central challenges of overview methodology (Cochrane Handbook, Chapter V):
  1. Overlapping primary studies: Systematic reviews addressing the same question commonly include many of the same trials. If both umbrella reviews pool results across their included reviews without accounting for this overlap, they are effectively double-counting trials - inflating precision without adding information.
  2. Heterogeneity of review quality: An umbrella review that weights a Critically Low AMSTAR-2 review equally with a High-quality Cochrane review will produce a misleading synthesis.
  3. Currency: An umbrella review published in 2022 that includes a 2019 Cochrane review as its most recent source will miss any subsequent evidence.
The umbrella reviews are structurally the highest-level synthesis in Irina's set. But if they reach conflicting conclusions, it may be because they included different subsets of the 11 lower-level reviews - making it essential that she map which primary reviews each umbrella review incorporated.
Evidence hierarchy for Irina's weighting purposes:
High-confidence Cochrane SRs with recent searches > High-quality NMAs with demonstrated consistency > Moderate-confidence Cochrane SRs > High-quality umbrella reviews > Moderate-quality NMAs > Rapid reviews (supportive role only) > Low/Critically Low rated reviews (flagged but largely discounted)

Question 3: Sources of Discordance and Communicating Uncertainty

Why Reviews Addressing the Same Question Reach Different Conclusions

This is the diagnostic core of Irina's task. The Jadad algorithm - one of the earliest systematic frameworks for adjudicating between conflicting reviews - identifies six primary sources of discordance: differences in the question asked, inclusion/exclusion criteria, data extracted, methodological quality assessments, methods of combining data, and statistical analysis approaches (Lunny et al., BMC Med Res Methodol 2022). Each deserves separate scrutiny.
1. Inclusion criteria differences - the most common driver
Inclusion criteria encode many implicit decisions that systematically shift results. Critical variables include:
  • Population - a review requiring a minimum disease duration of 6 months will include a more severe population than one requiring 3 months; if intervention A works better in established disease, this alone creates divergent conclusions
  • Intervention definition - dose, formulation, duration, and co-interventions vary substantially in trials of chronic conditions; a review that includes all doses of intervention A versus one restricted to licensed doses will pool a different set of trials
  • Comparator - placebo-controlled versus active-controlled designs are not directly comparable; a review including both will have more data but answer a different question than one restricted to placebo comparators
  • Outcome - the choice of primary outcome is particularly consequential; if review 1 uses a validated disease-specific questionnaire at 12 months as its primary endpoint and review 2 uses a generic quality-of-life measure at 6 months, they may legitimately reach different conclusions about the same intervention
2. Search strategy comprehensiveness
Search strategy differences explain a substantial proportion of review discordance (Cochrane Handbook, Chapter III). A review searching MEDLINE and PsycINFO only versus one searching MEDLINE, Embase, CENTRAL, CINAHL, clinicaltrials.gov, and WHO ICTRP will miss unpublished trials. Unpublished trials are disproportionately null or negative (publication bias). The review with the more comprehensive search will therefore show smaller, more conservative effect sizes - which may appear to conflict with the more restricted review's positive finding when in fact it is the more reliable estimate.
3. Quality assessment methodology and its downstream consequences
The three Cochrane reviews likely used Cochrane RoB 2; the rapid reviews may have used Newcastle-Ottawa Scale, SIGN checklists, or no validated instrument. More importantly, reviews differ not just in which tool they use but in how they act on the results. A review that identifies 40% of trials as high risk of bias but proceeds with a sensitivity analysis including and excluding high-risk trials will reach a different conclusion than one that pools all trials regardless, and different again from one that restricts its primary analysis to low-risk studies only. These downstream decisions about how risk-of-bias findings are incorporated into conclusions are rarely transparently reported.
4. Statistical methods and heterogeneity handling
Three statistical decisions are particularly consequential:
  • Fixed versus random effects: Fixed-effects models assume a single true effect size across all studies - appropriate only when the evidence base is genuinely homogeneous. Random-effects models acknowledge between-study variance. When heterogeneity is substantial (I² > 50-75%), random-effects models will produce wider, more honest confidence intervals that may cross the null, shifting a conclusion from "statistically significant benefit" to "insufficient evidence." Different model choices for the same data can produce different headline conclusions.
  • Handling of heterogeneity: Some reviews conduct subgroup analyses to explain heterogeneity (e.g., by dose, duration, population severity); others simply report an I² value and pool regardless. Reviews that explore heterogeneity sources often find that intervention A works in a specific subpopulation even when the overall pooled estimate is non-significant - leading to "insufficient evidence" in one review and "evidence in defined subgroup" in another.
  • Threshold for conducting meta-analysis: Some reviews synthesise narratively when fewer than three trials are available; others pool two trials. Small meta-analyses are highly susceptible to the influence of a single trial's result.

Communicating Uncertainty to NICE

The Cochrane overview methodology (Chapter V) recommends a structured approach to presenting discordant evidence that Irina should adapt:
Step 1: Map the overlap. Create a matrix showing which primary trials appear in which reviews. This immediately reveals whether discordant reviews are reaching different conclusions from the same trials (suggesting interpretive differences) or from different trial sets (suggesting inclusion criteria or search strategy differences).
Step 2: Apply GRADE to the body of evidence. GRADE (Grading of Recommendations Assessment, Development and Evaluation) provides the standard framework that NICE committees use to interpret evidence certainty (Very Low / Low / Moderate / High). For each intervention comparison, Irina should use GRADE to assess: risk of bias in the included studies, inconsistency between studies, indirectness of the evidence, imprecision of the estimates, and publication bias. Critically, where reviews disagree, the appropriate GRADE certainty will often be lower than any individual review suggests, because disagreement itself is evidence of uncertainty.
Step 3: Transparent uncertainty language. For the NICE committee, Irina should resist false precision. Phrases such as "the evidence suggests a possible benefit" or "findings are consistent with an effect but the evidence is insufficient to exclude a clinically important benefit of the comparator" communicate calibrated uncertainty more honestly than a binary "evidence supports/does not support." The committee can make a policy decision under uncertainty - that is their function - but they cannot do so well if the uncertainty is hidden behind spuriously confident language.
Step 4: Identify the specific question that would resolve the discordance. Rather than simply reporting that evidence is conflicting, Irina should identify what type of new evidence - a specific trial design, in a specific population, with a specific outcome measure - would resolve the key uncertainties. This converts an apparently unresolvable evidential conflict into a tractable research question and demonstrates the value of her expert synthesis to the committee.

Question 4: Expert Evidence Synthesis - Moving Beyond Description to Interpretation

The Role of the Expert Synthesiser

Irina's role before the NICE committee is not that of a librarian cataloguing reviews. It is that of a calibrated interpretive expert - someone who can weight evidence according to methodological rigour, contextualise statistical results within clinical realities, and communicate actionable conclusions despite remaining uncertainty. This is a distinct skill from conducting a systematic review, and it requires a different framework.

Available Frameworks for Expert Evidence Synthesis

1. The Consolidated Framework for Multiple Reviews (CFMR)
When multiple reviews address the same question but with varying quality and overlapping evidence, the expert synthesiser must first establish a review-level quality hierarchy (using AMSTAR-2), then work downward: conclusions from the highest-quality reviews take precedence; lower-quality reviews are examined to determine whether they offer any new evidence not captured by higher-quality reviews, or whether their discordant conclusions can be explained by methodological deficiencies already identified.
2. Evidence maps
An evidence map - a visual representation of the volume and quality of evidence across review types, interventions, and outcomes - allows the NICE committee to see at a glance where the evidence base is strong, sparse, or contested. For Irina's 11 reviews, a two-dimensional map with intervention (A, B, combinations) on one axis and review quality (AMSTAR-2 confidence rating) on the other axis, with each review represented as a bubble sized by the number of included trials, makes the evidence landscape immediately legible.
3. Structured narrative synthesis
Where meta-analytic pooling across reviews is inappropriate (due to overlap, heterogeneity, or incomparable outcomes), structured narrative synthesis with explicit direction-of-effect tables provides a transparent alternative. For each outcome of clinical importance to NICE, Irina should tabulate the direction of findings (favours A / favours B / no difference / inconclusive) from each review, weighted by its AMSTAR-2 rating. Convergent high-quality evidence in one direction is actionable; divergent evidence signals the need for GRADE certainty downgrading.
4. The GRADE EtD (Evidence to Decision) Framework
NICE formally uses the GRADE Evidence to Decision framework for guideline development. Irina's presentation should be structured around its domains: (a) desirable effects, (b) undesirable effects, (c) certainty of evidence, (d) values and preferences, (e) balance of effects, (f) resource implications, (g) equity considerations. This framework is the language the NICE committee already speaks - structuring her expert synthesis within it ensures her interpretation is directly usable rather than requiring translation.
5. The Jadad Algorithm
For adjudicating between specific pairs of conflicting reviews, the Jadad algorithm provides a stepwise decision procedure examining: Is the PICO question identical? If not, which is more applicable to the clinical decision? Are the searches comparable? Are the same primary studies included? Is the quality assessment comparable? Was the statistical synthesis appropriate? (Lunny et al., BMC Med Res Methodol 2022). This algorithm operationalises what an expert synthesiser does intuitively, but with explicit, auditable steps - important for testimony before a national body.

Communicating Uncertainty to Decision-Makers: Practical Principles

Acknowledge irreducible uncertainty explicitly. When the body of evidence is genuinely insufficient to favour intervention A or B, Irina must say so directly and resist institutional pressure to deliver false certainty. A premature recommendation based on conflicted evidence could harm 250,000 patients if it directs treatment toward the less effective or less safe intervention.
Distinguish types of uncertainty. There is a difference between: (a) aleatory uncertainty - inherent variability in biological response that no amount of further research will eliminate; (b) epistemic uncertainty - gaps in current knowledge that could be resolved by further research; and (c) model uncertainty - uncertainty about which synthesis methodology best represents the evidence. For the NICE committee, (b) and (c) are actionable through research commissioning; (a) is managed through individualised clinical decision-making.
Scenario analysis. Rather than presenting a single point estimate, Irina should present best-case, central-estimate, and worst-case scenarios based on different reasonable weighting of the available reviews. If all three scenarios point toward the same policy direction, the committee can act with confidence. If scenarios diverge dramatically, this itself is evidence that the evidence base is insufficient for a definitive recommendation.
Conditional recommendations. The NICE framework accommodates conditional recommendations - for instance, "Consider intervention A for patients with [specific characteristic] where the evidence from Cochrane review X (high confidence) is most applicable; for patients with [other characteristic], the evidence is insufficient to recommend either intervention over the other." This is more honest and more useful than either a blanket recommendation or a non-recommendation.
Framing the cost-effectiveness context. With £50 million in annual NHS spending at stake, Irina should note that evidence uncertainty directly increases uncertainty in health economic models. NICE cost-effectiveness thresholds (£20,000-30,000 per QALY) apply to expected values - but when the underpinning evidence is rated Low or Critically Low, probabilistic sensitivity analyses will produce wide credible intervals around the ICER, which the committee must factor into its decision.

Summary Matrix for Professor Volkov's NICE Presentation

Review TypeMethodological StrengthKey VulnerabilityAMSTAR-2 Priority ItemsWeight in Evidence Hierarchy
Cochrane SRs (2019-2023)Pre-registered; dual-review; Cochrane RoB 2May be outdated; narrow eligibility criteriaItems 2, 4, 7, 9, 11, 13, 15Primary
NMAs (n=4)Enables indirect comparisons; comparative rankingTransitivity and consistency may be unverifiedTransitivity justification; consistency tests; GRADE-NMAPrimary if consistency confirmed
Umbrella reviews (n=2)Highest-level synthesis; synthesises across reviewsOverlap not always addressed; quality weighting variableAMSTAR-2 of included reviews; overlap matrixContextualising
Rapid reviews (n=2)Currency; policy-relevant framingSingle reviewer; restricted search; abbreviated QAItems 4, 5, 9 (high vulnerability)Supplementary/corroborative

Key References

  • Shea, B.J. et al. (2017) 'AMSTAR 2: a critical appraisal tool for systematic reviews that include randomised or non-randomised studies of healthcare interventions', BMJ, 358, j4008. https://doi.org/10.1136/bmj.j4008
  • Salanti, G. (2012) 'Indirect and mixed-treatment comparison, network, or multiple-treatments meta-analysis: many names, many benefits, one statistical framework', Research Synthesis Methods, 3(2), pp. 80-97. https://doi.org/10.1002/jrsm.1037
  • De Santis, K.K. et al. (2023) 'User experience of applying AMSTAR 2 to appraise systematic reviews of healthcare interventions: a commentary', BMC Medical Research Methodology, [PMID 36927334]
  • Lunny, C. et al. (2022) 'How can clinicians choose between conflicting and discordant systematic reviews?', BMC Medical Research Methodology, 22:276
  • Cochrane Handbook for Systematic Reviews of Interventions, Chapter V: Overviews of Reviews (Cochrane.org, current edition)
  • Higgins, J.P.T. et al. (2019) Cochrane Handbook for Systematic Reviews of Interventions, Version 6. Cochrane. Available at: www.training.cochrane.org/handbookHere is a consolidated overview of what the analysis covers, Ahmed, so your group can see the architecture clearly before engaging with each section.

Overview for the Group

The response is structured around four interconnected analytical arguments:
Q1 - AMSTAR-2 Application: The analysis distinguishes AMSTAR-2's seven critical items from its nine non-critical ones, following Shea et al. (2017) directly. The most important domains when reviews reach conflicting conclusions are protocol pre-registration (Item 2), search comprehensiveness (Item 4), exclusion justification (Item 7), risk-of-bias assessment AND incorporation into conclusions (Items 9 and 13), appropriate statistical methods (Item 11), and publication bias assessment (Item 15). A comparative AMSTAR-2 matrix across all 11 reviews is proposed as the foundation of Irina's presentation. Critically, AMSTAR-2 produces a confidence rating - not a score - and treating it as a percentage (e.g., "13/16") is a specific misuse the tool's authors warn against (De Santis et al., 2023, PMID 36927334).
Q2 - Methodological weighting: Drawing on Salanti (2012), NMAs are positioned as high-value when three conditions are met: transitivity is justified (patient populations and contexts are comparable across the network), consistency is confirmed (direct and indirect estimates agree), and common heterogeneity is reasonable. Where these are unverified, NMAs should be treated as hypothesis-generating rather than conclusive. Cochrane reviews take precedence as primary evidence; rapid reviews serve a supplementary and corroborative role only.
Q3 - Sources of discordance: Six primary drivers are analysed: population/intervention/comparator/outcome definition differences, search comprehensiveness gaps (and their link to publication bias), choice of quality assessment instrument and - crucially - how risk-of-bias findings are acted upon, and statistical model choices (fixed vs. random effects, heterogeneity handling, the minimum-study threshold for pooling). The Jadad algorithm is introduced as a stepwise tool for adjudicating between specific conflicting review pairs.
Q4 - Expert synthesis frameworks: The GRADE Evidence to Decision framework is identified as the appropriate vehicle for structuring Irina's testimony since it is the language NICE already uses. Scenario analysis (best/central/worst case), evidence maps, and conditional recommendations are proposed as mechanisms for communicating calibrated uncertainty without sacrificing actionability. The analysis concludes with a direct connection to the NHS spending context: Low-certainty evidence directly widens cost-effectiveness credible intervals, which the NICE committee must factor into its threshold decision.
This is a shared conversation. Sign in to Orris to start your own chat.