These are the two article for today’s journal club i am supposed to hive comment as an audience for each study separately. Tell me what are the setbacks of these studies individually other than those already mentioned in the article… and also tell how they could have avoided or improved those setbacks

Reading File
Reading File
Searching the Web

"Choudhary" "July September" 2025 original article medical journal

Browsing Directory

~/

Reading File
Reading File
Running Command

mkdir -p "/tmp/workspace/journal-club/src" "/tmp/workspace/journal-club/output" && python -c "from pypdf import PdfReader; from pathlib import Path; files=[Path('/tmp/workspace/attachments/561adb86-2b09-4612-9a32-720012ac57d9/7_Original_article__Choudhary_et_al_July_September_2025.pdf'),Path('/tmp/workspace/attachments/561adb86-2b09-4612-9a32-720012ac57d9/QVNTTWNpV3pSM2t6aU83TndRaWVRUT09.pdf')]; out=Path('/tmp/workspace/journal-club/output'); [((out/(f.stem+'.txt')).write_text('\\n\\n'.join((p.extract_text() or '') for p in PdfReader(str(f)).pages),errors='ignore'), print(f.name, len(PdfReader(str(f)).pages))) for f in files]"

Reading File
Reading File
Searching PubMed

clinical research methodology bias limitations

Searching Content

"Materials" in ~/journal-club/output

Searching Content

"Methods" in ~/journal-club/output

Searching Content

"Limitations" in ~/journal-club/output

Finding Sources
Reading File
Reading File
Searching the Web

IL-6 erythema nodosum leprosum biomarker serum study treatment confounding corticosteroids

Searching PubMed

pemphigus rituximab relapse biomarkers

Reading File
Reading File
Reading Web Page

https://pmc.ncbi.nlm.nih.gov/articles/PMC5346883

Reading Article

PMID: 39075718

Reading File
Reading File
Reading Article

PMID: 40826847

Searching PubMed

erythema nodosum leprosum biomarkers IL-6

Reading File
Reading File

1. Choudhary et al.

“A Pilot Study to Measure Serum IL-6 Levels in ENL Cases and Correlate These Levels with Severity of Reaction”

Limitations already acknowledged by the authors

They already note the single-biomarker approach, lack of longitudinal follow-up, need for larger multicentre samples, and absence of leprosy disease controls without reaction.

Additional limitations you can raise

Additional setbackWhy it mattersHow it could be improved
Inappropriate or insufficiently justified use of the t-test for IL-6IL-6 values were markedly skewed: mean 33.02 pg/mL, SD 37.77, median 11.7, range 3.6-136 pg/mL in cases. This distribution makes a parametric t-test questionable. A very high value in a few patients may drive the apparent case-control difference.Test normality, present individual data and use a Mann-Whitney U test or log-transform IL-6 before parametric modelling. Report effect sizes with 95% confidence intervals, not only p-values.
The severity cut-off was derived and tested in the same small sampleThe proposed IL-6 cut-off of 24.15 pg/mL, AUC 0.86, sensitivity 85.7%, and specificity 95.2% may be overly optimistic because it was selected after examining only 35 ENL cases, with relatively few moderate-severe cases. This is a classic risk of overfitting.Pre-specify the cut-off or derive it in one cohort and validate it prospectively in an independent external cohort. Use bootstrapping or cross-validation at minimum.
No adjustment for important clinical confoundersIL-6 can vary with bacillary burden, leprosy spectrum, duration and recurrence of ENL, time since onset of reaction, MDT phase, subclinical infection, nutritional status, and smoking. Matching only for age and sex does not address these factors. Therefore, IL-6 may be associated with a more inflammatory phenotype rather than independently reflecting ENL severity.Record these variables and use multivariable regression. Stratify by BL versus LL disease, first versus recurrent ENL, ENL duration, bacterial index, and MDT status.
Potential misclassification of the severity outcomeThe 16-item EESS includes some subjective components, including pain and wellbeing. Also, women were classified using a different maximum possible score because orchitis is not applicable. This may reduce comparability between male and female severity scores.Use sex-standardised severity scores, analyse the continuous score rather than only dichotomising severity, and ensure blinded or independently assessed EESS scoring.
Temporal ambiguity from the cross-sectional designA single IL-6 measurement cannot establish whether raised IL-6 precedes worsening ENL, results from it, or simply accompanies acute inflammation. Thus, calling IL-6 an “early detection” biomarker or therapeutic target goes beyond the design.Obtain serial samples at ENL onset, during treatment, remission, and recurrence. Assess whether an IL-6 rise predicts a subsequent clinical flare.
Biological specificity of IL-6 is limitedIL-6 is a general acute-phase cytokine, not an ENL-specific marker. Although the authors excluded overt infections and inflammatory conditions, a single IL-6 result may have limited real-world diagnostic usefulness. The literature itself is not fully consistent about serum IL-6 in ENL. A prior systematic review found mixed associations across ENL studies, as summarised in this ENL immunology review.Assess IL-6 as part of a clinically useful panel and report incremental value beyond clinical EESS assessment. Evaluate whether IL-6 changes a clinician’s decision or improves patient outcomes.
No information on assay reproducibility and pre-analytical handlingCytokine assays are sensitive to sample timing, storage, freeze-thaw cycles, laboratory batch effects, and assay detection limits. Without intra-assay/inter-assay variation or blinding of laboratory staff, reproducibility is difficult to judge.Report specimen processing protocol, assay manufacturer and detection range, duplicate testing, coefficients of variation, and whether laboratory personnel were blinded to case/severity status.

A concise audience comment

“This pilot study supports an association between IL-6 and ENL severity, but the ROC cut-off appears to have been developed internally in a very small sample with highly skewed IL-6 values. Could the authors clarify whether non-parametric or log-transformed analyses were performed, and whether IL-6 remains independently associated with severity after accounting for leprosy spectrum, bacterial load, ENL duration, and MDT status? A longitudinal external-validation cohort would be needed before using 24.15 pg/mL clinically.”

An especially strong discussion point

Do not challenge the finding that IL-6 is elevated. Challenge the clinical claim: the study shows association, not validation of a diagnostic, prognostic, or treatment-guiding biomarker.

2. Ahuja et al.

“Longitudinal trends in immunological biomarkers, and predictors of relapse in pemphigus treated with rituximab: A prospective cohort study”

Limitations already acknowledged by the authors

They acknowledge the single-centre, small and heterogeneous sample; mixed pemphigus vulgaris/pemphigus foliaceus and treatment-naive/previously treated cases; variability in steroid adjuvant treatment; and potentially low baseline biomarker levels in previously treated patients.

Additional limitations you can raise

Additional setbackWhy it mattersHow it could be improved
Post-enrolment attrition and competing-risk biasOf 52 enrolled patients, 3 died within the first month and 5 were lost to follow-up within 6 months. The final analysis used 44 patients. Excluding early deaths, including deaths related to sepsis and disease flare, may bias relapse estimates toward patients well enough to remain under observation.Report outcomes for the full enrolled cohort, describe baseline characteristics of excluded patients, and use time-to-event methods that treat death as a competing risk. Apply sensitivity analyses for loss to follow-up.
Risk of overfitting in biomarker cut-offsThe anti-DSG3 threshold of 120 RU/mL and the combined predictive rule were derived from the same small cohort, with only 18 early relapses. Multiple biomarkers, time points, percentage changes, and ROC analyses were assessed. This increases false-positive findings and inflates apparent predictive accuracy.Pre-specify the main predictor and cut-off, restrict the number of candidate predictors according to event count, adjust for multiple testing, and validate the model in an independent cohort.
The reported PPV is modestThe combined criterion had a PPV of 64.7%. In practical terms, about one in three patients labelled “high risk” would not relapse early. Giving maintenance rituximab to all such patients could expose some to unnecessary immunosuppression, infection risk, and cost.Present full 2x2 diagnostic tables with confidence intervals, decision-curve analysis, and clinical consequences. Test biomarker-guided maintenance rituximab against standard follow-up in a randomised trial before recommending it.
Informative censoring in the longitudinal biomarker analysisBlood sampling stopped at relapse. Therefore, patients who relapse early contribute fewer later measurements, whereas those who remain well contribute more. Comparisons at 9-12 months may represent a selected group of patients who had already remained relapse-free long enough to be sampled.Use mixed-effects longitudinal models jointly with survival analysis, or time-dependent Cox models. Clearly report the denominator available for every biomarker at each time point.
Potential “look-ahead” bias in grouping patients by eventual relapse timePatients were classified as early relapsers, late relapsers, or non-relapsers based on future outcomes, then their earlier biomarker trajectories were compared. This is useful descriptively but may overstate real-time predictive performance.Use landmark analyses at 3, 6, 9, and 12 months. At each landmark, include only patients still relapse-free and test whether the biomarker predicts future relapse.
Peripheral CD19 percentage may not be an adequate measure of B-cell biologyCD19+ cells were expressed as a percentage of total lymphocytes. Steroids, infection, and changes in total lymphocyte counts can alter this percentage without equivalent changes in absolute B-cell number. The authors themselves state that absolute B-cell counts could not be estimated. Also, peripheral blood B-cell measurements may not represent pathogenic tissue-resident or autoreactive B-cell populations.Measure absolute CD19+ counts alongside percentages, lymphocyte counts, and clinically relevant B-cell subsets such as transitional, naïve, memory, plasmablast, and regulatory B cells.
Relapse ascertainment may have observer biasRelapse requires judging new lesions and whether they persist without spontaneous healing. If clinicians knew biomarker results, even unconsciously, this could influence assessment or follow-up intensity.Predefine relapse adjudication, blind outcome assessors to biomarker results, and use standardised photographs and PDAI/POLIS scoring at suspected relapse.
Treatment recommendation exceeds the evidenceThe conclusion suggests that high-risk patients “may benefit” from maintenance rituximab at 6 months. But this was an observational prognostic study, not an intervention study. It cannot establish that biomarker-guided maintenance treatment prevents relapse or has an acceptable safety profile.Phrase this as a hypothesis. Conduct a prospective randomised trial comparing biomarker-guided maintenance rituximab with usual care, including infection, hypogammaglobulinaemia, cost, and quality-of-life outcomes.
Limited generalisability of the rituximab regimenThe cohort received a biosimilar using the rheumatoid arthritis protocol, in a tertiary Indian referral centre, mostly after prior systemic treatment. Findings may not translate to first-line rituximab, lymphoma-protocol dosing, other biosimilars, or different health-care settings.Validate across centres, ethnic/geographic settings, treatment lines, rituximab regimens, and originator versus biosimilar products.

A concise audience comment

“The prospective serial sampling is a major strength, but the proposed anti-DSG3 cut-off and combined prediction rule were both derived in only 44 analysable patients, with 18 early relapses and many repeated biomarker comparisons. Has the model been internally validated, and could the authors provide confidence intervals for sensitivity, specificity, PPV, and NPV? Before recommending maintenance rituximab, the strategy needs external validation and ideally a biomarker-guided randomised trial.”

Another high-yield comment

“Because biomarkers were no longer sampled after relapse, later time-point comparisons may be affected by informative censoring. A landmark or time-dependent survival analysis would better establish whether a 3- or 6-month biomarker result predicts subsequent relapse in patients who are still clinically in remission at that point.”

Bottom line for journal club

  • Choudhary et al.: Good pilot signal, but the main weaknesses are skewed data analysis, unadjusted confounding, internally derived ROC cut-off, and inability of a cross-sectional design to establish prediction.
  • Ahuja et al.: Better design and clinically relevant question, but the predictive model is vulnerable to small-event overfitting, attrition, repeated testing, and informative censoring. The proposed maintenance-rituximab action is not yet supported by causal evidence.
For context, a 2024 systematic review of pemphigus biomarkers found many candidate biomarkers but concluded that further validation is needed before clinical utility can be assumed (Geng et al., PMID 39075718).

More comments in first study

Reading File
Searching the Web

erythema nodosum leprosum IL-6 serum biomarker diagnostic studies limitations controls disease severity

Searching PubMed

erythema nodosum leprosum biomarkers

Finding Sources
Here are additional audience comments for the first study, phrased so you can use them directly during journal club.

Study: Serum IL-6 and ENL severity

1. No sample-size or power calculation

“As this is described as a pilot study, was a prior sample-size calculation performed? With only 35 cases, the estimate of correlation and especially the ROC-derived cut-off may be imprecise.”
Why it matters: A statistically significant p-value does not ensure a stable or clinically reliable estimate. The reported sensitivity, specificity, and AUC can change substantially in another small cohort.
Better approach: Define the intended precision for the correlation or AUC beforehand and recruit accordingly. Report 95% confidence intervals for correlation, AUC, sensitivity, specificity, and the 24.15 pg/mL cut-off.

2. Convenience sampling may introduce selection bias

“The study included all willing ENL patients presenting to one hospital and willing healthy volunteers as controls. Could this preferentially select more symptomatic or referral-level ENL cases and healthier-than-average controls?”
Why it matters: A tertiary dermatology centre can receive disproportionately severe, recurrent, or complicated ENL. The IL-6 difference may therefore be larger than in routine peripheral or community practice.
Better approach: Consecutive enrolment with a screening log of eligible, included, and excluded patients. Include multiple levels of care and recruit controls from the same underlying population.

3. The healthy-control comparison does not answer the key clinical question

“The comparison with healthy controls establishes that ENL is inflammatory, but in practice the diagnostic challenge is distinguishing ENL from non-reactional multibacillary leprosy, type 1 reaction, infection, or another inflammatory event. Why were these clinically relevant disease-control groups not included?”
Why it matters: IL-6 is a non-specific inflammatory cytokine. It may distinguish ENL from health but fail to distinguish ENL from other common inflammatory states.
A prior study that evaluated IL-6 included untreated multibacillary leprosy without ENL as a disease-control group, not only healthy controls, as described in the 2021 ENL IL-6 study.
Better approach: Include:
  • BL/LL leprosy without reaction, matched by bacterial index and treatment stage
  • Type 1 reaction
  • Intercurrent infection or inflammatory dermatoses when clinically relevant
  • Healthy controls only as a secondary reference group

4. No matching or adjustment for leprosy-specific disease burden

“Age, sex, BMI, and rural/urban background were reported, but did the analysis account for the underlying leprosy spectrum, bacterial index, disease duration, or current MDT phase?”
Why it matters: ENL occurs primarily in BL/LL disease. Bacillary burden and underlying disease phenotype can affect immune activation and cytokine levels. Without adjustment, IL-6 may partly reflect underlying multibacillary disease rather than ENL severity itself.
Better approach: Match or stratify cases by:
  • BL versus LL classification
  • Bacterial index and morphological index
  • Duration of leprosy and ENL
  • First episode versus recurrent/chronic ENL
  • Before, during, or after MDT
Then use multivariable regression to test whether IL-6 independently predicts EESS score.

5. ENL timing and disease phenotype are insufficiently characterised

“Was blood collected during the same phase of ENL for all patients, for example within a defined number of days from onset? Were acute, recurrent, and chronic ENL analysed separately?”
Why it matters: IL-6 may be highest early in an acute inflammatory episode and decline as the episode evolves. A patient sampled on day 1 and another sampled after several weeks of symptoms may have a different IL-6 value despite similar clinical scores.
Better approach: Record the exact onset-to-sampling interval and analyse acute, recurrent, and chronic ENL separately. Repeat samples at standardised intervals.

6. Severity is measured once and includes subjective items

“EESS includes pain and wellbeing visual analogue scales, which are patient-reported and potentially influenced by mood, analgesic use, and social factors. Was EESS assessment standardised or performed by assessors blinded to IL-6 results?”
Why it matters: Subjectivity in the reference severity measure can weaken or distort the correlation between IL-6 and severity.
Better approach: Use trained assessors with a standard operating protocol, inter-rater reliability assessment, and blinding of clinical assessors and laboratory personnel to each other’s results. Consider reporting objective severity components separately, such as fever, lesion count, neuritis, eye inflammation, orchitis, proteinuria, and nerve-function impairment.

7. Sex-specific severity classification is problematic

“Since orchitis is not applicable in women, males and females have different maximum possible EESS scores and different thresholds for mild versus moderate-severe ENL. Could this affect comparability of severity classification across sexes?”
Why it matters: A raw total score can be structurally lower in female participants because one component cannot occur. With only eight women, this may not be easily detected but is still a measurement issue.
Better approach: Use a percentage of the applicable maximum score, a sex-standardised score, or analyse disease severity by components that apply to both sexes.

8. The ROC analysis may be circular

“The ROC curve classifies moderate-severe versus mild ENL using a threshold derived from the same EESS score with which IL-6 was correlated. Does this add clinical information beyond the EESS itself?”
Why it matters: The analysis may be statistically impressive but clinically circular. The biomarker is being tested against a severity categorisation created from the same bedside clinical scale. It does not show that IL-6 improves management beyond careful EESS assessment.
Better approach: Compare a clinical model based on EESS alone with EESS plus IL-6. Report whether IL-6 improves discrimination, calibration, reclassification, or treatment decisions.

9. A single IL-6 result cannot support “early detection”

“The authors suggest earlier detection and enhanced management, but no pre-ENL samples were taken. Does the study demonstrate severity association rather than early prediction?”
Why it matters: This is an important distinction:
  • Associated biomarker: high during ENL
  • Severity biomarker: correlates with EESS at the same time
  • Predictive biomarker: rises before ENL or before worsening
  • Treatment-response biomarker: declines reliably with clinical recovery
This study only addresses the first two.
Better approach: Follow BL/LL patients prospectively from before reaction onset and assess whether IL-6 predicts ENL, worsening disease, recurrence, or treatment response.

10. One-month treatment exclusion does not remove all treatment-related confounding

“The study excluded patients who had received ENL treatment in the previous month, but were details of MDT exposure, previous corticosteroid or thalidomide exposure, and other anti-inflammatory medicines recorded?”
Why it matters: Immunomodulatory therapy may alter cytokine levels beyond one month in some patients. Also, MDT stage and prior reaction treatment can influence the inflammatory state.
Better approach: Collect detailed medication history, including dose, duration, timing of last steroid/thalidomide/NSAID exposure, MDT phase, adherence, and previous ENL episodes. Pre-specify a washout period appropriate for each therapy.

11. The maximum IL-6 value may be influencing the findings

“The case IL-6 range was wide, from 3.6 to 136 pg/mL, and the SD exceeded the mean. Could a few high values be driving the mean difference and the correlation?”
Why it matters: This suggests a skewed distribution and possible influential outliers. The study does report the median, which is helpful, but the analytical approach should reflect this distribution.
Better approach: Show a scatter plot or box-and-whisker plot with all individual values. Use median and interquartile range, non-parametric tests, robust regression, and sensitivity analysis excluding influential values.

12. Missing reporting of confidence intervals and effect size

“The paper reports p-values, but what are the confidence intervals for the IL-6 group difference, Spearman correlation of 0.61, AUC, sensitivity, and specificity?”
Why it matters: P-values indicate compatibility with a null hypothesis, whereas confidence intervals show the likely range and precision of the estimated clinical effect. This is particularly important in a small pilot study.
Better approach: Report:
  • Median difference or Hodges-Lehmann estimate with 95% CI
  • Spearman rho with 95% CI
  • AUC with 95% CI
  • Sensitivity and specificity with 95% CIs
  • Calibration measures for the proposed cut-off

13. No assessment of laboratory reliability

“Were samples tested in duplicate? What were the assay’s lower detection limit and intra-assay/inter-assay coefficients of variation? Were all samples analysed in one batch or across multiple batches?”
Why it matters: Cytokine assays are susceptible to pre-analytical and analytical variation. Without these details, it is difficult to know whether small differences in IL-6 are biological or technical.
Better approach: Standardise blood collection time, specimen processing, storage, and freeze-thaw cycles. Analyse in duplicate, report assay reliability, use blinded laboratory analysis, and include quality-control samples.

14. The clinical utility of the threshold is uncertain

“Even if 24.15 pg/mL separates mild from moderate-severe ENL in this cohort, how would this change treatment when EESS already provides clinical severity assessment?”
Why it matters: A biomarker should improve a real clinical decision, not just correlate with a clinical score. In a resource-limited setting, the extra cost and turnaround time need justification.
Better approach: Define a clinical use case, for example:
  • Identifying occult severe systemic ENL
  • Predicting deterioration before clinical signs
  • Monitoring treatment response
  • Selecting patients for more intensive follow-up
Then compare patient outcomes and costs with and without IL-6 testing.

Best 3 comments if you only have time to speak once

“My concern is that healthy controls are not the clinically relevant comparator. Since IL-6 is a non-specific inflammatory cytokine, the key comparison should be with non-reactional multibacillary leprosy and type 1 reaction, ideally matched for bacillary burden and MDT stage.”
“The IL-6 distribution appears highly skewed, with a standard deviation greater than the mean and a wide range. Were non-parametric analyses, log transformation, and outlier sensitivity analyses performed? The internally derived ROC cut-off also needs external validation.”
“This cross-sectional study shows association with concurrent ENL severity, but does not establish early prediction, treatment responsiveness, or a role as a therapeutic target. Serial sampling before, during, and after ENL would be needed for those claims.”
A useful evidence-based caution is that the ENL literature has reported both positive and null associations between serum IL-6 and ENL. The systematic immunology review highlights that heterogeneity in ENL definitions, control groups, treatments, and methods makes isolated biomarker findings difficult to generalise.

More in second study

Reading File
Searching the Web

pemphigus rituximab relapse anti DSG3 CD19 biomarkers predictors limitations prospective cohort

Searching PubMed

pemphigus rituximab relapse biomarkers

More audience comments for the second study

Ahuja et al. - Biomarker trends and relapse prediction after rituximab in pemphigus

1. Biomarkers were measured only until 12 months, but relapse was assessed for 24 months

“The biomarker measurements stop at 12 months, whereas patients were followed for relapse until 24 months. How can the study fully characterise biomarker trajectories before late relapse, especially in patients relapsing between 12 and 24 months?”
Why this matters: A late relapser may have had a rise in anti-DSG3 or B cells after month 12, but this could not be detected. The study can therefore support prediction of early relapse better than the biological interpretation of late relapse.
Improvement: Continue serial biomarker measurements through 24 months, or until relapse, with a predefined schedule.

2. “Incomplete B-cell depletion at 3 months” may actually represent early repopulation

“Since the first post-rituximab B-cell measurement was at 3 months, how can the authors distinguish true failure of initial B-cell depletion from successful depletion followed by early B-cell repopulation?”
Why this matters: This is an important biological distinction. A patient could have achieved complete depletion immediately after rituximab and then repopulated before the 3-month sample. Calling both situations “incomplete depletion” may misclassify the mechanism of relapse risk.
Improvement: Measure CD19+ cells at baseline, 2-4 weeks after infusion, then at 3, 6, 9, 12, 18, and 24 months. This would identify:
  • initial depletion failure
  • early repopulation
  • delayed repopulation

3. Anti-DSG3 may be confounded by clinical phenotype

“Could anti-DSG3 be predicting relapse partly because it is a marker of mucosal-predominant pemphigus vulgaris rather than an independent relapse biomarker?”
Why this matters: Anti-DSG3 is strongly linked to mucosal disease. The paper reports a trend toward more mucosal onset and higher baseline disease activity among early relapsers. Thus, the observed association between anti-DSG3 and relapse may be partly explained by underlying phenotype and baseline severity.
Improvement: Use multivariable modelling that adjusts for:
  • baseline PDAI and POLIS
  • mucosal versus cutaneous phenotype
  • anti-DSG3 baseline titre
  • pemphigus type
  • disease duration
  • prior rituximab exposure
  • steroid dose and conventional immunosuppressants
A stratified analysis among mucosal-predominant PV alone would also help.

4. Multiple testing increases the chance of false-positive findings

“The investigators tested several biomarkers at multiple time points, compared several relapse groups, examined percentage changes, and derived ROC cut-offs. Was any adjustment made for multiple comparisons?”
Why this matters: With many statistical tests, some p-values below 0.05 can occur by chance. This is particularly important in a cohort of 44 patients.
Improvement: Pre-specify one or two primary biomarkers and time points. For exploratory analyses, apply a false-discovery-rate correction or clearly label results as hypothesis-generating.

5. The prediction model lacks a full multivariable model

“The study identifies anti-DSG3 and incomplete B-cell depletion as predictors, but were they tested together with clinical predictors in a multivariable survival model?”
Why this matters: A biomarker should add predictive information beyond routine clinical assessment. If high baseline PDAI, mucosal disease, previous relapse, and steroid requirement predict relapse equally well, the additional value of anti-DSG3 needs to be demonstrated.
Improvement: Build a parsimonious Cox proportional-hazards or landmark prediction model including clinical and biomarker variables. Then report:
  • adjusted hazard ratios
  • calibration
  • discrimination, such as C-statistic
  • internal validation by bootstrapping
  • incremental predictive gain from biomarkers over clinical variables alone

6. Confidence intervals are needed for the proposed clinical rule

“The combined predictor has a PPV of 64.7% and NPV of 87.5%, but what are the 95% confidence intervals around these estimates?”
Why this matters: In a small cohort, PPV, NPV, sensitivity, and specificity can be unstable. PPV and NPV also change with the relapse prevalence in a different centre or population.
Improvement: Present a 2x2 table and 95% confidence intervals for all predictive parameters. Report likelihood ratios and calibration rather than relying only on PPV and NPV.

7. The proposed anti-DSG3 cut-off may not be transferable between laboratories

“The anti-DSG3 threshold of 120 RU/mL is assay-specific. Can it be applied to a different ELISA kit, laboratory, or calibration standard?”
Why this matters: “RU/mL” values can vary across commercial assays, batches, and laboratories. A cut-off derived in one laboratory may not work elsewhere.
Improvement: Report the exact kit, assay range, calibration procedures, inter-assay variation, and whether samples were analysed in duplicate. Validate the threshold across centres and, if possible, use standardised assay platforms.

8. Biomarker results may have been influenced by ongoing treatment changes

“Were steroid tapering, adjuvant immunosuppression, rescue treatment, or treatment adherence recorded longitudinally and incorporated into the analysis?”
Why this matters: Steroids and conventional immunosuppressants can alter disease activity, B-cell populations, and antibody titres. A patient with rising anti-DSG3 may receive treatment escalation before a formal relapse, obscuring the natural relationship between biomarker change and relapse.
Improvement: Record all treatment changes as time-varying covariates. Use standardised taper protocols, or adjust analyses for cumulative steroid exposure and concurrent immunosuppressive therapy.

9. Exclusion of early deaths may create survivor bias

“Three patients died within the first month, including patients with sepsis and paradoxical disease flare, and were excluded from the final analysis. Could this remove patients with the most severe disease biology from the biomarker analysis?”
Why this matters: The analysed cohort represents those who survived and remained available long enough for repeat testing. Relapse estimates and biomarker associations may not apply to the sickest patients.
Improvement: Present baseline biomarker and clinical characteristics of the excluded patients. Analyse the complete enrolled cohort where possible and treat death as a competing event rather than simply excluding it.

10. Loss to follow-up can bias relapse estimates

“Five patients were lost to follow-up within six months. Were their baseline disease activity and biomarkers different from those who completed follow-up?”
Why this matters: If patients lost to follow-up had severe disease, poor response, financial barriers, or treatment adverse effects, relapse risk may be underestimated.
Improvement: Compare completers and non-completers at baseline, document reasons for loss, and conduct sensitivity analyses assuming different outcomes for those lost.

11. Timing of relapse may be imprecise

“Patients were reviewed every three months after month 3. Could a relapse have occurred and partially resolved between scheduled visits, or could the recorded relapse date be later than the true onset?”
Why this matters: This affects correlation between biomarker timing and clinical relapse timing.
Improvement: Use patient diaries, teledermatology/photo reporting, interim telephone monitoring, and standardised instructions for immediate reporting of new lesions. Record symptom onset date separately from clinic-confirmed relapse date.

12. The relapse outcome could be affected by observer bias

“Were clinicians who determined relapse blinded to anti-DSG and B-cell results?”
Why this matters: Knowledge of a high biomarker level can unintentionally lower the threshold for labelling minor lesions as relapse or prompt more intensive examination.
Improvement: Use blinded outcome adjudication with standardised clinical photographs, PDAI/POLIS scoring, and predefined relapse criteria.

13. Peripheral blood B cells may not reflect pathogenic B cells in tissue

“Could peripheral CD19+ counts underestimate the relevant B-cell activity in lymphoid tissue, skin, or mucosa?”
Why this matters: Rituximab depletes circulating B cells effectively, but tissue B-cell depletion may differ. Also, total CD19+ B cells do not identify autoreactive B cells, plasmablasts, or long-lived plasma cells, which may continue producing pathogenic antibodies.
Improvement: Include absolute B-cell counts and deeper phenotyping, such as naïve, memory, transitional B cells, plasmablasts, and regulatory B cells. If feasible, correlate blood markers with lesional or mucosal tissue findings.

14. Percentage B-cell counts can be misleading without absolute counts

“CD19+ cells were expressed as a percentage of total lymphocytes. Could changes in lymphocyte count, steroid exposure, or intercurrent infection change the percentage without a true increase in B-cell number?”
Why this matters: A patient can appear to have B-cell repopulation by percentage while absolute B-cell numbers remain low, or vice versa.
Improvement: Report both:
  • CD19+ B cells as percentage of lymphocytes
  • absolute CD19+ B-cell count per microlitre

15. The clinical recommendation is stronger than the design permits

“The authors suggest that patients with high-risk biomarkers may benefit from maintenance rituximab at six months. However, can an observational prognostic study determine whether this intervention prevents relapse and is safe?”
Why this matters: A predictive biomarker does not automatically prove that biomarker-guided treatment improves outcome. Additional rituximab may reduce relapse but also increase infection, hypogammaglobulinaemia, vaccination problems, and cost.
Improvement: Phrase this as a hypothesis. Test it in a prospective biomarker-stratified randomised trial comparing:
  • usual follow-up and treatment
  • biomarker-guided maintenance rituximab
The trial should include relapse-free survival, serious infections, immunoglobulin levels, cumulative steroid exposure, quality of life, and cost.

16. Safety outcomes are not integrated into the proposed strategy

“If the biomarker rule is intended to trigger extra rituximab, were immunoglobulin levels, infections, hospitalisations, and vaccine responses assessed in relation to repeat B-cell depletion?”
Why this matters: The benefit-risk balance of maintenance rituximab cannot be judged using relapse alone.
Improvement: Prospectively monitor IgG, IgM, recurrent infections, serious infections, neutropenia, hepatitis B screening/reactivation, and need for immunoglobulin replacement.

High-impact comments to use in the discussion

Comment 1: strongest methodological point

“The first post-rituximab CD19 measurement was at three months. Therefore, ‘incomplete depletion’ may combine true failure of initial depletion with early repopulation. Measuring B cells at two to four weeks after infusion would clarify whether this is a pharmacodynamic failure or early recovery.”

Comment 2: strongest statistical point

“The anti-DSG3 cut-off of 120 RU/mL and the combined predictive rule were developed after examining several biomarkers at multiple time points in only 44 patients. This makes external validation and confidence intervals essential before applying the rule clinically.”

Comment 3: strongest clinical point

“The study supports closer monitoring of high-risk patients, but it cannot by itself establish that giving maintenance rituximab at six months prevents relapse safely. That needs a biomarker-guided interventional trial.”
A recent systematic review also concludes that pemphigus biomarkers need further validation before routine clinical implementation (Geng et al., 2024).
This is a shared conversation. Sign in to Orris to start your own chat.