These are the two article for today’s journal club i am supposed to hive comment as an audience for each study separately. Tell me what are the setbacks of these studies individually other than those already mentioned in the article… and also tell how they could have avoided or improved those setbacks
"Choudhary" "July September" 2025 original article medical journal
~/
mkdir -p "/tmp/workspace/journal-club/src" "/tmp/workspace/journal-club/output" && python -c "from pypdf import PdfReader; from pathlib import Path; files=[Path('/tmp/workspace/attachments/561adb86-2b09-4612-9a32-720012ac57d9/7_Original_article__Choudhary_et_al_July_September_2025.pdf'),Path('/tmp/workspace/attachments/561adb86-2b09-4612-9a32-720012ac57d9/QVNTTWNpV3pSM2t6aU83TndRaWVRUT09.pdf')]; out=Path('/tmp/workspace/journal-club/output'); [((out/(f.stem+'.txt')).write_text('\\n\\n'.join((p.extract_text() or '') for p in PdfReader(str(f)).pages),errors='ignore'), print(f.name, len(PdfReader(str(f)).pages))) for f in files]"
clinical research methodology bias limitations
"Materials" in ~/journal-club/output
"Methods" in ~/journal-club/output
"Limitations" in ~/journal-club/output
IL-6 erythema nodosum leprosum biomarker serum study treatment confounding corticosteroids
pemphigus rituximab relapse biomarkers
https://pmc.ncbi.nlm.nih.gov/articles/PMC5346883
PMID: 39075718
PMID: 40826847
erythema nodosum leprosum biomarkers IL-6
| Additional setback | Why it matters | How it could be improved |
|---|---|---|
| Inappropriate or insufficiently justified use of the t-test for IL-6 | IL-6 values were markedly skewed: mean 33.02 pg/mL, SD 37.77, median 11.7, range 3.6-136 pg/mL in cases. This distribution makes a parametric t-test questionable. A very high value in a few patients may drive the apparent case-control difference. | Test normality, present individual data and use a Mann-Whitney U test or log-transform IL-6 before parametric modelling. Report effect sizes with 95% confidence intervals, not only p-values. |
| The severity cut-off was derived and tested in the same small sample | The proposed IL-6 cut-off of 24.15 pg/mL, AUC 0.86, sensitivity 85.7%, and specificity 95.2% may be overly optimistic because it was selected after examining only 35 ENL cases, with relatively few moderate-severe cases. This is a classic risk of overfitting. | Pre-specify the cut-off or derive it in one cohort and validate it prospectively in an independent external cohort. Use bootstrapping or cross-validation at minimum. |
| No adjustment for important clinical confounders | IL-6 can vary with bacillary burden, leprosy spectrum, duration and recurrence of ENL, time since onset of reaction, MDT phase, subclinical infection, nutritional status, and smoking. Matching only for age and sex does not address these factors. Therefore, IL-6 may be associated with a more inflammatory phenotype rather than independently reflecting ENL severity. | Record these variables and use multivariable regression. Stratify by BL versus LL disease, first versus recurrent ENL, ENL duration, bacterial index, and MDT status. |
| Potential misclassification of the severity outcome | The 16-item EESS includes some subjective components, including pain and wellbeing. Also, women were classified using a different maximum possible score because orchitis is not applicable. This may reduce comparability between male and female severity scores. | Use sex-standardised severity scores, analyse the continuous score rather than only dichotomising severity, and ensure blinded or independently assessed EESS scoring. |
| Temporal ambiguity from the cross-sectional design | A single IL-6 measurement cannot establish whether raised IL-6 precedes worsening ENL, results from it, or simply accompanies acute inflammation. Thus, calling IL-6 an “early detection” biomarker or therapeutic target goes beyond the design. | Obtain serial samples at ENL onset, during treatment, remission, and recurrence. Assess whether an IL-6 rise predicts a subsequent clinical flare. |
| Biological specificity of IL-6 is limited | IL-6 is a general acute-phase cytokine, not an ENL-specific marker. Although the authors excluded overt infections and inflammatory conditions, a single IL-6 result may have limited real-world diagnostic usefulness. The literature itself is not fully consistent about serum IL-6 in ENL. A prior systematic review found mixed associations across ENL studies, as summarised in this ENL immunology review. | Assess IL-6 as part of a clinically useful panel and report incremental value beyond clinical EESS assessment. Evaluate whether IL-6 changes a clinician’s decision or improves patient outcomes. |
| No information on assay reproducibility and pre-analytical handling | Cytokine assays are sensitive to sample timing, storage, freeze-thaw cycles, laboratory batch effects, and assay detection limits. Without intra-assay/inter-assay variation or blinding of laboratory staff, reproducibility is difficult to judge. | Report specimen processing protocol, assay manufacturer and detection range, duplicate testing, coefficients of variation, and whether laboratory personnel were blinded to case/severity status. |
“This pilot study supports an association between IL-6 and ENL severity, but the ROC cut-off appears to have been developed internally in a very small sample with highly skewed IL-6 values. Could the authors clarify whether non-parametric or log-transformed analyses were performed, and whether IL-6 remains independently associated with severity after accounting for leprosy spectrum, bacterial load, ENL duration, and MDT status? A longitudinal external-validation cohort would be needed before using 24.15 pg/mL clinically.”
| Additional setback | Why it matters | How it could be improved |
|---|---|---|
| Post-enrolment attrition and competing-risk bias | Of 52 enrolled patients, 3 died within the first month and 5 were lost to follow-up within 6 months. The final analysis used 44 patients. Excluding early deaths, including deaths related to sepsis and disease flare, may bias relapse estimates toward patients well enough to remain under observation. | Report outcomes for the full enrolled cohort, describe baseline characteristics of excluded patients, and use time-to-event methods that treat death as a competing risk. Apply sensitivity analyses for loss to follow-up. |
| Risk of overfitting in biomarker cut-offs | The anti-DSG3 threshold of 120 RU/mL and the combined predictive rule were derived from the same small cohort, with only 18 early relapses. Multiple biomarkers, time points, percentage changes, and ROC analyses were assessed. This increases false-positive findings and inflates apparent predictive accuracy. | Pre-specify the main predictor and cut-off, restrict the number of candidate predictors according to event count, adjust for multiple testing, and validate the model in an independent cohort. |
| The reported PPV is modest | The combined criterion had a PPV of 64.7%. In practical terms, about one in three patients labelled “high risk” would not relapse early. Giving maintenance rituximab to all such patients could expose some to unnecessary immunosuppression, infection risk, and cost. | Present full 2x2 diagnostic tables with confidence intervals, decision-curve analysis, and clinical consequences. Test biomarker-guided maintenance rituximab against standard follow-up in a randomised trial before recommending it. |
| Informative censoring in the longitudinal biomarker analysis | Blood sampling stopped at relapse. Therefore, patients who relapse early contribute fewer later measurements, whereas those who remain well contribute more. Comparisons at 9-12 months may represent a selected group of patients who had already remained relapse-free long enough to be sampled. | Use mixed-effects longitudinal models jointly with survival analysis, or time-dependent Cox models. Clearly report the denominator available for every biomarker at each time point. |
| Potential “look-ahead” bias in grouping patients by eventual relapse time | Patients were classified as early relapsers, late relapsers, or non-relapsers based on future outcomes, then their earlier biomarker trajectories were compared. This is useful descriptively but may overstate real-time predictive performance. | Use landmark analyses at 3, 6, 9, and 12 months. At each landmark, include only patients still relapse-free and test whether the biomarker predicts future relapse. |
| Peripheral CD19 percentage may not be an adequate measure of B-cell biology | CD19+ cells were expressed as a percentage of total lymphocytes. Steroids, infection, and changes in total lymphocyte counts can alter this percentage without equivalent changes in absolute B-cell number. The authors themselves state that absolute B-cell counts could not be estimated. Also, peripheral blood B-cell measurements may not represent pathogenic tissue-resident or autoreactive B-cell populations. | Measure absolute CD19+ counts alongside percentages, lymphocyte counts, and clinically relevant B-cell subsets such as transitional, naïve, memory, plasmablast, and regulatory B cells. |
| Relapse ascertainment may have observer bias | Relapse requires judging new lesions and whether they persist without spontaneous healing. If clinicians knew biomarker results, even unconsciously, this could influence assessment or follow-up intensity. | Predefine relapse adjudication, blind outcome assessors to biomarker results, and use standardised photographs and PDAI/POLIS scoring at suspected relapse. |
| Treatment recommendation exceeds the evidence | The conclusion suggests that high-risk patients “may benefit” from maintenance rituximab at 6 months. But this was an observational prognostic study, not an intervention study. It cannot establish that biomarker-guided maintenance treatment prevents relapse or has an acceptable safety profile. | Phrase this as a hypothesis. Conduct a prospective randomised trial comparing biomarker-guided maintenance rituximab with usual care, including infection, hypogammaglobulinaemia, cost, and quality-of-life outcomes. |
| Limited generalisability of the rituximab regimen | The cohort received a biosimilar using the rheumatoid arthritis protocol, in a tertiary Indian referral centre, mostly after prior systemic treatment. Findings may not translate to first-line rituximab, lymphoma-protocol dosing, other biosimilars, or different health-care settings. | Validate across centres, ethnic/geographic settings, treatment lines, rituximab regimens, and originator versus biosimilar products. |
“The prospective serial sampling is a major strength, but the proposed anti-DSG3 cut-off and combined prediction rule were both derived in only 44 analysable patients, with 18 early relapses and many repeated biomarker comparisons. Has the model been internally validated, and could the authors provide confidence intervals for sensitivity, specificity, PPV, and NPV? Before recommending maintenance rituximab, the strategy needs external validation and ideally a biomarker-guided randomised trial.”
“Because biomarkers were no longer sampled after relapse, later time-point comparisons may be affected by informative censoring. A landmark or time-dependent survival analysis would better establish whether a 3- or 6-month biomarker result predicts subsequent relapse in patients who are still clinically in remission at that point.”
More comments in first study
erythema nodosum leprosum IL-6 serum biomarker diagnostic studies limitations controls disease severity
erythema nodosum leprosum biomarkers
“As this is described as a pilot study, was a prior sample-size calculation performed? With only 35 cases, the estimate of correlation and especially the ROC-derived cut-off may be imprecise.”
“The study included all willing ENL patients presenting to one hospital and willing healthy volunteers as controls. Could this preferentially select more symptomatic or referral-level ENL cases and healthier-than-average controls?”
“The comparison with healthy controls establishes that ENL is inflammatory, but in practice the diagnostic challenge is distinguishing ENL from non-reactional multibacillary leprosy, type 1 reaction, infection, or another inflammatory event. Why were these clinically relevant disease-control groups not included?”
“Age, sex, BMI, and rural/urban background were reported, but did the analysis account for the underlying leprosy spectrum, bacterial index, disease duration, or current MDT phase?”
“Was blood collected during the same phase of ENL for all patients, for example within a defined number of days from onset? Were acute, recurrent, and chronic ENL analysed separately?”
“EESS includes pain and wellbeing visual analogue scales, which are patient-reported and potentially influenced by mood, analgesic use, and social factors. Was EESS assessment standardised or performed by assessors blinded to IL-6 results?”
“Since orchitis is not applicable in women, males and females have different maximum possible EESS scores and different thresholds for mild versus moderate-severe ENL. Could this affect comparability of severity classification across sexes?”
“The ROC curve classifies moderate-severe versus mild ENL using a threshold derived from the same EESS score with which IL-6 was correlated. Does this add clinical information beyond the EESS itself?”
“The authors suggest earlier detection and enhanced management, but no pre-ENL samples were taken. Does the study demonstrate severity association rather than early prediction?”
“The study excluded patients who had received ENL treatment in the previous month, but were details of MDT exposure, previous corticosteroid or thalidomide exposure, and other anti-inflammatory medicines recorded?”
“The case IL-6 range was wide, from 3.6 to 136 pg/mL, and the SD exceeded the mean. Could a few high values be driving the mean difference and the correlation?”
“The paper reports p-values, but what are the confidence intervals for the IL-6 group difference, Spearman correlation of 0.61, AUC, sensitivity, and specificity?”
“Were samples tested in duplicate? What were the assay’s lower detection limit and intra-assay/inter-assay coefficients of variation? Were all samples analysed in one batch or across multiple batches?”
“Even if 24.15 pg/mL separates mild from moderate-severe ENL in this cohort, how would this change treatment when EESS already provides clinical severity assessment?”
“My concern is that healthy controls are not the clinically relevant comparator. Since IL-6 is a non-specific inflammatory cytokine, the key comparison should be with non-reactional multibacillary leprosy and type 1 reaction, ideally matched for bacillary burden and MDT stage.”
“The IL-6 distribution appears highly skewed, with a standard deviation greater than the mean and a wide range. Were non-parametric analyses, log transformation, and outlier sensitivity analyses performed? The internally derived ROC cut-off also needs external validation.”
“This cross-sectional study shows association with concurrent ENL severity, but does not establish early prediction, treatment responsiveness, or a role as a therapeutic target. Serial sampling before, during, and after ENL would be needed for those claims.”
More in second study
pemphigus rituximab relapse anti DSG3 CD19 biomarkers predictors limitations prospective cohort
pemphigus rituximab relapse biomarkers
“The biomarker measurements stop at 12 months, whereas patients were followed for relapse until 24 months. How can the study fully characterise biomarker trajectories before late relapse, especially in patients relapsing between 12 and 24 months?”
“Since the first post-rituximab B-cell measurement was at 3 months, how can the authors distinguish true failure of initial B-cell depletion from successful depletion followed by early B-cell repopulation?”
“Could anti-DSG3 be predicting relapse partly because it is a marker of mucosal-predominant pemphigus vulgaris rather than an independent relapse biomarker?”
“The investigators tested several biomarkers at multiple time points, compared several relapse groups, examined percentage changes, and derived ROC cut-offs. Was any adjustment made for multiple comparisons?”
“The study identifies anti-DSG3 and incomplete B-cell depletion as predictors, but were they tested together with clinical predictors in a multivariable survival model?”
“The combined predictor has a PPV of 64.7% and NPV of 87.5%, but what are the 95% confidence intervals around these estimates?”
“The anti-DSG3 threshold of 120 RU/mL is assay-specific. Can it be applied to a different ELISA kit, laboratory, or calibration standard?”
“Were steroid tapering, adjuvant immunosuppression, rescue treatment, or treatment adherence recorded longitudinally and incorporated into the analysis?”
“Three patients died within the first month, including patients with sepsis and paradoxical disease flare, and were excluded from the final analysis. Could this remove patients with the most severe disease biology from the biomarker analysis?”
“Five patients were lost to follow-up within six months. Were their baseline disease activity and biomarkers different from those who completed follow-up?”
“Patients were reviewed every three months after month 3. Could a relapse have occurred and partially resolved between scheduled visits, or could the recorded relapse date be later than the true onset?”
“Were clinicians who determined relapse blinded to anti-DSG and B-cell results?”
“Could peripheral CD19+ counts underestimate the relevant B-cell activity in lymphoid tissue, skin, or mucosa?”
“CD19+ cells were expressed as a percentage of total lymphocytes. Could changes in lymphocyte count, steroid exposure, or intercurrent infection change the percentage without a true increase in B-cell number?”
“The authors suggest that patients with high-risk biomarkers may benefit from maintenance rituximab at six months. However, can an observational prognostic study determine whether this intervention prevents relapse and is safe?”
“If the biomarker rule is intended to trigger extra rituximab, were immunoglobulin levels, infections, hospitalisations, and vaccine responses assessed in relation to repeat B-cell depletion?”
“The first post-rituximab CD19 measurement was at three months. Therefore, ‘incomplete depletion’ may combine true failure of initial depletion with early repopulation. Measuring B cells at two to four weeks after infusion would clarify whether this is a pharmacodynamic failure or early recovery.”
“The anti-DSG3 cut-off of 120 RU/mL and the combined predictive rule were developed after examining several biomarkers at multiple time points in only 44 patients. This makes external validation and confidence intervals essential before applying the rule clinically.”
“The study supports closer monitoring of high-risk patients, but it cannot by itself establish that giving maintenance rituximab at six months prevents relapse safely. That needs a biomarker-guided interventional trial.”