This is an article for a journal club… tell me what are the setbacks of this study(other than those already mentioned) and how we can correct them, basically what we could have done differently to improve this study
excimer laser vitiligo
308 nm excimer laser vitiligo systematic review randomized trial limitations
~/
mkdir -p "/tmp/workspace/journal-club/extracted" && python -c "import pdfplumber; p='/tmp/workspace/attachments/0ac3f1d0-cdb2-413f-a0e8-27f25cce7233/308 nm Excimer Las Int J Dermatology - 2025 - Qi - _260820_232742.pdf'; out='/tmp/workspace/journal-club/extracted/article.txt'; pdf=pdfplumber.open(p); open(out,'w',encoding='utf-8').write('\\n\\n'.join('--- PAGE %d ---\\n%s'%(i+1,(page.extract_text() or '')) for i,page in enumerate(pdf.pages))); print('pages',len(pdf.pages),'out',out)"
"Limitations" in ~/journal-club/extracted
"random" in ~/journal-club/extracted
"positive control" in ~/journal-club/extracted
"2 | Methods" in ~/journal-club/extracted
"Table 1" in ~/journal-club/extracted
PMID: 38230424
https://pubmed.ncbi.nlm.nih.gov/25428573
PMID: 40988292
"Table 1" in ~/journal-club/extracted
"CONSORT" in ~/journal-club/extracted
"intention" in ~/journal-club/extracted
| Issue | Why it matters | What could have been done differently |
|---|---|---|
| 1. Inadequate reporting of randomization and allocation concealment | The paper calls itself a multicenter RCT, but does not state the random-sequence method, block size, stratification, allocation concealment, or who enrolled and assigned participants. This leaves risk of selection bias. | Report randomization according to CONSORT: computer-generated sequence, central web-based allocation or opaque sequentially numbered envelopes, and clearly state who was blinded to allocation. Stratify randomization by center, facial versus non-facial site, baseline VASI, and disease duration. |
| 2. Major numerical inconsistency in the sample | The abstract states 251 randomized participants, whereas Methods states 240 were enrolled and randomized, 60 per arm, after exclusions for poor adherence or missed follow-ups. This is confusing and threatens confidence in the trial flow. | Include a CONSORT flow diagram showing screened, excluded, randomized, treated, completed, lost to follow-up, and analyzed patients for each arm. Reconcile the 251 versus 240 discrepancy explicitly. |
| 3. Excluding patients for poor adherence or missed follow-up before randomization can introduce selection bias | The study says people with at least three missed follow-ups or poor adherence were excluded before the final randomized cohort was formed. This preferentially selects highly adherent patients and can overestimate effectiveness in routine practice. | Randomize eligible patients first and analyze all randomized participants by intention-to-treat. Adherence should be measured and reported, not used to create an artificially adherent study population. A per-protocol analysis can be presented secondarily. |
| 4. The analysis appears to use completers rather than intention-to-treat | Response rate was defined using the number of completers as the denominator (p. 2285). Yet patients were lost during the 1-year post-treatment follow-up. Excluding them can bias results if withdrawal is related to lack of benefit or toxicity. | Use an intention-to-treat primary analysis, retain all randomized patients in their assigned group, specify how missing VASI data were handled, and use multiple imputation or mixed models under stated assumptions. Present per-protocol results only as sensitivity analyses. |
| 5. No placebo or sham control, and no participant blinding | Oral baricitinib versus no oral treatment is obvious to participants. Knowledge of treatment can affect adherence, expectations, reporting of adverse effects, sun-exposure behavior, and even attendance at laser sessions. Blinded dermatologists reduce observer bias but cannot eliminate performance bias. | Use double-dummy design where feasible: baricitinib plus placebo laser/sham procedures, or matched oral placebo for groups without baricitinib. At minimum, standardize contact time, counseling, and follow-up intensity across arms. |
| 6. The “positive control” is not a clean comparator for the main question | The positive-control group received tacrolimus plus excimer laser, while the intervention group received baricitinib plus excimer laser. Thus it tests two different combination regimens, but does not cleanly isolate whether benefit is due to baricitinib, the treatment intensity, or another difference. | Predefine one primary estimand: for example, “Does adding baricitinib to excimer laser improve VASI compared with excimer laser alone?” The factorial four-arm design is reasonable, but formal interaction testing should be used to test synergy: baricitinib effect, laser effect, and the baricitinib-by-laser interaction. |
| 7. “Synergy” is claimed but not formally demonstrated | A superior combination arm does not automatically prove biological or statistical synergy. The paper reports better outcomes with the combination but does not describe a factorial interaction test. | Analyze the 2 x 2 factorial structure using a mixed-effects model with treatment terms for baricitinib, laser, and their interaction. Report the interaction estimate, 95% confidence interval, and p value. If it is not significant, call the result “superior combination therapy,” not “synergy.” |
| 8. Incomplete control of important baseline prognostic factors | Facial lesions respond much better than acral or non-facial lesions. Disease duration, leukotrichia, lesion size, initial VASI, activity, skin phototype, and prior treatment history may also strongly affect response. The paper reports broad comparability, but detailed arm-level balance and subgroup distribution are not clearly shown in the main report. | Present full baseline data for each group, including anatomical subsite rather than only facial/non-facial classification, baseline target-lesion VASI, leukotrichia, skin type, prior therapies, and duration of stability. Stratify randomization and adjust the primary model for prespecified prognostic variables. |
| 9. Center effects are not adequately addressed | This was a four-center study. Variation in laser delivery, MED testing, photography, counseling, adherence support, dermoscopy technique, and patient mix can influence outcomes. “Multicenter” alone does not solve this. | Include study center as a random effect in the mixed model, report recruitment and results by center, and standardize laser calibration, photography, assessor training, and protocol adherence. A central reading panel for photographs and dermoscopy would help. |
| 10. The statistical approach is not fully appropriate for repeated dermoscopy measurements | Pigment-island density was compared at many time points using separate Kruskal-Wallis tests. Repeated observations within the same patient are correlated, and multiple testing increases false-positive risk. | Use a longitudinal mixed-effects model or generalized estimating equation for dermoscopic counts. Include time, treatment, and treatment-by-time interaction. Predefine one or two clinically relevant time points, and adjust secondary multiple comparisons. |
| 11. Multiple outcomes and subgroup analyses raise false-positive risk | The paper evaluates VASI over time, categorical response, dermoscopic island density at repeated visits, site, sex, disease duration, age, pigment density, stability, recurrence, and safety. There is no clear hierarchy or multiplicity adjustment. | Specify a single primary endpoint, such as mean percent VASI improvement at week 52, and a limited hierarchy of secondary outcomes. Adjust for multiplicity, for example with Holm procedures or a gatekeeping strategy. Label exploratory analyses as exploratory. |
| 12. Outcome measurement may not fully capture clinically meaningful benefit | VASI is useful, but the study gives no patient-reported quality-of-life outcome, satisfaction score, or assessment of cosmetic acceptability. A statistically large VASI improvement may not equal an important improvement for the patient. | Add validated patient-centered endpoints such as DLQI, Vitiligo Impact Patient Scale, treatment satisfaction, and patient global assessment. Photograph-based blinded global assessments could complement VASI. |
| 13. Dermoscopic “melanocyte islands” may be a promising marker but is not yet validated as a predictive biomarker | The threshold of at least 5 islands/cm² at week 4 was associated with improved response. But it was assessed and tested within the same dataset, so it may be an optimistic, data-derived cut-off rather than a validated predictor. | Prespecify the threshold based on prior evidence, or derive it in one cohort and validate it in an independent external cohort. Report discrimination and calibration, not only an association p value. |
| 14. Safety conclusions are too strong for the sample size and exposure period | Sixty patients per baricitinib arm and 52 weeks of exposure cannot reliably exclude uncommon but important JAK-inhibitor harms such as serious infection, herpes zoster, venous thromboembolism, major cardiovascular events, malignancy, or substantial laboratory abnormalities. “No serious events” means none were observed, not that risk is absent. | Report adverse events with denominators, event rates, severity, timing, laboratory changes, confidence intervals, and prespecified adverse-event definitions. Use a larger safety-focused cohort or longer registry follow-up before making broad safety claims. |
| 15. No formal sample-size or power calculation is reported | It is unclear whether 60 participants per group was chosen to detect a clinically meaningful difference in VASI, recurrence, or adverse effects. The study may be underpowered for safety, subgroup comparisons, and recurrence. | State the expected between-group VASI difference, standard deviation, alpha, target power, anticipated attrition, and resulting sample size. Power the trial for its stated primary outcome. |
| 16. Follow-up is still short for a chronic relapsing disease | A 1-year follow-up is helpful, but vitiligo can relapse later. Also, the abstract says “no recurrence,” whereas the text reports an 8.8% relapse rate in the combination group (p. 2289). This internal inconsistency needs clarification. | Clearly define relapse, report its denominator and timing, use Kaplan-Meier time-to-relapse analysis, and extend follow-up to at least 2 years if possible. Correct the abstract so it matches the main results. |
| 17. Treatment exposure and co-interventions are insufficiently characterized | Laser dose changes were allowed in response to erythema or blisters. Actual cumulative dose, number of completed sessions, missed sessions, baricitinib adherence, tacrolimus adherence, sun exposure, and use of cosmetics or other therapies could differ by group and influence response. | Report delivered laser sessions, cumulative fluence, dose reductions, adherence to oral and topical drugs, protocol deviations, sun-exposure advice, and prohibited/rescue treatments for each group. |
| 18. External validity is narrower than simply “adult localized disease” | All sites were in China, and participants had stable, treatment-resistant disease affecting no more than 10% body surface area. Results may not apply to children, active disease, generalized vitiligo, acral disease, people with major comorbidity, or more ethnically diverse populations. | Run a larger pragmatic international trial with diverse skin phototypes and prespecified strata for disease activity, extent, body site, age, and prior therapy exposure. |
This is an article for a journal club… tell me what are the setbacks of this study(other than those already mentioned) and how we can correct them, basically what we could have done differently to improve this study
JAK inhibitors vitiligo systematic review
Lasheras-Perez JAK inhibitors International Journal Dermatology 2025 vitiligo study
~/
mkdir -p "/tmp/workspace/journal-club/jaks-extracted" && python -c "import pdfplumber; p='/tmp/workspace/attachments/0ac3f1d0-cdb2-413f-a0e8-27f25cce7233/jaks Lasheras‐Pérez - ImmunoInt J Dermatology - 2025 - _260820_232823.pdf'; out='/tmp/workspace/journal-club/jaks-extracted/article.txt'; pdf=pdfplumber.open(p); open(out,'w',encoding='utf-8').write('\\n\\n'.join('--- PAGE %d ---\\n%s'%(i+1,(page.extract_text() or '')) for i,page in enumerate(pdf.pages))); print('pages',len(pdf.pages),'out',out)"
"Limitations" in ~/journal-club/jaks-extracted
"random" in ~/journal-club/jaks-extracted
"Methods" in ~/journal-club/jaks-extracted
PMID: 39713946
"2.1" in ~/journal-club/jaks-extracted
"3.1" in ~/journal-club/jaks-extracted
"Discussion" in ~/journal-club/jaks-extracted
"retrospective" in ~/journal-club/jaks-extracted
"Table 1" in ~/journal-club/jaks-extracted
lichen planopilaris JAK inhibitor
PMID: 37852211
https://pubmed.ncbi.nlm.nih.gov/37852211
| Additional setback | Why it matters | How the study could be improved |
|---|---|---|
| 1. The molecular and clinical parts cannot establish a treatment-response relationship | The IHC samples were explicitly independent of the clinical cohort. Therefore, demonstrating TYK2 or JAK3 staining in one set of patients cannot show that those particular patients respond to JAK inhibitors, nor that staining predicts response. | Use a prospective translational cohort: biopsy the same patients before treatment, during therapy, and after therapy, then correlate baseline and change in TYK2/JAK3 activity with LPPAI, itch, hair counts, and progression. |
| 2. IHC detects protein presence, not necessarily pathway activation | “Overexpression” on IHC does not prove that the JAK-STAT pathway is actively signaling. JAK protein may be present but inactive. The authors mention absence of phospho-JAKs, but the broader implication is that the mechanistic conclusion should be cautious. | Measure phosphorylated JAK/STAT proteins, especially p-STAT1 and p-STAT3, plus downstream interferon-responsive genes such as CXCL9, CXCL10, and IFN-gamma-associated transcripts. Confirm IHC with RNA sequencing, quantitative PCR, western blotting, or spatial transcriptomics. |
| 3. No cell-type identification in the inflammatory infiltrate | The paper states that TYK2-positive inflammatory cells occur in lesions. But it does not establish whether these cells are CD4+ T cells, CD8+ T cells, macrophages, dendritic cells, neutrophils, follicular epithelial cells, or fibroblasts. This weakens the biological interpretation and does not identify the therapeutic target cell. | Use double or multiplex immunofluorescence: TYK2/JAK3 with CD3, CD4, CD8, CD68, CD11c, MPO, cytokeratin, and fibroblast markers. This would show which cells express the relevant pathway components. |
| 4. Healthy controls alone are an insufficient comparator | Healthy scalp differs from diseased scalp in inflammation, fibrosis, follicular destruction, and tissue remodeling. Therefore, a difference from healthy controls does not establish that TYK2/JAK3 is specific to LPP, FFA, or FD rather than a nonspecific marker of scalp inflammation or scarring. | Add disease controls: psoriasis, seborrheic dermatitis, discoid lupus erythematosus, alopecia areata, central centrifugal cicatricial alopecia, nonscarring inflammatory scalp disorders, and other neutrophilic cicatricial alopecias. |
| 5. Potential demographic and anatomical mismatch of controls | The IHC sample is very small, with seven healthy controls and different sex distributions across disease groups. Scalp region, sex, age, hair type, and menopausal status can influence follicular biology and immune milieu. FFA, in particular, predominantly affects women. | Match controls to cases by age, sex, scalp site, ethnicity, and, where relevant, menopausal status. Report these factors and adjust for them in analysis. |
| 6. Unclear disease activity and treatment status at the time of biopsy | JAK expression may differ between active inflammatory disease and burnt-out scarred disease. Previous corticosteroids, hydroxychloroquine, tetracyclines, or other immunomodulators could alter inflammatory-cell density and pathway expression. | Prospectively collect standardized biopsies from clearly active lesion borders and specify disease activity, duration, histologic stage, and all treatment exposures before biopsy. Ideally require a washout period when ethically feasible. |
| 7. Possible measurement bias in H-score assessment | H-score is semi-quantitative and can be affected by tissue processing, staining batch, image selection, threshold setting, and reader interpretation. The report does not clearly describe blinded scoring, inter-rater reliability, digital image quantification, or batch controls. | Blind pathologists to diagnosis, use at least two independent readers, report intra-class correlation or kappa, digitize whole-slide images, and use a prespecified image-analysis pipeline. Include positive and negative controls for every staining batch. |
| 8. Multiple comparisons with very small groups | The IHC study compares several proteins, several diseases, positivity rates, and H-scores with groups of 7 to 12 samples. This creates a risk of chance findings, especially for marginal results. | Predefine a primary molecular endpoint, reduce exploratory comparisons, and adjust secondary analyses for multiplicity. Replicate the findings in an independent validation cohort. |
| 9. The clinical cohort contains different drugs with different selectivity and doses | Upadacitinib, baricitinib, and abrocitinib have different JAK selectivity, dosing, pharmacology, and risk profiles. Pooling all 19 patients as “oral JAK inhibitors” makes it impossible to know which agent was associated with benefit or toxicity. | Study one JAK inhibitor at one standardized dose first. If several agents are included, power the study for prespecified drug-specific comparisons and present outcomes separately rather than pooled. |
| 10. LPP and FFA were combined despite potentially different disease behavior | Sixteen patients had classic LPP and three had LPP plus FFA. These phenotypes can differ in pattern, progression, symptoms, response, and patient demographics. A combined result may obscure clinically important differences. | Enroll adequate numbers of each phenotype and stratify randomization or analysis by classic LPP, FFA, and overlap disease. Report individual phenotype-specific outcomes. |
| 11. Confounding by indication is likely | In retrospective practice, clinicians may preferentially prescribe an oral JAK inhibitor to patients with more severe, more active, refractory, or rapidly progressive disease. Those treatment decisions are not random. | Conduct a randomized controlled trial. If retrospective data must be used, create a comparator group and use propensity-score matching or inverse-probability weighting based on baseline severity, activity, disease duration, phenotype, prior treatment failures, and center. |
| 12. Regression to the mean and natural fluctuation could account for part of the improvement | Patients often receive a new systemic treatment when disease is particularly symptomatic or active. Symptoms and LPPAI can later improve naturally, even without the tested treatment. Without a control group, the observed decline cannot confidently be attributed to JAK inhibition. | Include placebo or active-standard-care control. A randomized withdrawal design could also be useful after initial disease stabilization. |
| 13. Concomitant treatment history is more problematic than simply “present” | Twelve of 19 patients continued concomitant treatment that had started at least 3 months earlier. The authors infer that lack of prior efficacy makes those treatments unimportant, but delayed effects, treatment interaction, or improved adherence cannot be excluded. | Either discontinue nonessential therapies before enrollment, standardize background therapy across all groups, or randomize JAK inhibitor versus placebo as an add-on to the same background regimen. |
| 14. Variable treatment and follow-up duration may distort results | Median JAK-inhibitor exposure was 7 months, with a range of 3 to 14 months. Patients with longer exposure may be more likely to improve or more likely to have adverse events detected. A “last follow-up” analysis is not a standardized endpoint. | Use fixed assessment times, for example baseline, weeks 12, 24, 36, and 52. Analyze repeated measures with a mixed-effects model and report time-to-response, time-to-flare, and treatment persistence. |
| 15. The study mainly demonstrates reduced activity and itch, not convincing restoration of lost hair | LPPAI and itch-NRS are useful activity outcomes. However, LPP is a scarring alopecia, and a reduction in inflammation does not necessarily mean meaningful hair regrowth or reversal of permanent follicular destruction. The reduction in %SCALP may be influenced by clinical estimation. | Add standardized global photographs, dermoscopic follicular counts, phototrichograms, hair density measurements, and blinded investigator/patient global assessments. Separate outcomes into disease arrest, reduction in symptoms, prevention of further loss, and actual regrowth. |
| 16. Diagnostic confirmation was incomplete in the clinical cohort | Three of 19 patients did not undergo biopsy and were diagnosed clinically. In a small sample, even a few misclassified cases could affect results, especially because cicatricial alopecias can overlap clinically. | Require baseline histopathologic confirmation for all participants, or use an independent diagnostic adjudication panel combining clinical, dermoscopic, and histologic information. |
| 17. Safety conclusions are underpowered | The paper reports mostly mild adverse events but one severe neutropenia and liver enzyme elevations. With only 19 patients and short, variable follow-up, it cannot establish that oral JAK inhibitors have a favorable long-term safety profile in LPP. Rare but serious risks may not appear. | Use prespecified safety definitions, standardized laboratory monitoring, event rates with confidence intervals, longer follow-up, and a larger prospective registry or RCT. Avoid broad safety claims from this sample. |
| 18. Patient-important outcomes are still incompletely represented | The authors appropriately note that DLQI was absent and may be imperfect for scarring alopecia. But this should not mean quality of life is omitted altogether. Hair loss can affect psychological well-being, social functioning, camouflage behavior, and treatment burden. | Use both generic and alopecia-specific instruments, such as the Alopecia Areata Symptom Impact Scale only if validated for the population, or preferably a validated cicatricial-alopecia-specific patient-reported measure. Include patient satisfaction, perceived progression, and acceptability of systemic treatment. |
| 19. No preregistration or clearly prespecified hierarchy of outcomes is apparent | Exploratory retrospective studies are susceptible to selective outcome selection and emphasis on statistically significant findings. | Prospectively register the study, define a primary outcome, specify the statistical plan before data analysis, and report all planned outcomes, including negative findings. |