P-value viva questions based on your thesis
Your thesis compares sitting versus lateral decubitus position during spinal anaesthesia for LSCS, with PDPH incidence as the primary outcome, in 100 patients, 50 per group.
1. What is a p-value?
Answer:
A p-value is the probability of obtaining results at least as extreme as those observed,
assuming that the null hypothesis is true. It does not tell us the probability that the null hypothesis is true. In my study, the null hypothesis for the primary outcome was that PDPH incidence is the same in the sitting and lateral-position groups. This interpretation is consistent with guidance on
p-value reporting.
2. What level of statistical significance did you use?
Answer:
I used a significance level of 5%, so p < 0.05 was considered statistically significant. This corresponds to accepting up to a 5% Type I error risk for an individual hypothesis test.
3. What was your null hypothesis for the primary outcome?
Answer:
The null hypothesis was that there is no difference in the incidence of post-dural puncture headache between the sitting and lateral decubitus groups.
The alternative hypothesis was that PDPH incidence differs between the two groups.
4. What was the key p-value for your primary outcome?
Answer:
For PDPH incidence, 9 of 50 patients (18%) in the sitting group developed PDPH compared with 2 of 50 patients (4%) in the lateral group. The thesis reports χ² = 5.01 and p = 0.025.
Using the pre-specified cutoff of 0.05, this is statistically significant and provides evidence against the null hypothesis.
5. How would you interpret p = 0.025 in Table 11?
Answer:
If there were truly no difference in PDPH incidence between positions, the probability of observing a difference at least this large, or more extreme, would be about 2.5%, under the assumptions of the test.
It does not mean that there is a 2.5% probability that the result occurred by chance, nor that there is a 97.5% probability that the lateral position is better.
6. Can you say that lateral position “prevents” PDPH because p = 0.025?
Answer:
Not solely from the p-value. I can say the data show a statistically significant association between lateral positioning and lower PDPH incidence. A causal statement depends on study design, allocation method, control of confounders, and bias.
Also, the clinical effect should be described, not only the p-value: PDPH was 14 percentage points lower in the lateral group, from 18% to 4%.
7. What are the effect measures for the primary outcome?
Answer:
From the observed data:
- Risk in sitting group: 9/50 = 18%
- Risk in lateral group: 2/50 = 4%
- Absolute risk reduction: 18% - 4% = 14%
- Relative risk: 4% / 18% ≈ 0.22
- Relative risk reduction: approximately 78%
- Number needed to treat or position laterally: 1/0.14 ≈ 7
Thus, approximately seven patients would need to receive spinal anaesthesia in the lateral position instead of sitting to prevent one PDPH event, based on this sample.
8. Why should you report confidence intervals in addition to p-values?
Answer:
A p-value indicates compatibility of the data with the null hypothesis, but it does not show the size or precision of the effect. A confidence interval gives a plausible range for the treatment or exposure effect and helps determine clinical importance.
For example, the 14% absolute difference in PDPH may be clinically important, but its confidence interval would show how precise that estimate is. P-values should not be used as the only basis for clinical interpretation, as explained in this
review of p-value misunderstandings.
9. Why did you use the chi-square test in Table 11?
Answer:
Table 11 compares two categorical variables:
- Position: sitting or lateral
- PDPH: yes or no
The chi-square test assesses whether there is an association between these categorical variables.
10. Is chi-square the best test for Table 11?
Answer:
This needs careful consideration. In Table 11, the PDPH-positive counts are small, especially only 2 patients in the lateral group. The expected count in one cell is below 5. Therefore, Fisher’s exact test is more appropriate than the ordinary chi-square test.
The thesis reports p = 0.025 using chi-square. With a two-sided Fisher’s exact test, the p-value is approximately 0.051, which is just above 0.05. Therefore, the primary finding should be interpreted cautiously and reported with an effect size and confidence interval rather than a simple significant/non-significant label.
This is an important viva point: recognize the limitation rather than defend an unsuitable test.
11. What should you say if the examiner asks why p = 0.025 and Fisher’s exact p-value differ?
Answer:
The chi-square test uses an asymptotic approximation, which can be unreliable when cell frequencies are small. Fisher’s exact test calculates an exact probability conditional on the marginal totals and is preferred when expected cell frequencies are low.
In my primary outcome table, only 11 total PDPH events occurred and one group had only two events. Therefore, Fisher’s exact test gives a more defensible result.
12. What is the interpretation of p = 0.979 for age distribution?
Answer:
The thesis reports p = 0.979 for baseline age-category distribution. This means the observed age distributions in the two groups are highly compatible with the null hypothesis of no association.
It does not prove the groups are identical in age. It only indicates that this study did not detect evidence of a difference in the categorized age distribution.
13. Why is it not ideal to use p-values alone to prove baseline comparability?
Answer:
A non-significant baseline p-value does not prove balance. It may be non-significant because of small sample size and limited power. Baseline characteristics should mainly be shown descriptively using means, standard deviations, proportions, and preferably standardized differences.
In the thesis, age, BMI, ASA status, gravida, parity, gestational age, LSCS type, and puncture level were broadly similar between groups, but “no statistically significant difference” should not be equated with “proven equal.”
14. How do you interpret p = 0.118 for mean NRS score?
Answer:
Among the 11 PDPH-positive patients, mean NRS was 3.9 ± 1.2 in the sitting group and 2.5 ± 0.7 in the lateral group, with p = 0.118.
This means that, under the t-test assumptions and the null hypothesis of equal mean NRS score, the data did not provide statistically significant evidence of a difference at the 5% level. It does not establish that headache severity was equal between groups.
Because there were only 9 and 2 patients in the two groups, the analysis has very low precision and the t-test assumptions are difficult to verify.
15. Is Student’s t-test appropriate for Table 14 and Table 15?
Answer:
It is questionable because the analyses include only 11 PDPH cases, with only 2 patients in the lateral group. With such a small group, normality and equal-variance assumptions cannot be reliably assessed.
For duration and NRS score, I would report individual data or medians and interquartile ranges, and consider a non-parametric or permutation-based comparison. Any p-value from the t-test should be interpreted cautiously.
16. How would you interpret p = 0.041 for PDPH duration?
Answer:
In the thesis, mean PDPH duration was 34.8 ± 12.6 hours in the sitting group and 14.0 ± 5.7 hours in the lateral group, with reported p = 0.041.
At face value, it suggests a statistically significant difference in mean duration. However, this result is based on 9 versus 2 affected patients. It is highly sensitive to test choice and assumptions. Therefore, I would describe it as a preliminary finding requiring cautious interpretation rather than definitive evidence.
17. What does “not statistically significant” mean?
Answer:
It means that the observed data did not cross the pre-specified threshold for rejecting the null hypothesis. It does not mean there is no effect, no difference, or clinical equivalence.
For example, a non-significant NRS comparison with p = 0.118 may reflect inadequate sample size, especially because only two lateral-position patients developed PDPH.
18. What is a Type I error?
Answer:
A Type I error is falsely rejecting a true null hypothesis, also called a false-positive result. With α = 0.05, the nominal risk is 5% for one test.
In this thesis, declaring a difference in PDPH incidence when no true difference exists would be a Type I error.
19. What is a Type II error?
Answer:
A Type II error is failing to reject a false null hypothesis, also called a false-negative result. It is denoted by β.
In this study, some secondary outcomes may be non-significant because there are too few events or too small a sample, not necessarily because there is no true difference.
20. What is statistical power?
Answer:
Statistical power is the probability of detecting a true effect of a specified size when it actually exists. It equals 1 - β.
Power depends on sample size, expected effect size, outcome variability, significance level, and event rate. With only 11 PDPH events overall, analyses restricted to PDPH-positive patients have limited power.
21. What is the problem of multiple comparisons in your thesis?
Answer:
The thesis tests many baseline variables, procedural factors, secondary outcomes, and risk factors. When many hypothesis tests are performed at α = 0.05, some statistically significant p-values may occur by chance.
The primary outcome should receive the greatest emphasis. Secondary and exploratory findings should be clearly labelled and interpreted with caution. A multiple-comparison adjustment can be considered depending on whether the analyses were pre-specified.
22. Explain p < 0.001 for traumatic tap and PDPH.
Answer:
The thesis reports traumatic tap in 4 of 11 PDPH-positive patients versus 2 of 89 PDPH-negative patients, with p < 0.001.
This indicates a strong association in the sample. However, because there are very small cell counts, Fisher’s exact test is the appropriate method. The association remains statistically strong with an exact test, but the effect estimate should be reported with a wide confidence interval because only six traumatic taps occurred.
23. Can traumatic tap be called the “strongest predictor” based on p < 0.001?
Answer:
Not reliably. A smaller p-value does not necessarily mean a factor is the strongest predictor. P-values are influenced by sample size, variability, and event frequency.
To identify independent predictors, the study would need an appropriately powered multivariable model, usually logistic regression for PDPH occurrence, with adjusted odds ratios and confidence intervals. With only 11 PDPH events, a standard multivariable model would be unstable and likely overfitted.
24. How do you interpret p = 0.022 for number of attempts and PDPH?
Answer:
The thesis found that two attempts were more frequent in PDPH-positive patients: 36.4% versus 12.4%. The reported chi-square p-value was 0.022.
However, because counts are small, Fisher’s exact test is more appropriate and gives a p-value around 0.058. Thus, the finding is suggestive of an association but should not be presented as definitively statistically significant at the 5% level without specifying the test used.
25. How do you interpret p = 0.018 for prior PDPH history?
Answer:
The thesis reports prior PDPH history in 2 of 11 PDPH-positive patients and 2 of 89 PDPH-negative patients, with p = 0.018 by chi-square.
These are sparse data, so Fisher’s exact test is preferable and gives a p-value of approximately 0.059. The observed association may be clinically plausible, but the estimate is imprecise and requires cautious wording.
26. Why was Fisher’s exact test used for readmission?
Answer:
Only one patient was readmitted and none in the lateral group were readmitted. Because expected cell counts are very low and there is a zero cell, chi-square assumptions are violated. Fisher’s exact test is appropriate.
The reported p = 0.315 means that the study does not show statistically significant evidence of a difference in readmission rates.
27. How should you interpret p = 0.029 for patient satisfaction?
Answer:
The thesis reports mean satisfaction scores of 4.32 ± 0.75 in the sitting group and 4.62 ± 0.57 in the lateral group, with p = 0.029.
This suggests evidence of a difference in mean scores. However, satisfaction on a 5-point scale is ordinal data. A t-test is often used pragmatically for multi-category scores, but an ordinal comparison such as Mann-Whitney U test, or ordinal logistic regression, may be more suitable. The clinical importance of a 0.30-point difference should also be considered.
28. What is the difference between statistical significance and clinical significance?
Answer:
Statistical significance asks whether the observed data are sufficiently inconsistent with the null hypothesis under a statistical model. Clinical significance asks whether the magnitude of difference is meaningful for patients and practice.
For PDPH, an absolute reduction from 18% to 4% is potentially clinically meaningful. For satisfaction, a mean difference of 0.30 on a 5-point scale may be statistically significant but should be interpreted in terms of patient-relevant importance.
29. Why should exact p-values be reported?
Answer:
Exact p-values, such as p = 0.025 or p = 0.029, give more information than simply reporting p < 0.05. They allow readers to judge the strength of evidence.
For very small p-values, p < 0.001 is acceptable, but “highly significant” should not be used as a substitute for effect size and confidence intervals.
30. What is the best final statistical conclusion for your thesis?
Answer:
A careful conclusion would be:
In this sample of 100 women undergoing LSCS under spinal anaesthesia, PDPH occurred less frequently in the lateral decubitus group than in the sitting group, with an observed absolute risk difference of 14 percentage points. However, the number of PDPH events was small, several contingency tables had sparse cell counts, and exact methods are preferable for the main comparison. Therefore, the findings suggest a possible benefit of lateral positioning but should be confirmed in a larger adequately powered study with effect estimates and confidence intervals.
High-yield corrections to remember in the viva
-
Do not say: “p = 0.025 means there is a 2.5% chance that the result is due to chance.”
Say: “Assuming no true difference, the probability of observing results this extreme or more extreme is 2.5%.”
-
Do not say: “p > 0.05 proves there is no difference.”
Say: “The study did not find statistically significant evidence of a difference.”
-
Do not use p-values alone. State the event rates, absolute difference, relative effect, and confidence interval.
-
Use Fisher’s exact test when expected cell frequencies are small, especially in Tables 11, 20, 22, 24, and the sparse risk-factor analyses.
-
The thesis text calls the final risk-factor summary “Table 27”, although it is labeled Table 25 in the extracted results. Be prepared to correct this formatting inconsistency if asked.