Hypothesis and Underlying Principles Behind Estimating Sample Size and Power
1. Introduction
Sample size estimation means deciding the minimum number of participants needed before starting a research study.
Power means the ability of a study to detect a real difference or real effect, if it actually exists.
In MPT research, proper sample size is needed to avoid:
- Very small study with unreliable results
- Missing a useful physiotherapy intervention effect
- Wasting time, money and participants by recruiting unnecessarily large numbers
- Ethical problems due to exposing subjects to research without adequate chance of getting meaningful results
Key idea: Sample size is calculated from the study hypothesis, expected effect, variability, chosen significance level and desired power.
2. Hypothesis
Definition
A hypothesis is a clear, testable statement or prediction about the relationship or difference between variables.
Example
“Task-specific balance training improves Berg Balance Scale score more than conventional exercise in stroke patients.”
Types of Hypothesis
A. Null Hypothesis (H₀)
The null hypothesis states that there is no real difference, no association or no treatment effect.
Example:
H₀: There is no difference in Berg Balance Scale score between task-specific balance training and conventional exercise.
B. Alternative Hypothesis (H₁ or Ha)
The alternative hypothesis states that a real difference, association or effect exists.
Example:
H₁: Task-specific balance training produces greater improvement in Berg Balance Scale score than conventional exercise.
Types of Alternative Hypothesis
1. Non-directional hypothesis
States that a difference exists, but does not say which group will be better.
H₁: There is a difference in pain reduction between ultrasound therapy and TENS.
This usually needs a two-tailed test.
2. Directional hypothesis
States the expected direction of difference.
H₁: Ultrasound therapy reduces pain more than TENS.
This may use a one-tailed test, if justified before data collection.
Exam point: In most clinical studies, a two-tailed test is preferred because the intervention may produce benefit or harm in either direction.
3. Hypothesis Testing
Hypothesis testing is a statistical method used to decide whether the observed study result is likely due to a real effect or only due to chance.
Basic Steps
- State H₀ and H₁
- Select the significance level, usually α = 0.05
- Collect data from the calculated sample
- Apply the appropriate statistical test
- Obtain the p value
- Compare p value with α
- Reject or fail to reject H₀
Interpretation
- If p < 0.05, result is statistically significant. Reject H₀.
- If p ≥ 0.05, result is not statistically significant. Fail to reject H₀.
Important: Correct wording is “fail to reject H₀”, not “prove H₀ is true.”
A null hypothesis usually represents the default condition of no difference, while H₁ represents a difference or effect (
Schwartz's Principles of Surgery, p. 1710).
4. Errors in Hypothesis Testing
When a sample is used to make conclusions about a population, errors can occur.
| Actual situation | Research decision | Result |
|---|
| H₀ is true | Fail to reject H₀ | Correct decision |
| H₀ is true | Reject H₀ | Type I error |
| H₀ is false | Fail to reject H₀ | Type II error |
| H₀ is false | Reject H₀ | Correct decision |
A. Type I Error (α error)
Definition
Type I error occurs when the researcher rejects H₀ even though H₀ is actually true.
It is a false-positive result.
Example
A researcher concludes that a new physiotherapy protocol improves knee pain, but in reality it has no extra benefit.
Alpha level
- Represented by α
- Usually set at 0.05
- α = 0.05 means accepting a 5% risk of false-positive conclusion
B. Type II Error (β error)
Definition
Type II error occurs when the researcher fails to reject H₀ even though H₀ is false.
It is a false-negative result.
Example
A new gait-training programme is truly effective, but the study concludes that it is not effective because the sample was too small.
Beta level
- Represented by β
- Commonly set at 0.20
- β = 0.20 means 20% risk of missing a true effect
Type I error is rejection of a true H₀, whereas Type II error is failure to reject a false H₀ (
Schwartz's Principles of Surgery, p. 1711-1713).
5. Statistical Power
Definition
Statistical power is the probability that a study will correctly detect a true effect.
[
\textbf{Power = 1 - β}
]
Common Values
| Beta (β) | Power (1 - β) | Meaning |
|---|
| 0.20 | 80% | Common minimum power |
| 0.10 | 90% | Better power, needs larger sample |
| 0.05 | 95% | Very high power, needs still larger sample |
Example
If a study has:
[
β = 0.20
]
Then:
[
Power = 1 - 0.20 = 0.80 = 80%
]
This means there is an 80% chance that the study will detect the treatment effect if that effect truly exists.
Power is defined as 1 minus β and is affected by sample size, effect size and variance (
The Harriet Lane Handbook, p. 958).
6. Diagram: Relationship Between H₀, Errors and Power
TRUE SITUATION IN POPULATION
H₀ is true H₀ is false
------------------------------------------------
Decision: | Correct decision | Type II error
Fail to reject | True negative | False negative
H₀ | | β
------------------------------------------------
Decision: | Type I error | Correct decision
Reject H₀ | False positive | True positive
| α | Power = 1 - β
------------------------------------------------
Memory point:
α = False positive
β = False negative
Power = Ability to find a true positive result
7. Sample Size
Definition
Sample size is the number of participants required to answer the research question with acceptable precision, significance level and power.
It should be calculated before data collection.
Why Sample Size Estimation is Important
- Ensures adequate power
- Reduces Type II error
- Gives reliable and interpretable findings
- Prevents underpowered studies
- Avoids unnecessary recruitment
- Protects participant time and safety
- Helps in ethical approval and research proposal writing
- Supports proper budgeting and planning
8. Underlying Principles of Sample Size Estimation
Sample size is not selected randomly. It depends on the following main factors.
1. Research Question and Primary Outcome
The first step is to clearly define:
- Population
- Intervention
- Comparison group
- Outcome measure
- Study design
- Main or primary outcome
Example
| Component | Example |
|---|
| Population | Patients with chronic low back pain |
| Intervention | Core stabilization exercise |
| Comparison | Conventional exercise |
| Outcome | Oswestry Disability Index score |
| Design | Randomized controlled trial |
Sample size should be calculated mainly for the primary outcome, not for every outcome measure.
2. Type of Outcome Variable
The formula changes according to the type of data.
| Type of outcome | Example | Common statistical approach |
|---|
| Continuous | VAS pain score, ROM, 6-minute walk distance | t test, ANOVA |
| Categorical | Improved/not improved, fall/no fall | Chi-square test, proportion test |
| Correlation | Relationship between pain and disability | Correlation analysis |
| Time-to-event | Time to return to work | Survival analysis |
3. Level of Significance (α)
Alpha is the probability of Type I error.
- Usually α = 0.05
- In highly sensitive studies, α may be 0.01
- Lower alpha means stricter evidence is required
- Lower alpha generally increases required sample size
Relationship
Lower α → stricter significance level → larger sample required
4. Power (1 - β)
Power indicates the probability of detecting a true effect.
- Commonly selected power: 80%
- Higher power: 90%
- Higher desired power requires a larger sample
Relationship
Higher power → lower β → larger sample required
5. Effect Size
Definition
Effect size is the magnitude of difference expected between groups or before-after measurements.
Example
If expected mean reduction in VAS pain is:
- Experimental group: 4 cm
- Control group: 2 cm
Then expected difference:
[
\Delta = 4 - 2 = 2 \text{ cm}
]
This difference is the expected effect.
Important Rule
Small effect size → large sample needed
Large effect size → small sample may be enough
Effect size should be clinically meaningful and preferably based on:
- Previous studies
- Pilot study
- Published literature
- Expert opinion
- Minimum clinically important difference, or MCID
A sample-size calculation should use a realistic effect size and should state the primary outcome, variability, beta and expected difference. Arbitrary assumptions can make the calculation misleading (
Rheumatology, p. 252).
6. Variability or Standard Deviation (SD)
Variability means how widely individual values differ from the mean.
- High variability means participants are very different from each other.
- Low variability means participants are more similar.
Relationship
High SD / high variability → larger sample required
Low SD / low variability → smaller sample required
Example
Pain score may vary widely due to:
- Severity of condition
- Age
- Duration of symptoms
- Psychological factors
- Compliance with exercise
- Different medication use
Therefore, studies with highly variable pain scores need more participants.
7. One-tailed or Two-tailed Test
One-tailed test
- Tests effect in only one direction
- Requires a comparatively smaller sample
- Used only when opposite direction is not considered relevant
Two-tailed test
- Tests effect in both directions
- More commonly used in health research
- Usually requires a larger sample than one-tailed testing
MPT exam point: A two-tailed test is generally safer and more acceptable in clinical research.
8. Study Design
Required sample size differs according to design.
| Study design | Sample size consideration |
|---|
| Cross-sectional study | Prevalence, allowable error |
| Case-control study | Exposure difference, odds ratio |
| Cohort study | Incidence and expected risk |
| RCT | Difference between intervention and control groups |
| Pre-post study | Expected mean change and SD of change |
| Correlation study | Expected correlation coefficient |
| Diagnostic study | Expected sensitivity, specificity and prevalence |
9. Number of Groups
More groups usually need more participants.
Example
- Two-group RCT: exercise versus control
- Three-group RCT: exercise A versus exercise B versus control
If there are more groups, enough subjects should be present in each group for meaningful comparison.
10. Allocation Ratio
Allocation ratio means the distribution of participants between groups.
Equal allocation
[
1:1
]
Example: 30 participants in experimental group and 30 in control group.
This is usually statistically efficient.
Unequal allocation
[
2:1
]
Example: 40 participants in experimental group and 20 in control group.
Unequal allocation may be used when:
- Treatment is expensive
- Control data are easily available
- Ethical reasons require more participants to receive the better intervention
However, unequal allocation often increases the total sample required.
11. Dropout or Attrition
Some participants may leave the study due to:
- Loss to follow-up
- Non-compliance
- Travel difficulty
- Adverse event
- Withdrawal of consent
- Change in health status
Therefore, sample size should be increased for expected dropout.
Formula for dropout adjustment
[
\text{Final sample size} =
\frac{\text{Calculated sample size}}
{1 - \text{Expected dropout proportion}}
]
Example
Calculated sample size = 50 participants
Expected dropout = 20% = 0.20
[
\text{Final sample size} =
\frac{50}{1-0.20}
]
[
= \frac{50}{0.80} = 62.5
]
Therefore, recruit approximately 63 participants.
12. Cluster Sampling and Design Effect
In cluster studies, participants are recruited in groups such as:
- Hospitals
- Villages
- Schools
- Physiotherapy centres
- Wards
Participants within one cluster may be similar. Therefore, the usual sample size is multiplied by the design effect.
[
\text{Adjusted sample} =
\text{Usual sample size} \times \text{Design effect}
]
9. Important Relationships for Examination
Sample size ↑ → Power ↑ → β ↓
Effect size ↓ → Sample size ↑
Variability / SD ↑ → Sample size ↑
Desired power ↑ → Sample size ↑
Alpha ↓ → Sample size ↑
Dropout expected ↑ → Recruitment target ↑
10. Flowchart: Steps in Sample Size Estimation
Start
↓
Define research problem and primary outcome
↓
Frame H₀ and H₁
↓
Choose study design and statistical test
↓
Decide α level, usually 0.05
↓
Decide desired power, usually 80% or 90%
↓
Estimate clinically meaningful effect size
↓
Estimate SD / variability or expected proportion
↓
Calculate sample size using formula or software
↓
Adjust for number of groups, design effect and dropout
↓
Round up to nearest whole number
↓
Final recruitment target
11. Common Formulae
A. Sample Size for Estimating a Single Proportion
Used in prevalence studies.
[
n = \frac{Z^2pq}{d^2}
]
Where:
- n = required sample size
- Z = standard normal value, 1.96 for 95% confidence level
- p = expected prevalence
- q = 1 - p
- d = allowable error or precision
B. Sample Size for Comparing Two Means
Used when comparing continuous outcomes, such as VAS, ROM, walking distance or disability score.
[
n = \frac{2\sigma^2 (Z_{\alpha/2} + Z_{\beta})^2}{\Delta^2}
]
Where:
- n = sample required per group
- σ = standard deviation
- Δ = minimum meaningful difference between groups
- Zα/2 = Z value for selected alpha level
- Zβ = Z value for selected beta level
For α = 0.05 and power = 80%:
[
Z_{\alpha/2} = 1.96
]
[
Z_{\beta} = 0.84
]
Note: The formula should be selected according to study design and outcome variable. In practical research, G*Power, OpenEpi, nMaster or consultation with a statistician may be used.
12. Worked Example
Research Question
Does a 6-week core stabilization programme reduce pain more than conventional physiotherapy in chronic low back pain?
Assumptions
- Primary outcome: VAS pain score
- Expected difference between groups: 2 cm
- Standard deviation: 2.5 cm
- Alpha: 0.05
- Power: 80%
- Two groups
- Expected dropout: 15%
Interpretation
- A difference of 2 cm is the effect size of clinical interest
- SD represents expected variability in pain score
- Alpha protects against false-positive conclusion
- Power protects against missing a true treatment benefit
- Final sample should be increased to compensate for dropout
13. Underpowered and Overpowered Studies
Underpowered Study
An underpowered study has insufficient sample size.
Problems
- High chance of Type II error
- Real treatment benefit may be missed
- Non-significant result may be wrongly interpreted as “no effect”
- Results may be inconclusive
Overpowered Study
An overpowered study has unnecessarily very large sample size.
Problems
- Wastes resources and time
- May detect a statistically significant but clinically unimportant difference
- May expose more subjects than needed
Key point: Statistical significance does not always mean clinical significance.
For example, a very large study may show a statistically significant 0.2 cm reduction in VAS pain. This difference may not be important for the patient.
14. Difference Between Statistical Significance and Clinical Significance
| Statistical significance | Clinical significance |
|---|
| Based mainly on p value | Based on benefit to patient |
| Indicates result is unlikely due to chance | Indicates change is meaningful in practice |
| Influenced by sample size | Influenced by MCID and patient outcome |
| Large sample can make very small effects significant | A small effect may still be unimportant clinically |
Example
An intervention improves walking distance by 3 metres with p < 0.05.
- It may be statistically significant
- But may not be clinically meaningful for a stroke patient
15. Short Note: A Priori and Post Hoc Power
A Priori Power Analysis
Done before the study begins.
Purpose:
- To calculate required sample size
- To plan adequate recruitment
- To avoid underpowered research
This is the preferred method.
Post Hoc Power Analysis
Done after completion of the study.
Purpose:
- To estimate observed power based on obtained results
Limitation
Post hoc power has limited value when interpreting a non-significant study result. A confidence interval and effect estimate are usually more useful.
16. Common Mistakes in Sample Size Estimation
- Selecting sample size by convenience only
- Not identifying a primary outcome
- Using an unrealistic effect size
- Ignoring variability or SD
- Not accounting for dropout
- Using a one-tailed test without justification
- Not mentioning alpha and power
- Using a formula not suitable for study design
- Confusing statistical significance with clinical significance
- Concluding “no effect” only because p > 0.05 in a small study
17. One-Minute Revision Box
H₀ = No difference / no association
H₁ = Difference / association exists
Type I error = False positive = α
Type II error = False negative = β
Power = 1 - β
Common alpha = 0.05
Common power = 80% or 90%
Small effect size = larger sample needed
High variability = larger sample needed
Expected dropout = increase final recruitment target
Sample size must be calculated before data collection.
18. Conclusion
Sample-size estimation is based on the study hypothesis and aims to obtain a study with adequate statistical power. The investigator must decide the primary outcome, effect size, variability, alpha, beta, power, type of test, study design and expected dropout. A properly calculated sample size improves the scientific, ethical and clinical value of physiotherapy research.
For further reading, the
NIH review on sample size and power explains the relation among alpha, beta, effect size, power and required participant number.