Sample Size Calculator for Experiments

Plan a two-group experiment to detect a meaningful difference with specified power. Input your pilot data or literature estimates to calculate the sample size you need.

Power of 0.80 (80%) is a common convention. Higher power requires larger samples.

Estimated mean difference between treatment and control groups. Use conservative estimates from pilot studies (see note below).

Pooled standard deviation from your pilot or prior research.

Typical rate is 10–20% for in-person studies, lower for online studies.

N per group
Total N
Recruit target
z(α)
z(1−β)
Fill in the inputs to see your sample size.
Sample size by effect size
Fig. 1 — Sample size per group required across a range of effect sizes (Cohen's d).

§2From pilot study to full experiment: a worked example

This is the workflow most experiments follow: run a small pilot to learn what effect size and variability you actually observe in your population, then use those numbers to plan your full study.

The pilot. You run a small two-group experiment with 15 participants per group (30 total). You randomly assign them to a "cognitive load" condition (5 minutes of mental arithmetic before a memory task) versus a control (5 minutes of rest). You measure recall accuracy (0–100 points). Results:
  • Control group mean: 62 points, SD = 9
  • Cognitive load group mean: 59 points, SD = 7
  • Observed difference: Δ = 3 points (cognitive load hurts recall by 3 points on average)
  • Pooled SD: σ ≈ 8 (the average of the two group SDs)
Your pilot was small and the observed difference might be inflated (winner's curse), but now you have a data-driven starting point, not a guess.
Planning the full study. You enter your pilot results into this calculator, but you apply a conservative shrinkage: instead of using Δ = 3, you assume the true effect is only Δ = 2.25 (75% of the observed pilot difference), and you stick with σ = 8. You want 80% power and α = .05. The calculator tells you that you need 128 participants per group (256 total). Accounting for a 10% dropout rate in in-person studies, you should recruit 142 per group (284 total).
Why be conservative? Pilot estimates are noisy. A small pilot is more likely to find a large effect by chance (sampling variability). If you plan your full study assuming the pilot's observed Δ is the true effect, you risk an underpowered full experiment. By shrinking the pilot Δ by 25–50%, you hedge against this bias. Better to recruit slightly more participants than to run an underpowered study.

The winner's curse: why pilot estimates overstate effects

When you run a small study, random sampling variability is large. The observed effect in your pilot is a noisy estimate of the true effect. Among all the noise, you see the one that happened to be lucky—hence "winner's curse." Additionally, researcher degrees of freedom (choice of analysis, exclusion criteria, multiple outcomes tested) can inflate pilot estimates. Solution: treat your pilot Δ as an upper bound. Use a shrinkage estimator (e.g., 50–75% of the observed Δ) when planning your full study. This is conservative but principled.

§3The formula

n = 2 × (z₍α/₂₎ + z₍β₎)² × σ² / Δ²
n
Sample size per group
z₍α/₂₎
Critical z-value for significance level α (two-tailed)
z₍β₎
Critical z-value for power (1 − β)
σ
Estimated population standard deviation (assumed equal in both groups)
Δ
Expected difference between the means (treatment − control)
Cohen's d
Effect size: d = Δ / σ. Small ≈ 0.2, medium ≈ 0.5, large ≈ 0.8

§4Statistical power and Type II error

What power means

Power (1 − β) is the probability that you will detect a true effect of the size you specified. A power of 0.80 means you have an 80% chance of rejecting H₀ if the true effect is Δ as you entered it. The other 20% is the risk of a Type II error — failing to detect a real effect.

Why 80% power is a convention

Jacob Cohen proposed 0.80 as a rule of thumb in 1969, balancing the cost of Type I error (false positive) and Type II error (false negative). It is not a law. Your field may use different standards; some clinical trials demand 90% power. But absent specific guidance, 80% is reasonable.

The role of effect size

Smaller true effects require larger samples to detect with the same power. If your expected Δ is small, you will need many more participants. This is why accurate effect size estimation from pilots or literature is so important.

§5Two-group experimental design assumptions

Random assignment

These formulas assume you randomly assign participants to treatment and control groups. If assignment is not random (e.g., volunteers choose which group), confounding variables can bias your results and inflate apparent effect sizes. Your effect size estimate will be unreliable.

Equal group sizes

The formula above assumes you put the same number of participants in each group. If you must use unequal allocation (e.g., control is cheap, treatment is expensive), you will need slightly more total participants. Consult a statistician for the adjustment.

Independence of observations

Each participant's outcome must be independent of others. If you test two family members or run multiple trials per person without accounting for clustering, you inflate Type I error (false positives). Use mixed-model or hierarchical methods if your design has dependencies.

Normality and heterogeneity of variance

The t-test is robust to moderate departures from normality with moderate sample sizes, but if your outcome is highly skewed or variances differ drastically between groups, consider a transformation or a non-parametric alternative. For proportions far from 0.5, Fisher's exact test may be more appropriate.

§6After data collection: what to do next

Once you have collected your data, you will need to run the statistical test. Use these tools to compute t-statistics, p-values, and confidence intervals:

  • t-test Calculator — Enter your raw data or group means, SDs, and n values to test whether two groups differ significantly.
  • ANOVA Calculator — If you are comparing three or more groups, use this tool instead.

Then cross-link back to this page if you run a follow-up study and need to re-plan your sample size.

§7FAQ

Should I run a pilot study before calculating my real sample size?

Yes, a pilot study is often valuable. Run a small pilot (10–30 participants, depending on your resources) to estimate the actual effect size and standard deviation in your population. These estimates are much better than guessing from the literature or conventions. However, pilot estimates are noisy due to small sample size. Use them as a starting point, but be conservative—apply a shrinkage factor (e.g., assume your pilot Δ is only 50–75% of the true effect) when planning your full study. This guards against the winner's curse: pilot estimates tend to overstate true effects.

How many participants do I need for a pilot study?

A rough rule of thumb is 10–30 participants per group, depending on how much measurement error you expect. The goal is not statistical significance; it is to get a stable estimate of the standard deviation (σ) and a rough idea of the treatment effect (Δ). Once you have σ from your pilot, you can use this calculator to plan your full study. If your pilot is too small, your estimate of σ will be unstable; if it is too large, you may not have resources for a much larger full study. Consult your advisor or a statistician for your specific field.

My pilot study showed a big effect. Why doesn't my full sample size come out smaller?

This is the winner's curse. Pilot estimates of effect size tend to overstate the true effect because of random sampling variability and researcher degrees of freedom. If your pilot was small (e.g., n = 20 per group), the observed effect Δ was lucky—the true effect is probably smaller. When planning your full study, be conservative: either (a) use only 50–75% of the pilot's observed Δ, or (b) conduct a more formal meta-analysis of your pilot plus any published similar studies. This protects you from underpowered studies.

What if I can't find effect size estimates in the literature for my specific question?

Run a pilot study first. Even a small one (10–15 per group) will give you a local estimate of effect size and variability tailored to your population and measurement methods. Alternatively, use a medium effect size (Cohen's d ≈ 0.5) as a conservative default in behavioral science, but be explicit that you are guessing and be prepared to justify this in your methods section. When in doubt, plan for a smaller effect than you hope for—this leads to adequate power.

Should I use a one-tailed or two-tailed test?

Use a two-tailed test unless you have a very strong, pre-registered theoretical reason to predict the direction of the effect. Two-tailed tests are more conservative and are the default in most fields. One-tailed tests require stronger justification and offer a smaller critical value (z), which looks tempting but often reflects researcher wishful thinking rather than prior evidence. Stick with two-tailed unless your advisor and field consensus clearly support a one-tailed approach.

Does this calculator assume equal sample sizes in both groups?

Yes. The formulas shown assume you allocate the same number of participants to the treatment and control groups. If you must use unequal group sizes (e.g., control is cheap but treatment is expensive), you can modify the allocation, but you will need slightly more total participants. Consult a statistician for the adjustment formula if your study requires unequal groups.

§8Sources

  • OpenStax Introductory Statistics, 2e — Open-access textbook covering power analysis, hypothesis testing, and experimental design.
  • Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Lawrence Erlbaum. — Foundational work defining power, effect size conventions, and the 0.80 threshold.
  • Kraemer, H. C., & Thiemann, S. (1987). How Many Subjects? Statistical Power Analysis in Research. Sage. — Practical guide to power analysis in behavioral research.
  • UCLA OARC: Sample Size — University of California tutorial on sample size calculation for various designs.
  • Schoenfeld, D. A. (1980). Statistical considerations for pilot studies. International Journal of Technology Assessment in Health Care, 15(3), 613–619. — On pilot study design and effect size estimation.
  • Sample Size Calculator (base page) — For survey and other designs.
  • Sample Size Calculator for Surveys — Variant for survey proportions.