Replicability

Is the field running studies that can find the effects it studies? Bergmann, Tsuji, Piccinini, Lewis, Braginsky, Frank, & Cristia (2018) used MetaLab to diagnose replicability across language acquisition research: typical sample sizes, statistical power, method sensitivity, and small-study effects. This page recomputes those analyses live from the current data release for every dataset — the 2018 paper’s snapshot covered 12 meta-analyses; the numbers below update as MetaLab grows.

Loading data…

Power of typical studies, dataset by dataset

For each dataset: the multilevel meta-analytic effect size (Cohen’s d), the median per-study sample size, and the power of that median-sized study to detect that effect (two-sided α = .05, within-subjects normal approximation). The 2018 paper’s headline — a median power of 44% — reflected a field standard of n ≈ 15–20 regardless of effect size.

Do methods differ in the effects they yield?

Effect sizes by experimental method, pooled across datasets (methods with at least 10 effect sizes). The 2018 paper found conditioned head-turn yields d > 1 while several passive-looking methods hover near d ≈ 0.2 — method choice is a first-order design decision.

Small-study effects: sample size vs effect size

If studies with smaller samples report systematically larger effects — because small studies only reach significance (and publication) with big effects — the literature overestimates. Kendall’s τ between n and effect size per dataset (the 2018 paper’s Table 5 found significant negative correlations in 4 of 12 meta-analyses).


Analyses follow Bergmann et al. (2018) (scripts), recomputed in the browser from the current MetaLab release. Power uses the same normal approximation as the power analysis tool; the multilevel effect-size estimates match the visualization page.