childes-db

A flexible and reproducible interface to CHILDES

The childes-db project is an open database storing child language datasets from CHILDES in a well-documented, easily accessible, tabular format. It also provides a versioning system for corpora and tools to facilitate reproducible research with child language corpora. Researchers can interface with CHILDES through the interactive visualizations on this site, the childesr R package, the childespy Python package, or directly through SQL.

For a complete overview along with examples, refer to our paper in Behavior Research Methods:

*Sanchez, A., *Meylan, S. C., Braginsky, M., MacDonald, K. E., Yurovsky, D., & Frank, M. C. (2019). childes-db: A flexible and reproducible interface to the Child Language Data Exchange System. Behavior Research Methods, 51(4), 1928–1941. (* indicates co-first authorship)

New in 2026.1

  • More data: 24.2 million utterances (+24%) and 89 million tokens, including ~73 corpora new to childes-db.
  • New annotation tiers: Universal Dependencies morphology (with a morpheme-level table) and dependency parses, Japanese romanization, speech acts, and phonology.
  • Redivis-hosted, versioned data: every release browsable, queryable, and downloadable on Redivis.
  • Corrected child identities: target-child identities corrected across corpora (Day corrections).