Skip to main content
v2026.11,610 entries · CC-BY 4.0
LAC HealthLaboratory & ResearchLab & research supplies.Reagents, consumables, PPE & instruments — documented, fast, chain-of-custody shipping.Shop lac.us lac.us

TalkBank and CHILDES: The Domain Repository for Language Acquisition and Communication Research

TalkBank is a research infrastructure for studying human communication; CHILDES, its flagship child-language-acquisition component, is the field-standard repository for language development data. This guide covers governance, the CHAT transcription format, licensing, and how it compares to CLARIN.

TalkBank is a research infrastructure for studying human communication, with a particular emphasis on spoken language. It is organized as a family of specialized “banks” of shared, transcribed communication data, each covering a different research domain. Its oldest and best-known component is CHILDES (the Child Language Data Exchange System), which serves as TalkBank’s dedicated repository for child language acquisition data. For research administrators and data stewards supporting linguistics, developmental psychology, communication sciences and disorders (CSD), and speech-language pathology researchers, TalkBank/CHILDES is the default disciplinary repository that funder data management plans (DMPs) in these fields typically point to.

What TalkBank and CHILDES Are

TalkBank describes its mission as fostering fundamental research into human communication by maintaining open, standardized corpora of transcribed (and often audio- or video-linked) interaction data. Rather than a single dataset, it is a system of discipline-specific repositories, each built around the same underlying transcription conventions and analysis tools so that data contributed by different research groups can be searched, compared, and reused consistently.

CHILDES is the child-language component of that system: a repository of transcripts, recordings, and associated tools focused on how children acquire language. It predates TalkBank as a standalone project (Brian MacWhinney and Catherine Snow began CHILDES in the 1980s) and today operates as the flagship, most heavily used bank within the broader TalkBank infrastructure.

Governance and Institutional Home

TalkBank and CHILDES are directed by Brian MacWhinney, professor of psychology at Carnegie Mellon University, where the project is based. A TalkBank Advisory Board oversees the broader system. TalkBank also designates an affiliated journal, Language Development Research (LDR), hosted using open-source publishing software through the Carnegie Mellon University Library.

Per TalkBank’s own documentation, CHILDES has received federal funding through the National Institute of Child Health and Human Development (NICHD), including grant HD082736. Research administrators citing TalkBank/CHILDES in a proposal or acknowledgments section should verify current award numbers directly on talkbank.org, since federal grant support for shared infrastructure is typically renewed under new award numbers over time rather than held fixed indefinitely.

How TalkBank Is Organized: CHILDES and the Other Component Banks

Beyond CHILDES, TalkBank groups its component corpora by research area. As documented on talkbank.org, these include:

  • Child language: CHILDES (general child language acquisition), PhonBank (phonological development), FluencyBank (stuttering and fluency), and HomeBank (naturalistic home audio recordings)
  • Conversation and discourse: CABank (conversation analysis), SamtaleBank, and ClassBank (classroom discourse)
  • Multilingualism: BilingBank and SLABank (second language acquisition)
  • Clinical and communication disorders: AphasiaBank, ASDBank (autism spectrum disorder), DementiaBank, RHDBank (right hemisphere damage), TBIBank (traumatic brain injury), MotorSpeechBank, and PsychosisBank

Each bank applies the same shared transcription and analysis conventions described below, which is what allows a researcher to, for example, compare a clinical aphasia transcript against a typically-developing child language sample using the same tooling.

The CHAT Transcription Format and CLAN Software

Data across TalkBank’s component banks is transcribed using CHAT (Codes for the Human Analysis of Transcripts), a standardized transcription manual and file format that specifies how utterances, speaker tiers, morphosyntactic coding, and timing links to audio/video should be represented in a consistent, machine-processable way. Because every contributed corpus follows the same CHAT conventions, transcripts from different research groups, institutions, and even different component banks can be searched and analyzed with a shared toolset rather than requiring bespoke parsing for each dataset.

The companion analysis software is CLAN (Computerized Language Analysis), a free program distributed by TalkBank for querying, coding, and running automated analyses (word counts, mean length of utterance, morphosyntactic profiles, and similar measures) directly against CHAT-formatted transcripts. CLAN and CHAT together are the technical backbone that makes TalkBank function as a genuinely interoperable repository rather than a loose collection of independently formatted files.

Data Sharing Rules, Ethics, and Licensing

Depositing data into TalkBank/CHILDES is governed by the project’s published “Ground Rules,” which set expectations for contributors around consent, de-identification, and appropriate use of shared corpora, alongside a stated expectation that depositors have followed applicable IRB (institutional review board) principles for human-subjects research before contributing. Corpora are made available under a Creative Commons Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license, meaning reuse must credit the original contributors, is restricted to non-commercial purposes, and derivative works must be shared under the same terms.

For a research administrator reviewing a data management plan that names TalkBank or CHILDES as the repository of record, the practical compliance checkpoints are: (1) confirming human-subjects approval and consent language cover the eventual public deposit of transcripts and recordings, (2) confirming any necessary de-identification has been applied before deposit, since CHAT transcripts and linked audio/video can otherwise be individually identifying, and (3) confirming the CC BY-NC-SA license is compatible with the funder’s own data-sharing terms (some funder or publisher mandates require fully open, unrestricted licensing, which CC BY-NC-SA’s non-commercial and share-alike conditions do not satisfy).

TalkBank/CHILDES vs. CLARIN: How They Differ

TalkBank/CHILDES and CLARIN are sometimes conflated because both serve language researchers, but they operate at different scales and with different scope:

  • Scope: TalkBank/CHILDES is a specific, corpus-based domain repository organized around communication and language-acquisition research questions (how children learn language, how conversation unfolds, how communication breaks down in clinical populations). CLARIN is a broad, pan-European research infrastructure spanning many kinds of language resources and tools (corpora, lexicons, annotation tools, natural-language-processing services) across dozens of languages and member countries, coordinated through a network of national CLARIN centers.
  • Governance: TalkBank/CHILDES is directed from a single institutional home, Carnegie Mellon University, under Brian MacWhinney. CLARIN operates as a distributed European Research Infrastructure Consortium (ERIC) with member-country nodes and shared governance.
  • Format standardization: TalkBank enforces a single shared transcription format (CHAT) and toolset (CLAN) across all its component banks. CLARIN federates many different repositories and formats behind a common discovery and access layer rather than requiring one uniform transcription standard.
  • Typical use case: A researcher studying child language development, clinical communication disorders, or conversational structure will most often deposit in or query TalkBank/CHILDES directly. A researcher needing broader European language-resource discovery, or tools spanning many resource types and languages, is more likely to work through CLARIN’s federated infrastructure.

The two are not competitors so much as operating at different levels: TalkBank/CHILDES is a deep, standardized domain repository; CLARIN is a broad, federated infrastructure layer that can itself reference domain repositories like TalkBank as one of many resource providers.

Using TalkBank/CHILDES in Research Data Management

For linguistics, psychology, and communication sciences researchers, naming TalkBank or CHILDES as the target repository in a data management plan is a defensible, field-standard choice precisely because the repository has an established governance structure, a documented data format, and published deposit rules — the same criteria funders and repository-evaluation frameworks such as CoreTrustSeal look for when assessing whether a named repository is a credible long-term home for shared data. Administrators should still confirm, on a per-project basis, that:

  • The specific component bank (CHILDES, AphasiaBank, etc.) matches the study population and research question
  • Consent and IRB documentation explicitly anticipate public data sharing under CC BY-NC-SA terms
  • Any funder mandate for fully open licensing is reconciled against TalkBank’s non-commercial license before the DMP is finalized
  • Metadata and de-identification requirements specific to child or clinical populations are met before deposit

Frequently Asked Questions

Is CHILDES part of TalkBank, or a separate system?

CHILDES is a component of TalkBank — specifically, the child-language-acquisition bank within the broader TalkBank system. CHILDES existed first and is still commonly referred to by its own name because of its long history and heavy use, but administratively and technically it now operates within the same CHAT/CLAN infrastructure as TalkBank’s other component banks.

Who can deposit data in TalkBank/CHILDES?

Researchers who have collected communication or language data under appropriate human-subjects approval can contribute a corpus, subject to TalkBank’s published Ground Rules covering consent, de-identification, and IRB compliance. Contributors should review talkbank.org’s current contribution guidance directly, since exact submission mechanics can change.

What file format does TalkBank use?

Transcripts follow the CHAT (Codes for the Human Analysis of Transcripts) format, a standardized transcription convention analyzed using the free CLAN (Computerized Language Analysis) software that TalkBank distributes.

Can TalkBank/CHILDES data be used commercially?

Generally no. TalkBank corpora are distributed under a Creative Commons Attribution-NonCommercial-ShareAlike license, which restricts reuse to non-commercial purposes and requires derivative works to carry the same license terms.

How does TalkBank/CHILDES relate to general-purpose data repositories?

Unlike a generalist repository, TalkBank/CHILDES is a domain-specific system: it only accepts communication/language data, and it enforces a shared transcription format and analysis toolset across everything it holds. This makes it a stronger fit for a DMP than a generalist repository whenever the underlying data is transcribed spoken or signed communication, since reviewers and funders in linguistics and communication sciences recognize it as the field’s standard repository.

This guide summarizes governance, format, and licensing information published by TalkBank at talkbank.org as of 2026. Grant numbers, contribution procedures, and specific corpus holdings should be verified directly against the current talkbank.org documentation before citing them in a funding proposal or data management plan.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →