A team led by Shenghan Gao and senior author Glennis A. Logsdon (University of Pennsylvania) published A global view of human centromere variation and evolution in Nature on 29 July 2026. The study assembled and characterised 2,110 centromeres from individuals representing 5 continental and 28 population groups, identifying 226 distinct centromere haplotypes and 1,870 alpha-satellite higher-order-repeat variants. It is the largest population-scale look yet at the one part of the human genome that conventional sequencing and variant-calling pipelines have never actually resolved.
Why this is a tooling story, not just a biology story
Centromeres are built from megabase-scale tandem repeats — arrays of near-identical alpha-satellite sequence that standard short-read aligners and variant callers cannot uniquely place. Long-read, telomere-to-telomere (T2T) assembly made it possible to resolve individual centromere sequences at all, but resolving thousands of them across a diverse population required going further: the authors report building bioinformatic tools specifically tailored to centromere sequence, because general-purpose genomic tooling does not work reliably in these regions. That is the detail research-data teams should register: any pipeline validated against a reference that treats centromeres as an unresolved gap has been silently blind to this variation, not because the data didn’t exist, but because nothing could read it.
How the study defines a functioning centromere
Rather than inferring centromere location from sequence alone, the authors anchored kinetochore position empirically, using CENP-A chromatin profiling — CENP-A is the histone H3 variant that marks the nucleosomes where the kinetochore actually assembles for chromosome segregation. Combining that functional readout with long-read methylation data, multigenerational inheritance patterns, and archaic (ancient-hominin) sequence comparisons, the team modelled centromere evolution as an ongoing “arms race” between centromeric DNA sequence and the proteins that bind it — frequent mutation at the kinetochore site itself, driving rapid turnover in both the genetic and epigenetic landscape of the region. The paper also reports that most centromeres carry a single kinetochore site, but around 6% show two active sites (di-kinetochores) and fewer than 1% show three.
What changes for reference-genome pipelines
The authors compared their assemblies against the 5,747 centromeres already assembled by the Human Pangenome Reference Consortium (HPRC) and report up to a 20-fold variation in mutation rate across different centromeres — evidence that centromere evolutionary rate is itself a variable worth tracking, not a constant. For any group whose reference-genome, structural-variant, or population-genomics pipeline was built on an assembly that masked or collapsed satellite regions, this dataset is the first population-scale benchmark against which those pipelines can be checked — and, in many cases, rebuilt. The practical consequence is less “here is new biology” and more “here is a reference resource your existing tools were never validated against.”
Where this fits alongside other genomic-data infrastructure work
CASRAI has covered the broader push toward representative, well-governed reference-genome data elsewhere: the launch of AfriGen-D, a hub built specifically to address the underrepresentation of African genomic diversity in reference datasets, gains a concrete data point from this study — population-scale centromere variation is not a marginal correction to the reference genome, it is 226 haplotypes’ worth of structure that most pipelines have never seen. Data managers overseeing genomic datasets should also revisit the NIH Genomic Data Sharing (GDS) Policy and general FAIR Data Principles guidance against this new resource, and see CASRAI’s guide on genome projects and research data management for how reference-assembly consortia structure data-sharing and versioning obligations.
Key figures from the study
- 2,110 centromeres assembled, spanning 5 continental and 28 population groups
- 226 centromere haplotypes identified
- 1,870 alpha-satellite higher-order-repeat variants identified
- 5,747 centromeres from the Human Pangenome Reference Consortium used as a comparison set
- 20-fold variation in mutation rate observed across centromeres
- ~6% of centromeres carry two active kinetochore sites (di-kinetochores); <1% carry three
The paper is: Gao, S. et al. “A global view of human centromere variation and evolution.” Nature, published online 29 July 2026 (DOI: 10.1038/s41586-026-10841-9).







