The human genome’s most variable and clinically important regions (centromeres, telomeres, and acrocentric short arms) have been the hardest to study at scale. Thrilled to share KaryoScope, our new preprint that brings them within reach.
Built on k-mer matching (short, fixed-length DNA fragments), KaryoScope annotates a complete diploid human genome at base-pair resolution in ~2 minutes, across repeats, satellite families, genes, and chromosome-end structure. That is ~300× faster than RepeatMasker.

Robertsonian translocations are complex chromosomal rearrangements of acrocentric chromosomes. Adam Phillippy, Erik Garrison and Jennifer Gerton showed SST1 is the fusion substrate; KaryoScope confirms this from k-mers alone.

FSHD1, a muscular dystrophy, is caused by structural changes in D4Z4, a complex subtelomeric repeat array on chromosomes 4q and 10q. Across hundreds of Human Pangenome Reference Consortium haplotypes, KaryoScope first catalogs D4Z4 diversity, including configurations previously not described.

Karen Miga and Glennis Logsdon have charted centromeres. KaryoScope adds pangenome-scale variation: a chr9 megabase inversion and chr3/chr5 repeat losses, FISH-validated.

KaryoScope works on any sequence input, beyond diploid assemblies: long reads, short reads, Hi-C, RNA-seq, metagenomics. It detects SVs from individual long reads (manuscripts forthcoming) and dissects cancer genome assemblies.

KaryoScope analyzes sequence data faster than instruments from PacBio, Oxford Nanopore, Illumina and Element Biosciences (and others) produce it, on laptop-grade hardware. Real-time, on-instrument sequence annotation is within reach.

Annotation is the pangenome era’s bottleneck, and KaryoScope is our step toward dissolving it: a framework that any annotation source can plug into. Tremendous thanks to Rhyker for leading this work, and to all co-authors.
Read more: the paper · the preprint on bioRxiv · the code on GitHub