Document Type
Article
Publication Date
2026
DOI
10.1016/j.cell.2026.06.016
Publication Title
Cell
Volume
189
Issue
16
Pages
4857-4875.e31
Abstract
Human genome sequencing typically relies on mapping reads to a reference genome to call variants, but this approach introduces technical biases, excluding duplicated and structurally polymorphic regions of the genome. To overcome this, we present a telomere-to-telomere genome benchmark with near-perfect accuracy across 99.4% of the diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), which were absent from prior benchmarks. We annotated genes and repeats on both haplotypes, including 19,956 protein-coding genes on the maternal haplotype and 19,190 on the paternal haplotype, and developed new methods to measure the accuracy of reads, phased variant call sets, and assemblies against a diploid reference. Genome-wide analyses show that de novo assembly resolves 2%-7% more sequence and outperforms variant calling accuracy by an order of magnitude, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.
Rights
© 2026 The Authors.
This is an open access article under the Creative Commons Attribution 4.0 International (CC-BY 4.0) License.
ORCID
0000-0002-7548-4375 (Adam), 0000-0003-4626-4733 (Riethman)
Original Publication Citation
Hansen, N. F., Dwarshuis, N., Ji, H. J., Rhie, A., Loucks, H., Logsdon, G. A., Vollger, M. R., Storer, J. M., Kim, J., Adam, E., Altemose, N., Antipov, D., Asri, M., Barreira, S., Bohaczuk, S. C., Bzikadze, A. V., Carioscia, S. A., Carroll, A., Chao, K. H.,…Phillippy, A. M. (2026). A complete diploid human genome benchmark for personalized genomics. Cell, 189(16), 4857-4875.e4831. https://doi.org/10.1016/j.cell.2026.06.016
Repository Citation
Hansen, Nancy F.; Dwarshuis, Nathan; Ji, Hyun Joo; Rhie, Arang; Loucks, Hailey; Logsdon, Glennis A.; Vallger, Mitchell R.; Storer, Jessica M.; Kim, Juhyun; Adam, Eleni; Alternose, Nicolas; Antipov, Dmitry; Asri, Mobin; Barreira, Sofia; Bohaczuk, Stephanie C.; Bzikadze, Andrey V.; Carioscia, Sara A.; Carroll, Andrew; Chao, Kuan-Hao; Chu, Yanan; Das, Arun; Ebert, Peter; English, Adam; Fleharty, Mark; Fleming, Laura E.; Formenti, Giulio; Guarracino, Andrea; Hartley, Gabrielle A.; Jenike, Katharine; Kalleberg, Jenna; Kang, Yu; King, Robert; Lipovac, Josipa; Mastoras, Mira; Mitchell, Matthew W.; Negi, Shloka; Olson, Nathan D.; Oshima, Keisuke K.; Paulin, Luis F.; Pickett, Brandon D.; Porubsky, David; Ranchalis, Jane; Ranjan, Desh; Rautiainen, Mikko; Riethman, Harold; Schnabel, Robert D.; Sedlazeck, Fritz J.; Shafin, Kishwar; Sikic, Mile; Solar, Steven J.; Sweeten, Alexander P.; Timp, Winston; Wagner, Justin; Yoo, DongAhn; Zhou, Ying; Garrison, Erik; Eichler, Evan E.; Schatz, Michaeel C.; Stergachis, Andrew B.; O'Neill, Rachel J.; Miga, Karen H.; Salzberg, Steven L.; Koren, Sergey; Zook, Justin M.; and Phillippy, Adam M., "A Complete Diploid Human Genome Benchmark for Personalized Genomics" (2026). School of Medical Diagnostics & Translational Sciences Publications. 51.
https://digitalcommons.odu.edu/medicaldiagnostics_fac_pubs/51
All Supplementary files included with this article
1-s2.0-S0092867426007038-mmc1.pdf (2598 kB)
Document S1. Figures S1-S13 and Tables S3, S18, and S19
1-s2.0-S0092867426007038-mmc2.xlsx (26 kB)
Table S1: Sequencing data used for assembly, polishing, and validation, related to STAR Methods
1-s2.0-S0092867426007038-mmc3.xlsx (12 kB)
Table S2. Variant calls used in polishing, related to STAR Methods
1-s2.0-S0092867426007038-mmc4.xlsx (26 kB)
Table S4. K-mer-based evaluation, Related to Figure 3
1-s2.0-S0092867426007038-mmc5.xlsx (16 kB)
Table S5. K-mer-based evaluation of HG002v1.1, related to STAR Methods
1-s2.0-S0092867426007038-mmc6.xlsx (23 kB)
Table S6. Counts of suspicious regions/bases in T2T-HG002v1.1*, related to STAR Methods
1-s2.0-S0092867426007038-mmc7.xlsx (131 kB)
Table S7. Issues and excluded regions in T2T-HG002v1.1, related to STAR Methods
1-s2.0-S0092867426007038-mmc8.xlsx (60 kB)
Table S8. Discrepancies between T2T-HG002v1.1 and the GIAB v4.2.1 benchmark, related to STAR Methods
1-s2.0-S0092867426007038-mmc9.xlsx (4617 kB)
Table S9. List of distinct HUGO Gene Symbols Found in the T2T-CHM13 annotation, related to STAR Methods
1-s2.0-S0092867426007038-mmc10.xlsx (2490 kB)
Table S10. Broken genes in T2T-HG002v1.1, related to STAR Methods
1-s2.0-S0092867426007038-mmc11.xlsx (1955 kB)
Table S11: Number of gene copies for each HUGO gene annotated on the T2T-CHM13 Reference Genome and Both Haplotypes (MAT and PAT), related to STAR Methods
1-s2.0-S0092867426007038-mmc12.xlsx (75 kB)
Table S12. Haplotype-specific gene instances in T2T-HG002v1.1, related to STAR Methods
1-s2.0-S0092867426007038-mmc13.xlsx (12 kB)
Table S13. Assemblies and sequencing reads used for benchmarking and validation, related to Figures 3 and 4.
1-s2.0-S0092867426007038-mmc14.xlsx (46 kB)
Table S14. GQC benchmarking results for assemblies, related to Figure 3
1-s2.0-S0092867426007038-mmc15.xlsx (31 kB)
Table S15. GQC benchmarking statistics for read datasets
1-s2.0-S0092867426007038-mmc16.xlsx (9 kB)
Table S16. Variant classes and descriptions, related to STAR Methods
1-s2.0-S0092867426007038-mmc17.xlsx (16 kB)
Table S17. GIABv4.2.1-Covered Regions of HG002v1.1, related to Figure 2
1-s2.0-S0092867426007038-mmc18.pdf (6767 kB)
Document S2. Article plus supplemental information