By capturing both inherited chromosome sets and reaching genomic regions missed by conventional benchmarks, T2T-HG002 could help shift genomics from reference-based variant calling toward complete personalized genomes.

Study: A complete diploid human genome benchmark for personalized genomics
In a recent study published in the journal Cell, researchers developed a near-perfect, complete diploid human genome benchmark for personalized genomics. Including autosomal and sex chromosome sequences, this genome benchmarking framework addresses limitations of conventional reference-based variant benchmarking.
While conventional approaches typically rely on an existing genome sequence for variant identification, this new strategy provides a telomere-to-telomere (T2T) benchmark that helps distinguish sequencing and assembly errors from genuine genetic variation while avoiding biases caused by errors, gaps, and structural differences in external reference genomes. This advancement could help researchers analyze previously inaccessible genomic regions and ultimately support more comprehensive interpretation of disease-associated genetic variation.
Conventional resequencing approaches often struggle to detect structural changes and highly variable regions of the genome linked to human genetic diseases. Short-read sequencing often struggles to distinguish highly similar duplicated regions or assign genetic variants to their respective parental copies, while structural differences between genomes can further affect accuracy. Improved strategies are needed to enhance genome sequencing precision and accelerate personalized genomic medicine.
About the study
In the present study, researchers developed a near-perfect, nearly complete T2T genome benchmark for evaluating the widely used human diploid HG002 genome under the Q100 Project. Most sequencing data were derived from cultured HG002 lymphoblastoid cells. Unlike traditional variant benchmarks that compare variants against a reference genome, this approach relies on a nearly complete, highly accurate diploid genome sequence as the benchmark standard, separating sequencing and assembly errors from genuine genetic differences between the sample and an external reference genome. This genome benchmarking approach avoids reference biases and directly evaluates genome inference and haplotype accuracy.
The researchers mapped genes and repetitive deoxyribonucleic acid (DNA) regions across both parental copies of the T2T-HG002 assembly, allowing direct comparison between the maternal and paternal haplotypes. They also established approaches to assess sequencing accuracy, determine allele-resolved genetic differences, and benchmark genome assemblies against the diploid benchmark sequence. This enabled evaluation of base-level accuracy, haplotype resolution, and large-scale genomic changes while addressing limitations of earlier variant benchmarking methods.
The HG002 genome was assembled using long-read sequencing data from PacBio HiFi and Oxford Nanopore Technologies (ONT), supplemented with parental Illumina sequencing for phasing and Strand-seq and Hi-C data for scaffolding. Fluorescence in situ hybridization (FISH) validated ribosomal DNA (rDNA) scaffolding and helped estimate rDNA array sizes.
The team improved the T2T-HG002 assembly through successive rounds of error correction and validation against independent sequencing datasets. The resulting v1.1 version contained 38,037 small fixes and 14.18 million base pairs of patched consensus regions. Assembly accuracy and phasing were assessed using Merqury, while immunoglobulin loci were manually validated. Researchers also developed GQC software to evaluate assemblies against the benchmark and compared multiple HG002 assemblies generated using different sequencing technologies over five years.
Results
The benchmark was free of detectable errors across 99.4% of the diploid genome. It incorporated an additional 701.4 Mb (11.7%) of high-confidence autosomal DNA and 216.8 Mb of sex chromosome sequences missing from the earlier v4.2.1 benchmark. Unlike conventional human reference sequences, which generally represent a single haplotype, T2T-HG002 contains both maternal and paternal haplotypes, providing a diploid, two-haplotype representation. The GQC tool analyzes these genome assemblies by distinguishing between the two parental copies, enabling a more accurate assessment of complex genomic regions.
Advances in T2T sequencing and genome reconstruction have improved the characterization of genomic regions that were previously difficult to resolve, including duplicated sequences, expanded gene clusters, satellite DNA regions, and repetitive DNA structures. A genome-wide comparison revealed that de novo genome assemblies recovered two to seven percent more sequence than genomes reconstructed from variant calls and achieved approximately tenfold greater genome-wide accuracy.
Following refinement, the assembly’s quality score increased from Q63.1 in version 0.7 to Q68.9 in version 1.1. Within approximately 2.66 Gb reliably covered by both benchmarks, discrepancies with the earlier Genome in a Bottle (GIAB) benchmark fell from 5,972 to 219 variants, with the authors attributing nearly all remaining differences to errors in the older benchmark.
The HG002 annotation identified 13 maternal-only and 12 paternal-only autosomal genes, including copy-number variable genes such as dual specificity phosphatase 22 (DUSP22), complement factor H-related 1 (CFHR1), complement factor H-related 3 (CFHR3), glutathione S-transferase theta 1 (GSTT1), and glutathione S-transferase mu 1 (GSTM1). The researchers also integrated long-read functional genomics datasets, including methylation, chromatin accessibility, and transcriptome profiles, providing genome-wide coverage across both haplotypes.
Benchmarking of previous HG002 assemblies showed progressive improvements in completeness and reductions in substitution, indel, and phase-switch errors, with T2T-HG002 providing the ground truth for comparison. Variant-constructed genomes represented a smaller fraction of the HG002 genome than de novo assemblies, demonstrating that complete-genome benchmarking can reduce reference bias in evaluation and reveal opportunities to improve personalized genome reconstruction.
Conclusion
The study presents a telomere-to-telomere human diploid HG002 genome benchmark that is free of detectable errors across 99.4% of the genome, providing greater completeness and precision than previous benchmarks. Although all 46 chromosomes achieve T2T continuity, the interiors of some rDNA arrays remain unfinished. As a National Institute of Standards and Technology (NIST) reference material with established variant benchmarks, HG002 offers an ideal resource for developing personalized genomics approaches based on complete diploid genomes. The benchmark enables evaluation of genome assemblies, phased variant calls, sequencing reads, and pangenome-level haplotypes across current and emerging technologies.
In future studies, researchers should establish standardized genome benchmarking metrics and identify clinically relevant variant sites in diploid genomes without reference bias. Variant calling remains the primary approach in clinical genomics, and methods for translating whole-genome benchmarking performance into clinical interpretation are not yet standardized. Advances in sequencing may help resolve complex regions such as rDNA arrays, while applying GQC across diverse genomes could enable more comprehensive whole-genome analysis. The authors also note that HG002 cells contain low levels of somatic variation and that some repetitive genomic regions remain difficult to validate.