Repetitive sequences are a major component of eukaryotic genomes and play important roles in genome organization, chromosome stability and species evolution. Yet highly repetitive regions have long remained among the least understood parts of the genome, because they are difficult to assemble and resolve accurately. Previous studies of repetitive DNA have focused mainly on relatively tractable classes, such as transposable elements. By contrast, many non-transposable repeats, especially satellite repeats, remain poorly characterized in terms of their sequence composition, structural diversity and evolutionary history.
On August 6, 2026, researchers from the laboratories of Dr. Mao at Shanghai Jiao Tong University and Dr. Sun at the Chinese Academy of Sciences Center for Excellence in Brain Science and Intelligence Technology (now at the Sun Yat-sen University Hong Kong Institute of Advanced Studies and Zhongshan School of Medicine), published a study in Cell. The article is titled "Complex subtelomeric architectures in a complete rhesus macaque reference genome". By developing and optimizing assembly strategies for highly repetitive DNA, the team identified and systematically resolved satellite repeats, exemplified by SATR arrays, and built the first primate reference genome that approaches near-perfect accuracy. The study revealed that SATR exemplifies a previously unresolved class of satellite DNA and forms complex arrays that are highly enriched in macaque subtelomeres. The authors also identified gene copies within these repetitive regions that show evidence of transcription. Using the new genome, they substantially improved the accuracy and resolution of population-genetic and single-cell multi-omics analyses. Overall, this study provides a robust foundation for studying primate genome evolution and advancing research in evolutionary medicine.
In the same issues, Cell and Cell Genomics published eight related research articles as a collection on complete genome assembly, repeat resolution and the biological significance of repetitive DNA, alongside related commentaries. Together, these studies show that previously intractable repetitive sequences are becoming a valuable window into genome evolution, species-specific biology and biomedical applications.

Figure 1. The first page of the published article.
Sequencing biases in long-read platforms
Achieving complete assemblies of highly repetitive regions, including subtelomeres, remains one of the most persistent technical challenges in genomics. Long-read sequencing platforms can span large repeat structures and have therefore become essential for complete genome assembly. However, even when read lengths are sufficient to cover complex regions, different sequencing platforms can still be affected by the composition and architecture of these repeats. Such biases may be modest in ordinary genomic regions but can accumulate in highly repetitive DNA, becoming a major barrier to moving from gap-free assemblies to high-accuracy genomes.
In this study, the researchers summarized four common classes of long-read sequencing bias in complex repeat regions: sequencing-depth dropout, homopolymer errors, strand-orientation bias and sequence-context errors. These biases can lead to sequence collapse and misassembly across large repetitive regions. To resolve these problems, the researchers constructed sparse de Bruijn graphs from ONT data and used anchor placement, graph pruning and path disentanglement to recover the correct assembly paths. They then combined locally anchored k-mer graphs with short-read recruitment to polish the resolved regions at single-base resolution.

Figure 2. Sequencing biases in long-read platforms and newly resolved sequences.
Using this strategy, the team generated T2T-MMU8v2.0, a complete and near-perfect reference genome for the rhesus macaque. The assembly showed uniform sequencing coverage, reached a genome continuity index of 100 and achieved a Merqury-estimated QV of 100 at k = 21. In T2T-MMU8v2.0, the team identified 268 previously unannotated repeat families, providing a high-quality reference for studying the structure, potential function and evolutionary divergence of complex primate genomic regions.
Macaque subtelomeric architectures and epigenetic features
Subtelomeres are transition regions adjacent to telomeres, connecting chromosome-specific sequences with terminal repeats. They often contain interleaved segmental duplications and satellite sequences, and their structures can be reshaped by duplication, translocation and ectopic recombination. As a result, subtelomeres frequently show high copy-number and structural diversity among individuals and species. They may provide sequence material for gene-family expansion and divergence, while also increasing the risk of abnormal rearrangements at chromosome ends.
The study found that SATR exemplifies a previously unresolved class of satellite DNA in the macaque genome. These sequences form long tandem arrays composed of short repeat units, including the major families SATR1, SATR1v and SATR2. The team resolved approximately 8 Mbp of SATR satellite arrays and found that, although SATR sequences also occur within internal chromosomal regions, they are enriched about 99-fold in p-arm subtelomeres.
Importantly, macaque subtelomeres are not random accumulations of satellite DNA. Instead, SATR arrays form core structures interspersed with segmental-duplication spacers and composite repeats in a recognizable order. Based on the combinations of SATR families and the organization of their flanking sequences, the team classified four SATR architectural types. Their distribution across subtelomeric and internal chromosomal loci indicates that SATR-enriched regions have a recognizable and relatively ordered architecture.
At the sequence level, subtelomeric segmental-duplication spacers showed an average identity of 97.88%, higher than that of comparable internal chromosomal regions. This pattern suggests recent expansion or the influence of recombination and gene conversion among chromosome ends. At the epigenetic level, the cores of different SATR architectures were generally highly methylated, while their GC content varied with satellite sequence composition. The team further validated SATR1 signals in the subtelomeric regions of 17 chromosomes by fluorescence in situ hybridization, supporting both the assembly and the chromosomal distribution of SATR arrays.
Compared with subtelomeric satellite structures of African great apes, macaque SATR architectures differ markedly in sequence composition, structural organization and methylation pattern. These differences reveal lineage-specific divergence in primate subtelomeric architecture.

Figure 3. Structural features of macaque subtelomeric sequences.
Transcriptionally active gene copies in highly repetitive regions
Previous reference genomes often failed to resolve these regions because of sequence gaps and misassembled repeats. In T2T-MMU8v2.0, the team recovered 225 additional gene copies, including copies related to ZNF669, SH3TC1, PFKP and PITRM1. Comparative analysis showed that these newly resolved copies were significantly enriched in SATR-associated regions. Notably, ZNF669L, a truncated copy of ZNF669, was consistently located at the boundaries of subtelomeric SATR architectures, with at least one ZNF669L copy detected at the boundary of each of the 17 architectures analyzed. Full-length transcriptome and epigenetic data supported the presence of four ZNF669L copies with full-length transcripts and active chromatin marks.
Across SATR-associated regions, the researchers identified 58 gene copies with evidence of transcriptional activity. Some transcripts were detectable in multiple tissues, including prefrontal cortex, spleen and ovary. These findings provide candidate loci and starting points for further research on the functional potential and evolutionary fate of gene copies embedded within repeat arrays.

Figure 4. Gene-family expansion and transcriptional activity within repeat arrays.
In summary, this study developed optimized assembly and correction strategies to overcome long-read sequencing biases and improve the resolution of highly repetitive genomic regions. The resulting T2T-MMU8v2.0 genome is a complete, highly accurate reference genome for the rhesus macaque. Using this resource, the team systematically resolved the complex organization and epigenetic features of macaque SATR satellite arrays. The analysis covered subtelomeric and internal chromosomal regions, bringing formerly inaccessible genomic dark regions into clearer focus at both the sequence and structural levels.
A related study from the same research groups further showed that the fusion of human chromosome 2 involved the replacement of segmental duplications and subtelomeric repeats. It also found that incomplete lineage sorting of related segmental duplications in African great apes may have promoted the formation of the fusion site. Together, these findings suggest that dynamic changes in subtelomeric repeats may be closely linked to the structural evolution of primate chromosomes. The precise mechanisms will require comparative complete genomes from more species and additional functional studies.
Yafei Mao of Shanghai Jiao Tong University and Qiang Sun of the Chinese Academy of Sciences Center for Excellence in Brain Science and Intelligence Technology (now at the Sun Yat-sen University Hong Kong Institute of Advanced Studies and Zhongshan School of Medicine) are co-corresponding authors of this study. Shilong Zhang of Shanghai Jiao Tong University and Ning Xu of the Chinese Academy of Sciences Center for Excellence in Brain Science and Intelligence Technology are co-first authors. The study was conducted in collaboration with multiple primate consortia in China and abroad, with support provided through data sharing among the participating consortia.
The Mao laboratory at Shanghai Jiao Tong University focuses on primate evolutionary medicine. The group combines evolutionary biology, computational biology, neurobiology and large-scale functional screens to investigate the genetic mechanisms underlying primate-specific adaptive traits and the genetic basis of human disease-risk loci. The laboratory is recruiting graduate students, postdoctoral fellows and research assistants with backgrounds in evolutionary biology, bioinformatics, cell biology, neuroscience, computational science and related fields. More information is available at https://www.yafmao.org/.
Original article
https://doi.org/10.1016/j.cell.2026.02.018


