Comparative analysis of fully sequenced Y chromosomes reveals genomic insights.

CWSF · 2026 Health & Wellness

Thumbnail supplied by the source for Comparative analysis of fully sequenced Y chromosomes reveals genomic insights.

Overview

The human Y chromosome is one of the least understood parts of the human genome, mainly because it is very repetitive, making it difficult to sequence accurately. With the release of new human and primate genomes based on T2T (Telomere-to-Telomere) sequencing technology, clearer resolution of these regions is now possible. This project used supercomputing resources and bioinformatics tools to analyze fully sequenced T2T genomes across multiple species, focusing on repetitive element distribution in the Y chromosome. Results revealed that the second half of the human Y chromosome, which is considered genetically inactive, contains a high concentration of non-coding genes. Further analysis showed a tandem duplication event involving two repetitive elements (AluY and HSATII). 602 out of 604 non-coding genes in this region contain these elements. This pattern is unique to the human Y chromosome, suggesting that repetitive elements may play a larger role than previously recognized in human genetic evolution.

Video

Why?

The genome defines who we are as human beings. However, it is a complex structure with over 3 billion base pairs separated into 22 of autosomes and 1 sex chromosome from each parent. In the past, our DNA was sequenced by chopping it into numerous small parts, obtaining short reads, and then assembling them into the genome with computing tools. This technique worked for the most part, but it ran into issues within repetitive regions of the genome.

Recently, new techniques have emerged, making it much easier to sequence these difficult-to-read regions. By using nanopore sequencing technology, longer reads are possible, thereby resolving challenges with assembling repetitive regions of the genome. The human Y chromosome (the male sex chromosome) in particular contains repetitive regions and therefore represents an opportunity for insights based on improved sequencing approaches.

I decided to look into newly sequenced Y chromosomes in humans and primates. My hypothesis was that trends concerning repetitive elements would emerge in the human Y chromosome. This project's goal was to analyze the human Y chromosome as relatively little is known about its repetitive regions.

My inspiration for this project and its hypothesis came from discussions with my mentor Dr. Ping Liang. He introduced me to genomic informatics and provided me guidance for analyses.

This project provides insights in the field of human genetics, and evolutionary genome biology. Specifically, we shed new light on genomic repetition and how it contributes to the human male sex chromosome.

How?

To carry out my project, I first built a background understanding of the topic. With guidance from Dr. Liang at Brock University and weekly meetings with him, I learned key concepts about DNA and genome analysis. I also read three peer-reviewed scientific articles on the topic. I obtained genomic information from trusted databases like NCBI to make sure my information was reliable.

For the analyses, I used a computational approach. I was able to obtain a user account at SHARCNET (the Canadian high-performance computing network providing supercomputing resources for researchers) for the analysis. My work was done mainly on the NIBI node of SHARCNET, located at the University of Waterloo. This allowed me to download and analyze large DNA datasets which would not be feasible with regular computers. I downloaded raw genomic data (FNA files, which contain DNA sequences) from NCBI, focusing on the most recent “telomere-to-telomere” (T2T) genome assemblies for human and primate species.

I then analyzed the data using computations tools. I scripted most of the commands myself in the Linux shell (Bash), using tools like awk and grep to filter and organize the data. I used RepeatMasker with the Dfam library, which are publicly accessible software tools, to identify repetitive DNA elements. I also wrote and ran commands in Python. I also created some of my own scripts to improve the workflow.

To test my ideas, I compared patterns across different chromosomes and species, making sure to use the same methods each time to control variables. I collected data from multiple chromosomes and species. I visualized my results using Excel and Interactive Genome Viewer (IGV), which helped visualize the results and patterns in the data.

What?

Figure 1 shows coding, non-coding, and pseudo- gene distributions among human and chimpanzee X and Y chromosomes. In human chromosome Y (chrY hereafter) coding and pseudogenes are completely absent from the second half of the chrY, but non-coding genes are abundant. This second half of chrY in humans is genetically inactive (no coding genes) so it was not expected that there would be as many non-coding genes are observed. The plots also demonstrate that this chrY pattern was only observed in human chrY, and not in Chimpanzee. The plots also show that the pattern was absent in the X chromosome.

Next, TEs (transposable elements) were assessed for any correlation with these non-coding genes on the human chrY. TEs make up 50% of the genome, so we hypothesized there may be a relationship. Figure 2 (top) shows a large amount of same age TEs of the SINE subfamily in chrY. The age of a TE can be estimated based on the number of its mutations from consensus. This was unexpected because TEs would be assumed to duplicate gradually, not suddenly. The bottom left and bottom right graphs show that this spike in genetic age was not perceived anywhere else in other species or other chromosomes. Although there are spikes present they are clearly more gradual and not sharp.

In Figure 3, the top panel shows the SINE subfamily contributes to the aformentioned spike. SINE is a main type of TE. AluY, a type of SINE, contributes over 90% to this spike. Normally, AluY contributes only 5% to a chromosome. The bottom panel shows the age distribution of the HSATII satellite, an independent repetitive element different from AluY. Although they are independent from each other, they have a close to identical age spike.

Figure 4 shows a repeating pattern between several repetitive elements in close proximity in the blue box. This is an example of tandem duplication, when a block of DNA is duplicated over and over by genetic error in close proximity, which is observable from the pattern. This explains why there are so many AluY and HSATII of one age and why the two types of TEs they have the same genetic age; because they were not duplicated separately but duplicated over and over again together at the same time.

The bottom half of Figure 4 shows the non-coding gene pattern in chrY against the tandem duplication pattern in the same spot. There is a significant amount of overlap by visual inspection. Out of 604 long non-coding genes, 602 have at least 1 HSATII or AluY inside of them, showing that the tandem duplicated repetitive elements contribute to actual noncoding genes.

So What?

In this project, we found new patterns in fully sequenced Y chromosomes with respect to its non-coding genes and repetitive elements. Until recently, the Y chromosome has been difficult to sequence due to its repetitive nature and these new T2T genomes have better resolution of these compared to previous ones. The results revealed that the second half of the human Y chromosome, which is known to be genetically inactive actually had a high concentration of non-coding genes that correlated with AluY and HSATII repeats. Our analyses suggest they were duplicated together as they share a close to identical age signature. This is important because they caused noncoding genes to appear. In the second half 602 out of 604 non coding genes contain AluY or HSATII -- which suggests that this tandem duplication shaped these genes. Importantly none of these patterns appeared in chimpanzee, our closest relative or in any other human chromosomes.

Our findings improve our understanding of the human Y chromosome by showing that regions once thought to be inactive may actually have meaningful structure and function. Our ongoing studies indicate that these non-coding genes show tissue-specific expression and may play a more important role than previously realized. This suggests that the transposable elements we found may not be just “junk,” but can play an active role in shaping the genome and driving human-specific evolution. More broadly, it highlights how newly completed genomes can uncover hidden features that were previously missed, changing how we interpret non-coding DNA.

What's Next?

My project has found numerous interesting findings in the repeating DNA of the human Y chromosome. Future work aims to determine what the mechanism behind the tandem duplication was. Although we found that tandem duplication happened, we are unsure if there were any specific DNA elements that encouraged it to happen. Additional studies we are currently working on are to determine the tissue-level expression of the non-coding genes that were part of the tandem duplications. Preliminary results suggest these non-coding genes are expressed differentially in human testes from males of different ages.

Thanks

Although I did all the work in this project myself, this project would not be possible without the support and guidance of Dr. Liang. Dr. Liang is a renowned bioinformatics professor at Brock University. He patiently taught me coding and bioinformatics. He was able to guide me in weekly meetings and provide advice on troubleshooting when I encountered unforeseen problems.

I would also like to thank the SHARCNET community. If they did provide me with a volunteer access account to their servers and hardware, I would not have been able to do any of the analyses.

I am thanking my family for being there. Encouraging me, supporting me and giving me advice even when they didn't know what they were talking about.

Lastly, I would like to thank Niagara team (NRSEF) for having a strong community and supporting science in the Niagara region.

References

References

Journal Articles

Nurk, S., Koren, S., Rhie, A., Rautiainen, M., Bzikadze, A. V., Mikheenko, A., & Phillippy, A. M. (2022). The complete sequence of a human genome. Science, 376(6588), 44–53. https://doi.org/10.1126/science.abj6987

Yoo, D., Rhie, A., Hebbar, P., et al. (2025). Complete sequencing of ape genomes. Nature, 641, 401–418. https://doi.org/10.1038/s41586-025-08816-3

Colonna Romano, N., & Fanti, L. (2022). Transposable elements: Major players in shaping genomic and evolutionary patterns. Cells, 11(6), 1048. https://doi.org/10.3390/cells11061048

Wanxiangfu T., Ping L., Comparative Genomics Analysis Reveals High Levels of Differential Retrotransposition among Primates from the Hominidae and the Cercopithecidae Families, Genome Biology Evolution, 11, 11, 2019, 33093325, https://doi.org/10.1093/gbe/evz234

Software and resources

Smit, A. F. A., Hubley, R., & Green, P. (n.d.). RepeatMasker. Retrieved April 2025, from https://www.repeatmasker.org/

Storer, J., Hubley, R., Rosen, J., & Smit, A. (n.d.). Dfam: Repetitive DNA element database. Retrieved April 2025, from https://dfam.org

Alliance Canada. (n.d.). The Digital Research Alliance of Canada. Retrieved April 2025, from https://www.alliancecan.ca

SHARCNET. (n.d.). SHARCNET high performance computing. Retrieved April 2025, from https://www.sharcnet.ca

OpenAI. (2024). ChatGPT [Large language model]. https://chat.openai.com, used for coding help

Images

Why slide

https://doe-humangenomeproject.ornl.gov/chromosome-y/ picture showing Y chromosome (slide 2).

https://www.veritasint.com/blog/en/genes-and-chromosomes-how-do-they-determine-our-life-and-health/ picture demonstrating all chromosomes side by side (slide 3).

How slide

https://datacenternews.ca/story/nokia-hypertec-power-new-nibi-supercomputer-in-canada an image of the nibi supercomputer (slide 2)

https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_009914755.1/ screenshot of one of genomes used (slide 3)

Thanks slide

https://brocku.ca/mathematics-science/biology/directory/ping-liang/ photo of my mentor Dr. Liang (slide 2)

Images (19)

Awards (1)

  • Selected for CWSF 2026

Competition history

  • CWSF 2026 Health & Wellness Qualified through Niagara, ON

Related projects

Closest projects by meaning, across every fair and year in the corpus.

Browse more like this

Source: ProjectBoard / Youth Science Canada

Save projects to your library

Sign in with Google to keep track of projects you find interesting, organized into folders. An account also raises your daily allowance for “Has this been done?”, and lets you create a key for the MCP server with a much higher limit than anonymous use. Browsing stays public.

Continue with Google