Efficient identification of Y chromosome sequences in the human and Drosophila genomes

Genome Res. 2013 Nov;23(11):1894-907. doi: 10.1101/gr.156034.113. Epub 2013 Aug 6.

Abstract

Notwithstanding their biological importance, Y chromosomes remain poorly known in most species. A major obstacle to their study is the identification of Y chromosome sequences; due to its high content of repetitive DNA, in most genome projects, the Y chromosome sequence is fragmented into a large number of small, unmapped scaffolds. Identification of Y-linked genes among these fragments has yielded important insights about the origin and evolution of Y chromosomes, but the process is labor intensive, restricting studies to a small number of species. Apart from these fragmentary assemblies, in a few mammalian species, the euchromatic sequence of the Y is essentially complete, owing to painstaking BAC mapping and sequencing. Here we use female short-read sequencing and k-mer comparison to identify Y-linked sequences in two very different genomes, Drosophila virilis and human. Using this method, essentially all D. virilis scaffolds were unambiguously classified as Y-linked or not Y-linked. We found 800 new scaffolds (totaling 8.5 Mbp), and four new genes in the Y chromosome of D. virilis, including JYalpha, a gene involved in hybrid male sterility. Our results also strongly support the preponderance of gene gains over gene losses in the evolution of the Drosophila Y. In the intensively studied human genome, used here as a positive control, we recovered all previously known genes or gene families, plus a small amount (283 kb) of new, unfinished sequence. Hence, this method works in large and complex genomes and can be applied to any species with sex chromosomes.

Publication types

  • Research Support, N.I.H., Extramural
  • Research Support, Non-U.S. Gov't

MeSH terms

  • Animals
  • Databases, Genetic
  • Drosophila / genetics*
  • Euchromatin / genetics
  • Evolution, Molecular
  • Female
  • Genes, Y-Linked*
  • Genome, Human*
  • Genome, Insect*
  • Genomics / methods*
  • Humans
  • Male
  • Molecular Sequence Data
  • Phylogeny
  • Repetitive Sequences, Nucleic Acid
  • Segmental Duplications, Genomic
  • Y Chromosome / genetics*

Substances

  • Euchromatin

Associated data

  • GENBANK/BK008736
  • GENBANK/BK008737
  • GENBANK/BK008738
  • GENBANK/BK008739
  • GENBANK/BK008740
  • GENBANK/BK008741
  • GENBANK/BK008742
  • GENBANK/BK008743
  • GENBANK/BK008744