Large-Scale Genome Comparison Based on Cumulative Fourier Power and Phase Spectra: Central Moment and Covariance Vector

Shaojun Pei Department of Mathematical Sciences, Tsinghua University, Beijing, PR China Rui Dong Department of Mathematical Sciences, Tsinghua University, Beijing, PR China Rong Lucy He Department of Biological Sciences, Chicago State University, Chicago, IL 60628, USA Stephen S.-T. Yau Department of Mathematical Sciences, Tsinghua University, Beijing, PR China

Data Analysis, Bio-Statistics, Bio-Mathematics mathscidoc:1909.42001

Computational and Structural Biotechnology Journal, 17, 982-994, 2019.7
Genome comparison is a vital research area of bioinformatics. For large-scale genome comparisons, the Multiple Sequence Alignment (MSA) methods have been impractical to use due to its algorithmic complexity. In this study, we propose a novel alignment-free method based on the one-to-one correspondence between a DNA sequence and its complete central moment vector of the cumulative Fourier power and phase spectra. In addition, the covariance between the four nucleotides in the power and phase spectra is included. We use the cumulative Fourier power and phase spectra to define a 28-dimensional vector for each DNA sequence. Euclidean distances between the vectors can measure the dissimilarity between DNA sequences. We perform testing with datasets of different sizes and types including simulated DNA sequences, exon-intron and complete genomes. The results show that our method is more accurate and efficient for performing hierarchical clustering than other alignment-free methods and MSA methods.
Cumulative fourier transform; Power and phase spectra; Central moments; Covariance
[ Download ] [ 2019-09-04 17:31:31 uploaded by Stephenyau ] [ 119 downloads ] [ 0 comments ]
@inproceedings{shaojun2019large-scale,
  title={Large-Scale Genome Comparison Based on Cumulative Fourier Power and Phase Spectra: Central Moment and Covariance Vector},
  author={Shaojun Pei, Rui Dong, Rong Lucy He, and Stephen S.-T. Yau},
  url={http://archive.ymsc.tsinghua.edu.cn/pacm_paperurl/20190904173131221833485},
  booktitle={Computational and Structural Biotechnology Journal},
  volume={17},
  pages={982-994},
  year={2019},
}
Shaojun Pei, Rui Dong, Rong Lucy He, and Stephen S.-T. Yau. Large-Scale Genome Comparison Based on Cumulative Fourier Power and Phase Spectra: Central Moment and Covariance Vector. 2019. Vol. 17. In Computational and Structural Biotechnology Journal. pp.982-994. http://archive.ymsc.tsinghua.edu.cn/pacm_paperurl/20190904173131221833485.
Please log in for comment!
 
 
Contact us: office-iccm@tsinghua.edu.cn | Copyright Reserved