?
t-SNE Highlights Phylogenetic and Temporal Patterns of SARS-CoV-2 Spike and Nucleocapsid Protein Evolution
Since the beginning of the COVID-19 pandemic, whole-genome sequences of SARS-CoV-2 have been continuously added to public databases, such as NCBI Virus and GISAID. As of July 2022, the SARS-CoV-2 Data Hub of the
NCBI Virus database stored more than one million complete whole-genome sequences of the coronavirus. For navigating the SARS-CoV-2 genome sequences, the Pango nomenclature and the Pangolin software were developed.
The nomenclature and the software have been also extensively used for tracking the coronavirus evolution and rapidly classifying new genomes. To supplement the Pango nomenclature, we propose applying modern manifold learning techniques to protein sequences of SARS-CoV-2 to construct, visualize and study the global evolutionary space of the coronavirus. The basic idea is to explore the COVID-19 evolution space by trying to find geometric structure hidden in the evolutionary distances between variants. The adequate visualization of such a geometric structure might provide more information on evolution compared to the conventional phylogeny methods. For instance, phylogenetic trees do not provide any information on possible evolutionary patterns, nor they may reveal the “empty spaces” in the global evolutionary space left by extinguished specimens. Moreover, phylogenetic trees, especially very large ones, may be not well adapted for cluster analysis of the species diversity, which was also the motivation under the application of manifold learning techniques, among all, to COVID-19 genomic data.