In this section, we outline the method we used to assign genomes to each species cluster. We also verify the robustness of the species assignment by using three different methods and show that they give identical results.
Comparisons of genomes from bacterial populations frequently enough reveal clusters at different length and divergence scales (Jain et al.,2018). But the interpretation of such clusters is controversial (Cohan, 2002; Fraser et al., 2009). There are three main possible interpretations for such patterns. one interpretation is that the clusters represent the bacterial analog of eukaryotic species, which are sexually and ecologically isolated from each other. In practice, a whole-genome divergence cutoff of 5% is frequently used to define bacterial species (Jain et al., 2018). An alternative interpretation is that clusters are the result of clonal expansions of particular strains from a much more diverse and possibly more genetically homogeneous population.The third interpretation is that the clusters are the result of uneven sampling from a larger and more diverse population. The unique combination of having a large collection of genomes, with minimal compositional bias, from a geographically isolated population allows us to test these alternatives. We will mainly aim to distinguish between the first two interpretations and the third one by checking whether there are robust genomic clusters in our sample. More discussion on the evidence concerning the first two hypotheses can be found in Appendices 3 and 5.
We used three different methods to check whether the genomes form robust clusters: the phylogenetic tree of the 16S rRNA sequences, average genome-wide divergence from the two reference genomes OS-A and OS-B’, and patterns of divergences between triplets formed by the SAG and the two reference sequences at different loci.The phylogenetic tree of the 16S is the simplest and still the most commonly used method to characterize the diversity of natural microbial communities (Yarza et al., 2014). Because of its practical importance, we included it for comparison with the other two methods. Though, in a highly recombining population like the Yellowstone Synechococcus, it is indeed not clear to what extent the 16S divergence accurately reflects the whole-genome diversity. To address this question, we used the divergences across the entire genome as a comparison. The main limitation of using summary statistics such as the average divergence across the entire genome is that it obscures heterogeneities along the genome that could have crucial ecological and evolutionary implications.There are different ways to address this concern. Here, we used an extension of a method first used in Rosen et al., 2018, which provides a coarse description of the variation in the diversity patterns across different loci. We present the results from each of these analyses next.
We first constructed the phylogenetic tree of 16S sequences. Predicted 16S rRNA sequences (117 from 331 filtered SAGs) were extracted and aligned using MAFFT (v7.453) (Katoh and Standley, 2013). Two sequences (SAGs MA02M11_3 and MuA02C9) were incomplete and were removed from the analysis. The remaining (115) sequences were hierarchically clustered using the average linkage (UPGMA) method implemented in Scipy.Results are shown in Appendix 2-figure 3, with the colored background showing the species assignment using whole-genome divergences (see below) for comparison. We found perfect agreement between the whole-genome assignment and the largest three clusters in the 16S phyl
“`
for the diversity within species. We then classify the shape of the triangle into the 5 distinct patterns shown in Appendix 2-figure 1, depending on whether each side of the triangle is longer or shorter than . Such as, denote the divergence triplet by , representing the divergences between the SAG sequence and OS-A, the SAG sequence and OS-B’, and OS-A and OS-B’, respectively. Then the first pattern in Appendix 2-figure 1A (labeled ‘XA-B’) contains loci where , , and , the second pattern (labeled ‘XAB’) contains loci where , , and , and so on. The cutoff of is the same as the one used in Rosen et al.,
“`## Appendix 2-Figure 1. Fingerprints of SAGs.
Each row represents a SAG, and each column represents a gene family. The color of each cell indicates the presence (blue) or absence (white) of a gene from that family in the SAG’s genome. The SAGs are ordered according to their average divergence from OS-A and OS-B’ as shown in Figure 1A. The three distinct patterns of gene family presence/absence correspond to the three species assignments.
Gene Triplet Patterns Show Distinct Clusters
The proportion of gene triplet patterns shows three distinct clusters.
The proportion of genes in each of the five patterns shown in Appendix 2-figure 1 is shown for each cell passing our quality control criteria (see Appendix A). Lines represent different cells and are colored according to the species classification based on whole-genome divergences.
Related reading