close
Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 2016 Mar 7;26(5):692-7.
doi: 10.1016/j.cub.2016.01.016. Epub 2016 Feb 25.

The Effect of Local Sequence Context on Mutational Bias of Genes Encoded on the Leading and Lagging Strands

Affiliations

The Effect of Local Sequence Context on Mutational Bias of Genes Encoded on the Leading and Lagging Strands

Jeremy W Schroeder et al. Curr Biol. .

Abstract

All organisms must replicate their genetic information accurately to ensure its faithful transmission. DNA polymerase errors provide an important source of genetic variation that can drive evolution. Understanding the origins of genetic variation will inform our understanding of evolution and the development of genetic diseases. A number of factors have been proposed to influence mutagenesis [1-10]. Here, we used mutation accumulation lines, whole-genome sequencing, and whole-transcriptome analysis to study the locations and rate at which mutations arise in bacteria with as little selection bias as possible [11, 12]. Our analysis of greater than 7,000 replication errors in over 180 sequenced lines that underwent a total of more than 370,000 generations has provided new insights into how DNA polymerase errors sculpt genetic variation and drive evolution. Homopolymer run enrichment outside of genes causes insertions and deletions in these regions. Genes encoded in the lagging strand are transcribed such that RNA polymerase and DNA polymerase collide head-on. Head-on genes have been proposed to mutate at a higher rate than genes transcribed codirectionally with DNA polymerase progression due to conflicts between transcription and DNA replication [6, 10]. We did not detect associations between the number of base pair substitutions in genes and their orientation or expression. Strikingly, any higher mutation rate for head-on genes can be explained by differing sequence composition between the leading and lagging strands and the error bias for DNA polymerase in specific sequence contexts. Therefore, we find local sequence context is the major determinant of mutagenesis in bacteria.

PubMed Disclaimer

Conflict of interest statement

The authors have no conflict of interest to declare.

Figures

Figure 1
Figure 1. Indels in homopolymer runs are enriched outside of coding regions
All starts and ends of CDSs were aligned at relative position zero. Indels were counted in 50 bp bins, offset by 25. Negative distances indicate the indel was 5′ to a CDS start site or 3′ to a CDS end site. The lines in each plot represent the locally weighted polynomial regression (loess) fit to the data. (A) The number of indels found in each bin without correcting for homopolymer run bias. (B) Uncorrected indel counts separated by the homopolymer run length in which the indel was produced. (C) Expected indel count in each bin after applying a correction for homopolymer run bias (see Supplemental Experimental Procedures). See also Figure S3.
Figure 2
Figure 2. Transitions display complementary symmetry between replichores
(A) Schematic representation of the B. subtilis chromosome. DNA replication initiates at oriC, which is at position zero in the reference genome, and proceeds bidirectionally toward terC, which is at 1.97 Mb in the reference genome. The left and right replichores of the chromosome are in green and black, respectively. (B) Cumulative distributions of the indicated types of transitions along the genome. The origin of replication is indicated by the vertical dashed red line and the terminus is at both ends. (C) A barplot displaying the mutation rate for the indicated types of transitions binned by replichore. The mutation rate is normalized to the number of each base in each replichore as described in Supplemental Experimental Procedures. Error bars represent 95% confidence intervals determined by bootstrapping. All comparisons between left and right replichores are statistically significant with the exception of T → C transitions in MMR+ lines. MMR intact refers to wild type data and MMR deficient refers to the pooled data for ΔmutSL, ΔwalJ, and mutL[E468K].
Figure 3
Figure 3. Regression of base pair substitution count against coding sequence length, expression and orientation
(A) A graphical representation of linear regression analysis. Blue triangles indicate head-on CDSs and red circles represent codirectional CDSs. Lines indicate the linear fit to the data, and the shaded region around each line indicates the 95% confidence interval for the fit. The plot on the left includes all CDSs and the plot on the right excludes CDSs greater or less than three standard deviations from either the mean length or expression (RPKM). See Table S4 for a summary of each CDS including which were determined to be outliers. (B) A table listing the variables determined to be significantly associated with the average number of BPSs found in CDSs either with or without outliers. See Table S3 for detailed results and Equation S5 for the regression model.
Figure 4
Figure 4. Increased mutation rate in head-on genes due to sequence composition
(A) The transition rate for the focal base in each of the 64 possible triplet nucleotide sequences in MMR-lines is shown. Rates are normalized to the number of times each triplet is present in the leading strand. (B) The leading strand triplet composition of head-on CDSs is plotted versus that of codirectional CDSs. “Fraction of triplets” in the axis labels refers to the number of a given triplet divided by the total number of all triplets present in the leading strand of either head-on or codirectional CDSs. The triplets with the highest transition rates are plotted as larger red dots and indicated by arrows. The red dashed line indicates a slope of one so that differences between head-on and codirectional genes may be easily noticed. (C) Monte Carlo simulations were performed to generate transitions using the MMR-context-dependent transition rates shown in panel A. The boxplots on the left indicate the distribution of transition rates for codirectional (red) and head-on (blue) CDSs using the MMR-context-dependent transition rates. Boxplots on the right indicate the distribution of transition rates with the 5′-CCG-3′ triplet rate set artificially to zero. Each pair of boxplots represents the distribution of mutation rates resulting from 1,000 independent simulations of 500,000 generations. (D) The simulation performed in C was carried out at a range of generations per iteration. For each number of generations, 1,000 iterations was performed and hypothesis testing was carried out to test whether head-on genes had a mutation rate greater than that of codirectional genes. The proportion of those 1,000 p-values less than or equal to 0.05 is plotted in the y-axis against the number of generations per iteration. See also Figure S3.

References

    1. Fijalkowska IJ, Jonczyk P, Tkaczyk MM, Bialoskorska M, Schaaper RM. Unequal fidelity of leading strand and lagging strand DNA replication on the Escherichia coli chromosome. Proceedings of the National Academy of Sciences of the United States of America. 1998;95:10020–10025. - PMC - PubMed
    1. Lind PA, Andersson DI. Whole-genome mutational biases in bacteria. Proceedings of the National Academy of Sciences of the United States of America. 2008;105:17878–17883. - PMC - PubMed
    1. Ma X, Rogacheva MV, Nishant KT, Zanders S, Bustamante CD, Alani E. Mutation hot spots in yeast caused by long-range clustering of homopolymeric sequences. Cell Rep. 2012;1:36–42. - PMC - PubMed
    1. Lee H, Popodi E, Tang H, Foster PL. Rate and molecular spectrum of spontaneous mutations in the bacterium Escherichia coli as determined by whole-genome sequencing. Proceedings of the National Academy of Sciences of the United States of America. 2012;109:E2774–E2783. - PMC - PubMed
    1. Schaibley VM, Zawistowski M, Wegmann D, Ehm MG, Nelson MR, St Jean PL, Abecasis GR, Novembre J, Zollner S, Li JZ. The influence of genomic context on mutation patterns in the human genome inferred from rare variants. Genome Res. 2013;23:1974–1984. - PMC - PubMed

Publication types

MeSH terms

Substances

LinkOut - more resources