Direct estimation of genotype fitness from time series
Direct estimation of genotype fitness from time series
Mohanty, V.; Shakhnovich, E.
AbstractHeterogeneous adapting populations, whether in laboratory evolution experiments or global-scale pandemics, experience complex evolutionary dynamics due to the interplay of selection, mutation, and stochasticity. Inference of individual genotypes' fitnesses therefore becomes difficult, especially when many lineages are competing and data are noisy. Existing fitness inference methods tend to rely on assumptions on the fitness landscape's maximum order of epistasis, or they require complicated iterative optimization algorithms to converge on fitness estimates. Here, we show that fitness landscapes can be computed from time series data, without any restrictions on epistatic order or iterative optimization, using a simple, closed-form mathematical expression that is easily implemented with standard matrix operations used commonly in linear algebra. We demonstrate successful fitness inference from noisy in silico evolutionary dynamics from four different noisy microscopic processes, including Wright-Fisher, Moran, ProSeD (serial dilution), and barcoded passage simulations. Then, we illustrate the broad applicability of the equation to five experimental time series datasets, including barcoded yeast evolution experiments, murine norovirus-1 serial passage experiments, and SARS-CoV-2 global genomic prevalence data. Our formula successfully infers fitnesses for even for rare genotypes several orders of magnitude less prevalent than top lineages, works with both laboratory evolution and epidemiological data, and can be implemented in most modern scientific programming languages.