Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
66
datasets available to search
ShareScore release 0.7.1
Dataset results
66 results for “fitness landscapes”
Evolutionary "crowdsourcing": alignment of fitness landscapes allows for cross-species adaptation of a horizontally transferred gene
<p>This repository accompanies the publication of <i><strong>Evolutionary "crowdsourcing": alignment of fitness landscapes allows for cross-species adaptation of a horizontally transferred gene</strong></i> by Kosterlitz et. al. This research project explores the cross-species adaptation of a horizontally transferred gene through evolutionary "crowdsourcing." The repository provides all relevant data, code, and figures associated with the publication, enabling users to replicate the results and explore the findings in-depth.</p>
FLIGHTED: Inferring Fitness Landscapes from Noisy High-Throughput Experimental Data (Part 1)
<p>Data for FLIGHTED (Inferring Fitness Landscapes from Noisy High-Throughput Experimental Data). This data contains the model weights for FLIGHTED-Selection and FLIGHTED-DHARMA, the training data for both, and fits on the GB1 landscape. It does not contain anything related to the TEV protease landscape.</p> <p>The data is arranged in the following folders:</p> <ol> <li>DHARMA_Input: contains the input for the DHARMA models, with the canvas sequence, the DHARMA reads, and the FACS data.</li> <li>DHARMA_Models: contains the weights, hyperparameters, and model training history for the FLIGHTED-DHARMA model.</li> <li>Fitness_Landscapes: the GB1 landscape, with and without FLIGHTED, as well as splits published by FLIP.</li> <li>Landscape_Models: models trained on the GB1 landscape with and without FLIGHTED under the various FLIP splits. Each model folder contains hyperparameters, training history, and predictions on the test set which can be used to evaluate model performance. Raw model parameters are not provided for fine-tuned models due to size; contact us if you want them.</li> <li>FLIGHTED_Selection: contains the weights, hyperparameters, and model training history for the FLIGHTED-Selection model.</li> <li>Selection_Simulations: contains the simulated training data for the FLIGHTED-Selection model.</li> </ol>
FLIGHTED: Inferring Fitness Landscapes from Noisy High-Throughput Experimental Data (Part 2)
<p>Data for FLIGHTED (Inferring Fitness Landscapes from Noisy High-Throughput Experimental Data). This data contains the TEV protease landscape and models trained on it. All other FLIGHTED data is in the other Zenodo repository (refer to the paper for details).</p> <p>The data is arranged in the following folders:</p> <ol> <li>TEV_Landscape: contains the TEV landscape (in flighted_fitnesses.csv) and splits thereof in Splits/. The main files for model training are flighted_fitnesses.csv and the files labeled one_vs_rest, two_vs_rest, and three_vs_rest. The files labeled three_vs_rest_control within Splits/ and the read count CSV files refer to further information about the read count in the landscape; see the Supplement for details. The dictionary files are the original raw data prior to processing with FLIGHTED.</li> <li>TEV_Models: contains models trained on the TEV landscape under the various splits. Each model folder contains hyperparameters, training history, and predictions on the test set which can be used to evaluate model performance. Raw model parameters are not provided for fine-tuned models due to size; contact us if you want them. The control_run/ refers to the run described in the supplement on just read counts.</li> </ol>
Embeddings from FLIP: benchmark tasks in fitness landscape inference for proteins
<p>ESM-1b embeddings used in FLIP. </p>
Fig. 8 Fitness landscape for models 7 and 8 in Modelling sympatric speciation by means of biologically plausible mechanistic processes as exemplified by threespine stickleback species pairs
Fig. 8 Fitness landscape for models 7 and 8. Relative fitness is a function of trait T1 and trait T2. Epistasis is modelled as follows:
Data from: Analysis of statistical correlations between properties of adaptive walks in fitness landscapes
Open the record for dataset details and reuse information.
Data from: Pollen diversity matters: revealing the neglected effect of pollen diversity on fitness in fragmented landscapes
Open the record for dataset details and reuse information.
Data from: Contrasting habitat and landscape effects on the fitness of a long-lived grassland plant under forest encroachment: do they provide evidence for extinction debt?
Open the record for dataset details and reuse information.
Data from: Crossing fitness valleys: empirical estimation of adaptive landscape associated with polymorphic mimicry
Open the record for dataset details and reuse information.
Data from: Constraints imposed by a natural landscape override offspring fitness effects to shape oviposition decisions in wild forked fungus beetles
Open the record for dataset details and reuse information.
Data from: Mapping the fitness landscape of gene expression uncovers the cause of antagonism and sign epistasis between adaptive mutations
How do adapting populations navigate the tensions between the costs of gene expression and the benefits of gene products to optimize the levels of many genes at once? Here we combined independently-arising beneficial mutations that altered enzyme levels in the central metabolism of Methylobacterium extorquens to uncover the fitness landscape defined by gene expression levels. We found strong antagonism and sign epistasis between these beneficial mutations. Mutations with the largest individual benefit interacted the most antagonistically with other mutations, a trend we also uncovered through analyses of datasets from other model systems. However, these beneficial mutations interacted multiplicatively (i.e., no epistasis) at the level of enzyme expression. By generating a model that predicts fitness from enzyme levels we could explain the observed sign epistasis as a result of overshooting the optimum defined by a balance between enzyme catalysis benefits and fitness costs. Knowledge of the phenotypic landscape also illuminated that, although the fitness peak was phenotypically far from the ancestral state, it was not genetically distant. Single beneficial mutations jumped straight toward the global optimum rather than being constrained to change the expression phenotypes in the correlated fashion expected by the genetic architecture. Given that adaptation in nature often results from optimizing gene expression, these conclusions can be widely applicable to other organisms and selective conditions. Poor interactions between individually beneficial alleles affecting gene expression may thus compromise the benefit of sex during adaptation and promote genetic differentiation.
Data from: The fitness landscape of the codon space across environments
Fitness landscapes map the relationship between genotypes and fitness. However, most fitness landscape studies ignore the genetic architecture imposed by the codon table and thereby neglect the potential role of synonymous mutations. To quantify the fitness effects of synonymous mutations and their potential impact on adaptation on a fitness landscape, we use a new software based on Bayesian Monte Carlo Markov Chain methods and re-estimate selection coefficients of all possible codon mutations across 9 amino-acid positions in Saccharomyces cerevisiae Hsp90 across 6 environments. We quantify the distribution of fitness effects of synonymous mutations and show that it is dominated by many mutations of small or no effect and few mutations of larger effect. We then compare the shape of the codon fitness landscape across amino-acid positions and environments, and quantify how the consideration of synonymous fitness effects changes the evolutionary dynamics on these fitness landscapes. Together these results highlight a possible role of synonymous mutations in adaptation and indicate the potential mis-inference when they are neglected in fitness landscape studies.
Data from: Fisher's geometrical model of fitness landscape and variance in fitness within a changing environment
The fitness of an individual can be simply defined as the number of its offspring in the next generation. However, it is not well understood how selection on the phenotype determines fitness. In accordance with Fisher's fundamental theorem, fitness should have no or very little genetic variance, whereas empirical data suggest that is not the case. To bridge these knowledge gaps, we follow Fisher's geometrical model and assume that fitness is determined by multivariate stabilizing selection towards an optimum that may vary among generations. We assume random mating, free recombination, additive genes, and uncorrelated stabilizing selection and mutational effects on traits. In a constant environment, we find that genetic variance in fitness under mutation-selection balance is a U-shaped function of the number of traits (i.e. of the so-called "organismal complexity"). Because the variance can be high if the organism is of either low or high complexity, this suggests that complexity has little direct costs. Under a temporally varying optimum, genetic variance increases relative to a constant optimum and increasingly so when the mutation rate is small. Therefore mutation and changing environment together can maintain high genetic variance. These results therefore lend support to Fisher's geometric model of a fitness landscape.
Data from: On the (un)predictability of a large intragenic fitness landscape
The study of fitness landscapes, which aims at mapping genotypes to fitness, is receiving ever-increasing attention. Novel experimental approaches combined with next-generation sequencing (NGS) methods enable accurate and extensive studies of the fitness effects of mutations, allowing us to test theoretical predictions and improve our understanding of the shape of the true underlying fitness landscape and its implications for the predictability and repeatability of evolution. Here, we present a uniquely large multiallelic fitness landscape comprising 640 engineered mutants that represent all possible combinations of 13 amino acid-changing mutations at 6 sites in the heat-shock protein Hsp90 in Saccharomyces cerevisiae under elevated salinity. Despite a prevalent pattern of negative epistasis in the landscape, we find that the global fitness peak is reached via four positively epistatic mutations. Combining traditional and extending recently proposed theoretical and statistical approaches, we quantify features of the global multiallelic fitness landscape. Using subsets of the data, we demonstrate that extrapolation beyond a known part of the landscape is difficult owing to both local ruggedness and amino acid-specific epistatic hotspots and that inference is additionally confounded by the nonrandom choice of mutations for experimental fitness landscapes.
Data from: Host use dynamics in a heterogeneous fitness landscape generates oscillations in host range and diversification
Colonization of novel hosts is thought to play an important role in parasite diversification, yet little consensus has been achieved about the macroevolutionary consequences of changes in host use. Here we offer a mechanistic basis for the origins of parasite diversity by simulating lineages evolved in silico. We describe an individual-based model in which (i) parasites undergo sexual reproduction limited by genetic proximity, (ii) hosts are uniformly distributed along a one-dimensional resource gradient, and (iii) host use is determined by the interaction between the phenotype of the parasite and a heterogeneous fitness landscape. We found two main effects of host use on the evolution of a parasite lineage. First, the colonization of a novel host allowed parasites to explore new areas of the resource space, increasing phenotypic and genotypic variation. Second, hosts produced heterogeneity in the parasite fitness landscape, which led to reproductive isolation and therefore, speciation. As a validation of the model, we analyzed empirical data from Nymphalidae butterflies and their host plants. We then assessed the number of hosts used by parasite lineages and the diversity of resources they encompass. In both simulated and empirical systems, host diversity emerged as the main predictor of parasite species richness.
Data from: Male competition fitness landscapes predict both forward and reverse speciation
Speciation is facilitated when selection generates a rugged fitness landscape such that populations occupy different peaks separated by valleys. Competition for food resources is a strong ecological force that can generate such divergent selection. However, it is unclear whether intrasexual competition over resources that provide mating opportunities can generate rugged fitness landscapes that foster speciation. Here we use highly variable male F2 hybrids of benthic and limnetic threespine sticklebacks, Gasterosteus aculeatus Linnaeus, 1758, to quantify the male competition fitness landscape. We find that disruptive sexual selection generates two fitness peaks corresponding closely to the male phenotypes of the two parental species, favouring divergence. Most surprisingly, an additional region of high fitness favours novel hybrid phenotypes that correspond to those observed in a recent case of reverse speciation after anthropogenic disturbance. Our results reveal that sexual selection through male competition plays an integral role in both forward and reverse speciation.
Data from: Shifting fitness landscapes in response to altered environments
The role of adaptation in molecular evolution has been contentious for decades. Here, we shed light on the adaptive potential in Saccharomyces cerevisiae by presenting systematic fitness measurements for all possible point mutations in a region of Hsp90 under four environmental conditions. Under elevated salinity, we observe numerous beneficial mutations with growth advantages up to 7% relative to the wild type. All of these beneficial mutations were observed to be associated with high costs of adaptation. We thus demonstrate that an essential protein can harbor adaptive potential upon an environmental challenge, and report a remarkable fit of the data to a version of Fisher's geometric model that focuses on the fitness trade-offs between mutations in different environments.
Data from: Comprehensive experimental fitness landscape and evolutionary network for small RNA
The origin of life is believed to have progressed through an RNA world, in which RNA acted as both genetic material and functional molecules. The structure of the evolutionary fitness landscape of RNA would determine natural selection for the first functional sequences. Fitness landscapes are the subject of much speculation, but their structure is essentially unknown. Here we describe a comprehensive map of a fitness landscape, exploring nearly all of sequence space, for short RNAs surviving selection in vitro. With the exception of a small evolutionary network, we find that fitness peaks are largely isolated from one another, highlighting the importance of historical contingency and indicating that natural selection would be constrained to local exploration in the RNA world.
Data from: Efficient escape from local optima in a highly rugged fitness landscape by evolving RNA virus populations
Open the record for dataset details and reuse information.
Data from: Comprehensive experimental fitness landscape and evolutionary network for small RNA
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.