Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
150
datasets available to search
ShareScore release 0.9.0
Dataset results
150 results for “Evolutionary modelling”
Data from: Digging through model complexity: using hierarchical models to uncover evolutionary processes in the wild
The growing interest for studying questions in the wild requires acknowledging that eco-evolutionary processes are complex, hierarchically structured and often partially observed or with measurement error. These issues have long been ignored in evolutionary biology, which might have led to flawed inference when addressing evolutionary questions. Hierarchical modelling (HM) has been proposed as a generic statistical framework to deal with complexity in ecological data and account for uncertainty. However, to date, HM has seldom been used to investigate evolutionary mechanisms possibly underlying observed patterns. Here, we contend the HM approach offers a relevant approach for the study of eco-evolutionary processes in the wild by confronting formal theories to empirical data through proper statistical inference. Studying eco-evolutionary processes requires considering the complete and often complex life histories of organisms. We show how this can be achieved by combining sequentially all life histories components and all available sources of information through HM. We demonstrate how eco-evolutionary processes may be poorly inferred or even missed without using the full potential of HM. As a case study, we use the Atlantic salmon and data on wild marked juveniles. We assess a reaction norm for migration and two potential trade-offs for survival. Overall, HM has a great potential to address evolutionary questions and investigate important processes that could not previously be assessed in laboratory or short time-scale studies.
Orthrus: Towards Evolutionary and Functional RNA Foundation Models
<p>Orthrus is a mature RNA model for RNA property prediction. It uses a Mamba encoder backbone, a variant of state-space models specifically designed for long-sequence data, such as RNA.</p> <p> </p> <p>Two versions of Orthrus are available:</p> <ul> <li>4-track base version: Encodes the mRNA sequence with a simplified one-hot approach.</li> <li>6-track large version: Adds biological context by including splice site indicators and coding sequence markers, which is crucial for accurate mRNA property prediction such as RNA half-life, ribosome load, and exon junction detection.</li> </ul> <p>This repository contains the annotations used to train Orthrus, as well as processed datasets used to evaluate Orthrus's ability to perform RNA property prediction. The datasets are taken from the following sources:</p> <ul> <li>Protein Subcellular Localization: Thul, P. J. et al. A subcellular map of the human proteome. Science 356 (2017).</li> <li>Mean Ribosome Load: Sugimoto, Y. & Ratcliffe, P. J. Isoform-resolved mRNA profiling of ribosome load defines interplay of HIF and mTOR dysregulation in kidney cancer. Nature Structural Molecular Biology 29, 871–880 (2022).</li> <li>RNA Halflife: Agarwal, V. & Kelley, D. R. The genetic and biochemical determinants of mRNA degradation rates in mammals. Genome Biol 23, 245 (2022).</li> <li>GO Molecular Function: Consortium, T. G. O. et al. The Gene Ontology knowledgebase in 2023. Genetics 224, iyad031 (2023).</li> </ul> <p>Dataset files encode sequence using a one-hot encoding using the vocabulary: [A, C, G, T].</p>
Figure 1 from: Sánchez-Fernández D, Rizzo V, Bourdeau C, Cieslak A, Comas J, Faille A, Fresneda J, Lleopart E, Millán A, Montes A, Pallares S, Ribera I (2018) The deep subterranean environment as a model system in ecological, biogeographical and evolutionary research. Subterranean Biology 25: 1-7. https://doi.org/10.3897/subtbiol.25.23530
Figure 1 Relationship between the temperature inside the cave and the surface (Mean Annual Temperature (°C) of each pixel (0.08° cells).
Data from: Mechanistic model of evolutionary rate variation en route to a nonphotosynthetic lifestyle in plants
Because novel environmental conditions alter the selection pressure on genes or entire subgenomes, adaptive and nonadaptive changes will leave a measurable signature in the genomes, shaping their molecular evolution. We present herein a model of the trajectory of plastid genome evolution under progressively relaxed functional constraints during the transition from autotrophy to a nonphotosynthetic parasitic lifestyle. We show that relaxed purifying selection in all plastid genes is linked to obligate parasitism, characterized by the parasite's dependence on a host to fulfill its life cycle, rather than the loss of photosynthesis. Evolutionary rates and selection pressure coevolve with macrostructural and microstructural changes, the extent of functional reduction, and the establishment of the obligate parasitic lifestyle. Inferred bursts of gene losses coincide with periods of relaxed selection, which are followed by phases of intensified selection and rate deceleration in the retained functional complexes. Our findings suggest that the transition to obligate parasitism relaxes functional constraints on plastid genes in a stepwise manner. During the functional reduction process, the elevation of evolutionary rates reaches several new rate equilibria, possibly relating to the modified protein turnover rates in heterotrophic plastids.
Demographic inferences and climatic niche modeling shed light on the evolutionary history of the emblematic cold-adapted Apollo butterfly at regional scale
<p>Cold-adapted species escape climate warming by latitudinal and/or altitudinal range shifts, and currently occur in Southern Europe in isolated mountain ranges within 'sky islands.</p> <p>Here we studied the genetic structure of the Apollo butterfly in five such alpine islands (above 1000 m) in France, and infer its demographic history since the last interglacial, using single nucleotide polymorphisms (ddRADseq SNPs). The Auvergne and Alps populations show strong genetic differentiation but not alpine massifs, although separated by deep valleys. Combining three complementary demographic inference methods and species distribution models (SDMs) we show that the LIG period was highly defavorable for Apollo that probably survived in small population in the highest summits of Auvergne. The population shifted downslope and expanded eastward between LIG and LGM throughout the large climatically suitable Rhône valley between the glaciated summits of Auvergne and Alps. The Auvergne and Alps populations started diverging before the LGM but remained largely connected till the mid-Holocene. Population decline in Auvergne was more gradual but started before (~7 kya versus 800 ya), and was much stronger with current population size ten times lower than in the Alps. In the Alps, the low genetic structure and limited evidence for isolation by distance suggest a non-equilibrium metapopulation functioning. The core Apollo population experienced cycles of contraction-expansion with climate fluctuations with largely inter-connected populations over time according to a 'metapopulation-pulsar' functioning. This study demonstrates the power of combining demographic inferences and SDMs to determine past and future evolutionary trajectories of an endangered species at a regional scale.</p>
Data from: When scientific experts come to be media stars: an evolutionary model tested by analysing coronavirus media coverage across Italian newspapers
<p>This dataset includes metadata of the newspaper articles used for the paper "When scientific experts come to be media stars: an evolutionary model tested by analysing coronavirus media coverage across Italian newspapers". The dataset is in JSON format. The metadata includes: "uuid" (unique identifier we associated to an article), "URLs" (the URLs where the article was published), "sources" (newspaper and feed/section where the article was published), "datesPublished" (dates when the article was published/updated).</p> <p>License: Attribution-ShareAlike 4.0 International (<a href="https://creativecommons.org/licenses/by-sa/4.0/legalcode">https://creativecommons.org/licenses/by-sa/4.0/legalcode</a>)</p> <p> </p>
Data and code for, "Large language models design sequence-defined macromolecules via evolutionary optimization"
<div> <pre># Codes and data for "Large language models design sequence-defined macromolecules via evolutionary optimization"<br><br>Note this repository contains codes and data files for the manuscript. This is a snapshot of the repository, frozen at the time of submission.<br><br># Codes<br><br>## LLM codes<br>- `run_claude.py` - the routine for performing LLM-based rollouts; intended for command line execution using argparse<br>- `message_utils.py` - utilities for constructing and parsing messages for LLM I/O<br>- `model_utils.py` - lightweight utilities for retrieving formatted predictions from the RNN ensemble<br>- `target_defs.py` - defines the sequence, locations, and natural language descriptions of the target structures<br>- `ask_about_oracle.ipynb` - asks the LLM to speculate about the nature of the optimization task<br><br>## other algorithms<br>- `active_learning.ipynb` - use EI acquisition with RF surrogate to label new sequences; includes an unused tokenization scheme<br>- `evolutionary_algorithm.ipynb` - use DEAP library to perform evolutionary optimization<br>- `random_sampling.ipynb` - sample sequences randomly from all possible sequences<br><br>## postprocessing<br>- `process_aggregated_logs.py` - reads data from the raw log files and prepares them for visualization<br>- `process_sample_rollouts.py` - reads data from the raw log files and prepares individual rollouts<br><br>## visualization<br>- `figure1b.ipynb` - renders panel b of Fig. 1<br>- `figure1efg.ipynb` - renders the last row of Fig. 1 (panels e-g)<br>- `figure2.ipynb` - renders all of Fig. 2<br>- `figure_si.ipynb` - renders Figs. S1 and S2<br>- `figure_md_validation.ipynb` - renders Fig. S3<br><br># Data files<br><br>- `prompts/`<br> - `prompt-scientific-v4.4.yml` - the full text of the scientific prompt, to be read by `run_claude.py`<br> - `prompt-oracle-v4.4.yml` - the full text of the oracle prompt, to be read by `run_claude.py`<br>- `models/` - the TorchScript RNN models used to make predictions<br>- `data/`<br> - `embeddings` - calculated embeddings for a collection of sequences from our prior work<br> - `llm-logs` - the raw logs obtained from the Claude 3.5 Sonnet LLM (other algorithms made to look like the LLM logs after the fact)<br> - `llm-logs-opus` - the raw logs obtained from the Claude 3.0 Opus LLM (used in the first draft of the article, replaced by Claude 3.5 Sonnet) <br> - `all-rollouts-kltd.csv` - postprocessed logs for all the rollouts using the "top $k < d^*$" metric<br> - `all-rollouts-topkd.csv` - postprocessed logs for all the rollouts using the "mean $d$ for top $k$" metric<br> - `sample-rollout-membranes-x-3.csv` - postprocessed logs for a single rollout replica, `x` = each algorithm type<br> - `snapshots` - png snapshots of MD simulation results at different locations in the manifold</pre> </div>
Demographic inferences and climatic niche modeling shed light on the evolutionary history of the emblematic cold-adapted Apollo butterfly at regional scale
Open the record for dataset details and reuse information.
Data from: Avian malaria: a new lease of life for an old experimental model to study the evolutionary ecology of Plasmodium
Open the record for dataset details and reuse information.
Data from: Integrating fossils, phylogenies, and niche models into biogeography to reveal ancient evolutionary history: the case of Hypericum (Hypericaceae)
Open the record for dataset details and reuse information.
Data from: A general and efficient algorithm for the likelihood of diversification and discrete-trait evolutionary models
Open the record for dataset details and reuse information.
Data from: Patterns of male fitness conform to predictions of evolutionary models of late-life
Open the record for dataset details and reuse information.
Data from: A new Bayesian method for fitting evolutionary models to comparative data with intraspecific variation
Open the record for dataset details and reuse information.
Supplementary code for: Polygenic local adaptation in metapopulations: a stochastic eco-evolutionary model
Open the record for dataset details and reuse information.
Data from: Bacterial competition and quorum-sensing signalling shapes the eco-evolutionary outcomes of model in vitro phage therapy
Open the record for dataset details and reuse information.
Data from: An evolutionary modelling approach to understanding the factors behind plant invasiveness and community susceptibility to invasion
Open the record for dataset details and reuse information.
Data from: Digging through model complexity: using hierarchical models to uncover evolutionary processes in the wild
Open the record for dataset details and reuse information.
Data from: An explicit model for the inbreeding load in the evolutionary analysis of selfing
Open the record for dataset details and reuse information.
Data from: How does evolutionary variation in basal metabolic rates arise? A statistical assessment and a mechanistic model
Open the record for dataset details and reuse information.
Data from: When should we expect early bursts of trait evolution in comparative data? Predictions from an evolutionary food web model
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.