Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

5,805

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

5,805 results for “Data model”

Learn how ShareScore rates datasets ↗
zenodo40/100

Input data of GridPath model for South America's MERCOSUR sub-region

<p><strong>This repository contains input data of the GridPath model for South America&rsquo;s MERCOSUR sub-region, presented in the forthcoming paper entitled &ldquo;Exploring sustainable electricity system development pathways in South America&rsquo;s MERCOSUR sub-region&rdquo; by the same authors. These data can be used in conjunction with the GridPath model (version v0.8.0), available at <a href="https://github.com/blue-marble/gridpath](https://github.com/blue-marble/gridpath)">https://github.com/blue-marble/gridpath</a>. Description of the key input data and the scenarios presented in the paper are provided in the&nbsp;<a href="https://docs.google.com/spreadsheets/d/1l7ImQtKHpYTPp9cMmTGHCdJB-wJvq7klwrwKmrtCrqw/edit?usp=sharing">Readme file</a>.</strong></p>

opencc-by-4.0Sep 2022View details →
dryad40/100

Data from: Learning of probabilistic punishment as a model of anxiety produces changes in action but not punisher encoding in the dmPFC and VTA

<p>Previously, we developed a novel model for anxiety during motivated behavior by training rats to perform a task where actions executed to obtain a reward were probabilistically punished and observed that after learning, neuronal activity in the ventral tegmental area (VTA) and dorsomedial prefrontal cortex (dmPFC) represent the relationship between action and punishment risk (Park &amp; Moghaddam, 2017). Here we used male and female rats to expand on the previous work by focusing on neural changes in the dmPFC and VTA that were associated with the learning of probabilistic punishment, and anxiolytic treatment with diazepam after learning. We find that adaptive neural responses of dmPFC and VTA during the learning of anxiogenic contingencies are independent from the punisher experience and occur primarily during the peri-action and reward period. Our results also identify peri-action ramping of VTA neural calcium activity, and VTA-dmPFC correlated activity, as potential markers for the anxiolytic properties of diazepam.</p>

opencc-zeroSep 2022View details →
zenodo40/100

Data and Code for "Quantifying spatio-temporal risk of Harmful Algal Blooms and their impacts on bivalve shellfish mariculture using a data-driven modelling approach"

<p>This is a zipped file of all associated code and data for the submitted paper entitled &quot;Quantifying spatio-temporal risk of Harmful Algal Blooms and their impacts on bivalve shellfish mariculture using a data-driven modelling approach&quot;.</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Data and code related to publication "Migration pulsedness alters patterns of allele fixation and local adaptation in a mainland-island model" - Aubree et al. 2021

<p>Those data sets and codes are related to the manuscript &quot;Migration pulsedness alters patterns of allele fixation and local adaptation in a mainland-island model&quot; available on BioRXiv.</p> <p>All the information that are necessary to use those data sets and codes are contained in the file &quot;readme.txt&quot;.</p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

CMIP6 model sea-ice concentration data

<p>Sea-ice concentration data from Climate System Models that participated in the sixth Coupled Model Intercomparison Project (CMIP6).&nbsp;<br> <br> Historical (1850-2014) and future projections (2015-2100) of sea-ice concentration were downloaded on the native model grid of 29 models, after which this data was regridded to a regular 1x1 degree latitude by longitude grid for multi-model comparison and calculation of statistics (mean, median, etc.).</p> <p>Future projections include the SSP1-2.6, SSP2-4.5 and SSP5-8.5 scenarios.</p> <p>Models included in the scenarios are shown in the excel spreadsheet.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Pre-processed ex vivo MRI data for manuscript titled "Neuroanatomical and cognitive biomarkers of alpha-synuclein propagation in a mouse model of synucleinopathy prior to onset of motor symptoms""

<p>Repository for <em>ex vivo</em> magnetic resonance imaging data from the project&nbsp;titled &quot;Presymptomatic neuroanatomical and cognitive biomarkers of alpha-synuclein propagation in a mouse model of synucleinopathy&quot;</p> <p>Contains the pre-processed <em>ex vivo</em> T1-weighted images (Bruker 7T; 70&nbsp;micron isotropic voxel resolution) for M83 alpha-synuclein A53T hemizygous mice that received either a phosphate buffered saline (PBS) or alpha-synuclein pre-formed fibrils (PFF) injection in the right dorsal striatum. Full subject list can be viewed with the &quot;subject_list.csv&quot; file. More details are available in the manuscript.&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

"Chirality and accurate structure models by exploiting dynamical effects in continuous-rotation 3D ED data". Raw data and JANA refinement files.

<p><strong>Chirality and accurate structure models by exploiting dynamical effects in continuous-rotation 3D ED data</strong><br> 3D ED data sets of 5 compounds and JANA refinement files of 12 compounds</p> <p><strong>Relevant tools</strong><strong>:</strong></p> <ul> <li>PETS2: data reduction and analysis of electron diffraction patterns <ul> <li>Download program and access step-by-step tutorials at <a href="http://pets.fzu.cz/">http://pets.fzu.cz/</a></li> <li>Palatinus, L. <em>et al.</em> Specifics of the data processing of precession electron diffraction tomography data and their implementation in the program PETS2.0. <em>Acta Cryst. B</em><strong>75</strong>, 512&ndash;522 (2019). <a href="https://doi.org/10.1107/S2052520619007534">DOI: 10.1107/S2052520619007534</a></li> </ul> </li> <li>JANA2006: crystal structure model refinement program <ul> <li>Download program from <a href="http://jana.fzu.cz/">http://jana.fzu.cz/</a> and access step-by-step tutorials at <a href="http://pets.fzu.cz/">http://pets.fzu.cz/</a></li> <li>Results here were obtained with JANA2006. We recommend using JANA2020.</li> <li>Petricek, V., Dusek, M. &amp; Palatinus, L. Crystallographic Computing System JANA2006: General features. <em>Z. Kristallogr.</em> <strong>229</strong>, 345&ndash;352 (2014). <a href="https://doi.org/10.1515/zkri-2014-1737">DOI: 10.1515/zkri-2014-1737</a></li> </ul> </li> <li>DYNGO: Bloch wave program, calculates dynamical diffraction intensities and derivatives <ul> <li>Program automatically included in JANA2006/JANA2020</li> <li>Palatinus, L., Petř&iacute;ček, V. &amp; Corr&ecirc;a, C. A. Structure refinement using precession electron diffraction tomography and dynamical diffraction: theory and implementation. <em>Acta Cryst. A</em><strong>71</strong>, 235&ndash;244 (2015). <a href="https://doi.org/10.1107/S2053273315001266">DOI: 10.1107/S2053273315001266</a></li> </ul> </li> </ul> <p><strong>3D ED data sets:</strong></p> <p>STW_HPM-1 (RT) was measured on a JEOL JEM-2100-LaB6 and diffraction patterns were recorded with an ASI Timepix detector. Another sample of STW_HPM-1 was measured at a temperature of 100 K after cryotransfer with a Titan Krios (CETA-D detector). The other data sets were measured on an FEI Tecnai G2 20 (Olympus SIS Veleta, CCD). Each data set contains the raw diffraction patterns (*.tif) and the basic input files needed to reproduce the data reduction with PETS2 as used in the associated publication (*.pts2, *.celllist, *.cenloc). Step-by-step tutorials are provided for quartz and glycine (and selected steps for abiraterone acetate) at <a href="http://pets.fzu.cz/">http://pets.fzu.cz/</a>.</p> <ul> <li>&alpha;-quartz, stepwise continuous-rotation and precession-assisted (2 data sets from the same crystal)</li> <li>natrolite, stepwise continuous-rotation and precession-assisted (2 data sets from the same crystal)</li> <li>cobalt aluminophosphate (CAP), static ED patterns recorded in 0.1&deg; steps (3 data sets from 2 crystals)</li> <li>abiraterone acetate, stepwise continous-rotation (5 data sets from 5 crystals)</li> <li>STW_HPM-1, continuous-rotation (1 data set, room temperature)</li> <li>STW_HPM-1, continuous-rotation (1 data set, <em>T</em> = 100 K, cryotransfer)</li> </ul> <p><strong>JANA refinement and CIF files:</strong></p> <p>CIF (Crystallographic Information Framework) files include two data items. The first is related to the dynamical and the second to the kinematical refinement. Relevant parameters and statistics specific for dynamical refinement are found in the field _refine_special_details.</p> <p>JANA files are provided for the dynamical and kinematical refinement at the stage after the final refinement cycle together with the original input files generated by PETS2. For quartz and natrolite, relevant files for the refinements against precession-assisted 3D ED data are included. For abiraterone acetate and limaspermidine, relevant files for the absolute structure determination are included.</p> <ul> <li>&alpha;-quartz</li> <li>albite</li> <li>mordenite</li> <li>natrolite</li> <li>STW_HPM-1</li> <li>cobalt aluminophosphate (CAP)</li> <li>CAU-36</li> <li>&alpha;-glycine</li> <li>carbamazepine</li> <li>(+)-limaspermidine</li> <li>abiraterone acetate</li> <li>MBBF4</li> </ul> <p>For the kinematical refinements based on more than one data set, the self-written tool &quot;CompInt&quot; (unpublished) was used. The tool can be found in the file &quot;tool_scalehkl_compint.zip&quot;. Input (*.hkl, *.compint) and output files (*.scalehkl) are provided in the respective folder with the JANA files.</p> <p>Raw data sources of other data sets relevant for the associated publication are given in the SI of the associated publication.</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

data sets of "Dynamic Linear Modeling estimates of long-term ozone trends from homogenized Dobson Umkehr profiles at Arosa, Switzerland"

<p>data sets from&nbsp;&quot;Dynamic Linear Modeling estimates of long-term ozone trends from homogenized Dobson Umkehr profiles at Arosa, Switzerland&quot;</p> <p>Monthly means&nbsp;ozone profiles data sets of MCH homogenized Dobson D051 and of Brewer B040 used in the article entitled: &quot;Dynamic Linear Modeling estimates of long-term ozone trends from homogenized Dobson Umkehr profiles at Arosa, Switzerland&quot;&nbsp;by Eliane&nbsp;Maillard Barras, Alexander Haefele, Ren&eacute; St&uuml;bi, Achille Jouberton, Herbert Schill, Irina Petropavlovskikh, Koji Miyagawa, Martin Stanek, and Lucien Froidevaux.</p> <p><a href="https://doi.org/10.5194/acp-2022-344">https://doi.org/10.5194/acp-2022-344</a></p>

opencc-by-4.0Oct 2022View details →
dryad40/100

Data: Applying stochastic and Bayesian integral projection modeling to amphibian population viability analysis

<p>Integral projection models (IPMs) can estimate the population dynamics of species for which both discrete life stages and continuous variables influence demographic rates. Stochastic IPMs for imperiled species, in turn, can facilitate population viability analyses (PVAs) to guide conservation decision-making. Biphasic amphibians are globally distributed, often highly imperiled, and ecologically well-suited to the IPM approach. Herein, we present the first stochastic size- and stage-structured IPM for a biphasic amphibian, the U.S. federally threatened California tiger salamander (<em>Ambystoma</em> <em>californiense</em>; CTS). This Bayesian model reveals that CTS population dynamics show the greatest elasticity to changes in juvenile and metamorph growth and that populations are likely to experience rapid growth at low density. We integrated this IPM with climatic drivers of CTS demography to develop a PVA and examined CTS extinction risk under the primary threats of habitat loss and climate change. The PVA indicates that long-term viability is possible with surprisingly high (20–50%) terrestrial mortality, but simultaneously identified likely minimum terrestrial buffer requirements of 600–1000 m while accounting for numerous parameter uncertainties through the Bayesian framework. These analyses underscore the value of stochastic and Bayesian IPMs for understanding both climate-dependent taxa and those with cryptic life histories (e.g., biphasic amphibians) in service of ecological discovery and biodiversity conservation. In addition to providing guidance for CTS recovery, the contributed IPM and PVA supply a framework for applying these tools to investigations of ecologically-similar species.</p>

opencc-zeroOct 2022View details →
zenodo40/100

Data supplement to "From Grains to Plastics: Modeling Nourishment Patterns and Hydraulic Sorting of Fluvially Transported Materials in Deltas"

<p>Supplementary data and codes for&nbsp;&quot;From Grains to Plastics: Modeling Nourishment Patterns and Hydraulic Sorting of Fluvially Transported Materials in Deltas&quot;. Zipped files contain ANUGA hydrodynamic model outputs, dorado particle-routing simulation outputs, Python scripts for running additional dorado simulations, and other metadata used in the analysis of dorado outputs.&nbsp;See README for additional details about directory contents.&nbsp;Note that this directory does not contain the model software itself, which is available on GitHub and has been archived elsewhere&nbsp;(relevant links can be found in README).</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Supporting data sets for "Estimating Carbon Fixation of Plant Organs for Afforestation Monitoring using a Process-based Ecosystem Model and Ecophysiological Parameter Optimization". (the survey of tree breast diameter and tree height in 11-year old Eucommia ulmoides plantation, values of simulation results used in figures and tables.)

<p>Supporting data sets for Miyauchi et al., Ecology and Evolution, 2019 (accepted).</p> <p>The files store:&nbsp;</p> <p>(1) The survey of tree breast diameter and tree height in <em>Eucommia ulmoides</em> plantation<em>.</em> The ring and stem analysis and dry weight&nbsp;of&nbsp;seven harvested sample trees in the plantation.</p> <p>(2) Values of&nbsp;optimization result used fig.7.</p> <p>(3) Values of prediction result used fig.8. and table 4.</p> <p>(4)&nbsp;Values of optimized parameters by optimization methods, parameter range and&nbsp;constrain.</p>

opencc-by-4.0May 2019View details →
zenodo40/100

Making the Most out of a Hydrological Model Dataset: Sensitivity Analyses to Open the Model Black-Box (data and code)

<p>This is "data and code" repository for the Water Resources Research Article 2017WR020401 by Borgonovo et al. (2017): "Making the most out of a hydrological model data set: Sensitivity analyses to open the model black-box". Each sub-directory contains the Matlab or R scripts to reproduce all paper plots. </p> <p>Note, that the data of this repository (i.e. under ./data_input ) are identical to the data analysed by Rakovec et al. (2014).</p> <p>References:</p> <ul> <li>Borgonovo, E., Lu, X., Plischke, E., Rakovec, O. and Hill, M. C. (2017), Making the most out of a hydrological model data set: Sensitivity analyses to open the model black-box. Water Resour. Res.. Accepted Author Manuscript. doi:10.1002/2017WR020767</li> <li>Rakovec, O., M. C. Hill, M. P. Clark, A. H. Weerts, A. J. Teuling, and R. Uijlenhoet (2014), Distributed Evaluation of Local Sensitivity Analysis (DELSA), with application to hydrologic models, Water Resour. Res., 50, 409–426, doi:10.1002/2013WR014063.</li> </ul>

opencc-by-4.0Sep 2017View details →
zenodo40/100

Supplementary material 1 from: Motloung R, Robertson M, Rouget M, Wilson J (2014) Forestry trial data can be used to evaluate climate-based species distribution models in predicting tree invasions. NeoBiota 20: 31-48. https://doi.org/10.3897/neobiota.20.5778

Current and potential distributions of sixteen species that are not widespread in southern Africa arranged on the basis of their suitable range size : a) Acacia paradoxa, b) A. cultriformis, c) A. falciformis, d) A. pendula, e) A. rubida, f) A. stricta, g) A. retinodes, h) A. fimbriata, i) A. aneura, j) A. viscidula, k) A. acuminata, l) A. adunca, m) A. binervata, n) A. schinoides, o) A. prominens, p) A. mangium. The grey shading indicates areas that SDMs have identified as suitable by SDMs while the white ones are unsuitable.

opencc-by-4.0Jan 2014View details →
zenodo40/100

Supplementary data: Modelling of future changes in seasonal snowpack and impacts on summer low flows in Alpine catchments

<p>The files in this record represent&nbsp;supplementary data for the article titled &ldquo;Modelling of future changes in seasonal snowpack and impacts on summer low flows in Alpine catchments&rdquo; in&nbsp;Water Resources Research. The files&nbsp;contain&nbsp;simulations of the HBV rainfall-runoff model for 14 alpine catchments in Switzerland. The model simulated different water balance components (such as runoff, snow water equivalent and evapotranspiration) for the reference period 1980-2009 and the three scenario periods (2020-2049, 2045-2074 and 2070-2099) using the A1B emission scenario.</p>

opencc-by-4.0Oct 2017View details →
zenodo40/100

Modelling the influence of parental effects on gene network evolution : Data and Program

<p>The C++ software &quot;Simevolv&quot; is a population genetics tool able to simulate the evolution of complex genetic architectures.</p> <p>This version is a development version from which the results of the manuscript &quot;Modelling the influence of parental effects on gene network evolution&quot; (in prep) have been obtained.</p> <p>The &quot;Data&quot; file corresponds to the simulation results used in the above mentioned article.</p>

opengpl-2.0Jun 2017View details →
zenodo40/100

Laboratory data - physical modelling of gravel bed rivers under unsteady flow conditions

<p>Sediment transport and topography data from laboratory experiments under unsteady flow conditions.</p> <p>Data supporting the manuscript:<br> Redolfi, M., Bertoldi, W., Tubino, M., &amp; Welber, M. (2018). Bed load variability and morphology of gravel bed rivers<br> subject to unsteady flow: A laboratory investigation. Water Resources Research, 54, 842&ndash;862. https://doi.org/<br> 10.1002/2017WR021143<br> Details about the physical model and the experimental procedure can be found in the paper.</p>

opencc-by-4.0Dec 2017View details →
zenodo40/100

Data format figures-DATA MINING LEARNING MODELS AND ALGORITHMS ON A SCADA SYSTEM DATA REPOSITORY

<p>The original data set included noisy, missing and inconsistent data. Data<br> preprocessing improved the quality of the data and facilitated e&plusmn;cient data<br> mining tasks.<br> Before the experiment, we prepared data suitable to next operation as<br> following steps:<br> &sup2; Delete or replace missing values;<br> &sup2; Delete redundant properties (columns);<br> &sup2; Data Transformation;<br> &sup2; Data Discretization;<br> &sup2; Export data to a required .ar&reg; or .csv format &macr;le [11].<br> The original and modi&macr;ed formats of data set are shown in Figure 1 and<br> Figure 2.<br> Data visualization is also a very useful technique because it helps to deter-<br> mine the di&plusmn;culty of the learning problem. We visualized with Weka single<br> attributes (1-d) and pairs of attributes (2-d). The &macr;gure 3 shows the variation<br> of the temperature in time.</p>

opencc-by-4.0Jun 2010View details →
zenodo40/100

Figure 3. Data visualization-DATA MINING LEARNING MODELS AND ALGORITHMS ON A SCADA SYSTEM DATA REPOSITORY

<p>Data visualization is also a very useful technique because it helps to deter-<br> mine the di&plusmn;culty of the learning problem. We visualized with Weka single<br> attributes (1-d) and pairs of attributes (2-d). The &macr;gure 3 shows the variation<br> of the temperature in time.</p>

opencc-by-4.0Jun 2010View details →
zenodo40/100

Figure 4: The estimated impedance within the measured frequency range for the symmetric (*) and the asymmetric (o) case against averaged data from healthy subjects-THE RESPIRATORY IMPEDANCE IN AN ASYMMETRIC MODEL OF THE LUNG STRUCTURE

<p>It is significant to observe that in the frequency interval of clinical interest,<br> ! 2 [25; 300] rad/s, the two impedances tend to behave similarly. For the<br> asymmetric case, we have a decrease of about -10dB/dec and a phase of ap-<br> proximately &iexcl;50o, resulting in a fractional order of n &raquo;=</p> <p>This observation suggests that a combined efect of more than one fractal order is present in the<br> lungs and that it leads naturally to values closer to measured data in the low<br> frequency range. In other words, the symmetric tree representation does not<br> suffice to obtain a good&nbsp; fit between the model and the measured impedance<br> data. Another observation is that the constant-phase behavior is emphasized<br> at frequencies below those evaluated standardly in clinical practice, i.e. below<br> 5Hz. However, in the standard clinical range of frequencies for the forced oscil-<br> lation technique, namely 4-48Hz, both symmetric and asymmetric tree models<br> give similar results, as depicted in &macr;gure 4</p>

opencc-by-4.0Oct 2010View details →
zenodo40/100

Data for 'VespaG: Expert-guided protein language models enable accurate and blazingly fast fitness prediction'

<div>Datasets used for development of VespaG and VespaG predictions generated with <a href="https://github.com/JSchlensok/VespaG">https://github.com/JSchlensok/VespaG</a>.&nbsp;</div> <div>&nbsp;</div> <div>Uploads contain:</div> <div> <ol> <li><strong>Performance</strong> summaries for ProteinGym [1]:<br>- Spearman and Pearson correlation for VespaG:&nbsp;<em>proteingym_performance_vespag.csv&nbsp;</em>(columns: 'DMS_id', 'Spearman', 'Pearson')<br>- Spearman correlation for evaluated methods VespaG, GEMME [2], VESPA [3], TranceptEVE [4], AlphaMissense [5], PoET [6]: <em>proteingym_spearman_allmethods.csv&nbsp;</em>(columns: 'DMS_id', 'Trancept EVE-L', 'VESPA', 'VespaG', 'GEMME', 'AlphaMissense', 'PoET', 'UniProt_ID', 'coarse_selection_type' (function), 'taxon')</li> <li><strong>Fasta</strong> files with sequences for all train sets (<em>vespag_fasta_training_datasets.zip</em> with seq_all9k.fasta, seq_human5k.fasta, seq_droso4k.fasta, seq_ecoli2k.fasta, seq_virus1k.fasta) and test set (<em>proteingym_217.fasta</em>)</li> <li><strong>VespaG</strong> <strong>Predictions</strong> for test set:&nbsp;<em>vespag_proteingym_rawpreds_by_training_dataset.zip</em> with raw_preds_ecoli.csv, raw_preds_human.csv, raw_preds_virus.csv, raw_preds_all.csv, raw_preds_droso.csv (columns: 'DMS_id', 'mutation', 'DMS_score', 'VespaG'). Predictions are based on different training data, the final model VespaG was trained on a subset of the human proteome and <strong>raw VespaG predictions</strong> <strong>for</strong> <strong>the</strong> <strong>ProteinGym benchmark are in&nbsp;raw_preds_human.csv </strong>(used to calculate the performances above).</li> <li><strong>GEMME predictions</strong> for train sets:&nbsp;<em>vespag_proteingym_rawpreds_by_training_dataset.zip&nbsp;</em>with folders 'human', 'droso', 'ecoli', 'virus', 'all' for respective fasta file (each containing GEMME mutational landscape output files named '<em>ID' + '</em>_normPred_evolCombi.txt')</li> <li><strong>ESM-2</strong> <strong>embeddings</strong> [7] for test set (<em>proteingym_217_esm2.h5</em>)</li> </ol> </div> <div>For details on VespaG see:</div> <div> <div> <div>VespaG: Expert-guided protein Language Models enable accurate and blazingly fast fitness prediction</div> </div> <div>Celine Marquet, Julius Schlensok, Marina Abakarova, Burkhard Rost, Elodie Laine</div> <div>bioRxiv 2024.04.24.590982; doi: https://doi.org/10.1101/2024.04.24.590982</div> <div>&nbsp;</div> <div>For more information on data usage and generation please see&nbsp;<a href="https://github.com/JSchlensok/VespaG">https://github.com/JSchlensok/VespaG</a>.</div> <div>&nbsp;</div> <div>Abstract:</div> <div>Exhaustive experimental annotation of the effect of all known protein variants remains daunting and expensive, stressing the need for scalable effect predictions. We introduce VespaG, a blazingly fast single amino acid variant effect predictor, leveraging embeddings of protein Language Models as input to a minimal deep learning model. To overcome the sparsity of experimental training data, we created a dataset of 39 million single amino acid variants from the human proteome applying the multiple sequence alignment-based effect predictor GEMME as a pseudo standard-of-truth. Assessed against the ProteinGym Substitution Benchmark (217 multiplex assays of variant effect with 2.5 million variants), VespaG achieved a mean Spearman correlation of 0.48 +/- 0.01, matching state-of-the-art methods such as GEMME, TranceptEVE, PoET, AlphaMissense, and VESPA. VespaG reached its top-level performance several orders of magnitude faster, predicting all mutational landscapes of the human proteome in 30 minutes on a consumer laptop (12-core CPU, 16 GB RAM).</div> <div>&nbsp;</div> <div>[1] Notin, Pascal, et al. "ProteinGym: large-scale benchmarks for protein fitness prediction and design." <em>Advances in Neural Information Processing Systems</em> 36 (2024).<br>[2] Laine, Elodie, Yasaman Karami, and Alessandra Carbone. "GEMME: a simple and fast global epistatic model predicting mutational effects." <em>Molecular biology and evolution</em> 36.11 (2019): 2604-2619.</div> <div>[3] Marquet, C&eacute;line, et al. "Embeddings from protein language models predict conservation and variant effects." <em>Human genetics</em> 141.10 (2022): 1629-1647.</div> <div>[4] Notin, Pascal, et al. "TranceptEVE: Combining family-specific and family-agnostic models of protein sequences for improved fitness prediction." <em>bioRxiv</em> (2022): 2022-12.</div> <div>[5] Cheng, Jun, et al. "Accurate proteome-wide missense variant effect prediction with AlphaMissense." <em>Science</em> 381.6664 (2023): eadg7492.</div> <div>[6] Truong Jr, Timothy, and Tristan Bepler. "PoET: A generative model of protein families as sequences-of-sequences." <em>Advances in Neural Information Processing Systems</em> 36 (2024).</div> <div>[7] Lin, Zeming, et al. "Evolutionary-scale prediction of atomic-level protein structure with a language model." <em>Science</em>379.6637 (2023): 1123-1130.</div> </div>

opencc-by-4.0Apr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record