Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

159

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

159 results for “code prediction”

Learn how ShareScore rates datasets ↗
zenodo40/100

Data and code of "Post-trauma behavioral phenotype predicts the degree of vulnerability to fear relapse after extinction in male rats"

<p>This dataset contains&nbsp;behavioral and&nbsp;transcriptomic data, and the original code related to the following article:</p> <p>Post-trauma behavioral phenotype predicts the degree of vulnerability to fear relapse after extinction in male rats. Fanny Demars, Ralitsa Todorova, Gabriel Makdah, Antonin Forestier, Marie-Odile Krebs, Bill P Godsil, Th&eacute;r&egrave;se M Jay, Sidney I Wiener, &amp; Marco N Pompili&nbsp;(2022) Current Biology <em>32.&nbsp;https://doi.org/10.1016/j.cub.2022.05.050</em></p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Code and data for "Global warming generates predictable extinctions of warm- and cold-water marine benthic invertebrates via thermal habitat loss"

<pre>This repository contains the following information: Datasets S1 to S4 can all be loaded, manipulated, and analysed in R using script provided in Data S5 to obtain the results of the paper, Reddin et al. 2022, &quot;Global warming generates predictable extinctions of warm and cold-water marine benthic invertebrates via thermal habitat loss&quot;. Data S1. (separate file) The original downloaded PaleoDB dataset. Data S2. (separate file) The pre-prepared dataset of occurrences. Data S3. (separate file) The finished environmental dataset. Data S4. (separate file) Additional environmental dataset. Data S5. (separate file) The R-code for the main analysis. Data S6. (compressed directory) Output data and code from the simulations. Table S7 (separate file). List of data source publications for PaleoDB data used in our study. Listed are the data source author list (ref_author), year (ref_pubyr), and reference number as appears in the PaleoDB (reference_no). </pre>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Codes for Purgar et al. 2022: Investigating the ability of growth models to predict in situ Vibrio spp. abundances

<p>Model simulations and analysis of the Vibrio spp. growth models.&nbsp;<br> Prepared to accompany the publication, Purgar et al. 2022 &quot;Investigating the ability of growth models to predict in situ<br> Vibrio spp. abundances&quot; in Microorganisms, Special Issue &bdquo;Microbial Communities in Changing Aquatic Environments&ldquo;.&nbsp;</p> <p>Description of the files&nbsp;can be found in the Readme file.</p>

opencc-by-4.0Aug 2022View details →
dryad40/100

Data and R code from: Spatiotemporal risk factors predict landscape-scale survivorship for a northern ungulate

<p>These data and computer code (written in R, https://www.r-project.org) were created to statistically evaluate a suite of spatiotemporal covariates that could potentially explain pronghorn (Antilocapra americana) mortality risk in the Northern Sagebrush Steppe (NSS) ecosystem (50.0757<sup>o</sup> N, −108.7526<sup>o</sup> W). Known-fate data were collected from 170 adult female pronghorn monitored with GPS collars from 2003-2011, which were used to construct a time-to-event (TTE) dataset with a daily timescale and an annual recurrent origin of 11 November. Seasonal risk periods (winter, spring, summer, autumn) were defined by median migration dates of collared pronghorn. We linked this TTE dataset with spatiotemporal covariates that were extracted and collated from pronghorn seasonal activity areas (estimated using 95% minimum convex polygons) to form a final dataset. Specifically, average fence and road densities (km/km2), average snow water equivalent (SWE; kg/m2), and maximum decadal normalized difference vegetation index (NDVI) were considered as predictors. We tested for these main effects of spatiotemporal risk covariates as well as the hypotheses that pronghorn mortality risk from roads or fences could be intensified during severe winter weather (i.e., interactions: SWE*road density and SWE*fence density). We also compare an analogous frequentist implementation to estimate model-averaged risk coefficients. Ultimately, the study aimed to develop the first broad-scale, spatially explicit map of predicted annual pronghorn survivorship based on anthropogenic features and environmental gradients to identify areas for conservation and habitat restoration efforts.</p> <p> </p>

opencc-zeroAug 2022View details →
zenodo40/100

Supplementary data and code: Human–Wildlife Interactions Predict Febrile Illness in Park Landscapes of Western Uganda

<p>Data at code used for analyses.</p> <p>Salerno J, Ross N, Ghai R, et al (2017) Human–Wildlife Interactions Predict Febrile Illness in Park Landscapes of Western Uganda. EcoHealth. doi: 10.1007/s10393-017-1286-1<br>  </p>

opencc-by-4.0Nov 2017View details →
zenodo40/100

Results and code associated with «Predictability of ecological and evolutionary dynamics in a changing world»

<p>Below you'll find results and code associated with the following article (*):</p> <p>Bozzuto, C, Ives, AR (2024): Predictability of ecological and evolutionary dynamics in a changing world. <em>Proceedings of the Royal Society B</em>, <strong>291</strong>:&nbsp;20240980. https://doi.org/10.1098/rspb.2024.0980</p> <p>ABSTRACT: Ecological and evolutionary predictions are being increasingly employed to inform decision-makers confronted with intensifying pressures on biodiversity. For these efforts to effectively guide conservation actions, knowing the limit of predictability is pivotal. In this study, we provide realistic expectations for the enterprise of predicting changes in ecological and evolutionary observations through time. We begin with an intuitive explanation of predictability (the extent to which predictions are possible) employing an easy-to-use metric, predictive power <em>PP</em>(<em>t</em>). To illustrate the challenge of forecasting, we then show that among insects, birds, fishes and mammals, (i) 50% of the populations are predictable at most 1 year in advance and (ii) the median 1-year-ahead predictive power corresponds to a prediction <em>R</em><sup>2</sup> of only 20%. Predictability is not an immutable property of ecological systems. For example, different harvesting strategies can impact the predictability of exploited populations to varying degrees. Moreover, incorporating explanatory variables, accounting for time trends and considering multivariate time series can enhance predictability. To effectively address the challenge of biodiversity loss, researchers and practitioners must be aware of the information within the available data that can be used for prediction and explore efficient ways to leverage this knowledge for environmental stewardship.</p> <p>(*) previously a preprint on <em>bioRxiv</em>: https://doi.org/10.1101/2023.11.01.565089</p>

opencc-by-4.0Nov 2023View details →
dryad40/100

Data and code from: Evaluating genomic offset predictions in a forest tree with high population genetic structure

<p>Predicting how tree populations will respond to climate change is an urgent societal concern. An increasingly popular way to make such predictions is the genomic offset (GO) approach, which aims to use genomic and climate data to identify populations that may experience climate maladaptation in the near future. More precisely, GO tries to represent the change in allele frequencies required to maintain the current gene-climate relationships under climate change. However, the GO approach has major limitations and, despite promising validation of its predictions using height data from common gardens, it still lacks broad empirical testing. In the present study, we evaluated the consistency and empirical validity of GO predictions in maritime pine (<em>Pinus pinaster</em> Ait.), a tree species from southwestern Europe and North Africa with a marked population genetic structure. First, gene-climate relationships were estimated using 9,817 SNPs genotyped in 454 trees from 34 populations; and candidate SNPs potentially involved in climate adaptation were identified. Second, GO was predicted using four methods, namely Gradient Forest (GF), Redundancy Analysis (RDA), latent factor mixed model (LFMM) and Generalised Dissimilarity Modeling (GDM), two sets of SNPs (candidate and control SNPs) and five climate general circulation models (GCMs) to account for uncertainty in future climate predictions. Last, the empirical validity of GO predictions was evaluated within a Bayesian framework by estimating the associations between GO predictions and two independent data sources: mortality data from National Forest Inventories (NFI), and mortality and height data from five common gardens in contrasting environments. We found high variability in GO predictions across methods, SNP sets and GCMs. Regarding validation, GO predictions with GDM and GF (and to a lesser extent RDA) based on the candidate SNPs showed the strongest and most consistent associations with mortality rates in common gardens and NFI plots. We found almost no association between GO predictions and tree height in common gardens, most likely due to the overwhelming effect of population genetic structure on tree height in this species. Our study demonstrates the imperative to validate GO predictions with a range of independent data sources before they can be used as informative and reliable metrics in conservation or management strategies.</p>

opencc-zeroMay 2024View details →
zenodo40/100

Data and code for: Predicting rapid adaptation in time from adaptation in space: a 30-year field experiment in marine snails

<p>Scripts and data used in the research study <strong>Predicting rapid adaptation in time from adaptation in space: a 30-year field experiment in marine snails</strong>.</p> <p><a href="https://doi.org/10.1101/2023.09.27.559715" target="_blank" rel="noopener">https://doi.org/10.1101/2023.09.27.559715</a></p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

SLDeep: Statement-Level Software Defect Prediction Using Deep-Learning Models on Static Code Features

<p>Software defect prediction (SDP) seeks to estimate fault-prone areas of the code to focus testing activities on more suspicious portions. Consequently, high-quality software is released with less time and effort. The current SDP techniques however work at coarse-grained units, such as a module or a class, putting some burden on the developers to locate the fault. To address this issue, we propose Statement-Level software defect prediction using Deep-learning model (SLDeep). To reify our proposal, we defined a suite of 32 statement-level metrics, such as the number of binary and unary operators used in a statement. Then, we applied as learning model, long short-term memory (LSTM). The significance of SLDeep for intelligent and expert systems is that it demonstrates a novel use of deep-learning models to the solution of a practical problem faced by software developers. We conducted experiments using more than 100,000 C/C++ programs within the Code4Bench. The programs total 2,356,458 lines of code with 292,064 faulty lines. The benchmark comprises diverse set of programs and versions, written by thousands of developers. Therefore, it tends to give a model that can be used for cross-project SDP. In the experiments, our trained model could successfully classify the unseen data with average performance measures 0.945, 0.971, and 0.976 in terms of recall, precision, and accuracy, respectively. These experimental results suggest that SLDeep is effective for statement-level SDP. The impact of this work is twofold. Working at statement-level further alleviates developer&rsquo;s burden in pinpointing the fault locations. Second, cross-project feature of SLDeep helps defect prediction research become more industrially-viable</p> <p>for more information visit&nbsp;<a href="https://github.com/sldeep/SLDeep">https://github.com/sldeep/SLDeep</a></p>

opencc-by-4.0Jul 2019View details →
zenodo40/100

Supplementary data and code for "Predicting anthropogenic food supplementation from individual tracking data"

<p><span>This repository contains data and R-Scripts to reproduce the results presented in the paper &ldquo;</span><span>Predicting anthropogenic food supplementation from individual tracking data</span><span>&rdquo; published in Ibis. Please note that some data (basemap of Switzerland, forests, buildings) are publicly available from </span><span><a href="https://map.geo.admin.ch/"><span>https://map.geo.admin.ch/</span></a></span><span> and are therefore not provided here. The abstract of the article is as follows: </span><span>Many wildlife species consume food or refuse provided by humans. To understand the effect of anthropogenic food subsidies on wildlife populations, we first need to quantify where and when individuals can access such food sources. The red kite <em>Milvus milvus</em> is an opportunistic raptor species and uses both inadvertent and deliberate food subsidies provided by citizens. Here we present a new approach using GPS-tracking data to predict where anthropogenic food subsidies likely occur. We tracked 497 individuals with solar-powered GPS transmitters over an average of 3.2 (range 1 &ndash; 9) breeding seasons in Switzerland, and combined these data with locations of 125 known feeding sites obtained through interviews. We used two sequential random forest models, at both individual movement and population levels, to predict where anthropogenic food subsidies were attended by red kites. The first model classified locations that were frequently and regularly revisited, and successfully predicted 85% of locations that were within 50 m of an externally validated feeding site. These predicted locations were aggregated in 500 m grid cells to calculate the proportion of individuals and locations associated with predicted food subsidy. A second model related the presence of known food subsidies to the aggregated predictions. In our study area, 80% of known anthropogenic food provision locations could be correctly identified using red kite tracking data, but data sparsity beyond the core range of tracked individuals limits predictions of anthropogenic food subsidies at larger geographic scales. Nonetheless, biologging data can identify ephemeral food sources, and facilitate an assessment of the importance of anthropogenic food subsidies on the fitness of individuals in tracked populations.</span></p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

R code to accompany 'A novel initialisation technique for decadal climate predictions'

<p>R code for the diagnostic published in &#39;A novel initialisation technique for decadal climate predictions&#39; <a href="https://doi.org/10.3389/fclim.2021.681127">https://doi.org/10.3389/fclim.2021.681127</a><br> The data loaded follow the cmor standard.</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Dataset and codes of the article "Neural correlates of hierarchical predictive processes in autistic adults"

<p>Data and code related to the article&nbsp;&quot;Neural correlates of hierarchical predictive processes in autistic adults&quot;&nbsp; by Laurie-Anne Sapey-Triomphe, Lauren Pattyn, Veith Weilnhammer, Philipp Sterzer and Johan Wagemans (Nature Communications):</p> <p>-&nbsp;Behavioral dataset&nbsp;of the 26 neurotypical participants (NT_behavioral_data.zip) and of the 26 autistic participants (ASD_behavioral_data.zip)</p> <p>- Source data of the graphics appearing in the article (Source data.xls)</p> <p>- Matlab codes used to run the experiment (Codes_to_run_experiment.zip)</p> <p>- Matlab codes to perform&nbsp;the main behavioral analyses (Codes_behavioral_analyses.zip) and to analyze the behavioral data with the HGF models (Codes_comput_model_analyses.zip)</p> <p>- Matlab codes to preprocess (Codes_fMRI_preprocessing.zip) and run the main fMRI analyses (Codes_fMRI_analyses.zip)</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Training patches and prediction codes of deep learning (LANA) model for Landsat 8/9 cloud/shadow mask

<p>This dataset includes (i) the image patches dataset and (ii) application/prediction (not training) codes for Landsat 8 cloud and cloud shadow masking used in a paper in review and uploaded here: &nbsp;</p> <p>Hankui Zhang, Dong Luo, David Roy, A learning attention network algorithm (LANA) for accurate Landsat-8 cloud and shadow masking,&nbsp;<em>Remote Sensing of Environment</em>&nbsp;</p> <p>The documentation is in&nbsp;<a href="https://zenodo.org/api/files/5462baa5-2bba-4b0f-92aa-c17681b6464b/l8_training_data_readme_new.pdf?versionId=95253efb-6447-46d8-ae4b-ac0c93b43532">l8_training_data_readme_new.pdf</a>.&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

CPT-1 pre-computed whole-proteome variant effect predictions and model source code

<p><strong>Cross-protein transfer learning for variant effect prediction</strong></p><p>This repository contains the variant effect predictions of CPT-1 for 18,602 human proteins, initially released with the manuscript "Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects". The proteins are split into three files.</p><p><i>CPT1_score_EVE_set.zip</i>: Proteins in the EVE set (<a href="https://www.nature.com/articles/s41586-021-04043-8">Frazer et al., 2021</a>)</p><p><i>CPT1_score_no_EVE_set_1.zip</i> &amp; <i>CPT1_score_no_EVE_set_2.zip</i>: Proteins not in the EVE set. Predictions for these proteins use imputed values for features depending on the EVE MSA.</p><p>The protein names are UniProt gene names.</p><p>We also provide source code to train CPT-1 model and reproduce results in the manuscript :</p><p><i>source_code.zip </i>(corresponds to GitHub repository&nbsp;songlab-cal/CPT version as of Jul 12, 2023)</p><p>&nbsp;</p><p><strong>Citation</strong></p><p>Jagota, M.*, Ye, C.*, Albors, C., Rastogi, R., Koehl, A., Ioannidis, N., and Song, Y.S.†<br>"Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects", bioRxiv (2022)</p><p>*These authors contributed equally to this work.<br>†To whom correspondence should be addressed:&nbsp;<a href="mailto:yss@berkeley.edu">yss@berkeley.edu</a></p><p>DOI:&nbsp;<a href="https://doi.org/10.1101/2022.11.15.516532">https://doi.org/10.1101/2022.11.15.516532</a></p><p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Submitted forecasts and analysis code for "Predicting spring phenology in deciduous broadleaf forests: NEON Phenology Forecasting Community Challenge"

<p>Submitted forecasts for the 2021 Ecological Forecasting Initiative NEON Phenology Forecast Challenge and the analysis code for the accompanying manuscript.&nbsp;</p>

opencc-by-4.0Jul 2023View details →
dryad40/100

Data and code for: Veterinary Expert System for Outcome (VESOP) Prediction

<p>Timely detection and understanding of causes for population decline are essential for effective wildlife management and conservation. Assessing trends in population size has been the standard approach but we propose that monitoring population health could prove more effective. We collated data from seven bottlenose dolphin (<em>Tursiops</em> <em>truncatus</em>) populations in the southeastern U.S. to develop the Veterinary Expert System for Outcome Prediction (VESOP), which estimates survival probability using a suite of health measures identified by experts as indices for inflammatory, metabolic, pulmonary, and neuroendocrine systems. VESOP was implemented using logistic regression within a Bayesian analysis framework, and parameters were fit using records from five of the sites that had robust stranding network and frequent photographic identification (photo-ID) surveys to document definitive survival outcomes. We also conducted capture-mark-recapture (CMR) analyses of photo-ID data to obtain separate estimates of population survival rates for comparison with VESOP survival estimates.  VESOP analyses found multiple measures of health, particularly markers of inflammation, were predictive of 1- and 2-year individual survival. The highest mortality risk one year following health assessment related to low alkaline phosphatase, with an odds ratio of 10.2 (95% CI 3.41–26.8), while 2-year mortality was most influenced by elevated globulin (9.60; 95% CI 3.88–22.4); both are markers of inflammation. The VESOP model predicted population-level survival rates that correlated with estimated survival rates from CMR analyses for the same populations (1-year Pearson's r=0.99; p=1.52e<sup>-05</sup>, 2-year r=0.94; p=0.001). While our proposed approach will not detect acute mortality threats that are largely independent of animal health, such as harmful algal blooms, it is applicable for detecting chronic health conditions that increase mortality risk. Random sampling of the population is important and advancement in remote sampling methods could facilitate more random selection of subjects, obtainment of larger sample sizes, and extension of the approach to other wildlife species.</p>

opencc-zeroAug 2023View details →
dryad40/100

Data and code for: Spatial cell type enrichment predicts mouse brain connectivity

<p>A fundamental neuroscience topic is the link between the brain's molecular, cellular and cytoarchitectonic properties and structural connectivity (SC). Recent studies relate inter-regional connectivity to gene expression, but the relationship to regional cell-type distributions remains understudied. Here, we utilize whole-brain mapping of neuronal and non-neuronal subtypes via the Matrix Inversion and Subset Selection (MISS) algorithm to model inter-regional connectivity as a function of regional cell-type composition with machine learning. We deployed random forest algorithms for predicting connectivity from cell type densities, demonstrating surprisingly strong prediction accuracy of cell types in general and particular cells like oligodendrocytes. We found evidence of a strong distance-dependency in the cell-connectivity relationship, with layer-specific excitatory neurons contributing the most for long-range connectivity, while vascular and astroglia are salient for short-range connections. Our results demonstrate a link between cell types and connectivity, providing a roadmap for examining this relationship in other species, including humans.</p>

opencc-zeroAug 2023View details →
zenodo40/100

Data and analytical codes for: Learning beyond-pairwise interactions enables the bottom-up prediction of microbial community structure

<p>Data and analytical codes for: Ishizawa et al. (2023) Learning beyond-pairwise interactions enables the bottom-up prediction of microbial community structure, bioRxiv, 2023.07.04.546222</p> <p>&nbsp;https://www.biorxiv.org/content/10.1101/2023.07.04.546222v1</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
dryad40/100

Data and code from: A timid choice: Risk-taking behavior predicts individualized niche in a varying landscape of safety

Open the record for dataset details and reuse information.

publicMay 2025View details →
dryad40/100

Predicting arrhythmia recurrence post-ablation in atrial fibrillation using explainable machine learning: Code repository

Open the record for dataset details and reuse information.

publicJul 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record