Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

159

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

159 results for “code prediction”

Learn how ShareScore rates datasets ↗
dryad40/100

Data and code from: Evaluating genomic offset predictions in a forest tree with high population genetic structure

Open the record for dataset details and reuse information.

publicApr 2025View details →
dryad40/100

Data and code for: Veterinary Expert System for Outcome (VESOP) Prediction

Open the record for dataset details and reuse information.

publicAug 2023View details →
dryad40/100

Data from: Cortico-Fugal regulation of predictive coding

Open the record for dataset details and reuse information.

publicApr 2022View details →
dryad40/100

Data and R code from: Spatiotemporal risk factors predict landscape-scale survivorship for a northern ungulate

Open the record for dataset details and reuse information.

publicAug 2022View details →
dryad40/100

Data and code from: Predicting the fundamental thermal niche of ectotherms

Open the record for dataset details and reuse information.

publicMar 2024View details →
dryad40/100

Data and code for: Spatial cell type enrichment predicts mouse brain connectivity

Open the record for dataset details and reuse information.

publicSep 2023View details →
dryad40/100

Code and data for Bayesian joint species distribution model selection for community-level prediction

Open the record for dataset details and reuse information.

publicNov 2023View details →
dryad40/100

Source code for StrVCTVRE: a supervised learning method to predict the pathogenicity of human genome structural variants

Open the record for dataset details and reuse information.

publicOct 2021View details →
zenodo36/100

Replication code and data for: "Machine Learning Predicts Large Scale Declines in Native Plant Phylogenetic Diversity."

<p>Replication code and data for the paper: &quot;Machine Learning Predicts Large Scale Declines in Native Plant Phylogenetic Diversity.&quot; The following files are included in this repository:</p> <p>1) R scripts (numbered 0 through 9) include replication code for data analysis</p> <p>2) Datasets (6 zip folders) contain the data analyzed in&nbsp;the R scripts</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2020View details →
zenodo36/100

Data and code for training and evaluating machine learning models for thunderstorm prediction from reanalysis data

<p>FIXED Data and Python code for training and evaluating machine learning models for predicting thunderstorms, associated with the paper:</p> <p>&quot;Evaluation of machine learning classifiers for predicting deep convection&quot;</p> <p>by Peter Ukkonen and Antti M&auml;kel&auml;&nbsp;(to appear in JAMES)</p> <p>The data (preprocessed inputs and outputs)&nbsp;is stored as netCDF files and .mat files which can be loaded with Python.&nbsp;</p>

opencc-by-4.0Nov 2018View details →
zenodo36/100

Dataset and source code for ICSME2017 paper "Supervised vs Unsupervised Models: A Holistic Look at Effort-Aware Just-in-Time Defect Prediction"

<p>Dataset and source code for ICSME2017 paper “Supervised vs Unsupervised Models: A Holistic Look at Effort-Aware Just-in-Time Defect Prediction”</p> <p>There are four different models in the paper (i.e., EALR, LT, CBS and OneWay). Each model was implemented in a single Java file in the model package. To reproduce the experiment results of each model in the paper, just run the main method in the corresponding Java file. </p> <p> </p>

opencc-by-4.0Jul 2017View details →
zenodo36/100

Data and Code for Yeager et al., 2022: Enhanced Skill and Signal-to-noise in an Eddy-Resolving Decadal Prediction System

The sensitivity of decadal prediction system performance to model resolution is examined by comparing results from low- and high-resolution (LR and HR) predictions conducted with the Community Earth System Model (CESM). The primary difference between the two systems is the horizontal grid spacing of the ocean and atmosphere models (1° for both in LR; 0.1° and 0.25°, respectively, in HR), permitting a first direct comparison of how skill and signal-to-noise characteristics change when moving to the ocean eddy-resolved modeling regime. HR exhibits significantly increased skill and enhanced signal-to-noise for atmospheric fields compared to LR. This result suggests that mesoscale atmosphere-ocean interaction, which is present in HR but absent in LR, is a key mechanism involved in the transmission of predictable signals from the ocean to the atmosphere. Climate predictions can potentially be improved (and the signal-to-noise paradox alleviated) through explicit representation of ocean eddies and their interactions with the atmosphere.

opencc-by-4.0Dec 2021View details →
zenodo36/100

Data and codes for Landslide hazard spatiotemporal prediction based on data-driven models: Estimating where, when and how large landslide may be

<p>Data and codes for Landslide hazard spatiotemporal prediction based on data-driven models: Estimating where, when and how large landslide may be</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Spatiotemporal brain hierarchies of auditory memory recognition and predictive coding - Nature Communications

<p>The full analysis pipeline used in this study is available at the following link:&nbsp;</p> <p>https://doi.org/10.5281/zenodo.11072410</p> <p>Additional in-house-built code and functions used in this study are part of the LBPD repository which is available at the following link:&nbsp;</p> <p>https://doi.org/10.5281/zenodo.10701724</p> <p>&nbsp;</p> <p>Abstract</p> <p>Our brain is constantly extracting, predicting, and recognising key spatiotemporal features of the physical world in order to survive. While neural processing of visuospatial patterns has been extensively studied, the hierarchical brain mechanisms underlying conscious recognition of auditory sequences and the associated prediction errors remain elusive. Using magnetoencephalography (MEG), we describe the brain functioning of 83 participants during recognition of previously memorised musical sequences and systematic variations. The results show feedforward connections originating from auditory cortices, and extending to the hippocampus, anterior cingulate gyrus, and medial cingulate gyrus. Simultaneously, we observe backward connections operating in the opposite direction. Throughout the sequences, the hippocampus and cingulate gyrus maintain the same hierarchical level, except for the final tone, where the cingulate gyrus assumes the top position within the hierarchy. The evoked responses of memorised sequences and variations engage the same hierarchical brain network but systematically differ in terms of temporal dynamics, strength, and polarity. Furthermore, induced-response analysis shows that alpha and beta power is stronger for the variations, while gamma power is enhanced for the memorised sequences. This study expands on the predictive coding theory by providing quantitative evidence of hierarchical brain mechanisms during conscious memory and predictive processing of auditory sequences.</p> <p>&nbsp;</p> <p>The data is provided after pre-processing (Maxfilter, ICA for removing eye blink and heart beat, co-registration with the individual MRI T1) and epoching.</p> <p>Source data files are provided for Supplementary Figures and Tables (referred to as Supplementary Data in the paper).</p> <p>Supplementary Figures are provided in PDF format in high resolution.</p>

opencc-by-4.0Feb 2024View details →
zenodo36/100

Code and data repository for the role of topography, geotechnical layering, and attenuation on ground motion prediction

<p>Data and scripts to reproduce research on the influence of topography, geotechnical layer, and attenuation on ground motion prediction in Salton Trough.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Data and code of "Germline mutation rate predicts cancer mortality across 37 vertebrate species" article

<p>Data and code used in this article.</p> <p>The "Data.csv" file contains information about each species class, common name, average yearly mutation rate, trophic level, number of animals per species that died from cancer, records of mortality per species, the minimum confidence interval (CI) for the cancer mortality values, the maximum CI for the cancer mortality values, the standard error (SE) of the records of mortality per species, cancer mortality, average parental age (measured in months), maximum lifespan (measured in months), average parental age (measured in months), and whether a species is reported as domesticated/semidomesticated or not. The way the authors obtained these data, and the original sources of these data, are described in the methodology section of the article.</p> <p>The "Regression_analyses.R" file is the code we used in the regression analyses conducted in this article. The .phy file is the phylogenetic tree used in the regression analyses.</p> <p>The .txt, .nwk, .Rproj, and "germFit.R" files are the data and code we used for comparing the different models of phenotype evolution (comparison of 3 different models: Ornstein&ndash;Uhlenbeck, Brownian, and Early Burst).</p>

opencc-by-4.0Feb 2024View details →
zenodo36/100

Code and data from: Towards a more dynamic metabolic theory of ecology to predict climate change effects on biological systems

<p>This repository contains data and code used to produce Figures 1, 2, and S1 in <em>Towards a more dynamic metabolic theory of ecology to predict climate change effects on biological systems.</em></p> <p>The file "empirical_mte_database.csv" contains a database of peer-reviewed articles pulled from a Web of Science search of papers that empirically tested the temperature dependence predictions of the metabolic theory of ecology (Brown et al. 2004) between 2004 and 2024.&nbsp;</p> <p>The .zip file contains a jupyter notebook, julia project toml file and a README in order to reproduce the simulations illustrating how different temperature dynamics can lead to different inferred thermal performance curves in populations.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Datasets and Code: Revealing the Role of Redox Reaction Selectivity and Mass Transfer in Current–Voltage Predictions for Ensembles of Photocatalysts

<ul> <li>Raw datasets (.mat and .fig files) and codes (.mlx and .m files) used in our manuscript of the same title.&nbsp;</li> <li>Figure numbers correspond with the figure numbers in the corresponding&nbsp; manuscript. <ul> <li>Figure 4: Effects of kinetic parameters on&nbsp;Solar-to-chemical (STC) efficiencies and reaction selectivity</li> <li>Figure 5: Solar-to-chemical (STC) efficiencies for a model incorporating competing undesired redox reactions implemented for different redox shuttle pairs</li> <li>Figure 7: Solar-to-chemical efficiencies for an ensemble of light absorbers</li> <li>Figure 8: Maximum solar-to-chemical (STC) efficiencies and corresponding number of light absorbers as a function of asymmetry factors in limiting current density for redox shuttle reduction</li> <li>Figure 9: Solar-to-chemical efficiencies for an increasing number of light absorbers for different total absorptance values (99%, 75%, 50%).</li> <li>Figure 10: Qualitative comparisons between experimental measurements and model predictions for a photocatalytic suspension reactor</li> </ul> </li> <li>The main piece of the code developed is provided as an interactive .mlx file; not all subfunction calls within the main code is included, and can be shared upon reasonable request via email from the lead (luisab@umich.edu) and the corresponding authors (rbchan@umich.edu) of this paper.&nbsp;</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Yields, Cannabinoids Quantification, and Predictive Programming Codes Using Machine Learning for Non-Psychoactive Cannabis Flowers and Extracts (Cannabis sativa L.) Cultivated in Ecuador.

<p>This publication presents data from various extraction methods, including maceration, Soxhlet, and supercritical fluids, performed on different cannabis flower varieties (Cannabis sativa L.) under varying operating conditions. We quantified the amounts of CBD, THC, CBG, and CBN in the extracts produced by each method using High-Performance Liquid Chromatography (HPLC). Using this data, we developed a machine learning algorithm in RStudio to make predictions and determine the best conditions and yields for each extraction method. The analysis focuses on different varieties of non-psychoactive cannabis cultivated in Ecuador at altitudes over 2,450 m.a.s.l.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code

<p>Despite their success, large language models (LLMs) face the critical challenge of hallucinations, generating plausible but incorrect content. While much research has focused on hallucinations in multiple modalities including images and natural language text, less attention has been given to hallucinations in source code, which leads to incorrect and vulnerable code that causes significant financial loss. To pave the way for research in LLMs' hallucinations in code, we introduce Collu-Bench, a benchmark for predicting code hallucinations of LLMs across code generation (CG) and automated program repair (APR) tasks. Collu-Bench includes 13,234 code hallucination instances collected from five datasets and 11 diverse LLMs, ranging from open-source models to commercial ones.&nbsp;<br>To better understand and predict code hallucinations, Collu-Bench provides detailed features such as the per-step log probabilities of LLMs' output, token types, and the execution feedback of LLMs' generated code for in-depth analysis. In addition, we conduct experiments to predict hallucination on Collu-Bench, using both traditional machine learning techniques and neural networks, which achieves 22.03 -- 33.15% accuracy.&nbsp;Our experiments draw insightful findings of code hallucination patterns, reveal the challenge of accurately localizing LLMs' hallucinations, and highlight the need for more sophisticated techniques.</p>

opencc-by-4.0Oct 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record