Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
159
datasets available to search
ShareScore release 0.9.0
Dataset results
159 results for “code prediction”
Improving Build Outcome Prediction usingTextual Analysis of Source Code
<p>The uploaded data sets contain code deltas and their corresponding feature vectors. Those were derived from 13 open-source Java-based projects. </p>
Code and data sets for "DeepLC can predict retention times for peptides that carry as-yet unseen modifications"
<p>Code used to prepare the data sets, calibrate retention times, generate DeepLC models, make predictions, and generate the figures. See README.md for more information on how to use these files and reproduce the results reported in the manuscript titled "DeepLC can predict retention times for peptides that carry as-yet unseen modifications".</p>
Data and R code from: Nature calls: intelligence and natural foraging style predict welfare problems in captive parrots
<p>Around half of all parrots (a highly threatened order) live in captivity. Here, some species thrive. Others, however, breed poorly or display stereotypic behaviours indicating stress. Using data on the prevalence of three types of stereotypic behaviour in pet (50 species; 1,378 individuals) and aviculture hatch rates (115 species; 10,255 breeding pairs), we applied Phylogenetic Comparative Methods (PCMs) to test hypothesised causes of this variation (relating to species' rarity and constraints on natural behaviour). In the first empirical evidence that high intelligence increases vulnerability to poor captive welfare, species with large relative brain sizes were found to be most at risk of oral and whole-body stereotypic behaviour. This suggests that if they are to be kept in private homes, such parrots must be offered substantially more cognitive stimulation. Self-harming behaviours involving feather damage were predicted by naturally relying on food items that require substantial handling, highlighting inadequacies in captive diets (often highly processed); while relatively low hatch rates in aviculture were predicted by small captive population sizes, potentially due to genetic bottlenecks, inbreeding, and/or low availability of compatible mates. These novel findings should help advance captive parrot husbandry, and inspire further research applying PCMs to understand and improve animal welfare.</p>
Data and codes of "Odor descriptive ratings can predict some odor-color associations in different color features of hue or lightness"
<p>This project included data and codes used in our original article titled "Odor descriptive ratings can predict some odor-color associations in different color features of hue or lightness", written by Kaori Tamura and Tsuyoshi Okamoto.</p> <p>Please see https://gitlab.com/tamurak415/olfqr</p>
Data and code from: Dynamics of cortical contrast adaptation predict perception of signals in noise
<p>Neurons throughout the sensory pathway adapt their responses depending on the statistical structure of the sensory environment. Contrast gain control is a form of adaptation in the auditory cortex, but it is unclear whether the dynamics of gain control reflect efficient adaptation, and whether they shape behavioral perception. Here, we trained mice to detect a target presented in background noise shortly after a change in the contrast of the background. The observed changes in cortical gain and behavioral detection followed the dynamics of a normative model of efficient contrast gain control; specifically, target detection and sensitivity improved slowly in low contrast, but degraded rapidly in high contrast. Auditory cortex was required for this task, and cortical responses were not only similarly affected by contrast but predicted variability in behavioral performance. Combined, our results demonstrate that dynamic gain adaptation supports efficient coding in auditory cortex and predicts the perception of sounds in noise.</p>
Data and codes: Speech-recognition in landlide predictive modelling
<p>This is the data and codes for the manuscript "Speech-recognition in landlide predictive modelling"</p>
Data and code from: Life-history traits predict ability of British wild bees to fill their climate envelopes
Open the record for dataset details and reuse information.
Data and code from: Head morphology predicts prey traits and drives individual dietary specialization in generalist anurans
Open the record for dataset details and reuse information.
Data and code from: Dynamics of cortical contrast adaptation predict perception of signals in noise
Open the record for dataset details and reuse information.
Data and code from: Predicting population genetic change in an autocorrelated random environment: insights from a large automated experiment
Open the record for dataset details and reuse information.
Data and R code from: Nature calls: intelligence and natural foraging style predict welfare problems in captive parrots
Open the record for dataset details and reuse information.
Data and Code for: Food distribution, but not market forces, predict behavioral social tolerance in rhesus macaques
Open the record for dataset details and reuse information.
Predictive coding and internal error correction in speech production
Open the record for dataset details and reuse information.
Dataset and Code for "Mining and Predicting Micro-Process Patterns of Issue Resolution for Open Source Software Projects"
<p>Dataset and Code for "Mining and Predicting Micro-Process Patterns of Issue Resolution for Open Source Software Projects" with README included</p>
CodiEsp Silver Standard: Participant predictions in eHealth CLEF2020 - Spanish clinical cases coded in ICD10 (CIE10)
<p><strong>Introduction</strong></p> <p>Predictions in the background set of <a href="https://temu.bsc.es/codiesp/">eHealth CLEF 2020 Task 1</a> participants.</p> <p> </p> <p><strong>Zip structure</strong></p> <p>One directory per CodiEsp subtask. Within each CodiEsp subtask directory, there is one directory per team that contains the prediction runs.</p> <p> </p> <p><strong>Format</strong><br> The text documents are distributed in plain text files, UTF-8 encoding.<br> The CodiEsp Silver Standard annotations have the following format:</p> <p>For the sub-tracks CodiEsp-Diagnostic and CodiEsp-Procedure, the file files have the following fields:</p> <pre>articleID ICD10-code </pre> <p>Tab-separated files for the sub-track CodiEsp-X (explainability) contain extra fields that provide the text-reference and its position:</p> <pre>articleID label ICD10-code text-reference reference-position</pre> <p> </p> <p><strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/codiesp/">Web</a></strong></li> <li><strong>Citation: </strong>Miranda-Escalada, A., Farré, E., & Krallinger, M. (2020). Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results. In <em>Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2020), CEUR Workshop Proceedings</em>.</li> <li><strong><a href="https://zenodo.org/record/3837305#.X7T9KVlKg5k">Gold Standard corpus</a></strong></li> <li><strong><a href="https://doi.org/10.5281/zenodo.3878178">Annotation guidelines</a></strong></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhC24g5dsp5eVMp8BZFWCraX"><strong>YouTube presentations</strong></a></li> <li><a href="https://temu.bsc.es/cantemist/?p=4606"><strong>Participant codes</strong></a></li> </ul> <p> </p> <p>All credit to CodiEsp participants</p>
Code and data underlying "Interictal intracranial EEG for predicting surgical success: the importance of space and time"
<p>Code and data underlying our publication "Interictal intracranial EEG for predicting surgical success: the importance of space and time" in Epilepsia 2020</p>
Dynamic Load Balancing for Predictions of Storm Surge and Coastal Flooding-Model setup and source code
<p>Source code and model setup/inputs for the paper titled "Dynamic Load Balancing for Predictions of Storm Surge and Coastal Flooding" article. Simulations were conducted using a modified version of ADCIRC+DLB (ADCIRC + Dynamic Load Balancing) on unstructured triangular meshes.</p> <p>Contains:</p> <ol> <li>Model input files. <ol> <li>ADCIRC model input files for the ideal channel setup and Hurricane Irene simulation (*.13, *.14, *.15)</li> </ol> </li> <li>Zipped archive of the ADCIRC code (adcirc-cg-DLB.zip) used to produce the simulations for the paper.</li> <li>Step-by-step compilation and usage instructions for ADCIRC+DLB. <ol> <li>Installation.html </li> <li>Usage.html</li> </ol> </li> </ol>
Cantemist Silver Standard: Participant predictions in SEPLN IberLEF2020 - Spanish oncology clinical cases coded in ICD-O
<p><strong>Introduction</strong></p> <p>Predictions in the background set of Cantemist participants.</p> <p> </p> <p><strong>Zip structure</strong></p> <p>One directory per Cantemist subtask. Within each Cantemist subtask directory, there is one directory per team that contains the prediction runs.</p> <p> </p> <p><strong>Format</strong></p> <p>The text documents are distributed in plain text files, UTF-8 encoding.<br> The CodiEsp Silver Standard annotations have the following format:</p> <p>For the sub-tracks Cantemist-NER and Cantemist-Norm, the files are in Brat format.</p> <p>For the sub-track Cantemist-Coding files have the following fields:</p> <pre>articleID ICDO-code </pre> <p> </p> <p><strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/cantemist/">Web</a></strong></li> <li><strong>Citation: </strong>Miranda-Escalada, A., Farré, E., & Krallinger, M. (2020). Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results. In <em>Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2020), CEUR Workshop Proceedings</em>.</li> <li><a href="https://doi.org/10.5281/zenodo.3773228"><strong>Gold Standard corpus</strong></a></li> <li><strong><a href="https://doi.org/10.5281/zenodo.3878178">Annotation guidelines</a></strong></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhC24g5dsp5eVMp8BZFWCraX"><strong>YouTube presentations</strong></a></li> <li><a href="https://temu.bsc.es/cantemist/?p=4606"><strong>Participant codes</strong></a></li> </ul> <p> </p> <p>All credit to Cantemist participants. </p> <p> </p> <p>For more information, visit the track webpage: <a href="http://temu.bsc.es/cantemist/">http://temu.bsc.es/cantemist/</a> or email us at encargo-pln-life@bsc.es</p>
Codes and data: Community size predicts temporal β-diversity at local but not regional scales
<p>UPDATED VERSION 2025-08-30 (models were updated)</p> <p>This zip file contains the codes demonstrating how I analyzed and selected publicly available and globally extensive data on fish composition and environmental variables to test the hypothesis that random fluctuations caused by demographic stochasticity in small populations might extend to communities and metacommunities, potentially affecting stability propagation across biological levels and spatial scales. The READ_ME file contains additional details about the steps I took to develop this analysis.</p> <p>This study was financed by the Coordenação de Aperfeiçoamento de Pessoal de Nível Superior-Brasil (CAPES) - Finance Code 001.</p>
Intracranial and behavioral data from "Asymmetric coding of reward prediction errors in human insula and dorsomedial prefrontal cortex"
<p>Preprocessed intracranial EEG and behavioral data from Hoy, Quiroga-Martinez, et al. manuscript titled "Asymmetric coding of reward prediction errors in human insula and dorsomedial prefrontal cortex" published in Nature Communications (2023). Source data files for figures are included as well.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.