Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,773
datasets available to search
ShareScore release 0.7.1
Dataset results
1,773 results for “predictive modeling”
Figure 9. A model that helps the diagnosis prediction-Modern Tools in Patient-Centred Speech Therapy for Romanian Language
<p>We have already implemented many modules of Logo-DM, such as: data cleaning module, data transformation module, feature extraction module, data clustering module and a classification module for diagnosis prediction. Figure 9 shows the model achieved using a decision tree built on complex examination data that aims to predict the patient’s diagnosis. Currently, we are testing the built models on new cases in order to estimate their quality.</p>
BRAIN Journal-Prediction of Thyroid Disease Using Data Mining Techniques-Figure 2. Attributes of the classification models used in the experiments
<p>The authors used for their experiments a data set (UCI, 2016) containing 756 records about persons with thyroid dysfunctions. The classification model has 22 attributes; the class attribute is the target and it has three possible values: hypothyroidism, hyperthyroidism and normal. The current data set was extracted and preprocessed from the original file. A description of the attributes used in the experiments is given in Figure 2 (an extract from thyroid.arff test file). </p>
Distinguishing between pan assay interference compounds (PAINS) that are promiscuous or represent dark chemical matter - data set and prediction models
<p>Data sets of promiscuous PAINS (PROM_PAINS) and dark chemical matter PAINS (DCM_PAINS) are provided and support vector machine models built on the basis of original and balanced training data (see readme.txt).<br> </p>
SLDeep: Statement-Level Software Defect Prediction Using Deep-Learning Models on Static Code Features
<p>Software defect prediction (SDP) seeks to estimate fault-prone areas of the code to focus testing activities on more suspicious portions. Consequently, high-quality software is released with less time and effort. The current SDP techniques however work at coarse-grained units, such as a module or a class, putting some burden on the developers to locate the fault. To address this issue, we propose Statement-Level software defect prediction using Deep-learning model (SLDeep). To reify our proposal, we defined a suite of 32 statement-level metrics, such as the number of binary and unary operators used in a statement. Then, we applied as learning model, long short-term memory (LSTM). The significance of SLDeep for intelligent and expert systems is that it demonstrates a novel use of deep-learning models to the solution of a practical problem faced by software developers. We conducted experiments using more than 100,000 C/C++ programs within the Code4Bench. The programs total 2,356,458 lines of code with 292,064 faulty lines. The benchmark comprises diverse set of programs and versions, written by thousands of developers. Therefore, it tends to give a model that can be used for cross-project SDP. In the experiments, our trained model could successfully classify the unseen data with average performance measures 0.945, 0.971, and 0.976 in terms of recall, precision, and accuracy, respectively. These experimental results suggest that SLDeep is effective for statement-level SDP. The impact of this work is twofold. Working at statement-level further alleviates developer’s burden in pinpointing the fault locations. Second, cross-project feature of SLDeep helps defect prediction research become more industrially-viable</p> <p>for more information visit <a href="https://github.com/sldeep/SLDeep">https://github.com/sldeep/SLDeep</a></p>
GTEx v8 Elastic Net prediction models
<p>Elastic Net prediction models and LD reference (PrediXcan/MultiXcan support, individual-level or summary-level versions)</p> <p># Data usage policy</p> <p>When using this data, you must acknowledge the source by citing the publication "Widespread dose-dependent effects of RNA expression and splicing on complex diseases and traits" (https://doi.org/10.1101/814350).</p> <p> </p> <p># Disclaimer</p> <p>The data is provided "as is", and the authors assume no responsibility for errors or omissions. <br> The User assumes the entire risk associated with its use of these data. <br> The authors shall not be held liable for any use or misuse of the data described and/or contained herein. <br> The User bears all responsibility in determining whether these data are fit for the User's intended use. </p> <p>The information contained in these data is not better than the original sources from which they were derived,<br> and both scale and accuracy may vary across the data set. <br> These data may not have the accuracy, resolution, completeness, timeliness, or other characteristics<br> appropriate for applications that potential users of the data may contemplate. <br> <br> The user is responsible to comply with any data usage policy from the original GWAS studies;<br> refer to the list of traits described [here](https://www.biorxiv.org/content/10.1101/814350v1)<br> to identify their respective Consortia's requirements.</p> <p><br> THE DATA IS PROVIDED WITHOUT WARRANTY OF ANY KIND,<br> EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,<br> FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.<br> IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY,<br> WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,<br> OUT OF OR IN CONNECTION WITH THE DATA OR THE USE OR OTHER DEALINGS IN THE DATA.</p>
Model output used in the manuscript "Micro and macro parametric uncertainty in climate change prediction: a large ensemble perspective"
<p>This *.zip file contains the model output from ensemble simulations for the Lorenz 84-Stommel 61 model (hereafter L84-S61; <a href="https://doi.org/10.3402/tellusa.v53i5.12229" target="_blank" rel="noopener">Van Veen et al, 2001</a>; <a href="https://doi.org/10.1088/1748-9326/8/3/034021" target="_blank" rel="noopener">Daron and Stainforth, 2013</a>). To run these simulations, we used the Low-EFFourth ensemble generator (<a href="https://doi.org/10.48550/arXiv.2506.03313" target="_blank" rel="noopener">de Melo Viríssimo, 2025a</a>; <a href="https://doi.org/10.5281/zenodo.15566109" target="_blank" rel="noopener">de Melo Viríssimo, 2025b</a>), which is a MATLAB-based framework that allows for large ensembles of low-dimensional dynamical systems to be run and studied in a systematic way (<a href="https://doi.org/10.5194/egusphere-egu23-14755" target="_blank" rel="noopener">de Melo Viríssimo and Stainforth, 2023</a>).</p> <p>These model outputs are presented and discussed in the manuscript "<em>Micro and macro parametric uncertainty in climate change prediction: a large ensemble perspective</em>", published by the Bulletin of the American Meteorological Society (<a href="https://doi.org/10.1175/BAMS-D-24-0064.1" target="_blank" rel="noopener">de Melo Viríssimo and Stainforth, 2025</a>). The manuscript describes the experiments performed, the parameter values used and the modifications done to the original L84-S61 model. For this matter, we also refer you to <a href="https://doi.org/10.1088/1748-9326/8/3/034021" target="_blank" rel="noopener">Daron and Stainforth (2013)</a> and <a href="https://doi.org/10.1063/5.0180870" target="_blank" rel="noopener">de Melo Viríssimo et al. (2024)</a>.</p> <p>All files uploaded were generated from simulations run by the lead author.</p> <p>For specific information about each file uploaded, please refer to the README file. The details of each experiment are also presented in the supplementary materials of the manuscript. If you have any questions, please feel free to contact me.</p>
Figure 6 in Maxent modeling for predicting potential distribution of goitered gazelle in central Iran the effect of extent and grain size on performance of the model
Figure 6. Total cross-validation AUC (CV-AUC) and spatial congruence AUC (SC- AUC) for a range of grain sizes.
Figure 4 in Maxent modeling for predicting potential distribution of goitered gazelle in central Iran the effect of extent and grain size on performance of the model
Figure 4. Divergence between the uncorrelated and pruned models estimated through Parolo divergence index at 250-m resolution. As is shown, there was little divergence (0– 0.2) between models in most of the study area.
Figure 1 in Maxent modeling for predicting potential distribution of goitered gazelle in central Iran the effect of extent and grain size on performance of the model
Figure 1. The location of the study area on a map of western Asia (right). Inset shows DEM of study area with polygons indicating the protected areas where populations of goitered gazelle occur.
Figure 5 in Maxent modeling for predicting potential distribution of goitered gazelle in central Iran the effect of extent and grain size on performance of the model
Figure 5. The change in performance index (AUC, left) and habitat suitability area (right) of the output model with increasing extent size (open circles) and grain size (filled circles) from 250 to 3000 m.
Figure 3 in Maxent modeling for predicting potential distribution of goitered gazelle in central Iran the effect of extent and grain size on performance of the model
Figure 3. Goitered gazelle distribution maps based on the uncorrelated model (left) and the pruned model (right) for the 250-m grid size.
Sample data for "Classification Modeling for Hazardous Rip Current Prediction" Notebook
<p>This sample dataset is used in the notebook "Classification Modeling for Hazardous Rip Current Prediction" to demonstrate the application of using machine learning to identify hazardous rip current. The notebook is available in the NOAA Center for Artificial Intelligence GitHub Learning Journey repository (https://github.com/noaa-ncai/learning-journey). The full dataset is available via NOAA.</p>
Winter Precipitation-Type Models for "Evidential Deep Learning: Enhancing Predictive Uncertainty Estimation for Earth System Science Applications"
<p>This contains trained model weights, scalers, and evaluation metrics for the winter precipitation-type models trained as part of the paper "Evidential Deep Learning: Enhancing Predictive Uncertainty Estimation for Earth System Science Applications". </p>
A3D Model Organism Database (A3D-MODB): a database for proteome aggregation predictions in model organisms
<p>The unified and integrated metadata accompanied by referencing identifiers from the A3D database is available for download in CSV format.</p> <p>Aleksandra E Badaczewska-Dawid, Aleksander Kuriata, Carlos Pintado-Grima, Javier Garcia-Pardo, Michał Burdukiewicz, Valentín Iglesias, Sebastian Kmiecik, Salvador Ventura, A3D Model Organism Database (A3D-MODB): a database for proteome aggregation predictions in model organisms, <em>Nucleic Acids Research</em>, Volume 52, Issue D1, 5 January 2024, Pages D360–D367, <a href="https://doi.org/10.1093/nar/gkad942">https://doi.org/10.1093/nar/gkad942</a></p>
Evaluating mesoscale model predictions of diurnal speedup events in the Altamont Pass Wind Resource Area of California
<p>This dataset contains input files for the Weather Research and Forecasting (WRF) model related to the manuscript "Evaluating mesoscale model predictions of diurnal speedup events in the Altamont Pass Wind Resource Area of California," to be submitted to the <em>Journal of Applied Meteorology and Climatology</em> by Arthur, et al. Included are:</p> <ul> <li><strong>namelist.wps</strong>: used by the WRF preprocessing system (WPS) to configure the model domain and initial/boundary conditions</li> <li><strong>Vtable.HRRR</strong>: used by WPS to process data from the High-Resolution Rapid Refresh (HRRR) model for WRF initial/boundary conditions</li> <li><strong>namelist.input.mynn</strong>: used to run the MYNN PBL simulation</li> <li><strong>namelist.input.3dpbl</strong>: used to run the 3D PBL simulation</li> <li><strong>windturbines.txt</strong>: used to define the location and type of wind turbines included in the simulations</li> <li><strong>wind-turbine-*.tbl</strong>: used to define the parameters of each turbine type (see Table 1 in Arthur et al.) <ul> <li><strong>1</strong>: NREL 1.7MW, H=80m, D=103m</li> <li><strong>2</strong>: NREL 2.3MW, H=80m, D=107m</li> <li><strong>3</strong>: NREL 2.3MW, H=80m, D=116m</li> <li><strong>4</strong>: Vestas V47 0.66MW, H=60m, D=47m</li> <li><strong>5</strong>: Bonus B54 1.0MW, H=55m, D=54m</li> </ul> </li> </ul> <p>This work was prepared by LLNL under Contract DE-AC52-07NA27344.</p>
Four spatial prediction datasets of susceptibility to gully erosion, comparing machine learning models, in the Piraí drainage basin, southeastern Brazil
<p>The data in this repository refer to the article published in the journal Land, entitled: Machine Learning Models for the Spatial Prediction of Gully Erosion Susceptibility in the Piraí Drainage Basin, Paraíba do Sul Middle Valley, Southeast Brazil.</p>
Development and Comparison of Model-Based and Data-Driven Approaches for the Prediction of the Mechanical Properties of Lattice Structures
<p>This dataset comes from the following paper:</p> <p>Chiara Pasini, Oscar Ramponi, Stefano Pandini, Luciana Sartore, Giulia Scalet, Development and Comparison of Model-Based and Data-Driven Approaches for the Prediction of the Mechanical Properties of Lattice Structures, J. of Materi Eng and Perform, 2024. <a href="https://doi.org/10.1007/s11665-024-10199-x">https://doi.org/10.1007/s11665-024-10199-x</a></p> <p>It contains:</p> <ul> <li>"Notes.pdf" describing all the files uploaded</li> <li>. m of the neural network</li> <li>. inp of the Abaqus finite element simulations</li> </ul>
The raw data for the research "Comparing Neural Network Models Based on Macro Perspective Economic and Environmental Indicators with ARIMA Model in predicting Construction Cost Index in UK"
<p>The raw data for the research "Comparing Neural Network Models Based on Macro Perspective Economic and Environmental Indicators with ARIMA Model in predicting Construction Cost Index in UK".</p> <p>Data collector: Runda Zheng</p>
Online example database generated representing nuclear astrophysics models predictions of correlations between stable/stable abundances of specific isotopes.
<p>Library of figures created using the SIMPLE code (Stellar Interpretation for Meteoritic data and PLotting). The SIMPLE stellar database includes 18 core-collapse supernova models with 3 different initial masses of 15, 20, 25 solar masses, all of solar metallicity and non-rotating stars. The 6 sets are the following:</p> <ul> <li>Rauscher et al. 2002 [Ra02]<br>(<a href="https://ui.adsabs.harvard.edu/abs/2002ApJ...576..323R/abstract">https://ui.adsabs.harvard.edu/abs/2002ApJ...576..323R/abstract</a>),</li> <li>Pignatari et al. 2016 [Pi16]<br>(<a href="https://ui.adsabs.harvard.edu/abs/2016ApJS..225...24P/abstract">https://ui.adsabs.harvard.edu/abs/2016ApJS..225...24P/abstract</a>),</li> <li>Sieverdin et al. 2018 [Si18]<br>(<a href="https://ui.adsabs.harvard.edu/abs/2018ApJ...865..143S/abstract">https://ui.adsabs.harvard.edu/abs/2018ApJ...865..143S/abstract</a>),</li> <li>Limongi & Chieffi 2018 [LC18]<br>(<a href="https://ui.adsabs.harvard.edu/abs/2018ApJS..237...13L/abstract">https://ui.adsabs.harvard.edu/abs/2018ApJS..237...13L/abstract</a>),</li> <li>Ritter et al. 2018 [Ri18]<br>(<a href="https://ui.adsabs.harvard.edu/abs/2018MNRAS.480..538R/abstract">https://ui.adsabs.harvard.edu/abs/2018MNRAS.480..538R/abstract</a>),</li> <li>Lawson et al. 2022 [La22]<br>(<a href="https://ui.adsabs.harvard.edu/abs/2022MNRAS.511..886L/abstract">https://ui.adsabs.harvard.edu/abs/2022MNRAS.511..886L/abstract</a>)</li> </ul> <p>The figures can be divided into two types. The first shows the structure of the ejecta and the abundance of the selected isotopes. The layers are automatically detected using SIMPLE based on the abundances of the main fuels (H-1, He-4, C-12, O-16, Ne-20, Si-28) from the supernova model ejecta. The code names the different layers based on the schematic diagram in Schofield et al. 2022 (<a href="https://ui.adsabs.harvard.edu/abs/2022MNRAS.517.1803S/abstract">https://ui.adsabs.harvard.edu/abs/2022MNRAS.517.1803S/abstract</a>). <br>Ni and Fe isotopes are plotted in the figures. The abundances shown include the radiogenic contribution from unstable isotopes.</p> <p>The SIMPLE code is designed to compare stellar data with measurements from meteorites. To achieve this, abundances in mass fractions need to be converted into isotopic ratios using specific units. See Lugaro et al. 2023 (<a href="https://ui.adsabs.harvard.edu/abs/2023EPJA...59...53L/abstract">https://ui.adsabs.harvard.edu/abs/2023EPJA...59...53L/abstract</a>) for details. For the specific case of Ni64 ratios, in comparison with model data we report the measured meteoritic anomaly by Steele et al 2012 (<a href="https://ui.adsabs.harvard.edu/abs/2012ApJ...758...59S/abstract">https://ui.adsabs.harvard.edu/abs/2012ApJ...758...59S/abstract</a>) as a continuous horizontal line. The same is done for the Fe54 ratios, with reference measurements by Hopp et al 2022 (<a href="https://ui.adsabs.harvard.edu/abs/2022E%26PSL.57717245H/abstract">https://ui.adsabs.harvard.edu/abs/2022E%26PSL.57717245H/abstract</a>). </p> <p>In the database the abundance plots are identified as <strong>structure_<model refe<em>rence>_<initial mass>_<element>_<decayed or undecayed>.png</em></strong><em>. In particular, the available reference model options are Ra02, Pi16, Si18, LC18, Ri18, La22; the initial mass of the progenitors are 15, 20 or 25 (solar masses). The third part of the filenames are the plotted elements (in this case Ni or Fe) and then if they are decayed or undecayed. In this database we only consider the decayed species, which means that the radioactive isotopes whose decay can add to the abundance of the selected isotopes were considered. In this plots the x-axis represents the total mass from the core and the y-axis is the mass fraction on a logarithmic scale. For the slopes the same name scheme applies, but they are identified as <strong>slopes_<model reference>_<initial mass>_<element>_<decayed or undecayed></strong></em><strong>.png,</strong> and on the y-axis there are the slope values insteas of abundances.</p>
Transfer learning and DNA language models enhance transcription factor binding predictions
<p>This is the dataset for replicating the results of the paper called "Transfer learning and DNA language models enhance transcription factor binding predictions" by Ekin Deniz Aksu and Martin Vingron.</p> <p>See https://github.com/ekinda/tfbs_prediction_paper</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.