Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

38

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

38 results for “predictive design”

Learn how ShareScore rates datasets ↗
zenodo40/100

MiRoR4_P1_A systematic review describes models for recruitment prediction at the design stage of a clinical trial

<p>This dataset is related to the publication &quot;A systematic review describes models for recruitment prediction at the design stage of a clinical trial&quot;.</p>

opencc-by-4.0Feb 2020View details →
zenodo40/100

Exploring Design Smells for Smell-Based Defect Prediction

<p>The archived file datasets.zip includes the datasets used for supporting the conclusions in the article <em>Exploring Design Smells for Smell-Based Defect Prediction.</em></p> <p>In this paper, we answer two research questions:</p> <p><strong>RQ1.</strong> Do Design code smells contribute to the performance of defect prediction models trained with Traditional code smells?</p> <p><strong>RQ2. </strong>How do the different categories of Design smells impact the performance of the defect prediction models?</p> <p>Therefore, after extracting the archived file documents, you will find two sub-directories, respectively named &quot;RQ1&quot; and &quot;RQ2&quot;. They include the results obtained for each one of the research questions, thus supporting our conclusions.</p> <p>(You will also find a README.pdf file with these same instructions regarding the datasets.)</p> <p>Inside &quot;RQ1,&quot; you will find two directories, respectively named &quot;configuration_1&quot; and &quot;configuration_2&quot;. They represent the different configurations for the experiments. <strong>&quot;configuration_1&quot;</strong> contains the datasets with results for the ten classifiers configurations with the highest scores and <strong>&quot;configuration_2&quot; </strong>contains the datasets with the results classifier configuration with the overall best results - Support Vector Machine with C=0.1. Furthermore, within each directory, there are three sub-directories, respectively named &quot;designite,&quot; &quot;designite_traditional,&quot; and &quot;traditional.&quot; These have the datasets for each of the considered smell sets in our study. Inside &quot;RQ2,&quot; you will find four directories. Each corresponds to a category from the design smells for the dataset &quot;designite_traditional.&quot; These datasets were build from the same configuration as &quot;configuration_2&quot;.</p> <p>Then, within every directory, there are 97 sub-directories representing the 97 projects analyzed in this study.</p> <p>Every project folder follows the same structure, which we define as follows.</p> <ul> <li>The &quot;dataset&quot; directory contains the original training and testing dataset used.</li> <li>The &quot;oversamples&quot; directory contains the training dataset after oversampling for each of the feature selection approaches.</li> <li>The &quot;score_summary&quot; directory contains all classifier configurations considered, not only the 10 with the highest scores.</li> <li>The &quot;scores.csv&quot; file contains all the scores for the main classifier configurations studied in the particular experiment.</li> <li>The &quot;selected_features&quot; directory contains the selected features&#39; information and the selected features dataset for each feature_selection method.</li> <li>The &quot;selected_testing_X&quot; directory contains the testing datasets.</li> <li>The &quot;top_scores_summary&quot; directory contains the classifier configurations and hyper-parameter scores for the top 10 highest scores.</li> </ul>

opencc-by-4.0Oct 2020View details →
zenodo40/100

Dataset of the manuscript: Predictive design of plasmonic color

<p>This data publication is based on the metadata and datasets underlying the manuscripts "Predictive design of plasmonic color"</p> <p>Folder "Figure 1" contains:</p> <ul> <li>Simulated Extinction spectra of 120 nm silica particle coated with 20 nm gold shell</li> <li>Simulated Transmission spectra of 120 nm silica particle coated with 20 nm gold shell scaled for different concentrations + Lab and RGB color coordinates</li> <li>3D plots of accessible color gamut for Ag, Au, SiO2@Au and SiO2@Ag nanoparticles</li> </ul> <p>Figure "Figure 2" contains:</p> <ul> <li>Simulated colors of targeted for varying concentrations</li> <li>Color-map for color of the year 2022</li> <li>Color coordinates (Lab and RGB) of structure in spot 1 and spot 4 scaled for different concentrations</li> </ul> <p>Figure "Figure 3" contains:</p> <ul> <li>Photographs of realized particle dispersion (spot1: 0@52, spot2: 113@46, spot3: 137@24, spot 4: 398@37) at high and low concentrations, approx. concentrations in number of particles per ml</li> <li>TEM images of realized particles (for core-shell particle: after seeding (x@NP) and after shell growth)</li> <li>UV-Vis:</li> <ul> <li>Measured absorbance spectra and scaled transmission spectra of realized particles spot</li> <li>Color coordinates of simulated, experimentally determined and target color</li> </ul> </ul> <p>Figure "Figure 4" contains:</p> <ul> <li>Photographs of realized particle dispersion (CRC1411yellow: Ag0@3-24, CRC1411red: Au0@13, CRC1411blue: Au113@35), approx. concentrations in number of particles per ml</li> <li>TEM images of realized particles (for core-shell particle: after seeding (x@NP) and after shell growth)</li> <li>UV-Vis:</li> <ul> <li>Measured absorbance spectra and scaled transmission spectra of realized particles</li> <li>Color coordinates of simulated, experimentally determined and target color</li> </ul> <li>color maps for CRC1411 target colors</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Galaxy Training Data for "Designing plasmids encoding predicted pathways by using the BASIC assembly method"

<p>This dataset provides the data needed for the Galaxy BASIC assembly workflow training tutorial (<a href="https://galaxy-synbiocad.org">https://galaxy-synbiocad.org</a>). This workflow provides a pathway to design plasmids encoding predicted metabolic pathways using the BASIC assembly method (<a href="https://doi.org/10.1021/sb500356d">https://doi.org/10.1021/sb500356d</a>). It generates scripts allowing the automatic construction of these plasmids using an Opentrons liquid handling robot. After downloading these scripts on a computer connected to an Opentrons (<a href="https://opentrons.com">https://opentrons.com</a>), the user can perform the automatic construction of the plasmids on the bench.</p> <p>The content of the dataset is as follows:</p> <ul> <li> <p>an SBML file modeling a heterologous pathway producing lycopene such as those produced by the Pathway Analysis Workflow (<a href="https://galaxy-synbiocad.org">https://galaxy-synbiocad.org</a>).</p> </li> <li> <p>a CSV file listing the parts to be used (linkers, backbone and promoters) in the constructions.</p> </li> <li> <p>two YAML files providing two examples of settings, i.e. providing the identifiers of the laboratory equipment and the parameters of the DNA robot.</p> </li> </ul>

opencc-by-4.0Feb 2022View details →
dryad40/100

How to quantify factors degrading DNA in the environment and predict degradation for effective sampling design

<p>Extra-organismal DNA (eoDNA) from material left behind by organisms (non-invasive DNA: e.g., faeces, hair) or from environmental samples (eDNA: e.g., water, soil) is a valuable source of genetic information. However, the relatively low quality and quantity of eoDNA, which can be further degraded by environmental factors, results in reduced amplification and sequencing success. This is often compensated for through cost- and time-intensive replications of genotyping/sequencing procedures. Therefore, system- and site-specific quantifications of environmental degradation are needed to maximize sampling efficiency (e.g., fewer replicates, shorter sampling durations), and to improve species detection and abundance estimates. Using ten environmentally diverse bat roosts as a case study, we developed a robust modelling pipeline to quantify the environmental factors degrading eoDNA, predict eoDNA quality, and estimate sampling-site-specific ideal exposure duration. Maximum humidity was the strongest eoDNA-degrading factor, followed by exposure duration and then maximum temperature. We also found a positive effect when hottest days occurred later. The strength of this effect fell between the strength of the effects of exposure duration and maximum temperature. With those predictors and information on sampling period (before or after offspring were born), we reliably predicted mean eoDNA quality per sampling visit at new sites with a mean squared error of 0.0349. Site-specific simulations revealed that reducing exposure duration to 2-8 days could substantially improve eoDNA quality for future sampling. Our pipeline identified high humidity and temperature as strong drivers of eoDNA degradation even in the absence of rain and direct sunlight. Furthermore, we outline the pipeline's utility for other systems and study goals, such as estimating sample age, improving eDNA-based species detection, and increasing the accuracy of abundance estimates.</p>

opencc-zeroMar 2023View details →
zenodo40/100

Microclimate simulation output: "Between vision and action: the predicted effects of co-designed green infrastructure solutions on environmental burdens"

<p>The following microclimate simulation dataset&nbsp;supports the paper &quot;Between vision and action: the predicted effects of co-designed green infrastructure solutions on environmental burdens&quot; by Mathias Schaefer, published in&nbsp;Urban Ecosystems (2022).</p> <p>&quot;T0Simulation_11082020_output&quot; contains data about the status quo simulation of the area of interest (500 m x 500 m x 60 m), whereas &quot;T1Simulation_11082020_output&quot; shows the results of the Green Infrastructure scenario described in the research article&nbsp;above.&nbsp;Please ensure enough memory space on your device, as both files have a size of approximately 25 GB (unzipped).</p> <p>The output files can be visualized with the ENVI-met Leonardo extension. The ENVI-met LITE-version&nbsp;is freely available and can be downloaded at the&nbsp;<a href="https://envi-met.info/doku.php?id=files:download">ENVI-met homepage</a>. Alternatively, the included .NETCDF files can be imported&nbsp;as a multidimensional raster dataset in ArcGIS Pro.</p> <p>Files in the folder &quot;atmosphere&quot; represent meteorological parameters such as potential air temperature [&deg;C], relative humidity [%], or wind speed [m/s]. Air pollution calculations like particulate matter concentrations [&micro;g/m&sup3;] can be found in the folder &quot;pollutants&quot;. The folder &quot;buildings&quot; contains building data for 3D visualizations of surface temperatures&nbsp;[&deg;C].</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

INSPIRED: Inelastic Neutron Scattering Prediction for Instantaneous Results and Experimental Design

<p>INSPIRED is a graphic user interface (GUI) that performs rapid prediction and calculation of phonons and inelastic neutron scattering (INS) spectra. It consists of three modules. The "Predictor" module uses a symmetry-aware neural network (coupled with an autoencoder) [1-3] to perform direct prediction of total/partial phonon density of states and powder 1D/2D INS spectra from a given structure. The "DFT database" module uses pre-calculated force constants from density functional theory (DFT) [4] to perform INS simulations for single crystals and powders (for the crystals available in the database). The "MLFF" module uses pre-trained universal force fields [8-12] to perform structural optimization, phonon calculation, and INS simulations for single crystals and powders for any given crystal. The predicted/calculated results are saved in CSV files and can be visualized with the GUI. INSPIRED is developed to be a convenient tool for INS experimental planning, steering, and quick data analysis.</p> <p>This repository contains two files as an update to the previous version:</p> <p>1. A tarball file (dftdb.tar.gz) containing the DFT database (currently with 12734 crystals)</p> <p>2. A VirtualBox appliance file (inspired_vm.ova) to run INSPIRED as a virtual machine.</p> <p>The ML model file (model.tar.gz) remains the same and can be obtained from the previous version.</p> <p>Instructions on how to use these files, as well as the rest part of the software, can be found on the&nbsp;<a href="https://github.com/cyqjh/inspired">GitHub page</a>.&nbsp;</p>

openmit-licenseMar 2024View details →
zenodo40/100

Dataset of the manuscript: Predictive design to determine optimal absorber placement in colloidal photonic crystals

<p><span>This data publication is based on the metadata and datasets underlying the manuscript "Predictive design to determine optimal absorber placement in colloidal photonic crystals". The Data is roughly organized by the figure of appearance.</span></p> <p><span>Figure 1 contained no result data</span></p> <p><span>Figure "Figure 2" contains:</span></p> <ul> <li><span>Simulated and experimental reflectance spectra of bare PS colloidal crystal and CIELab color coordinates of simulated bare PS colloidal crystal.</span></li> </ul> <p><span>Figure "Figure 3+4" contains:</span></p> <ul> <li><span>Data for particle based and layer based designs for chroma optimization according to Eq. 2 </span></li> <ul> <li><span>convergence history J(steps)</span></li> <li><span>Optimized design absorber distributions (average of layers)</span></li> <li><span>CIELab color coordinates</span></li> <li><span>Spectra</span></li> </ul> <li><span>CIELab color coordinates and Chroma of all predictive designs sorted by threshold L value according to Eq. 3 + comparative designs: bottom absorber, top absorber, homogeneous.</span></li> </ul> <p><span>Figure "Figure 5" contains:</span></p> <ul> <li><span>Chemdraw File containing chemical structures</span></li> <li><span>Pendant drop surface tension measurements&nbsp;</span></li> <li><span>Surface pressure increase on Langmuir-Blodgett trough</span></li> </ul> <p><span>Figure "Figure 6" contains:</span></p> <ul> <li><span>SEM images of mono- and multilayers labeled in accordance to design and composition</span></li> </ul> <p><span>Figure "Figure 7+8" contains:</span></p> <ul> <li><span>Photographs of fabricated multilayers</span></li> <ul> <li><span>Homogeneous designs labeled in accordance to composition</span></li> <li><span>Layered designs labeled in accordance to design type (XBA: bottom absorber with X absorbing numbers; XTA: bottom absorber with X absorbing numbers; Ld_XX: predictive design with L threshold of XX)</span></li> </ul> <li><span>Spectra of all samples including their error determined from 2 measurements</span></li> <li><span>Average spectra of all designs (averaged from all samples of that design) including their error estimated using gaussian error propagation</span></li> <li><span>Color data of all samples calculated from spectra </span></li> <ul> <li><span>CIELab coordinates and Chroma</span></li> <li><span>xyz values</span></li> <li><span>RGB values</span></li> </ul> </ul> <p><span>Figure "Figure 9" contains:</span></p> <ul> <li><span>Optimized design absorber distributions, spectra and CIELab color coordinates and chroma for colloidal crystals of varying primary particle size</span></li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo40/100

In silico design, docking simulation, and ANN-QSAR model for predicting the anticoagulant activity of thiourea isosteviol compounds as FXa inhibitors

<p>The present work combined molecular modeling and docking approach for searching and designing novel thiourea isosteviol-based compounds as potential FXa inhibitors. Elaborated regression model establishes the relationships between experimentally determined anticoagulant activity and molecular descriptors and enables the prediction of FXa inhibitory activity for novel compounds. The obtained results proved that the Artificial Neural Network algorithm facilitates the search for the most promising isosteviol derivatives incorporating thiourea fragments as FXa inhibitors. Moreover, docking simulation confirms the prominent binding of the newly in silico designed molecules with the active sites of the protein, which may be the lead molecules and can be further optimized for the efficient pharmacodynamic and pharmacokinetic profiles.&nbsp;The enclosed files are representations of&nbsp;molecular structures of thiourea isosteviol compounds with experimentally tested FXa inhibitory activity (i20-i39) geometrically optimized in hyperchem, newly in silico designed thiourea isosteviol compounds geometrically optimized in hyperchem (e1-e11), one file contains molecular descriptors for optimized structures calculated in Dragon and there is also a code for ANN QSAR model for predicting activity of novel thiourea isosteviol compounds.&nbsp;</p>

opencc-by-4.0Jun 2023View details →
dryad40/100

How to quantify factors degrading DNA in the environment and predict degradation for effective sampling design

Open the record for dataset details and reuse information.

publicMar 2023View details →
zenodo36/100

Supplementary Data associated with the article "LukProt: A database of eukaryotic predicted proteins designed for investigations of animal origins"

<p>Supplementary Data associated with the article "LukProt: A database of eukaryotic predicted proteins designed for investigations of animal origins" - BUSCO example.</p> <p>&nbsp;</p> <p>National Science Centre of Poland is acknowledged for funding of the project 2020/36/C/NZ8/00081, "The role of glycosylation in the emergence of animal multicellularity", which enabled the creation of this dataset.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Testability prediction dataset: Design level

<p>This dataset contains design metrics of the software projects in the SF110 corpus. The metrics have been extracted from the extended class diagram of these projects. The last column is the design testability value of each class in the dataset.</p> <p>The dataset is part of the paper entitled &quot;<strong>Measuring and improving software testability at the design level</strong>&quot;.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Dataset for Peptide binder design with inverse folding and protein structure prediction

<p>Dataset for a paper on peptide design</p> <p>&nbsp;</p> <p><br> mutated_peptides - results for randomly intriduced mutations in protein-peptide complexes that can be predicted at 2 &Aring; (Figure 1)<br> pdb_peptide&nbsp; - variation in the number of recycles (1-10) for 96 peptides (Figure 1)<br> minibinder - results for the minibinder set (Figure 2)<br> Pfam - results for the Pfam set (Figures 4+5)<br> protein_mpnn - results on protein_mpnn test set (Figure 6)</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo32/100

Predictive Design of Ultrastretchable Electrodes with Strain-Insensitive Performance via Robotics- and Machine Learning-Integrated Workflow

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo32/100

DSMBind: SE(3) denoising score matching for unsupervised binding energy prediction and nanobody design

<p>This record provides the training and evaluation dataset for DSMBind</p> <p>Paper link: https://www.biorxiv.org/content/10.1101/2023.12.10.570461v1.abstract</p> <p>Github repo: https://github.com/wengong-jin/DSMBind</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

Source data for manuscript(De novo protein design with a denoising diffusion network independent of pre-trained structure prediction models)

<p>This respository contains the source data for figure and supplementary figure in manuscript(SCUBA-D).</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Source data for manuscript(De novo protein design with a denoising diffusion network independent of pre-trained structure prediction models)

<p>This respository contains the source data for figure and supplementary figure in manuscript(SCUBA-D).</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

A Deep Learning Approach to the Forward Prediction and Inverse Design of Plasmonic Metasurface Structural Color - Raw Data

<p>Reflection spectra of PDMS - Al nanorod metamaterials&nbsp;were collected using LUMERICAL&nbsp;FDTD simulations. PDMS material properties were defined using a refractive index of 1.41 and Al material properties were defined using frequency selective permittivities from the handbook of Palik. A total of 4620 structures were simulated, sweeping the following dimension parameters:</p> <ul> <li>Aluminium thickness (t)</li> <li>Pillar height (h)</li> <li>Pillar diameter (d)</li> </ul> <p>Reflectance spectra were converted into CIE 1931 chromaticity values (x,y). This dataset is comprehensive and allows for the development of deep learning models for the forward and inverse design of the given metamaterial structure as detailed in the associated manuscript.&nbsp;The associated manuscript and supporting documentation provide extensive details of data collection and processing methods.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo32/100

Data and Code from "Structure-based prediction of Ras-effector binding affinities and design of 'branchegetic' interface mutations"

<p>Data, data generation and data analysis for manuscript &quot;Structure-based prediction of Ras-effector binding affinities and design of &lsquo;branchegetic&rsquo; interface mutations&quot;, currently available as a preprint <a href="https://doi.org/10.1101/2022.09.04.506480">here</a>.</p> <p>Contains the following directories:</p> <ul> <li>01_models: Contains all scripts for model generation and selection, as well as some of the generated and selected models. <ul> <li>01_inputs: The different inputs for the homology modelling pipeline. This includes AlphaFold single and complex templates, PDB templates and sequence alignments.</li> <li>02_validation: Model generation and initial selection for validation models, based on AF2 single models and PDB complex models.</li> <li>03_production1: Model generation and initial selection for Ras effector complexes, based on AF2 single models and PDB complex models.</li> <li>04_production2: Model generation and initial selection for Ras effector complexes, based on AF2 single models and AF2 complex models.</li> <li>05_selection_optics: Code and analysis for selection by unsupervised learning using OPTICS.</li> </ul> </li> <li>02_selected_models: The three representative models selected for each complex.</li> <li>03_affinity_prediction: Contains code and data for the prediction of binding affinities for Ras effector complexes.</li> <li>04_branch_pruning: Contains code and data for branch pruning analysis.</li> <li>05_systems_analysis: Contains code and data for the analysis of Ras effector systems based on affinities derived from affinity prediction and branch pruning analysis.</li> <li>06_visualization: Information on where in the raw data the panels for the figures in the manuscript can be found.</li> </ul>

opencc-by-4.0Oct 2022View details →
ClinicalTrials.gov32/100

Design of a Genetic Score to Predict the Response to a Dietary Intervention in Adults With Metabolic Syndrome

ClinicalTrials.gov study NCT05495074. IPD Sharing: NO. Countries: 1. Publications: 5.

closedIPD-NOFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record