Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

38

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

38 results for “predictive design”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: Predictive design of crystallographic chiral separation

Open the record for dataset details and reuse information.

publicAug 2025View details →
dryad28/100

Data from: Experimental design in phylogenetics: testing predictions from expected information

Taxon and character sampling is central to phylogenetic experimental design yet we lack general rules. Goldman introduced a method to construct efficient sampling designs in phylogenetics, based on the calculation of expected Fisher information given a probabilistic model of sequence evolution. The considerable potential of this approach remains largely unexplored. In an earlier study, we applied Goldman's method to a problem in the phylogenetics of caecilian amphibians and made an a priori evaluation and testable predictions of which taxon additions would increase information about a particular weakly supported branch of the caecilian phylogeny by the greatest amount. Using mitogenomic and rag1 sequences (some newly determined for this study) from additional caecilian species we studied how information (both expected and observed) and bootstrap support varies as each new taxon is individually added, providing the first empirical test of specific predictions made using Goldman's method for phylogenetic experimental design. Our results empirically validate the top three (more intuitive) taxon addition predictions made in our previous study, but only information results validate unambiguously the fourth (less intuitive) prediction. This highlights a complex relationship between information and support, reflecting that each measures different things: information is related to the ability to estimate branch length accurately, and support to the ability to estimate the tree topology accurately. Thus, an increase in information may be correlated with but does not necessitate an increase in support Our results also provide the first empirical validation of the widely held intuition that additional taxa that join the tree proximal to poorly supported internal branches are more informative and enhance support more than additional taxa that join the tree more distally. Our work supports the view that adding more data for a single (well chosen) taxon may increase phylogenetic resolution and support in weakly supported parts of the tree without adding more characters/genes while illustrating that less well chosen taxon additions can have the opposite effect. Altogether our results corroborate that, although still underexplored, Goldman's method offers a powerful tool for experimental design in molecular phylogenetic studies. However, there are still several drawbacks to overcome, and further assessment of the method is needed in order to make it better understood, more accessible, and able to assess additions of multiple taxa.

opencc-zeroDec 2011View details →
zenodo28/100

Prediction of designer-recombinases for DNA editing with generative deep learning

<p>Sequence data from &quot;Prediction of designer-recombinases for DNA editing with generative deep learning&quot; publication.<br> Tyrosine site-specific recombinase gene sequences with the corresponding target sequences from already published projects. Recombinases were produced with directed evolution and sequenced with PacBio HiFi.</p> <p>The evolution of the recombinases sequences in this dataset were published in the following publications:</p> <ul> <li>https://doi.org/10.1126/science.1141453</li> <li>https://doi.org/10.1038/nbt.3467</li> <li>https://doi.org/10.1093/nar/gkz1078</li> <li>https://doi.org/10.1038/s41467-022-28080-7</li> </ul>

openOct 2022View details →
dryad28/100

Data from: Experimental design in phylogenetics: testing predictions from expected information

Open the record for dataset details and reuse information.

publicFeb 2012View details →
dryad28/100

Data from: Predicting rice hybrid performance using univariate and multivariate GBLUP models based on North Carolina mating design II

Open the record for dataset details and reuse information.

publicAug 2016View details →
geo24/100

DeepSTARR predicts enhancer activity from DNA sequence and enables the de novo design of synthetic enhancers [Human oligo UMI-STARR-seq]

GEO Series GSE183938. Homo sapiens; synthetic construct. 4 samples. Type: Other.

openGEO-OpenFeb 2022View details →
geo24/100

DeepSTARR predicts enhancer activity from DNA sequence and enables the de novo design of synthetic enhancers [Drosophila genome-wide UMI-STARR-seq]

GEO Series GSE183936. Drosophila melanogaster; synthetic construct. 6 samples. Type: Other.

openGEO-OpenFeb 2022View details →
geo24/100

DeepSTARR predicts enhancer activity from DNA sequence and enables the de novo design of synthetic enhancers [Drosophila oligo UMI-STARR-seq]

GEO Series GSE183937. Drosophila melanogaster; synthetic construct. 12 samples. Type: Other.

openGEO-OpenFeb 2022View details →
geo24/100

Human 5′ UTR design and variant effect prediction from a massively parallel translation assay

GEO Series GSE114002. Homo sapiens. 10 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMay 2018View details →
geo24/100

Design of an unbiased machine learning workflow to predict Multiple Sclerosis staging from blood transcriptome

GEO Series GSE136411. Homo sapiens. 336 samples. Type: Expression profiling by array.

openGEO-OpenSep 2020View details →
dryad24/100

Data from: A polymer dataset for accelerated property prediction and design

Emerging computation- and data-driven approaches are particularly useful for rationally designing materials with targeted properties. Generally, these approaches rely on identifying structure-property relationships by learning from a dataset of sufficiently large number of relevant materials. The learned information can then be used to predict the properties of materials not already in the dataset, thus accelerating the materials design. Herein, we develop a dataset of 1,073 polymers and related materials and make it available at http://khazana.uconn.edu/. This dataset is uniformly prepared using first-principles calculations with structures obtained either from other sources or by using structure search methods. Because the immediate target of this work is to assist the design of high dielectric constant polymers, it is initially designed to include the optimized structures, atomization energies, band gaps, and dielectric constants. It will be progressively expanded by accumulating new materials and including additional properties calculated for the optimized structures provided.

opencc-zeroDec 2015View details →
zenodo24/100

Dataset for Crashworthiness in preliminary design: Mean crushing force prediction for closed-section thin-walled metallic structures

<p>This dataset is the official implementation of the following paper published in the International Journal of Impact Engineering journal:</p> <blockquote> <p>Shreyas Anand, Ren&eacute; Alderliesten, Saullo G. P. Castro. "Crashworthiness in preliminary design: Mean crushing force prediction for closed-section thin-walled metallic structures". International Journal of Impact Engineering, 2024.&nbsp;<a href="https://doi.org/10.1016/j.ijimpeng.2024.104946" target="_blank" rel="noopener">10.1016/j.ijimpeng.2024.104946</a>&nbsp;</p> </blockquote> <h2>Highlights</h2> <div> <div> <ul> <li> <div>Evaluation of analytical models to predict mean crushing force</div> </li> <li> <div>Analysis of both extensional and inextensional analytical crushing models.</div> </li> <li>Use of numerical and experimental dataset to improve existing analytical models.</li> <li>Generalized expression for predicting mean crushing force</li> </ul> </div> </div> <h2>Abstract</h2> <div> <div>To design crash structures for disruptive aircraft designs, it is required to have fast and accurate methods that can predict crashworthiness of aircraft structures early in the design phase. Axial crushing is one of the key energy absorbing mechanisms during a crash event. In this study, various analytical models proposed for calculation of mean crushing force for thin-walled tubular structures are compared with a database of numerical and experimental values to ascertain their accuracy. Improvements to some of the models have also been proposed. Finally a generalized model based on the studied and improved analytical models for prediction of mean crushing force for closed section thin-walled tubular structures is introduced. The generalized model demonstrates high accuracy when compared against experimental/numerical dataset as evidenced by a high coefficient of determination (R^2) value of 0.97 and can therefore be used to estimate the mean crushing force for closed-section thin-walled metallic tubular structures with various cross-sectional shapes and crushing modes early in the design phase.</div> </div>

opencc-by-4.0Jan 2024View details →
zenodo24/100

Database of "Dynamic Design and Performance Prediction of Tuned Particle Damper Based on Co-simulation"

<p>These data are obtained based on the co-simulation of ADAMS and EDEM. The first 7 columns are inputs and the last column is output.</p>

opencc-by-4.0Jul 2024View details →
ClinicalTrials.gov24/100

A I in the Prediction of Clinical Performance, Marginal Fit and Fracture Resistance of Vertical Versus Horizontal Margin Designs Fabricated With 2 Ceramic Materials

ClinicalTrials.gov study NCT06164002. IPD Sharing: NO. Countries: 1. Publications: 0.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov24/100

Design of a Predictive Score for Contamination of Pediatric Blood Cultures

ClinicalTrials.gov study NCT06300736. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad24/100

Data from: A polymer dataset for accelerated property prediction and design

Open the record for dataset details and reuse information.

publicFeb 2017View details →
geo20/100

Clinical evaluation of a functional combinatorial precision medicine platform to predict treatment outcomes and enhance combination therapy design in soft tissue sarcomas

GEO Series GSE282752. Homo sapiens. 24 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMar 2025View details →
geo20/100

DeepSTARR predicts enhancer activity from DNA sequence and enables the de novo design of synthetic enhancers

GEO Series GSE183939. Homo sapiens; synthetic construct; Drosophila melanogaster. 22 samples. Type: Other.

openGEO-OpenFeb 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record