Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
38
datasets available to search
ShareScore release 0.9.0
Dataset results
38 results for “predictive design”
Data from: Predictive design of crystallographic chiral separation
Open the record for dataset details and reuse information.
Data from: Experimental design in phylogenetics: testing predictions from expected information
Taxon and character sampling is central to phylogenetic experimental design yet we lack general rules. Goldman introduced a method to construct efficient sampling designs in phylogenetics, based on the calculation of expected Fisher information given a probabilistic model of sequence evolution. The considerable potential of this approach remains largely unexplored. In an earlier study, we applied Goldman's method to a problem in the phylogenetics of caecilian amphibians and made an a priori evaluation and testable predictions of which taxon additions would increase information about a particular weakly supported branch of the caecilian phylogeny by the greatest amount. Using mitogenomic and rag1 sequences (some newly determined for this study) from additional caecilian species we studied how information (both expected and observed) and bootstrap support varies as each new taxon is individually added, providing the first empirical test of specific predictions made using Goldman's method for phylogenetic experimental design. Our results empirically validate the top three (more intuitive) taxon addition predictions made in our previous study, but only information results validate unambiguously the fourth (less intuitive) prediction. This highlights a complex relationship between information and support, reflecting that each measures different things: information is related to the ability to estimate branch length accurately, and support to the ability to estimate the tree topology accurately. Thus, an increase in information may be correlated with but does not necessitate an increase in support Our results also provide the first empirical validation of the widely held intuition that additional taxa that join the tree proximal to poorly supported internal branches are more informative and enhance support more than additional taxa that join the tree more distally. Our work supports the view that adding more data for a single (well chosen) taxon may increase phylogenetic resolution and support in weakly supported parts of the tree without adding more characters/genes while illustrating that less well chosen taxon additions can have the opposite effect. Altogether our results corroborate that, although still underexplored, Goldman's method offers a powerful tool for experimental design in molecular phylogenetic studies. However, there are still several drawbacks to overcome, and further assessment of the method is needed in order to make it better understood, more accessible, and able to assess additions of multiple taxa.
Prediction of designer-recombinases for DNA editing with generative deep learning
<p>Sequence data from "Prediction of designer-recombinases for DNA editing with generative deep learning" publication.<br> Tyrosine site-specific recombinase gene sequences with the corresponding target sequences from already published projects. Recombinases were produced with directed evolution and sequenced with PacBio HiFi.</p> <p>The evolution of the recombinases sequences in this dataset were published in the following publications:</p> <ul> <li>https://doi.org/10.1126/science.1141453</li> <li>https://doi.org/10.1038/nbt.3467</li> <li>https://doi.org/10.1093/nar/gkz1078</li> <li>https://doi.org/10.1038/s41467-022-28080-7</li> </ul>
Data from: Experimental design in phylogenetics: testing predictions from expected information
Open the record for dataset details and reuse information.
Data from: Predicting rice hybrid performance using univariate and multivariate GBLUP models based on North Carolina mating design II
Open the record for dataset details and reuse information.
DeepSTARR predicts enhancer activity from DNA sequence and enables the de novo design of synthetic enhancers [Human oligo UMI-STARR-seq]
GEO Series GSE183938. Homo sapiens; synthetic construct. 4 samples. Type: Other.
DeepSTARR predicts enhancer activity from DNA sequence and enables the de novo design of synthetic enhancers [Drosophila genome-wide UMI-STARR-seq]
GEO Series GSE183936. Drosophila melanogaster; synthetic construct. 6 samples. Type: Other.
DeepSTARR predicts enhancer activity from DNA sequence and enables the de novo design of synthetic enhancers [Drosophila oligo UMI-STARR-seq]
GEO Series GSE183937. Drosophila melanogaster; synthetic construct. 12 samples. Type: Other.
Human 5′ UTR design and variant effect prediction from a massively parallel translation assay
GEO Series GSE114002. Homo sapiens. 10 samples. Type: Expression profiling by high throughput sequencing.
Design of an unbiased machine learning workflow to predict Multiple Sclerosis staging from blood transcriptome
GEO Series GSE136411. Homo sapiens. 336 samples. Type: Expression profiling by array.
Data from: A polymer dataset for accelerated property prediction and design
Emerging computation- and data-driven approaches are particularly useful for rationally designing materials with targeted properties. Generally, these approaches rely on identifying structure-property relationships by learning from a dataset of sufficiently large number of relevant materials. The learned information can then be used to predict the properties of materials not already in the dataset, thus accelerating the materials design. Herein, we develop a dataset of 1,073 polymers and related materials and make it available at http://khazana.uconn.edu/. This dataset is uniformly prepared using first-principles calculations with structures obtained either from other sources or by using structure search methods. Because the immediate target of this work is to assist the design of high dielectric constant polymers, it is initially designed to include the optimized structures, atomization energies, band gaps, and dielectric constants. It will be progressively expanded by accumulating new materials and including additional properties calculated for the optimized structures provided.
Dataset for Crashworthiness in preliminary design: Mean crushing force prediction for closed-section thin-walled metallic structures
<p>This dataset is the official implementation of the following paper published in the International Journal of Impact Engineering journal:</p> <blockquote> <p>Shreyas Anand, René Alderliesten, Saullo G. P. Castro. "Crashworthiness in preliminary design: Mean crushing force prediction for closed-section thin-walled metallic structures". International Journal of Impact Engineering, 2024. <a href="https://doi.org/10.1016/j.ijimpeng.2024.104946" target="_blank" rel="noopener">10.1016/j.ijimpeng.2024.104946</a> </p> </blockquote> <h2>Highlights</h2> <div> <div> <ul> <li> <div>Evaluation of analytical models to predict mean crushing force</div> </li> <li> <div>Analysis of both extensional and inextensional analytical crushing models.</div> </li> <li>Use of numerical and experimental dataset to improve existing analytical models.</li> <li>Generalized expression for predicting mean crushing force</li> </ul> </div> </div> <h2>Abstract</h2> <div> <div>To design crash structures for disruptive aircraft designs, it is required to have fast and accurate methods that can predict crashworthiness of aircraft structures early in the design phase. Axial crushing is one of the key energy absorbing mechanisms during a crash event. In this study, various analytical models proposed for calculation of mean crushing force for thin-walled tubular structures are compared with a database of numerical and experimental values to ascertain their accuracy. Improvements to some of the models have also been proposed. Finally a generalized model based on the studied and improved analytical models for prediction of mean crushing force for closed section thin-walled tubular structures is introduced. The generalized model demonstrates high accuracy when compared against experimental/numerical dataset as evidenced by a high coefficient of determination (R^2) value of 0.97 and can therefore be used to estimate the mean crushing force for closed-section thin-walled metallic tubular structures with various cross-sectional shapes and crushing modes early in the design phase.</div> </div>
Database of "Dynamic Design and Performance Prediction of Tuned Particle Damper Based on Co-simulation"
<p>These data are obtained based on the co-simulation of ADAMS and EDEM. The first 7 columns are inputs and the last column is output.</p>
A I in the Prediction of Clinical Performance, Marginal Fit and Fracture Resistance of Vertical Versus Horizontal Margin Designs Fabricated With 2 Ceramic Materials
ClinicalTrials.gov study NCT06164002. IPD Sharing: NO. Countries: 1. Publications: 0.
Design of a Predictive Score for Contamination of Pediatric Blood Cultures
ClinicalTrials.gov study NCT06300736. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Data from: A polymer dataset for accelerated property prediction and design
Open the record for dataset details and reuse information.
Clinical evaluation of a functional combinatorial precision medicine platform to predict treatment outcomes and enhance combination therapy design in soft tissue sarcomas
GEO Series GSE282752. Homo sapiens. 24 samples. Type: Expression profiling by high throughput sequencing.
DeepSTARR predicts enhancer activity from DNA sequence and enables the de novo design of synthetic enhancers
GEO Series GSE183939. Homo sapiens; synthetic construct; Drosophila melanogaster. 22 samples. Type: Other.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.