Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
252
datasets available to search
ShareScore release 0.9.0
Dataset results
252 results for “Synthetic data”
Orthogonal light-activated DNA for patterned biocomputing within synthetic cells (Source Data)
<p>Source data for the published version of "Orthogonal light-activated DNA for patterned biocomputing within synthetic cells": Preprint (https://chemrxiv.org/engage/chemrxiv/article-details/63b55bb6ff4651ef52429534)</p>
Data from: Classifying interactions in a synthetic bacterial community is hindered by inhibitory growth medium
<p>Predicting the fate of a microbial community and its member species relies on understanding the nature of their interactions. However, designing simple assays that distinguish between interaction types can be challenging. Here, we performed spent media assays based on the predictions of a mathematical model to decipher the interactions between four bacterial species: <em>Agrobacterium</em> <em>tumefaciens</em> (<em>At</em>), <em>Comamonas</em> <em>testosteroni</em> (<em>Ct</em>), <em>Microbacterium</em> <em>saperdae</em> (<em>Ms</em>) and <em>Ochrobactrum</em> <em>anthropi</em> (<em>Oa</em>). While most experimental results matched model predictions, the behavior of <em>Ct</em> did not: its lag phase was reduced in the pure spent media of <em>At</em> and <em>Ms</em>, but prolonged again when we replenished with our growth medium. Further experiments showed that the growth medium actually delayed the growth of <em>Ct</em>, leading us to suspect that <em>At</em> and <em>Ms</em> could alleviate this inhibitory effect. There was, however, no evidence supporting such "cross-detoxification" and instead, we identified metabolites secreted by <em>At</em> and <em>Ms</em> that were then consumed or "cross-fed" by <em>Ct</em>, shortening its lag phase. Our results highlight that even simple, defined growth media can have inhibitory effects on some species and that such negative effects need to be included in our models. Based on this, we present new guidelines to correctly distinguish between different interaction types, such as cross-detoxification and cross-feeding.</p>
Data for: Engineering tRNA abundances for synthetic cellular systems
<p>Data for publication:<strong> Engineering tRNA abundances for synthetic cellular systems</strong></p> <p><strong>Abstract</strong></p> <p>Routinizing the engineering of synthetic cells requires specifying determining beforehand how many of each molecule are needed. First-principles tools for specifying molecular abundances enabling whole-cell synthetic biology are missing. We use a colloidal dynamics simulator to make predictions for how tRNA abundances impact protein synthesis rates. We use rational design and direct RNA synthesis to make 21 synthetic tRNA surrogates from scratch. We use evolutionary algorithms within a computer aided design framework to design engineer translation systems predicted to work faster or slower depending on tRNA abundance differences. We build and test the so-specified synthetic systems and find that qualitative agreement between expected and observed systems performance matchqualitatively match. First-principles modeling combined with bottom-up experiments can help molecular-to-cellular scale synthetic biology realize “design, build, work” frameworks that transcend tinker-and-test.</p> <p><strong>Data description</strong></p> <p>The data here consists of (1) All Colloidal Smoldyn & CD-CAD simulation input parameter and output files & (2) experimental data used for the associated publication. Simulation data was produced using Colloidal Dynamics modeling and Colloidal Dyamics-CAD (CD-CAD) as described in the associated manuscript. Data folders should be used directly with modeling and analysis code provided on Github: https://github.com/EndyLab/tRNACAD.</p>
Modified Fuchs et al. model Synthetic Data Sets
<p>Synthetic data sets used for machine learning of laser acceleration of protons</p>
Statistical error estimation from residual statistics of multiple collocated datasets: Data from synthetic experiments
<p>This archive contains the data used in the paper "How far can the statistical error estimation problem be closed by collocated data?" by A.Vogel and R.Menard (preprint available at: https://doi.org/10.5194/egusphere-2022-996) accepted for publication in Nonlinear Processes in Geophysics (NPG). </p> <p>The data refers to the synthetic experiments in Sect.5 of the paper which demonstrate the general ability to estimate statistical error covariances and cross-statistics from residual covariances, as well as the effects of inaccurate assumptions with respect to different setups.</p> <p>Further information on the data can be found in the README.txt file.</p>
Synthetic fuel scenario data
<p>Scenario data generated by AIM/Technology model for the global synfuel scenario analysis.</p>
Simulated and Synthetic Health Data: Improving Clinical Research on Rare Diseases. A Real-World Data Simulation of Autosomal Dominant Polycystic Kidney Disease (ADPKD) Trials. A Retrospective, Observa
ClinicalTrials.gov study NCT07016282. IPD Sharing: NO. Countries: 2. Publications: 27.
Development of Synthetic Medical Data Generation Technology to Predict Postoperative Complications
ClinicalTrials.gov study NCT05986474. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Data from: Analyzing negative feedback using a synthetic gene network expressed in the Drosophila melanogaster embryo
Open the record for dataset details and reuse information.
Data from: Two-scale dispersal estimation for biological invasions via synthetic likelihood
Open the record for dataset details and reuse information.
Data from: Classifying interactions in a synthetic bacterial community is hindered by inhibitory growth medium
Open the record for dataset details and reuse information.
Data from: Increases and fluctuations in nutrient availability do not promote dominance of alien plants in synthetic communities of common natives
Open the record for dataset details and reuse information.
Data from: Subgenome dominance in an interspecific hybrid, synthetic allopolyploid, and a 140-year-old naturally established neo-allopolyploid monkeyflower
Open the record for dataset details and reuse information.
Data from: Varying the spatial arrangement of synthetic herbivore-induced plant volatiles and companion plants to improve conservation biological control
Open the record for dataset details and reuse information.
Data for: Synthetic red supergiant explosion model grid for systematic characterization of Type II supernovae
Open the record for dataset details and reuse information.
Data from: Biomimicry of iridescent, patterned insect cuticles: comparison of biological and synthetic, cholesteric microcells using hyperspectral imaging
Open the record for dataset details and reuse information.
Data from: The effects of synthetic estrogen exposure on pre-mating and post-mating episodes of selection in sex-role-reversed Gulf pipefish
Open the record for dataset details and reuse information.
Data for: Grounding zone of Amery Ice Shelf, Antarctica, from differential synthetic-aperture radar interferometry
Open the record for dataset details and reuse information.
Automatic delineation of glacier grounding lines in differential interferometric synthetic-aperture radar data using deep learning
Open the record for dataset details and reuse information.
Synthetic Data Set for Uplift Modeling (One Trial)
<p>This dataset is designed and simulated for evaluating uplift modeling and feature selection methods.</p> <p>This dataset contains 10,000 samples and 36 features (one trial).</p> <p>The samples are equally split for control and treatment group.</p> <p>The generated data has three types of features: (1) uplift features influencing the treatment effect on the conversion probability; (2) classification features affecting the conversion probability but independent of the treatment effect; and (3) irrelevant features that are independent of both conversion probability and the treatment effect. To model the relationship between uplift features and the treatment effect and classification features and outcome probability, we implement six types of association patterns in the data generation process: linear, quadratic, cubic, ReLU (Rectified Linear Unit), trigonometric function sine, and cosine.</p> <p>In this data set, there are 36 features in total, including 10 classification features, 6 uplift features, and 20 irrelevant features.</p> <p>Column names:</p> <ul> <li>Experiment group label: 'treatment_group_key'</li> <li>Feature names: ['x1_informative',<br> 'x2_informative',<br> 'x3_informative',<br> 'x4_informative',<br> 'x5_informative',<br> 'x6_informative',<br> 'x7_informative',<br> 'x8_informative',<br> 'x9_informative',<br> 'x10_informative',<br> 'x11_irrelevant',<br> 'x12_irrelevant',<br> 'x13_irrelevant',<br> 'x14_irrelevant',<br> 'x15_irrelevant',<br> 'x16_irrelevant',<br> 'x17_irrelevant',<br> 'x18_irrelevant',<br> 'x19_irrelevant',<br> 'x20_irrelevant',<br> 'x21_irrelevant',<br> 'x22_irrelevant',<br> 'x23_irrelevant',<br> 'x24_irrelevant',<br> 'x25_irrelevant',<br> 'x26_irrelevant',<br> 'x27_irrelevant',<br> 'x28_irrelevant',<br> 'x29_irrelevant',<br> 'x30_irrelevant',<br> 'x31_uplift_increase',<br> 'x32_uplift_increase',<br> 'x33_uplift_increase',<br> 'x34_uplift_increase',<br> 'x35_uplift_increase',<br> 'x36_uplift_increase']</li> <li>Outcome variable: 'conversion'</li> <li>True underlying control conversion probability: 'control_conversion_prob'</li> <li>True underlying treatment conversion probability: 'treatment1_conversion_prob'</li> <li>True treatment effect: 'treatment1_true_effect'</li> <li>Note columns names with '_transformed' suffix are feature variables used in the intermediate steps during the data generation, that should be excluded for model training.</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.