Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
252
datasets available to search
ShareScore release 0.9.0
Dataset results
252 results for “Synthetic data”
Data from: A synthetic biology and green bioprocess approach to recreate agarwood sesquiterpenoid mixtures
<p>Certain endangered Thymelaeaceous trees are major sources of the fragrant and highly valued resinous agarwood, comprised of hundreds of oxygenated sesquiterpenoids (STPs). Despite growing pressure on natural agarwood sources, the chemical complexity of STPs severely limits synthetic production. Here, we catalogued the chemical diversity in 58 agarwood samples by two-dimensional gas chromatography–mass spectrometry and partially recreated complex STP mixtures through synthetic biology. We improved STP yields in the unicellular alga <em>Chlamydomonas reinhardtii </em>by combinatorial engineering to biosynthesise nine macrocyclic STP backbones found in agarwood. A bioprocess following green-chemistry principles was developed that exploits 'milking' of STPs without cell lysis, solvent–solvent STP extraction, solvent–STP nanofiltration, and bulk STP oxy-functionalisation to obtain terpene mixtures like those of agarwood. This process occurs with total solvent recycling and enables continuous production. Our synthetic-biology approach offers a sustainable alternative to harvesting agarwood trees to obtain mixtures of complex, fragrant, oxygenated STPs.</p>
Raw data of experiments on the Bisexual lures and their comparison with synthetic sex attractants for trapping Orthosia species (Lepidoptera: Noctuidae)
<p>Raw data of experiments on the manuscript of the "Bisexual lures and their comparison with synthetic sex attractants for trapping Orthosia species (Lepidoptera: Noctuidae)."</p>
Supplementary data for 'Ferrofluid impregnation efficiency and its spatial variability in natural and synthetic porous media: Implications for magnetic pore fabric studies'
<p>Supplementary data for the manuscript 'Ferrofluid impregnation efficiency and its spatial variability in natural and synthetic porous media: Implications for magnetic pore fabric studies'</p>
A synthetic fraud detection data set.
<p>A synthetic fraud detection data set created using sklearn's make_blob for use in a blog.</p> <p>X, y = datasets.make_blobs(n_samples=[800000,200000], centers=None, cluster_std=[10.0, 2],random_state=42,n_features=4)</p>
Training dataset for "A deep learned nanowire segmentation model using synthetic data augmentation"
<p>This image dataset contains synthetic structure images used for training the deep-learning based nanowire segmentation model presented in our work "A deep learned nanowire segmentation model using synthetic data augmentation" to be published in <em>npj Computational materials. </em>Detailed information can be found in the corresponding article.</p>
Untargeted metabolomics data for the publication Weiss et al. 2022 "In vitro interaction network of a synthetic gut bacterial community"
<p>This dataset contains the untargeted metabolomics data for the publication Weiss et al. 2022 "In vitro interaction network of a synthetic gut bacterial community". The dataset has also been submitted to MetaboLights repository with ID "MTBLS3535". Please refer to the MetaboLights repository for the most up-to-date datasets. </p> <p>Publication abstract:</p> <p>A key challenge in microbiome research is to predict the functionality of microbial communities based on community membership and (meta)-genomic data. As central microbiota functions are determined by bacterial community networks, it is important to gain insight into the principles that govern bacteria-bacteria interactions. Here, we focused on the growth and metabolic interactions of the Oligo-Mouse-Microbiota (OMM<sup>12</sup>) synthetic bacterial community, which is increasingly used as a model system in gut microbiome research. Using a bottom-up approach, we uncovered the directionality of strain-strain interactions in mono- and pairwise co-culture experiments as well as in community batch culture. Metabolic network reconstruction in combination with metabolomics analysis of bacterial culture supernatants provided insights into the metabolic potential and activity of the individual community members. Thereby, we could show that the OMM<sup>12</sup> interaction network is shaped by both exploitative and interference competition in vitro in nutrient-rich culture media and demonstrate how community structure can be shifted by changing the nutritional environment. In particular, <em>Enterococcus faecalis</em> KB1 was identified as an important driver of community composition by affecting the abundance of several other consortium members in vitro. As a result, this study gives fundamental insight into key drivers and mechanistic basis of the OMM<sup>12</sup> interaction network in vitro, which serves as a knowledge base for future mechanistic in vivo studies.</p>
synthetic 3DBOS test data
<p>This data is test data for the code accompanying the paper</p> <p><strong>Atcheson, Ihrke, Heidrich, Tevs, Bradley, Magnor, Seidel<br> "Time-resolved 3d capture of non-stationary gas flows"<br> Siggraph Asia 2008</strong></p> <p><br> Author: Ivo Ihrke (2007)</p> <p>Contact: ivo [dot] ihrke [at] uni-siegen [dot] de</p> <p>Note that the format is custom; the necessary code for processing it is in preparation of open-sourcing.</p>
Synthetic data for R training
<p>Synthetic data for R training</p>
Data for "Magic Angle Spinning Solid-State 13C Photochemically Induced Dynamic Nuclear Polarization by a Synthetic Donor–Chromophore–Acceptor System at 9.4 T"
<p>NMR data and photo-CIDNP-enhanced NMR data for "Magic Angle Spinning Solid-State 13C Photochemically Induced Dynamic Nuclear Polarization by a Synthetic Donor–Chromophore–Acceptor System at 9.4 T".</p> <p>All data but those relative to the spectra in Figure S13 are provided in Bruker format. For the low field experiments (Figure S13), raw free induction decay NMR data are provided, together with a processing script written in Wolfram Mathematica.</p>
Data pertaining to the published article "Detection of pathological contrast enhancement with synthetic brain imaging from quantitative multiparametric MRI" by Donatelli et al., 2024
<p>Data pertaining to the published article "Detection of pathological contrast enhancement with synthetic brain imaging from quantitative multiparametric MRI" by Donatelli et al., 2024. <a href="https://doi.org/10.1111/jon.13201">https://doi.org/10.1111/jon.13201</a></p>
Synthetic benchmarking data from "Comprehensive Benchmarking of SNV Callers for Highly Admixed Tumor Data"
<p>Synthetic benchmarking data simulating heterogeneous and admixed tumor data with implanted SNVs and indels at different admixture levels (0% to 90%) for targeted sequencing (exome and targeted gene panel).</p>
Research data supporting "Bioinspired fabrication of DNA-inorganic hybrid composites using synthetic DNA"
<p>Raw data supporting the publication:</p> <p>Kim E., et al., "Bio-inspired fabrication of DNA-inorganic hybrid composites using synthetic DNA.", ACS Nano, 2019, DOI: 10.1020/acsnano.8b06492.</p>
DCASE2019_task4_synthetic_data
<p><strong>Synthetic data for DCASE 2019 task 4</strong></p> <p>Freesound dataset [1,2]: A subset of <a href="https://datasets.freesound.org/fsd/">FSD</a> is used as foreground sound events for the synthetic subset of the dataset for DCASE 2019 task 4. FSD is a large-scale, general-purpose audio dataset composed of Freesound content annotated with labels from the AudioSet Ontology [3].</p> <p>SINS dataset [4]: The derivative of the SINS dataset used for DCASE2018 task 5 is used as background for the synthetic subset of the dataset for DCASE 2019 task 4.<br> The SINS dataset contains a continuous recording of one person living in a vacation home over a period of one week.<br> It was collected using a network of 13 microphone arrays distributed over the entire home.<br> The microphone array consists of 4 linearly arranged microphones.</p> <p>The synthetic set is composed of 10 sec audio clips generated with <a href="https://github.com/justinsalamon/scaper">Scaper</a> [5]. The foreground events are obtained from FSD. Each event audio clip was verified manually to ensure that the sound quality and the event-to-background ratio were sufficient to be used an isolated event. We also verified that the event was actually dominant in the clip and we controlled if the event onset and offset are present in the clip. Each selected clip was then segmented when needed to remove silences before and after the event and between events when the file contained multiple occurrences of the event class.</p> <p><strong>License:</strong></p> <p>All sounds comming from FSD are released under Creative Commons licences. <strong>Synthetic sounds can only be used for competition purposes until the full CC license list is made available at the end of the competition.</strong></p> <p> </p> <p>Further information on <a href="http://dcase.community/challenge2019/task-sound-event-detection-in-domestic-environments">dcase website.</a></p> <p>References:</p> <ul> <li>[1] F. Font, G. Roma & X. Serra. Freesound technical demo. In Proceedings of the 21st ACM international conference on Multimedia. ACM, 2013.</li> <li> [2] E. Fonseca, J. Pons, X. Favory, F. Font, D. Bogdanov, A. Ferraro, S. Oramas, A. Porter & X. Serra. Freesound Datasets: A Platform for the Creation of Open Audio Datasets. In Proceedings of the 18th International Society for Music Information Retrieval Conference, Suzhou, China, 2017.</li> <li>[3] Jort F. Gemmeke and Daniel P. W. Ellis and Dylan Freedman and Aren Jansen and Wade Lawrence and R. Channing Moore and Manoj Plakal and Marvin Ritter. Audio Set: An ontology and human-labeled dataset for audio events. In Proceedings IEEE ICASSP 2017, New Orleans, LA, 2017.</li> <li> <p>[4] Gert Dekkers, Steven Lauwereins, Bart Thoen, Mulu Weldegebreal Adhana, Henk Brouckxon, Toon van Waterschoot, Bart Vanrumste, Marian Verhelst, and Peter Karsmakers.<br> The SINS database for detection of daily activities in a home environment using an acoustic sensor network.<br> In Proceedings of the Detection and Classification of Acoustic Scenes and Events 2017 Workshop (DCASE2017), 32–36. November 2017.</p> </li> <li>[5] J. Salamon, D. MacConnell, M. Cartwright, P. Li, and J. P. Bello. Scaper: A library for soundscape synthesis and augmentation In IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), New Paltz, NY, USA, Oct. 2017.</li> </ul>
Fully synthetic longitudinal real-world data from hearing aid wearers for public health policy modeling
<p>Real-world data from hearing aids and Bluetooth connected smartphones. The associated data report can be found here: <a href="https://doi.org/10.3389/fnins.2019.00850">https://doi.org/10.3389/fnins.2019.00850</a></p>
SH component synthetic seismograms (SE_Supplementary_data)
<p>Includes synthetic data files generated with pseudospectral method. See README for file format details.</p>
Synthetic Pelagic Biomass Size Spectra of the Tropical and Subtropical Atlantic - biovolume and carbon biomass data
<p>Synthetic Pelagic Biomass Size Spectra of the Tropical and Subtropical Atlantic</p> <p>Normalized size spectra data are presented for (a) biovolume and (b) for carbon contents for the following ecosystem components: Phytoplankton, zooplankton and micronekton</p> <p>Dataset Biovolume_NBSS contains the following variables, three heading lines</p> <table> <tbody> <tr> <td> <p>COLUMN</p> </td> <td> <p>HEADER</p> </td> <td> <p>DESCRIPTION</p> </td> </tr> <tr> <td> <p>1</p> </td> <td> <p>ConsecutiveNumber</p> </td> <td> <p>Control number</p> </td> </tr> <tr> <td> <p>2</p> </td> <td> <p>Cruise</p> </td> <td> <p>Cruise name</p> </td> </tr> <tr> <td> <p>3</p> </td> <td> <p>Reference</p> </td> <td> <p>Cruise/data record reference</p> </td> </tr> <tr> <td> <p>4</p> </td> <td> <p>CruiseID</p> </td> <td> <p>ID</p> </td> </tr> <tr> <td> <p>5</p> </td> <td> <p>StationID</p> </td> <td> <p>ID</p> </td> </tr> <tr> <td> <p>6</p> </td> <td> <p>Net_index</p> </td> <td> <p>ID, opt.</p> </td> </tr> <tr> <td> <p>7</p> </td> <td> <p>Date</p> </td> <td> <p>YYYY-MM-DD</p> </td> </tr> <tr> <td> <p>8</p> </td> <td> <p>Longitude</p> </td> <td> <p>Position, decimal degrees</p> </td> </tr> <tr> <td> <p>9</p> </td> <td> <p>Latitude</p> </td> <td> <p>Position, decimal degrees</p> </td> </tr> <tr> <td> <p>10</p> </td> <td> <p>Target organisms</p> </td> <td> <p>Phytoplankton</p> <p>Zooplankton</p> <p>Detritus + zooplankton (only for UVP)</p> <p>Mesopelagic fishes</p> <p>invMicronekton – invertebrate micronekton only</p> <p>totMicronekton – mesopelagic fishes + invertebrate Micronekton</p> </td> </tr> <tr> <td> <p>11</p> </td> <td> <p>Gear</p> </td> <td> <p>Gear applied</p> </td> </tr> <tr> <td> <p>12</p> </td> <td> <p>Operation mode</p> </td> <td> <p>Gear operation mode</p> </td> </tr> <tr> <td> <p>13</p> </td> <td> <p>Catching Depth max [m]</p> </td> <td> <p>Catching depth maximum</p> </td> </tr> <tr> <td> <p>14</p> </td> <td> <p>Catching Depth min [m]</p> </td> <td> <p>Catching depth minimum</p> </td> </tr> <tr> <td> <p>15</p> </td> <td> <p>Biomass determination</p> </td> <td> <p>Biomass Determination method</p> </td> </tr> <tr> <td> <p>16</p> </td> <td> <p>Contact Person</p> </td> <td> <p>Contact person</p> </td> </tr> <tr> <td> <p>17</p> </td> <td> <p>Day/night</p> </td> <td> <p>Sampling time</p> </td> </tr> <tr> <td> <p>18</p> </td> <td> <p>Region</p> </td> <td> <p>Region affiliation</p> </td> </tr> <tr> <td> <p>19-74</p> </td> <td>mm3 m-3 mm-3</td> <td> <p>Biovolume data normalized, additionally with reference to "Size class interval (mm-mm]" and "log mm3/individual"</p> </td> </tr> </tbody> </table> <p> </p> <p> </p> <p>Dataset Carbon_NBSS contains the following variables, three heading lines</p> <table> <tbody> <tr> <td> <p>COLUMN</p> </td> <td> <p>HEADER</p> </td> <td> <p>DESCRIPTION</p> </td> </tr> <tr> <td> <p>1</p> </td> <td> <p>ConsecutiveNumber</p> </td> <td> <p>Control number</p> </td> </tr> <tr> <td> <p>2</p> </td> <td> <p>Cruise</p> </td> <td> <p>Cruise name</p> </td> </tr> <tr> <td> <p>3</p> </td> <td> <p>Reference</p> </td> <td> <p>Cruise/data record reference</p> </td> </tr> <tr> <td> <p>4</p> </td> <td> <p>CruiseID</p> </td> <td> <p>ID</p> </td> </tr> <tr> <td> <p>5</p> </td> <td> <p>StationID</p> </td> <td> <p>ID</p> </td> </tr> <tr> <td> <p>6</p> </td> <td> <p>Net_index</p> </td> <td> <p>ID, opt.</p> </td> </tr> <tr> <td> <p>7</p> </td> <td> <p>Date</p> </td> <td> <p>YYYY-MM-DD</p> </td> </tr> <tr> <td> <p>8</p> </td> <td> <p>Longitude</p> </td> <td> <p>Position, decimal degrees</p> </td> </tr> <tr> <td> <p>9</p> </td> <td> <p>Latitude</p> </td> <td> <p>Position, decimal degrees</p> </td> </tr> <tr> <td> <p>10</p> </td> <td> <p>target organisms</p> </td> <td> <p>Phytoplankton</p> <p>Zooplankton</p> <p>Mesopelagic fishes</p> <p>invMicronekton – invertebrate micronekton only</p> <p>totMicronekton – mesopelagic fishes + invertebrate Micronekton</p> </td> </tr> <tr> <td> <p>11</p> </td> <td> <p>Gear</p> </td> <td> <p>Gear applied</p> </td> </tr> <tr> <td> <p>12</p> </td> <td> <p>Operation mode</p> </td> <td> <p>Gear operation mode</p> </td> </tr> <tr> <td> <p>13</p> </td> <td> <p>Catching Depth max [m]</p> </td> <td> <p>Catching depth maximum</p> </td> </tr> <tr> <td> <p>14</p> </td> <td> <p>Catching Depth min [m]</p> </td> <td> <p>Catching depth minimum</p> </td> </tr> <tr> <td> <p>15</p> </td> <td> <p>Biomass determination</p> </td> <td> <p>Biomass Determination method</p> </td> </tr> <tr> <td> <p>16</p> </td> <td> <p>Contact Person</p> </td> <td> <p>Contact person</p> </td> </tr> <tr> <td> <p>17</p> </td> <td> <p>filtered volume [m3]</p> </td> <td> <p> </p> </td> </tr> <tr> <td> <p>18</p> </td> <td> <p>Day/night</p> </td> <td> <p>Sampling time</p> </td> </tr> <tr> <td> <p>19</p> </td> <td> <p>SST (C)</p> </td> <td> <p>In situ SST</p> </td> </tr> <tr> <td> <p>20</p> </td> <td> <p>Temperature (C)</p> </td> <td> <p>In situ temperature</p> </td> </tr> <tr> <td> <p>21</p> </td> <td> <p>Salinity</p> </td> <td> <p>In situ salinity (PSU)</p> </td> </tr> <tr> <td> <p>22</p> </td> <td> <p>Oxygen (umol/kg)</p> </td> <td> <p>In situ oxygen (µmol/kg)</p> </td> </tr> <tr> <td> <p>23</p> </td> <td> <p>Region</p> </td> <td> <p>Region affiliation</p> </td> </tr> <tr> <td> <p>24-79</p> </td> <td>gC m-3 g-1C</td> <td> <p>Carbon biomass data normalized, additionally with reference to " Size class number" and "</p> <p>Exponent"</p> </td> </tr> </tbody> </table> <p> </p>
Raw ptychographic synthetic data for manuscript with title 'Purity-based self-calibration in ptychography'
<ul> <li>Here given are the dataset for purity scan, numerically created based on the synthetic setup in the manuscript titled "Purity-based self-calibration in ptychography". The dataset includes raw diffraction patterns, preprocessed diffraction patterns for reconstruction, as well as the preprocessed script.</li> <li>For the two experimental verificatoin cases, the preprocessed diffraction patterns are provided, where the scanning grid is included.</li> <li>The reconstruction, calculation of purity, and zPIE were conducted at open-source PtyLab framework.</li> <li>Experimental data were measured at the Institute of Applied Physics in Jena using a Fiber Laser driven High-order harmonic source, which can be referenced in </li> </ul> <p>Please contact me (liu.chang@uni-jena.de) for additional support.</p>
synthetic data for scDOT https://www.biorxiv.org/content/10.1101/2023.08.16.553591v1
<p>h5ad files are for synthetic data 1, csv files are for synthetic data 2. The code for processing these data is on GitHub: https://github.com/namtk/scDOT/blob/main/src/data.py</p> <p>For synthetic data 1: data=='simu_spatial'</p> <p>For synthetic data 2: data=='synthetic'</p>
Benchmark datasets to study fairness in synthetic data generation
<p>The traveltime dataset is based on the Folktables project covering US census data. The target is a binary variable encoding whether or not the individual needs to travel more than 20 minutes for work; here, having a shorter travel time is the desirable outcome. We use a subset of data from the states of California, Florida, Maine, New York, Utah, and Wyoming states in 2018. Although the folktables dataset does not have any missing values, there are some values recorded as NaN due to the Bureau's data collection methodology. We remove the "esp" column, which encodes the employment status of parents, and has 99.55% missing values. We encode the missing values in the povpip, income to poverty ratio (0.85%), to -1 in accordance to the methodology in Ding et al.. See https://arxiv.org/pdf/2108.04884 for metadata.</p> <p>The cardio (a) dataset contains patient data recorded during medical examination, including 3 binary features supplied by the patient. The target class denotes the presence of cardiovascular disease. This dataset represents predictive tasks that allocate access to priority medical care for patients, and has been used for fairness evaluations in the domain.</p> <p>The credit dataset contains historical financial data of borrowers, including past non-serious delinquencies. Here, a serious delinquency is considered to be 90 days past due, and this is the target variable.</p> <p>The German Credit dataset (https://archive.ics.uci.edu/dataset/144/statlog+german+credit+data) contains financial and personal information regarding loan-seeking applicants.</p>
Data from: Genome-wide CRISPR synthetic lethality screen identifies a role for the ADP-ribosyltransferase PARP14 in replication fork stability controlled by ATR
<p>The DNA damage response is essential to maintain genomic stability, suppress replication stress, and protect against carcinogenesis. The ATR-CHK1 pathway is an essential component of this response, which regulates cell cycle progression in the face of replication stress. PARP14 is an ADP-ribosyltransferase with multiple roles in transcription, signaling, and DNA repair. To understand the biological functions of PARP14, we catalogued the genetic components that impact cellular viability upon loss of PARP14 by performing an unbiased, comprehensive, genome-wide CRISPR knockout genetic screen in PARP14-deficient cells. We uncovered the ATR-CHK1 pathway as essential for viability of PARP14-deficient cells, and identified regulation of DNA replication dynamics as an important mechanistic contributor to the synthetic lethality observed. Our work shows that PARP14 is an important modulator of the response to ATR-CHK1 pathway inhibitors.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.