Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
250
datasets available to search
ShareScore release 0.9.0
Dataset results
250 results for “Synthetic Dataset”
Synthetic Dataset of Emergency Healthcare Services
<p>Synthetic dataset of emergency services comprised of several CSV files that we have generated using a simulation software. This dataset is open for public use; please cite our work if used in research or applications.</p>
ChemProp2 Dataset: Investigating potential biotransformations with a panel of 45 Drugs on Synthetic Community (Com20) to elucidate Drug-Microbiome-Host Dynamics
<p>This dataset contains the examination of each of the 45 drugs on the synthetic community Com20, conducted at two distinct timepoints: initially (time 0) and after a 2-hour interval.<br>List of drugs: Ethopropazine, Methotrexate, Felodipine, Miconazole, Floxuridine, Nalidixic acid, Fluconazole, Niclosamide, Ketoconazole, Omeprazole, Lacidipine, Pentamidine isothionate, DMSO, Promethazine, Lansoprazole, Protriptyline, Loratadine, Sertindole, Loxapine, Simvastatin, L-Thyroxine, Streptozotocin, Metformin, Tamoxifen, Tazobactam, Doxorubicin, Telmisartan, Ofloxacin, Terfenadine, Oxolinic acid, Thioguanosine, Tobramycin, Tiratricol, Vancomycin, Water, Rivaroxoban, Metronidazole, Tolfenamic acid, Clarithromycin, Tribenoside, Norfloxacin, Zafirlukast, Montelukast, Sertraline, Fluoxetine, Clomipramin, Novobiocin</p>
ChemProp2 Dataset: Investigating potential biotransformations with a panel of 23 Drugs on Synthetic Community (Com20) to elucidate Drug-Microbiome-Host Dynamics
<p>Tested 23 drugs against Synthetic community (Com20) at two different timepoints (t=0, t= 2 hrs) to look for potential biotransformations<br><br>List of drugs tested:<br>Acarbose, Clemizole, Amlodipine, Clindamycin, Amoxicillin, Clomifen, Aprepitant, Clotrimazole, Avermectin B1, Diacerein, Azithromycin, Dicumarol, Control, Dienestrol, Benzbromarone, Dienogest, Ceterizin, Doxycycline hyclate, Chlorpromazine, Duloxetin, Chlorprothixene, Erythromycin, Cilnidipine, Ethinylestradiol<br><br>Samples were measured (in March 18/19, 2023) using an UHPLC-Q Exactive HF Orbitrap mass spectrometer, equipped with a C18 column, following the LC-MS/MS method.</p>
Synthetic ECG dataset
<p>Synthetic dataset with 6888 clean and 6888 noisy ElectroCardioGram (ECG) signals with various levels of strong drifts and random noise (SNR-Signal-to-noise-ratio=-7dB). </p> <p>ECG signals (i.e. duration of 10 seconds) with 30000 samples per ECG signal and which vary between 60 heart beats per minute to 100 heart beats per minute. The voltage varies between 1 mV to 3 mV.</p>
Synthetic dataset for the testing of an MT-Mag geophysical integration workflow
<p>This datasets is related to the manuscript "<strong>Utilisation of probabilistic MT inversions to constrain magnetic data inversion: proof-of-concept and field application</strong>", by Jérémie Giraud, Hoël Seillé, Gerhard Visser, Mark D. Lindsay, Vitaliy Ogarko, and Mark W. Jessell, intended for publication in Solid Earth. </p> <p>It is organised as follows. <br> <br> MT FOLDER<br> model subfolder: contains the synthetic resistivity model, 2 formats available:<br> - ModEM format (.mod)<br> - WinGLink format (.out)<br> <br> responses subfolder: contains the synthetic model responses, 2 formats available:<br> - ModEM format (.dat): This data has not been perturbed by synthetic noise. <br> - EDI format (.edi): This data has been perturbed by 5% Gaussian noise, this is the data used in the synthetic part of the study.<br> <br> coordinates of the synthetic MT sites with respect to the model: coordinates.txt <br> - it assumes the origin (0,0) in the center of the synthetic 3D model.<br> <br> Mag FOLDER<br> model subfolder: contains the magnetic susceptibility model at the Tomofast format<br> - mag_voxet_true_model.txt <br> <br> responses subfolder: contains the synthetic model responses, for both the noisy and clean data. <br> - fwd_mag_data_clean.txt<br> - fwd_mag_data_with_noise.txt<br> <br> Rock units FOLDER: contains the file with indices of the rock unit model, in 3D, of the modified Mansfield model. <br> It is of dimensions 128 * 128 * 36. The indices are stored as a column vector. </p>
ACPAS dataset: Aligned Classical Piano Audio and Score (synthetic subset)
<p><strong>ACPAS</strong> is a dataset with aligned audio and scores for classical piano music containing 497 distinct music scores aligned with 2189 performances, in total 179.77 hours. For each performance, we provide the corresponding performance audio (real recording or synthesized recording), performance MIDI, and MIDI score, together with rhythm and key annotations.</p> <p>This is the <strong>Synthetic subset</strong> of the ACPAS dataset. To download the full dataset and for dataset details, please refer to the dataset webpage at <a href="https://cheriell.github.io/research/ACPAS_dataset">https://cheriell.github.io/research/ACPAS_dataset</a></p> <p>For any questions, suggestions, or comments, please do not hesitate to contact <a href="mailto:lele.liu@qmul.ac.uk">lele.liu@qmul.ac.uk</a></p> <p><strong>How to cite:</strong></p> <p>- Lele Liu, Veronica Morfi, and Emmanouil Benetos, "ACPAS: A Dataset of Aligned Classical Piano Audio and Scores for Audio-to-Score Transcription," in ISMIR Late-breaking Demo, 2021.</p> <p><strong>Funding:</strong></p> <p>L. Liu is a research student at the UKRI Centre for Doctoral Training in Artificial Intelligence and Music, supported jointly by the China Scholarship Council and Queen Mary University of London.</p>
Qualified Synthetic Dataset 2.0 for Semiconductor Order Lead Times
<p>This is the updated version of the prior uploaded dataset. It contains the most recent data.</p> <p>The created data is collected from an algorithm for measuring Infineon's customer Order Lead Times. Due to the frequent changes in e.g. volumes or Confirmed Delivery dates, an algorithm is necessary to calculate and thereby measure the correct Order Lead Times since the SAP data is misleading here. As the Lead Time data is confidential, Qualified Synthetic data will be created that share the same distribution and characteristics, but do not depict sensible customer information. The distribution of products and customers will be taken into account as well, but in encoded form both for security and confidentiality reasons. The final table of data will include Product Line, Business Month, Order Entry date, Requested and confirmed order lead time, customer name encoded, product name encoded, Order number and order volume. This is the updated dataset, containing additional data points.</p> <p> </p> <p>This project has received funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement No 825225. <a href="https://safe-deed.eu/">https://safe-deed.eu/</a></p>
Qualified Synthetic Dataset 3.0 for Semiconductor Order Lead Times
<p>This is the updated version (3.0) of the prior uploaded dataset. It contains the most recent data.</p> <p>The created data is collected from an algorithm for measuring Infineon's customer Order Lead Times. Due to the frequent changes in e.g. volumes or Confirmed Delivery dates, an algorithm is necessary to calculate and thereby measure the correct Order Lead Times since the SAP data is misleading here. As the Lead Time data is confidential, Qualified Synthetic data will be created that share the same distribution and characteristics, but do not depict sensible customer information. The distribution of products and customers will be taken into account as well, but in encoded form both for security and confidentiality reasons. The final table of data will include Product Line, Business Month, Order Entry date, Requested and confirmed order lead time, customer name encoded, product name encoded, Order number and order volume. This is the updated dataset, containing additional data points.</p> <p> </p> <p> </p> <p>This project has received funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement No 825225. <a href="https://safe-deed.eu/">https://safe-deed.eu/</a></p>
Synthetic datasets used for numerical testing of geology-geophyiscs integration
<p>This datasets is a companion dataset to the manuscript "<strong>Integration of automatic implicit geological modelling in geophysical inversion with posterior topological analysis</strong>", by Jérémie Giraud, Guillaume Caumon, Lachlan Grose, Vitaliy Ogarko, and Paul Cupillard, for publication in Solid Earth. <br> </p> <p>It contains models and data shown in the paper that are not available elsewhere.<br> <br> The folder organisation is as follows, where <strong>bold</strong> refers to folders and subfolders, and text in <em>italic</em> corresponds to a succinct description of the contents.</p> <p> </p> <p>|-- <strong>synthetic 1 </strong>> <em>synthetic dataset and results for the first synthetic example</em><br> | |-- <strong>geol_data_layered_model.pckl </strong>> <em>geological data and model</em><br> | |-- <strong>grav_data_synthetic1.pckl </strong>> <em>gravity data produced by the true model</em><br> | |-- <strong>inverted_model_no_correction.txt </strong>> <em> inversion results</em><br> | |-- <strong>inverted_model_with_correction.txt </strong>> <em> inversion results</em><br> | |--<strong> note.txt </strong>> <em> metadata</em><br> | |--<strong> starting_model.txt </strong>> <em> starting model for inversion</em><br> | |--<strong> true_model.txt </strong>> <em> true model</em><br> <br> |-- <strong>synthetic 2 </strong>> <em>synthetic dataset and results for the second synthetic example</em><br> | |-- <strong>case<num>.pckl </strong>> <em>inversion results for case with number 1..5 as in the manuscript</em><br> | |-- <strong>grav_data_synthetic2.pckl </strong>> <em>gravity data produced by the true model</em><br> | |-- <strong>inverted_model_no_correction.txt </strong>> <em> inversion results</em><br> | |-- <strong>inverted_model_with_correction.txt </strong>> <em> inversion results</em><br> | |--<strong> note.txt </strong>> <em> metadata</em><br> | |--<strong> starting_model_case5.txt </strong>> <em> starting_model_case5</em><br> | |--<strong> model_start_unconformity.pckl </strong>> <em> starting model for inversion, cases 1..5.</em><br> | |--<strong> true_mod_geol_data.pckl </strong>> <em> true model and geological data</em><br> | |--<strong> start_model_and_geol_data.pckl </strong>> <em> starting geological model and corresponding geological data</em></p> <p> </p>
Insights from Synthetic Star-forming Regions: Synthetic dataset "observed" in IRAC1, IRAC2 & MIPS1 at 1 to 15 kpc
<p>Synthetic dataset "observed" in IRAC1, IRAC2 & MIPS1 at 1 to 15 kpc for 3 different orientations. The meaning of the abbreviations in the file names is the same as described in the appendix of Koepferl et al. (2016; https://arxiv.org/abs/1603.02270). </p> <p> </p> <p>Please cite the following papers: </p> <p>http://adsabs.harvard.edu/abs/2017ApJ...849….3K<br> http://adsabs.harvard.edu/abs/2017ApJS..233....1K</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.