Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

250

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

250 results for “Synthetic dataset”

Learn how ShareScore rates datasets ↗
zenodo52/100

HD-SIM-RBV: a synthetic dataset with model-based simulations of blood volume changes during hemodialysis

<p>The HD-SIM-RBV dataset is a synthetic (model-based) dataset generated to enable the study of blood volume (BV) or relative blood volume (RBV) changes during hemodialysis (HD).</p> <p>The dataset includes the profiles of BV changes during a standard 4-hour HD session simulated using a lumped-parameter, physiologically-based model of the cardiovascular system and the whole-body water and solute kinetics in 5,000 virtual patients with randomly adjusted values of 90 physiological parameters.</p> <p>For each of the 90 selected parameters, a random value was drawn from a normal distribution with the mean equal to the baseline value used originally in the model (with a few exceptions) and the standard deviation (SD) assumed at the level of 10%, 20%, or 40% of the baseline value, depending on the nature of the given parameter and the likelihood of its variation in the population (for some parameters, SD was set below 10% - see Parameters.xlsx). Only values within &plusmn;2SD from the mean were accepted. &nbsp;</p> <p>Ultrafiltration was set randomly within &plusmn;1 L from the assigned fluid overload. &nbsp;All other parameters as well as dialysis settings were kept constant for all virtual patients (at the levels used in our previous work - see the references below).</p> <p>&nbsp;</p> <p>When using the dataset, please cite the associated conference paper:</p> <p>Pstras L, Waniewski J. A Model-Based Dataset for In-Silico Exploration of the Patterns of Relative Blood Volume Changes During Hemodialysis. 2023 IEEE EMBS Special Topic Conference on Data Science and Engineering in Healthcare, Medicine and Biology, 149-150, 2023, doi: 10.1109/IEEECONF58974.2023.10404528.</p>

opencc-zeroOct 2023View details →
zenodo52/100

Synthetic Dataset of Citation Strings in 12 Styles

<p>This dataset was produced in the aim of testing different tools for citation string parsing, as part of the experiment reported in the paper:</p> <blockquote> <p>Iana Atanassova and Marc Bertin, 2024. "Breaking Boundaries in Citation Parsing: A Comparative Study of Generative LLMs and Traditional Out-of-the-box Citation Parsers", Bibliometric-enhanced Information Retrieval workshop (BIR), collocated with ECIR 2024, Glasgow, Scotland.&nbsp;</p> </blockquote> <h2><br>Data</h2> <p>The data that is provided here is organised as follows:</p> <ul> <li>the file <strong>citation-strings.zip</strong> contains raw citation strings that were generated for each of the 12 citation styles in txt format</li> <li>the file <strong>parsers-output.csv</strong> contains the output that was produced from the parsers: ChatGPT, Llama, and Neural ParsCit</li> </ul> <h2><br>To cite this work</h2> <p>To use this dataset and/or the results produced in the experiment, please cite the following article:</p> <blockquote> <p>@inproceedings{atanassova2024citparse,<br>&nbsp; &nbsp; title = {{Breaking Boundaries in Citation Parsing: A Comparative Study of Generative LLMs and Traditional Out-of-the-box Citation Parsers}},&nbsp;<br>&nbsp; &nbsp; author = {Iana Atanassova and Marc Bertin},<br>&nbsp; &nbsp; year = {2024},<br>&nbsp; &nbsp; booktitle = {{International Workshop on Bibliometric-enhanced Information Retrieval (BIR 2024) co-located with the 46\textsuperscript{st} European Conference on Information Retrieval (ECIR 2024)}},<br>&nbsp; &nbsp; address = {Glasgow, Scotland}<br>}</p> </blockquote> <h3>Authors information</h3> <ul> <li>Iana Atanassova, ORCID https://orcid.org/0000-0003-3571-4006 URL https://iana-atanassova.github.io/</li> <li>Marc Bertin, ORCID https://orcid.org/0000-0003-1803-6952 URL https://elico-recherche.msh-lse.fr/membres/marc-bertin</li> </ul> <h3>Related github repository</h3> <p>https://github.com/iana-atanassova/citation-parsers-bir2024.git&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo52/100

Dataset of "MoO3-xNiMoO4 nanorods synthetized using NiO nanoparticles for hydrogen evolution in anion exchange membrane water electrolysis"

<p>Novel method of Mo-Ni catalyst for hydrogen evolution reaction in anion exchange membrane water electrolysis was used. Complete physico-chemical and electrochemical characterization was done. Prepared material showed enhanced performance when compared to the similar Ni based materials. Physico-chemical characterization showed, that final material is formed by NiMoO4 nanorods coverd on the surface by the layer of the MoO3-x.</p>

opencc-by-4.0Oct 2024View details →
zenodo52/100

Dataset - Generating reliable estimates of tropical cyclone induced coastal hazards along the Bay of Bengal for current and future climates using synthetic tracks

<p>This data is complementary to the paper by Leijnse et al. 2022 &quot;Generating reliable estimates of tropical cyclone induced coastal hazards along the Bay of Bengal for current and future climates using synthetic tracks&quot;&nbsp;<br> https://doi.org/10.5194/nhess-2021-181</p> <p>This data is made available in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE</p> <p>For questions about the data ask: tim.leijnse@deltares.nl</p> <p>For more information about the tool to generate the used synthetic tracks TCWiSE see:&nbsp;<a href="https://www.deltares.nl/en/software/tcwise/">https://www.deltares.nl/en/software/tcwise/</a></p> <p>&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo52/100

AraucanaXAI - HEPAR synthetic datasets

<p>Supporting datasets (iid and ood) used in the evaluation experiments of the paper &quot;Why did AI get this one wrong? - tree-based explanations of machine learning model predictions&quot; by Parimbelli, Buonocore, Nicora, Michalowski, Wilk and Bellazzi.</p>

opencc-by-4.0Jun 2022View details →
zenodo52/100

The extrAIM dataset: A merged satellite-based daily precipitation dataset for the Mediterranean region (including an ensemble of 20 synthetic realisations)

<p><strong>extrAIM </strong>dataset is a <strong>new merged daily precipitation product</strong> (extraim_merged_data.nc) for the Mediterranean region with the following characteristics:</p> <ul> <li><strong>Dataset format:</strong> NetCDF</li> <li><strong>Spatial resolution:</strong> 25 x 25 km</li> <li><strong>Temporal resolution:</strong> 1 day</li> <li><strong>Spatial coverage:</strong> Longitude: from -6.25 to 38.25, Latitude: 27.75 to 49</li> <li><strong>Temporal coverage:&nbsp;</strong>01-01-2007 to 30-09-2021</li> <li><strong>Merging approach:&nbsp;</strong>Two-step merging (classification and regression) <ul> <li><strong>Algorithm:&nbsp;</strong>Random Forest for both classification and regression</li> <li><strong>Training strategy:</strong> Full training strategy</li> </ul> </li> <li><strong>Merged precipitation products: </strong>SM2Rain-ASCAT and GPM Late Run</li> <li><strong>Reference precipitation product:</strong> EMO5</li> <li><strong>Static covariates: </strong>Longitude, Latitude and Elevation, in both classification and regression step <ul> <li><strong>Classification step:</strong> probability dry and probability dry of the 5 neighboring points around the target locations</li> <li><strong>Regression step:</strong> mean, standard deviation and skewness of daily precipitation, of the entire series and non-zero amounts, as well as mean precipitation of the 5 neighboring points around the target locations</li> </ul> </li> </ul> <p>In addition, an <strong>ensemble of 20 synthetic realizations</strong> (equiprobable and bias-adjusted) of the merged dataset is provided (files named: &ldquo;extraim_realisation_XX.nc&rdquo;). The synthetic realisations were produced using the extrAIM&rsquo;s uncertainty-quantification approach and the associated conditional sampling method.</p>

opencc-by-4.0Jul 2024View details →
zenodo52/100

Synthetic and real EEG datasets for closed-loop neuroscience

<p>The dataset is made primarily for the task of real-time low latency filtering of the EEG data in the closed loop neuroscience experiments and for EEG forecasting task. The dataset consists of a real data and 5 options of the synthetic data of varying difficulty.</p><p>The real dataset consists of 25 people involved into the P4 alpha neurofeedback training. Its total size is about 16.3 hours. A more detailed instruction for this file is provided in the file Real dataset instructions.txt.</p><p>Synthetic data is generated in 5 different ways: sine wave with white noise, sine wave with pink noise, narrow-band filtered pink noise sample with pink noise, state-space model with white noise and&nbsp;state-space model with pink noise.&nbsp;Each of these datasets has about 34.5 hours of data. It is generated similarly to (Wodeyar&nbsp;et al, 2021). A more detailed instruction for the synthetic dataset can be found in the file&nbsp;Synthetic datasets instructions.txt.<br>&nbsp;</p><p>In LowLatencyEEGFiltering.zip one can find a code for the models used in our paper for low-latency filtering with this data.</p><p>NOTE: Code is also published in the following GitHub repository: https://github.com/ivsemenkov/LowLatencyEEGFiltering</p><p>&nbsp;</p><p>If you use our data or code please cite:&nbsp;https://www.doi.org/10.1088/1741-2552/acf7f3</p>

opencc-by-4.0Jul 2023View details →
zenodo48/100

Synthetic Dataset of Emergency Healthcare Services

<p>Synthetic dataset of emergency services comprised of several CSV files that we have generated using a simulation software. This dataset is open for public use; please cite our work if used in research or applications.</p> <p>## File Overview</p> <ol> <li>**CheckBloodPressure.csv** - (9 KB): Contains blood pressure Server records of patients.</li> <li>**CheckPatientType.csv** - (19 KB): Identifies the type of each patient (e.g., 1 or 3).</li> <li>**Fill_Information.csv** - (2 KB): Fill information records for new patients.</li> <li>**MedicalRecord1.csv** - (10 KB): Medical record dataset for patient type 1.</li> <li>**MedicalRecord2.csv** - (4 KB): Medical record dataset for patient type 2.</li> <li>**MedicalRecord3.csv** - (2 KB): Medical record dataset for patient type 3.</li> <li>**MedicalRecord4.csv** - (13 KB): Medical record dataset for patient type 4.</li> <li>**OutPatientDepartment.csv** - (18 KB): Data related to the satisfaction and length of stay of an given patient.</li> <li>**Triage.csv** - (13 KB): Data related to the triage process.</li> <li>**README.txt** - (4 KB): Documentation of the dataset, including structure, metadata, and usage.</li> </ol> <p>&nbsp;</p> <p>## Common Fields Across Files</p> <ol> <li>**Patient ID** *(Integer)*: Unique identifier for each patient.</li> <li>**Patient Type** *(Integer)*: Classification of patient (e.g., 1, 4).</li> <li>**Medical Records Arrival Time** *(DateTime)*: Timestamp of the patient's first arrival in the medical record department.</li> <li>**Exiting Time** *(DateTime)*: Timestamp when the patient exits a Server.</li> <li>**Waiting Time (min)** *(Real)*: Total waiting time before being attended to.</li> <li>**Resource Used** *(String)*: Resource (e.g., Operator) allocated to the patient.</li> <li>**Utilization %** *(Real)*: Utilization rate of the resource as a percentage.</li> <li>**Queue Count Before Processing** *(Integer)*: Number of patients in the queue before processing begins.</li> <li>**Queue Count After Processing** *(Integer)*: Number of patients in the queue after processing ends.</li> <li>**Queue Difference** *(Integer)*: Difference between the before and after queue counts.</li> <li>**Length of Stay (min)** *(Real)*: Total time spent in the simulation by the patient.</li> <li>**LOS without Queues (min)** *(Real)*: Length of stay excluding any queuing time.</li> <li>**Satisfaction %** *(Real)*: Patient satisfaction rating based on their experience.</li> <li>**New Patient?** *(String)*: Indicates if this is a new patient or a returning one.</li> </ol> <p>Ferreira, M. (2024). Synthetic Dataset of Emergency Healthcare Services [Data set]. Zenodo. https://doi.org/10.5281/zenodo.14212812</p> <p>&nbsp;</p>

restrictedcc-by-sa-4.0Nov 2024View details →
zenodo48/100

Synthetic Simulations Of Extracellular Recordings (SSOER) Dataset

<p>This dataset contains synthetic data from simulations&nbsp;(for a total duration of 10 minutes) including the activity of one multi-unit and two single-units for different firing rates and signal-to-noise ratio levels. It is intended to be used as a standardized dataset&nbsp;to evaluate spike sorting algorithms.</p> <p>Recordings were&nbsp;taken using a sampling rate of 24 kHz, and are comprised of spikes from a database with 594 different average spike shapes, taken from real recordings from monkey neocortex and basal ganglia.</p> <p>This dataset is comprised of two files: <em>data.npy</em> and <em>labels.csv</em>.</p> <ul> <li><em>data.npy</em> contains&nbsp;14,400,000 sampled voltage values, from a single channel, taken at&nbsp;a sampling rate of&nbsp;24 kHz.&nbsp;</li> <li><em>labels.csv</em>&nbsp;contains the timestep, spike class, amplitude (SNR), and firing rate associated with each spiking event.</li> </ul> <p>The original samples used to construct this dataset where previously constructed and made available in [1]. This dataset is an amalgamation of&nbsp;simulation files, which were previously publicly accessible at:&nbsp;<a href="http://www2.le.ac.uk/departments/engineering/research/bioengineering/neuroengineering-lab/software">http://www2.le.ac.uk/departments/engineering/research/bioengineering/neuroengineering-lab/software</a>. Consequently, when using or making modifications to this dataset, in addition to&nbsp;acknowledging this record, [1] must also be acknowledged, as per the original author&#39;s request.</p> <p>[1] J. Martinez, C. Pedreira, M. J. Ison, and R. Quian Quiroga, &ldquo;Realistic simulation of extracellular recordings,&rdquo; Journal of Neuroscience Methods, vol. 184, no. 2, pp. 285&ndash;293, Nov. 2009, doi: 10.1016/j.jneumeth.2009.08.017.</p>

opencc-by-4.0Mar 2022View details →
zenodo48/100

Synthetic datasets reflecting the shRNA-seq knockdown ENCODE data for HepG2 and K562 with coresponding GRN

<p>Synthetic data correspond to the ENCODE data for cell lines HepG2 (https://www.encodeproject.org/biosamples/ENCBS282XVK/) and K562 (https://www.encodeproject.org/biosamples/ENCBS023XVB/). The data and networks were generated using GeneSPIDER (publicly available at https://bitbucket.org/sonnhammergrni/genespider/).</p> <p>&nbsp;</p> <p><strong>Table.1 </strong>Description of the files</p> <table> <tbody> <tr> <td>data_HepG2like_SNR_L=0.0054699_diff=1.6188e-05.txt</td> <td>Synthetic gene expression knockdown (shRNA-seq) data immitating the ENCODE data for HepG2 cell line. Data size: 232 RBPs vs 464 experiments (2 replicates). SNR_L is the value of signal to noise ratio. Difference (diff) value tells the difference between replicate correlation coefficients of real and synthetic ENCODE data. Columns represent experiments, rows represent genes.</td> </tr> <tr> <td>data_K562like_SNR_L=0.0028692_diff=0.00017339.txt</td> <td>Synthetic gene expression knockdown (shRNA-seq) data immitating the ENCODE data for K562 cell line. Data size: 232 RBPs vs 464 experiments (2 replicates). SNR_L is the value of signal to noise ratio. Difference (diff) value tells the difference between replicate correlation coefficients of real and synthetic ENCODE data. Columns represent experiments, rows represent genes.</td> </tr> <tr> <td>network_HEPG2like_sparsity4.txt</td> <td>Synthetic scale-free gene regulatory network compatibile with data_HepG2like_SNR_L=0.0054699_diff=1.6188e-05.txt. Sparsity (average node degree) is 4 including selfloops. Direction should be read from columns to rows.</td> </tr> <tr> <td>network_K562like_sparsity4.txt</td> <td>Synthetic scale-free gene regulatory network compatibile with data_K562like_SNR_L=0.0028692_diff=0.00017339.txt. Sparsity (average node degree) is 4 including selfloops. Direction should be read from columns to rows.</td> </tr> <tr> <td>perturbations_HepG2&amp;K562_2replicates.txt</td> <td>Perturbation matrix including information about knockeddown RBPs. Data size: 232 RBPs vs 464 experiments (2 replicates).</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>Created by Garbulowski et al. (2024) as a part of the work entitled "Comprehensive analysis of the RBP regulome reveals functional modules and drug candidates in liver cancer"</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Synthetic dataset accompanying Neural Image Compression for Gigapixel Histopathology Image Analysis

<p>This dataset was used to develop and evaluate&nbsp;the main method proposed in the paper &quot;Neural Image Compression for Gigapixel Histopathology Image Analysis&quot; published in&nbsp;IEEE Transactions on Pattern Analysis and Machine Intelligence with DOI&nbsp;10.1109/TPAMI.2019.2936841. Please refer to the paper for a detailed description of the dataset.</p> <p>The&nbsp;dataset&nbsp;consists of a set of 50000 images and 50000 associated ground truth masks, distributed into training and test partitions. The name of each file follows the&nbsp;pattern &quot;{id}_{tilted_label}_{nontilted_label}_{tilted_size}_{nontilted_size}_{kind}.png&quot; where:<br> &nbsp; * id: unique identifier within each partition.<br> &nbsp; * tilted_label: image-level label corresponding to the tilted rectangle.<br> &nbsp; * nontilted_label: image-level label corresponding to the non-tilted rectangle.<br> &nbsp; * tilted_size: longest size of the tilted rectangle.<br> &nbsp; * nontilted_size: longest size of the non-tilted rectangle.<br> &nbsp; * kind: either &quot;tile&quot; or &quot;mask&quot; image type.</p> <p>The images are distributed into several data partitions used during cross-validation and fully described in &quot;mnist_folds_set.json&quot;. Please rename &quot;mnist_folds_set.json.removethis&quot; into &quot;mnist_folds_set.json&quot;.</p> <p>The code to recreate this dataset can be found in https://github.com/davidtellez/neural-image-compression.</p>

opencc-by-4.0Aug 2019View details →
zenodo48/100

Synthetic Collision Dataset for Spacecraft Collision Avoidance

<p>This dataset is intended to be used as a banchmark for testing collision avoidance strategies.</p> <p>It is made of 21000000 relative geometries between LEO space objects, 1000 of which are true collision (miss-distance smaller than combined hard body radius).<br>These relative geometries are expressed as target and chaser 6-dimensional state vectors (cartesian coordinates) at time of closest approach.</p> <p>The relative geometry of the encounters are statistically matched to the ESA's Kelvins dataset for the collision avoidance challenge through statistical fitting methods.</p> <p>The collision proportion is tuned to reflect a 1year mission in LEO orbit with an a-priori collision probability of 1e-3 (yearly) and a 21 collision warnings per year.</p> <p>*<em><strong> Implementation Description *</strong></em></p> <p>The dataset is made of a series of .mat files storing the following variables:</p> <div> <ul> <li>'rv_t', target's cartesian state at TCA (km, km/s, in ECI) 6xN vector</li> <li>'rv_c', chaser's cartesian state at TCA (km, km/s, in ECI) 6xN vector&nbsp;</li> <li>'Ct', target's position covariance matrix at TCA (km^2, in target's RTN at TCA) 3x3xN&nbsp;</li> <li>'Cc', chaser's position covariance matrix at TCA (km^2, in chaser's RTN at TCA) 3x3xN&nbsp;</li> <li>'Rc', combined hard body radius (m) 1xN</li> <li>'CollFlag', logic value of collision 1xN (0: no-collision, 1: collision)</li> <li>'missDistance', miss distance at TCA (km) Nx1</li> </ul> <p>The name of the .mat file is formatted as:</p> <p>batch_&lt;batch start index&gt;.mat</p> <p>Each batch file has a maximum dimension of N = 1e5.</p> </div>

opencc-by-4.0Oct 2024View details →
zenodo48/100

UDP Synthetic Dataset for training ML time series models

<p>The dataset available has been produced by the &quot;Next-Generation IoT solutions for the universal supply chain&quot; (iNGENIOUS) project&rsquo;s consortium under EC grant agreement 957216, &nbsp;made publicly available as part of the Horizon 2020 Open Research Data Pilot (<a href="https://www.openaire.eu/what-is-the-open-research-data-pilot">ORD pilot</a>).<br> The European Commission is not liable for any use that may be made of the information contained herein.</p> <p>The available dataset is in csv format and contains synthetic data of UDP packets received and sent by a single User Plane Function (UPF) covering a span of 6 weeks. The format of the datafile is:</p> <ul> <li>index</li> <li>timestamp&nbsp;</li> <li>UDP packets_rcvd - Total number of UDP packets received</li> <li>UDP packets sent - Total number of UDP packets sent</li> </ul> <p>The simulation was performed based on behavior of UPF and 5GC Network functions inferred from stress tests performed in the iNGENIOUS project&#39;s Automated Robots with Heterogeneous Networks Use Case, as well as patterns in urban mobility taken from available UE datasets [NCS+19].</p> <p>More information on the iNGENIOUS project can be found on the project&rsquo;s website: <a href="https://ingenious-iot.eu/">https://ingenious-iot.eu/</a></p> <p>[NCS+19] Noussan M, Carioni G, Sanvito FD, Colombo E. Urban Mobility Demand Profiles:<br> Time Series for Cars and Bike-Sharing Use as a Resource for Transport and Energy<br> Modeling. Data. 2019; 4(3):108. https://doi.org/10.3390/data4030108</p>

opencc-by-4.0Apr 2023View details →
zenodo48/100

Synthetic Particle Image Dataset (SPID)

<p>SPID is a comprehensive dataset composed of synthetic particle image velocimetry (PIV) image pairs and their corresponding exact optical flow computations. It serves as a valuable resource for researchers and practitioners in the field. The dataset is organized into three subsets: training, validation, and test, distributed in a ratio of 70%, 15%, and 15%, respectively.</p><p>Each subset within SPID consists of an input denoted as "x", which comprises synthetic image pairs. These image pairs provide the necessary context for the optical flow computations. Additionally, an output termed "y" is provided, which represents the exact optical flow calculated for each image pair. Notably, the images within the dataset are single-channel, and the optical flow is decomposed into its u and v components.</p><p>The shape of the input subsets in SPID is given by (number of samples, number of frames, image width, image height, number of channels), representing the dimensions of the input data. On the other hand, the shape of the output subsets is given by (number of samples, velocity components, image width, image height), denoting the shape of the optical flow data.</p><p>It is important to mention that SPID dataset is a preprocessed version of the Raw Synthetic Particle Image Dataset (RSPID), ensuring improved usability and reliability. Moreover, the dataset is packaged as a NumPy compressed NPZ file, which conveniently stores the inputs and outputs as separate NumPy NPZ files with the labels train, validation and test as acess keys. This format simplifies data extraction and integration into machine learning frameworks and libraries, facilitating seamless usage of the dataset.</p><p>SPID incorporates various factors that impact PIV analysis to provide a comprehensive and realistic simulation. The dataset includes image pairs with an image width of 665 pixels and an image height of 630 pixels, ensuring a high level of detail and accuracy with an 8-bit depth. It incorporates different particle radii (1, 2, 3, and 4 pixels) and particle densities (15, 17, 20, 23, 25, and 32 particles) to capture diverse particle configurations.</p><p>To simulate real-world scenarios, SPID introduces displacement variations through the delta x factor, ranging from 0.05% to 0.25%. Noise levels (1, 5, 10, and 15) are also incorporated to mimic practical PIV measurements with varying degrees of noise. Furthermore, out-of-plane motion effects are considered with standard deviations of 0.01, 0.025, and 0.05 to assess their impact on optical flow accuracy.</p><p>The dataset covers a wide range of flow patterns encountered in fluid dynamics. It includes Rankine uniform, Rankine vortex, parabolic, stagnation, shear, and decaying vortex flows, allowing for comprehensive testing and evaluation of PIV algorithms across different scenarios.</p><p>By leveraging the SPID dataset, researchers can develop and validate PIV algorithms and techniques under various challenging conditions. Its realistic and diverse simulation of particle image velocimetry scenarios makes it an invaluable tool for advancing the field and improving the accuracy and reliability of optical flow computations.</p><p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo48/100

Synthetic Multimodal Dataset for Daily Life Activities

<p><strong>Outline</strong></p> <ul> <li>This dataset is originally created for the&nbsp;<a href="https://challenge.knowledge-graph.jp/2022/">Knowledge Graph Reasoning Challenge for Social Issue</a>s (KGRC4SI)</li> <li>Video data that simulates daily life actions in a virtual space from Scenario Data.</li> <li>Knowledge graphs, and transcriptions of the Video Data content (&quot;who&quot; did what &quot;action&quot; with what &quot;object,&quot; when and where, and the resulting &quot;state&quot; or &quot;position&quot; of the object).</li> <li>Knowledge Graph Embedding Data are created for reasoning based on machine learning&nbsp;</li> <li>This data&nbsp;is open to the public as open data</li> </ul> <p><strong>Details</strong></p> <ul> <li> <p><a href="https://github.com/KnowledgeGraphJapan/KGRC-RDF/blob/kgrc4si/Movie">Videos</a></p> <ul> <li>mp4 format</li> <li>203&nbsp;action scenarios</li> <li>For each scenario, there is a character rear view (file name ending in 0), an indoor camera switching view (file name ending in 1), and a fixed camera view placed in each corner of the room (file name ending in 2-5). Also, for each action scenario, data was generated for a minimum of 1 to a maximum of 7 patterns with different room layouts (scenes). A total of 1,218&nbsp;videos</li> <li>Videos with slowly moving characters simulate the movements of elderly people.</li> </ul> </li> <li> <p><a href="https://github.com/KnowledgeGraphJapan/KGRC-RDF/blob/kgrc4si/RDF">Knowledge Graphs</a></p> <ul> <li>RDF format</li> <li>203&nbsp;knowledge graphs corresponding to the videos</li> <li>Includes schema and location supplement information</li> <li>The schema is described below</li> <li><a href="http://kgrc4si.ml:7200/sparql">SPARQL endpoints</a>&nbsp;and&nbsp;<a href="https://github.com/KnowledgeGraphJapan/KGRC-RDF/tree/kgrc4si#%E3%83%8A%E3%83%AC%E3%83%83%E3%82%B8%E3%82%B0%E3%83%A9%E3%83%95%E3%81%AE%E4%BD%BF%E7%94%A8%E6%96%B9%E6%B3%95">query examples</a>&nbsp;are available</li> </ul> </li> <li> <p><a href="https://github.com/KnowledgeGraphJapan/KGRC-RDF/blob/kgrc4si/Program">Script Data</a></p> <ul> <li>txt format</li> <li>Data provided to VirtualHome2KG to generate videos and knowledge graphs</li> <li>Includes the action title and a brief description in text format.</li> </ul> </li> <li>Embedding <ul> <li>Embedding Vectors in TransE, ComplEx, and RotatE. Created with DGL-KE (<a href="https://dglke.dgl.ai/doc/">https://dglke.dgl.ai/doc/</a>)</li> <li>Embedding Vectors created with jRDF2vec (<a href="https://github.com/dwslab/jRDF2Vec">https://github.com/dwslab/jRDF2Vec</a>).</li> </ul> </li> </ul> <p><strong>Specification of Ontology</strong></p> <ul> <li>Please refer to the&nbsp;specification for descriptions of all classes, instances, and properties:&nbsp;<a href="https://aistairc.github.io/VirtualHome2KG/vh2kg_ontology.html">https://aistairc.github.io/VirtualHome2KG/vh2kg_ontology.htm</a></li> </ul> <p><strong>Related Resources</strong></p> <ul> <li><a href="https://www.youtube.com/watch?v=Ajbn8hNXiZ8&amp;list=PLHaRK-B0LUwjvrPgmIBTrf3DsPhmdnFTW">KGRC4SI Final Presentations with automatic English subtitles (YouTube)</a></li> <li><a href="https://github.com/aistairc/VirtualHome2KG">VirtualHome2KG (Software)</a></li> <li><a href="https://github.com/aistairc/virtualhome_unity_aist">VirtualHome-AIST (Unity</a>)</li> <li><a href="https://github.com/aistairc/virtualhome_aist">VirtualHome-AIST (Python API</a>)</li> <li><a href="https://github.com/aistairc/virtualhome2kg_visualization">Visualization Tool</a>&nbsp;(Software)</li> <li><a href="https://github.com/aistairc/virtualhome2kg_generation">Script Editor</a>&nbsp;(Software)</li> </ul>

opencc-by-4.0Jun 2023View details →
zenodo48/100

Dataset of reported synthetic conditions for ZIF-8 since 2006 and final characteristics and properties of the particles obtained

<p>This is a dataset generated by mining available literature between 2006 and August 2021. Initial literature search was done using Scopus database, employing a combination of keywords such:&nbsp;&nbsp;&quot;ZIF-8&quot;, &quot;Zeolitic+ZIF-8&quot;, &quot;Framework+ZIF-8&quot;, &quot;Zeolitic+MOF&quot;.</p> <p>First, removal of duplicates and review articles was done and the remaining documents were chosen by title+abstract analysis. The dataset contains a total of 254 entries, i.e., 254 individual reported synthesis.</p> <p>Throughout this dataset, the following parameters can be found:</p> <p>-About synthesis conditions: Zinc source and the amount employed for the synthesis in mmol (milli-moles); 2-methylimidazole in mmol (HmIm); solvent and quantity employed (in mmol). Modulator and quantity employed (in mmol). Please note that quantities in milli-moles were calculated by hand in most of the cases, since reported data was expressed in different units. Reaction temperature (in &deg;C), reaction time (min) and stirring condition (YES-NO-Initial-time). Finally, reports were classified as &quot;systematic&quot; (or not) based on wether the scope of the work was to explore different synthetic conditions.</p> <p>-About ZIF-8 characteristics: information is mainly focusing on structure-related characteristic, namely the&nbsp;particle morphology (classified as Faceted, poor-faceted, quasispherical, aggregated) and the particle&acute;s size. Reported sizes where classified&nbsp;by the different techniques employed. Finally, Surface area determined by BET formalism was also included.</p>

opencc-by-4.0Jul 2023View details →
zenodo48/100

3D-rhi-synth-2000- Synthetic Rhinophyma Visual Dataset

<p>In the real world, only a handful of data is available for the Rhinophyma skin condition, typically numbering in the hundreds. This repository contains a Synthetic Dataset of Rhinophyma, generated through 3D head models of one male and one female. The purpose of this data generation is to address the data scarcity of the Rhinophyma skin condition within the medical visual data and computer vision community. By generating such data, we aim to bridge the gap in data scarcity for this disease condition, as well as introduce a proof-of-concept methodology for generating synthetic data for specialized disease conditions.</p> <p>The <code>highlight</code> folder &#39;highlight_female_male_rendered&#39; provides a glimpse of the entire dataset. The file <code>&#39;2000_deformations.npy</code>&#39; contains the 2000 values of deformations applied during rendering.</p> <p>The dataset is divided into two main folders: &#39;female_rendered&#39; and &#39;male_rendered&#39;. Within each of these folders, there are three subfolders: &#39;configu&#39;, &#39;images&#39;, and &#39;points.</p> <p>1. &#39;configu&#39;: This subfolder contains `.json` files with configuration details for each model. The files include various parameters, such as:<br> &nbsp;&nbsp; - &quot;total_num_cameras&quot;: the total number of cameras.<br> &nbsp;&nbsp; - &quot;active_camera_name&quot;: the name of the active camera.<br> &nbsp;&nbsp; - &quot;camera_focal_len&quot;: the camera&#39;s focal length.<br> &nbsp;&nbsp; - &quot;camera_loc&quot;: the camera&#39;s location.<br> &nbsp;&nbsp; - &quot;camera_rot&quot;: the camera&#39;s rotation.<br> &nbsp;&nbsp; - &quot;nose_deformation_severity&quot;: a measure of the severity of nose deformation.<br> &nbsp;&nbsp; - &quot;label&quot;: the label for the model (e.g., &quot;Severe&quot;).<br> &nbsp;&nbsp; - &quot;nose_variants&quot;: additional details about nose variants.</p> <p>2. &#39;images&#39;: This subfolder contains the rendered images in resolution 960x540. They are named according to the following convention e.g.&#39;Nose_Deformation_Severity_0_2.716669764843742_Camera_00001&#39;, with specific details related to the deformation severity and camera number. There are images for 10 different cameras.</p> <p>3. &#39;points&#39;: This subfolder contains polygon files corresponding to each model. These files represent the deformations applied to the models during rendering.</p> <p>|-- Dataset Root<br> &nbsp;&nbsp;&nbsp; |-- female_rendered<br> &nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |-- configu<br> &nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |&nbsp;&nbsp; |-- Nose_Deformation_Severity_0_2.212930927821943_Camera_00001.json<br> &nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |&nbsp;&nbsp; |-- Nose_Deformation_Severity_0_2.212930927821943_Camera_00002.json<br> &nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |&nbsp;&nbsp; |-- ...<br> &nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |<br> &nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |-- images<br> &nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |&nbsp;&nbsp; |-- Nose_Deformation_Severity_0_2.212930927821943_Camera_00001.png<br> &nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |&nbsp;&nbsp; |-- Nose_Deformation_Severity_0_2.212930927821943_Camera_00002.png<br> &nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |&nbsp;&nbsp; |-- ...<br> &nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |<br> &nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |-- points<br> &nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |&nbsp;&nbsp; |-- Nose_Deformation_Severity_0_2.212930927821943.ply<br> &nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |&nbsp;&nbsp; |-- ...<br> &nbsp;&nbsp;&nbsp; |&nbsp; &nbsp;<br> &nbsp;&nbsp;&nbsp; |-- male_rendered<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; |-- configu<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |-- Nose_Deformation_Severity_0_2.716669764843742_Camera_00001.json<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |-- Nose_Deformation_Severity_0_2.716669764843742_Camera_00002.json<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |-- ...<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; |<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; |-- images<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |-- Nose_Deformation_Severity_0_2.716669764843742_Camera_00001.jpg<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |-- Nose_Deformation_Severity_0_2.716669764843742_Camera_00002.jpg<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; |&nbsp;&nbsp; |-- ...<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; |<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; |-- points<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; |-- Nose_Deformation_Severity_0_2.716669764843742.ply<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; |-- ...</p> <p>Together, these folders and files comprise a dataset designed to represent and analyze the Rhinophyma condition in both male and female 3D head models that we have created. These 3D models will be made available upon a genuine request to the authors of this dataset.</p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

Synthetic AIS Dataset of Vessel Proximity Events

<p>The Automatic Identification System (AIS) allows vessels to share identification, characteristics, and location data through self-reporting. This information is periodically broadcast and can be received by other vessels with AIS transceivers, as well as ground or satellite sensors. Since the International Maritime Organisation (IMO) mandated AIS for vessels above 300 gross tonnage, extensive datasets have emerged, becoming a valuable resource for maritime intelligence.</p> <p>Maritime collisions occur when two vessels collide or when a vessel collides with a floating or stationary object, such as an iceberg. Maritime collisions hold significant importance in the realm of marine accidents for several reasons:</p> <ol> <li>Injuries and fatalities of vessel crew members and passengers.</li> <li>Environmental effects, especially in cases involving large tanker ships and oil spills.</li> <li>Direct and indirect economic losses on local communities near the accident area.</li> <li>Adverse financial consequences for ship owners, insurance companies and cargo owners including vessel loss and penalties.</li> </ol> <p>As sea routes become more congested and vessel speeds increase, the likelihood of significant accidents during a ship&#39;s operational life rises. The increasing congestion on sea lanes elevates the probability of accidents and especially collisions between vessels.</p> <p>The development of solutions and models for the analysis, early detection and mitigation of vessel collision events is a significant step towards ensuring future maritime safety. In this context, a synthetic vessel proximity event dataset is created using real vessel AIS messages. The synthetic dataset of trajectories with reconstructed timestamps is generated so that a pair of trajectories reach simultaneously their intersection point, simulating an unintended proximity event (collision close call).&nbsp;The dataset aims&nbsp;to provide a basis for the development of methods for the detection and mitigation of maritime collisions and proximity events, as well as the study and training of vessel crews in simulator environments.</p> <p>The dataset consists of 4658 samples/AIS messages of 213 unique vessels from the Aegean Sea. The steps that were followed to create the collision dataset are:</p> <p>Given 2 vessels X (vessel_id1) and Y (vessel_id2) with their current known location (LATITUDE [lat], LONGITUDE [lon]):&nbsp;</p> <ol> <li>Check if the trajectories of vessels X and Y are spatially intersecting.</li> <li>If the trajectories of vessels X and Y are intersecting, then align temporally the timestamp of vessel Y at the intersect point according to X&rsquo;s timestamp at the intersect point. The temporal alignment is performed so the spatial intersection (nearest proximity point) occurs at the same time for both vessels.</li> <li>Also for each vessel pair the timestamp of the proximity event is different from a proximity event that occurs later so that different vessel trajectory pairs do not overlap temporarily.</li> </ol> <p>Two csv files are provided. vessel_positions.csv includes the AIS positions vessel_id, t, lon, lat, heading, course, speed of all vessels. Simulated_vessel_proximity_events.csv includes the id, position and timestamp of each identified proximity event along with the vessel_id number of the associated vessels.&nbsp;The final sum of unintended&nbsp;proximity events in the dataset is 237. Examples of unintended vessel proximity events are visualized in the respective png and gif files.</p> <p>The research leading to these results has received funding from the European Union&#39;s Horizon&nbsp;Europe&nbsp;Programme&nbsp;under the CREXDATA Project, grant agreement n&deg;&nbsp;101092749.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Datasets of synthetic task graphs for evaluating a reliability and latency multi-objective task allocation framework

<p>These datasets of synthetic task graphs were generated to evaluate the performance and scalability of a multi-objective task allocation approach for workflow applications of various structures and sizes in a system based on the edge-hub-cloud paradigm. The targeted architecture comprised an edge device (e.g., a single-board computer attached to an unmanned aerial vehicle (UAV)) interacting with a hub device (e.g., a laptop), which in turn communicated with a more computationally capable cloud server. The objectives were the maximization of the overall reliability and the minimization of the overall latency of the application, under memory, storage, energy, and task precedence constraints. We considered that a percentage of the tasks required fixed allocation on the edge or hub device. Each task had a different vulnerability factor (i.e., probability of failure) on each device.</p> <p>We generated nine task graphs of serial, parallel, and mixed (a combination of serial and parallel) structure with 10, 100, and 1000 nodes, utilizing the Task Graphs For Free (TGFF) random task graph generator [1]. Additional task parameters (e.g., execution time, power consumption, vulnerability factor, memory, storage, output data size) were included post-generation, using representative random values. More details are provided in README.txt.</p> <p>Note: These datasets are released under a Creative Commons Attribution license. If you utilize these datasets in your work, please cite us using the corresponding Zenodo DOI https://doi.org/10.5281/zenodo.10357101.</p> <p>References:</p> <p>[1] R. P. Dick, D. L. Rhodes and W. Wolf, "TGFF: Task graphs for free," Proceedings of the Sixth International Workshop on Hardware/Software Codesign (CODES/CASHE'98), Seattle, WA, USA, 1998, pp. 97-101, doi: 10.1109/HSC.1998.666245.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

TERMINET eHealth post-operation complications synthetic dataset

<p><strong>1. Introduction</strong></p> <p>Older adults with cancer often need to undergo operations. Post-surgery complications may arise, and Real-World Data (RWD) collected from such patients during a pre-operation monitoring period of two weeks can help identify risk for post-surgery complications. The involved RWD span behavioral data (measured or reported) as well as clinical data (collected during clinical tests). This dataset is synthesized by <a href="https://innovationsprint.eu/">Innovation Sprint</a>, using actual data collected from eligible <a href="https://www.policlinicogemelli.it/en/">Fondazione Policlinico Gemelli</a>&nbsp;patients participating to the SUPERO study. The clinical data is collected by the hospital, while&nbsp;the behavioral is collected using <a href="https://innovationsprint.eu/healthentia/">Healthentia</a>, a medical decision support software developed by Innovation Sprint, facilitating the collection, analysis and presentation of behavioral data.</p> <p><strong>2. Dataset description</strong></p> <p>The provided TERMINET eHealth post-operation complications synthetic dataset contains 10,000 synthetic patients, provided in an equal number of rows in the CSV file containing the dataset. The different attributes of the dataset are organized in columns.</p> <p>The attributes are summarized as follows:</p> <ul> <li>6 columns of step data statistics</li> <li>20 columns of clinical attributes</li> <li>2 columns of demographics attributes</li> <li>12 columns of questionnaire attributes</li> <li>1 column of outcome attribute</li> </ul> <p><em>2.1. Step statistics</em></p> <p>Step data is collected per day of the pre-hospitalization period. The final two weeks of that period are used to derive the step statistics. For each of the weeks, the mean, standard deviation and slope of the linear regression of the step data is reported, 3 attributes per week, 6 attributes in total.</p> <p><em>2.2. Clinical attributes</em></p> <p>The 20 clinical attributes collected at the hospital are ALT, Hematocrit (%), AST, Lymphocytes, Hepatitis B, Neutrophils (%), Hepatitis C, Neutrophils, INR (%), INR, White Blood Cells, INR (seconds), Platelets, Sodium, Hemoglobin, Potassium, Lymphocytes (%), Creatinine, Bilirubin and Urea Nitrogen.</p> <p><em>2.3. Demographic attributes</em></p> <p>The sex and age are the two demographic attributes collected.</p> <p><em>2.4. Questionnaire attributes</em></p> <p>Three questionnaires are involved in the SUPERO study are:</p> <ul> <li>G8, spanning the categories of food intake, weight loss, movement, neuropsychological, BMI, multiple medication, health and age.</li> <li>SPPB, spanning the categories of balance, speed and strength</li> <li>MiniCog, where only the clock drawing capabilities are assessed</li> </ul> <p>2.5. Outcome attribute</p> <p>The single outcome attribute is the existence of any post-surgery complications. Please note that the dataset is quite imbalanced, since complications are very rare.</p> <p><strong>3. Data synthesis</strong></p> <p>This dataset is synthesized from the early data of the SUPERO study. Currently there are 21 patients registered, with the decision to operate them being reached for 20 of them. 16&nbsp;of the patients have already been operated, 2 of them having exhibited post-surgery complications. More vectors have been generated by adding Gaussian noise to the original 16&nbsp;vectors, resulting to 128 vectors. The resulting vectors have been clustered into 16 clusters using Agglomerative clustering. Every cluster has been modelled via Gaussian Mixture Models. The resulting set of GMMs has been used to generate the 10,000 synthetic vectors of the dataset.</p> <p><strong>4. Acknowledgement</strong></p> <p>The development of this dataset has received funding from the European Union&rsquo;s Horizon 2020 research and innovation programme under grant agreement No. 957406 (TERMINET).</p>

opencc-by-4.0May 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record