Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
250
datasets available to search
ShareScore release 0.9.0
Dataset results
250 results for “Synthetic dataset”
SYNTHETIC dataset attached to the paper "Grasp Pre-shape Selection by Synthetic Training: Eye-in-hand Shared Control on the Hannes Prosthesis"
<p>SYNTHETIC dataset to replicate the results in "Grasp Pre-shape Selection by Synthetic Training: Eye-in-hand Shared Control on the Hannes Prosthesis", accepted to IEEE/RSJ IROS 2022.</p> <p>In order to fully reproduce the experiments, download also the REAL dataset. </p> <p>To automatically download the REAL and SYNTHETIC dataset, run the script provided at the link below.</p> <p>Code to replicate the results available at: https://github.com/hsp-iit/prosthetic-grasping-experiments</p>
REAL dataset attached to the paper "Grasp Pre-shape Selection by Synthetic Training: Eye-in-hand Shared Control on the Hannes Prosthesis"
<p>REAL dataset to replicate the results in "Grasp Pre-shape Selection by Synthetic Training: Eye-in-hand Shared Control on the Hannes Prosthesis", accepted to IEEE/RSJ IROS 2022.</p> <p>In order to fully reproduce the experiments, download also the SYNTHETIC dataset. </p> <p>To automatically download the REAL and SYNTHETIC dataset, run the script provided at the link below.</p> <p>Code to replicate the results available at: https://github.com/hsp-iit/prosthetic-grasping-experiments</p>
Shotgun metagenomic sequencing dataset of a synthetic mock community containing 20 genomes spiked-in at even and staggered concentrations.
<p>Shotgun metagenomics (SM) sequencing is a popular method used in microbial ecology to obtain insights on microbial community structure and function potential in a given biological system without the need to cultivate microorganisms. The dataset described in this article describes technical triplicates of shotgun metagenomic sequence libraries generated from two purified and titrated mixes of 20 distinct reference bacterial genomes for which key characteristics such as genome size, sequence and spiked-in concentrations are known. In one of the genomic DNA mix, each genome is spiked-in at similar concentrations (representing an even microbial community) and in the other, genomes are spiked-in at different concentrations with some genomes highly abundant and other in low quantity, mimicking an uneven microbial community DNA extract. In order to be interpretable, SM sequencing data needs to be properly analyzed by complex analytical bioinformatic pipelines. Environments investigated with this method can range from simple to very complex. Typically, microbial communities contain microbes that are ubiquitous and some others much rarer. Analysis of rare microbes in a complex microbial community are challenging to perform as their sequencing signals get submerged by the microbial genomes that are more abundant. In this context, it is critical to have access to sequencing data of simple mock communities of mixes of well characterized genomes in order to develop and validate bioinformatic methods that aim to accurately analyze microbial communities.</p>
Synthetic Hand Recognition Dataset
<p>This is the synthetic hand recognition dataset generated with the Synthetic Human Dataset Generator (available on github). It contains 90,000 images of 10 different virtual humans (5 female, 5 male) in various poses. All images are annotated with the color coded segmented hands.</p>
Beyond microplastics: Water soluble synthetic polymers exert sublethal adverse effects in the freshwater cladoceran Daphnia magna - experimental Dataset
<p>This Dataset contains the raw experimental data for the article "Beyond microplastics: Water soluble synthetic polymers exert sublethal adverse effects in the freshwater cladoceran Daphnia magna" by Simona Mondellini, Matthias Schott, Martin G.J. Löder, Seema Agarwal, Andreas Greiner, Christian Laforsch. Published on Science of the Total Environment (2022) <a href="https://doi.org/10.1016/j.scitotenv.2022.157608">https://doi.org/10.1016/j.scitotenv.2022.157608</a><br> The file "dataset information" contains a description of the other files.</p>
DATA7: A dataset that uses synthetic trajectories of vehicles and real cellular tower locations to simulate the workload of Edge nodes in the city of Pisa
<p><strong>Description</strong></p> <p>The dataset contains observations of vehicles in the range of edge nodes (cellular towers). The trajectories of vehicles are synthetically generated with <a href="https://www.eclipse.org/sumo/">SUMO</a>. The cellular tower positions have been taken from <a href="https://opencellid.org/">OpenCelliD</a>. The dataset is in the comma-separated values (CSV) format, and is around 220MB decompressed.</p> <p><br> The CSV contains the following fields:<br> * edge_id: unique identifier of the edge devices<br> * edge_lat: latitude coordinate of the edge device<br> * edge_lon: longitude coordinate of the edge device<br> * time: simulation step of the observation<br> * vehicle_id: unique identifier of the vehicle<br> * vehicle_lat: latitude coordinate of the vehicle<br> * vehicle_lon: longitude coordinate of the vehicle<br> * distance: geodesic distance in meters from the vehicle and the edge device</p>
Dataset of "Theoretical and practical aspects of the design and production of synthetic holograms for transmission electron microscopy"
<p>Dataset with script and article images published in https://doi.org/10.1063/5.0067528</p>
SynthRAD2023 Grand Challenge validation dataset: synthetizing computed tomography for radiotherapy
<p><strong>Version 1.1</strong>, updated on 2023-06-04 --> the task2_val.zip has been modified with a new cbct file for patient 2BA078.<br> <br> The dataset can be downloaded from <a href="https://doi.org/10.5281/zenodo.7260705">https://doi.org/</a><a href="https://doi.org/10.5281/zenodo.7868169">10.5281/zenodo.7868169</a> and a detailed description is offered at <a href="https://doi.org/10.5281/zenodo.7260704">https://doi.org/10.5281/zenodo.7260704</a> in the "synthRAD2023_dataset_description.pdf".</p> <p>The<strong> </strong>input of the<strong> validation datasets</strong> for Task1 is in Task1_val.zip, while for Task2 in Task2_val.zip. After unzipping, each Task is organized according to the following folder structure:</p> <p>Task1_val.zip/</p> <p>├── Task1</p> <p> ├── brain</p> <p> ├── 1Bxxxx</p> <p> ├── mr.nii.gz</p> <p> └── mask.nii.gz</p> <p> ├── ...</p> <p>└── overview</p> <p> ├── 1_brain_val.xlsx</p> <p> ├── 1Bxxxx_val.png</p> <p> └── ... </p> <p> └── pelvis</p> <p> ├── 1Pxxxx</p> <p> ├── mr.nii.gz</p> <p> ├── mask.nii.gz</p> <p> ├── ...</p> <p>└── overview</p> <p> ├── 1_pelvis_val.xlsx</p> <p> ├── 1Pxxxx_val.png</p> <p> └── ....</p> <p>Task2_val.zip/</p> <p>├──Task2</p> <p> ├── brain</p> <p> ├── 2Bxxxx</p> <p> ├── cbct.nii.gz</p> <p> └── mask.nii.gz</p> <p> ├── ...</p> <p>└── overview</p> <p> ├── 2_brain_val.xlsx</p> <p> ├── 2Bxxxx_val.png</p> <p> └── ... </p> <p> └── pelvis</p> <p> ├── 2Pxxxx</p> <p> ├── cbct.nii.gz</p> <p> ├── mask.nii.gz</p> <p>├── ...</p> <p>└── overview</p> <p> ├── 2_pelvis_val.xlsx</p> <p> ├── 2Pxxxx_val.png</p> <p> └── ....</p> <p>Each patient folder has a unique name that contains information about the task, anatomy, center and a patient ID. The naming follows the convention below:</p> <p>[Task] [Anatomy] [Center] [PatientID]</p> <p>1 B A 001</p> <p>In each patient folder, two files can be found: </p> <ul> <li> <p>mr.nii.gz or cbct.nii.gz (depending on the task): CBCT/MR image</p> </li> <li> <p>mask.nii.gz: image containing a binary mask of the dilated patient outline </p> </li> </ul> <p>For each task and anatomy, an overview folder is provided which contains the following files:</p> <ul> <li> <p>[task]_[anatomy]_val.xlsx: This file contains information about the image acquisition protocol for each patient.</p> </li> <li> <p>[task][anatomy][center][PatientID]_val.png: For each patient a png showing axial, coronal and sagittal slices of CBCT/MR, CT, mask and the difference between CBCT/MR and CT is provided. These images are meant to provide a quick visual overview of the data.</p> </li> </ul> <p><strong>DATASET DESCRIPTION</strong></p> <p>This challenge dataset contains imaging data of patients who underwent radiotherapy in the brain or pelvis region. Overall, the population is predominantly adult and no gender restrictions were considered during data collection. For Task 1, the inclusion criteria were the acquisition of a CT and MRI during treatment planning while for task 2, acquisitions of a CT and CBCT, used for patient positioning, were required. Datasets for task 1 and 2 do not necessarily contain the same patients, given the different image acquisitions for the different tasks.</p> <p>Data was collected at 3 Dutch university medical centers:</p> <ul> <li> <p>Radboud University Medical Center;</p> </li> <li> <p>University Medical Center Utrecht;</p> </li> <li> <p>University Medical Center Groningen.</p> </li> </ul> <p>For anonymization purposes, from here on, institution names are substituted with A, B and C, without specifying which institute each letter refers to.</p> <p>The following number of patients is available in the validation set.</p> <p><strong>Validation</strong></p> <table> <tbody> <tr> <td> </td> <td> <p><strong>Brain</strong></p> </td> <td> <p><strong>Pelvis</strong></p> </td> </tr> <tr> <td> </td> <td> <p><strong>Center A</strong></p> </td> <td> <p><strong>Center B</strong></p> </td> <td> <p><strong>Center C</strong></p> </td> <td> <p><strong>Total</strong></p> </td> <td> <p><strong>Center A</strong></p> </td> <td> <p><strong>Center B</strong></p> </td> <td> <p><strong>Center C</strong></p> </td> <td> <p><strong>Tota</strong>l</p> </td> </tr> <tr> <td> <p><strong>Task 1</strong></p> </td> <td> <p>10</p> </td> <td> <p>10</p> </td> <td> <p>10</p> </td> <td> <p>30</p> </td> <td> <p>20</p> </td> <td> <p>0</p> </td> <td> <p>10</p> </td> <td> <p>30</p> </td> </tr> <tr> <td> <p><strong>Task 2</strong></p> </td> <td> <p>10</p> </td> <td> <p>10</p> </td> <td> <p>10</p> </td> <td> <p>30</p> </td> <td> <p>10</p> </td> <td> <p>10</p> </td> <td> <p>10</p> </td> <td> <p>30</p> </td> </tr> </tbody> </table> <p>In total, for all tasks and anatomies combined, 120 image pairs are available in this dataset. <strong>This repository only contains the validation data. </strong>The training data is provided at: h<a href="https://doi.org/10.5281/zenodo.7260704">ttps://doi.org/10.5281/zenodo.7260704</a>.</p> <p>All images were acquired with the clinically used scanners and imaging protocols of the respective centers and reflect typical images found in clinical routine. As a result, imaging protocols and scanner can vary between patients. A detailed description of the imaging protocol for each image, can be found in spreadsheets that are part of the dataset release (see dataset structure).</p> <p>Data was acquired with the following scanners:</p> <ul> <li> <p>Center A:</p> <ul> <li> <p>MRI: Philips Ingenia 1.5T/3.0T</p> </li> <li> <p>CT: Philips Brilliance Big Bore or Siemens Biograph20 PET-CT</p> </li> <li> <p>CBCT: Elekta XVI</p> </li> </ul> </li> <li> <p>Center B:</p> <ul> <li> <p>MRI: Siemens MAGNETOM Aera 1.5T or MAGNETOM Avanto_fit 1.5T</p> </li> <li> <p>CT: Siemens SOMATOM Definition AS</p> </li> <li> <p>CBCT: IBA Proteus+ or Elekta XVI</p> </li> </ul> </li> <li> <p>Center C:</p> <ul> <li> <p>MRI: Siemens Avanto fit 1.5T or Siemens MAGNETOM Vida fit 3.0T</p> </li> <li> <p>CT: Philips Brilliance Big Bore</p> </li> <li> <p>CBCT: Elekta XVI</p> </li> </ul> </li> </ul> <p>For task 1, MRIs were acquired with a T1-weighted gradient echo or an inversion prepared - turbo field echo (TFE) sequence and collected along with the corresponding planning CTs for all subjects. The exact acquisition parameters vary between patients and centers. For centers B and C, selected MRIs were acquired with Gadolinium contrast, while the selected MRIs of center A were acquired without contrast.</p> <p>For task 2, the CBCTs used for image-guided radiotherapy ensuring accurate patient position were selected for all subjects along with the corresponding planning CT.</p> <p>The following pre-processing steps were performed on the data:</p> <ul> <li> <p>Conversion from dicom to compressed nifti (nii.gz)</p> </li> <li> <p>Rigid registration between CT and MR/CBCT</p> </li> <li> <p>Anonymization (face removal, only for brain patients)</p> </li> <li> <p>Patient outline segmentation (provided as a binary mask)</p> </li> <li> <p>Crop MR/CBCT, CT and mask to remove background and reduce file sizes</p> </li> </ul> <p>The code used to preprocess the images can be found at: <a href="https://github.com/SynthRAD2023/">https://github.com/SynthRAD2023/</a>. Detailed information about the dataset are provided in SynthRAD2023_dataset_description.pdf published here along with the data and will also be submitted to Medical Physics.</p> <p><strong>ETHICAL APPROVAL</strong></p> <p>Each institution received ethical approval from their internal review board/Medical Ethical committee:</p> <ul> <li> <p>UMC Utrecht approved not-WMO on 4/03/2022 with number 22/474 entitled: “Synthetizing computed tomography for radiotherapy Grand Challenge (SynthRAD)”.</p> </li> <li> <p>UMC Groningen approved not-WMO on 20/07/2022 with number 202200310 entitled: “Synthesizing computed tomography for radiotherapy - Grand Challenge”.</p> </li> <li> <p>Radboud UMC declared the study not-WMO on 17/10/2022 with number 2022-15950 entitled “Synthetizing computed tomography for radiotherapy Grand Challenge”.</p> </li> </ul> <p><strong>CHALLENGE DESIGN</strong></p> <p>The overall challenge design can be found at <a href="https://doi.org/10.5281/zenodo.7746019">https://doi.org/10.5281/zenodo.7746019</a>.</p>
Synthetic Turfgrass Dataset
<p><em>Synthetic Turfgrass Dataset created in blender, top view of a golf course surface.</em></p> <p> </p> <table> <tbody> <tr> <td><strong>Image Type</strong></td> <td><strong>Folder name</strong></td> <td><strong>Count</strong></td> </tr> <tr> <td>RGB Training Images</td> <td><strong>Images</strong></td> <td>6510</td> </tr> <tr> <td> <p>Mono 8bit Annotation images:</p> <p>- Class grass, Value :30</p> <p>- Class divot anomaly, Value: 10</p> </td> <td><strong>ImageLables</strong></td> <td>6510</td> </tr> </tbody> </table> <p> </p> <p> </p> <p> </p>
Synthetic noisy datasets for submarine cable magnetic anomaly locaition
<p>Synthetic noisy datasets for submarine cable magnetic anomaly locaition, including a training, a validation and two test data sets.</p>
Hotels and Guests Dataset; synthetic and Non-synthetic
<p>Check this link for the algorithm</p> <p><a href="https://docs.sdv.dev/sdv/multi-table-data/modeling/synthesizers/hmasynthesizer">https://docs.sdv.dev/sdv/multi-table-data/modeling/synthesizers/hmasynthesizer</a></p>
Turfgrass Divot Dataset (Synthetic ) for divot detection object detection system
<p>The dataset provided below has been synthetically created using Blender. A fundamental analysis on this data was conducted utilizing the YOLO V3 object detection technique to identify divots or areas of damage.</p> <p>Used for paper: </p> <p>Advancing Turfgrass Maintenance with Synthetic Data for Divot Detection</p> <p>https://github.com/stevefoy/Turfgrass-Divot-Object-Detection</p> <pre><span>@inproceedings</span>{<span>IMVIP2024</span>, <span>author</span> = <span><span>{</span>Stephen Foy and Simon McLoughlin<span>}</span></span>, <span>title</span> = <span><span>{</span>Advancing Turfgrass Maintenance with Synthetic Data for Divot Detection<span>}</span></span>, <span>booktitle</span> = <span><span>{</span>Irish Machine Vision and Image Processing Conference (IMVIP)<span>}</span></span>, <span>year</span> = <span><span>{</span>2024<span>}</span></span> }</pre> <p><strong>Contents of the Zip File</strong>:</p> <ul> <li> <p><strong>synthDivot_416x416 Folder</strong>:</p> <ul> <li>Train and validation subfolders</li> <li>1200 RGB PNG images</li> <li>Corresponding masks for each image</li> <li>Bounding box data in YOLO <code>.txt</code> format</li> </ul> </li> <li> <p><strong>synthDivot_608x608 Folder</strong>:</p> <ul> <li>Train and validation subfolders</li> <li>1200 RGB PNG images</li> <li>Bounding box data in YOLO <code>.txt</code> format</li> </ul> </li> </ul> <p> </p> <p> </p>
NIFECG synthetic signals generated with fecgsym by PhysioNet: Dataset 1/2
<p>First part</p> <p>https://zenodo.org/records/8415709</p> <p>Second part</p> <p>https://zenodo.org/records/8429286</p> <p>Non-invasive fetal electrocardiogram (NIFECG) signals.</p> <p>Fetal's heart rate: 60 - 200 bpm</p> <p>Mother's heart rate: 65 - 120 bpm</p> <p>Sample frequency 1000 Hz</p> <p>8,008 signals in total</p> <p>Download the files and join them as follows:</p> <p>cat tmp2.tar_part* > nifecg_signals.tar</p> <p>Untar the file with the following command:</p> <p>tar xvf nifecg_signals.tar</p>
NIFECG synthetic signals generated with fecgsym by PhysioNet: Dataset 2/2
<p>First part</p> <p>https://zenodo.org/records/8415709</p> <p>Second part</p> <p>https://zenodo.org/records/8429286</p> <p>Non-invasive fetal electrocardiogram (NIFECG) signals.</p> <p>Fetal's heart rate: 60 - 200 bpm</p> <p>Mother's heart rate: 65 - 120 bpm</p> <p>Sample frequency 1000 Hz</p> <p>8,008 signals in total</p> <p>Download the files and join them as follows:</p> <p>cat tmp2.tar_part* > nifecg_signals.tar</p> <p>Untar the file with the following command:</p> <p>tar xvf nifecg_signals.tar</p>
Dataset for Phosphite as an Engineered Niche for Pseudomonas veronii in a Synthetic Soil Bacterial Community
<p>Files containing the source data and analyses used in the manuscript "Phosphite as an Engineered Niche for <em>Pseudomonas veronii </em>in a Synthetic Soil Bacterial Community".</p> <p> </p> <p>Clara Bailey (1), Philip Gwyther (2), Senka Čaušević (2), Brandon L. Greene (1), and Jan Roelof van der Meer (2)</p> <p>1) Department of Chemistry and Biochemistry, University of California, Santa Barbara, Santa Barbara, California, United States</p> <p>2) Department of Fundamental Microbiology, University of Lausanne, Lausanne, Switzerland</p> <p> </p> <p>This dataset contains 16S rRNA gene amplicon sequencing data (in the form of an abundance table, "abund.csv" and combined with CFU counts in "abund_cfu.csv"), toluene quantification data, and CFU counts. All data analysis, statistical tests, and generated figures are contained in the relevant .R script. Refer to README files for data tables as well as the README section at the header of the R script. The dataset has been updated from version 1 to include manuscript revisions, and updated calculations and figures. </p> <p> </p>
Synthetic dataset for the PFC3D simulation of fragmentable blocks sliding experiments
<p>Dataset 1 - The raw data of numerical simulations, including deposit parameters of all simulations, fragments size distribution of all simulations, impact force of D1, breakage bond number of D1, velocity of D1.</p> <p>Simulation code - all the PFC3D simulation code for the fragmentable blocks sliding experiments.</p>
EmoSynth: The Emotional Synthetic Audio Dataset
<p>The ability of sound to enhance human wellbeing has been known since ancient civilisations, and methods can be found today across domains of health and within a variety of cultures. </p> <p>EmoSynth is a dataset of 144 audio files which have been labelled by 40 listeners for their the perceived emotion, in regards to the dimensions of Valence and Arousal.</p> <p>Results on the dataset show that Arousal does correlate moderately to fundamental frequency, and that the sine waveform is perceived as significantly different to square and sawtooth waveforms when evaluating perceived Arousal. The general results suggest that isolated synthetic audio can be modelled as a means of evoking affective states of emotion.</p> <p>When using this version of the data for a publication, please cite the following study: </p> <pre>@inproceedings{Baird:2018:PEI:3243274.3243277, author = {Baird, Alice and Parada-Cabaleiro, Emilia and Fraser, Cameron and Hantke, Simone and Schuller, Bj\"{o}rn}, title = {The Perceived Emotion of Isolated Synthetic Audio: The EmoSynth Dataset and Results}, booktitle = {Proceedings of the Audio Mostly 2018 on Sound in Immersion and Emotion}, series = {AM'18}, year = {2018}, isbn = {978-1-4503-6609-0}, location = {Wrexham, United Kingdom}, pages = {7:1--7:8}, articleno = {7}, numpages = {8}, url = {http://doi.acm.org/10.1145/3243274.3243277}, doi = {10.1145/3243274.3243277}, acmid = {3243277}, publisher = {ACM}, address = {New York, NY, USA}, keywords = {affective computing, perception, sound healing, synthetic audio, wellbeing}, } </pre> <p> </p>
Groove2Groove MIDI Dataset: synthetic accompaniments in 3k styles
<p>The <em>Groove2Groove MIDI Dataset</em> is a parallel corpus of synthetic MIDI accompaniments in almost 3000 different styles, created as described in the paper <em><a href="https://doi.org/10.1109/TASLP.2020.3019642">Groove2Groove: One-Shot Accompaniment Style Transfer with Supervision from Synthetic Data</a></em> [<a href="https://groove2groove.telecom-paris.fr/data/paper.pdf">pdf</a>]. See the <code>README.md</code> file or the <em><a href="https://groove2groove.telecom-paris.fr/#Dataset">Groove2Groove website</a></em> for more information.</p> <p>The dataset is split into the following sections:</p> <ul> <li><code>train</code> contains 5744 MIDI files in 2872 styles (exactly 2 files per style). Each file contains 252 measures following a 2 measure count-in.</li> <li><code>val</code> and <code>test</code> each contain 1200 files in 40 styles (exactly 30 files per style, 16 bars per file after the count-in). The sets of styles are disjoint from each other and from those in <code>train</code>.</li> <li><code>itest</code> is generated from the same chord charts as <code>test</code>, but in 40 styles from the training set.</li> </ul> <p>Chord charts for all MIDI files are provided in the ABC format and the Band-in-a-Box (MGU) format. Each chord chart corresponds to at least 2 MIDI files in different styles.</p> <p>The code used to automate Band-in-a-Box is available in the <a href="https://github.com/cifkao/pybiab">pybiab</a> package.</p> <p>If you use the data in your research, please reference the paper (not just the Zenodo record):</p> <pre><code>@article{groove2groove, author={Ond\v{r}ej C\'{i}fka and Umut \c{S}im\c{s}ekli and Ga\"{e}l Richard}, title={{Groove2Groove}: One-Shot Music Style Transfer with Supervision from Synthetic Data}, journal={IEEE/ACM Transactions on Audio, Speech, and Language Processing}, publisher={IEEE}, year={2020}, volume={28}, pages={2638--2650}, doi={10.1109/TASLP.2020.3019642}, url={https://doi.org/10.1109/TASLP.2020.3019642} }</code></pre>
Perturbed Synthetic SWOT Datasets for Testing and Development of a Kalman Filter Approach to Estimate Daily Discharge
<p><strong>1. Introduction</strong></p> <p>Datasets are used to evaluate the performance of a Kalman filter approach to estimate daily discharge. This is a perturbed version of synthetic SWOT datasets consisting of 15 river sections, which are commonly agreed datasets for evaluating the performance of SWOT discharge algorithms (Frasson et al., 2020, 2021). The benchmarking manuscript entitled “A Kalman Filter Approach for Estimating Daily Discharge Using Space-based Discharge Estimates” is currently under review at Water Resources Research. Once the manuscript is accepted, its DOI will be included here.</p> <p> </p> <p><strong>2. </strong><strong>File description</strong></p> <p>The datasets are generally divided into two categories: river information (River_Info) and time series data (Timeseries_Data). River information provides fundamental and general river characteristics, whereas time series data offers daily reach-averaged data for each reach. In time series data, the data mainly contains three components: true data, perturbed measurements, and true and perturbed flow law parameters (A0, an, and b). For each reach, there are 10000 realizations of perturbed measurements per time step and there are 100 realizations of time-invariant perturbed flow law parameters through a Monte Carlo simulation (Frasson et al., 2023). Moreover, to support our proposed Kalman filter approach to estimate daily discharge, the datasets provide the median of the perturbed discharge, river width, water surface slope, and change in the cross-sectional area, as well as the uncertainty of the perturbed discharge and change in the cross-sectional area based on the interquartile range (Fox, 2015).</p> <p>To support reproducibility and facilitate example usage, we now include a MATLAB code package (<code>KalmanFilter_Code.zip</code>) that demonstrates how to run the Kalman filter approach using the Missouri Downstream case as an example. </p> <p>Datasets are contained in a .mat file per river. The detailed groups and variables are in the following:</p> <p><strong>River_Info</strong></p> <p>Name: River name, data type: char</p> <p>QWBM: Mean annual discharge from the water balance model WBMsed (Cohen et al., 2014)</p> <p>rch_bnd: Reach boundaries measured in meters from the upstream end of the model</p> <p>gdrch: Good reaches in the study. They were used to exclude small reaches defined around low-head dams and other obstacles where Manning’s equation should not be applied.</p> <p><strong>Timeseries_Data</strong></p> <p>t: Time measured in days since the first day or “0-January-0000” for cases when specific dates were available. Dimension: 1, time step.</p> <p>A: Reach-averaged cross-sectional area of flow in m<sup>2</sup>. Dimension: Reach, time step.</p> <p>Q_true: True reach-averaged discharge (m<sup>3</sup>/s). Dimension: Reach, time step.</p> <p>Q_ptb: Perturbed discharge (m<sup>3</sup>/s), including 10000 realizations for each measurement. Dimension: Good reach, time step, 10000.</p> <p>med_Q_ptb: Median perturbed discharge (m<sup>3</sup>/s) across the 10000 realizations. Dimension: Good reach, time step.</p> <p>sigma_Q_ptb: Uncertainty of the perturbed discharge (m<sup>3</sup>/s), calculated based on the interquartile range. Dimension: Good reach, time step.</p> <p>W_true: True reach-averaged river width (m). Dimension: Reach, time step.</p> <p>W_ptb: Perturbed river width (m), including 10000 realizations for each measurement. Dimension: Good reach, time step, 10000.</p> <p>med_W_ptb: Median perturbed river width (m) across the 10000 realizations. Dimension: Good reach, time step.</p> <p>H_true: True reach-averaged water surface elevation (m). Dimension: Reach, time step.</p> <p>H_ptb: Perturbed water surface elevation (m), including 10000 realizations for each measurement. Dimension: Good reach, time step, 10000.</p> <p>S_true: True reach-averaged water surface slope (m/m). Dimension: Reach, time step.</p> <p>S_ptb: Perturbed water surface slope (m/m), including 10000 realizations for each measurement. Dimension: Good reach, time step, 10000.</p> <p>med_S_ptb: Median perturbed water surface slope (m/m) across the 10000 realizations. Dimension: Good reach, time step.</p> <p>dA_true: True reach-averaged change in the cross-sectional area (m<sup>2</sup>). Dimension: Good reach, time step.</p> <p>dA_ptb: Perturbed change in the cross-sectional area (m<sup>2</sup>), including 10000 realizations for each measurement. Dimension: Good reach, time step, 10000.</p> <p>med_dA_ptb: Median perturbed change in the cross-sectional area (m<sup>2</sup>) across the 10000 realizations. Dimension: Good reach, time step.</p> <p>sigma_dA_ptb: Uncertainty of the perturbed change in the cross-sectional area (m<sup>2</sup>), calculated based on the interquartile range. Dimension: Good reach, time step.</p> <p>A0_true: True baseline cross-sectional area (m<sup>2</sup>). Dimension: Good reach, 1.</p> <p>A0: Perturbed baseline cross-sectional area (m<sup>2</sup>), including 100 realizations for each parameter. Dimension: Good reach, 100.</p> <p>na_true: True friction coefficient. Dimension: Good reach, 1.</p> <p>na: Perturbed friction coefficient, including 100 realizations for each parameter. Dimension: Good reach, 100.</p> <p>b_true: True exponent coefficient. Dimension: Good reach, 1.</p> <p>b: Perturbed exponent coefficient, including 100 realizations for each parameter. Dimension: Good reach, 100.</p>
A dataset of synthetically generated code blocks for the learning of WCET on Cortex A53
<p>WE-HML Dataset WCET code block with different pollution factors on data cache for Cortex A53</p><p>Git : https://gitlab.inria.fr/puaut/we-hml</p><p>Paper : https://ieeexplore.ieee.org/document/9545301</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.