Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,049
datasets available to search
ShareScore release 0.9.0
Dataset results
1,049 results for “Robustness”
Trees and alignments for: A robust phylogenomic framework for the calamoid palms
<p>Target file, alignments, gene trees and species trees from phylogenomic analyses in Kuhnhäuser et al. (2021), A robust phylogenomic framework for the calamoid palms, Molecular Phylogenetics and Evolution. <a href="https://doi.org/10.1016/j.ympev.2020.107067">https://doi.org/10.1016/j.ympev.2020.107067</a>.</p> <p>Raw sequence data are deposited in the European Nucleotide Archive of the European Bioinformatics Institute (<a href="https://www.ebi.ac.uk/ena">https://www.ebi.ac.uk/ena</a>) under project number PRJEB40689. Scripts for all phylogenetic analyses are available at <a href="https://github.com/BenKuhnhaeuser/PhyloFrame">https://github.com/BenKuhnhaeuser/PhyloFrame</a>.</p>
Datasets for testing the robustness of LiDAR vegetation metrics to varying point densities
<p><span>The calculation of vegetation metrics from LiDAR point clouds might be affected by the available point density of a dataset. Testing how the same LiDAR vegetation metrics differ with different point densities can therefore inform about their robustness for upscaling metrics to other areas or other LiDAR point clouds. The datasets made available here were generated to test the robustness of LiDAR vegetation metrics to varying point densities and spatial resolutions (i.e., plots of 1 × 1 m, 2 × 2 m, 5 × 5 m and 10 × 10 m size). A total of 25 LiDAR vegetation metrics representing different aspects of vegetation height, vegetation cover and structural complexity were tested (see metric definition in Kissling et al. 2023, </span><span><a href="https://doi.org/10.1016/j.dib.2022.108798"><span>https://doi.org/10.1016/j.dib.2022.108798</span></a></span><span>). The metric calculation was similar to the metric calculation in the Laserchicken software (Meijer et al. 2020, </span><span><a href="https://doi.org/10.1016/j.softx.2020.100626"><span>https://doi.org/10.1016/j.softx.2020.100626</span></a></span><span>) and the Laserfarm workflow (Kissling et al. 2022, https://doi.org/10.1016/j.ecoinf.2022.101836). The Dutch AHN4 dataset from the years 2020–2022 with a point density of 20–30 points/m<sup>2</sup> was used. Initially, 100 plots (i.e., squared polygons around centre points) were randomly placed across the Netherlands in Dutch Natura 2000 sites that predominantly contain woodland habitats (using shapefiles from the European Environmental Agency). For each centre point, square polygons of the desired resolutions (i.e., 1 × 1 m, 2 × 2 m, 5 × 5 m or 10 × 10 m plot size) were generated. The square polygons were subsequently used to clip the LiDAR point clouds from the Dutch AHN4 point cloud dataset. Since not all locations of the 100 randomly placed plots contained points, the actual sample sizes were slightly smaller than 100, i.e., 94 plots for the 1 × 1 m, 2 × 2 m and 5 × 5 m resolution and 95 plots for the 10 × 10 m resolution. Metrics were calculated with the original point density of the Dutch AHN4 dataset (20–30 points/m2) and with six systematically down-sampled point clouds for the same plots (i.e., keeping 5%, 10%, 20%, 40%, 60% and 80% of the points in the original point clouds). For each clipped point cloud of a plot at a given resolution, the points were first sorted according to their GPS acquisition time (from earliest to latest). Points were then systematically discarded and only 5%, 10%, 20%, 40%, 60% and 80% of the points in the original point clouds were kept. The kept points were used for calculating the 25 LiDAR vegetation metrics. </span></p>
Robust joint registration of multiple stains and MRI for multimodal 3D histology reconstruction: Application to the Allen human brain atlas
Open the record for dataset details and reuse information.
Robust framework and software implementation for fast speciation mapping
<p>R script and raw data to test the sparse excitation energy XAS procedure.</p>
VSR Databases used in article "Standardization of noisy volcano-seismic waveforms as a key step towards station-independent, robust automatic recognition"
<p>This dataset contains required volcano-seismic waveform DBs (<em>dec.95M.16c</em> and <em>dec.09U.4c</em>) used in the article:</p> <p>"<em>Standardization of noisy volcano-seismic waveforms as a key step towards station-independent, robust automatic recognition</em>",</p> <p>published in the Seismological Research Letters (<a href="https://doi.org/10.1785/0220180334">https://doi.org/10.1785/0220180334</a>). The authors want to thank everyone at the Instituto Andaluz of Geofísica (<a href="http://iagpds.ugr.es">http://iagpds.ugr.es</a>), precisely to Prof. Jesús Ibáñez and Dr. Javier Almendros, IPs of several research projects which </p> <p>have made possible the monitoring of Deception Island since early 1990s.</p> <p>This project has received funding from the European Union’s Horizon 2020 research and innovation programme under the Marie Sklodowska-Curie Grant Agreement No.[749249] (VULCAN.ears).</p>
Robust functional mapping of layer-selective responses in human lateral geniculate nucleus with high-resolution 7T fMRI
Open the record for dataset details and reuse information.
An Empirical Comparison of Meta-Modeling Techniques for Robust Design Optimization
<p>This is the data and source code used in the paper below:</p> <p>Sibghat Ullah, Hao Wang, Stefan Menzel, Bernhard Sendhoff and Thomas Bäck, “An Empirical Comparison of Meta-Modeling Techniques for Robust Design Optimization”, in 2019 IEEE Symposium Series on Computational Intelligence (SSCI), Xiamen, China, 6-9 December 2019, doi: 10.1109/SSCI44817.2019.9002805</p> <p>This research investigates the potential of using meta-modeling techniques in the context of robust optimization namely optimization under uncertainty/noise. A systematic empirical comparison is performed for evaluating and comparing different meta-modeling techniques for robust optimization. The experimental setup includes three noise levels, six meta-modeling algorithms, and six benchmark problems from the continuous optimization domain, each for three different dimensionalities. Two robustness definitions: robust regularization and robust composition, are used in the experiments. The meta-modeling techniques are evaluated and compared with respect to the modeling accuracy and the optimal function values. The results clearly show that Kriging, Support Vector Machine and Polynomial regression perform excellently as they achieve high accuracy and the optimal point on the model landscape is close to the true optimum of test functions in most cases.</p>
Data for: Machine learning identifies robust matrisome markers and regulatory mechanisms in cancer
<p>The expression and regulation of matrisome genes - the ensemble of extracellular matrix, ECM, ECM-associated proteins and regulators as well as cytokines, chemokines and growth factors - is of paramount importance for the many biological processes and signals within the tumor microenvironment. The availability of large and diverse multi-omics data enables mapping and understanding the regulatory circuitry governing the tumor matrisome to an unprecedented level, though such a volume of information requires robust approaches to data analysis and integration. In this study, we show that combining Pan-Cancer expression data from The Cancer Genome Atlas (TCGA) with genomics, epigenomics and microenvironmental features from TCGA and other sources enables the identification of “landmark” matrisome genes and machine learning-based reconstruction of their regulatory networks in 74 clinical and molecular subtypes of human cancers and approx. 6700 patients. These results, enriched for prognostic genes and cross-validated markers at the protein level, unravel the role of genetic and epigenetic programs in governing the tumor matrisome and allow the prioritization of tumor-specific matrisome genes (and their regulators) for the development of novel therapeutic approaches.</p>
DWCox: A Density-Weighted Cox Model for Outlier-Robust Prediction of Prostate Cancer Survival
<p>This package, <strong>DWCox</strong>, implements a <strong>d</strong>ensity-<strong>w</strong>eighted <strong>Cox</strong> regression model that is more robust against outliers in the training data. DWCox gives more accurate predictions than the standard Cox regression on prostate cancer survival, especially in cases where the training data are expected to contain a lot of outliers. More details can be found in our paper (coming soon) and the README file inside this package.</p>
Robustness assessment of a C++ implementation of a quantized (int8) version of the LeNet-5 convolutional neural network
<p>The architecture of the LeNet-5 convolutional neural network (CNN) was defined by LeCun in its paper "Gradient-based learning applied to document recognition" (<a href="https://ieeexplore.ieee.org/document/726791">https://ieeexplore.ieee.org/document/726791</a>) to classify images of hand written digits (MNIST dataset).</p><p>This architecture has been customized to use Rectified Linear Unit (ReLU) as activation functions instead of Sigmoid, and 8-bit integers for weights and activations instead of floating-point.</p><p>It consists of the following layers:</p><ul><li><strong>conv1</strong>: Convolution 2D, 1 input channel (28x28), 3 output channels (28x28), kernel size 5, stride 1, padding 2.</li><li><strong>relu1</strong>: Rectified Linear Unit (3@28x28).</li><li><strong>max1</strong>: Subsampling buy max pooling (3@14x14).</li><li><strong>conv2</strong>: Convolution 2D, 3 input channels (14x14), 6 output channels (14x14), kernel size 5, stride 1, padding 2.</li><li><i><strong>relu2</strong></i>: Rectified Linear Unit (6@14x14).</li><li>max2: Subsampling buy max pooling (6@7x7).</li><li><i><strong>fc1</strong></i>: Fully connected (294, 147)</li><li><i><strong>fc2</strong></i>: Fully connected (147, 10)</li></ul><p>The fault hypotheses for this work include the occurrence of:</p><ul><li><strong>BF</strong>: single, double-adjacent and triple-adjacent bit-flip faults</li><li><strong>S0</strong>: single, double-adjacent and triple-adjacent stuck-at-0 faults</li><li><strong>S1</strong>: single, double-adjacent and triple-adjacent stuck-at-1 faults</li></ul><p>In the memory cells containing all the parameters of the CNN: </p><ul><li><strong>w</strong>: weights (int8)</li><li><strong>zw</strong>: zero point of the weights (int8)</li><li><strong>b</strong>: biases (int32)</li><li><strong>z</strong>: zero point (int8)</li><li><strong>m</strong>: m (int32)</li></ul><p>Images 200 to 249 from the MNIST dataset have been used as workload.</p><p>This dataset contains the raw data obtained from running exhaustive fault injection campaigns for all considered fault models, targeting all considered locations and for all the images in the workload.</p><p>In addition, the raw data have been lightly processed to obtain global data related to the particular bits and parameters affected by the faults, and the obtained failure modes.</p><h3>Files information</h3><ul><li><i>golden_run.csv</i>: Prediction obtained for all the images considered in the workload in the absence of faults (Golden Run). This is intended to act as oracle to determine the impact of injected faults. </li><li><i>single_faults/bit_flip</i> folder: Prediction obtained for all the images considered in the workload in presence of single bit-flip faults. There is one file for each parameter of each layer.</li><li><i>single_faults/stuck_at_0</i> folder: Prediction obtained for all the images considered in the workload in presence of single stuck-at-0 faults. There is one file for each parameter of each layer.</li><li><i>single_faults/stuck_at_1</i> folder: Prediction obtained for all the images considered in the workload in presence of single stuck-at-1 faults. There is one file for each parameter of each layer.</li><li><i>double_adjacent_faults/bit_flip</i> folder: Prediction obtained for all the images considered in the workload in presence of double adjacent bit-flip faults. There is one file for each parameter of each layer.</li><li><i>double_adjacent_faults/stuck_at_0</i> folder: Prediction obtained for all the images considered in the workload in presence of double adjacent stuck-at-0 faults. There is one file for each parameter of each layer.</li><li><i>double_adjacent_faults/stuck_at_1</i> folder: Prediction obtained for all the images considered in the workload in presence of double adjacent stuck-at-1 faults. There is one file for each parameter of each layer.</li><li><i>triple_adjacent_faults/bit_flip</i> folder: Prediction obtained for all the images considered in the workload in presence of triple adjacent bit-flip faults. There is one file for each parameter of each layer.</li><li><i>triple_adjacent_faults/stuck_at_0</i> folder: Prediction obtained for all the images considered in the workload in presence of triple adjacent stuck-at-0 faults. There is one file for each parameter of each layer.</li><li><i>triple_adjacent_faults/stuck_at_1</i> folder: Prediction obtained for all the images considered in the workload in presence of triple adjacent stuck-at-1 faults. There is one file for each parameter of each layer.</li></ul><h3>Methodology information</h3><p>First, the CNN was used to classify all the images of the workload in the absence of faults to get a reference to determine the impact of faults. This is golden_run.csv file.</p><p>After that, one fault injection experiment was executed for each bit of each element of each parameter of the CNN.</p><p>Each experiment consisted in:</p><ul><li>Affecting the bits (inverting it in case of bit-flip faults, setting it to 0 or 1 in case of stuck-at-0 or atuck-at-1 faults) identified by the mask.</li><li>Classifying all the images of the workload in the presence of this fault. The obtained output was stored in a given .csv file.</li><li>Removing the fault from the CNN by restoring the affected bits to its previous value.</li></ul><h3>List of variables (Name : Description (Possible values))</h3><ul><li><strong>IMGID</strong>: Integer number identifying the considered image (200-249).</li><li><strong>TENSORID</strong>: Integer number identiying the parameter affected by the fault (0 - No fault, 1 - conv1.w, 2 - conv1.zw, 3 - conv1.m, 4 - conv1.b, 5 - conv1.z, 6 - conv2.w, 7 - conv2.zw, 8 - conv2.m, 9 - conv2.b, 10 - conv2.z, 11 - fc1.w, 12 - fc1.zw, 13 - fc1.m, 14 - fc.b, 15 - fc1.z, 16 - fc2.w, 17 - fc2.zw, 18 - fc2.m, 19 - fc2.b, 20 - fc2.z)</li><li><strong>ELEMID</strong>: Integer number identiying the element of the parameter affected by the fault (-1 - No fault, [0-2] - {conv1.b, conv1.m, conv1.zw}, [0-74] - conv1.w, 0 - conv1.z, [0-5] - {conv2.b, conv2.m, conv2.zw}, [0-149] - conv2.w, 0 - {conv1.z, conv2.z, fc1.z, fc2.z}, [0-146] - {fc1.b, fc1.m, fc1.zw}, [0-43217] - fc1.w, [0-9] - {fc2.b, fc2.m, fc2.zw}, [0-1469] - fc2.w)</li><li><strong>MASK</strong>: 8-digit hexadecimal number identifying those bits affected by the fault ([00000000 - No fault, FFFFFFFF - all 32 bits faulty])</li><li><strong>FAULT</strong>: String identiying the type of fault (NF - No fault, BF - bit-flip, S0 - Stuck-at-0, S1 - Stuck-at-1)</li><li><strong>OUTPUT</strong>: 10 integer numbers provided by the CNN as output after processing the image. The highest value identifies the selected category for classification.</li><li><strong>SOFTMAX</strong>: 10 decimal numbers obtained after applying the softmax function to the provided output. They represent the probability of the image of belonging to the corresponding category for classification.</li><li><strong>PRED</strong>: Integer number representing the category predicted for the processed image.</li><li><strong>LABEL</strong>: integer number representing the actual category for the processed image.</li></ul>
Dataset related to publication: Robust radiative cooling via surface phonon coupling-enhanced emissivity from SiO2 micropillar arrays
<p>Dataset related to the publication:</p><p>Zhenmin Ding, Xin Li, Hulin Zhang, Dukang Yan, Jérémy Werlé, Ying Song, Lorenzo Pattelli, Jiupeng Zhao, Hongbo Xu, Yao Li. Robust radiative cooling via surface phonon coupling-enhanced emissivity from SiO2 micropillar arrays. <i>International Journal of Heat and Mass Transfer</i>, 220, 125004 (2024). doi: <a href="https://doi.org/10.1016/j.ijheatmasstransfer.2023.125004">10.1016/j.ijheatmasstransfer.2023.125004</a></p><p>The repository contains MATLAB/Octave scripts to perform rigorous coupled-wave analysis (RCWA) simulations for a SiO2 layer decorated with micropillars.</p><p>The main script runs a series of rigorous electromagnetic simulations over the atmospheric transparency window wavelength range (8-13 µm) for all combination of three main structural parameters (pillar diameter, spacing and height), within a user-defined range.</p><p>Running the code requires the RETICOLO v9 RCWA code:</p><blockquote><p>Jean-Paul Hugonin, & Philippe Lalanne. (2021). Light-in-complex-nanostructures/RETICOLO: V9. Zenodo. <a href="https://doi.org/10.5281/zenodo.4419063">https://doi.org/10.5281/zenodo.4419063</a></p></blockquote>
R Code and Re-analyzed Datasets for: Robust approaches for the quantitative analysis of genome formula variation in multipartite and segmented viruses
<p>This submission includes all the scripts and data analyzed in the manuscript "Robust approaches for the quantitative analysis of genome formula variation in multipartite and segmented viruses". This manuscript is a technical note on how genome formula data can be analyzed. There are no new experimental data in the manuscript, as published datasets are re-analyzed. Here we reproduce those datasets as formatted for our analysis, for the convenience of the reader. Please consult the README.txt file first.</p> <p>The corresponding paper was published in Viruses <em>16</em>(2): 270. (<a href="https://doi.org/10.3390/v16020270">https://doi.org/10.3390/v16020270</a>).</p> <p>This is the second version of the code, corresponding to the final version of the paper. The intial restricted version for review had a DOI 10.5281/zenodo.10355273.</p> <p> </p>
Robust genetic codes enhance protein evolvability
<p>The repository contains all data for our manuscript on protein evolvability under rewired genetic codes: https://www.biorxiv.org/content/10.1101/2023.06.20.545706v1</p> <p>The corresponding code is available on GitHub: https://github.com/parizkh/rewired_codes_landscapes</p>
Collider Bias Correction for Multiple Covariates in GWAS Using Robust Multivariable Mendelian Randomization
<p>This repository contains the data underlying the figures in paper "Collider Bias Correction for Multiple Covariates in GWAS<br>Using Robust Multivariable Mendelian Randomization".</p> <p> </p> <p> </p> <p>The file names and sheet names in the xlsx file indicate the corresponding figures of data. </p> <p><br>The underlying data of manhattan plots and QQ plots are in text file. For other figures, the underlying data are in the spreadsheet.</p> <p>In each file, column names indicate the MVMR method used to obtain the result. </p> <p>For example: </p> <p>In text files:</p> <p>The abbreviation "mPC" refers to metabolomic principle components.</p> <p>beta_no_correction: the SNP effect estimate without bias correction.</p> <p>beta_cml or beta_MVMR_cml: the standard error of SNP effect estimate after the bias correction of MVMR-cML.</p> <p>SE_UVMR_cml: the standard error of SNP effect estimate after the bias correction of UVMR-cML.</p> <p>p_value_Egger or p_value_MVMR_Egger: the p-value of SNP effect estimate after the bias correction of MVMR-Egger regression.</p> <p><br>In the spreadsheet, column names follow the same style. </p> <p>The GWAS data is also available. The column names follows the plink output file. The detailed explanations are available at https://www.cog-genomics.org/plink/2.0/formats#glm_linear</p>
Robust Method for Property Prediction via Artificial Neural Networks: Incorporating Key Structural Features for Carbon Dioxide – Ionic Liquid Mixtures
<p>This Dataset comprises two sub-sets of information:</p> <ul> <li>Database and Results of the work present in the paper "Robust Method for Property Prediction via Artificial Neural Networks: Incorporating Key Structural Features for Carbon Dioxide – Ionic Liquid Mixtures" published in The Journal of Physical Chemistry B (https://doi.org/10.1021/acs.jpcb.4c04432).</li> <li>Sample of the code used, in order to reproduce any of the results presented above. This can be found in the previous version of this Dataset (v1.0 https://zenodo.org/records/11216901)</li> </ul> <p> </p> <p>Regarding the sample code, an example for all ANN Models used in this work is provided. This includes the three models used:</p> <ol> <li>One based only on Critical Properties of Ionic Liquids (CRT Model)</li> <li>One based only on Structural Properties of Ionic Liquids (STR Model)</li> <li>One combination of the previous models, taking into account both Critical and Structural Properties (COMB Model)</li> </ol> <p>In this manner, it is possible to observe the differences between the performance of the different models, either through statiscal analysis or using graphical representation. This allows for the benchmarking to be done in a more concise way.</p>
Robust Damage Estimation of Typhoon Goni on Coconut Crops with Sentinel-2 Imagery
<p>Damage estimation status of coconut trees plantation in the Phillippines derived from Sentinel-2 Imagery. Overall we estimated that 14.1 M coconut trees were affected by the typhoon inside our area of study. Please refer to our <a href="https://www.mdpi.com/2072-4292/13/21/4302">original</a> publication for more details.</p> <p>0: Uncertain</p> <p>1: No Data</p> <p>2: Background class</p> <p>3: Unchanged coconut plantation</p> <p>4: Damaged coconut plantation</p> <p>5: New coconut plantation</p> <p> </p> <p> </p> <p> </p>
Common Phone: A Multilingual Dataset for Robust Acoustic Modelling
<p><em>Release Date: 17.01.22</em></p> <p><strong>Welcome to Common Phone 1.0</strong></p> <p><strong>Legal Information</strong></p> <p><em>Common Phone</em> is a subset of the <em>Common Voice</em> corpus collected by <em>Mozilla Corporation</em>. By using <em>Common Phone</em>, you agree to the <a href="https://commonvoice.mozilla.org/en/terms">Common Voice Legal Terms</a>. <em>Common Phone</em> is maintained and distributed by speech researchers at the <a href="https://lme.tf.fau.de/">Pattern Recognition Lab</a> of Friedrich-Alexander-University Erlangen-Nuremberg (<a href="https://www.fau.de/">FAU</a>) under the <a href="https://creativecommons.org/publicdomain/zero/1.0/">CC0 license</a>.</p> <p>Like for <em>Common Voice</em>, you must not make any attempt to identify speakers that contributed to <em>Common Phone</em>.</p> <p><strong>About <em>Common Phone</em></strong></p> <p>This corpus aims to provide a basis for Machine Learning (ML) researchers and enthusiasts to train and test their models against a wide variety of speakers, hardware/software ecosystems and acoustic conditions to improve generalization and availability of ML in real-world speech applications.<br> The current version of <em>Common Phone</em> comprises 116,5 hours of speech samples, collected from 11.246 speakers in 6 languages:</p> <table align="center"> <thead> <tr> <th> <p><strong>Language</strong></p> </th> <th> <p><strong>Speakers</strong></p> </th> <th> <p><strong>Hours</strong></p> </th> </tr> </thead> <tbody> <tr> <td> </td> <td> <p><code>train</code> / <code>dev</code> / <code>test</code></p> </td> <td> <p><code>train</code> / <code>dev</code> / <code>test</code></p> </td> </tr> <tr> <td> <p>English</p> </td> <td> <p>4716 / 771 / 774</p> </td> <td> <p>14.1 / 2.3 / 2.3</p> </td> </tr> <tr> <td> <p>French</p> </td> <td> <p>796 / 138 / 135</p> </td> <td> <p>13.6 / 2.3 / 2.2</p> </td> </tr> <tr> <td> <p>German</p> </td> <td> <p>1176 / 202 / 206</p> </td> <td> <p>14.5 / 2.5 / 2.6</p> </td> </tr> <tr> <td> <p>Italian</p> </td> <td> <p>1031 / 176 / 178</p> </td> <td> <p>14.6 / 2.5 / 2.5</p> </td> </tr> <tr> <td> <p>Spanish</p> </td> <td> <p>508 / 88 / 91</p> </td> <td> <p>16.5 / 3.0 / 3.1</p> </td> </tr> <tr> <td> <p>Russian</p> </td> <td> <p>190 / 34 / 36</p> </td> <td> <p>12.7 / 2.6 / 2.8</p> </td> </tr> <tr> <td> <p><strong>Total</strong></p> </td> <td> <p>8417 / 1409 / 1420</p> </td> <td> <p>85.8 / 15.2 / 15.5</p> </td> </tr> </tbody> </table> <p> </p> <p>Presented <code>train</code>, <code>dev</code> and <code>test</code> splits are <strong>not identical</strong> to those shipped with <em>Common Voice</em>. Speaker separation among splits was realized by only using those speakers that had provided age and gender information. This information can only be provided as a registered user on the website. When logged in, the session ID of contributed recordings is always linked to your user, thus we could easily link recordings to individual speakers. Keep in mind this would not be possible for unregistered users, as their session ID changes if they decide to contribute more than once.<br> During speaker selection, we considered that some speakers had contributed to more than one of the six <em>Common Voice</em> datasets (one for each language). In <em>Common Phone</em>, a speaker will only appear in one language.<br> The dataset is structured as follows:</p> <ul> <li>Six top-level directories, one for each language.</li> <li>Each language folder contains: <ul> <li>[train|dev|test].csv files listing audio files, respective speaker ID and plain text transcript.</li> <li>meta.csv provides speaker information: age group, gender, language, accent (if available) and which of the three splits this speaker was assigned to. File names match corresponding audio file names except their extension.</li> <li>/grids/ contains phonetic transcription for every audio file in Praat TextGrid format.</li> <li>/mp3/ contains audio files in mp3, identical to those of <em>Common Voice</em>, e.g., sampling rates have been preserved and may vary for different files.</li> <li>/wav/ contains raw audio files in 16 bits/sample, 16 kHz single channel. They had been created from the original mp3 audios. We provide them for convenience, keep in mind that their source had undergone MP3-compression.</li> </ul> </li> </ul> <p><strong>Where does the phonetic annotation come from?</strong></p> <p>Phonetic annotation was computed via <a href="https://clarin.phonetik.uni-muenchen.de/BASWebServices/interface/Pipeline">BAS Web Services</a>. We used the regular Pipeline (G2P-MAUS) without ASR to create an alignment of text transcripts with audio signals. We chose International Phonetic Alphabet (IPA) output symbols as they work well even in a multi-lingual setup. <em>Common Phone</em> annotation comprises 101 phonetic symbols, including silence.</p> <p><strong>Why <em>Common Phone</em>?</strong></p> <ul> <li>Large number of speakers and varying acoustic conditions to improve robustness of ML models</li> <li>Time-aligned IPA phonetic transcription for every audio sample</li> <li>Gender-balanced and age-group-matched (equal number of female/male speakers in every age group)</li> <li>Support for six different languages to leverage multi-lingual approaches</li> <li>Original MP3 files plus standard WAVE files</li> </ul> <p><strong>Is there any publication available?</strong></p> <p><em>Yes, a paper describing Common Phone in detail is currently under revision for LREC </em><em>2022. You can access a pre-print version on arXiv entitled “<a href="https://arxiv.org/abs/2201.05912">Common Phone: A Multilingual Dataset for Robust Acoustic Modelling</a>”.</em></p>
Supporting data for "Robust spin squeezing from the tower of states of U(1)-symmetric spin Hamiltonians"
<p>Supporting data for "Robust spin squeezing from the tower of states of U(1)-symmetric spin Hamiltonians" (<a href="https://link.aps.org/doi/10.1103/PhysRevA.105.022625">https://link.aps.org/doi/10.1103/PhysRevA.105.022625</a>, <a href="https://arxiv.org/abs/2103.07354">https://arxiv.org/abs/2103.07354</a>), by Tommaso Comparin, Fabio Mezzacapo and Tommaso Roscilde. If you use these data in a scientific work, please cite the corresponding article.<br> For additional details, please contact Tommaso Comparin (tommaso.comparin@ens-lyon.fr).</p> <p>This dataset includes tVMC results for several values of the coupling exponent alpha and of the system size N (see file names). Each file includes a set of relevant observables (see file header).</p> <p>These results are directly shown in Figures 1, 3, 7, 9. Further data processing leads to Fig. 4.</p> <p><br> Details on the variational Ansatz for tVMC<br> - We employ the pair-product Ansatz defined in the main text. In general, there are N*(N-1)/2 independent spin pairs with 4 possible states for each pair, leading to 4*N complex coefficients.<br> - Thanks to translational invariance, we can use coefficients which only depend on the distance between the two spins, reducing the number of independent pairs to N/2 (for a one-dimensional chain with periodic boundary conditions).<br> - We also impose the spin-inversion symmetry for each spin pair, so that configurations like {up,down} and {down,up} have the same coefficient.<br> - Therefore the total number of variational coefficients in our tVMC simulations is equal to N.</p> <p>NOTE:<br> Data are provided without error bars. An analysis of the statistical/systematic errors is included in the folder Errors, for some representative cases.</p>
Modeling robust COVID-19 intensive care unit occupancy thresholds for imposing mitigation to prevent exceeding capacities
<p>Simulation output files for 'Modeling robust COVID-19 intensive care unit occupancy thresholds for imposing mitigation to prevent exceeding capacities'.</p> <p>Simulating COVID-19 transmission and hospital burden to assess at which intensive care unit (ICU) occupancies mitigation, that reduces transmission, needs to be triggered to avoid exceeding ICU capacity limits, using the city of Chicago, Illinois as an example.</p> <p>Manuscript is under review for scientific publication, (see <a href="https://www.medrxiv.org/content/10.1101/2021.06.27.21259530v1">preprint on medRxiv</a>) and scripts are available from the GitHub repository at https://github.com/numalariamodeling/ICUtrigger_covid_chicago_paper_2021. </p> <p>Simulation output files uploaded per scenario including projected COVIID-19 transmission and burden trajectories for Chicago city for March 2020 to May 2021 per day.</p> <p>Simulation scenarios:</p> <p><reopening % above ICU capacity>_<delay after reaching ICU threshold>_<%mitigation>_<common simulation name> i.e. `50perc_1daysdelay_pr6_triggeredrollback_reopen`</p> <ul> <li>`emodl` file <ul> <li>required file for COVID-19 transmission model in the <a href="https://docs.idmod.org/projects/cms/en/latest/index.html">Compartmental Modeling Software</a> (see <a href="https://github.com/numalariamodeling/ICUtrigger_covid_chicago_paper_2021">GitHub repository</a> for details)</li> </ul> </li> <li>sampled_parameters.csv <ul> <li>simulation input and scenario parameters, (nrow=4400, 400 unique parameter combinations * 11 scenario values)</li> </ul> </li> <li>rt_trajectoriescovidregion_11.csv <ul> <li>estimated reproductive numbers per trajectory for complete timeline per day</li> </ul> </li> <li>trajectoriesDat_region_11_traces.csv <ul> <li>filtered to include top 100 trajectories fitted to ICU data</li> </ul> </li> <li>trajectoriesDat_region_trimfut.csv <ul> <li>truncated to only include projections after September 1st 2020</li> </ul> </li> </ul> <p>The folder `mainfigures_csvs.zip` includes processed simulation output data for the publication figures.</p>
Generated Data for the Manuscript "Nonideality-Aware Training for Accurate and Robust Low-Power Memristive Neural Networks"
<p>The file contains data generated and referred to in the text and the figures of the manuscript.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.