Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

376

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

376 results for “scalability”

Learn how ShareScore rates datasets ↗
zenodo56/100

Size scalability of Monte Carlo simulations applied to oxidized polypyrrole systems: Data and Codes

<p>This work generalizes our recently proposed coarse grained force field (CGFF) for halogen oxidized PPy in the condensed phases and introduces a novel implementation of the Nettropolis Monte Carlo (MMC) simulation based on the CGFF that enables simulations of polymer systems with more than<br>100000 particles. The MMC implementation utilizes a combination of CPU and GPUs and exploits a numerical approximation based on polynomial piecewise interpolation for the calculation of the CGFF pairwise additive terms. Our simulations evidence that the oxidized PPy thermodynamic and structural properties are consistent as the system size is scaled up. Predicted properties include density, enthalpy, potential energy, heat capacity, coefficient of thermal expansion, caloric curve, glass transition temperature range, compressibility, bulk modulus, radial distribution functions, and polymer chain characteristics.</p>

opencc-by-4.0Nov 2024View details →
zenodo52/100

Mappings for "Developing a Scalable Annotation Method for Large Datasets That Enhances Alarms With Actionability Data to Increase Informativeness: Mixed Methods Approach"

<p>Studies identified false and non-actionnable alarms as a factor for alarm fatigue in intensive care units.</p> <p>To annotate patient alarms, and analyse the alarm situation in intensive care units, we conceptualized and performed data mappings related to airway management and medication interventions. The mappings were based on information retrieved from the patient data management system (PDMS) and clinical expertise. For the airway management mappings, we used additional resources such as ISO 19223:2019 or ventilator instruction manuals. The mappings do not include patient data.</p> <p>As the mappings are generic, they could be used in other contexts than alarm annotation and research.</p> <p><strong>1. Respiratory Management Mappings:</strong></p> <ul> <li>General tables summarizing the 1) categories based on ISO 19223:2019 to describe respiratory support therapies (RSTs), 2) defining the invasiveness level of a RST and 3) listing the abbreviations used in the mappings</li> <li> <p>Tables including PDMS entries for airway devices (ADs), ventilation devices (VDs), and ventilation modes (VMs)</p> </li> <li> <p>Mapping of AD entries (from the PDMS) to defined categories</p> </li> <li> <p>Mapping of VDs, VMs, and ADs to defined RSTs, including information on invasiveness</p> </li> <li> <p>Table specifying suitable ventilation parameters in the context of each RST</p> </li> </ul> <p><strong>2. Medication Mappings:</strong></p> <ul> <li> <p>General tables providing information on physiological alarm conditions (PACs), interventions, routes, and techniques of administration of interest</p> </li> <li> <p>Mapping of routes of administration to techniques of administration including PDMS entries</p> </li> <li> <p>Mapping of active ingredients (including SNOMED CT Fully Specified Names and Identifiers), related PDMS information, and routes and techniques of administration to defined PAC and interventions</p> </li> </ul>

opencc-by-4.0Jan 2023View details →
zenodo52/100

Datasets for evaluating scalable supervised learning for synthesize-on-demand chemical libraries

<p>This repository contains datasets for the manuscript &quot;Evaluating scalable supervised learning for synthesize-on-demand chemical libraries&quot;:</p> <ul> <li><strong>ams_all_preds.csv.gz</strong>: The AMS dataset predictions when using an RF or baseline model trained on the training dataset. Includes the predicted score and rank from each model for each compound. We started with 8,434,707 AMS compounds and detected that 247,025 were in the LC or MLPCN training data. These were removed from the AMS list, leaving 8,187,682 compounds to score. The compound matching was done on the SMILES that we canonicalized in rdkit.</li> <li><strong>ams_order_results.csv.gz</strong>: Information about the 1,024 compounds purchased from the AMS library. Excludes the 4 AMS compounds that were incompletely dissolved. Includes the chemical feature representation, information from the vendor, RF and baseline model predictions, screening results, and clustering results.</li> <li><strong>baseline_weight.npy</strong>: The saved Similarity Baseline model, which consists of the active compounds in the training data. This model was used to score the AMS library. See the <a href="https://github.com/gitter-lab/pria-ams-enamine">GitHub repository</a>&nbsp;for code to load the model and make predictions on new compounds.</li> <li><strong>cdd_training_data.tar.gz</strong>: The LC1234 and MLPCN PriA-SSB screening data exported from CDD.</li> <li><strong>enamine_costs_clustered_v3_with_nneighbor.csv.gz</strong>: Contains 5,620 Enamine compounds that were selected based on the RF prediction score and availability. This file also contains the Taylor-Butina cluster ID when clustering the training compounds, 1,024 tested AMS compounds, and top-ranked Enamine compounds at a 0.4 threshold. The nearest neighbor compounds in the training and AMS sets are also included along with compound information from Enamine, RF model scores, and chemical feature representations.</li> <li><strong>enamine_dose_response_curve_plots.xlsx</strong>: Images of the dose response curves from all three runs on the 68 Enamine compounds. If a compound was tested multiple times, multiple curves are shown in the same plot. The compound structure images and SMILES are exported from CDD, not generated with RDKit.</li> <li><strong>enamine_dose_response_curves.tsv</strong>: The dose response curve summaries from all three runs on the 68 Enamine compounds. If a compound was tested multiple times, only the highest-quality dose response curve was used.</li> <li><strong>enamine_final_list.csv.gz</strong>: The final 100 filtered compounds from&nbsp;<code>enamine_top_10000.csv.gz</code>. Contains compound information from Enamine as well as RF model scores, chemical feature representations, and clustering results.</li> <li><strong>enamine_PriA-SSB_dose_response_data.tar.gz</strong>: The dose response screening data from all three runs on the 68 Enamine compounds. The 2021-06-16 run was originally screened on 2020-08-24. 2021-06-16 is the date the compound identities were corrected. This run contains two 1,536 well plates.</li> <li><strong>enamine_top_10000.csv.gz</strong>: Top 10,000 predictions from the Enamine REAL dataset using the selected RF model. Contains compound information from Enamine as well as RF model scores, chemical feature representations, and clustering results.</li> <li><strong>master_df.csv.gz</strong>: The output of preprocessing the files in&nbsp;<code>cdd_training_data.tar.gz</code>. Contains 441,900 rows.</li> <li><strong>random_forest_classification_139.pkl</strong>: The saved RF classification model with&nbsp;hyperparameter ID 139. This model was used to score the AMS and Enamine REAL libraries. See the <a href="https://github.com/gitter-lab/pria-ams-enamine">GitHub repository</a> directory for code to load the model and make predictions on new compounds.</li> <li><strong>train_ams_real_cluster.csv.gz</strong>: Contains cluster IDs for Taylor-Butina clustering at a 0.4 threshold applied to the training compounds, 1,024 tested AMS compounds, and top-ranked compounds from Enamine. Includes the chemical features, dataset to which the compound belongs, leader compound for each cluster, and whether the compound is a known hit.</li> <li><strong>training_df_single_fold.csv.gz</strong>: This is all ten folds in&nbsp;<code>training_folds.tar.gz</code>&nbsp;merged for convenience. Contains 427,300 compounds.</li> <li><strong>training_df_single_fold_with_ams_clustering.csv.gz</strong>: Contains cluster IDs for Taylor-Butina clustering applied to the 427,300 training compounds and the 1,024 tested AMS compounds. Different clustering results are shown at the 0.2, 0.3, and 0.4 thresholds. Includes the leader compound for each cluster. Although the training and AMS compounds were clustered jointly, only the training compounds&#39; clusters are shown. The AMS compounds&#39; clusters are in&nbsp;<code>ams_order_results.csv.gz</code>.</li> <li><strong>training_folds.tar.gz</strong>: The LC1234 and MLPCN training data split into ten folds. This dataset with 427,300 compounds was used for cross validation and model selection. This dataset is derived from&nbsp;<code>master_df.csv.gz.</code></li> </ul> <p>If you use&nbsp;these&nbsp;datasets in a publication, please cite:</p> <p>Moayad Alnammi, Shengchao Liu, Spencer S. Ericksen, Gene E. Ananiev, Andrew F. Voter, Song Guo, James L. Keck, F. Michael Hoffmann, Scott A. Wildman, Anthony Gitter.&nbsp;<a href="https://doi.org/10.1021/acs.jcim.3c00912">Evaluating scalable supervised learning for synthesize-on-demand chemical libraries</a>.&nbsp;<em>Journal of Chemical Information and Modeling</em>&nbsp;2023.</p> <p>See&nbsp;PubChem AID&nbsp;<a href="https://pubchem.ncbi.nlm.nih.gov/bioassay/1272365">1272365</a>, AID&nbsp;<a href="https://pubchem.ncbi.nlm.nih.gov/bioassay/1918986">1918986</a>,&nbsp;and the associated publications for details about the PriA-SSB screening data. The screening datasets were compiled from three separate sources that should all be cited if the training dataset is used in a publication:</p> <ul> <li>Moayad Alnammi, Shengchao Liu, Spencer S. Ericksen, Gene E. Ananiev, Andrew F. Voter, Song Guo, James L. Keck, F. Michael Hoffmann, Scott A. Wildman, Anthony Gitter.&nbsp;<a href="https://doi.org/10.1021/acs.jcim.3c00912">Evaluating scalable supervised learning for synthesize-on-demand chemical libraries</a>.&nbsp;<em>Journal of Chemical Information and Modeling</em>&nbsp;2023.</li> <li>Shengchao Liu<sup>+</sup>, Moayad Alnammi<sup>+</sup>, Spencer S. Ericksen, Andrew F. Voter, Gene E. Ananiev, James L. Keck, F. Michael Hoffmann, Scott A. Wildman, Anthony Gitter.&nbsp;<a href="https://doi.org/10.1021/acs.jcim.8b00363">Practical model selection for prospective virtual screening</a>.&nbsp;<em>Journal of Chemical Information and Modeling</em>&nbsp;2018.</li> <li>Andrew F. Voter<sup>+</sup>, Michael P. Killoran<sup>+</sup>, Gene E. Ananiev, Scott A. Wildman, F. Michael Hoffmann, James L. Keck.&nbsp;<a href="https://doi.org/10.1177/2472555217712001">A high-throughput screening strategy to identify inhibitors of SSB protein&ndash;protein interactions in an academic screening facility</a>.&nbsp;<em>SLAS Discovery</em>&nbsp;2018.</li> </ul> <ul> </ul>

opencc-by-4.0Oct 2021View details →
zenodo48/100

Supplementary Materials for "On the Scalability of Data Reduction Techniques in Current and Upcoming HPC Systems from an Application Perspective"

<p>Supplementary materials with all used benchmark scripts, plot scripts, benchmark results and PIConGPU example data for the submission to &quot;The 1st International Workshop on Data Reduction for Big Scientific Data (DRBSD-1)&quot; held in conjunction with ISC 2017 in Frankfurt, Germany.</p>

opencc-by-4.0Apr 2017View details →
zenodo48/100

Thermally switchable, bifunctional, scalable, mid-infrared metasurfaces with VO2 grids capable of versatile polarization manipulation and asymmetric transmission

<p>The data generated by CST Studio Suite that are used to plot a part of the figures, and sample CST scripts.&nbsp;</p> <p>Research supported by Narodowe Centrum Nauki, project no UMO-2020/39/I/ST3/02413.&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Model outputs from the study "A scalable framework for soil property mapping tested across a highly diverse tropical data-scarce region"

<p>Model outputs from the study "A scalable framework for soil property mapping tested across a highly diverse tropical data-scarce region". The study is published as open access and can be found at the following link: <a href="https://www.sciencedirect.com/science/article/pii/S2950289625000326">https://www.sciencedirect.com/science/article/pii/S2950289625000326</a></p> <p>&nbsp;</p> <p>The file "SWAT_USERSOIL.csv" was included to facilitate the assimilation of the soil mapping data into the Soil &amp; Water Assessment Tool (SWAT, https://swat.tamu.edu/) for hydrological modeling.&nbsp;</p> <p>&nbsp;</p> <p>Regarding the raster files, please note:</p> <p>a) All values in these datasets have been multiplied by 10,000 to optimize file sizes.</p> <p>b) Files are named using the variable acronym, followed by the corresponding soil layer. For outputs derived from pedotransfer functions (PTFs), the PTF reference is appended after the variable acronym.</p> <p>c) Available data decrease with increasing soil layer number. This occurs because not all locations (grid cells) have the same soil depth or number of soil layers.</p> <p>&nbsp;</p> <p>If you have any questions about the dataset or its use, please don't hesitate to contact us.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Data for Paper "Scalable Semantic 3D Mapping of Coral Reefs with Deep Learning"

<p><strong>Example Data for DeepReefMap</strong></p> <p>This dataset contains input videos in MP4 format taken with GoPro Hero 10 Cameras in Reefs in the Red Sea to demonstrate the DeepReefMap tool, which is described in the paper "Scalable Semantic 3D Mapping of Coral Reefs with Deep Learning" by Sauder et al.</p> <p>It contains a directory for model checkpoints for semantic segmentation, and for the 3D SLAM component:</p> <p>```<br>checkpoints/<br>&nbsp; &nbsp; &nbsp; &nbsp; segmentation_net.pth<br>&nbsp; &nbsp; &nbsp; &nbsp; sfm_net.pth<br>```</p> <p>It also contains videos to run the reconstruction with. See the detailed instructions for running reconstructions in https://github.com/josauder/mee-deepreefmap</p> <p>```<br>input_videos/<br>&nbsp; &nbsp; &nbsp; &nbsp; GX_SINGLE_VIDEO.MP4<br>&nbsp; &nbsp; &nbsp; &nbsp; GX_VIDEO_1_OF_2.MP4<br>&nbsp; &nbsp; &nbsp; &nbsp; GX_VIDEO_2_OF_2.MP4<br>```</p>

opencc-by-4.0Feb 2024View details →
zenodo44/100

Systematic reconstruction of molecular pathway signatures using scalable single-cell perturbation screens

<p>This repo contains Seurat objects, differential expression analysis results, and pathway gene lists for the manuscript "Systematic reconstruction of molecular pathway signatures using scalable single-cell perturbation screens"<br>List of files:</p> <p>1. Seurat_object_IFNB_Perturb_seq.rds: &nbsp; &nbsp; Seurat object of the Perturb-seq data for Interferon-beta pathway<br>2. Seurat_object_IFNG_Perturb_seq.rds: &nbsp; &nbsp;Seurat object of the Perturb-seq data for Interferon-gamma pathway<br>3. Seurat_object_TNFA_Perturb_seq.rds: &nbsp; Seurat object of the Perturb-seq data for TNF-alpha pathway<br>4. Seurat_object_TGFB1_Perturb_seq.rds: Seurat object of the Perturb-seq data for TGF-beta1 pathway<br>5. Seurat_object_INS_Perturb_seq.rds: &nbsp; &nbsp; &nbsp;Seurat object of the Perturb-seq data for insulin pathway<br>6. Pathway_genelist.rds: &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; The pathway gene lists from MultiCCA analysis<br>7. Pathway_Exclusive_genelist.rds: &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;The pathway exclusive gene lists generated from Pathway_genelist.rds<br>8. HClust_Pathway_celltype_specific_genelist.rds: &nbsp; &nbsp; The cell-line specific pathway gene lists from hierarchical clustering analysis independently done on each cell line<br>9. DE_results_all_pathway.zip: &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; The DE test results for all the regulators, cell lines, and pathways (from Mixscale weighted DE test.)<br>10. Bulk_RNAseq_Seurat_object_IFNG_and_TGFB_stim.rds: &nbsp; &nbsp; &nbsp; Seurat object for the bulk RNA-seq data for interferon-gamma and TGF-beta stimulation experiments<br>11. Parse_Guide_Capture_Protocol.pdf: &nbsp; &nbsp; &nbsp;The guide RNA capture protocol developed for Parse Evercode Whole Transcriptome kit</p>

opencc-by-4.0Jan 2024View details →
zenodo44/100

Data for Scalable parametric encoding of multiple modalities

<p>Two folders with figures created using Generative Encoding, using data from Stenbeck et al [1]:</p> <p>&nbsp;</p> <p>Folder 1 -&nbsp;epithelial_immune_transition</p> <p>This contains in-silico perturbations of epithelial and immune genes. Using generative encoders, the transformation begins with&nbsp;a histology to&nbsp;gene expression transformation, a perturbation of the gene expression vector, followed by an inverse transformation back to&nbsp;histology.</p> <p>&nbsp;</p> <p>Folder 2 -&nbsp;heatmap_he_et_al</p> <p>This contains in-silico transformation of histology tissue to gene expression. It is similar to the first folder, yet is only a single transformation (without an inverse) from histology to gene expression.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>[1] Stenbeck, Linnea; Bergenstr&aring;hle, Ludvig; Lundeberg, Joakim; Borg, &Aring;ke (2021), &ldquo;Human breast cancer in situ capturing transcriptomics&rdquo;, Mendeley Data, V5, doi: 10.17632/29ntw7sh4r.5</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Model Zoo Dataset Samples for Scalable Weight Space Learning

<p>This dataset contains small versions of model zoo datasets for our ICML 2024 paper "Towards Scalable and Versatile Weight Space Learning". These datasets are intended for testing and rapid pipeline evaluation of the code in the <a title="https://github.com/HSG-AIML/SANE" href="https://github.com/HSG-AIML/SANE">corresponding </a><a href="https://github.com/HSG-AIML/SANE">repository</a>. For full model zoos, please see&nbsp;<a href="modelzoos.cc">modelzoos.cc</a>.</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Mowgli: DBMS Performance & Scalability Evaluation Data Sets

<p>These data sets contain the performance and scalability&nbsp;&nbsp;evaluation data created by the&nbsp;<a href="https://omi-gitlab.e-technik.uni-ulm.de/mowgli/getting-started">Mowgli</a>&nbsp;framework for Apache Cassandra and Couchbase, operated on a private Openstack and the Amazon EC2 cloud.</p>

openapache2.0Oct 2019View details →
zenodo44/100

MAD (MAlicious Traffic Dataset) in home and commercial environments - Environment with scalability

<p>We have used the Internet environment: 01 Switch, 01 IP camera, 01 server for monitoring, 01&nbsp;server for honeypot and no firewall. This environment is directly connected to the Internet. We installed a server, functioning as a Monitoring Environment. The network traffic was obtained via&nbsp;Port Mirroring on the switch to the Monitoring Environment server.</p> <p>We added 08 virtual machines and performed the following test with a denial of service DoS attack:</p> <p>01 virtual machine from 04:00 pm to&nbsp;23:55 pm on 2019-12-04&nbsp;with an interval every 01 hour;<br> 02 virtual machines from 23:55 am on 2019-12-04&nbsp;to 08:50 am&nbsp;on 2019-12-05&nbsp;with an interval every 01 hour;<br> 04 virtual machines as of 08:55 am on 2019-12-05 to 05:25&nbsp;pm on 2019-12-06 with an interval every 5 minutes;<br> 08 virtual machines from 05:30 pm on 2019-12-06 to 23:59&nbsp;on 2019-12-06 with an interval every 5 minutes;<br> End of tests with shutdown of virtual machines at 23:59&nbsp;on 2019-12-06.</p> <p>The results were obtained from Suricata and Telegraf collections from the TICK stack. All evidence was performed by queries via EveBox, which received data from Suricata, Grafana or graphics with information extracted from the InfluxDB (Grafana) and PostgreSQL (EveBox) databases.</p> <p>events.csv.gz - Suricata / Evebox collections</p> <p>net.csv.gz -&nbsp;Telegraf collections from the TICK stack</p> <p>netstat.csv.gz -&nbsp;Telegraf collections from the TICK stack</p> <p>For correlation purposes, use the events.csv.gz file as a basis. The key to correlation is the &#39;timestamp&#39; column events.csv.gz with the &#39;time&#39; column in the net.csv.gz and netstat.csv.gz files.</p> <p>The interval between collections, non-consecutive, was from 2019-12-04 to 2019-12-06</p>

opencc-by-4.0Apr 2021View details →
zenodo44/100

Scalably learning quantum many-body Hamiltonians from dynamical data

<p>Our <a href="https://arxiv.org/abs/2209.14328">paper</a> on Hamiltonian learning for large quantum systems contains several numerical results. The results were produced with the <a href="https://github.com/frederikwilde/differentiable-tebd">differentiable-tebd package</a> which we developed for this study. The scripts and raw output data, as well as Jupyter notebooks for generating the plots shown in the paper are contained in this repository. For more information please refer to the <a href="https://github.com/frederikwilde/scalable-dynamical-hamiltonian-learning/">guiding repository</a>.</p> <p>For funding information please refer to the acknowledgement section of the <a href="https://arxiv.org/abs/2209.14328">paper</a>.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Dataset for "FlexTDOA: Robust and Scalable Time-Difference of Arrival Localization Using Ultra-Wideband Devices"

<p>Dataset for the paper &quot;FlexTDOA: Robust and Scalable Time-Difference of Arrival Localization Using Ultra-Wideband Devices&quot;</p> <p>The dataset contains localization measurements acquired with UWB devices. We compare the proposed localization method, called FlexTDOA, with a classic TDOA implementation, and with TWR-based localization. For more information about the localization methods, please refer to the paper.</p> <p>The dataset contains the measurements necessary to generate all the plots in the paper. For code examples on how to read and plot the data, please check out the associated Github repository: https://github.com/lauraflu/flextdoa</p> <p>If you find the dataset useful, please consider citing our work:</p> <blockquote> <p>Pătru, G. C., Flueratoru, L., Vasilescu, I., Niculescu, D., &amp; Rosner, D. (2023). FlexTDOA: Robust and Scalable Time-Difference of Arrival Localization Using Ultra-Wideband Devices. <em>IEEE Access</em>.</p> </blockquote>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Dataset for the evaluation of the scalability of a primary school Digital Education curricular reform

<p>Dataset for the evaluation of the scalability of a primary school Digital Education curricular reform<br> =======================================================</p> <p>&bull; If you publish material based on this dataset, please cite the following :</p> <p>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&bull; The Zenodo repository : Laila El-Hamamsy, Barbara Bruno, Jessica Dehler Zufferey, &amp; Francesco Mondada (2023). Dataset for the evaluation of the scalability of a primary school Digital Education curricular reform [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7912941</p> <p><br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&bull; The corresponding article : El-Hamamsy, L.*, Monnier, E.-C. *, Chessel-Lazzarotto F., Li&eacute;geois G., Bruno, B., Dehler Zufferey, J., and Mondada, F. (2023). An Adapted Cascade Model to Scale Primary School Digital Education Curricular Reforms and Teacher Professional Development Programs. arXiv. https://doi.org/10.48550/arXiv.2306.02751</p> <p><br> &bull; License: This work is licensed under a Creative Commons Attribution 4.0 International license (CC-BY-4.0)</p> <p>&bull; Creator: El-Hamamsy, L., Bruno, B., Dehler Zufferey, J., and Mondada, F.</p> <p>&bull; Date: May 9th 2023</p> <p>&bull; Subject: Educational change, Scalability, &nbsp;Professional Development, &nbsp;Digital Education, Curricular<br> Reform, &nbsp;Primary School</p> <p>&bull; Dataset format: CSV</p> <p>&bull; Dataset collection: September 2018 to September 2022</p> <p>&bull; Dataset size : &lt; 100 kB</p> <p>&bull; Dataset content : one excel file with detailed description below. Please note that the spreadsheet may contain missing values due to teachers either choosing not to respond to the questions or the questions not being presented at each of the training sessions. &nbsp;To have access to the specific survey questions please refer to the associated publication [a].</p> <p>&bull; Abbreviations :<br> &nbsp; - DE : Digital Education<br> &nbsp; - PD : Professional Development</p> <p>&bull; Funding : This work was funded by the the NCCR Robotics, a National Centre of Competence in Research, funded by the Swiss National Science Foundation (grant number 51NF40_185543)</p> <p># References</p> <p>[a] El-Hamamsy, L.*, Monnier, E.-C. *, Chessel-Lazzarotto F., Li&eacute;geois G., Bruno, B., Dehler Zufferey, J., and Mondada, F. (2023). An Adapted Cascade Model to Scale Primary School Digital Education Curricular Reforms and Teacher Professional Development Programs. arXiv. https://doi.org/10.48550/arXiv.2306.02751</p>

opencc-by-4.0May 2023View details →
zenodo44/100

A scalable, accurate, and universal analysis framework using individual-level allele frequency for large-scale genetic association studies in an admixed population

<p>Inclusion of individuals with diverse or admixed genetic ancestries is crucial to discover novel findings that may be missed by genomics analyses rooted solely in Caucasian population. Here, we present an analysis framework, SPAmix, which is scalable to a large-scale biobank data analysis including hundreds of thousands of admixed individuals and is universally applicable to various types of complex traits including binary trait, quantitative trait, time-to-event trait, longitudinal traits, etc. For each genetic variant, SPAmix uses genotype data and genetic principal components (PCs) to estimate individual-level allele frequency, which is subsequently used to calibrate p values via a retrospective analysis. A hybrid strategy including saddlepoint approximation (SPA) can greatly increase the accuracy to analyze rare genetic variants, especially if the phenotypic distribution is unbalanced or extremely unbalanced. Compared to Tractor, SPAmix does not require local ancestry information and can be straightforwardly applicable to a multi-way admixed population. Meanwhile, SPAmix can also be extended to SPAmix<sub>local</sub> in which the local ancestry can be incorporated if available. In addition, we propose SPAmix<sub>CCT</sub> to combine the p values of SPAmix and SPAmix<sub>local</sub> via Cauchy combination (CCT). SPAmix<sub>local</sub> performs close to Tractor when analyzing quantitative traits and is more accurate when analyzing binary traits with an unbalanced case-control ratio. And SPAmix<sub>CCT </sub>is an optimal unified approach for various cross-ancestry genetic architectures. Extensive simulation studies and real data analyses of 369,314 UK Biobank individuals from multiple ancestries demonstrated that SPAmix is scalable and can discover novel hits while controlling type I error rates well.</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

Data for: Scalable and Live Trace Processing with Kieker Utilizing Cloud Computing

<p>Knowledge of the internal behavior of applications often gets lost over the years. This circumstance can arise, for example, from missing documentation. Application-level monitoring, e.g., provided by Kieker, can help with the comprehension of such internal behavior. However, it can have large impact on the performance of the monitored system. High-throughput processing of traces is required by projects where millions of events per second must be processed live. In the cloud, such processing requires scaling by the number of instances.</p> <p>In this paper, we present our performance tunings conducted on the basis of the Kieker monitoring framework to support high-throughput and live analysis of application-level traces. Furthermore, we illustrate how our tuned version of Kieker can be used to provide scalable trace processing in the cloud.</p> <p>This is the dataset containing the results of our conducted benchmarks.</p>

opencc-zeroNov 2013View details →
zenodo40/100

Data for: Improving Kieker's Scalability by Employing Linked Read-Optimized and Write-Optimized NoSQL Storage

<p>We show, how polyglot persistence increase Kieker&#39;s scalability by employing separate read-optimized and write-optimized noSQL storage. For this purpose we extended Kieker to store its monitoring output in Apache Cassandra , which is a write-optimized wide-column noSQL database. For the analysis of the generated monitoring output we are using ElasticSearch, a read-optimized document store noSQL storage. We are interlinking read-optimized and write-optimized noSQL storage within our Regression Benchmarking Execution Environment (RBEE). To ensure scalability we are employing a container infrastructure. The mentioned noSQL storages and the linker are operating within one single Docker container which scales horizontally.</p> <p>For generating reference values, we instrumented a Java SE application with Kieker&#39;s file system writer and measured throughput and method&#39;s execution times. Consecutively, we instrumented the same Java SE application with our Apache Cassandra writer and measured throughput and method&#39;s execution time. Finally, we compared measurement results of Kieker&#39;s file system writer with the measurement results of our Apache Cassandra writer.</p>

opencc-by-4.0Aug 2016View details →
zenodo40/100

Generic, Scalable and Decentralized Fault Detection for Robot Swarms

<p>This raw data archive includes the data on fault detection in a simulated swarm of 20 e-puck robots. The data was used in the paper Generic, Scalable and Decentralized Fault Detection for Robot Swarms by D. Tarapore et al. (2017).</p> <p>See readme.txt for more details.</p>

opencc-by-4.0Jul 2017View details →
zenodo40/100

Scalability, dynamicity and performance evaluation results of Mantus framework

<p>Datasets used for experimental results (Figure 5): (a) Compositional weaver efficiency; (b) incremental weaving efficiency; (c) relative overhead of weaving in workflow; (d) weaver efficiency vs. aspect complexity.</p> <p>Type of data: raw and processed</p> <p>Hardware/software used: Intel Xeon E5-2650 Haswell at 2.60GHz with 64 GB of RAM; Testing input for all Mantus benchmarks: OpenStack-based ORBITS template described in paper, composed of a controller node and of 3 different group instances of compute nodes (Xen, KVM, LXC), with two virtual networks and relative network resources.</p> <p>Data format: CSV</p> <p>Source: Experiments</p> <p> </p>

opencc-by-nc-4.0Aug 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record