Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7,505

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

7,505 results for “Generation”

Learn how ShareScore rates datasets ↗
zenodo44/100

Generated Data for the Manuscript "Nonideality-Aware Training for Accurate and Robust Low-Power Memristive Neural Networks"

<p>The file contains&nbsp;data generated and referred to in the text and the figures of the manuscript.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Data to generate the figures of: "Symmetry breaking of azimuthal waves: Slow-flow dynamics on the Bloch sphere"

<p>The folder contains the data and scripts to generate all the figures of the paper, with detailed instructions.</p> <p>No experimental data was used for this article.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

LARD: Large-scale Artificial Disfluency Generation

<p>This dataset contains 95,992 examples of utterances with 71,994 artificial inserted disfluencies using the LARD method. We use the&nbsp;<a href="https://arxiv.org/pdf/1801.04871.pdf">Schema-Guided Dialogue (SGD)&nbsp;</a>dataset as a base to construct the synthetic disfluencies. The LARD dataset contains three different types of disfluencies: repetitions, replacements, and restarts.&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

HIKARI-2021: Generating Network Intrusion Detection Dataset Based on Real and Encrypted Synthetic Attack Traffic

<p>Available datasets from the paper&nbsp;Generating Encrypted Network Traffic for Intrusion Detection Datasets.</p> <p>To produce the dataset follow the technical detail in <a href="https://github.com/andreysfc/generating-encrypted-network">github</a></p>

opencc-by-4.0May 2021View details →
zenodo44/100

GPT-2 generated form fields

<p>This dataset is a single json containing label-value form fields generated using GPT-2. This data was used to train <strong>Dessurt</strong> (<a href="https://arxiv.org/abs/2203.16618">https://arxiv.org/abs/2203.16618</a>). Details of the generation process can be found in Dessurt&#39;s Supplementary Materials and the script used to generate it is <em>gpt_forms.py</em> in&nbsp;<a href="https://github.com/herobd/dessurt">https://github.com/herobd/dessurt</a></p> <p>The data has groups of label-value pairs each with a &quot;title&quot; or topic (or null). Each label-value pair group was generated in a single GPT-2 generation and thus the pairs&nbsp;&quot;belong to the same form.&quot; The json structure is a list of tuples, where each tuple has the title or null as the first element and the list of label-value pairs of the group as the second element. Each label-value pair is another tuple with the first element being the label&nbsp;and the second being the value&nbsp;or a list of values.</p> <p>&nbsp;</p> <p>For example:</p> <p>[ [&quot;title&quot;,[ [&quot;first label&quot;, &quot;first value&quot;], [&quot;second label&quot;, [&quot;a label&quot;, &quot;another label&quot;] ] ] ], [null, [ [&quot;again label&quot;, &quot;again value&quot;] ] ] ]</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Data generated and analysed for Santos Neves, Lambert, Valente & Etienne 2022

<p>This repository contains the data and metadata for accompanying the publication of Santos Neves, P., Lambert, J. W., Valente, L., &amp; Etienne, R. S. (2022). The robustness of a simple dynamic model of island biodiversity to geological and sea-level change.</p> <p><br> All files apart from metadata contained within this repository were obtained via computation at University of Groningen Peregrine High Performance Computing Cluster (HPCC).<br> Data was generated using the pipeline implemented on the R package DAISIErobustness, which itself greatly depends on the R package DAISIE. The code for these packages is version controlled on GitHub and is freely available in open-source repositories. See the Related Identifiers section for links to relevant archived versions of both these packages.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Workplans generated during AWOPS tests

<p>The dataset is about workplans generated during the validation of&nbsp;the Automated Work Planning Services (AWOPS).The dataset includes:</p> <ul> <li>9 xml files regarding workplans automatically generated by AWOPS for renovation works&#39; activities during the following days: <ul> <li>March 7<sup>th</sup> 2022 (day 0);</li> <li>March 14<sup>th</sup> 2022 (day 2);</li> <li>March 17<sup>th</sup> 2022 (day 5);</li> <li>March 21<sup>st</sup>&nbsp;2022 (day 9);</li> <li>March 24<sup>th</sup> 2022 (day 12);</li> <li>March 28<sup>th</sup> 2022 (day 16);</li> <li>March 31<sup>st</sup>&nbsp;2022 (day 19);</li> <li>April 4<sup>th</sup> 2022 (day 23);</li> <li>April 7<sup>th</sup> 2022 (day 26).</li> </ul> </li> </ul>

opencc-by-4.0May 2022View details →
zenodo44/100

Dataset of 30 energy customers with flexibility data, and distributed generation, considering residential, small commerce, large commerce, and industrial customers

<p>The dataset has 30 customers: ten residential, ten small commerce, five large commerce, and five industrial customers. The combination of several energy customer types allows the creation of a dataset with different types of consumption profiles, generation, and flexibility, and, therefore, different values of participation in demand response events.</p> <p>The residential profiles of the considered customers use the data available in the Working Group on Intelligent Data Mining and Analysis (IDMA): https://site.ieee.org/pes-iss/data-sets/</p> <p>The values represent a week period using 15 minutes reading periods. All the values are expressed in kWh and the matrixes were created as [customer x time_period].</p> <p>&nbsp;</p> <p>We would be grateful if you could acknowledge the use of this dataset in your publications. Please use the Zenodo publication to cite this work.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

A Practical Tool-Chain for the Development of Coordination Scenarios - Graphical Modeler, DSL, Code Generators and Automaton-Based Simulator

<p>The Peer Model is a modeling tool for coordination based on blackboard-based collaboration.&nbsp;</p> <p>The tool-chain consists of a modeler, translator and simulator.</p> <p>Its goal is to help developers of distributed and concurrent coordination software better understand algorithms and identify deficiencies from the beginning.</p> <p><br> &nbsp;</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Data from microphone to measure the noise generated by the mobilefuge

<p>The two datasets uploaded are the measurement of noise generated when the mobilefuge is placed on table with a damping pad or without a damping pad.&nbsp;We found&nbsp;that with the use of the damping pad, the noise recorded in the microphone decreased by 13dB indicating the improved stable operation of the mobilefuge.</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Generators of Architectural Atmosphere Symposium

<p>This dataset is an output of the &lsquo;Generators of Architectural Atmosphere&rsquo; Symposium, an Interfaces event of the Academy of Neuroscience for Architecture (ANFA), sponsored by the EU&rsquo;s Horizon 2020 MSCA Program &mdash; RESONANCES Project, the Perkins Eastman Studio, and the 2020 Regnier Chair. The symposium was hosted in the College of Architecture, Planning and Design (APDesign), Kansas State University, Manhattan (Kansas, USA), on April 12, 2022. Speakers: Bob Condia (Kansas State University), Elisabetta Canepa (University of Genoa and Kansas State University), Kutay G&uuml;ler (Kansas State University), and Tiziana Proietti (Oklahoma University).</p> <p><br> Recent advances in science confirm many of the architect&rsquo;s expert intuitions opening new doors to the perception of space and the meaning of architectural and urban design. The symposium &lsquo;Generators of Architectural Atmosphere&rsquo; presented to an audience of students, educators, architects, and scientists a conversation about human perception of design and building, specifically speaking to the significance of atmosphere, mood, architectural proportion, and virtual reality.</p> <p><br> This dataset is made of six files:<br> no. 1 dataset summary (.pdf)<br> no. 1 symposium poster (.pdf)<br> no. 4 videos containing speakers&rsquo; presentations (.mp4).</p> <p><br> Recorded videos of each lecture are also available on the RESONANCES project website (www.resonances-project.com/harvest) and its YouTube channel (UCk32skDiT4Bz1AHnltT51Yg).</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

study of second harmonic generation in periodically poled fiber in double pass configuration

<p>This dataset includes the experimental measurements and the numerical simulations&nbsp;of the&nbsp;power of second harmonic generated inside&nbsp;a periodically poled fiber traversed in single and double pass by a fundamental signal whose wavelength is included in a certain range of values.&nbsp;This measurements are the preliminary study for situation where the PPSF can be exploited in multiple pass configuration, such as in a cavity.&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

ring fiber cavity for generation of solitons of short duration

<p>This dataset includes all the theoretical and experimental results obtained during the implementation of a fiber cavity, in ring configuration and including an active fiber, to generate solitons of duration shorter than the ones generated so far. The scope of this work has been acquiring familiarity with short solitons, which are the solitons that should be generated in nonlinear cavities of tents of cm of length (CAFR). The dataset includes also the metadata for each file uploaded.&nbsp;&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Landslide Susceptibility and 3-day Antecedent Rainfall generated by Artificial Neural Networks

<p>[ENGLISH]</p> <p>In this dataset, you can find:</p> <p>- Landslide susceptibility indexes for the Serra Geral geomorphic unit, from 0 (low susceptibility) to 1 (high susceptibility). Files starting in map_susc</p> <p>- 3-day Antecedent Rainfall Thresholds for rainfall-induced landslides&nbsp;in&nbsp;the Serra Geral geomorphic unit. Files starting in map_3day</p> <p>The figure tiles_location.png shows the locations of each tile within the states of Rio Grande do Sul and Santa Catarina, Brazil. Background map: OpenStreetMap contributors (2022)</p> <p>This dataset was produced within the research conducted for the&nbsp;PhD Thesis of Lu&iacute;sa Vieira Lucchese. The link to the Thesis will be added here when it is available. Reference:</p> <p>LUCCHESE, Lu&iacute;sa Vieira.&nbsp;Modelagem de Suscetibilidade e de Limiares de Precipita&ccedil;&atilde;o para Deslizamentos de Terra utilizando m&eacute;todos de Aprendizagem de M&aacute;quina.&nbsp;2022.&nbsp;PhD Thesis&nbsp;(Water Resources and Environmental Sanitation) &mdash; Instituto de Pesquisas Hidr&aacute;ulicas,&nbsp;Universidade Federal do Rio Grande do Sul, Porto Alegre,&nbsp;2022.</p> <p>&nbsp;</p> <p>[PORTUGU&Ecirc;S DO BRASIL]</p> <p>Neste conjunto de dados, voc&ecirc; encontra:</p> <p>- &Iacute;ndices de suscetibilidade a deslizamentos de terra para a unidade geomorfol&oacute;gica da Serra Geral, de 0 (baixa suscetibilidade) at&eacute; 1 (alta suscetibilidade). Os arquivos t&ecirc;m o prefixo&nbsp;map_susc</p> <p>- Precipita&ccedil;&atilde;o antecedente de 3 dias para a ocorr&ecirc;ncia de deslizamentos de terra na&nbsp;unidade geomorfol&oacute;gica da Serra Geral. Os arquivos t&ecirc;m o prefixo map_3day</p> <p>A figura&nbsp;tiles_location.png mostra a localiza&ccedil;&atilde;o de cada bloco dentro dos estados do Rio Grande do Sul e de Santa Catarina.&nbsp;Mapa de fundo: OpenStreetMap contributors (2022)</p> <p>Este conjunto de dados &eacute; produto da Tese de Doutorado de Lu&iacute;sa Vieira Lucchese. O link para a Tese ser&aacute; adicionado aqui, quando estiver dispon&iacute;vel. Refer&ecirc;ncia:</p> <p>LUCCHESE, Lu&iacute;sa Vieira. Modelagem de Suscetibilidade e de Limiares de Precipita&ccedil;&atilde;o para Deslizamentos de Terra utilizando m&eacute;todos de Aprendizagem de M&aacute;quina. 2022. Tese (Doutorado em Recursos H&iacute;dricos e Saneamento Ambiental) &mdash; Instituto de Pesquisas Hidr&aacute;ulicas, Universidade Federal do Rio Grande Sul, Porto Alegre, 2022.</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

Hourly generation and supply data - current mix and future scenarios

<p>Dataset on hourly generation, imports and exports of electricity in Italy for 2018, 2019 and 2020 (current mix) and two future scenarios (2030).</p> <p>Modelling materials and methods are described in the paper &quot;Life-cycle assessment of current and future electricity supply in Italy: addressing average and marginal hourly demand&quot;.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

Data for "Accelerating equilibrium spin-glass simulations using quantum annealers via generative deep learning"

<p>Datasets and material for replicating plots and results from the paper &quot;Accelerating equilibrium spin-glass simulations using quantum annealers via generative deep learning&quot; <a href="https://scipost.org/SciPostPhys.15.1.018">SciPost Phys. 15, 018 (2023)</a>.</p> <p>You will find three data&nbsp;files and a ReadMe.txt:</p> <ul> <li><strong>couplings.tar.gz&nbsp;</strong>contains the random couplings of the system&#39;s Hamiltonian&nbsp;<span class="math-tex">\(H = \sum_{\langle ij \rangle}{J_{ij} \sigma_i \sigma_j}\)</span>;</li> <li><strong>datasets.tar.gz&nbsp;</strong>contains all the datasets generated by the&nbsp;<a href="https://www.dwavesys.com/">D-Wave</a>&nbsp;quantum computer. They are already split&nbsp;into train and validation and divided for the type of model and annealing time;</li> <li><strong>data_for_fig.tar.gz&nbsp;</strong>contains files for reproducing the plots of the article, almost all of them are saved in double format, .csv and .npy or .npz.</li> </ul> <p>We encourage you to download the GitHub code linked below to open all the listed data.</p> <p>All the data are zip, so to unzip them using</p> <pre><code class="language-bash">tar -xvf datasets.tar.gz</code></pre> <p>The code for training the Neural Networks and reproducing all the results&nbsp;is open access at <a href="https://doi.org/10.5281/zenodo.7118502">zenodo.7118502</a>.</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

LLM generated Python Compiler Test Dataset

<p>This dataset is generated by integrating Large Language Models (LLMs) with AFL++ fuzzing to enhance compiler testing for CPython. It includes original Python test scripts created by LLMs such as Mistral 7B, Codellama 7B, and Gemma 7B, targeted at various compiler functionalities. These scripts were subjected to fuzzing, resulting in a rich collection of test cases that tests potential vulnerabilities. An optional minimization process with AFL-cmin refined the dataset, ensuring it focuses on test cases that significantly contribute to code coverage and bug discovery. This dataset serves as a valuable resource for improving compiler design and testing efficiency, supporting further research and development in AI-driven software testing methods.</p> <p>please see references for citations of software used in this development</p>

openmit-licenseApr 2024View details →
zenodo44/100

Data supporting "Transformer Model Generated Bacteriophage Genomes are Compositionally Distinct from Natural Sequences"

<p>Sequence and composition data supporting doi: <a href="https://doi.org/10.1101/2024.03.19.585716" target="_blank" rel="noopener">10.1101/2024.03.19.585716</a>.&nbsp;Uncompressed file size is ~5.8GB.</p> <p>Data in zip files is organized by sequence provenance (generRNA, natural, or transformer (megaDNA)). Common file types between folders include:</p> <ul> <li>Multi-record fasta file: Sequence data for all sequences of a given provenance. For generRNA sequences, these are found within the `seq` column of file "MFE_distribution_Fig4a.csv"</li> <li>Composition files: Individual sequence level compositional metrics for sliding 120 bp windows. Only structural metrics were used in this study.</li> <li>Genomad: Results from the genomad pipeline (https://portal.nersc.gov/genomad/)</li> <li>Stats: Aggregate statistics for all sequences of a given provenance.</li> </ul> <p>The natural folder also has a metadata file detailing the taxonomy for all natural sequences.<br><br>Figure datasets are the cleaned (sometimes aggregated) datasets that underly specific figures in the manuscript. The figure designations are based on the order in: https://www.biorxiv.org/content/10.1101/2024.03.19.585716v1.</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Proposal for a Next-Generation Metadata framework

<p>This figure illustrates a Next-generation Metadata framework and how it will provide the community a composable collection of metadata schemas to remove barriers to re-use and collaboration.&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Trypanosoma Epitope Dataset: Valid Epitopes and Randomly Generated Peptides with Biochemical Metrics and AI-Generated Scores

<p>This dataset contains information about valid linear B-Cell epitopes from the Trypanosoma genus, as well as randomly generated peptides. It includes biochemical metrics generated by the EpiBuilder-1.0 tool and scores generated by the BepiPred-3.0 software. The data was originally collected from the IEDB and UniProtKB platforms and has been processed and enhanced with these informations for researchers interested in understanding the molecular interactions between Trypanosoma protozoans and the immune system.</p>

opencc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record