Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
358
datasets available to search
ShareScore release 0.9.0
Dataset results
358 results for “dataset generation”
Dataset Strategic behaviour by wind generators: an empirical investigation- IJIO
<p>Dataset</p>
FIGURE 1. Maximum Likelihood tree generated from a combined dataset using ITS and 28S in A new species of Boletinellus (Boletinellaceae, Boletales) from India
FIGURE 1. Maximum Likelihood tree generated from a combined dataset using ITS and 28S sequences. Bootstrap values (>50 %) are indicated above/below branches. The new species is indicated in bold.
Maestro Platform-Generated Dataset of Classified Bird Images
<p>The bird dataset, mentioned in the publication titled "Evaluation of Maestro, an extensible general-purpose data gathering and data classification platform" comprises two files: a zip file containing the classified files and a JSON file that includes the corresponding classification results.</p>
Maestro Platform-Generated Dataset of Classified Bird Images
<p>The bird dataset, mentioned in the publication titled "Maestro: An Extensible General-Purpose Data Gathering and Data Classification Platform," comprises two files: a zip file containing the classified files and a JSON file that includes the corresponding classification results.</p>
Maestro Platform-Generated Dataset of Classified sound files
<p>The sounds dataset, mentioned in the publication titled "Maestro: An Extensible General-Purpose Data Gathering and Data Classification Platform," comprises two files: a zip file containing the classified sound files and a JSON file that includes the corresponding classification results.</p>
Datasets for Oktoberfest: Open-source spectral library generation and rescoring pipeline based on Prosit
<p>This sdataset contains two zip files, one for the HLA and one for metaproteomics data. These contain the results from individual oktoberfest runs for the publication including the required msms.txt files and percolator outputs.</p>
Dataset for automated image-based generation of finite element models for masonry buildings
<p>This repository contains the dataset used for computing finite element models for masonry buildings via image-based approach. The method that uses this data set was presented in the paper "Automated image-based generation of finite element models for masonry buildings" by Pantoja-Rosero et., al. (2023)" https://doi.org/10.1007/s10518-023-01726-7</p>
Image datasets used in the paper "Revealing invisible cell phenotypes with conditional generative modeling"
<p>- BBBC021_selection_128 is a selection of the BBBC021 image dataset from the Broad Bioimage Benchmarck Collection from the Broad Institute</p> <p>- golgi_256_subset is a subset (one plate) of the Golgi Dataset we used (which is about 3 times larger). It was generated by the Biophenics platform in Institut Curie, Paris, France</p> <p>- translocation_256 is the translocation Dataset we used. It was generated by the Biophenics platform in Institut Curie, Paris, France</p> <p>- LRKK2_256 is the Parkinson LRKK2 mutation dataset we used. It was generated by Ksilink, Strasbourg, France</p> <p>- smala_256 is the Malaria dataset we used. It was generated by IRD, Paris, France and acquired by the Histopathology Platform at Institut Pasteur in Paris, France. </p>
Photovoltaic Generation and Load Demand Datasets with 30 seconds resolution from an Actual Prosumer in Cyprus
<p>Real-life datasets regarding the photovoltaic generation (active and reactive power) and the load demand (active and reactive power) from an actual residential prosumer (consumer with a rooftop photovoltaic system) in Cyprus. </p> <p>The residential building, with two occupants and a 200 m<sup>2</sup> approximately indoor area, is located in Nicosia, Cyprus. The building is equipped with a rooftop photovoltaic system consist of a 5 kVA Solar Edge Inverter (SE5K), integrating 18 x REC310PE72 PV panels. The PV panels are installed with 3<sup>o</sup> inclination angle (almost flat) an 190<sup>o</sup> azimuth angle (almost south direction). The building is using split-unit air-conditioners to cover the cooling needs during the summer and a heat-pump underfloor heating system to cover the heating needs during winter. </p> <p>The datasets regarding the photovoltaic generation and the load consumption is captured through the WiseWire Energy Box (local hub) and WiseWire Cloud Platform (http://wisewiresolutions.com/) with a resolution of 30 seconds. It is noted that the photovoltaic generation is taken through the inverter's Modbus interface while the load consumption is obtained through the Modbus interface of Janitza UMG 604 fast reporting smart meter.</p> <p>A total of 24 daily profiles are provided (1 daily profile each month from October 2021 until September 2023, while the exact date is indicated by the ".csv" file name considering the following format yyyy-mm-dd). </p> <p>It should be noted that these profiles has been used for the integration of the Cyprus power system digital twin. In particular, an accurarate and high resolution simulation model of the Cyprus power system has been developed and executed in a real time simulator, where field data from various sources (e.g., PMUs, smart meters, SCADA) are fed in order to replicate the operating conditions of the actual system. Therefore, the datasets provided here, are examples of time-series profiles that have been used to replicate the photovoltaic generation and load consumption of a residential building emulated within the digital twin. More information about the digital twin can be found in Deliverable D8.3 of the OneNet project (<a href="https://onenet-project.eu/wp-content/uploads/2023/12/OneNet_D8.3_V1.0.pdf">OneNet_D8.3.pdf </a>). Moreover, these datasets have been used to develop realistic pre-piloting setups for the Smart5Grid project to preliminary examine pilot use cases in a digital twin based hardware in the loop environment in Deliverable D3.4 (<a href="https://smart5grid.eu/wp-content/uploads/2023/03/Smart5Grid_WP3__D3.4_PU_Smart5Grid-platform-integration-and-HIL-testing-activities_V1.0.pdf">Smart5Grid_D3.4.pdf</a>) of the corresponding project. </p> <p> </p> <p> </p>
Dataset for "Exploring the distribution of phylogenetic networks generated under a birth-death-hybridization process"
<p>Contains all simulation scripts, simulated data, and supplemental materials</p>
Raw dataset of "Supersonic: Learning to Generate Source Code Optimisations in C/C++"
<p>The raw unfiltered dataset of "<a href="https://arxiv.org/abs/2309.14846">Supersonic: Learning to Generate Source Code Optimisations in C/C++</a>".</p>
Ascorbic acid supports ex vivo generation of plasmacytoid dendritic cells from circulating hematopoietic stem cells: RNA-seq dataset
Open the record for dataset details and reuse information.
Dataset from: Changes in cell size and shape during 50,000 generations of experimental evolution with Escherichia coli
Open the record for dataset details and reuse information.
Psocodea Phylogenomic dataset from: Phylogenomics of parasitic and non-parasitic lice (Insecta: Psocodea): combining sequence data and Exploring compositional bias solutions in Next Generation Datasets
Open the record for dataset details and reuse information.
Dataset for the paper: Generating Question Titles for Stack Overflow from Mined Code Snippets
<p>This is the dataset for our paper: Generating Question Titles for Stack Overflow from Mined Code Snippets</p> <p>All the data are extracted from the Stack Overflow data dump, please feel free to use! :)</p>
Publicly archived datasets analyzed or generated during the study.
<p> Publicly archived datasets analyzed or generated during the study.</p>
A new dataset of rain cell generated from observations of the Tropical Rainfall Measuring Mission (TRMM) precipitation radar and visible and infrared scanner and microwave imager
<p>This new dataset (M.TRMM-1B01-1B11-2A25-PMD-Rain) contains orbit-level data with 5 km spatial resolution and 0.25 km vertical resolution. It is produced by merging TRMM PR, VIRS and TMI measurements at PR pixel resolution combined with rain cell identification. The near-surface rain rate, profiles of rain rate and precipitation reflectivity factor, visible and infrared signals and microwave signals can be obtained in the dataset. The dataset provides new important data for in-depth research on the structural characteristics of rain cells and supports the study of precipitation mechanisms.</p>
Proteo Dataset (2023) for article "Proteo: A Framework for the Generation and Evaluation of Malleable MPI Applications"
<p>This dataset is the one used by the publication "Proteo: A Framework for the Generation and Evaluation of Malleable MPI Applications". The data are the data processed after the experiments to obtain the graphs and perform the analysis of the two applications shown.</p><p>On the one hand, there are data from the conjugate gradient (CG) execution and on the other hand, data from the Proteo framework, configured to emulate the CG.</p><p>The set is divided into 4 sections, each representing the results obtained for each result section of the article:</p><ul><li><strong>5.3</strong>: Similarity results between CG and Proteo configured to emulate CG. These results were obtained from the previous paper "Configurable synthetic application for studying malleability in HPC" at the PDP2023 conference. It consists of 4 datasheets:<ul><li><i>Nasp CG-VS-Proteo.xlsx</i>: Datasheet showing the total execution, iteration and computation times of Nasp (System_1) with the CG and Proteo. Used to create Figures 7 and 8.</li><li><i>JupiterThor CG-VS-Proteo.xlsx: </i>Datasheet showing the total execution, iteration and computation times of Jupiter (System_2) and Thor (System_3) with the CG and Proteus. Used to create Figures 7 and 8.</li><li><i>Comparison Comms CG Only.xlsx</i>: Datasheet showing MPI_Allgatherv communication times for different sizes on Nasp, Jupiter and Thor systems (System_1, System_2 and System_3). Used to create Table 1.</li><li><i>Comms CV_VS_Proteo.xlsx</i>: Datasheet showing MPI_Allgatherv communication times between the CG and Proteo for different sizes in Nasp, Jupiter and Thor systems (System_1, System_2 and System_3). Used to create Table 2.</li></ul></li><li><strong>5.4.1</strong>: Evaluation of the reconfiguration techniques in isolation of Proteo for synchronous methods. Used to create Figure 9. It is divided into three files:<ul><li><i>dataG.pkl</i>: Contains data on the complete runtimes in Proteo. The description can be found in the file <i>Proteo_dataG_description.txt</i></li><li><i>dataM.pkl</i>: Contains data on reconfiguration times in Proteo. The description can be found in the file <i>Proteo_dataM_description.txt</i></li><li><i>dataL.pkl</i>: Contains data on iteration times in Proteo. The description can be found in the file <i>Proteo_dataL_description.txt</i></li></ul></li><li><strong>5.4.2</strong>: Evaluation of the reconfiguration in a malleable application, CG against Proteo. Used to create Figure 10 and 11. It is divided into two directories, differentiating between CG and Proteus. In total there are the following 5 files:<ul><li><i>dataG.pkl</i>: Contains data on the complete runtimes in Proteo. The description can be found in the file <i>Proteo_dataG_description.txt</i></li><li><i>dataM.pkl</i>: Contains data on reconfiguration times in Proteo. The description can be found in the file <i>Proteo_dataM_description.txt</i></li><li><i>dataL.pkl</i>: Contains data on iteration times in Proteo. The description can be found in the file <i>Proteo_dataL_description.txt</i></li><li><i>dataCG_G.pkl</i>: Contains data on the complete execution times in the CG. The description can be found in the CG<i>_dataG_description.txt</i></li><li><i>dataCG_M.pkl</i>: Contains data on reconfiguration times in the CG. The description can be found in the CG<i>_dataM_description.txt</i></li></ul></li><li><strong>5.4.3</strong>: Evaluation of a malleable emulated application, CG against Proteo. Used to create Figure 12, 13 and 14. It is divided into two directories, differentiating between CG and Proteus. In total there are the following 5 files:<ul><li><i>dataG.pkl</i>: Contains data on the complete runtimes in Proteo. The description can be found in the file <i>Proteo_dataG_description.txt</i></li><li><i>dataM.pkl</i>: Contains data on reconfiguration times in Proteo. The description can be found in the file <i>Proteo_dataM_description.txt</i></li><li><i>dataL.pkl</i>: Contains data on iteration times in Proteo. The description can be found in the file <i>Proteo_dataL_description.txt</i></li><li><i>dataCG_G.pkl</i>: Contains data on the complete execution times in the CG. The description can be found in the CG<i>_dataG_description.txt</i></li><li><i>dataCG_M.pkl</i>: Contains data on reconfiguration times in the CG. The description can be found in the CG<i>_dataM_description.txt</i></li></ul></li></ul>
6DOF pose estimation - synthetically generated dataset using BlenderProc
<p>Accurate and robust 6DOF (Six Degrees of Freedom) pose estimation is a critical task in various fields, including computer vision, robotics, and augmented reality. This research paper presents a novel approach to enhance the accuracy and reliability of 6DOF pose estimation by introducing a robust method for generating synthetic data and leveraging the ease of multi-class training using the generated dataset. The proposed method tackles the challenge of insufficient real-world annotated data by creating a large and diverse synthetic dataset that accurately mimics real-world scenarios. The proposed method only requires a CAD model of the object and there is no limit to the number of unique data that can be generated. Furthermore, a multi-class training strategy that harnesses the synthetic dataset's diversity is proposed and presented. This approach mitigates class imbalance issues and significantly boosts accuracy across varied object classes and poses. Experimental results underscore the method's effectiveness in challenging conditions, highlighting its potential for advancing 6DOF pose estimation across diverse applications. Our approach only uses a single RGB frame and is real-time.</p>
Aesthetic Indigenous Forms: Generative AI Imagery Dataset
<p>The past, present, and future of our city communities in the United States exist within colonization's ongoing violence - a perpetual state of lived aftermath to stolen lands, white supremacy, genocide, slavery, and anthropogenic climate change. Our land and water relations remember how we treat them with sewage, chemicals, and trash, and they influence the artistic expressions and world's of Black Philadelphia writers, philosophers, artists. The <strong>Aesthetic Indigenous Forms Dataset</strong> includes 29 generative artificial intelligence (Gen AI) images, historical research, and curatorial prose that complicates the stories of climate racism, Indigeneity, and climate change that Philadelphia's public art and environmental histories tell. This data is integrated within the Post Colonial Dreams Museum, one of two distinct, yet interconnected virtual museums in <a href="https://tinyurl.com/thecreativecollabproject"><i>Relational Possibilities: A Remix of Aesthetic Forms Through Indigeneity and Blackness</i></a>. <br><br>Curated by Dana Reijerkerk, B.A., M.I.S.<br>The Creative CoLab Project: Relational Possibilities, LEADING Fellow 2023-2024. <br>This work is licensed under: <a href="https://creativecommons.org/licenses/by-nc-nd/4.0/">CC BY-NC-ND 4.0</a>. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.