Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
358
datasets available to search
ShareScore release 0.9.0
Dataset results
358 results for “dataset generation”
Datasets: Urbanisation generates multiple trait syndromes for terrestrial animal taxa worldwide
<p>This repository contains the datasets used in the main article:</p> <ul> <li>There is one Excel file per taxonomic group (amphibians, bats, bees, birds, ground beetles, and reptiles).</li> <li>Each file consists of three Excel spreadsheets: "Species" = matrix of species by sites; "Sites" = ID and coordinates of sites + urban and forest land cover variables; "Traits" = matrix of species by traits.</li> <li>Each spreadsheet contains the raw data used for the analyses in the article. For more information on how to handle the data, see the "Method" section.</li> </ul>
The dataset of inline comment generation (ICG)
<p>This repository stores the dataset of inline comment generation. You can find more details, analyses, and baseline results in our paper "ICG: A Benchmark Dataset for Inline Comments Generation Task".</p> <p> </p>
Dataset of paper "New trends on photoelectrocatalysis (PEC): nanomaterials, wastewater treatment and hydrogen generation"
<p>Dataset of paper "New trends on photoelectrocatalysis (PEC): nanomaterials, wastewater treatment and hydrogen generation"</p> <ul> <li>Revision of nanomaterials investigated for different applications of PEC for water treatment and hydrogen generation with summary of information on their performance, key advantages of the material and the reference of publication.</li> </ul>
CommitChronicle dataset from the paper "From Commit Message Generation to History-Aware Commit Message Completion", ASE 2023
<pre>This is the CommitChronicle dataset for commit message generation (and/or completion), introduced in the paper "From Commit Message Generation to History-Aware Commit Message Completion", ASE 2023. For further details, see the attached README.md. <strong>Note.</strong> Also available on HuggingFace Hub: <a href="https://huggingface.co/datasets/JetBrains-Research/commit-chronicle">JetBrains-Research/commit-chronicle</a></pre>
Dataset underlying the paper Generating variability from motor primitives during infant locomotor development
<p>EMG data after basic preprocessing</p> <p>Each file is a matrix of M colomns and S*T raws with :</p> <p>M = number of muscles (from m1 to m10: right Rectus Femoris, left Rectus Femoris, right Tibialis Anterior, left Tibialis Anterior, right Biceps Femoris, left Biceps Femoris, right Soleus, left Soleus, right Gluteus Medius, left Gluteus Medius).</p> <p>S = number of steps (specific to each file)</p> <p>T = number of time points per step (always equals to 200)</p> <p> </p> <p>File names are ordered as follows:</p> <p>inputMat_S{subjectnumber}_C{conditionnumber}_a{agenumber}_200bins_10muscles</p> <p>subject number car take values between 1 and 18</p> <p>condition number can take the following values: 1 for stepping on a treadmill, 2 for stepping overground, 4 for kicking and 5 for walking </p> <p>age number can take the following values: 1 for around birth, 2 for around three months, 4 for around walking onset</p>
Time-Resolved Plasmon-Assisted Generation of Arbitrary Optical-Vortex Pulses- Dataset
<p>This folder contains raw data and information to reproduce the findings of the article titled 'Time-Resolved Plasmon-Assisted Generation of Arbitrary Optical-Vortex Pulses' . Each folder corresponds to a figure in the article.</p> <p>Raw data is provided for the calculations along with the input file and output log of the calculation. When raw data is too large, it is possible to reproduce the calculation from the provided input file. Calculations are performed with <a href="https://octopus-code.org/documentation/12/">Octopus code</a> , the related version and commit number of the code can be retrieved from output log file provided in calculation folder.</p>
GPRChinaSPEI1km: High spatial resolution and century-long SPEI datasets for China from 1901 to 2020 generated by machine learning
<p>The high spatial resolution and century-long Standardized Precipitation Evapotranspiration Index (SPEI) dataset with a spatial resolution of 0.0083 degrees (~1 km) was spatially downscaled from the global SPEI data with a 0.5 degrees spatial resolution (https://spei.csic.es/database.html) based on machine learning integrated with high spatial resolution climatic and topographic variables. The 1-km SPEI datasets are across the land areas of China from January 1901 to December 2020, including 1-month, 3-month, 6-month and 12-month SPEIs. The unit of the data is 0.01. The dataset was evaluated using the root zone soil moisture and the historical drought events, and the evaluation indicated that the high spatial resolution SPEI dataset is reliable.</p> <p>Data Information: </p> <p>GPRChinaSPEI1km: High spatial resolution and century-long SPEI datasets over China from 1901 to 2020 generated by machine learning</p> <p>Publication: </p> <p><span>He, Q., Wang, M., Liu, K., & Wang, B. (2025). High-resolution Standardized Precipitation Evapotranspiration Index (SPEI) reveals trends in drought and vegetation water availability in China. <em>Geography and Sustainability</em>, <em>6</em>(2), 100228. https://doi.org/10.1016/j.geosus.2024.08.007</span></p> <p></p> <p>----------------------------------------------------data description---------------------------------------------</p> <p>This is a gridded dataset for the Standardized Precipitation Evapotranspiration Index (SPEI) at a spatial resolution of 1 km over the main terrestrial lands of China for each month during 1901-2020, which is generated using the Gaussian process regression (GPR) based on the Global SPEI database (https://spei.csic.es/database.html) integrated with high spatial resolution climatic and topographic variables. Four timescales of SPEI were generated: 1-month (SPEI-1), 3-month (SPEI-3), 6-month (SPEI-6) and 12-month (SPEI-12). The details are as follows:</p> <p>Region: China</p> <p>Temporal Extent: January 1901 to December 2020</p> <p>Spatial resolution: 0.0083° (~1 km)</p> <p>Temporal resolution: month</p> <p>Timescales: 1-month, 3-month, 6-month and 12-month</p> <p>Data format: GeoTIFF</p> <p>Unit: unitless (0.01)</p> <p>Geographic coordinate system: WGS 1984</p> <p>---------------------------------------------------dataset filename---------------------------------------------</p> <p>The file name specifically shows the data information.</p> <p>For example,</p> <p>“SPEI_1_2020_1.tif” means “1-month SPEI of January 2020”.</p> <p>“SPEI_3_2020_1.tif” means “3-month SPEI of January 2020”.</p> <p>All the file names are formatted in “SPEI_timescale_year_month”</p> <p>timescale: 1, 3, 6 and 12 indicate 1-month, 3-month, 6-month and 12-month, respectively</p> <p>year: from 1901 to 2020</p> <p>month: from 1 to 12</p> <p>--------------------------------------------------storage information-------------------------------------------</p> <p>The high-resolution SPEI dataset is stored in TIFF format using WGS 1984 coordinate system. The data type is int16 with a scale factor of 0.01. The nodata value is -32768. The dataset requires multiplication by 0.01 during application to obtain the actual value ranges.</p> <p>The data were compressed into .rar format every 10 years for each timescale SPEI.</p>
Dataset for Posiform Planting: Generating QUBO Instances for Benchmarking
<p>Dataset for the paper titled Posiform Planting: Generating QUBO Instances for Benchmarking</p> <p>https://arxiv.org/abs/2308.05859</p> <p>LA-UR-23-29274</p>
EMHIRES dataset: wind and solar power generation
<p><strong>EMHIRES Wind</strong></p> <p>The first version of EMHIRES dataset releases four different files about the wind power generation hourly time series during 30 years (1986-2015), taking into account the existing wind fleet at the end of 2015, for each country (onshore and offshore), bidding zone and by NUTS 1 and NUTS 2 region. The time series are given as capacity factors. The installed capacity used accounted for calculating the capacity factors are summarised in the annexes of the report.</p> <p>https://setis.ec.europa.eu/emhires-dataset-part-i-wind-power-generation_en</p> <p><strong>EMHIRES Solar</strong></p> <p>EMHIRES provides RES-E generation time series for the EU-28 and neighbouring countries. The solar power time series are released at hourly granularity and at different aggregation levels: by country, power market bidding zone, and by the European Nomenclature of territorial units for statistics (NUTS) defined by EUROSTAT; in particular, by NUTS 1 and NUTS 2 level. The time series provided by bidding zones include special aggregations to reflect the power market reality where this deviates from political or territorial boundaries.</p> <p>The overall scope of EMHIRES is to allow users to assess the impact of meteorological and climate variability on the generation of solar power in Europe and not to mime the actual evolution of solar power production in the latest decades. For this reason, the hourly solar power generation time series are released for meteorological conditions of the years 1986-2015 (30 years) without considering any changes in the solar installed capacity. Thus, the installed capacity considered is fixed as the one installed at the end of 2015. For this reason, data from EMHIRES should not be compared with actual power generation data other than referring to the reference year 2015.</p> <p>https://setis.ec.europa.eu/emhires-dataset-part-ii-solar-power-generation_en</p>
A complete energy community dataset with photovoltaic generation, battery energy storage systems and electric vehicles (v1.5)
<p>This dataset represents a complete European energy community based on actual data. In this scenario, a community of 250 households was built using real energy consumption and solar generation data obtained in homes throughout Europe. In total, 200 community members were assigned solar generation, while 150 were assigned a battery storage system. From the acquired sample, new profiles were created and randomly assigned to each end-user while also receiving two electric cars with information on their capacity, state-of-charge, and usage. Furthermore, it is provided the electric vehicle chargers’ information on their location, type, and cost of operation.</p> <p> </p> <p>Version 1.5 update: <span>on the Sheet EVs, lines 29 (Capacity kW), 30 (Charge kW), and 31 (Discharge kW) were updated to the correct values.</span></p> <p> </p> <p>This work has been published in Elsevier's Data in Brief journal:<br><em> Ricardo Faia, Calvin Goncalves, Luis Gomes, Zita Vale<br> Dataset of an energy community with prosumer consumption, photovoltaic generation, battery storage, and electric vehicles<br> Data in Brief, 2023, 109218, ISSN 2352-3409<br> <a href="https://doi.org/10.1016/j.dib.2023.109218.">https://doi.org/10.1016/j.dib.2023.109218</a><br> (<a href="https://www.sciencedirect.com/science/article/pii/S2352340923003372)">https://www.sciencedirect.com/science/article/pii/S2352340923003372)</a></em></p> <p> </p> <p>We would be grateful if you could acknowledge the use of this dataset in your publications. Please use the Data in Brief publication to cite this work.</p> <p> </p> <p>Reference data used to create this dataset:</p> <ul> <li>Filtered energy profiles and renewable energy production profiles: <a href="../record/6778401">https://zenodo.org/record/6778401</a></li> </ul> <ul> <li>Battery storage systems and electric vehicles: <a href="../record/4737293">https://zenodo.org/record/4737293</a></li> </ul>
Dataset for:Doctoral students' reflection on Generative AI: a librarian outlook
<p>Data for practice paper:</p> <p><span><span>Doctoral </span></span><span><span>students’ </span></span><span><span>reflection on</span></span><span><span> Generative AI: a librarian outlook</span></span></p> <p>Contains:</p> <p>Instruction for written assignment used for analysis. </p> <p>Excel file with raw data from follow up questionnaire.</p> <p>Diagrams of answers in follow up questionnaire.</p> <p>Questionnaire form.</p> <p> </p>
Multi-contrast MRI and histology datasets used to train and validate MRH networks to generate virtual mouse brain histology
Open the record for dataset details and reuse information.
Data and code for: Generation and applications of simulated datasets to integrate social network and demographic analyses
Open the record for dataset details and reuse information.
Large‐scale genomic SNP dataset for central and southeast European Turkey oak (Quercus cerris L.) populations generated by ddRAD‐seq method
Open the record for dataset details and reuse information.
EOSC-Nordic climate 14 years dataset (generated from CESM 2.1.0) to test owncloud (WP5)
<p>This dataset has been created using CESM 2.1.0 and ran on Saga Norwegian HPC (<a href="https://www.sigma2.no/">https://www.sigma2.no/</a>).</p> <p>The dataset contains 14 years of CESM F2000climo compset at resolution f19_g17:</p> <pre><code> module use /cluster/projects/nn1000k/modulefiles module load cesm/2.1.0 create_newcase --case $HOME/cases/F2000climo-f19_g17 --res f19_g17 --compset F2000climo --mach saga --run-unsupported --project nn1000k</code></pre> <p> </p> <p>netCDF outputs have been converted to zarray format for testing access on owncloud (part of WP5 EOSC-Nordic project). No further post-processing has been applied. </p>
Dataset for "Metabolic rate and oxygen radical levels increase but radical generation rate decreases with male age in Drosophila melanogaster sperm"
<p>Metabolic rate, H<sub>2</sub>O<sub>2</sub> level, and ROS production data for <em>D. melanogaster </em>sperm and gut tissue</p>
Generative Fourier-based Auto-Encoders:Preliminary Results Dataset
<p><a href="https://urbansounddataset.weebly.com/urbansound8k.html">UrbanSound8K Dataset</a> snapshot used for the "Generative Fourier-based Auto-Encoders: Preliminary Results" paper. Only the sounds tagged "dog_bark" are present in this small dataset</p>
CGRE Framework Dataset - A Dataset automatically generated to evaluate OCR Software on Webdocuments
<p><strong>Description</strong><br> The provided dataset was generated by the <a href="https://github.com/Drizzy3D/CGRE">CGRE Framework.</a><br> It was generated as a part of a bachelor thesis and used to evaluate the Tesseract OCR Software on webdocuments.</p> <p><strong>CGRE_dataset.zip:</strong><br> <em>1. crawl.json</em><br> This file contains crawling results from the alexa.com Top 50 most used webpages in the US from the 7th June 2020.<br> The crawling was done specifically for styling information only.</p> <p><em>2. html</em><br> The generated webdocuments can be found in this directory.<br> They are based on the crawled styling information.<br> The levels of the directory are used to store the different styling attributes.<br> Every directory is named by the used value for a specific styling attribute.<br> Every word is placed in a span html element.</p> <p><em>3. dataset</em><br> The rendered webdocuments can be found in this directory as png files.<br> They were rendered using the Chromium Embedded Framework (CEF) and contain corresponding labels.<br> The labels are in the same directory with the same name as the corresponding rendered webdocument, just as txt files.<br> The labels contain "word\t(left,top,width,height)\n" lines.<br> "(left,top,width,height)" is the bounding box of a span element containing a word.<br> "word" is the word in the bounding box.</p> <p><em>4. dataset_tesseract_complete</em><br> This directory contains the Tesseract results on the dataset as txt files.<br> The structure is analogue to the dataset.<br> The txt files contain analogue to the dataset "word\t(left,top,width,height)\n" lines.</p> <p><em>5. evaluation</em><br> The results of the evaluation of Tesseract on the dataset.<br> To evaluate the localisation of words by Tesseract, the Intersection Over Union metric was used, with different threshold values (0.5, 0.6, 0.7, 0.8, 0.9).<br> To evaluate the determination of words by Tesseract, a normalized Levenshtein distance metric was used, with different threshold values (0.5, 0.6, 0.7, 0.8, 0.9).<br> The times were measured by using this system:<br> Ubuntu 20.04, AMD Ryzen 5 1600 CPU, AMD Radeon RX Vega 56 GPU, 16 GB DDR4 RAM with 2400 MHz<br> The different threshold values are stored in the filenames.<br> You can find the results in the csv files.<br> Every line contains the results for a specific webdocument.<br> The txt files contain calculated precision and recall values.</p>
Psocodea Phylogenomic dataset from: Phylogenomics of parasitic and non-parasitic lice (Insecta: Psocodea): combining sequence data and Exploring compositional bias solutions in Next Generation Datasets
<p>This dataset includes all alignments used for the phylogenomic analysis of Psocodea. In this dataset, includes all result files of phylogenomic analyses completed. This includes maximum likelihood, astral, MCMCtree, quartet sampling, and all gene trees. Any relevant input files are included, and any materials are available upon request.</p> <p>The insect order Psocodea is a diverse lineage comprising both parasitic (Phthiraptera) and non-parasitic members (Psocoptera). The extreme age and ecological diversity of the group may be associated with major genomic changes, such as base compositional biases expected to affect phylogenetic inference. Divergent morphology between parasitic and non-parasitic members has also obscured the origins of parasitism within the order. We conducted a phylogenomic analysis on the order Psocodea utilizing both transcriptome and genome sequencing to obtain a data set of 2,370 orthologous genes. All phylogenomic analyses, including both concatenated and coalescent methods suggest a single origin of parasitism within the order Psocodea, resolving conflicting results from previous studies. This phylogeny allows us to propose a stable ordinal level classification scheme that retains significant taxonomic names present in historical scientific literature and reflects the evolution of the group as a whole. A dating analysis, with internal nodes calibrated by fossil evidence, suggests an origin of parasitism that predates the K-Pg boundary. Nucleotide compositional biases are detected in third and first codon positions and result in the anomalous placement of the Amphientometae as sister to Psocomorpha when all nucleotide sites are analyzed. Likelihood-mapping and quartet sampling methods demonstrate that base compositional biases can also have an effect on quartet-based methods.</p>
Datasets for ``The effect of a dynamo-generated field on the Parker wind''
<pre>This directory contains an index.html file with links to the run directories for Runs A-C and idl plotting routines with secondary data for the other figures for the paper "The effect of a dynamo-generated field on the Parker wind" by Jakab & Brandenburg.</pre>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.