Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

358

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

358 results for “dataset generation”

Learn how ShareScore rates datasets ↗
zenodo32/100

(capsicum) deepNIR: Dataset for generating synthetic NIR images

<p>This dataset contains&nbsp;<strong>capsicum</strong>&nbsp;NIR+RGB dataset used in our paper; deepNIR: Dataset for generating synthetic NIR images and improved fruit detection system using deep learning techniques.</p> <p>Please refer to&nbsp;<a href="http://tiny.one/deepNIR">http://tiny.one/deepNIR</a>&nbsp;for more detail.</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Refinement of an Attacker Model for Attack Paths Generation | Test Dataset

<p>This repository contains the test models and associated tests of the evaluation of the bachelor thesis &quot;Refinement of an Attacker Model for Attack Paths Generation&quot; / German: &quot;Verfeinerung des Angreifermodells und F&auml;higkeiten in einer Angriffspfadgenerierung&quot;. The basis for the model used is the TravelPlaner CaseStudy. The tests for measuring the performance delta are also included.</p> <p>Palladio Bench 5, Java 11 and the following GitHub repositories are used for execution:<br> Analysis: https://github.com/Patrick-Spiesberger/Palladio-Addons-ContextConfidentialityAnalysis_Mitigation<br> Meta model: https://github.com/Patrick-Spiesberger/Palladio-Addons-ContextConfidentialityMetamodell_Mitigation<br> <br> Environment installation is included in the README files. See also the previous work by Maximilian Walter:<br> Analysis: https://github.com/FluidTrust/Palladio-Addons-ContextConfidentiality-Analysis<br> Meta Model: https://github.com/FluidTrust/Palladio-Addons-ContextConfidentiality-Metamodel</p>

opencc-by-4.0Mar 2022View details →
dryad32/100

Dataset from: Changes in cell size and shape during 50,000 generations of experimental evolution with Escherichia coli

<p>Bacteria adopt a wide variety of sizes and shapes, with many species exhibiting stereotypical morphologies. How morphology changes, and over what timescales, is less clear. Previous work examining cell morphology in an experiment with Escherichia coli showed that populations evolved larger cells and, in some cases, cells that were less rod-like. That experiment has now run for over two more decades. Meanwhile, genome sequence data are available for these populations, and new computational methods enable high-throughput microscopic analyses. In this study, we measured stationary-phase cell volumes for the ancestor and 12 populations at 2,000, 10,000, and 50,000 generations, including measurements during exponential growth at the last time point. We measured the distribution of cell volumes for each sample using a Coulter counter and microscopy, the latter of which also provided data on cell shape. Our data confirm the trend toward larger cells while also revealing substantial variation in size and shape across replicate populations. Most populations first evolved wider cells but later reverted to the ancestral length-to-width ratio. All but one population evolved mutations in rod shape maintenance genes. We also observed many ghost-like cells in the only population that evolved the novel ability to grow on citrate, supporting the hypothesis that this lineage struggles with maintaining balanced growth. Lastly, we show that cell size and fitness remain correlated across 50,000 generations. Our results suggest that larger cells are beneficial in the experimental environment, while the reversion toward ancestral length-to-width ratios suggests partial compensation for the less favorable surface area-to-volume ratios of the evolved cells.</p>

opencc-zeroMar 2022View details →
zenodo32/100

Event Generation and Density Estimation with Surjective Normalizing Flows: Dataset

<p>The four-gluino and two-gluino data used in 2205.01697 .</p>

opencc-by-4.0May 2022View details →
zenodo32/100

(nirscene) deepNIR: Dataset for generating synthetic NIR images

<p>This dataset contains <strong>nirscene</strong> NIR+RGB dataset used in our paper; deepNIR: Dataset for generating synthetic NIR images and improved fruit detection system using deep learning techniques.</p> <p>Please refer to <a href="http://tiny.one/deepNIR">http://tiny.one/deepNIR</a> for more detail.</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Dataset for "Long-wavelength pulse generation via light-sail backscattering"

<p>Dataset for&nbsp;&quot;Long-wavelength pulse generation via light-sail backscattering&quot;, paper submitted to PPCF</p>

opencc-by-4.0Aug 2022View details →
zenodo32/100

Dataset for "Exploring the Verifiability of Code Generated by GitHub Copilot"

<p>Collection of Python implementations and translations to Dafny with verification attempts.</p>

opencc-by-4.0Aug 2022View details →
zenodo32/100

Part design aiming at dataset generation for training a vision based quality monitoring system of a circular economy cell

<p>Part design aiming at dataset generation</p> <p>&nbsp;</p> <p>Related to the publication: Stavropoulos, P., Papacharalampopoulos, A., Athanasopoulou, L., Kampouris, K., &amp; Lagios, P. (2022). Designing a digitalized cell for remanufacturing of automotive frames.&nbsp;<em>Procedia CIRP</em>,&nbsp;<em>109</em>, 513-519.</p> <p>&nbsp;</p> <p>Contents of the ZIP file: Different configurations as CATPart files</p> <p>&nbsp;</p> <p>It can open with the help of CATIA software</p>

opencc-by-4.0Jun 2022View details →
zenodo32/100

FIGURE 3. Dendrogram generated from PATN analyses using Czekanowski metric association measures the dataset comprising eight samples and 133 in Stolonochloa, a new Australian genus segregated from Panicum (Poaceae: Panicoideae: Paniceae: Boivinellinae) based on phenetic analysis of morphological data

FIGURE 3. Dendrogram generated from PATN analyses using Czekanowski metric association measures the dataset comprising eight samples and 133 morphological characters. Classification strategy set at flexible UPGMA agglomerative hierarchical fusion technique with Beta = -0.10.

opennotspecifiedOct 2022View details →
zenodo32/100

FIGURE 1. Dendrogram generated from PATN analyses using Gower association measures the dataset comprising 21 samples and 133 in Stolonochloa, a new Australian genus segregated from Panicum (Poaceae: Panicoideae: Paniceae: Boivinellinae) based on phenetic analysis of morphological data

FIGURE 1. Dendrogram generated from PATN analyses using Gower association measures the dataset comprising 21 samples and 133 morphological characters. Classification strategy set at flexible UPGMA agglomerative hierarchical fusion technique with Beta = -0.10.

opennotspecifiedOct 2022View details →
zenodo32/100

Dataset of AI-generated code created by various versions of GPT model

<p>This is the dataset used for the paper "<span>Human vs AI: Investigation of Security Risks in AI-generated </span><span>Code via Comparison with Human-written Code".</span></p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Dataset used for "End-to-end simulation of particle physics events with Flow Matching and generator Oversampling" , https://arxiv.org/abs/2402.13684

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →
zenodo32/100

Detailed dataset and code generation for Artificial intelligence-based modelling of compressive strength of slurry infiltrated fiber concrete

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →
zenodo32/100

Dataset for "MA-MGAN: Mixed Attention Markovian Generative Adversarial Network for Meteorological Downscaling"

<p>Dataset for "MA-MGAN: Mixed Attention Markovian Generative Adversarial Network for Meteorological Downscaling", including the training set and testing set used by the model, as well as the experimental results in the main text and supplementary Information.</p>

opencc-by-4.0May 2024View details →
zenodo32/100

In-house WGS Datasets Generated for GenRiskPro Project

<p>The dataset for the <strong>GenRiskPro </strong>Project manuscript.</p> <p>The new data generated in this project includes:</p> <ol> <li>Whole genome sequencing (WGS) for 275 project participants, sampled and sequenced from Turkiye.</li> <li>Analysis output from the analysis pipeline, which generates candidate variants lists for each individual from the in-house Turkish cohort (TR) and a Swedish cohort (SW), regarding their known pathogenic variants, potential pathogenic variants, GWAS significant and Pharmacogenetics associated variants.</li> </ol> <p>In this repository, we included all original result files used for generating reports from the TR and SW cohorts. These are output files from the GenRiskPro pipeline.</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Dataset: Data augmentation experiments with style-based quantum generative adversarial networks on trapped-ion and superconducting-qubit technologies

<p>Dataset for the following paper: <a href="https://arxiv.org/abs/2405.04401">"Data augmentation experiments with style-based quantum generative adversarial networks on trapped-ion and superconducting-qubit technologies", Julien Baglio, arXiv:2405.04401</a></p> <p>It contains:</p> <ul> <li>one folder named "data_for_all_plots" containing the raw data for the s, t, and y distributions for all the figures of the paper as well as a Jupyter notebook to generate the figures.</li> <li>one file named "variance_calculations_qGAN.txt" containing the data to calculate the errors for the KL divergences.</li> </ul>

opencc-by-4.0May 2024View details →
zenodo32/100

ATLAS+, Dataset for "An Empirical Study on Focal Methods in Deep-Learning-Based Approaches for Assertion Generation"

<p>Dataset for "An Empirical Study on Focal Methods in Deep-Learning-Based Approaches for Assertion Generation"</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Dataset's used to generate the figures in Sanchez et al. 2024 (Relative Roles...)

<p>This is a collection of the netcdfs and mat files used to generate figures in Sanchez et al. 2024 (Relative Roles...) in JGR Oceans. Note, this is just the netcdf to make the figures. The netcdfs of the original model output are found <a href="../records/8356594">here</a>. This dataset can be run in this<a href="https://github.com/bob-sanchez/Relative_Roles_JGR24"> notebook</a> to generate the figs.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

Dataset for the study - The Attractiveness of Employee Benefits in Agriculture from the Perspective of Generation Z

<p><strong>This is the dataset for the study: The Attractiveness of Employee Benefits in Agriculture from the Perspective of Generation Z</strong></p> <p>Data contains:&nbsp;<br><strong>Data from Job advertisements</strong> &nbsp;- &nbsp;content analysis of job advertisements. Benefits offered by agricultural companies to employees were identified from the job advertisements.<br><strong>Questionnaire data</strong> - In a questionnaire survey, it was determined how attractive the employee benefits are to representatives of Generation Z.</p> <p>The headers of tables are translated into English. The data are in the original (Czech) language. &nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

Dataset for Automated Unit Test Generation via Chain of Thought Prompt and Reinforcement Learning

<p>This is the replication package including three types datasets: training dataset with CoT prompts, reward dataset for training reward model, rl dataset for optimizing policy model. The training dataset includes filter_test_cot_rule_50k.csv, filter_train_cot_rule_50k.csv, and filter_valid_cot_rule_50k.csv. These three datasets includes multiple fields (i.e., src_fm, intention, plan, elaboration, gpt_test, src_fm_cot_gpt, target, src_fm_fc_ms_ff,src_fm_intention,src_fm_plan,src_fm_elaboration,idx,rule_cot,rule_cot_nlp,combine_cot,src_fm_rule_cot_nlp,src_fm_cot_nlp_gpt,gpt_cot_filter,src_fm_plan_intention). The reward dataset includes test_athena.json, train_athena.json, and valid_athena.json three files. The rl dataset includes three files: filter_test_cot_gpt_rl.csv, filter_train_cot_gpt_rl.csv, filter_valid_cot_gpt_rl.csv. These files include mulitple fields: src_fm,intention,plan,elaboration,gpt_test,src_fm_cot_gpt,target,src_fm_fc_ms_ff,src_fm_intention,src_fm_plan,src_fm_elaboration,gpt_cot_filter.</p>

opencc-by-4.0Jun 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record