Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

358

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

358 results for “dataset generation”

Learn how ShareScore rates datasets ↗
zenodo36/100

Datasets-Coated soot particles with tunable, well-controlled properties generated in the laboratory with a miniCAST BC and a micro smog chamber

<p>Datasets for manuscript &quot;Coated soot particles with tunable, well-controlled properties generated in the laboratory with a miniCAST BC and a micro smog chamber&quot; by Michaela N. Ess, Michele Berto, Alejandro Keller, Martin Gysel-Beer, Konstantina Vasilatou</p> <p>https://doi.org/10.1016/j.jaerosci.2021.105820</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Graphine: A Dataset for Graph-aware Terminology Definition Generation

<p>This&nbsp;is&nbsp;the&nbsp;dataset&nbsp;of&nbsp;our&nbsp;EMNLP&nbsp;2021&nbsp;paper:</p> <p>Graphine:&nbsp;A&nbsp;Dataset&nbsp;for&nbsp;Graph-aware&nbsp;Terminology&nbsp;Definition&nbsp;Generation.&nbsp;</p> <p>Please read the &quot;readme.md&quot; in it for the format of the dataset.</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Generating a Labeled Dataset to Train Machine Learning Algorithms for Lithological Classification of Drill Cuttings

<p>This dataset contains 16,700 fully labeled&nbsp;SEM&nbsp;images of rock chips isolated from&nbsp;14 thin sections of drill cutting samples.&nbsp;These samples come from a low-permeability reservoir in western Canada.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

3dcap-md: Sample datasets of Agisoft Metashape for generating metadata

<p>In this repository we provide Agisoft Metashape (SfM) projects and the 3D models processed from them with their metadata using the example of&nbsp; a souvenir of a terracotta warrior (sample 1) and a&nbsp; preserved wood sample (sample 2). The metadata was generated with our script for metadata generation.<br> &nbsp;<br> The first example is referenced with coded targets and imported coordinate list. The second example is referenced with coded targets and scales defined in between. For each project there is metadata in the format of a *.json file and a *.ttl file. In the metadata folders there are files with all selected metadata and additionally those where only with a Uri link were exported. We used Agisoft Metashape version 1.8.3 to create this data.</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Guided wave representations dataset for material property estimation and generation

<p>Material property identification in composite materials is necessary for material degradation as well as non-destructive characterization. The inverse problem needs a forward simulator. Ultrasonic-guided waves are sensitive to material properties and can be used for the purpose. The stiffness matrix method and group velocity calculation routine are used as the forward solver. The solver outputs polar group velocity curves of two fundamental Lamb wave modes for different material properties and ply layup sequences. The curves can be converted into binary images (black and white) named polar representations for image processing algorithms. The datasets contain polar representations corresponding to different material properties and ply layup sequences of a transversely isotropic laminate.</p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Datasets of "An Automatically Generated Annotated Corpus for Albanian Named Entity Recognition"

<p>This is an Albanian named entities annotation corpus generated automatically (silver-standard)&nbsp;from Wikipedia and WikiData. It is offered in Apache OpenNLP annotation format.</p> <p>Details of the generation approach may be found in the respective published paper:&nbsp;https://doi.org/10.2478/cait-2018-0009</p> <p>Attached are also the files that were used for generating the Albanian named entities gazetteer and the gazetteer itself in JSON format.</p>

opencc-by-4.0Feb 2018View details →
zenodo36/100

Performance and limits of a shallow-water model for landslide-generated tsunamis: from laboratory experiments to simulations of flank collapses at Montagne Pelée (Martinique) - DATASETS

<p>Datasets of the 6 presented experiments with 4mm beads presented in the paper &quot;Performance and limits of a shallow-water model for landslide-generated tsunamis: from laboratory experiments to simulations of flank collapses at Montagne Pel&eacute;e (Martinique)&quot;.</p> <p>Each datasets (csv file) corresponds to the hand picked profile of either the water free surface or the granular material at 0.1 second of interval.&nbsp;</p> <p>The third&nbsp;dataset for each experiment corespond of the gauges records.</p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Large-area periodically-poled lithium niobate wafer stacks optimized for high-energy narrowband terahertz generation - Dataset

<p>Dataset for the publication &quot;Large-area periodically-poled lithium niobate wafer stacks optimized for high-energy narrowband terahertz generation&quot;.</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Dataset for Comparing Three Generations of D-Wave Quantum Annealers for Minor Embedded Combinatorial Optimization Problems

<p>Dataset for the paper titled &quot;Comparing Three Generations of D-Wave Quantum Annealers for Minor Embedded Combinatorial Optimization Problems&quot;</p> <p>LA-UR-23-20367</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Datasets generated for KA-Search: rapid and exhaustive sequence identity search of known antibodies

<p>Datasets generated for testing KA-Search in the paper &quot;KA-Search: rapid and exhaustive sequence identity search of known antibodies&quot;.</p> <p>Includes OAS-test and the 100 randomly selected non-redundant heavy chains of therapeutics.</p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

Curated Dataset of Association Constants Between a Cyclodextrin and a Guest for Machine Learning: Raw Data and Generation Script

<p>Determining the association constant between a cyclodextrin and a guest molecule is an important task for various applications in various industrial and academical fields. However, such a task is time consuming, tedious and requires samples of both molecules. A significant number of association constants and relevant data is available from the literature. The availability of data makes the use of machine learning techniques to predict association constants possible. However, such data is mainly available from tables in articles or appendices. It is necessary to make them available in a computer friendly format and to curate them. Furthermore, the raw data need to be enriched with physicochemical information about each molecule and when such information does not allow to discriminate molecules, some additional data is needed. We present a dataset built from data gathered from the literature. The dataset contains both the original raw data from the articles and the enriched ones. We also provide the scripts used to curate and enrich the raw data.</p>

openbsd-3-clauseJan 2023View details →
dryad36/100

Data and code for: Generation and applications of simulated datasets to integrate social network and demographic analyses

<p class="MsoNormal"><span>Social networks are tied to population dynamics; interactions are driven by population density and demographic structure, while social relationships can be key determinants of survival and reproductive success. However, difficulties integrating models used in demography and network analysis have limited research at this interface. We introduce the R package genNetDem for simulating integrated network-demographic datasets. It can be used to create longitudinal social networks and/or capture-recapture datasets with known properties. It incorporates the ability to generate populations and their social networks, generate grouping events using these networks, simulate social network effects on individual survival, and flexibly sample these longitudinal datasets of social associations. By generating co-capture data with known statistical relationships it provides functionality for methodological research. We demonstrate its use with case studies testing how imputation and sampling design influence the success of adding network traits to conventional Cormack-Jolly-Seber (CJS) models. We show that incorporating social network effects in CJS models generates qualitatively accurate results, but with downward-biased parameter estimates when network position influences survival. Biases are greater when fewer interactions are sampled or fewer individuals are observed in each interaction. While our results indicate the potential of incorporating social effects within demographic models, they show that imputing missing network measures alone is insufficient to accurately estimate social effects on survival, pointing to the importance of incorporating network imputation approaches. genNetDem provides a flexible tool to aid these methodological advancements and help researchers test other sampling considerations in social network studies.</span></p>

opencc-zeroMar 2023View details →
zenodo36/100

Dataset for Trans-generational effects on diapause and life-history-traits of an aphid parasitoid

<p>Dataset for the following article:&nbsp;Transgenerational effects act on a wide range of insects&rsquo; life-history traits and can be involved in the control of developmental plasticity, such as diapause expression. Decrease in or total loss of winter diapause expression recently observed in some species could arise from inhibiting maternal effects. In this study, we explored transgenerational effects on diapause expression and traits in one commercial and one Canadian field strain of the aphid parasitoid&nbsp;<em>Aphidius ervi</em>. These strains were reared under short photoperiod (8:16&nbsp;h LD) and low temperature (14&nbsp;&deg;C) conditions over two generations.&nbsp;<a href="https://www.sciencedirect.com/topics/agricultural-and-biological-sciences/diapause">Diapause</a>&nbsp;levels, developmental times, physiological and morphological traits were measured. Diapause levels increased after one generation in the Canadian field but not in the commercial strain. For both strains, the second generation took longer to develop than the first one. Tibia length and wing surface decreased over generations while fat content increased. A crossed-generations experiment focusing on the industrial parasitoid strain showed that offspring from mothers reared at 14&nbsp;&deg;C took longer to develop, were heavier, taller with wider wings and with more fat reserves than those from mothers reared at 20&nbsp;&deg;C (8:16&nbsp;h LD). No effect of the mother rearing conditions was shown on diapause expression. Additionally to direct plasticity of the offspring, results suggest transgenerational plasticity effects on diapause expression, development time, and on the values of life-history traits. We demonstrated that populations showing low diapause levels may recover higher levels through transgenerational plasticity in response to diapause-induction cues, provided that environmental conditions are reaching the induction-thresholds specific to each population. Transgenerational plasticity is thus important to consider when evaluating how insects adapt to changing environments.</p>

opencc-by-4.0Jan 2020View details →
zenodo36/100

Datasets used in the paper "The Face of Deception: The Impact of AI-Generated Photos on Malicious Social Bots"

<p>Datasets used in the paper &quot;The Face of Deception: The Impact of AI-Generated Photos on Malicious Social Bots&quot;</p> <p>We changed the datasets&#39; titles and omitted authors&#39; names for the blind review process. After the review, we will upload it to GitHub in an unanonymised form.</p> <p>Check README.md for details.</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

The knowledge graphs generated from CRE and The Session datasets

<p>This dataset&nbsp;contains the knowledge graphs generated from CRE and the Session datasets. All the knowledge graphs are in TTL format.</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Molecular dynamics-generated ensemble dataset of ubiquitin; for "PROTHON: A Local Order Parameter-Based Method for Efficient Comparison of Protein Ensembles"

<p>The molecular dynamics-generated ensemble dataset (229Mb zip file) for ubiquitin, used in the manuscript &quot;PROTHON: A Local Order Parameter-Based Method for Efficient Comparison of Protein Ensembles&quot;, submitted to the Journal of Chemical Information and Modeling (JCIM). The dataset consists of 6 .dcd files, and one .pdb file.&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Dataset for "Deep Dive into the Verifiability of Code Generated by GitHub Copilot"

<p>A collection of GitHub-Copilot-generated Python solutions&nbsp;and their&nbsp;translations to Dafny with verification attempts.</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Pento-DIARef: A Diagnostic Dataset for Learning the Incremental Algorithm for Referring Expression Generation from Examples

<p>We present a&nbsp;<strong>D</strong>iagnostic dataset of&nbsp;<strong>IA</strong>&nbsp;<strong>Ref</strong>erences in a&nbsp;<strong>Pento</strong>mino domain (Pento-DIARef) that ties extensional and intensional definitions more closely together, insofar as the latter is the generative process creating the former.</p> <p>We create a novel synthetic dataset of examples that pairs visual scenes with generated referring expressions; examine two variants of the dataset, representing two different ways to exemplify the underlying task; and evaluate an LSTM-based baseline, a transformer and a modified version with region embeddings on them.</p> <p>See&nbsp;https://github.com/clp-research/pento-diaref for more information.</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Experimental Evidence for Shear-induced Melting and Generation of Stishovite in Granite at Low (<18 GPa) Shock Pressure [Dataset]

<p>Source data for Experimental Evidence for Shear-induced Melting and Generation&nbsp; of Stishovite in Granite at Low (&lt;18 GPa) Shock Pressure.</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Mbit/s-range alkali vapour spin noise quantum random number generators - DATASETS

<p>Two example datasets. Each dataset consists of 2 files;&nbsp;one, where spin noise is present in the spectrum (noise_on.npy) and one, where spin noise is absent (noise_off.npy). The former is used generate random numbers, while the latter is used to 1. find a baseline threshold \Sigma, as described in the paper, and 2. to find the spin noise spectrum.&nbsp;</p> <p>The spin noise spectrum can be found by subtracting the spectra of the first and second file, as it is only present in the first file. Every dataset should be filtered around the Larmor frequency prior to any bit generation. The width of the band-pass filter should be such that the entirety of the spin noise spectrum is contained.</p> <p>Bit generation is to be done by protocols described in the paper &quot;Mbit/s-range alkali vapour spin noise quantum random number generators&quot;. All bitrates are calculated at \Sigma = 5 \sigma, where \sigma is the standard deviation of noise_off.npy for the given experiment.</p> <p>Cs12: Sample rate 50 MHz.&nbsp;&nbsp;Measured at&nbsp;a temperature of approximately&nbsp;140 degrees C.&nbsp; The Larmor frequency is approximately 670 kHz. There is a large induced gradient present in the system, so the spin noise spectrum is not a Lorentzian shape, but is spread with a full-width half max of 500 kHz . Protocol 1 raw bitrates are approximately 40 kb/s, whereas protocol 2 raw bitrates are approximately 2.40 Mb/s.</p> <p>Cs14: Sample rate 100 MHz. Spin noise measurement (noise_on.npy) has 1G samples (16 bit signed). Background measurement (noise_off.npy) has 10M samples (16 bit signed). Measured at&nbsp;a temperature of approximately 140 degrees C. The Larmor frequency is approximately 440kHz &nbsp;,T2&nbsp;is approximately 2e-6 s. This is the dataset presented in the paper. Protocol 1 raw bitrates are approximately 40 kb/s, protocol 2 raw bitrates are approximately 1.97 Mb/s.</p>

opencc-by-4.0May 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record