Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
358
datasets available to search
ShareScore release 0.9.0
Dataset results
358 results for “dataset generation”
Datasets-Coated soot particles with tunable, well-controlled properties generated in the laboratory with a miniCAST BC and a micro smog chamber
<p>Datasets for manuscript "Coated soot particles with tunable, well-controlled properties generated in the laboratory with a miniCAST BC and a micro smog chamber" by Michaela N. Ess, Michele Berto, Alejandro Keller, Martin Gysel-Beer, Konstantina Vasilatou</p> <p>https://doi.org/10.1016/j.jaerosci.2021.105820</p> <p> </p>
Graphine: A Dataset for Graph-aware Terminology Definition Generation
<p>This is the dataset of our EMNLP 2021 paper:</p> <p>Graphine: A Dataset for Graph-aware Terminology Definition Generation. </p> <p>Please read the "readme.md" in it for the format of the dataset.</p>
Generating a Labeled Dataset to Train Machine Learning Algorithms for Lithological Classification of Drill Cuttings
<p>This dataset contains 16,700 fully labeled SEM images of rock chips isolated from 14 thin sections of drill cutting samples. These samples come from a low-permeability reservoir in western Canada.</p>
3dcap-md: Sample datasets of Agisoft Metashape for generating metadata
<p>In this repository we provide Agisoft Metashape (SfM) projects and the 3D models processed from them with their metadata using the example of a souvenir of a terracotta warrior (sample 1) and a preserved wood sample (sample 2). The metadata was generated with our script for metadata generation.<br> <br> The first example is referenced with coded targets and imported coordinate list. The second example is referenced with coded targets and scales defined in between. For each project there is metadata in the format of a *.json file and a *.ttl file. In the metadata folders there are files with all selected metadata and additionally those where only with a Uri link were exported. We used Agisoft Metashape version 1.8.3 to create this data.</p>
Guided wave representations dataset for material property estimation and generation
<p>Material property identification in composite materials is necessary for material degradation as well as non-destructive characterization. The inverse problem needs a forward simulator. Ultrasonic-guided waves are sensitive to material properties and can be used for the purpose. The stiffness matrix method and group velocity calculation routine are used as the forward solver. The solver outputs polar group velocity curves of two fundamental Lamb wave modes for different material properties and ply layup sequences. The curves can be converted into binary images (black and white) named polar representations for image processing algorithms. The datasets contain polar representations corresponding to different material properties and ply layup sequences of a transversely isotropic laminate.</p>
Datasets of "An Automatically Generated Annotated Corpus for Albanian Named Entity Recognition"
<p>This is an Albanian named entities annotation corpus generated automatically (silver-standard) from Wikipedia and WikiData. It is offered in Apache OpenNLP annotation format.</p> <p>Details of the generation approach may be found in the respective published paper: https://doi.org/10.2478/cait-2018-0009</p> <p>Attached are also the files that were used for generating the Albanian named entities gazetteer and the gazetteer itself in JSON format.</p>
Performance and limits of a shallow-water model for landslide-generated tsunamis: from laboratory experiments to simulations of flank collapses at Montagne Pelée (Martinique) - DATASETS
<p>Datasets of the 6 presented experiments with 4mm beads presented in the paper "Performance and limits of a shallow-water model for landslide-generated tsunamis: from laboratory experiments to simulations of flank collapses at Montagne Pelée (Martinique)".</p> <p>Each datasets (csv file) corresponds to the hand picked profile of either the water free surface or the granular material at 0.1 second of interval. </p> <p>The third dataset for each experiment corespond of the gauges records.</p>
Large-area periodically-poled lithium niobate wafer stacks optimized for high-energy narrowband terahertz generation - Dataset
<p>Dataset for the publication "Large-area periodically-poled lithium niobate wafer stacks optimized for high-energy narrowband terahertz generation".</p>
Dataset for Comparing Three Generations of D-Wave Quantum Annealers for Minor Embedded Combinatorial Optimization Problems
<p>Dataset for the paper titled "Comparing Three Generations of D-Wave Quantum Annealers for Minor Embedded Combinatorial Optimization Problems"</p> <p>LA-UR-23-20367</p>
Datasets generated for KA-Search: rapid and exhaustive sequence identity search of known antibodies
<p>Datasets generated for testing KA-Search in the paper "KA-Search: rapid and exhaustive sequence identity search of known antibodies".</p> <p>Includes OAS-test and the 100 randomly selected non-redundant heavy chains of therapeutics.</p>
Curated Dataset of Association Constants Between a Cyclodextrin and a Guest for Machine Learning: Raw Data and Generation Script
<p>Determining the association constant between a cyclodextrin and a guest molecule is an important task for various applications in various industrial and academical fields. However, such a task is time consuming, tedious and requires samples of both molecules. A significant number of association constants and relevant data is available from the literature. The availability of data makes the use of machine learning techniques to predict association constants possible. However, such data is mainly available from tables in articles or appendices. It is necessary to make them available in a computer friendly format and to curate them. Furthermore, the raw data need to be enriched with physicochemical information about each molecule and when such information does not allow to discriminate molecules, some additional data is needed. We present a dataset built from data gathered from the literature. The dataset contains both the original raw data from the articles and the enriched ones. We also provide the scripts used to curate and enrich the raw data.</p>
Data and code for: Generation and applications of simulated datasets to integrate social network and demographic analyses
<p class="MsoNormal"><span>Social networks are tied to population dynamics; interactions are driven by population density and demographic structure, while social relationships can be key determinants of survival and reproductive success. However, difficulties integrating models used in demography and network analysis have limited research at this interface. We introduce the R package genNetDem for simulating integrated network-demographic datasets. It can be used to create longitudinal social networks and/or capture-recapture datasets with known properties. It incorporates the ability to generate populations and their social networks, generate grouping events using these networks, simulate social network effects on individual survival, and flexibly sample these longitudinal datasets of social associations. By generating co-capture data with known statistical relationships it provides functionality for methodological research. We demonstrate its use with case studies testing how imputation and sampling design influence the success of adding network traits to conventional Cormack-Jolly-Seber (CJS) models. We show that incorporating social network effects in CJS models generates qualitatively accurate results, but with downward-biased parameter estimates when network position influences survival. Biases are greater when fewer interactions are sampled or fewer individuals are observed in each interaction. While our results indicate the potential of incorporating social effects within demographic models, they show that imputing missing network measures alone is insufficient to accurately estimate social effects on survival, pointing to the importance of incorporating network imputation approaches. genNetDem provides a flexible tool to aid these methodological advancements and help researchers test other sampling considerations in social network studies.</span></p>
Dataset for Trans-generational effects on diapause and life-history-traits of an aphid parasitoid
<p>Dataset for the following article: Transgenerational effects act on a wide range of insects’ life-history traits and can be involved in the control of developmental plasticity, such as diapause expression. Decrease in or total loss of winter diapause expression recently observed in some species could arise from inhibiting maternal effects. In this study, we explored transgenerational effects on diapause expression and traits in one commercial and one Canadian field strain of the aphid parasitoid <em>Aphidius ervi</em>. These strains were reared under short photoperiod (8:16 h LD) and low temperature (14 °C) conditions over two generations. <a href="https://www.sciencedirect.com/topics/agricultural-and-biological-sciences/diapause">Diapause</a> levels, developmental times, physiological and morphological traits were measured. Diapause levels increased after one generation in the Canadian field but not in the commercial strain. For both strains, the second generation took longer to develop than the first one. Tibia length and wing surface decreased over generations while fat content increased. A crossed-generations experiment focusing on the industrial parasitoid strain showed that offspring from mothers reared at 14 °C took longer to develop, were heavier, taller with wider wings and with more fat reserves than those from mothers reared at 20 °C (8:16 h LD). No effect of the mother rearing conditions was shown on diapause expression. Additionally to direct plasticity of the offspring, results suggest transgenerational plasticity effects on diapause expression, development time, and on the values of life-history traits. We demonstrated that populations showing low diapause levels may recover higher levels through transgenerational plasticity in response to diapause-induction cues, provided that environmental conditions are reaching the induction-thresholds specific to each population. Transgenerational plasticity is thus important to consider when evaluating how insects adapt to changing environments.</p>
Datasets used in the paper "The Face of Deception: The Impact of AI-Generated Photos on Malicious Social Bots"
<p>Datasets used in the paper "The Face of Deception: The Impact of AI-Generated Photos on Malicious Social Bots"</p> <p>We changed the datasets' titles and omitted authors' names for the blind review process. After the review, we will upload it to GitHub in an unanonymised form.</p> <p>Check README.md for details.</p>
The knowledge graphs generated from CRE and The Session datasets
<p>This dataset contains the knowledge graphs generated from CRE and the Session datasets. All the knowledge graphs are in TTL format.</p>
Molecular dynamics-generated ensemble dataset of ubiquitin; for "PROTHON: A Local Order Parameter-Based Method for Efficient Comparison of Protein Ensembles"
<p>The molecular dynamics-generated ensemble dataset (229Mb zip file) for ubiquitin, used in the manuscript "PROTHON: A Local Order Parameter-Based Method for Efficient Comparison of Protein Ensembles", submitted to the Journal of Chemical Information and Modeling (JCIM). The dataset consists of 6 .dcd files, and one .pdb file. </p>
Dataset for "Deep Dive into the Verifiability of Code Generated by GitHub Copilot"
<p>A collection of GitHub-Copilot-generated Python solutions and their translations to Dafny with verification attempts.</p>
Pento-DIARef: A Diagnostic Dataset for Learning the Incremental Algorithm for Referring Expression Generation from Examples
<p>We present a <strong>D</strong>iagnostic dataset of <strong>IA</strong> <strong>Ref</strong>erences in a <strong>Pento</strong>mino domain (Pento-DIARef) that ties extensional and intensional definitions more closely together, insofar as the latter is the generative process creating the former.</p> <p>We create a novel synthetic dataset of examples that pairs visual scenes with generated referring expressions; examine two variants of the dataset, representing two different ways to exemplify the underlying task; and evaluate an LSTM-based baseline, a transformer and a modified version with region embeddings on them.</p> <p>See https://github.com/clp-research/pento-diaref for more information.</p>
Experimental Evidence for Shear-induced Melting and Generation of Stishovite in Granite at Low (<18 GPa) Shock Pressure [Dataset]
<p>Source data for Experimental Evidence for Shear-induced Melting and Generation of Stishovite in Granite at Low (<18 GPa) Shock Pressure.</p>
Mbit/s-range alkali vapour spin noise quantum random number generators - DATASETS
<p>Two example datasets. Each dataset consists of 2 files; one, where spin noise is present in the spectrum (noise_on.npy) and one, where spin noise is absent (noise_off.npy). The former is used generate random numbers, while the latter is used to 1. find a baseline threshold \Sigma, as described in the paper, and 2. to find the spin noise spectrum. </p> <p>The spin noise spectrum can be found by subtracting the spectra of the first and second file, as it is only present in the first file. Every dataset should be filtered around the Larmor frequency prior to any bit generation. The width of the band-pass filter should be such that the entirety of the spin noise spectrum is contained.</p> <p>Bit generation is to be done by protocols described in the paper "Mbit/s-range alkali vapour spin noise quantum random number generators". All bitrates are calculated at \Sigma = 5 \sigma, where \sigma is the standard deviation of noise_off.npy for the given experiment.</p> <p>Cs12: Sample rate 50 MHz. Measured at a temperature of approximately 140 degrees C. The Larmor frequency is approximately 670 kHz. There is a large induced gradient present in the system, so the spin noise spectrum is not a Lorentzian shape, but is spread with a full-width half max of 500 kHz . Protocol 1 raw bitrates are approximately 40 kb/s, whereas protocol 2 raw bitrates are approximately 2.40 Mb/s.</p> <p>Cs14: Sample rate 100 MHz. Spin noise measurement (noise_on.npy) has 1G samples (16 bit signed). Background measurement (noise_off.npy) has 10M samples (16 bit signed). Measured at a temperature of approximately 140 degrees C. The Larmor frequency is approximately 440kHz ,T2 is approximately 2e-6 s. This is the dataset presented in the paper. Protocol 1 raw bitrates are approximately 40 kb/s, protocol 2 raw bitrates are approximately 1.97 Mb/s.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.