Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
358
datasets available to search
ShareScore release 0.9.0
Dataset results
358 results for “dataset generation”
Dataset for generating LOD3 building models from structure-from-motion and semantic segmentation
<p>This repository contains the codes for computing geometrical digital twins as LOD3 models for buildings, using a structure from motion and semantic segmentation. The methodology hereby implements was presented in the paper [Generating LOD3 building models from structure-from-motion and semantic segmentation" by Pantoja-Rosero et., al. (2022)] (<a href="https://doi.org/10.1016/j.autcon.2022.104430">https://doi.org/10.1016/j.autcon.2022.104430</a>)</p>
Lowest Common Ancestor Generations (LCAG) Phasespace Particle Decay Reconstruction Dataset
<p>This record contains the corresponding dataset for the paper <a href="https://dx.doi.org/10.1088/2632-2153/ac8de0"><strong>Learning Tree Structures from Leaves For Particle Decay Reconstruction</strong></a>. The dataset contains the resulting simulated particle physics decays, with information about the detected particle (leaves) to be used as input, and Lowest Common Ancestor Generations (LCAGs) to be used as training targets. The code used for the paper's experiments, which contains the PyTorch dataset/dataloader, can be found at: <a href="https://github.com/Helmholtz-AI-Energy/BaumBauen">github.com/Helmholtz-AI-Energy/BaumBauen</a>.</p> <p>The dataset contains simulated synthetic particle decays, simulated using the <a href="https://github.com/zfit/phasespace">PhaseSpace</a> library.<br> All simulated decay topologies have a common root particle of mass 100 (arbitrary units). Intermediate particles are selected at random with replacement from the following masses: [90, 80, 70, 50, 25, 20, 10].<br> Final state particles, which make up the leaf nodes of generated topologies, are drawn with replacement from the following masses: [1, 2, 3, 5, 12]. For each intermediate particle (including the root), we limit the minimum number of children to two, and the maximum five.</p> <p>Tree topology creation to generate the dataset was as follows:<br> starting from the root particle a set of children are selected from the available intermediate and final state particles such that the sum of their masses totals less than the root, this process is then repeated for each child particle which is not a final state particle and so on until only final state particles remain.</p> <p>This dataset consists of 200 topologies (unique decay processes) in total, with 16,000 samples per topology. In the paper's experiments, 2000 topologies for each of training, validation, and testing were used. Leaf node features are not normalized. We have not enforced any ordering of the nodes and leave them unsorted as created in the dataset.</p> <p>When unpacked, the dataset archive will have the following structure, with the labelling pattern <em>[data]_[subset].[topology].npy</em></p> <pre><code>└── phasespace_dataset/ ├── lcas_train.000.npy ├── leaves_train.000.npy ├── ... ├── lcas_train.199.npy ├── leaves_train.199.npy ├── lcas_val.000.npy ├── leaves_val.000.npy ├── ... ├── lcas_val.199.npy ├── leaves_val.199.npy ├── lcas_test.000.npy ├── leaves_test.000.npy ├── ... ├── lcas_test.199.npy └── leaves_test.199.npy </code></pre> <p> </p>
A Method for Generating a Quasi-Linear Convective System Suitable for Observing System Simulation Experiments: Dataset
<p>This repository provides data files to quickly run the QLCS observing system simulation experiments. The data file includes assimilated observations prepared for the data assimilation research testbed system (./data/obs/). The file also contains a restart files to initialize the nature run simulation (./data/nature_run/) and the initial prior ensemble at the time of the first data assimilation cycle (./data/initial_fcst_ensemble/).</p>
Estimation of axial loads in tie-rods: Dataset generated from Finite Element simulations for training Artificial Neural Network
<p>Dataset employed for training the Artificial Neural Networks (ANNs) presented in the cited journal article. The trained ANNs were used to estimate the tensile force in tie-rods installed in a historical structure (the church of the monastery of Sant Cugat close to Barcelona) from dynamic parameters obtained from vibration testing.</p> <p>The dataset consists of input-otput data generated using finite element (FE) simulations. A blank column has been used to separate input data from output data.</p> <p>More details on the nature of the data and how it was employed can be found in the following journal article, which is supplemented by this upload:<br> <em><strong>Makoond N, Pelà L, Molins C. Robust estimation of axial loads sustained by tie-rods in historical structures using Artificial Neural Networks. Structural Health Monitoring. 2022;0(0). doi:</strong></em><strong><a href="https://doi.org/10.1177/14759217221123326">10.1177/14759217221123326</a></strong></p> <p><a href="https://www.researchgate.net/publication/364098652_Robust_estimation_of_axial_loads_sustained_by_tie-rods_in_historical_structures_using_Artificial_Neural_Networks">Link to author's version of accepted manuscript</a></p> <p>This work was supported by the Servei del Patrimoni Arquitectònic of the Generalitat de Catalunya through a project (managed by the City Council of Sant Cugat) aimed at monitoring the church of the Monastery of Sant Cugat (grant number C-10764). Financial support is also acknowledged from the Ministry of Science, Innovation and Universities of the Spanish Government and the ERDF (European Regional Development Fund) through the SEVERUS project (Multilevel evaluation of seismic vulnerability and risk mitigation of masonry buildings in resilient historical urban centres) (grant number RTI2018-099589-B-I00).</p>
Learning to Generate Wasserstein Barycenters: datasets
<p>Datasets used in the paper "Learning to Generate Wasserstein barycenters" published in the JMIV (<a href="https://link.springer.com/article/10.1007/s10851-022-01121-y">https://link.springer.com/article/10.1007/s10851-022-01121-y</a>) and also available on arXiv (<a href="https://arxiv.org/abs/2102.12178">https://arxiv.org/abs/2102.12178</a>), with code on GitHub (<a href="https://github.com/jlacombe/learning-to-generate-wasserstein-barycenters">https://github.com/jlacombe/learning-to-generate-wasserstein-barycenters</a>).</p>
TWIGMA: A dataset of AI-Generated Images with Metadata From Twitter
<p><strong>Update May 2024: Fixed a data type issue with "id" column that prevented twitter ids from rendering correctly.</strong></p> <p>Recent progress in generative artificial intelligence (gen-AI) has enabled the generation of photo-realistic and artistically-inspiring photos at a single click, catering to millions of users online. To explore how people use gen-AI models such as DALLE and StableDiffusion, it is critical to understand the themes, contents, and variations present in the AI-generated photos. In this work, we introduce TWIGMA (TWItter Generative-ai images with MetadatA), a comprehensive dataset encompassing 800,000 gen-AI images collected from Jan 2021 to March 2023 on Twitter, with associated metadata (e.g., tweet text, creation date, number of likes).</p> <p>Through a comparative analysis of TWIGMA with natural images and human artwork, we find that gen-AI images possess distinctive characteristics and exhibit, on average, lower variability when compared to their non-gen-AI counterparts. Additionally, we find that the similarity between a gen-AI image and human images (i) is correlated with the number of likes; and (ii) can be used to identify human images that served as inspiration for the gen-AI creations. Finally, we observe a longitudinal shift in the themes of AI-generated images on Twitter, with users increasingly sharing artistically sophisticated content such as intricate human portraits, whereas their interest in simple subjects such as natural scenes and animals has decreased. Our analyses and findings underscore the significance of TWIGMA as a unique data resource for studying AI-generated images.</p> <p>Note that in accordance with the privacy and control policy of Twitter, <strong>NO raw content from Twitter is included</strong> in this dataset and users could and need to retrieve the original Twitter content used for analysis using the Twitter id. In addition, users who want to access Twitter data should consult and follow rules and regulations closely at the official Twitter developer policy at https://developer.twitter.com/en/developer-terms/policy. </p> <p> </p> <p> </p>
RU-AI: A Large Multimodal Dataset for Machine Generated Content Detection
<p>This repository contains all the collected and aligned data for RU-AI dataset. It is constructed based on three large publicly available datasets: Flickr8K, COCO, and Places205, by adding their corresponding machine-generated pairs from five different generative models in each modality. </p>
Nantes-MobileHDRVQA Dataset: Video Quality of User Generated Mobile HDR Videos
<div>Nantes-MobileHDRVQA Dataset contains 60 source videos (SRC), each compressed with AV1 codec at different bitrate and resolution pairs. More info in ReadMe file.</div> <div>A subjective experiment with Absolute Category Rating with Hidden Reference (ARC-HR) protocol was conducted to collect video quality ratings in the range of (1, 5) where higher values indicate better video quality.</div> <div>The experiments were conducted in laboratory conditions at the facilities of Nantes University, France. </div> <div>For each playlists, individual subjective opinion scores and MOS, DMOS, and 95% CI of the MOS is given for each playlists in corresponding playlists</div> <div>Moreover, playlists are combined in the plistall_ACR.csv file with their MOS, DMOS, and 95% CI of the MOS scores. </div>
Dataset: Greenidge Generation Holdings Inc. 8.50% Senior Notes due 2026 (GREEL) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Dataset: Greenidge Generation Holdings Inc. (GREE) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Dataset: Generation Income Properties, Inc. (GIPR) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Dataset: Generation Income Properties, Inc. (GIPRW) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Dataset: Generations Bancorp NY, Inc. (GBNY) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Dataset: Generation Bio Co. (GBIO) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Dataset: Generation Asia I Acquisition Limited (GAQ) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Dataset: Themes Generative Artificial Intelligence ETF (WISE) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Transfer learning with generative models for object detection on limited datasets
<p>The provided datasets are used for the analysis in the work "Transfer learning with generative models for object detection on limited datasets" (https://doi.org/10.1088/2632-2153/ad65b5). The availability of data is limited in some fields, especially for object detection tasks, where it is necessary to have correctly labeled bounding boxes around each object. A notable example of such data scarcity is found in the domain of marine biology, where it is useful to develop methods to automatically detect submarine species for environmental monitoring. To address this data limitation, the state-of-the-art machine learning strategies employ two main approaches. The first involves pretraining models on existing datasets before generalizing to the specific domain of interest. The second strategy is to create synthetic datasets specifically tailored to the target domain using methods like copy-paste techniques or ad-hoc simulators. The first strategy often faces a significant domain shift, while the second demands custom solutions crafted for the specific task. In response to these challenges, here we propose a transfer learning framework that is valid for a generic scenario. In this framework, generated images help to improve the performances of an object detector in a few-real data regime. This is achieved through a diffusion-based generative model that was pretrained on large generic datasets. With respect to the state-of-the-art, we find that it is not necessary to fine tune the generative model on the specific domain of interest. We believe that this is an important advance because it mitigates the labor-intensive task of manual labeling the images in object detection tasks. We validate our approach focusing on fishes in an underwater environment, and on the more common domain of cars in an urban setting. Our method achieves detection performance comparable to models trained on thousands of images, using only a few hundreds of input data. Our results pave the way for new generative AI-based protocols for machine learning applications in various domains, for instance ranging from geophysics to biology and medicine. The provided datasets are built with the help of Gligen and the already existing NuImages, Ozfish and Deepfish datasets. The file "CarGenerated.zip" contains images generated with Gligen and with provided bounding boxes around cars in an urban environment. The file "fishes_on_bkg.zip" provides fish images generated with fishes from Deepfish inpainted with Gligen on generated backgrounds. The file "fish_text.zip" contains images completely generated with Gligen containing fishes with annotated bounding boxes. Finally, the file "oz_masked_512.zip" contains a simpler dataset of copy paste images of Deepfish fishes on Ozfish backrounds. All the files contains the images saved in different folders for training and validation, plus an index file called gt_fish.csv for the bounding boxes.</p>
Shingle example self-consistent source dataset for global domain generation
<p>Self-consistent source dataset for the Shingle project -- an approach and software library for the generation of boundary representation from arbitrary geophysical fields and initialisation for anisotropic, unstructured meshing (see https://www.shingleproject.org for more information).</p>
Image Synthesis with a Convolutional Capsule Generative Adversarial Network- Dataset
<p><strong>Dataset of 152 two-photon images (512x512) of axons with segmentation labels. </strong></p> <p>We combined data from two published sources (Bass et al., 2017; Canty et al., 2018) to get 152 (512×512) 2D images (produced from 3D image stacks), and manually produced the corresponding labels. These images were collected using in-vivo two-photon microscopy from the mouse somatosensory cortex. To generate the 2D images, we used a max projection over the 3D stack. The labels are binary segmentation maps of the axons.</p> <p>This dataset is split into a train (132 images) and test (20 images) sets. The raw 2D images of axons are in /original folder, and the segmentation labels are in /mask folder.</p> <p><strong>Please cite the following paper when using this dataset:</strong></p> <p>Bass, C., Dai, T., Billot, B., Arulkumaran, K., Creswell, A., Clopath, C., De Paola, V., and Bharath, A. A., 2019. “Image synthesis with a convolutional capsule generative adversarial network,” <em>Medial Imaging with Deep Learning.</em></p> <p><strong>This dataset was complied from the following publications:</strong><br> Bass, C., Helkkula, P., De Paola, V., Clopath, C. and Bharath, A.A., 2017. Detection of axonal synapses in 3D two-photon images. PloS one, 12(9), p.e0183309.<br> Canty, A.J., Jackson, J.S., Huang, L., Trabalza, A., Bass, C., Little, G. and De Paola, V., 2018. Single-axon-resolution intravital imaging reveals a rapid onset form of Wallerian degeneration in the adult neocortex. <em>bioRxiv</em>, p.391425.</p>
Princeton Prosody Archive Dataset Generated from T. V. F. Brogan's Original Bibliography
<p><strong>Overview</strong></p> <p>The <a href="https://prosody.princeton.edu/">Princeton Prosody Archive</a> (PPA) is a full-text searchable database of thousands of digitized books in English published between 1570 and 1923. The Archive collects historical documents and highlights discourses about the study of language, the study of poetry, and where and how these intersect and diverge. Currently, the PPA is comprised of public domain texts held by the <a href="https://www.hathitrust.org/">HathiTrust Digital Library</a>. It is a <a href="https://cdh.princeton.edu/projects/princeton-prosody-archive/">Sponsored Project</a> of <a href="https://cdh.princeton.edu/">the Center for Digital Humanities at Princeton</a>.</p> <p><strong>Dataset</strong></p> <p>The spreadsheets in this dataset were generated using T. V. F. Brogan's 1981 bibliography <em><a href="http://oregonstate.edu/versif/resources/evrg/sepfiles.html">English Versification, 1570-1980: A Reference Guide With a Global Appendix</a></em>, which provides the foundation for the Archive's holdings. Because the PPA is an ongoing project, we are making the full Brogan dataset available to scholars here, as well as two curated versions that pose particular data problems for us: first, a list of HathiTrust-held excerpts cited in Brogan (HathiTrust does not index periodicals or journals to allow for excerpting, so we have yet to integrate these into the PPA); and second, a list of public domain items that are not held by HathiTrust with their digital and/or analog locations and hyperlinks where available (our platform does not yet support integrating non-HathiTrust material into the PPA). </p> <p><strong>Data fields</strong></p> <ul> <li><strong>ID:</strong> unique alphanumerical ID assigned by Brogan</li> <li><strong>PPA Checked:</strong> indicates whether or not the work is in the PPA</li> <li><strong>Hathi?</strong> indicates whether or not the work is in the HathiTrust Digital Library </li> <li><strong>YEAR:</strong> year of the work's publication</li> <li><strong>AUTHOR:</strong> author of the work</li> <li><strong>TITLE:</strong> title of the work</li> <li><strong>PUBLISHER: </strong>publisher of the work</li> <li><strong>EXCERPT:</strong> indicates whether or not the work is an excerpt inside of a larger work (such as an article, essay, book chapter, etc)</li> <li><strong>JOURNAL:</strong> indicates whether or not the work is contained within a journal</li> <li><strong>ENUM:</strong> indicates multivolume works</li> <li><strong>VOLUME IDs: </strong>unique alphanumerical IDs for digitized copies of the work assigned by HathiTrust </li> <li><strong>NOTES: </strong>discursive notes field for data collectors</li> </ul> <p><em>additional fields </em><em>for Not-in-HT_Brogan-List_06.24.19 only:</em></p> <ul> <li><strong>In PUL?</strong> indicates whether or not the Princeton University Library has a copy of the work</li> <li><strong>PUL LOCATION:</strong> call number for works held by the Princeton University Library or its shared collections (ReCAP)</li> <li><strong>ALTERNATE LOCATION: </strong>analog and/or digital locations of works not held by PUL</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.