Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
358
datasets available to search
ShareScore release 0.9.0
Dataset results
358 results for “dataset generation”
Dataset: Design and Self-Assembly of Second-Generation Dendrimer-Like Block Copolymers
<p>This dataset contains processed data (data) and plotting scripts (plots) of the simulation related to the paper: </p> <p><br>F. Hartmann, R. Dockhorn, S. Pusse, B.-J. Niebuur, M. Koch, T. Kraus, A. Schießer, B. N. Balzer, and M. Gallei, <br>"Design and Self-Assembly of Second-Generation Dendrimer-Like Block Copolymers"<br>Macromolecules <strong>2024</strong>; DOI: <a href="https://doi.org/10.1021/acs.macromol.4c00944">10.1021/acs.macromol.4c00944</a></p> <p>Please consult the ReadMe.md in the zip.</p>
Dataset for Generation of multiple user-defined dispersive waves in a silicon nitride waveguide
<p>This is the data set for paper: Generation of multiple user-defined dispersive waves in a silicon nitride waveguide published in Optica. DOI: https://doi.org/10.1364/OPTICA.521625</p>
TwitterGAN dataset: Fake Twitter accounts using GAN-generated faces as profiles
<p>We release the TwitterGAN dataset, which contains 1,420 fake Twitter accounts using GAN-generated faces as their profiles.</p> <p>See https://github.com/osome-iu/fake_gan_accounts for details.</p>
Dataset for generating TL;DR
<p>This is the dataset for the TL;DR challenge containing posts from the Reddit corpus, suitable for abstractive summarization using deep learning. The format is a json file where each line is a JSON object representing a post. The schema of each post is shown below:</p> <ul> <li>author: string (nullable = true)</li> <li>body: string (nullable = true)</li> <li>normalizedBody: string (nullable = true)</li> <li>content: string (nullable = true)</li> <li>content_len: long (nullable = true)</li> <li>summary: string (nullable = true)</li> <li>summary_len: long (nullable = true)</li> <li>id: string (nullable = true)</li> <li>subreddit: string (nullable = true)</li> <li>subreddit_id: string (nullable = true)</li> <li>title: string (nullable = true)</li> </ul> <p>Specifically, the <strong>content</strong> and <strong>summary</strong> fields can be directly used as inputs to a deep learning model (e.g. Sequence to Sequence model ). The dataset consists of 3,084,410 posts with an average length of 211 words for content, and 25 words for the summary.</p> <p><strong>Note : </strong>As this is the complete dataset for the challenge, it is up to the participants to split it into training and validation sets accordingly.</p>
Original NGS dataset from publication "Next-generation sequencing analysis of a cluster of hepatitis C virus infections in a haematology and oncology center".
<p>Original hepatitis C virus hypervariable region 1 NGS sequences in fastq format from patients analyzed in the study "Next-generation sequencing analysis of a cluster of hepatitis C virus infections in a haematology and oncology center". </p> <p> </p>
Artificially-generated Lecture Video Fragmentation Dataset and Ground Truth
<p>We provide a large-scale lecture video dataset consisting of artificially-generated lectures, and the corresponding ground-truth fragmentation, for the purpose of evaluating lecture video fragmentation techniques.</p> <p>For creating this dataset, 1498 speech transcript files (generated automatically by ASR software) were used from the world's biggest academic online video repository, the VideoLectures.NET. These transcripts correspond to lectures from various fields of science, such as Computer science, Mathematics, Medicine, Politics etc. In order to create the synthetic video lectures, all transcripts were randomly split in fragments, the duration of which ranges between 4 and 8 minutes. Each synthetic lecture was then assembled by combining (stitching) exactly 20 randomly selected fragments. 300 such artificially-generated lectures are included in the released dataset. Each such lecture file has a mean duration of about 120 minutes, thus the dataset contains altogether about 600 hours of artificially-generated lectures. Every pair of consecutive fragments in these lectures originally comes from different videos, consequently the point in time where such two fragments are joined is a known ground-truth fragment boundary. All these boundaries form the dataset's ground truth. We should stress that we do not generate the corresponding video files for the artificially-generated lectures (only the transcripts), and one should not try to reverse-engineer the dataset creation process so as to use in some way the visual modality for detecting the fragments in this dataset.</p> <p><strong>File format</strong></p> <p>After you download the provided .zip and unpack it, the extracted folder will contain two sub-folders:</p> <pre><code>1. ALV_srt 2. ALV_srt_GT </code></pre> <p>Each of them contains 300 files.</p> <p>The <strong>ALV_srt</strong> folder contains the transcripts of every artificially-generated lecture, in the standard SRT format:</p> <pre><code>1. A numeric counter identifying each sequential subtitle 2. The time that the subtitle should appear on the screen, followed by --> and the time it should disappear 3. Subtitle's text itself on one or more lines 4. A blank line containing no text </code></pre> <p>The <strong>ALV_srt_GT</strong> folder contains the ground truth (GT) fragments corresponding to the lectures (transcripts) of the <strong>ALV_srt</strong> folder. Each GT file consists of 3 tab-separated columns and 20 rows, in the following format:</p> <pre><code><Fragment_ID_1> <StartTime_1> <EndTime_1> <Fragment_ID_2> <StartTime_2> <EndTime_2> <Fragment_ID_3> <StartTime_3> <EndTime_3> . . . <Fragment_ID_20> <StartTime_20> <EndTime_20> </code></pre> <p>Each row indicates a fragment. The first column indicates the ID of a fragment while the second and the third column indicate the start and the end time of the fragment respectively.</p> <p><strong>License and Citation</strong></p> <p>This dataset is provided for academic, non-commercial use only. If you find this dataset useful in your work, please cite the following publication where the dataset is introduced:</p> <p><em>D. Galanopoulos, V. Mezaris, “Temporal Lecture Video Fragmentation using Word Embeddings”, Proc. 25th Int. Conf. on Multimedia Modeling (MMM2019), Thessaloniki, Greece, Jan. 2019.</em></p> <p><strong>Acknowledgements</strong></p> <p>This work was supported by the EU’s Horizon 2020 research and innovation programme under grant agreement No 693092 MOVING. We are grateful to JSI/VideoLectures.NET for providing the lectures’ transcripts.</p>
Dataset for Orderly Generation of Butson Hadamard Matrices
<p>These files contain Butson Hadamard matrices classified in the article</p> <p> "Orderly generation of Butson Hadamard matrices"</p> <p>by Pekka H.J. Lampio, Patric R.J. Östergård, and Ferenc Szöllősi.</p> <p>This dataset includes all matrices up to monomial equivalence for which full<br> classification is given in Table 2 of the article.</p>
Syntactical Carving of PNGs and Automated Generation of Reproducible Datasets
<p>This is the dataset we used in the evaluation of our paper “Syntactical Carving of PNGs and Automated Generation of Reproducible Datasets”.</p>
PianoMotion10M: Dataset and Benchmark for Hand Motion Generation in Piano Performance
<p>Recently, artificial intelligence techniques for education have been received increasing attentions, while it still remains an open problem to design the effective music instrument instructing systems. Although key presses can be directly derived from sheet music, the transitional movements among key presses require more extensive guidance in piano performance. In this work, we construct a piano-hand motion generation benchmark to guide hand movements and fingerings for piano playing. To this end, we collect an annotated dataset, PianoMotion10M, consisting of 116 hours of piano playing videos from a bird's-eye view with 10 million annotated hand poses. We also introduce a powerful baseline model that generates hand motions from piano audios through a position predictor and a position-guided gesture generator. Furthermore, a series of evaluation metrics are designed to assess the performance of the baseline model, including motion similarity, smoothness, positional accuracy of left and right hands, and overall fidelity of movement distribution. Despite that piano key presses with respect to music scores or audios are already accessible, PianoMotion10M aims to provide guidance on piano fingering for instruction purposes.</p>
Dataset of "3D generative adversarial networks for turbulent flow estimation from wall measurements"
<p>Dataset of the article 'Three-dimensional generative adversarial networks for turbulent flow estimation from wall measurements' (https://doi.org/10.1017/jfm.2024.432). The codes processing data here are on https://github.com/erc-nextflow/3D-GAN.</p> <p>This project has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme (grant agreement no. 949085, NEXTFLOW). Views and opinions expressed are, however, those of the authors only, and do not necessarily reflect those of the European Union or the ERC. Neither the European Union nor the granting authority can be held responsible for them. A.C.M. acknowledges financial support from the Spanish Ministry of Universities under the Formación de Profesorado Universitario (FPU) programme 2020. R.V. acknowledges financial support from ERC (grant agreement no. 2021-CoG-101043998, DEEPCONTROL).</p>
The Impact of Generative AI on Student Learning Outcomes: A Statistical Analytical Approach - Dataset
Open the record for dataset details and reuse information.
Dataset for "On the Extrapolation of Generative Adversarial Networks for downscaling precipitation extremes in warmer climates"
<h1>Code and Dataset for "On the Extrapolation of Generative Adversarial Networks for downscaling precipitation extremes in warmer climates"</h1> <p>This dataset accompanies the research paper titled <strong>"On the Extrapolation of Generative Adversarial Networks for downscaling precipitation extremes in warmer climates"</strong>, currently under review for the AGU Journal GRL. The study introduces a novel Regional Climate Model (RCM) emulator focusing on high-resolution climate downscaling for the New Zealand region. For additional insights and access to the codebase utilized in this research, please refer to our <a href="https://github.com/nram812/On-the-Extrapolation-of-Generative-Adversarial-Networks-for-downscaling-precipitation-extremes">Github Repository</a>.</p> <p>The code can also be found as a ".zip" file: *On-the-Extrapolation-of-Generative-Adversarial-Networks-for-downscaling-precipitation-extremes-main. </p> <h2>Aims</h2> <p>Our study focuses on two important gaps in the literature regarding the extrapolation of empirical downscaling algorithms. First, we examine how well relationships learned from a historical period extrapolate to future unobserved climates. We compare two widely used algorithms, a GAN and a deterministic CNN baseline, that use a similar architecture (i.e. convolutional layers) trained in a model-as-truth framework to downscale daily precipitation over New Zealand. We evaluate their accuracy in capturing climate change signals in mean and extreme precipitation. Second, we explore whether training on future vs. only historical periods combined with different-sized training datasets can improve extrapolation skill. </p> <h2>Geographic Focus</h2> <p>Our research focuses only on the New Zealand Region (165°E-184°W, 33°S-51°S).</p> <p> </p> <h2>Data Overview</h2> <h3>Training and Evaluation Data</h3> <p>The training data used in this study (for our RCM emulator) spans the historical period and future period (SSP370) of simulation. It comprises daily accumulated precipitation as the primary target variable, alongside large-scale predictor variables. </p> <ul> <li> <p><strong>Resolution:</strong> The target variable is presented at a 12km resolution, reflecting the highest resolution face of RCM for the New Zealand region. Predictor variables are coarsened to a 1.5-degree resolution from original CCAM outputs using conservative interpolation. </p> </li> <li> <p><strong>Period Coverage:</strong></p> <ul> <li>Training Data: 1960-2100 (Depending on Experiment, see Table 1 for list of experiment configurations)</li> <li>Validation Data: 1985-2014 + 2070-2099 (to compute the climate change signal)</li> </ul> </li> <li> <p><strong>Models:</strong></p> <ul> <li>Training on: ACCESS-CM2</li> <li>Validated on: EC-Earth3, NorESM2-MM, CNRM-CM6-1, AWI-MR-1 </li> </ul> </li> </ul> <h3>File Structure</h3> <ul> <li> <p><strong>Training Data:</strong></p> <ul> <li>Target/Ground Truth (Y): <code>target_ACCESS-CM2_hist_ssp370_pr.nc</code></li> <li>Predictor (X): <code>predictor_ACCESS-CM2_hist_ssp370.nc</code></li> </ul> </li> <li> <p><strong>Evaluation Data:<br></strong>All other GCMs can be accessed in one single file, predictor and target variables have the dimensions (time, lat, lon, GCM).</p> <ul> <li>Target/Ground Truth (Y): <code>Other_GCMs_hist_SSP370_target_fields_pr.nc</code></li> <li>Predictor (X): <code>Other_GCMs_hist_SSP370_predictor_fields.nc</code></li> </ul> </li> </ul> <h2>Methodological Insights</h2> <ul> <li> <p><strong>Regional Climate Model</strong>, Our Regional Climate Model training data is from the Conformal Cubic Atmospheric Model (CCAM) which is a global non-hydrostatic atmospheric model renowned for its variable-resolution cubic grid. . For more information about CCAM, please see the following <a href="https://agupubs.onlinelibrary.wiley.com/doi/abs/10.1029/2023JD038530">paper</a>.</p> </li> <li> <p><strong>Predictor and Target Variables:</strong> Daily-averaged large-scale prognostic variables, including zonal wind, meridional wind, temperature, and specific humidity, are employed as predictors at the 500mb and 850mb pressure levels. These are normalized (see the GitHub repository for the mean and standard deviation fields). Precipitation is taken as is from CCAM and accumulated for each given day. Static predictors are also used in our model, which is stored in a GitHub repository.</p> </li> <li> <p><strong>Training Framework:</strong> Our dataset benefits from the "perfect framework" training strategy, which uses CCAM-coarsened predictor variables. For more information about the perfect and imperfect training frameworks, see the following <a title="review" href="https://journals.ametsoc.org/view/journals/aies/3/2/AIES-D-23-0066.1.xml">review</a></p> </li> </ul> <table> <tbody> <tr> <td> <p><strong>Algorithm</strong></p> </td> <td> <p><strong>Training Data</strong></p> </td> <td> <p><strong>Period</strong></p> </td> </tr> <tr> <td> <p>Deterministic Baseline</p> </td> <td> <p>Historical</p> </td> <td> <p>1960-2014 (~21,000 days)</p> </td> </tr> <tr> <td> <p>Deterministic Baseline</p> </td> <td> <p>Future (SSP370)</p> </td> <td> <p>2044-2099 (~21,000 days)</p> </td> </tr> <tr> <td> <p>Deterministic Baseline</p> </td> <td> <p>Historical and Future (SSP370)</p> </td> <td> <p>1960-2099 (~51,000 days)</p> </td> </tr> <tr> <td> <p>Residual GAN</p> </td> <td> <p>Historical</p> </td> <td> <p>1960-2014</p> </td> </tr> <tr> <td> <p>Residual GAN</p> </td> <td> <p>Future (SSP370)</p> </td> <td> <p>2044-2099</p> </td> </tr> <tr> <td> <p>Residual GAN</p> </td> <td> <p>Historical and Future (SSP370)</p> </td> <td> <p>1960-2099</p> </td> </tr> </tbody> </table> <p><strong>Table 1:</strong> The six RCM emulator experiments performed in this study.</p>
Dataset for publication Harnessing Ti3C2-WS2 Nanostructures as Efficient Energy Scaffoldings for Photocatalytic Hydrogen Generation
<p>The dataset contains all relevant data and figures regarding the manuscript "Harnessing Ti3C2-WS2 Nanostructures as Efficient Energy Scaffoldings for Photocatalytic Hydrogen Generation".</p> <p>All Figures are in tiff format and all relevant data are in csv formats. </p> <p>The data in csv format are labelled as specified in the corresping images (e.g. Figure 1a csv file corresponds to data used to plot graphs from Figure 1a etc.). </p> <p>Axis labeling and units are always specified at the beginning of individual columns. If more than one curve was plotted from the csv file, the conditions can also be found at the beginning of corresponding columns.</p>
Supplemental Dataset Excel files and Source Data Excel file for "START domains generate paralog-specific regulons from a single network architecture"
<p>Supplemental Dataset Excel files and Source Data Excel file for "START domains generate paralog-specific regulons from a single network architecture" in Nat Comms</p>
Dataset for Next-Generation Self-Powered Photodetectors using 2D Bismuth Oxide Selenide Crystals
<p>The dataset contains relevant data and figures regarding the manuscript "Next-Generation Self-Powered Photodetectors using 2D Bismuth Oxide Selenide Crystals".</p> <p>All Figures are in jpg/tiff format and all relevant data are in csv formats. </p> <p>The data in csv format are labelled as specified in the corresping images (e.g. Figure 1a csv file corresponds to data used to plot graphs from Figure 1a etc.). </p> <p>Axis labeling and units are always specified at the beginning of individual columns. If more than one curve was plotted from the csv file, the conditions can also be found at the beginning of corresponding columns.</p>
Dataset for "Early Jurassic rift-related low-pressure-high–temperature granulite facies metamorphism generates widespread peraluminous crustal melts"
<p>This repository contains data for the manuscript "Early Jurassic rift-related low-pressure-high–temperature granulite facies metamorphism generates widespread peraluminous crustal melts" by Anthony Ramírez-Salazar, Mattia Parolari, Arturo Gómez-Tuena, Fernando Ortega-Gutiérrez, and Mariano Elías-Herrera submitted for publication to Geochemistry, Geophysics, Geosystem.</p> <p>The file contains geochemical, isotopic and geochronological data collected from the metapelitic xenoliths of Pepechuca, southern Mexico. All xenoliths were collected between the coordinates 18° 47´ 17.36´´N, 100° 04´ 07.92´´W and 18° 47´ 12.47´´N, 100° 04´ 05.00´´W. The dataset is divided in four tables:</p> <p>Table S1 Xenoliths' Bulk rock major elements<br>Table S2 Xenoliths' trace elements<br>Table S3 Xenoliths' Sr, Nd, Hf and Pb isotopic compositions<br>Table S4 U-Pb and trace element data of the xenoliths</p> <p><br>Anthony Ramírez-Salazar. e-mail: r.s.anthonyy@gmail.com; anthony@geologia.unam.mx</p>
Dataset: Towards a Knowledge Management Framework for LLM-Generated Personas in Collaborative Systems
Open the record for dataset details and reuse information.
Training datasets for "Multi-purpose controllable protein generation via prompted language models"
<div> <p>The datasets used to tune modular prompts of PROPEND fall into three main categories based on their design objectives: tertiary structure, secondary structure, and functional annotation.</p> </div>
Benchmark datasets to study fairness in synthetic data generation
<p>The traveltime dataset is based on the Folktables project covering US census data. The target is a binary variable encoding whether or not the individual needs to travel more than 20 minutes for work; here, having a shorter travel time is the desirable outcome. We use a subset of data from the states of California, Florida, Maine, New York, Utah, and Wyoming states in 2018. Although the folktables dataset does not have any missing values, there are some values recorded as NaN due to the Bureau's data collection methodology. We remove the "esp" column, which encodes the employment status of parents, and has 99.55% missing values. We encode the missing values in the povpip, income to poverty ratio (0.85%), to -1 in accordance to the methodology in Ding et al.. See https://arxiv.org/pdf/2108.04884 for metadata.</p> <p>The cardio (a) dataset contains patient data recorded during medical examination, including 3 binary features supplied by the patient. The target class denotes the presence of cardiovascular disease. This dataset represents predictive tasks that allocate access to priority medical care for patients, and has been used for fairness evaluations in the domain.</p> <p>The credit dataset contains historical financial data of borrowers, including past non-serious delinquencies. Here, a serious delinquency is considered to be 90 days past due, and this is the target variable.</p> <p>The German Credit dataset (https://archive.ics.uci.edu/dataset/144/statlog+german+credit+data) contains financial and personal information regarding loan-seeking applicants.</p>
PSYCHE-D: predicting change in depression severity using person-generated health data (DATASET)
<p>This dataset is made available under <a href="https://creativecommons.org/licenses/by-nc/4.0/">Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)</a>. See LICENSE.pdf for details.</p> <p><strong>Dataset description</strong></p> <p>Parquet file, with:</p> <ul> <li>35694 rows</li> <li>154 columns</li> </ul> <p>The file is indexed on [<em>participant</em>]_[<em>month</em>], such that 34_12 means month 12 from participant 34. All participant IDs have been replaced with randomly generated integers and the conversion table deleted.</p> <p>Column names and explanations are included as a separate tab-delimited file. Detailed descriptions of feature engineering are available from the linked publications.</p> <p>File contains aggregated, derived feature matrix describing person-generated health data (PGHD) captured as part of the DiSCover Project (<a href="https://clinicaltrials.gov/ct2/show/NCT03421223">https://clinicaltrials.gov/ct2/show/NCT03421223</a>). This matrix focuses on individual changes in depression status over time, as measured by PHQ-9.</p> <p>The DiSCover Project is a 1-year long longitudinal study consisting of 10,036 individuals in the United States, who wore consumer-grade wearable devices throughout the study and completed monthly surveys about their mental health and/or lifestyle changes, between January 2018 and January 2020.</p> <p>The data subset used in this work comprises the following:</p> <ul> <li>Wearable PGHD: step and sleep data from the participants’ consumer-grade wearable devices (Fitbit) worn throughout the study</li> <li>Screener survey: prior to the study, participants self-reported socio-demographic information, as well as comorbidities</li> <li>Lifestyle and medication changes (LMC) survey: every month, participants were requested to complete a brief survey reporting changes in their lifestyle and medication over the past month</li> <li>Patient Health Questionnaire (PHQ-9) score: every 3 months, participants were requested to complete the PHQ-9, a 9-item questionnaire that has proven to be reliable and valid to measure depression severity</li> </ul> <p>From these input sources we define a range of input features, both static (defined once, remain constant for all samples from a given participant throughout the study, e.g. demographic features) and dynamic (varying with time for a given participant, e.g. behavioral features derived from consumer-grade wearables).</p> <p>The dataset contains a total of 35,694 rows for each month of data collection from the participants. We can generate 3-month long, non-overlapping, independent samples to capture changes in depression status over time with PGHD. We use the notation ‘SM0’ (sample month 0), ‘SM1’, ‘SM2’ and ‘SM3’ to refer to relative time points within each sample. Each 3-month sample consists of: PHQ-9 survey responses at SM0 and SM3, one set of screener survey responses, LMC survey responses at SM3 (as well as SM1, SM2, if available), and wearable PGHD for SM3 (and SM1, SM2, if available). The wearable PGHD includes data collected from 8 to 14 days prior to the PHQ-9 label generation date at SM3. Doing this generates a total of 10,866 samples from 4,036 unique participants.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.