Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
414
datasets available to search
ShareScore release 0.9.0
Dataset results
414 results for “Generative Model”
A Generic Model for Benchmark Aerodynamic Analysis of Fifth-Generation High-Performance Aircraft (CGNS grid files)
<p>Openly available supplementary data to accompany paper https://doi.org/10.3390/aerospace10090746. This data set includes unstructured CGNS grid files to facilitate code comparison. When using this data, please cite:</p> <p>Giannelis, N.F.; Bykerk, T.; Vio, G.A. A Generic Model for Benchmark Aerodynamic Analysis of Fifth-Generation High-Performance Aircraft. Aerospace 2023, 10, 746.</p>
Image datasets used in the paper "Revealing invisible cell phenotypes with conditional generative modeling"
<p>- BBBC021_selection_128 is a selection of the BBBC021 image dataset from the Broad Bioimage Benchmarck Collection from the Broad Institute</p> <p>- golgi_256_subset is a subset (one plate) of the Golgi Dataset we used (which is about 3 times larger). It was generated by the Biophenics platform in Institut Curie, Paris, France</p> <p>- translocation_256 is the translocation Dataset we used. It was generated by the Biophenics platform in Institut Curie, Paris, France</p> <p>- LRKK2_256 is the Parkinson LRKK2 mutation dataset we used. It was generated by Ksilink, Strasbourg, France</p> <p>- smala_256 is the Malaria dataset we used. It was generated by IRD, Paris, France and acquired by the Histopathology Platform at Institut Pasteur in Paris, France. </p>
Molecules used to train or generated by chemical language models
<p>This upload contains training datasets or generated molecules from the paper “Invalid SMILES are helpful, not harmful, for chemical language models.”</p> <p>The contents of the directories are as follows:</p> <ul> <li>training_sets: sets of molecules from ChEMBL or GDB-13 used to train chemical language models, represented either as SMILES or SELFIES</li> <li>sampled-*: unprocessed samples of 10 million molecules from each model trained on ChEMBL or GDB-13</li> <li>prior_inputs: sets of molecules from LOTUS, COCONUT, FooDB and NORMAN, split into ten folds and used to train chemical language models</li> <li>priors-*: samples of 100 million molecules from chemical language models trained on each cross-validation fold, with unique molecules represented as canonical SMILES and sorted in descending order by their sampling frequency</li> </ul>
Hierarchical generative modelling for autonomous robots
<p>This is the dataset accompanying the paper "Hierarchical generative modelling for autonomous robots" by Yuan et al.</p>
Survey2Survey: A deep learning generative model approach for cross-survey image mapping
<p>During the last decade, there has been an explosive growth in survey data and deep learning techniques, both of which have enabled great advances for astronomy. The amount of data from various surveys from multiple epochs with a wide range of wavelengths, albeit with varying brightness and quality, is overwhelming, and leveraging information from overlapping observations from different surveys has limitless potential in understanding galaxy formation and evolution. Synthetic galaxy image generation using physical models has been an important tool for survey data analysis, while deep learning generative models show great promise. In this paper, we present a novel approach for robustly expanding and improving survey data through cross survey feature translation. We trained two types of neural networks to map images from the Sloan Digital Sky Survey (SDSS) to corresponding images from the Dark Energy Survey (DES). This map was used to generate false DES representations of SDSS images, increasing the brightness and S/N while retaining important morphological information. We substantiate the robustness of our method by generating DES representations of SDSS images from outside the overlapping region, showing that the brightness and quality are improved even when the source images are of lower quality than the training images. Finally, we highlight several images in which the reconstruction process appears to have removed large artifacts from SDSS images. While only an initial application, our method shows promise as a method for robustly expanding and improving the quality of optical survey data and provides a potential avenue for cross-band reconstruction.</p><p>------------</p><p>This repository contains the image files from Survey2Survey: a deep learning generative model approach for cross-survey image mapping. Please cite https://arxiv.org/abs/2011.07124 if you use this data in a publication. For more information, contact Brandon Buncher at buncher2(at)illinois.edu</p><p><strong>--- Directory structure ---</strong></p><p>tutorial.ipynb demonstrates how to load the image files (uploaded here as tarballs). Images were obtained from the SDSS DR16 cutout server (https://skyserver.sdss.org/dr16/en/help/docs/api.aspx) and DES DR1 cutout server (https://des.ncsa.illinois.edu/desaccess/</p><ul><li>./sdss_train/ and ./des_train/ contain the original SDSS and DES images used to train the neural network (Stripe82)</li><li>./sdss_test/ and ./des_test/ contain the original SDSS and DES images used for the validation dataset (Stripe82)</li><li>./sdss_ext/ contain images from the external SDSS dataset (SDSS images without a DES counterpart, outside Stripe82)</li><li>./cae and ./cyclegan contain images generated by the CAE and CycleGAN, respectively. train_decoded/ and test_decoded/ contain the reconstructions of the images from the training dataset and test dataset, respectively. external_decoded/ contain the DES-like image reconstructions of SDSS objects from the external dataset (outside Stripe82).</li></ul>
[Replication Package] Evaluating the Impact of Post-Training Quantization on Large Language Models for Code Generation
<p>This repository contains scripts, datasets, and results of the work <em>"Evaluating the Impact of Post-Training Quantization on Large Language Models for Code Generation"</em></p> <p><strong>Scripts contained in this Zenodo repository can also be visualized at the following link: <a href="https://anonymous.4open.science/r/lowbit-quantization-D070/README.md">https://anonymous.4open.science/r/lowbit-quantization-D070/README.md</a><br></strong></p>
Personalised Therapy for Metastatic ADPC Determined by Genetic Testing and Avatar Model Generation
ClinicalTrials.gov study NCT02795650. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.
Modelling Tau Distribution From DTI With Generative Adversarial Network for Alzheimer's Disease Diagnosis
ClinicalTrials.gov study NCT05020626. IPD Sharing: Not stated. Countries: 1. Publications: 33.
Building Regulation in Dual Generations - Telehealth Model
ClinicalTrials.gov study NCT04639557. IPD Sharing: YES. Countries: 1. Publications: 1.
Large Language Model-Generated Messages to Improve Guideline-Directed Medical Therapy in Heart Failure
ClinicalTrials.gov study NCT07337577. IPD Sharing: UNDECIDED. Countries: 1. Publications: 3.
Chest X-Ray Image Diagnosis and Report Generation Dedicated Model Based on Deepseek
ClinicalTrials.gov study NCT06874647. IPD Sharing: NO. Countries: 1. Publications: 1.
Generation of Marfan Syndrome and Fontan Cardiovascular Models Using Patient-specific Induced Pluripotent Stem Cells
ClinicalTrials.gov study NCT02815072. IPD Sharing: NO. Countries: 1. Publications: 2.
Measurement and modeling of the multi-wavelength optical properties of uncoated flame-generated soot: Data from the BC2, BC3, BC3+ and BC4 studies
Open the record for dataset details and reuse information.
Data from: "Genome-wide microsatellite marker development from next-generation sequencing of two non-model bat species impacted by wind turbine mortality: Lasiurus borealis and L. cinereus (Vespertilionidae)" in Genomic Resources Notes accepted 1 October 2013 to 30 November 2013
Open the record for dataset details and reuse information.
Data from: A model-derived short-term estimation method of effective size for small populations with overlapping generations
Open the record for dataset details and reuse information.
Data from: Validation of the What Matters Index: a brief, patient-reported index that guides care for chronic conditions and can substitute for computer-generated risk models
Open the record for dataset details and reuse information.
UVic earth system climate model data generated under MIS3 boundary conditions
Open the record for dataset details and reuse information.
Data from: Reduced incompatibility in the production of second generation hybrids between two Magnolia species revealed by Bayesian gene dispersal modeling
Open the record for dataset details and reuse information.
Data from: Genome-wide single nucleotide polymorphism (SNP) identification and characterization in a non-model organism, the African buffalo (Syncerus caffer), using next generation sequencing
Open the record for dataset details and reuse information.
A modified niche model for generating food webs with stage-structured consumers: The stabilizing effects of life-history stages on complex food webs
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.