Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

414

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

414 results for “Generative Model”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: A stochastic generative model for citation networks among academic papers

Open the record for dataset details and reuse information.

publicJun 2022View details →
zenodo28/100

Model runs generated by publication "Re-evaluating 14 C dating accuracy in deep-sea sediment archives"

<p>Model runs generated by following publication:</p> <p>B.C. Lougheed, P. Ascough, A. Dolman, L. L&ouml;wemark and B. Metcalfe, 2020. &ldquo;Re-evaluating 14C dating accuracy in deep-sea sediment archives.&rdquo; Geochronology, doi:10.5194/gchron-2019-10</p> <p>Contains .mat files that can be opened by the Matlab / Octave environment.</p>

opencc-by-4.0Mar 2020View details →
zenodo28/100

Investigating Distributional Robustness: Semantic Perturbations Using Generative Models (ImageNet Examples)

<p>This dataset contains examples of semantically-perturbed images, for NeurIPS 2020 submission #4915.</p> <p>There are four top-level folders, each containing results for semantic perturbations restricted to adjust the activation values at only certain layers of the BigGAN generative network: the first six layers, the middle six layers, the last six layers, and all layers.</p> <p>Within each top-level folder, there are a further four folders, each corresponding to a classifier neural network whose evaluation is being evaluated. These are&nbsp;EfficientNet-B4 with NoisyStudent training [1],&nbsp;the standard ResNet50 [2], a pixel-perturbation-robust ResNet50 trained by&nbsp;Engstrom et al. [3]&nbsp;and another trained by&nbsp;Wong et al. [4], using their &quot;Fast is better than free&quot; technique.</p> <p>Within each of these, there are many folders, named &#39;version_$N&#39;. Each one of these contains three images: the unperturbed generated image, named&nbsp;unpert_generated_x_grid_0.png; the semantically-perturbed generated image, named&nbsp;generated_x_grid_0.png; and an image named semantic_pert_diffs_grid_0.png showing the pixel-space effect of the semantic perturbation, that is, the diff between the perturbed and unperturbed images. Note that if the&nbsp;perturbed and unperturbed images are identical, and the classifier misclassifies the unperturbed images, and so we skip this example.</p> <p>Along with the &#39;version_$N&#39; folders containing the images, there exists a file for each classifier named results.json. Each top-level item in this JSON file corresponds to one &#39;version_$N&#39; example. There are 5 attributes: &#39;label&#39;, indicating the target label of the unperturbed image; &#39;magnitude&#39;, which gives the magnitude of the semantic perturbation found; &#39;skipped_cla&#39;, which is 1 if the example is skipped because the classifier did not correctly classify the unperturbed image; &#39;skipped_judge&#39;, which is 1 if the human judged that the unperturbed image did not match its label, so this example is skipped; and &#39;pert_judgement&#39;, which is 1 if the semantically-perturbed image is judged by the human to be of the same class as the unperturbed image. These judgements on these images were used to construct the main graphs in the paper.</p> <p>&nbsp;</p> <p>[1]&nbsp;Qizhe Xie, Eduard H. Hovy, Minh-Thang Luong, and Quoc V. Le. Self-training with Noisy Student improves ImageNet classification. CoRR, abs/1911.04252, 2019. URL http://arxiv.org/abs/1911.04252.</p> <p>[2] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pages 770&ndash;778. IEEE Computer Society, 2016. doi: 10.1109/CVPR.2016.90. URL https://doi.org/10.1109/CVPR.2016.90.</p> <p>[3] Logan Engstrom, Andrew Ilyas, Shibani Santurkar, and Dimitris Tsipras. Robustness (Python library), 2019. URL ttps://github.com/MadryLab/robustness.</p> <p>[4] Eric Wong, Leslie Rice, and J. Zico Kolter. Fast is better than free: Revisiting adversarial training. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL<br> https://openreview.net/forum?id=BJx040EFvH.</p>

opencc-by-4.0Jun 2020View details →
zenodo28/100

Experimental data and linear model for "Asymmetric internal tide generation in the presence of a steady flow"

<p>This dataset contain the experimental data, analyses, and linear model that are described in the manuscript</p> <p>&quot;Asymmetric internal tide generation in the presence of a steady flow&quot;. submitted to Journal of Geophysical Research - Oceans.</p> <p>The folder data_density_fields_fluxes contains the experimental density fields (itXX/results/densityfields) and energy fluxes (itXX/results/fluxes) for the five experiments described in the manuscript: exp I (it17), exp II (it16), exp III (it15), exp IV (it13) and exp (V) (it 18).</p> <p>The Dossmannetal_wavesolution.m file is the linear model described in the manuscript for internal wave generation over a ridge.</p>

opencc-by-4.0Jun 2020View details →
dryad28/100

Data from: In silico study of the role of cell growth factors in photosynthesis using a virtual leaf tissue generator coupled to a microscale photosynthesis gas exchange model

Computational tools that allow in silico analysis of the role of cell growth and division on photosynthesis are scarce. We present a freely available tool that combines a virtual leaf tissue generator and a two-dimensional microscale model of gas transport during C3 photosynthesis. A total of 270 mesophyll geometries were generated with varying degree of growth anisotropy, growth extent and extent of schizogenous airspace formation in the palisade mesophyll. The anatomical properties of the virtual leaf tissue and microscopic cross sections of actual leaf tissue of tomato (Solanum lycopersicum L.) were statistically compared. Model equations for transport of CO2 in the liquid phase of the leaf tissue were discretized over the geometries. The virtual leaf tissue generator produced a leaf anatomy of tomato that was statistically similar to real tomato leaf tissue. The response of photosynthesis to intercellular CO2 predicted by a model that used the virtual leaf tissue geometry compared well with measured values. The results indicate that the light-saturated rate of photosynthesis was influenced by interactive effects of extent and directionality of cell growth and degree of airspace formation through the exposed surface of mesophyll per leaf area. The tool could be used further in investigations of improving photosynthesis and gas exchange in relation to cell growth and leaf anatomy.

opencc-zeroOct 2020View details →
zenodo28/100

Generating Adaptation Plans Based on Quality Models for Cloud Platforms

<p>Apresenta&ccedil;&atilde;o referente ao artigo &quot;Generating Adaptation Plans Based on Quality Models for Cloud Platforms&quot;. SBES 2020 - Trilha de Ideias Inovadoras e Resultados Emergentes.</p>

opencc-by-4.0Oct 2020View details →
zenodo28/100

Generating Trustworthiness Adaptation Plans Based on Quality Models for Cloud Platforms

<p>Apresenta&ccedil;&atilde;o referente ao artigo &quot;Generating Trustworthiness Adaptation Plans Based on Quality Models for Cloud Platforms&quot;. SBCARS 2020 - Artigos Completos</p>

opencc-by-4.0Oct 2020View details →
zenodo28/100

4x and 10x Super Resolution Generator Models Trained With Planet CubeSat Satellite Imagery

<p>Final resampling generator models produced from the Enhanced Super Resolution Generative Adversarial Network (ESRGAN) (https://github.com/xinntao/ESRGAN). ESRGAN was trained at two different resampling factors, 4x and 10x,&nbsp;using a training data set of global Planet CubeSat satellite images. These generators can be used to resample Planet CubeSat satellite images from 30m and 12m to 3m resolution. Descriptions and results of training can be found at&nbsp;https://wandb.ai/elezine/pixelsmasher. In press at Canadian Journal of Remote Sensing:&nbsp;Super-resolution surface water mapping on the&nbsp;Canadian&nbsp;Shield using Planet CubeSat images and a Generative Adversarial Network, Ekaterina M. D. Lezine, Ethan D. Kyzivat, and Laurence C. Smith (2021).&nbsp;</p>

opencc-by-4.0Oct 2020View details →
dryad28/100

Data from: Multiple models generate a geographic mosaic of resemblance in a Batesian mimicry complex

Batesian mimics—benign species that receive protection from predation by resembling a dangerous species—often occur with multiple model species. Here we examine whether geographic variation in the number of local models generates geographic variation in mimic-model resemblance. In areas with multiple models, selection might be relaxed or even favour imprecise mimicry relative to areas with only one model. We test the prediction that model-mimic match should vary with the number of other model species in a broadly-distributed snake mimicry complex where a mimic and a model co-occur both with and without other model species. We found that the mimic resembled its model more closely when they were exclusively sympatric than when they were sympatric with other model species. Moreover, in regions with multiple models, mimic-model resemblance was positively correlated with the resemblance between the model and other model species. However, contrary to predictions, free-ranging natural predators did not attack artificial replicas of imprecise mimics more often when only a single model was present. Taken together, our results suggest that multiple models might generate a geographic mosaic in the degree of phenotype matching between Batesian mimics and their models.

opencc-zeroAug 2019View details →
dryad28/100

Data from: Sequence Capture using PCR-generated Probes (SCPP): a cost-effective method of targeted high-throughput sequencing for non-model organisms

Recent advances in high-throughput sequencing library preparation and subgenomic enrichment methods have opened new avenues for population genetics and phylogenetics of non-model organisms. To multiplex large numbers of indexed samples while sequencing predominantly orthologous, targeted regions of the genome, we propose modifications to an existing, in-solution capture that utilizes PCR products as target probes to enrich library pools for the genomic subset of interest. The sequence capture using PCR-generated probes (SCPP) protocol requires no specialized equipment, is highly flexible, and significantly reduces experimental costs for projects where a modest scale of genetic data is optimal (25-100 genomic loci). Our alterations enable application of this method across a wider phylogenetic range of taxa and result in higher capture efficiencies and coverage at each locus. Efficient and consistent capture over multiple SCPP experiments and at various phylogenetic distances is demonstrated, extending the utility of this method to both phylogeographic and phylogenomic studies.

opencc-zeroDec 2013View details →
zenodo28/100

Supporting information, WRF output, and processed WRF files used for the generation of the manuscript "Observational and modelling analysis of Canada's only F5/EF5 tornado"

<p>WRF simulation output and processed files used to generate the figures and calculations in the manuscript "Observational and modelling analysis of Canada's only F5/EF5 tornado". &nbsp;Two zipped folders are attached, one is from the original control simulation with the microphysics scheme on (MP), and the other is from the experimental simulation with the microphysics scheme turned off (NOMP).</p><p>Each simulation zipped folder includes sub-directories containing the processed observed and simulated surface station data, surface wet-bulb potential temperature, cross sections (the MP simulation only), and convective parameter fields at 2100 UTC 22 June 2007.</p><p>A supplemental material in the form of a movie showing the radar observation between 2000 UTC 22 June 2007 and 0000 UTC 23 June 2007 is also attached. See the manuscript's Figure 5 caption for more information on the data shown in the animation.</p>

opencc-by-4.0Nov 2023View details →
zenodo28/100

HALO: An Ontology for Representing Hallucinations in Generative Models

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo28/100

Data Set for Predicting the Performance of ATL Model Transformations Based on Generated Models

<p>Predicting the execution time of model transformations can help to understand how a transformation reacts to a given input model without creating and transforming the respective model.</p> <p>In our previous data set (https://doi.org/10.5281/zenodo.8385957), we have documented our experiments in which we predict the performance of ATL transformations using predictive models obtained from training linear regression, random forest and support vector regression. As input for the prediction, our approach uses a characterization of the input model. In these experiments, we only used data from real models.</p> <p>However, a common problem is that transformation developers do not have enough models available to use such a prediction approach. Therefore, in a new variant of our experiments, we investigated whether the three considered machine learning approaches can predict the performance of transformations if we use data from generated models for training. We also investigated whether it is possible to achieve good predictions with smaller training data. The dataset provided here offers the corresponding raw data, scripts, and results.</p> <p>A detailed documentation is available in documentaion.pdf.</p>

opencc-by-4.0Dec 2023View details →
zenodo28/100

Audio samples from generative models trained on the TIMIT speech data.

<p>The snippets include samples and reconstructions.&nbsp;All samples are completely unconditional and utilise only the prior&nbsp;internal representations learned by the models.&nbsp;Reconstructions are computed from a given test audio snippet by first encoding it to a learned representation and then decoding that&nbsp;to a reconstruction of the audio.</p> <p>All models are trained on the TIMIT speech dataset (<a href="https://catalog.ldc.upenn.edu/LDC93s1">https://catalog.ldc.upenn.edu/LDC93s1</a>). Some snippets are from models&nbsp;trained at different temporal resolutions denoted by `s1` and `s64`. We refer to the paper for details.</p> <p>The files include:</p> <ul> <li>`clockwork-vae-s64-reconstruction-*` <ul> <li>Four reconstructions using a&nbsp;two-layered Clockwork VAE trained with temporal resolution s=64.</li> </ul> </li> <li>`clockwork-vae-s64-sample-*` <ul> <li>Four samples from the prior of a Clockwork VAE trained with temporal resolution s=64.</li> </ul> </li> <li>`original-*` <ul> <li>Four original samples from TIMIT corresponding in pairs to the reconstructions.</li> </ul> </li> <li>`vrnn-s64-sample-*` <ul> <li>Two samples from the prior of a VRNN trained with temporal resolution s=64.</li> </ul> </li> <li>`vrnn-s1-sample-*` <ul> <li>Two samples from the prior of a VRNN trained with temporal resolution s=1.</li> </ul> </li> <li>`srnn-s64-sample-*` <ul> <li>Two samples from the prior of a SRNN trained with temporal resolution s=64.</li> </ul> </li> <li>`srnn-s1-sample-*` <ul> <li>Two samples from the prior of a SRNN trained with temporal resolution s=1.</li> </ul> </li> <li>`wavenet-s64-sample-*` <ul> <li>Four samples from a WaveNet trained with temporal resolution s=1.</li> </ul> </li> <li>`wavenet-s1-sample-*` <ul> <li>Two samples from a WaveNet trained with temporal resolution s=64.</li> </ul> </li> </ul>

opencc-by-4.0Jan 2022View details →
zenodo28/100

UFO model for Vector-like Quarks at NLO QCD with five flavor scheme (third generation)

<p>Vector-like Quark UFO Model at NLO QCD with five flavour scheme (3rd generation only)</p>

opencc-zeroAug 2022View details →
zenodo28/100

UFO model for Vector-like Quarks at NLO QCD with four flavor scheme (third generation)

<p>Vector-like Quark UFO Model at NLO QCD with four flavour scheme (3rd generation only)</p>

opencc-zeroAug 2022View details →
zenodo28/100

Systematic Mapping Appendix - Identifying approaches to generate test cases in model-based testing: a

<p>The&nbsp; appendix&nbsp; present the complete mapping of individual studies and its categories.</p>

opencc-by-4.0Feb 2019View details →
zenodo28/100

Understanding The Impact of Solver Choice in Model-Based Test Generation

<p><strong># Solver Experiment - Data Package</strong></p> <p>This repository is intended to allow replication and extension of the study<br> conducted in the following paper:</p> <p>Ying Meng and Gregory Gay. Understanding The Impact&nbsp;of Solver Choice in Model-Based Test Generation. To appear, 2020&nbsp;ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM). Available from&nbsp;<a href="https://greg4cr.github.io/pdf/20solvers.pdf">https://greg4cr.github.io/pdf/20solvers.pdf</a>.</p> <p>This package includes all data used or generated in this experiment.<br> &nbsp; * Models<br> &nbsp; * Mutants<br> &nbsp; * Test Suites<br> &nbsp; * Execution Traces<br> &nbsp; * Fault-detection Results</p> <p>If you have any questions on this data, please contact us at greg@greggay.com.&nbsp;<br> &nbsp;</p>

opencc-by-4.0Oct 2019View details →
zenodo28/100

Datasets accompanying Deep generative AI models analyzing circulating orphan non-coding RNAs enable accurate detection of early-stage non-small cell lung cancer

<p>These datasets accompany the manuscript "Deep generative AI models analyzing circulating orphan non-coding RNAs enable accurate detection of early-stage non-small cell lung cancer".</p> <p>The datasets include:</p> <ul> <li> <p><code>metadata.tsv.gz</code>&nbsp;Phenotype information of the samples</p> </li> <li> <p><code>mirna_counts.tsv.gz</code>&nbsp;MicroRNA count matrix (for data normalization)</p> </li> <li> <p><code>oncrna_counts.tsv.gz</code>&nbsp;oncRNA count matrix</p> </li> <li> <p><code>orion/</code> Ensemble of Orion models</p> </li> </ul>

opencc-by-nc-nd-4.0Aug 2024View details →
zenodo28/100

Pocket-based generation task data for Token-Mol 1.0 model finetuning

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record