Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

94

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

94 results for “Generative AI”

Learn how ShareScore rates datasets ↗
zenodo36/100

Generative AI for designing and validating easily synthesizable and structurally novel antibiotics: Data and Models

<p>This repository contains data and models used in the following paper.</p> <p>Swanson, K., Liu, G., Catacutan, D., Zou, J. &amp; Stokes, J. <a href="https://www.nature.com/articles/s42256-024-00809-7">Generative AI for designing and validating easily synthesizable and structurally novel antibiotics</a>. <em>Nature Machine Intelligence, </em>2024.</p> <p>The data and models are meant to be used with the <a href="https://github.com/swansonk14/SyntheMol">SyntheMol</a> code. More details about how to use the data and models with the code are available <a href="https://github.com/swansonk14/SyntheMol/tree/main/docs">here</a>.</p> <p>The Data.zip file has the following structure. Note that the numbers for the Data subdirectories correspond to the supplementary data numbers in the paper (e.g., 1_training_data corresponds to Supplementary Data 1).</p> <p>Data</p> <p>&nbsp; 1_training_data: The <em>Acinetobacter baumannii</em> inhibition data used to train antibiotic property prediction models.</p> <p>&nbsp; 2_chembl: Known antibiotic and antibacterial molecules from <a href="https://www.ebi.ac.uk/chembl/">ChEMBL</a>, which are used to compute the novelty of generated antibiotic candidates.</p> <p>&nbsp; 4_real_space: Data files and statistics for the <a href="https://enamine.net/compound-collections/real-compounds/real-space-navigator">Enamine REAL Space</a>. The molecular building blocks file is version 2021 q3-4 while all other REAL Space details are computed from the full enumerated REAL space version 2022 q1-2 (downloaded on August 30, 2022).</p> <p>&nbsp; 5_generations_clogp: Compounds generated by SyntheMol using Chemprop models trained to predict cLogP.</p> <p>&nbsp; 6_generations_chemprop: Compounds generated by SyntheMol using Chemprop models trained to predict <em>A. baumannii</em> inhibition.</p> <p>&nbsp; 7_generations_chemprop_rdkit: Compounds generated by SyntheMol using Chemprop-RDKit models trained to predict <em>A. baumannii</em> inhibition.</p> <p>&nbsp; 8_generations_random_forest: Compounds generated by SyntheMol using random forest models trained to predict <em>A. baumannii</em> inhibition.</p> <p>&nbsp; 9_synthesized: Information on the 58 SyntheMol-generated compounds that were successfully synthesized by Enamine.</p> <p>The Models.zip file contains one folder for each model used in the paper. Note that each model is technically an ensemble of ten individual models, so each directory contains ten model files.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Datasets used in the paper "The Face of Deception: The Impact of AI-Generated Photos on Malicious Social Bots"

<p>Datasets used in the paper &quot;The Face of Deception: The Impact of AI-Generated Photos on Malicious Social Bots&quot;</p> <p>We changed the datasets&#39; titles and omitted authors&#39; names for the blind review process. After the review, we will upload it to GitHub in an unanonymised form.</p> <p>Check README.md for details.</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Replication package for: "Corrupted by Algorithms? How AI-generated and Human-written Advice Shape (Dis)honesty"

<p>Package to the following paper:</p> <p>Leib, M; K&ouml;bis, N; Rilke, R M; Hagens, M; Irlenbusch, B (2023)&nbsp; Corrupted by Algorithms? How AI-generated and Human-written Advice Shape (Dis)honesty</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

Replication Package for "The Double-edged Sword of Banning Generative AI on Online Question-and-Answer Communities: Evidence from Stack Exchange"

<p>This is a replication package for &quot;The Double-edged Sword of Banning Generative AI on Online Question-and-Answer Communities: Evidence from Stack Exchange&quot;.</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

Dataset for:Doctoral students' reflection on Generative AI: a librarian outlook

<p>Data for practice paper:</p> <p><span><span>Doctoral </span></span><span><span>students&rsquo; </span></span><span><span>reflection on</span></span><span><span> Generative AI: a librarian outlook</span></span></p> <p>Contains:</p> <p>Instruction for written assignment used for analysis.&nbsp;</p> <p>Excel file with raw data from follow up questionnaire.</p> <p>Diagrams of answers in follow up questionnaire.</p> <p>Questionnaire form.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2024View details →
ClinicalTrials.gov36/100

Generative AI-Based Simulation for Diagnostic Communication in Type 2 Diabetes (DIALOGUE-DM2)

ClinicalTrials.gov study NCT07252193. IPD Sharing: YES. Countries: 1. Publications: 1.

controlledIPD-YESFeb 2026View details →
dryad36/100

Eyes don’t lie: Indifferent to AI-generated depression screenings

Open the record for dataset details and reuse information.

publicJul 2025View details →
dryad36/100

Who expands the human creative frontier with generative AI: Hiveminds or masterminds?

Open the record for dataset details and reuse information.

publicAug 2025View details →
zenodo32/100

Expert and AI-generated annotations of the tissue types for the RMS-Mutation-Prediction microscopy images

<div> <p>This dataset corresponds to a collection of images and/or image-derived data available from National Cancer Institute <a href="https://portal.imaging.datacommons.cancer.gov/">Imaging Data Commons (IDC)</a> [1]. This dataset was converted into DICOM representation and ingested by the IDC team. You can explore and visualize the corresponding images using IDC Portal here: <a href="https://portal.imaging.datacommons.cancer.gov/explore/filters/?analysis_results_id=RMS-Mutation-Prediction-Expert-Annotations">https://portal.imaging.datacommons.cancer.gov/explore/filters/?analysis_results_id=RMS-Mutation-Prediction-Expert-Annotations</a>.. You can use the manifests included in this Zenodo record to download the content of the collection following the&nbsp;<strong>Download instructions</strong>&nbsp;below.</p> <h3>Collection description</h3> </div> <div> <div> <p>This dataset contains 2 components:</p> <ol> <li>Annotations of multiple&nbsp; regions of interest performed by an expert pathologist with eight years of experience for a subset of hematoxylin and eosin (H&amp;E) stained images from the RMS-Mutation-Prediction image collection [1,2]. Annotations were generated manually, using the Aperio ImageScope tool, to delineate regions of alveolar rhabdomyosarcoma (ARMS), embryonal rhabdomyosarcoma (ERMS), stroma, and necrosis [3]. The resulting planar contour annotations were originally stored in ImageScope-specific XML format, and subsequently converted into Digital Imaging and Communications in Medicine (DICOM) Structured Report (SR) representation using the open source conversion tool [4].</li> <li>AI-generated annotations stored as probabilistic segmentations.</li> </ol> <p><strong>WARNING</strong>: After the release of IDC v20 (v2 of this data record), it was discovered that a mistake had been made during data conversion that affected the newly-released segmentations accompanying the "RMS-Mutation-Prediction" collection. Segmentations released in v20 for this collection have the segment labels for alveolar rhabdomyosarcoma (ARMS) and embryonal rhabdomyosarcoma (ERMS) switched in the metadata relative to the correct labels. Thus segment 3 in the released files is labelled in the metadata (the SegmentSequence) as ARMS but should correctly be interpreted as ERMS, and conversely segment 4 in the released files is labelled as ERMS but should be correctly interpreted as ARMS. This mistake was fixed in the version v3 of this record (IDC data release v21).</p> <p>Many pixels from the whole slide images annotated by this dataset are not contained inside any annotation contours and are considered to belong to the background class. Other pixels are contained inside only one annotation contour and are assigned to a single class.&nbsp; However,&nbsp; cases also exist in this dataset where annotation contours overlap.&nbsp; In these cases, the pixels contained in multiple contours could be assigned membership in multiple classes.&nbsp; One example is a necrotic tissue contour overlapping an internal subregion of an area designated by a larger ARMS or ERMS annotation.&nbsp; The ordering of annotations in this DICOM dataset preserves the order in the original XML generated using ImageScope.&nbsp; These annotations were converted, in sequence, into segmentation masks and used in the training of several machine learning models. Details on the training methods and model results&nbsp; are presented in [1].&nbsp; In the case of overlapping contours, the order in which annotations are processed may affect the generated segmentation mask if prior contours are overwritten by later contours in the sequence.&nbsp; It is up to the application consuming this data to decide how to interpret tissues regions annotated with multiple classes. The annotations included in this dataset are available for visualization and exploration from the National Cancer Institute Imaging Data Commons (IDC) [5] (also see IDC Portal at <a href="https://imaging.datacommons.cancer.gov/">https://imaging.datacommons.cancer.gov</a>) as of data release v18.&nbsp;Direct link to open the collection in IDC Portal: <a href="https://portal.imaging.datacommons.cancer.gov/explore/filters/?analysis_results_id=RMS-Mutation-Prediction-Expert-Annotations">https://portal.imaging.datacommons.cancer.gov/explore/filters/?analysis_results_id=RMS-Mutation-Prediction-Expert-Annotations</a>.</p> </div> <div> <h3>Files included</h3> <p>A manifest file's name indicates the IDC data release in which a version of collection data was first introduced. For example,&nbsp;<code>pan_cancer_nuclei_seg_dicom-collection_id-idc_v19-aws.s5cmd</code> corresponds to the annotations for th eimages in the <code>collection_id</code> collection introduced in IDC data release v19. DICOM Binary segmentations were introduced in IDC v20. If there is a subsequent version of this Zenodo page, it will indicate when a subsequent version of the corresponding collection was introduced.</p> <p>For each of the collections, the following manifest files are provided:</p> <ol> <li><code>rms_mutation_prediction_expert_annotations-idc_v20-aws.s5cmd</code>: manifest of files available for download from public IDC Amazon Web Services buckets</li> <li><code>rms_mutation_prediction_expert_annotations-idc_v20-gcs.s5cmd</code>: manifest of files available for download from public IDC Google Cloud Storage buckets</li> <li><code>rms_mutation_prediction_expert_annotations-idc_v20-dcf.dcf</code>: Gen3 manifest (for details see&nbsp;<a href="../records/Gen3%20manifest%20documentation">https://learn.canceridc.dev/data/organization-of-data/guids-and-uuids</a>)</li> </ol> <p>Note that manifest files that end in&nbsp;<code>-aws.s5cmd</code>&nbsp;reference files stored in Amazon Web Services (AWS) buckets, while&nbsp;<code>-gcs.s5cmd</code>&nbsp;reference files in Google Cloud Storage. The actual files are identical and are mirrored between AWS and GCP.</p> <h3>Download instructions</h3> <p>Each of the manifests include instructions in the header on how to download the included files.</p> <p>To download the files using&nbsp;<code>.s5cmd</code>&nbsp;manifests:</p> <ol> <li>install <a href="https://github.com/imagingdatacommons/idc-index" target="_blank" rel="noopener">idc-index</a> package: <code>pip install --upgrade idc-index</code></li> <li>download the files referenced by manifests included in this dataset by passing the&nbsp;<code>.s5cmd</code>&nbsp;manifest file:&nbsp;<code>idc download&nbsp;manifest.s5cmd</code></li> </ol> <p>To download the files using&nbsp;<code>.dcf</code> manifest, see manifest header.</p> <h3>Acknowledgments</h3> <p>Imaging Data Commons team has been funded in whole or in part with Federal funds from the National Cancer Institute, National Institutes of Health, under Task Order No. HHSN26110071 under Contract No. HHSN261201500003l.</p> <p>If you use the files referenced in the attached manifests, we ask you to cite this dataset, as well as the publication describing the original dataset&nbsp;<a href="https://paperpile.com/c/NHiBXI/njdR">[2]</a>&nbsp;and publication acknowledging IDC&nbsp;<a href="https://paperpile.com/c/NHiBXI/uJJZ">[5]</a>.</p> <h3>References</h3> </div> </div> <div> <p>[1] D. Milewski et al., "Predicting molecular subtype and survival of rhabdomyosarcoma patients using deep learning of H&amp;E images: A report from the Children's Oncology Group," Clin. Cancer Res., vol. 29, no. 2, pp. 364&ndash;378, Jan. 2023, doi: 10.1158/1078-0432.CCR-22-1663.</p> <p>[2] Clunie, D., Khan, J., Milewski, D., Jung, H., Bowen, J., Lisle, C., Brown, T., Liu, Y., Collins, J., Linardic, C. M., Hawkins, D. S., Venkatramani, R., Clifford, W., Pot, D., Wagner, U., Farahani, K., Kim, E., &amp; Fedorov, A. (2023). DICOM converted whole slide hematoxylin and eosin images of rhabdomyosarcoma from Children's Oncology Group trials [Data set]. Zenodo.&nbsp;<a href="https://doi.org/10.5281/zenodo.8225132" rel="noopener">https://doi.org/10.5281/zenodo.8225132</a></p> <p>[3] Agaram NP. Evolving classification of rhabdomyosarcoma. Histopathology. 2022 Jan;80(1):98-108. doi: 10.1111/his.14449. PMID: 34958505; PMCID: PMC9425116,https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9425116/</p> <p>[4] Chris Bridge. (2024). ImagingDataCommons/idc-sm-annotations-conversion: v1.0.0 (v1.0.0). Zenodo.&nbsp;<a href="https://doi.org/10.5281/zenodo.10632182" rel="noopener">https://doi.org/10.5281/zenodo.10632182</a></p> <p>[5] Fedorov, A., Longabaugh, W. J. R., Pot, D., Clunie, D. A., Pieper, S. D., Gibbs, D. L., Bridge, C., Herrmann, M. D., Homeyer, A., Lewis, R., Aerts, H. J. W. L., Krishnaswamy, D., Thiriveedhi, V. K., Ciausu, C., Schacherer, D. P., Bontempi, D., Pihl, T., Wagner, U., Farahani, K., Kim, E. &amp; Kikinis, R. National cancer institute imaging data commons: Toward transparency, reproducibility, and scalability in imaging artificial intelligence. Radiographics 43, (2023).</p> </div>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Exploring Generative AI Tools for Software Quality: Insights from a Rapid Multivocal Literature Review

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo32/100

Experts fail to reliably detect AI-generated histological data

<p>This repository contains material related to the paper "<em>Experts fail to reliably detect AI-generated histological data</em>":</p> <ul> <li>Dreambooth parameters (<em>dreambooth_parameters.zip</em>)</li> <li>Images displayed during the study <em>(images.zip)</em></li> <li>Data collected during the survey (<em>results-survey.xlsx</em>)</li> <li>R Code to reproduce results and figures (<em>analysis_code.zip</em>)</li> </ul> <p>Please find our associated work here:</p> <p>Hartung, J., Reuter, S., Kulow, V.A., F&auml;hling, M., Spreckelsen, C., and Mrowka, R. (2024). Experts fail to reliably detect AI-generated histological data. Sci Rep&nbsp;<em>14</em>, 28677. https://doi.org/10.1038/s41598-024-73913-8.</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Evaluating AI-generated Research Plans - Appendix

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo32/100

Considerations on baseline generation for Imaging AI studies illustrated on the CT-based prediction of empyema and outcome assessment.

<p><strong>Considerations on baseline generation for Imaging AI studies illustrated on the CT-based prediction of empyema and outcome assessment.</strong></p> <p><strong>Introduction</strong>: For AI-based classification tasks in computed tomography, a reference standard for evaluating the clinical diagnostic accuracy of individual classes is essential. To enable the implementation of an AI tool in clinical practice, this should be drawn from clinical routine data, using State-of-the-art scanners, evaluated in a blinded manner, and verified with a reference test.</p> <p><strong>Methods:&nbsp;</strong>2659 consecutive CTs performed between 01/2016 and 01/2021 with reported pleural effusion were retrospectively included. Pathology reports from thoracocentesis or biopsy within 7 days of CT were used as reference standard (n = 335). Two radiologists (4 and 10 PGY) blindly assessed chest CTs (n=335, 81 empyemas) for pleural CT features and ICC was determined. In addition, both pleural CT features and radiological diagnosis were extracted from written radiological reports. If needed, consensus was achieved using an experienced radiologist&#39;s opinion (29 PGY). We assessed the correlation of these findings with the following patient outcomes: mortality and median hospital stay.</p> <p><strong>Results:&nbsp;</strong>Specificity and sensitivity for clinical detection of empyema (N=81) were 90.94 (95%-CI 86.55-94.05) and 72.84 (95%-CI: 61.63-81.85%) in all effusions, with moderate to almost perfect interrater agreement for all pleural findings associated with empyema (Cohen&#39;s kappa = 0.41-0.82). Features describing pleural enhancement or thickening achieved the highest accuracy with 87.02% and 81.49%, respectively. Empyema was associated with a longer hospital stay (median= 20 versus 14 days), and findings consistent with pleural carcinosis impacted mortality.&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

Dataset of AI-generated code created by various versions of GPT model

<p>This is the dataset used for the paper "<span>Human vs AI: Investigation of Security Risks in AI-generated </span><span>Code via Comparison with Human-written Code".</span></p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Intersectional Analysis of Visual Generative AI

<p>This data set contains the set of 180 images we created and analysed towards creating an intersectional STS analysis of Stable Diffusion.</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Paper samples for the SLR "A systematic literature review on the impact of AI models on the security of code generation"

<p>Here we provide the whole list of papers that were queried for the SLR "A systematic literature review on the impact of AI models on the security of code generation" by Negri-Ribalta et al. The dataset provides all the information of all the papers gathered, their database of origin, and if it was accepted/rejected/duplicated.&nbsp;</p> <p>The file is in xls format .</p>

opencc-by-4.0Feb 2024View details →
dryad32/100

Generative AI enhances individual creativity but reduces the collective diversity of novel content

<p>Creativity is core to being human. Generative AI—made readily available by powerful large language models (LLMs)—holds promise for humans to be more creative by offering new ideas, or less creative by anchoring on generative AI ideas. We study the causal impact of generative AI ideas on the production of short stories in an online experiment where some writers obtained story ideas from an LLM. We find that access to generative AI ideas causes stories to be evaluated as more creative, better written, and more enjoyable, especially among less creative writers. However, generative AI-enabled stories are more similar to each other than stories by humans alone. These results point to an increase in individual creativity at the risk of losing collective novelty. This dynamic resembles a social dilemma: with generative AI, writers are individually better off, but collectively a narrower scope of novel content is produced. Our results have implications for researchers, policy-makers, and practitioners interested in bolstering creativity.</p>

opencc-zeroJun 2024View details →
zenodo32/100

AI and Political Marketing: How Indonesian React to The Presidential Candidate AI-Generated Advertisement

<p><span>The integration of Artificial Intelligence (AI) into political campaigns, particularly through social media, has revolutionized voter behavior influencing mechanisms. This research explores the impact of AI-generated content on public decision-making in the context of the 2024 Indonesian presidential elections, focusing on the Prabowo-Gibran pair. Drawing upon selective exposure theory, the study investigates how perceived source slant, perceived source bias, and perceived effects of opponents' news diet influence political participation likelihood. Through a quantitative survey of 150 Indonesian citizens active on social media, the study reveals significant positive correlations between perceived source slant, bias, and political participation likelihood. The findings underscore the pivotal role of AI-generated content in shaping political engagement and suggest implications for enhancing voter mobilization strategies in the digital age. </span></p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Exploration of humor in generative AI

<p>Twelve conversations with different open-source LLM-based chatbots (Cohere4AI, Gemma, Llama, Mistral) exploring humor understanding and production with zero-shot and n-shot strategies. Conversations were conducted via HuggingFace chat in Spring 2024, exported as HTML pages and bundled in a zip file.</p>

opencc-by-sa-4.0Jul 2024View details →
zenodo32/100

Automated rationale generation: a technique for explainable AI and its effects on human perceptions (Dataset)

<p>Explainable AI Dataset for paper published in the proceedings of IUI 2019 titled, &quot;<strong>Automated rationale generation: a technique for explainable AI and its effects on human perceptions</strong>&quot;. Consists of data collected from human participants via Mechanical Turk.<br> <strong>Abstract</strong>:<br> <em>Automated rationale generation</em> is an approach for real-time explanation generation whereby a computational model learns to translate an autonomous agent&#39;s internal state and action data representations into natural language. Training on human explanation data can enable agents to learn to generate human-like explanations for their behavior. In this paper, using the context of an agent that plays <em>Frogger</em>, we describe (a) how to collect a corpus of explanations, (b) how to train a neural rationale generator to produce different styles of rationales, and (c) how people perceive these rationales. We conducted two user studies. The first study establishes the plausibility of each type of generated rationale and situates their user perceptions along the dimensions of <em>confidence, humanlike-ness, adequate justification, and understandability.</em> The second study further explores user preferences between the generated rationales with regard to <em>confidence</em> in the autonomous agent, communicating <em>failure and unexpected behavior.</em> Overall, we find alignment between the intended differences in features of the generated rationales and the perceived differences by users. Moreover, context permitting, participants preferred detailed rationales to form a stable mental model of the agent&#39;s behavior.</p>

opencc-by-4.0Mar 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record