Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,025
datasets available to search
ShareScore release 0.9.0
Dataset results
2,025 results for “AIS”
Does the tail show when the nose knows? AI-enhanced analysis of tail kinematics outperforms human experts at predicting when detection dogs find their target odor
Open the record for dataset details and reuse information.
Durably reducing conspiracy beliefs through dialogues with AI
Open the record for dataset details and reuse information.
Nationwide real-world implementation of AI for cancer detection in population-based mammography screening (PRAIM)
Open the record for dataset details and reuse information.
Artificial Intelligence (AI), in the Breast Cancer Screening Programme Questionnaire and Data
<p>The goal of this study is to gain insight into the level of trust of women in the Netherlands, in the decisions made by radiologists with the support of different applications of Artificial Intelligence (AI), in the Breast Cancer Screening Programme. Gaining insight into your level of trust in the decisions made by radiologists and AI regarding whether you have breast cancer or not, is of high importance, as this will help to anticipate which steps can be taken by hospitals and software developers, in the near future. The survey consists of 10 introductory questions and 42 statements and takes approximately 10 minutes to complete.</p> <p> </p>
Training dataset and final models used in "Digging roots is easier with AI"
<p>See the paper ''Digging roots is easier with AI' for an explanation of how the models and datasets were created.</p> <p>The images are extracted from the larger datasets available from 10.5281/zenodo.4300067 using the RootPainter software.</p>
Dataset and manual counts used in "Digging roots is easier with AI"
<p>Please see the paper "Digging roots is easier with AI" how the images were captured and the manual counting was performed for the three destructive sampling procedures. </p>
AI results complementing the Annual Report on surveillance for Avian Influenza in poultry and wild birds in Member States of the European Union - Denmark
<p>This dataset contains the results of the EU co-funded surveillance activities conducted in 2019, which consisted of:</p> <ul> <li>Serological surveys to monitor the circulation of AIV subtypes H5 and H7 in poultry (active surveillance). These surveys should preferentially target poultry species or production systems with increased risk for introduction of avian influenza (AI).</li> <li>Passive surveillance aiming at the virological detection of AI in wild birds found dead or moribund.</li> </ul>
AI results complementing the Annual Report on surveillance for Avian Influenza in poultry and wild birds in Member States of the European Union - Austria
<p>This dataset contains the results of the EU co-funded surveillance activities conducted in 2019, which consisted of:</p> <ul> <li>Serological surveys to monitor the circulation of AIV subtypes H5 and H7 in poultry (active surveillance). These surveys should preferentially target poultry species or production systems with increased risk for introduction of avian influenza (AI).</li> <li>Passive surveillance aiming at the virological detection of AI in wild birds found dead or moribund.</li> </ul>
Characterizing Technical Debt and Antipatterns in AI-Based Systems: A Systematic Mapping Study
<p>All artifacts related to a systematic mapping study on technical debt and antipatterns in AI-based systems; the data consists of an Excel file (00-all-data.xlsx) which includes a tab for all important constructs plus a separate CSV file per construct:</p> <ul> <li>List of primary studies (01-primary-studies.csv)</li> <li>Established types of technical debt (02-established-td-types.csv)</li> <li>New types of technical debt (03-new-td-types.csv)</li> <li>Affected software quality attributes (04-affected-qas.csv)</li> <li>Identified antipatterns (05-antipatterns.csv)</li> <li>Reported solutions to address technical debt or antipatterns (06-solutions.csv)</li> </ul> <p> </p>
Extract of Local Notice to Mariners, 2018-2019, for Development of AIS Model of Texas Gulf Intracoastal Waterway Travel Times
<p>These files summarize mentions of restrictions or cautions for navigation on the Texas Gulf Intracoastal Waterway contained in Coast Guard Local Notice to Mariners files.</p>
Weekly travel times by direction and sample count for each link in Development of AIS Model of Texas Gulf Intracoastal Waterway Travel Times
<p>Excel spreadsheet containing all of the travel times and sample counts for each link (by direction).</p>
Generative AI in University Communication - Survey Data (June 2023)
<p>Der Datensatz mit dem Titel "Generative KI in der Hochschulkommunikation - Umfragedaten (Juni 2023)" erfasst Informationen zur Einführung und Nutzung von generativer künstlicher Intelligenz (KI) im Kontext der Hochschulkommunikation. Die Umfrage, die im Juni 2023 unter 318 deutschen Hochschulen durchgeführt wurde, von denen 101 geantwortet haben, untersucht verschiedene Aspekte, darunter Bekanntheit und Wissen über verschiedene KI-Tools (z.B. ChatGPT), Diskussionen in Gremien, das Vorhandensein von Richtlinien für die Nutzung, das Vorhandensein von Arbeitsgruppen für generative KI, strategische Ziele und Initiativen, Schulungsangebote für generative KI-Tools und die wahrgenommene Bedeutung von generativen KI-Tools in der Hochschulkommunikation. Ziel des Datensatzes ist es, Einblicke in die aktuelle Landschaft und Praxis der Integration generativer KI im universitären Umfeld zu geben.</p><p>The dataset, titled "Generative AI in University Communication - Survey Data (June 2023)," captures information related to the adoption and utilization of generative artificial intelligence (AI) in the context of university communication. This survey, conducted in June 2023 among 318 German universities of which 101 responded, explores various aspects, including awarenes and knowledge of various AI tools (e.g. ChatGPT), discussions in committees, the existence of guidelines for usage, the presence of working groups for generative AI, strategic goals and initiatives, training offerings for generative AI tools, and the perceived importance of generative AI tools in university communication. The dataset aims to provide insights into the current landscape and practices regarding the integration of generative AI within university settings.</p><p> </p><p>More information here: <a href="https://www.hof.uni-halle.de/projekte/hochki/">https://www.hof.uni-halle.de/projekte/hochki/</a></p>
Monumento ai Caduti
Situato nei pressi della cartiera di Pale nel comune di Foligno. Il monumento è stato eretto nel 1919, con l'elenco dei Caduti nel primo conflitto mondiale, ha visto successivamente l'aggiunta dei nomi dei caduti della seconda guerra. Iscrizione sul basamento: PALE UNANIME AI SUOI FIGLI CADUTI PER L'ONORE E LA GLORIA D'ITALIA III-XI-MCMXIX Source: Objaverse 1.0 / Sketchfab
Can We Trust AI Agents? - Supplementary Material
<p>This is the supplementary material of our paper entitled Can We Trust AI Agents? An Experimental Study Towards Trustworthy LLM-Based Multi-Agent Systems for AI Ethics. Published in TETHICS 2025, Conference on Technology Ethics, Vaasa, Finland, 11-12 November 2025.</p>
Datasets, models and demos associated to "Celldetective: an AI-enhanced image analysis tool for unraveling dynamic cell interactions"
<p>This repository contains datasets, models and demos associated to <a href="https://github.com/remyeltorro/celldetective">Celldetective</a>, a software for single-cell analysis from multimodal time lapse microscopy images. </p> <h1>Demos</h1> <h2>Cell-cell interaction assay: ADCC</h2> <p>We imaged a co-culture of MCF-7 breast cancer cells (targets) and human primary NK cells (effectors), interacting in the presence of bispecific antibodies, to measure antibody dependent cellular cytotoxicity (ADCC). The nuclei of all cells are marked with the Hoechst nuclear stain, the dead nuclei with the propidium iodide nuclear stain, the cytoplasm of the NK cells with CFSE. The system in epifluorescence and brightfield at either 20 or 40X magnification. We provide a single position demo for the ADCC assay, as "demo_adcc.zip". After unzipping, the demo_adcc folder can be loaded in Celldetective for testing. </p> <h2>Cell-surface interaction assay: RICM</h2> <p>We imaged human primary NK cells engaging in spreading with a surface coated with a bispecific antibody similar to the one used in the ADCC assay (replacing the target cells with a flat surface). The system is imaged using the RICM technique. Images are normalized using a median estimate of the background, pooled from all the positions in a well and dividing the images by this estimate. Here, we provide a single position demo for the cell-surface interactiona assay imaged in RICM, as "demo_ricm.zip". As above, after unzipping, the experiment can be tested and processed in Celldetective.</p> <h1>Datasets</h1> <h2>Image annotations for segmentation</h2> <h3>Cell-cell interaction assay: ADCC</h3> <p>We generated two sets of annotations from images of a co-culture of MCF-7 breast cancer cells and human primary NK cells, interacting in the presence of bispecific antibodies, to measure antibody dependent cellular cytotoxicity (ADCC). Since there are two separate cell populations of interest, the targets (MCF-7) and effectors (NK cells), we curated two datasets. Each sample in a dataset consists of a multichannel image (up to five channels in the context of ADCC, among brightfield , Hoechst nuclear stain, PI nuclear stain, CFSE, LAMP1), the associated instance segmentation annotation for the population of interest and a json file summarizing the content of each channel and the spatial calibration of the image. These sample data are generated directly in Celldetective, using a custom napari plugin.</p> <ul> <li>db_mcf7_nuclei_w_lymphocytes: MCF-7 cell nuclei are annotated specifically on images where primary NK cells (or rarely primary T cells), and RBCs co-exist. The annotation exploits up to four channels simultaneously.</li> <li>db_primary_NK_w_mcf7: human primary NK cells, with annotated cytoplasm (mostly from CFSE) but exploiting brightfield and Hoechst to segment out of focus or poorly labelled cells.</li> </ul> <p>These datasets are used to train several segmentation models to segment on one hand the MCF-7 nuclei and on the other hand the primary NK cells.</p> <h3>Cell-surface interaction assay: RICM</h3> <ul> <li>db_spreading_lymphocytes: we provide a dataset of primary NK cells (and occasionnaly mice T cells) imaged in RICM (with sometimes paired brightfield images). Cells are detected as soon as they start forming interferences on the image (hovering behavior). A pre-annotation was performed using a threshold based segmentation on the RICM modality. Manuel separation of cell-cell contacts and removal of false positive objects was performed by an expert annotator (using brightfield when available). RBCs are ignored in the annotations. </li> </ul> <h2>Single-cell signal annotations for classification and regression</h2> <h3>Cell-cell interaction assay: ADCC</h3> <p>We generated several signal classification/regression datasets with Celldetective to characterize the ADCC assay. Briefly, for a given event cells can be classified as "the event occured during the observation", "no event occured during the observation", "the event already occured prior to observation". If the event occurred during the observation, we can estimate when (the regression). Each single-cell is a dictionary with a collection of signals. The attribute "class" sets the class and "t0" the time of event (default is -1 for absence of event). </p> <ul> <li>db-si-NucPI: classification and regression of single-cells with respect to lysis events characterized by a strong PI increase upon lysis (also associated with decreasing nuclear area and sometimes a decreasing Hoechst)</li> <li>db-si-NucCondensation: classification and regression of single-cells with respect to nucleus shrinking events characterized by a decreasing nuclear area (UPDATE on 23/01/2024)</li> </ul> <h1>Models</h1> <h2>Segmentation models</h2> <h3>Generalist models</h3> <p>We integrated in Celldetective select published models for cellular segmentation from StarDist and Cellpose. We wraped the models with an input configuration to help Celldetective handle the normalization, rescaling and channel selection upon inference. </p> <ul> <li>Cellpose [1,2]: <em>cyto3</em>, <em>livecell</em>, <em>tissuenet</em>, <em>nuclei</em></li> <li>StarDist [3]: <em>versatile_fluo</em>, <em>versatile_he</em></li> </ul> <p>If you use any of these models your research, don't forget to cite the StarDist or Cellpose papers accordingly!</p> <h3>ADCC models</h3> <ul> <li>MCF-7 (in the presence of lymphocytes): <em>mcf7_nuc_multimodal, mcf7_nuc_stardist_transfer</em></li> <li>primary NKs (in the presence of MCF-7): <em>primNK_multimodal</em>, <em>primNK_SD</em>, <em>primNK_cfse</em></li> </ul> <h3>Spreading-assay models</h3> <ul> <li>Lymphocytes: <em>lymphocytes_ricm</em></li> </ul> <h2>Signal analysis models</h2> <p>We developed Deep Learning models that classify and regress the time of events from single-cell signals, applied to the ADCC assay.</p> <ul> <li> lysis detection: <em>lysis_H_PI</em>, <em>lysis_PI_area</em><em>. </em>Detect lysis events characterized at least by an increase of PI from one or more measurements (respectively PI+Hoechst and PI+nucleus area, trained on db-si-NucPI)</li> <li>nucleus shrinking detection:<em> NucCond</em>. Detect nucleus shrinking events from nuclear area signal (db-si-NucCondensation)</li> </ul> <h1>References</h1> <ol> <li>Stringer, C., Wang, T., Michaelos, M. & Pachitariu, M. Cellpose: a generalist algorithm for cellular segmentation. Nat Methods 18, 100–106 (2021).</li> <li>Pachitariu, M. & Stringer, C. Cellpose 2.0: how to train your own model. Nat Methods 19, 1634–1641 (2022).</li> <li>Schmidt, U., Weigert, M., Broaddus, C. & Myers, G. Cell Detection with Star-Convex Polygons. in Medical Image Computing and Computer Assisted Intervention – MICCAI 2018 (eds. Frangi, A. F., Schnabel, J. A., Davatzikos, C., Alberola-López, C. & Fichtinger, G.) 265–273 (Springer International Publishing, Cham, 2018). doi:10.1007/978-3-030-00934-2_30.</li> </ol> <p> </p> <p> </p>
Acceptance of medical AI in skin cancer screening: A Choice-based Conjoint Survey
<p><strong>Background</strong>: There is a great interest in using artificial intelligence (AI) to screen for skin cancer. This is fueled by a rising incidence of skin cancer and an increasing scarcity of trained dermatologists. AI systems, capable of identifying melanoma, could save lives, enable immediate access to screenings, reduce unnecessary care and healthcare costs. While such AI-based systems are useful from a public health perspective, past research has shown that individual patients are very hesitant about being examined by an AI system. <strong>Objective</strong>: The aim of the present study was twofold. First, to determine how important the attributes provider (in-person physician, physician via teledermatology, AI, vs. personalized AI), costs of screening (free, 10€, 25€, vs. 40€) and waiting time (immediate, 1 day, 1 week, 4 weeks) were for patients’ choices of a particular mode of skin cancer screening. Second, to investigate whether sociodemographic characteristics, especially, age, were systematically related to participants’ individual choices. <strong>Methods</strong>: The study used choice-based conjoint-analysis to examine the acceptance of medical AI for a skin cancer screening from the patient's perspective. Participants responded to twelve choice sets, each containing three screening-variants, where each variant was described through attributes; provider, costs and waiting time. Furthermore, sociodemographic characteristics (age, gender, income, job status, educational background) were assessed. <strong>Results</strong>: 126 (33%) respondents completed the online survey. The results from the conjoint analysis showed that the three attributes were more or less equal important for the participant’s choices, with provider being the most important. Inspecting the individual part worths showed that treatment by a physician was most preferred, followed by e-consultation with a physician and personalized AI. The three AI levels scored significantly lower. Concerning the relationship between sociodemographic characteristics and relative importances we found, that only age showed a significant positive association to the important of the attribute provider (r = 0.21; p < .02). Younger participants put a lesser importance on the provider than older participants. All other correlations were not significant. <strong>Conclusions</strong>: The present study adds to the growing body of research using choice-experiments to investigate the acceptance of artificial intelligence in health contexts. Future studies need to explore the reasons <em>why</em> AI is accepted or rejected and whether sociodemographic characteristics are associated this decision.</p> <p> </p>
EfficientBioAI: Making Bioimaging AI Models Efficient in Energy and Latency
<p>This dataset contains trained deep learning models, dataset and experiment files for the manuscript "EfficientBioAI: Making Bioimaging AI Models Efficient in Energy and Latency". Please find the software and more information including tutorials here: <a href="https://github.com/MMV-Lab/EfficientBioAI">MMV-Lab/EfficientBioAI (github.com)</a>.</p>
Radial velocities and broadening functions for AI Hya
<p>Radial velocity and broadening function measurements from CORALIE and HIDES spectra.</p>
Reproducibility of "Diffusion-based Generative AI for Exploring Transition States from 2D Molecular Graphs"
<p>This file is the source data to ensure reproducibility of the paper "Diffusion-based Generative AI for Exploring Transition States from 2D Molecular Graphs". It contains the logs and results of all DFT calculations associated with transition states generated using the model proposed in the paper. It also includes code to reproduce the core findings of the paper, which can be done by running reproduce.sh. To accurately reproduce the results of the paper, use the v1.0.0 virtual environment from "https://github.com/seonghann/tsdiff".</p>
ADMET-AI: A machine learning ADMET platform for evaluation of large-scale chemical libraries – Data and Models
<p>This repository contains data and models used in the following paper.</p> <p> </p> <p>Swanson, K., Walther, P., Leitz, J., Mukherjee, S., Wu, J. C., Shivnaraine, R. V., & Zou, J. ADMET-AI: A machine learning ADMET platform for evaluation of large-scale chemical libraries. In review.</p> <p> </p> <p>The data and models are meant to be used with the <a href="https://github.com/swansonk14/admet_ai">ADMET-AI</a> code, which runs the ADMET-AI web server at <a href="https://admet.ai.greenstonebio.com/">admet.ai.greenstonebio.com</a>.</p> <p> </p> <p>The data.zip file has the following structure.</p> <p>data</p> <p> drugbank: Contains files with drugs from the <a href="https://go.drugbank.com/">DrugBank</a> that have received regulatory approval. drugbank_approved.csv contains the full set of approved drugs along with ADMET-AI predictions, while the other files contain subsets of these molecules used for testing the speed of ADMET prediction tools.</p> <p> tdc_admet_all: Contains the data (.csv files) and RDKit features (.npz files) for all 41 single-task ADMET datasets from the <a href="https://tdcommons.ai/">Therapeutics Data Commons</a> (TDC).</p> <p> tdc_admet_multitask: Contains the data (.csv files) and RDKit features (.npz files) for the two multi-task datasets (one regression and one classification) constructed by combining the tdc_admet_all datasets.</p> <p> tdc_admet_all.csv: A CSV file containing all 41 ADMET datasets from tdc_admet_all. This can be used to easily look up all ADMET properties for a given molecule in the TDC.</p> <p> tdc_admet_group: Contains the data (.csv files) and RDKit features (.npz files) for the 22 TDC ADMET Benchmark Group datasets with five splits per dataset.</p> <p> tdc_admet_group_raw: Contains the raw data (.csv files) used to construct the five splits per dataset in tdc_admet_group.</p> <p> </p> <p>The models.zip file has the following structure. Note that the ADMET-AI website and Python package use the multi-task Chemprop-RDKit models below.</p> <p>models</p> <p> tdc_admet_all: Contains Chemprop and Chemprop-RDKit models trained on all 41 single-task TDC ADMET datasets.</p> <p> tdc_admet_all_multitask: Contains Chemprop and Chemprop-RDKit models trained on the two multi-task TDC ADMET datasets (one regression and one classification).</p> <p> tdc_admet_group: Contains Chemprop and Chemprop-RDKit models trained on the 22 TDC ADMET Benchmark Group datasets.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.