Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

75

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

75 results for “pre-training”

Learn how ShareScore rates datasets ↗
zenodo32/100

Replication Package for "Cracks in The Stack: Hidden Vulnerabilities and Licensing Risks in LLM Pre-Training Datasets"

<p>Replication Package for "Cracks in The Stack: Hidden Vulnerabilities and Licensing Risks in LLM Pre-Training Datasets"<br><br>Includes datasets, and bash scripts.</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Leveraging Search-Based and Pre-Trained Code Language Models for Automated Program Repair

<p>This page serves as supplementary material for the article: <strong>Leveraging Search-Based and Pre-Trained Code Language Models for Automated Program Repair</strong>. Here, we provide the ARJACLM code utilized in the study, enabling other researchers to replicate the experiments and further develop the tool.&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

CofeXHug: A curated dataset of HuggingFace pre-trained models exploited in the GitHub ecosystem

<p>&nbsp;Pre-trained models (PTMs) are becoming increasingly popular in the software engineering community. Their usage is facilitated by model repositories, e.g., HuggingFace, which collect, store, and maintain a wide range of PTMs. However, the actual adoption of these models in real-world projects is still an open question. In particular, many of them are used in toy projects or simply as a mirror for the HF repository. Thus, we see the need for a curated codebase related to PTMs to support developers and practitioners who are interested in using them in their projects.<br>This artifact contains CodeXHug, a curated dataset of HuggingFace PTMs exploited in the GitHub ecosystem. Starting from the latest HF dump, we first conduct a data curation to collect PTMs with a tag and a model card. Then, the GitHub platform has been queried to find actual usages of the identified PTMs, resulting in 7,325 different models and 372,063 Python files. We also present a statistical analysis of the dataset, highlighting the most popular PTMs and the most common tasks for which they are used. Finally, we discuss the research opportunities enabled by CodeXHug and the implications of our findings for the software engineering community.</p>

opencc-by-4.0Dec 2024View details →
zenodo32/100

Replication Package for "Impact of ML Optimization Tactics on Greener Pre-Trained ML Models"

<p>This repository contains the replication package for the paper titled "Impact of ML Optimization Tactics on Greener Pre-Trained ML Models". The README file and accompanying scripts provide detailed instructions to facilitate the replication of the analysis.</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Pre-trained DNN model data for pruning example code

<p>Pre-trained DNN model datasets for example codes of&nbsp;neural network pruning.</p> <p>Example pruning codes are published in &quot;https://github.com/FujitsuLaboratories/CAC/tree/main/cac/pruning&quot;.</p>

opencc-zeroNov 2021View details →
zenodo32/100

Pre-trained neural network model for pruning example code

<p>Pre-trained neural network&nbsp;models for example codes of&nbsp;neural network pruning.</p> <p>Example pruning codes are published in &quot;https://github.com/FujitsuResearch/automatic_pruning&quot;.</p>

opencc-zeroJan 2022View details →
zenodo32/100

Automating Code-Related Tasks Through Transformers: The Impact of Pre-training

<p>Datasets used in the work &quot;Automating Code-Related Tasks Through Transformers: The Impact of Pre-training&quot;.</p>

opencc-by-4.0Sep 2022View details →
zenodo32/100

Windy events detection in big bioacoustics datasets using a pre-trained Convolutional Neural Network

<p>This repository includes the code and all relevant files used throughout our study. These encompass everything from the initial sheets of the whole acoustic dataset utilised for selecting the annotated dataset to the notebook (.ipynb) and the recordings employed in training the model.</p> <p>The .wav file are compressed in the file: <a href="../api/records/11220741/draft/files/wind-noise-detection-main.7z/content" target="_blank" rel="noopener noreferrer">wind-noise-detection-main.7z</a>/data/208_file_recordings_paper_wind.tar</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo32/100

Replication Package for "Classifying Open-Source Pre-Trained Models and Datasets for Software Engineering"

<p>The replication package for the short paper titled 'Classifying Open-Source Pre-Trained Models and Datasets for Software Engineering' is provided. It includes a README file and accompanying scripts with comprehensive instructions to facilitate the replication of the analysis presented in the paper.</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

U-TAE pre-trained weights on PASTIS for Semantic segmentation

<p>Pre-trained weights of U-TAE for <strong>semantic segmentation.</strong></p> <p>See&nbsp;<a href="https://github.com/VSainteuf/utae-paps">companion GitHub repository</a>&nbsp;and&nbsp;<a href="https://arxiv.org/abs/2107.07933">paper</a>&nbsp;for more information.</p>

opencc-by-4.0Aug 2021View details →
zenodo32/100

Bridging Pre-trained Models and Downstream Tasks for Source Code Understanding

<p>Datasets for the paper &quot;Bridging Pre-trained Models and Downstream Tasks for Source Code Understanding&quot;</p>

opencc-by-4.0Sep 2021View details →
zenodo32/100

UVCGANv2 Pre-Trained Models

<p>Pre-trained models of the &quot;Rethinking CycleGAN: Improving Quality of GANs for Unpaired Image-to-Image Translation&quot; paper.</p> <p>&nbsp;</p> <p>Files ending with `_full.zip` contain full models, including states of the optimizers, schedulers, discriminators, etc.</p> <p>Filed ending with `_only_gen.zip` contain only the translation generators.</p>

opencc-by-4.0Apr 2023View details →
zenodo32/100

Pre-trained Model and Relation Extraction Dataset

<p>NIPA</p>

opencc-by-4.0Apr 2023View details →
zenodo32/100

The dataset of ESEC/FSE 2023 paper titled "CCT5: A Code-Change-Oriented Pre-Trained Model"

<p>This is the dataset of&nbsp;ESEC/FSE 2023 paper titled&nbsp;<strong>CCT5: A Code-Change-Oriented Pre-Trained Model.</strong></p> <p>Please refer to this <a href="https://github.com/Ringbo/CCT5">link</a> for a detailed implementation guide.</p>

opencc-by-4.0May 2023View details →
zenodo32/100

Pre-trained weights for Protein Workshop

<p>Pre-trained model weights for ProteinWorkshop benchmark.</p>

opencc-by-4.0Aug 2023View details →
zenodo28/100

OrganoIDNetData: A Curated Cell Life Imaging Dataset of Immune-enriched Pancreatic Cancer Organoids with Pre-trained AI Models

<p>OrganoIDNetData encompasses 180 images with 34113 organoids of human and murine Pancreatic Ductal Adenocarcinoma co-cultured with immune cells.</p> <p><strong>For further information and to cite this dataset, please refer to:</strong></p> <p>Kulkarni, A., Ferreira, N., Scodellaro, R.&nbsp;<em>et al.</em>&nbsp;A Curated Cell Life Imaging Dataset of Immune-enriched Pancreatic Cancer Organoids with Pre-trained AI Models.&nbsp;<em>Sci Data</em>&nbsp;<strong>11</strong>, 820 (2024). https://doi.org/10.1038/s41597-024-03631-3</p>

opencc-by-4.0Feb 2024View details →
zenodo28/100

Pre-trained model of DAR-MVSNet

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo28/100

Data for ICSE 2024 paper "Learning in the Wild: Towards Leveraging Unlabeled Data for Effectively Tuning Pre-trained Code Models".

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo28/100

dMRI-RCNN Pre-Trained Weights - 1D b1000 30q

<p>Pre-trained weights for&nbsp;<a href="http://github.com/m-lyon/dMRI-RCNN">dMRI-RCNN</a>. 1D RCNN model, for q_in = 30,&nbsp;b = 1000.</p>

opencc-by-4.0Mar 2022View details →
zenodo28/100

dMRI-RCNN Pre-Trained Weights - 3D b2000 10q

<p>Pre-trained weights for&nbsp;<a href="http://github.com/m-lyon/dMRI-RCNN">dMRI-RCNN</a>. 3D RCNN model, for q_in = 10,&nbsp;b = 2000.</p>

opencc-by-4.0Mar 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record