Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
75
datasets available to search
ShareScore release 0.9.0
Dataset results
75 results for “pre-training”
Replication Package for "Cracks in The Stack: Hidden Vulnerabilities and Licensing Risks in LLM Pre-Training Datasets"
<p>Replication Package for "Cracks in The Stack: Hidden Vulnerabilities and Licensing Risks in LLM Pre-Training Datasets"<br><br>Includes datasets, and bash scripts.</p>
Leveraging Search-Based and Pre-Trained Code Language Models for Automated Program Repair
<p>This page serves as supplementary material for the article: <strong>Leveraging Search-Based and Pre-Trained Code Language Models for Automated Program Repair</strong>. Here, we provide the ARJACLM code utilized in the study, enabling other researchers to replicate the experiments and further develop the tool. </p>
CofeXHug: A curated dataset of HuggingFace pre-trained models exploited in the GitHub ecosystem
<p> Pre-trained models (PTMs) are becoming increasingly popular in the software engineering community. Their usage is facilitated by model repositories, e.g., HuggingFace, which collect, store, and maintain a wide range of PTMs. However, the actual adoption of these models in real-world projects is still an open question. In particular, many of them are used in toy projects or simply as a mirror for the HF repository. Thus, we see the need for a curated codebase related to PTMs to support developers and practitioners who are interested in using them in their projects.<br>This artifact contains CodeXHug, a curated dataset of HuggingFace PTMs exploited in the GitHub ecosystem. Starting from the latest HF dump, we first conduct a data curation to collect PTMs with a tag and a model card. Then, the GitHub platform has been queried to find actual usages of the identified PTMs, resulting in 7,325 different models and 372,063 Python files. We also present a statistical analysis of the dataset, highlighting the most popular PTMs and the most common tasks for which they are used. Finally, we discuss the research opportunities enabled by CodeXHug and the implications of our findings for the software engineering community.</p>
Replication Package for "Impact of ML Optimization Tactics on Greener Pre-Trained ML Models"
<p>This repository contains the replication package for the paper titled "Impact of ML Optimization Tactics on Greener Pre-Trained ML Models". The README file and accompanying scripts provide detailed instructions to facilitate the replication of the analysis.</p>
Pre-trained DNN model data for pruning example code
<p>Pre-trained DNN model datasets for example codes of neural network pruning.</p> <p>Example pruning codes are published in "https://github.com/FujitsuLaboratories/CAC/tree/main/cac/pruning".</p>
Pre-trained neural network model for pruning example code
<p>Pre-trained neural network models for example codes of neural network pruning.</p> <p>Example pruning codes are published in "https://github.com/FujitsuResearch/automatic_pruning".</p>
Automating Code-Related Tasks Through Transformers: The Impact of Pre-training
<p>Datasets used in the work "Automating Code-Related Tasks Through Transformers: The Impact of Pre-training".</p>
Windy events detection in big bioacoustics datasets using a pre-trained Convolutional Neural Network
<p>This repository includes the code and all relevant files used throughout our study. These encompass everything from the initial sheets of the whole acoustic dataset utilised for selecting the annotated dataset to the notebook (.ipynb) and the recordings employed in training the model.</p> <p>The .wav file are compressed in the file: <a href="../api/records/11220741/draft/files/wind-noise-detection-main.7z/content" target="_blank" rel="noopener noreferrer">wind-noise-detection-main.7z</a>/data/208_file_recordings_paper_wind.tar</p> <p> </p> <p> </p>
Replication Package for "Classifying Open-Source Pre-Trained Models and Datasets for Software Engineering"
<p>The replication package for the short paper titled 'Classifying Open-Source Pre-Trained Models and Datasets for Software Engineering' is provided. It includes a README file and accompanying scripts with comprehensive instructions to facilitate the replication of the analysis presented in the paper.</p>
U-TAE pre-trained weights on PASTIS for Semantic segmentation
<p>Pre-trained weights of U-TAE for <strong>semantic segmentation.</strong></p> <p>See <a href="https://github.com/VSainteuf/utae-paps">companion GitHub repository</a> and <a href="https://arxiv.org/abs/2107.07933">paper</a> for more information.</p>
Bridging Pre-trained Models and Downstream Tasks for Source Code Understanding
<p>Datasets for the paper "Bridging Pre-trained Models and Downstream Tasks for Source Code Understanding"</p>
UVCGANv2 Pre-Trained Models
<p>Pre-trained models of the "Rethinking CycleGAN: Improving Quality of GANs for Unpaired Image-to-Image Translation" paper.</p> <p> </p> <p>Files ending with `_full.zip` contain full models, including states of the optimizers, schedulers, discriminators, etc.</p> <p>Filed ending with `_only_gen.zip` contain only the translation generators.</p>
Pre-trained Model and Relation Extraction Dataset
<p>NIPA</p>
The dataset of ESEC/FSE 2023 paper titled "CCT5: A Code-Change-Oriented Pre-Trained Model"
<p>This is the dataset of ESEC/FSE 2023 paper titled <strong>CCT5: A Code-Change-Oriented Pre-Trained Model.</strong></p> <p>Please refer to this <a href="https://github.com/Ringbo/CCT5">link</a> for a detailed implementation guide.</p>
Pre-trained weights for Protein Workshop
<p>Pre-trained model weights for ProteinWorkshop benchmark.</p>
OrganoIDNetData: A Curated Cell Life Imaging Dataset of Immune-enriched Pancreatic Cancer Organoids with Pre-trained AI Models
<p>OrganoIDNetData encompasses 180 images with 34113 organoids of human and murine Pancreatic Ductal Adenocarcinoma co-cultured with immune cells.</p> <p><strong>For further information and to cite this dataset, please refer to:</strong></p> <p>Kulkarni, A., Ferreira, N., Scodellaro, R. <em>et al.</em> A Curated Cell Life Imaging Dataset of Immune-enriched Pancreatic Cancer Organoids with Pre-trained AI Models. <em>Sci Data</em> <strong>11</strong>, 820 (2024). https://doi.org/10.1038/s41597-024-03631-3</p>
Pre-trained model of DAR-MVSNet
Open the record for dataset details and reuse information.
Data for ICSE 2024 paper "Learning in the Wild: Towards Leveraging Unlabeled Data for Effectively Tuning Pre-trained Code Models".
Open the record for dataset details and reuse information.
dMRI-RCNN Pre-Trained Weights - 1D b1000 30q
<p>Pre-trained weights for <a href="http://github.com/m-lyon/dMRI-RCNN">dMRI-RCNN</a>. 1D RCNN model, for q_in = 30, b = 1000.</p>
dMRI-RCNN Pre-Trained Weights - 3D b2000 10q
<p>Pre-trained weights for <a href="http://github.com/m-lyon/dMRI-RCNN">dMRI-RCNN</a>. 3D RCNN model, for q_in = 10, b = 2000.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.