Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
43
datasets available to search
ShareScore release 0.9.0
Dataset results
43 results for “Pre-trained models”
Leveraging Search-Based and Pre-Trained Code Language Models for Automated Program Repair
<p>This page serves as supplementary material for the article: <strong>Leveraging Search-Based and Pre-Trained Code Language Models for Automated Program Repair</strong>. Here, we provide the ARJACLM code utilized in the study, enabling other researchers to replicate the experiments and further develop the tool. </p>
CofeXHug: A curated dataset of HuggingFace pre-trained models exploited in the GitHub ecosystem
<p> Pre-trained models (PTMs) are becoming increasingly popular in the software engineering community. Their usage is facilitated by model repositories, e.g., HuggingFace, which collect, store, and maintain a wide range of PTMs. However, the actual adoption of these models in real-world projects is still an open question. In particular, many of them are used in toy projects or simply as a mirror for the HF repository. Thus, we see the need for a curated codebase related to PTMs to support developers and practitioners who are interested in using them in their projects.<br>This artifact contains CodeXHug, a curated dataset of HuggingFace PTMs exploited in the GitHub ecosystem. Starting from the latest HF dump, we first conduct a data curation to collect PTMs with a tag and a model card. Then, the GitHub platform has been queried to find actual usages of the identified PTMs, resulting in 7,325 different models and 372,063 Python files. We also present a statistical analysis of the dataset, highlighting the most popular PTMs and the most common tasks for which they are used. Finally, we discuss the research opportunities enabled by CodeXHug and the implications of our findings for the software engineering community.</p>
Replication Package for "Impact of ML Optimization Tactics on Greener Pre-Trained ML Models"
<p>This repository contains the replication package for the paper titled "Impact of ML Optimization Tactics on Greener Pre-Trained ML Models". The README file and accompanying scripts provide detailed instructions to facilitate the replication of the analysis.</p>
Pre-trained DNN model data for pruning example code
<p>Pre-trained DNN model datasets for example codes of neural network pruning.</p> <p>Example pruning codes are published in "https://github.com/FujitsuLaboratories/CAC/tree/main/cac/pruning".</p>
Pre-trained neural network model for pruning example code
<p>Pre-trained neural network models for example codes of neural network pruning.</p> <p>Example pruning codes are published in "https://github.com/FujitsuResearch/automatic_pruning".</p>
Replication Package for "Classifying Open-Source Pre-Trained Models and Datasets for Software Engineering"
<p>The replication package for the short paper titled 'Classifying Open-Source Pre-Trained Models and Datasets for Software Engineering' is provided. It includes a README file and accompanying scripts with comprehensive instructions to facilitate the replication of the analysis presented in the paper.</p>
Bridging Pre-trained Models and Downstream Tasks for Source Code Understanding
<p>Datasets for the paper "Bridging Pre-trained Models and Downstream Tasks for Source Code Understanding"</p>
UVCGANv2 Pre-Trained Models
<p>Pre-trained models of the "Rethinking CycleGAN: Improving Quality of GANs for Unpaired Image-to-Image Translation" paper.</p> <p> </p> <p>Files ending with `_full.zip` contain full models, including states of the optimizers, schedulers, discriminators, etc.</p> <p>Filed ending with `_only_gen.zip` contain only the translation generators.</p>
Pre-trained Model and Relation Extraction Dataset
<p>NIPA</p>
The dataset of ESEC/FSE 2023 paper titled "CCT5: A Code-Change-Oriented Pre-Trained Model"
<p>This is the dataset of ESEC/FSE 2023 paper titled <strong>CCT5: A Code-Change-Oriented Pre-Trained Model.</strong></p> <p>Please refer to this <a href="https://github.com/Ringbo/CCT5">link</a> for a detailed implementation guide.</p>
OrganoIDNetData: A Curated Cell Life Imaging Dataset of Immune-enriched Pancreatic Cancer Organoids with Pre-trained AI Models
<p>OrganoIDNetData encompasses 180 images with 34113 organoids of human and murine Pancreatic Ductal Adenocarcinoma co-cultured with immune cells.</p> <p><strong>For further information and to cite this dataset, please refer to:</strong></p> <p>Kulkarni, A., Ferreira, N., Scodellaro, R. <em>et al.</em> A Curated Cell Life Imaging Dataset of Immune-enriched Pancreatic Cancer Organoids with Pre-trained AI Models. <em>Sci Data</em> <strong>11</strong>, 820 (2024). https://doi.org/10.1038/s41597-024-03631-3</p>
Pre-trained model of DAR-MVSNet
Open the record for dataset details and reuse information.
Data for ICSE 2024 paper "Learning in the Wild: Towards Leveraging Unlabeled Data for Effectively Tuning Pre-trained Code Models".
Open the record for dataset details and reuse information.
Pre-trained model of SPC-MVSNet.
<p>Pre-trained model of SPC-MVSNet.</p>
Classifying Open-Source Pre-Trained Models and Datasets for Software Engineering
Open the record for dataset details and reuse information.
Dataset and pre-trained models for TopoMatch
<p>This contains datasets and pre-trained models for TopoMatch</p>
Automated Program Repair in the Era of Large Pre-trained Language Models
<p>Code used for the paper along with the generated outputs</p>
animal soup sample data, ground truth dataset, and pre-trained models
<p>- sample data and ground truth files for animal soup tests</p> <p>- pre-trained models</p>
SITS-Former: A pre-trained spatio-spectral-temporal representation model for Sentinel-2 time series classifcation
<p>This is the unlabeled dataset we introduced in the presented paper '<strong>SITS-Former: A pre-trained spatio-spectral-temporal representation model for Sentinel-2 time series classifcation</strong>'. This dataset can be used to pre-train a specified deep learning model (such as SITS-Former, CNN-Transformer, ConvLSTM. etc) for patch-based Sentinel-2 time series classification. </p> <p>In this dataset, each sample corresponds to an unlabeled image patch time series, which is stored as a separate numpy file named '<em>unlabeled_XXX.npz</em>'. You can use '<em>np.load</em>' to open a saved '<em>.npz</em>' file and get two arrays (querid by "ts" and "doy") from the returned dictionary. The code will be released at <em>https://github.com/linlei1214/SITS-Former</em> soon.</p>
UVCGAN Pre-Trained Models
<p>UVCGAN Pre-Trained Models</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.