Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

75

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

75 results for “pre-training”

Learn how ShareScore rates datasets ↗
zenodo28/100

Pre-trained Dataset of mRNABERT

Open the record for dataset details and reuse information.

openapache2.0Jun 2024View details →
zenodo28/100

Pre-trained model of SPC-MVSNet.

<p>Pre-trained model of SPC-MVSNet.</p>

opencc-by-4.0May 2024View details →
zenodo28/100

Twitter pre-trained word vectors

<p>Clean up of <strong>glove.twitter.27B.zip </strong>&lt;ODC Public Domain Dedication and Licence (PDDL) 1.0&gt;.</p> <p>&quot;2B tweets, 27B tokens, 1.2M vocab, uncased&quot;</p> <p><strong>Changes from original</strong></p> <ul> <li>Headers added to allow loading by gensim. [Added via scripts.glove2word2vec]</li> <li>Recompressed as individual gzip files [Instead of a combined zip].</li> </ul> <p>These changes make the files easier to work with and increase compatibility.</p> <p><strong>Headers</strong></p> <p>Example of added header line</p> <blockquote> <p>1193513 200</p> </blockquote> <p>Header gives number of tokens and&nbsp;dimensions.</p> <p><strong>Statistics</strong></p> <ul> <li>Entries: 1,193,513</li> <li>Token length (characters). Min:1, Max:140, Avg:6.73</li> <li>Number of words per token. Min:0, Max:17, Avg:1.00669200921984</li> <li>Tokens with more than one word: 4874 (0.41%)</li> <li>Twitter data collection date: <em>Unknown.</em><strong><em> </em></strong></li> </ul> <p>History:</p> <ul> <li><strong>?? Aug 2014 </strong>&mdash; GloVe v.1.0 released</li> <li><strong>16 Aug 2014&nbsp;</strong>&mdash; Files first appear as headerless .txt.gz files, some files have mislabeled linked (via <a href="https://web.archive.org/web/20140816165523/http://www-nlp.stanford.edu/projects/glove/">wayback machine)</a></li> <li><strong>?? Oct 2015</strong> &mdash; GloVe v.1.2 released</li> <li><strong>?? ??? ????</strong> &mdash; Files replaced with a .zip file</li> <li><strong>03 June 2019 </strong>&mdash; (These files) Repackaged like original as .txt.gz, plus added headers for increased compatibility</li> </ul> <p>Example of 17-word token:</p> <blockquote> <p>&nbsp;سكس_طيز_قحبه_عنيف_اغتصاب_سكسيه_فحل_زب_نيك_بنات_مكوه_شهوه_لحس_عنف_تومبوي_ليدي_سبورت</p> </blockquote> <p><strong>200d file:</strong></p> <ul> <li>Normalized: no</li> <li>Values: &nbsp;Min:-6.7986, Max:4.609, Avg:0.009065093</li> <li>Zero values (exactly zero): none</li> <li>Zero values (approx zero) per entry: Min:0 (0.00%), Max:2 of 200 (1.0%), Avg:0.00375865197949247 (0.00%)</li> </ul>

openodc-pddlJun 2019View details →
zenodo28/100

Classifying Open-Source Pre-Trained Models and Datasets for Software Engineering

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo28/100

Dataset and pre-trained models for TopoMatch

<p>This contains datasets and&nbsp;pre-trained models for TopoMatch</p>

opencc-by-4.0Nov 2022View details →
zenodo28/100

Automated Program Repair in the Era of Large Pre-trained Language Models

<p>Code used for the paper along with the generated outputs</p>

opencc-by-4.0May 2023View details →
zenodo28/100

animal soup sample data, ground truth dataset, and pre-trained models

<p>- sample data and ground truth files for animal soup tests</p> <p>- pre-trained models</p>

opencc-by-4.0Jul 2023View details →
zenodo24/100

SITS-Former: A pre-trained spatio-spectral-temporal representation model for Sentinel-2 time series classifcation

<p>This is the unlabeled dataset we introduced in the presented paper &#39;<strong>SITS-Former: A pre-trained spatio-spectral-temporal representation model for Sentinel-2 time series classifcation</strong>&#39;. This dataset can be used to&nbsp;pre-train&nbsp;a specified deep learning model (such as&nbsp;SITS-Former, CNN-Transformer, ConvLSTM. etc) for patch-based Sentinel-2 time series classification.&nbsp;&nbsp;</p> <p>In this dataset, each sample corresponds&nbsp;to an unlabeled image patch time series, which&nbsp;is stored as a separate numpy file named &#39;<em>unlabeled_XXX.npz</em>&#39;.&nbsp;You can use &#39;<em>np.load</em>&#39;&nbsp;to open a saved &#39;<em>.npz</em>&#39; file and get two arrays (querid by &quot;ts&quot; and &quot;doy&quot;) from the returned&nbsp;dictionary.&nbsp; The code will be released at&nbsp;<em>https://github.com/linlei1214/SITS-Former</em> soon.</p>

opencc-by-4.0Dec 2021View details →
geo24/100

Discovery of antimicrobial peptides targeting Acinetobacter baumannii via a pre-trained and fine-tuned few-shot learning-based pipeline

GEO Series GSE306268. Acinetobacter baumannii. 6 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenSep 2025View details →
zenodo20/100

UVCGAN Pre-Trained Models

<p>UVCGAN Pre-Trained Models</p>

openbsd-2-clause-netbsdMar 2022View details →
zenodo20/100

Pre-trained embeddings HybridFC

<p>Contains pre-trained embeddings of entities and relations</p>

opencc-by-4.0May 2022View details →
zenodo20/100

Windy events detection in big bioacoustics datasets using a pre-trained Convolutional Neural Network

<p>This repository icludes the code and all relevant files used throughout our study. These encompass everything from the initial sheets of the whole acoustic dataset utilised for selecting the annotated dataset to the notebook (.ipynb) and the recordings employed in training the model.</p>

restrictedcc-by-4.0May 2024View details →
zenodo16/100

Classifying Open-Source Pre-Trained Models and Datasets for Software Engineering

<p>The replication package for the short paper titled 'Classifying Open-Source Pre-Trained Models and Datasets for Software Engineering' is provided. It includes a README file and accompanying scripts with comprehensive instructions to facilitate the replication of the analysis presented in the paper.</p>

openNov 2024View details →
zenodo12/100

Pre-trained word2vec models for ``Easy over Hard: A Case Study on Deep Learning''

<p>Since the whole stack overflow dump is so big, we can't easily handle well. Here, we provide 10 pre trained word2vec models with different seeds.</p> <p> </p> <p>More details about how to use it, please see paper </p>

restrictedMar 2017View details →
zenodo8/100

Pre-trained word2vec model for ``Easy over Hard: A Case Study on Deep Learning''

<p>Since the whole stack overflow dump is so big, we can't easily handle well. Here, we provide 10 pre trained word2vec models with different seeds using skip-gram algorithms</p> <p> </p> <p>More details about how to use it, please see paper </p>

restrictedMar 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record