Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
75
datasets available to search
ShareScore release 0.9.0
Dataset results
75 results for “pre-training”
Pre-trained Dataset of mRNABERT
Open the record for dataset details and reuse information.
Pre-trained model of SPC-MVSNet.
<p>Pre-trained model of SPC-MVSNet.</p>
Twitter pre-trained word vectors
<p>Clean up of <strong>glove.twitter.27B.zip </strong><ODC Public Domain Dedication and Licence (PDDL) 1.0>.</p> <p>"2B tweets, 27B tokens, 1.2M vocab, uncased"</p> <p><strong>Changes from original</strong></p> <ul> <li>Headers added to allow loading by gensim. [Added via scripts.glove2word2vec]</li> <li>Recompressed as individual gzip files [Instead of a combined zip].</li> </ul> <p>These changes make the files easier to work with and increase compatibility.</p> <p><strong>Headers</strong></p> <p>Example of added header line</p> <blockquote> <p>1193513 200</p> </blockquote> <p>Header gives number of tokens and dimensions.</p> <p><strong>Statistics</strong></p> <ul> <li>Entries: 1,193,513</li> <li>Token length (characters). Min:1, Max:140, Avg:6.73</li> <li>Number of words per token. Min:0, Max:17, Avg:1.00669200921984</li> <li>Tokens with more than one word: 4874 (0.41%)</li> <li>Twitter data collection date: <em>Unknown.</em><strong><em> </em></strong></li> </ul> <p>History:</p> <ul> <li><strong>?? Aug 2014 </strong>— GloVe v.1.0 released</li> <li><strong>16 Aug 2014 </strong>— Files first appear as headerless .txt.gz files, some files have mislabeled linked (via <a href="https://web.archive.org/web/20140816165523/http://www-nlp.stanford.edu/projects/glove/">wayback machine)</a></li> <li><strong>?? Oct 2015</strong> — GloVe v.1.2 released</li> <li><strong>?? ??? ????</strong> — Files replaced with a .zip file</li> <li><strong>03 June 2019 </strong>— (These files) Repackaged like original as .txt.gz, plus added headers for increased compatibility</li> </ul> <p>Example of 17-word token:</p> <blockquote> <p> سكس_طيز_قحبه_عنيف_اغتصاب_سكسيه_فحل_زب_نيك_بنات_مكوه_شهوه_لحس_عنف_تومبوي_ليدي_سبورت</p> </blockquote> <p><strong>200d file:</strong></p> <ul> <li>Normalized: no</li> <li>Values: Min:-6.7986, Max:4.609, Avg:0.009065093</li> <li>Zero values (exactly zero): none</li> <li>Zero values (approx zero) per entry: Min:0 (0.00%), Max:2 of 200 (1.0%), Avg:0.00375865197949247 (0.00%)</li> </ul>
Classifying Open-Source Pre-Trained Models and Datasets for Software Engineering
Open the record for dataset details and reuse information.
Dataset and pre-trained models for TopoMatch
<p>This contains datasets and pre-trained models for TopoMatch</p>
Automated Program Repair in the Era of Large Pre-trained Language Models
<p>Code used for the paper along with the generated outputs</p>
animal soup sample data, ground truth dataset, and pre-trained models
<p>- sample data and ground truth files for animal soup tests</p> <p>- pre-trained models</p>
SITS-Former: A pre-trained spatio-spectral-temporal representation model for Sentinel-2 time series classifcation
<p>This is the unlabeled dataset we introduced in the presented paper '<strong>SITS-Former: A pre-trained spatio-spectral-temporal representation model for Sentinel-2 time series classifcation</strong>'. This dataset can be used to pre-train a specified deep learning model (such as SITS-Former, CNN-Transformer, ConvLSTM. etc) for patch-based Sentinel-2 time series classification. </p> <p>In this dataset, each sample corresponds to an unlabeled image patch time series, which is stored as a separate numpy file named '<em>unlabeled_XXX.npz</em>'. You can use '<em>np.load</em>' to open a saved '<em>.npz</em>' file and get two arrays (querid by "ts" and "doy") from the returned dictionary. The code will be released at <em>https://github.com/linlei1214/SITS-Former</em> soon.</p>
Discovery of antimicrobial peptides targeting Acinetobacter baumannii via a pre-trained and fine-tuned few-shot learning-based pipeline
GEO Series GSE306268. Acinetobacter baumannii. 6 samples. Type: Expression profiling by high throughput sequencing.
UVCGAN Pre-Trained Models
<p>UVCGAN Pre-Trained Models</p>
Pre-trained embeddings HybridFC
<p>Contains pre-trained embeddings of entities and relations</p>
Windy events detection in big bioacoustics datasets using a pre-trained Convolutional Neural Network
<p>This repository icludes the code and all relevant files used throughout our study. These encompass everything from the initial sheets of the whole acoustic dataset utilised for selecting the annotated dataset to the notebook (.ipynb) and the recordings employed in training the model.</p>
Classifying Open-Source Pre-Trained Models and Datasets for Software Engineering
<p>The replication package for the short paper titled 'Classifying Open-Source Pre-Trained Models and Datasets for Software Engineering' is provided. It includes a README file and accompanying scripts with comprehensive instructions to facilitate the replication of the analysis presented in the paper.</p>
Pre-trained word2vec models for ``Easy over Hard: A Case Study on Deep Learning''
<p>Since the whole stack overflow dump is so big, we can't easily handle well. Here, we provide 10 pre trained word2vec models with different seeds.</p> <p> </p> <p>More details about how to use it, please see paper </p>
Pre-trained word2vec model for ``Easy over Hard: A Case Study on Deep Learning''
<p>Since the whole stack overflow dump is so big, we can't easily handle well. Here, we provide 10 pre trained word2vec models with different seeds using skip-gram algorithms</p> <p> </p> <p>More details about how to use it, please see paper </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.