Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

5

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

5 results for “word vectors”

Learn how ShareScore rates datasets ↗
zenodo36/100

Hacker News lda2vec pretrained word vectors

<p>See also: https://zenodo.org/record/45901 and https://zenodo.org/record/49899 and https://zenodo.org/record/49900&nbsp;and https://zenodo.org/record/49902</p>

opencc-zeroApr 2016View details →
zenodo36/100

Hacker News lda2vec model pretrained word vectors

<p>See also: https://zenodo.org/record/45901 and https://zenodo.org/record/49899 and https://zenodo.org/record/49900</p>

opencc-zeroApr 2016View details →
zenodo36/100

Training word vectors on text from The Physics Teacher using Word2Vec

<p>This notebook and dataset allows one to play around with word vectors trained on text from articles in the journal <a href="https://pubs.aip.org/aapt/pte">The Physics Teacher</a> published between 1963 (the start of publication) and 2020, around 15000 articles in total.</p> <p>The primary datafile, "TPT_word2vec_words_bigrams_V1.pkl", is a list of cleaned text from these articles. It contains a list, within which each paper is a sub-list. Each sentence in that paper is yet another sub-list which contains the words in that sentence in order. However, in the data cleaning process we have removed &ldquo;stop words&rdquo; (like if, and, but, etc.), punctuation, symbols, and numbers, as well as lowercased all words and combined words that frequently go together into one (like &ldquo;high&rdquo; and "school&rdquo; to &ldquo;high_school&rdquo;). Here is an example of 3 sentences taken from a random paper:&nbsp;<br>&nbsp;<br>[['magnet', 'spin', &nbsp;'tape', &nbsp;'magnetize', &nbsp;'strongly', &nbsp;'time', &nbsp;'pole', &nbsp;'approach'],<br>['magnet', &nbsp;'place', &nbsp;'center', &nbsp;'counterweight', &nbsp;'period', &nbsp;'magnetize', &nbsp;'pulse', &nbsp;'twice', &nbsp;'long'],<br>['trial', 'tape', 'examine', 'sprinkle_iron', 'filing', 'length'], ... ]</p> <p>In the notebook, we create a set of word vectors from these sentences using the Word2Vec technique, first published by Mikolov et al. (2013):</p> <p>Mikolov, T., Chen, K., Corrado, G., &amp; Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space (No. arXiv:1301.3781). arXiv. https://doi.org/10.48550/arXiv.1301.3781</p> <p>The notebook includes code for both loading in word vectors from a trained model (also included, "TPT_word2vec.model") and creating the same model from the TPT text dataset. Note that word vectors are randomly initialized, so we include a random seed to make this training replicable. Changing the seed will alter some of the results (although the changes seem fairly minor).</p> <p>With a trained model, we demonstrate some applications of word vectors: adding and subtracting meanings <br>(for example "experiment" - "uncertainty" = "demonstration") and visualizing low-dimensional representations of word vectors.</p> <p>In order to install the required packages, you can use the requirements.txt file. &nbsp;If using pip, run "pip install requirements.txt". Or, if using Anaconda (recommended), you can use "conda install --file requirements.txt". You will also need the software to run jupyter notebooks, which can be installed with Anaconda or pip.</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

Hacker News lda2vec model word vectors

<p>See also: https://zenodo.org/record/45901 and https://zenodo.org/record/49899</p>

opencc-zeroApr 2016View details →
zenodo28/100

Twitter pre-trained word vectors

<p>Clean up of <strong>glove.twitter.27B.zip </strong>&lt;ODC Public Domain Dedication and Licence (PDDL) 1.0&gt;.</p> <p>&quot;2B tweets, 27B tokens, 1.2M vocab, uncased&quot;</p> <p><strong>Changes from original</strong></p> <ul> <li>Headers added to allow loading by gensim. [Added via scripts.glove2word2vec]</li> <li>Recompressed as individual gzip files [Instead of a combined zip].</li> </ul> <p>These changes make the files easier to work with and increase compatibility.</p> <p><strong>Headers</strong></p> <p>Example of added header line</p> <blockquote> <p>1193513 200</p> </blockquote> <p>Header gives number of tokens and&nbsp;dimensions.</p> <p><strong>Statistics</strong></p> <ul> <li>Entries: 1,193,513</li> <li>Token length (characters). Min:1, Max:140, Avg:6.73</li> <li>Number of words per token. Min:0, Max:17, Avg:1.00669200921984</li> <li>Tokens with more than one word: 4874 (0.41%)</li> <li>Twitter data collection date: <em>Unknown.</em><strong><em> </em></strong></li> </ul> <p>History:</p> <ul> <li><strong>?? Aug 2014 </strong>&mdash; GloVe v.1.0 released</li> <li><strong>16 Aug 2014&nbsp;</strong>&mdash; Files first appear as headerless .txt.gz files, some files have mislabeled linked (via <a href="https://web.archive.org/web/20140816165523/http://www-nlp.stanford.edu/projects/glove/">wayback machine)</a></li> <li><strong>?? Oct 2015</strong> &mdash; GloVe v.1.2 released</li> <li><strong>?? ??? ????</strong> &mdash; Files replaced with a .zip file</li> <li><strong>03 June 2019 </strong>&mdash; (These files) Repackaged like original as .txt.gz, plus added headers for increased compatibility</li> </ul> <p>Example of 17-word token:</p> <blockquote> <p>&nbsp;سكس_طيز_قحبه_عنيف_اغتصاب_سكسيه_فحل_زب_نيك_بنات_مكوه_شهوه_لحس_عنف_تومبوي_ليدي_سبورت</p> </blockquote> <p><strong>200d file:</strong></p> <ul> <li>Normalized: no</li> <li>Values: &nbsp;Min:-6.7986, Max:4.609, Avg:0.009065093</li> <li>Zero values (exactly zero): none</li> <li>Zero values (approx zero) per entry: Min:0 (0.00%), Max:2 of 200 (1.0%), Avg:0.00375865197949247 (0.00%)</li> </ul>

openodc-pddlJun 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record