Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5
datasets available to search
ShareScore release 0.9.0
Dataset results
5 results for “word vectors”
Hacker News lda2vec pretrained word vectors
<p>See also: https://zenodo.org/record/45901 and https://zenodo.org/record/49899 and https://zenodo.org/record/49900 and https://zenodo.org/record/49902</p>
Hacker News lda2vec model pretrained word vectors
<p>See also: https://zenodo.org/record/45901 and https://zenodo.org/record/49899 and https://zenodo.org/record/49900</p>
Training word vectors on text from The Physics Teacher using Word2Vec
<p>This notebook and dataset allows one to play around with word vectors trained on text from articles in the journal <a href="https://pubs.aip.org/aapt/pte">The Physics Teacher</a> published between 1963 (the start of publication) and 2020, around 15000 articles in total.</p> <p>The primary datafile, "TPT_word2vec_words_bigrams_V1.pkl", is a list of cleaned text from these articles. It contains a list, within which each paper is a sub-list. Each sentence in that paper is yet another sub-list which contains the words in that sentence in order. However, in the data cleaning process we have removed “stop words” (like if, and, but, etc.), punctuation, symbols, and numbers, as well as lowercased all words and combined words that frequently go together into one (like “high” and "school” to “high_school”). Here is an example of 3 sentences taken from a random paper: <br> <br>[['magnet', 'spin', 'tape', 'magnetize', 'strongly', 'time', 'pole', 'approach'],<br>['magnet', 'place', 'center', 'counterweight', 'period', 'magnetize', 'pulse', 'twice', 'long'],<br>['trial', 'tape', 'examine', 'sprinkle_iron', 'filing', 'length'], ... ]</p> <p>In the notebook, we create a set of word vectors from these sentences using the Word2Vec technique, first published by Mikolov et al. (2013):</p> <p>Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space (No. arXiv:1301.3781). arXiv. https://doi.org/10.48550/arXiv.1301.3781</p> <p>The notebook includes code for both loading in word vectors from a trained model (also included, "TPT_word2vec.model") and creating the same model from the TPT text dataset. Note that word vectors are randomly initialized, so we include a random seed to make this training replicable. Changing the seed will alter some of the results (although the changes seem fairly minor).</p> <p>With a trained model, we demonstrate some applications of word vectors: adding and subtracting meanings <br>(for example "experiment" - "uncertainty" = "demonstration") and visualizing low-dimensional representations of word vectors.</p> <p>In order to install the required packages, you can use the requirements.txt file. If using pip, run "pip install requirements.txt". Or, if using Anaconda (recommended), you can use "conda install --file requirements.txt". You will also need the software to run jupyter notebooks, which can be installed with Anaconda or pip.</p>
Hacker News lda2vec model word vectors
<p>See also: https://zenodo.org/record/45901 and https://zenodo.org/record/49899</p>
Twitter pre-trained word vectors
<p>Clean up of <strong>glove.twitter.27B.zip </strong><ODC Public Domain Dedication and Licence (PDDL) 1.0>.</p> <p>"2B tweets, 27B tokens, 1.2M vocab, uncased"</p> <p><strong>Changes from original</strong></p> <ul> <li>Headers added to allow loading by gensim. [Added via scripts.glove2word2vec]</li> <li>Recompressed as individual gzip files [Instead of a combined zip].</li> </ul> <p>These changes make the files easier to work with and increase compatibility.</p> <p><strong>Headers</strong></p> <p>Example of added header line</p> <blockquote> <p>1193513 200</p> </blockquote> <p>Header gives number of tokens and dimensions.</p> <p><strong>Statistics</strong></p> <ul> <li>Entries: 1,193,513</li> <li>Token length (characters). Min:1, Max:140, Avg:6.73</li> <li>Number of words per token. Min:0, Max:17, Avg:1.00669200921984</li> <li>Tokens with more than one word: 4874 (0.41%)</li> <li>Twitter data collection date: <em>Unknown.</em><strong><em> </em></strong></li> </ul> <p>History:</p> <ul> <li><strong>?? Aug 2014 </strong>— GloVe v.1.0 released</li> <li><strong>16 Aug 2014 </strong>— Files first appear as headerless .txt.gz files, some files have mislabeled linked (via <a href="https://web.archive.org/web/20140816165523/http://www-nlp.stanford.edu/projects/glove/">wayback machine)</a></li> <li><strong>?? Oct 2015</strong> — GloVe v.1.2 released</li> <li><strong>?? ??? ????</strong> — Files replaced with a .zip file</li> <li><strong>03 June 2019 </strong>— (These files) Repackaged like original as .txt.gz, plus added headers for increased compatibility</li> </ul> <p>Example of 17-word token:</p> <blockquote> <p> سكس_طيز_قحبه_عنيف_اغتصاب_سكسيه_فحل_زب_نيك_بنات_مكوه_شهوه_لحس_عنف_تومبوي_ليدي_سبورت</p> </blockquote> <p><strong>200d file:</strong></p> <ul> <li>Normalized: no</li> <li>Values: Min:-6.7986, Max:4.609, Avg:0.009065093</li> <li>Zero values (exactly zero): none</li> <li>Zero values (approx zero) per entry: Min:0 (0.00%), Max:2 of 200 (1.0%), Avg:0.00375865197949247 (0.00%)</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.