Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

15

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

15 results for “Fasttext”

Learn how ShareScore rates datasets ↗
zenodo44/100

German CBOW FastText embeddings with min count 250

<p>FastText embeddings built from Common Crawl german dataset</p> <table> <caption>Parameters</caption> <thead> <tr> <th scope="col">Parameters</th> <th scope="col">Value(s)</th> </tr> </thead> <tbody> <tr> <td>Dimensions</td> <td>256 and 384</td> </tr> <tr> <td>Context window</td> <td>5</td> </tr> <tr> <td>Negative sampled</td> <td>10</td> </tr> <tr> <td>Epochs</td> <td>1</td> </tr> <tr> <td>Number of buckets</td> <td>131072 or 262144</td> </tr> <tr> <td>Min n</td> <td>3</td> </tr> <tr> <td>Max n</td> <td>6</td> </tr> </tbody> </table>

opencc-by-sa-3.0Oct 2021View details →
zenodo40/100

Quantized versions of fastText embeddings for 158 languages

<p>These are compressed (quantized) versions of word embeddings for 158 languages originally available from https://fasttext.cc/docs/en/crawl-vectors.html . Following steps were performed to reduce file size:</p> <ol> <li>the output matrix is discarded</li> <li>the quantization is performed with parameters &quot;-qnorm -dsub 1&quot;</li> </ol> <p>The average cosine similarity between original and quantized vectors for frequent words is 0.99. The file size is 4-6 times smaller.</p>

opencc-by-sa-3.0Jan 2020View details →
zenodo40/100

FastText Spanish Medical Embeddings

<p>[Plan TL/medicine/word embeddings] Word embeddings generated from Spanish corpora that include: (a) the full-text in Spanish available in SciELO.org (until December/2018), (b) all articles from the following Wikipedia categories: Pharmacology, Pharmacy, Medicine and Biology (during December/2018) and (c) the concatenation of the previous two corpora.</p> <p>We used fastText to train the word embeddings.</p> <p>For more information, we refer to the corresponding article:&nbsp;<a href="https://www.aclweb.org/anthology/W19-1916/">https://www.aclweb.org/anthology/W19-1916/</a></p> <p>Copyright (c) 2021 Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p>

opencc-by-4.0Apr 2020View details →
zenodo40/100

French CBOW FastText embeddings with min count 250

<p>FastText embeddings built from Common Crawl French dataset</p> <table> <caption>Parameters</caption> <thead> <tr> <th scope="col">Parameters</th> <th scope="col">Value(s)</th> </tr> </thead> <tbody> <tr> <td>Dimensions</td> <td>512</td> </tr> <tr> <td>Context window</td> <td>5</td> </tr> <tr> <td>Negative sampled</td> <td>10</td> </tr> <tr> <td>Epochs</td> <td>1</td> </tr> <tr> <td>Number of buckets</td> <td>262144</td> </tr> <tr> <td>Min n</td> <td>3</td> </tr> <tr> <td>Max n</td> <td>6</td> </tr> </tbody> </table>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Social Sciences Word Embeddings in FastText

<p>These social science word embeddings in FastText have been created from 37,604 open access social science research papers from the social science access repository (https://www.gesis.org/ssoar/home). They are available in German and English.</p> <p>(skipgram model, n-grams with n&ge;3 and n&le;6, different dimensions (100, 150, 200, 300, 500), five epochs, learning rate 0.05, five negative examples)</p> <p>Please cite:</p> <p>Schiffers, Ricardo, Dagmar Kern, and Daniel Hienert. 2022. &quot;Evaluation of Word Embeddings for the Social Sciences.&quot; In <em>Proceedings of the 6th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature</em>, edited by Stefania Degaetano, Anna Kazantseva, Nils Reiter, and Stan Szpakowicz, 1-6. Gyeongju: Association for Computational Linguistics. <a href="https://aclanthology.org/2022.latechclfl-1.1">https://aclanthology.org/2022.latechclfl-1.1</a>.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Spanish Skip-Gram Word Embeddings in FastText

<p>These Spanish word embeddings in FastText have been generated from the largest corpus ever made in Spanish till&nbsp;date. The corpus has more than 2TB of high-quality text,&nbsp;compiled from the different web crawlings done by the National Library of Spain from 2009 to 2019.&nbsp;</p> <p>These are the&nbsp;SKIP-GRAM&nbsp;embeddings, for the CBOW embeddings see:&nbsp;https://zenodo.org/record/5044988</p> <p><strong>Citation</strong></p> <pre><code>@article{gutierrezfandino2022, author = {Asier Gutiérrez-Fandiño and Jordi Armengol-Estapé and Marc Pàmies and Joan Llop-Palao and Joaquin Silveira-Ocampo and Casimiro Pio Carrino and Carme Armentano-Oller and Carlos Rodriguez-Penagos and Aitor Gonzalez-Agirre and Marta Villegas}, title = {MarIA: Spanish Language Models}, journal = {Procesamiento del Lenguaje Natural}, volume = {68}, number = {0}, year = {2022}, issn = {1989-7553}, url = {http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6405}, pages = {39--60} }</code></pre> <p><strong>Copyright </strong></p> <p>Copyright (c) 2021 Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

Spanish CBOW Word Embeddings in FastText

<p>These Spanish word embeddings in FastText have been generated from the largest corpus ever made in Spanish till&nbsp;date. The corpus has more than 2TB of&nbsp;high-quality text,&nbsp;compiled from the different web crawlings done by the National Library of Spain from 2009 to 2019.&nbsp;</p> <p>These are the CBOW embeddings, for the SKIP-GRAM embeddings see:&nbsp;https://zenodo.org/record/5046525</p> <p><strong>Citation</strong></p> <pre><code>@article{gutierrezfandino2022, author = {Asier Gutiérrez-Fandiño and Jordi Armengol-Estapé and Marc Pàmies and Joan Llop-Palao and Joaquin Silveira-Ocampo and Casimiro Pio Carrino and Carme Armentano-Oller and Carlos Rodriguez-Penagos and Aitor Gonzalez-Agirre and Marta Villegas}, title = {MarIA: Spanish Language Models}, journal = {Procesamiento del Lenguaje Natural}, volume = {68}, number = {0}, year = {2022}, issn = {1989-7553}, url = {http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6405}, pages = {39--60} }</code></pre> <p><strong>Copyright </strong></p> <p>Copyright (c) 2021 Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

German CBOW FastText embeddings with min count 100

<p>FastText embeddings built from Common Crawl german dataset</p> <table> <caption>Parameters</caption> <thead> <tr> <th scope="col">Parameters</th> <th scope="col">Value(s)</th> </tr> </thead> <tbody> <tr> <td>Dimensions</td> <td>256 and 384</td> </tr> <tr> <td>Context window</td> <td>5</td> </tr> <tr> <td>Negative sampled</td> <td>10</td> </tr> <tr> <td>Epochs</td> <td>1</td> </tr> <tr> <td>Number of buckets</td> <td>131072 or 262144</td> </tr> <tr> <td>Min n</td> <td>3</td> </tr> <tr> <td>Max n</td> <td>6</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-sa-3.0Oct 2021View details →
zenodo40/100

Ancient Greek Fasttext Word Embeddings

<p>Word embeddings generated with Fasttext and 1 GB of Ancient Greek texts. These embeddings were produced for the study of social networks and social semantics in ancient Greece by the Diogenet project at the University of San Diego, California.&nbsp;</p>

opencc-by-4.0Oct 2021View details →
zenodo36/100

Wbbyyr: FastText language models for Mandarin Chinese, trained on 14m Sina Weibo posts for each year in 2012-2018 (Fold 1 of 10)

<p>Wbbyyr: FastText language models for Mandarin Chinese, trained on 14,440,000 Sina Weibo posts for each year in 2012-2018.</p> <p>The&nbsp;14,440,000&nbsp;posts from&nbsp;each year are&nbsp;split into 10 folds. Due to Zenodo size limit, this dataset contains only the first fold from each year.</p> <p>Each model is trained for 20 iterations. Each vector is 300 dimensions long.</p>

opencc-by-4.0Jan 2020View details →
zenodo36/100

English and Portuguese CBOW Models from Europarl Corpus, version 7, using FastText with Subwords Option

<p>The models were trained&nbsp;using FastText, model CBOW, 40 epochs, and subwords. Each *.BIN file has its *.VEC file with the vocabulary ordered by frequency. The *.BIN file can return a vector to represent an out-of-vocabulary (OOV) word if the necessary parts of the OOV word were&nbsp;used in training. FastText and Gensim can use these&nbsp;files. The English and Portuguese models are identified in the file name,&nbsp;&nbsp;<strong>_en_</strong>&nbsp; and&nbsp;<strong>_pt_</strong>&nbsp;respectively.</p> <p>An Excel file has the neighborhood changes of some selected words during training on each epoch.&nbsp;A previous exercise to find words with more than one meaning.</p>

opencc-by-4.0May 2022View details →
zenodo36/100

FastText embeddings obtained from SNOMED CT - Spanish and International

<p>FastText model trained in a corpus that was obtained from SNOMED CT by performing random walks. There are three models: one trained using all relations available in the international version of SNOMED CT, another one trained using only is_a relations from the international version of SNOMED and the last one is trained using all relations from the Spanish version of SNOMED CT. This was developed for the end project of the Master Degree in Artificial Intelligence in the Universidad Polit&eacute;cnica de Madrid. More information and the document of the project can be found&nbsp;in&nbsp;<a href="https://github.com/JavierCastellD/SemanticFormalizationSNOMED">https://github.com/JavierCastellD/SemanticFormalizationSNOMED</a>.</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Italian FastText models

<p>Italian FastText models trained from scratch&nbsp;on a dataset composed of:</p> <p>- <strong>wiki</strong>: a&nbsp;dump of Italian Wikipedia (as of December 15, 2022), comprising 25,548,651 sentences and 526,640,982 words (3.2 GB of raw text);<br> - <strong>webz</strong>: a&nbsp;dataset of Italian news (159,226 documents) from the webz.io platform, crawled in October 2015, containing 44,041,823 sentences and 44,544,385 words (244 MB);<br> - a dataset of 5,510 Italian news articles from the newspaper ModenaToday&nbsp;(<strong>MT</strong>) or 15,115 documents from the Italian version of Reuters (<strong>RCV2</strong>).</p> <p><strong>ft_wiki_wbz_mt_20_epochs</strong>: FastText model trained on the&nbsp;dataset consisting of wiki, webz, and MT for 20 epochs</p> <p><strong>ft_wiki_wbz_mt_50_epochs</strong>: FastText model trained on the&nbsp;dataset consisting of wiki, webz, and MT for 50 epochs</p> <p><strong>ft_wiki_wbz_reut_20_epochs</strong>: FastText model trained on the&nbsp;dataset consisting of wiki, webz, and RCV2 for 20 epochs</p> <p><strong>ft_wiki_wbz_reut_50_epochs</strong>: FastText model trained on the&nbsp;dataset consisting of wiki, webz, and RCV2 for 50 epochs</p>

opencc-by-4.0Oct 2023View details →
zenodo32/100

Clean-Clean ER Datasets with FastText embeddings

<p>Please see <a href="https://zenodo.org/records/8433873">version 5 </a>for more datasets. This one contains only a print-out of the DBPedia dataset in CSV format.</p><p>Each line corresponds to a different entity profile and has the following structure (where n is number of attributes):</p><p>numerical_id , uri , n , &nbsp;aname_0 , aval_0 , aname_1 , aval_1 ,...aname_n , aval_n &nbsp;</p><p>That is, the separator is "space,space".</p><p>Special thanks to <a href="https://github.com/moemode">moemode </a>for providing the dataset.</p>

opencc-by-4.0Mar 2022View details →
zenodo28/100

English fastText Wikipedia embeddings

<p>This repository contains the fastText embeddings, pre-trained on English Wikipedia.</p>

opencc-by-4.0Apr 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record