Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
475
datasets available to search
ShareScore release 0.9.0
Dataset results
475 results for “word”
Basis words in Uzbek language(For Primary school students)
<p>Basis words in Uzbek language for Primary school students.</p> <p>The "1 sinf" file is a set of basic words prepared on the basis of unique lemmas from all 1st grade school textbooks.</p> <p>The "2 sinf" file is a set of basic words prepared on the basis of unique lemmas from all 1st grade school textbooks.</p> <p>The "3 sinf" file is a set of basic words prepared on the basis of unique lemmas from all 1st grade school textbooks.</p> <p>The "4 sinf" file is a set of basic words prepared on the basis of unique lemmas from all 1st grade school textbooks.</p>
H̶a̶n̶d̶w̶r̶i̶t̶i̶n̶g̶ - A collection of struck through handwritten word images using various styles
<p># H̶a̶n̶d̶w̶r̶i̶t̶i̶n̶g̶ - A collection of handwritten word images, each struck through using various styles</p> <p>This database may be used for non-commercial research purposes only. If you publish material based on this database - please cite:<br>Zesch, T., & Gold, C. (2024). H̶a̶n̶d̶w̶r̶i̶t̶i̶n̶g̶ - various struck-through handwritten words [Data set]. Zenodo.</p> <h3><br>Structure:</h3> <p>There are two types of sources: <strong>genuine</strong> and <strong>semi-genuine</strong><br>While genuine includes struck-through handwritten images that are made accidentally, semi-genuine were made on purpose and conducted as the main part of this dataset. </p> <h3>Genuine:</h3> <p>For the genuine part, only the strike-out_genuine.txt file exists. It links to struck-through words of the datasets: Handwritten ASAP Short Answer Scoring (published at: https://zenodo.org/records/8088866), IAM, and GoBo (published at: https://zenodo.org/records/8085511).<br>The struck-through types were annotated by 2 annotators and a gold version was created. </p> <p>The file has the following structure:<br><em>path status gold a1 a2 a3</em><br>e.g.:<br>genuine/Handwritten ASAP/SAS_3_6818_0.png ok wa wa? wa? wa<br>The image is part of the Handwritten ASAP dataset and refers to image "AS_3_6818_0.png". The status is ok and the gold label is "wa" which stands for wavy (see list below). The "?" indicates unsure annotation of both annotators independently.</p> <h3><br>Semi-genuine: </h3> <p><br>17 writers participated. 9 male, 8 female<br>The writers were asked to write 12 words for each struck-through type. <br>In sum over 2000 images of handwritten struck-through words (including none) were collected, with 204 images each type. <br>An instruction was presented to the writers including examples. </p> <p><strong>types for strike-out:</strong><br> no: none<br> sh: single-horizontal<br> so: single-oblique<br> mh: multiple-horizontal<br> mo: multiple-oblique<br> cr: crossed<br> ci: circled<br> wa: wavy<br> zi: zigzag<br> bl: blackened</p> <p>Afterward, the writers struck through the handwritten words according to the stated type.<br>The images are published in "boxes", "gray images" and "raw images" in color as scanned. The description file "Strike-out_semi-genuine.txt" references the "boxes" only and describes the path, status, and struck-through type of each image. </p> <h3><br>Sidenote:</h3> <p>The writers were asked to note if the struck-through type was their preferred type and if it felt natural to use this type or unnatural. More details regarding this topic can be found in the instruction pdf file.</p> <p> </p>
DWUG EN: Diachronic Word Usage Graphs for English
<p>This data collection contains diachronic Word Usage Graphs (WUGs) for English. Find a description of the data format, code to process the data and further datasets on the <a href="https://www.ims.uni-stuttgart.de/data/wugs">WUGsite</a>.</p> <p>See previous versions for additional testsets.</p> <p>Please find more information on the provided data in the papers referenced below.</p> <h3>Reference</h3> <p>Dominik Schlechtweg, Nina Tahmasebi, Simon Hengchen, Haim Dubossarsky, Barbara McGillivray. 2021. <a href="https://aclanthology.org/2021.emnlp-main.567/">DWUG: A large Resource of Diachronic Word Usage Graphs in Four Languages</a>. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing.</p> <p>Dominik Schlechtweg, Pierluigi Cassotti, Bill Noble, David Alfter, Sabine Schulte im Walde, Nina Tahmasebi. <a href="https://aclanthology.org/2024.emnlp-main.796/">More DWUGs: Extending and Evaluating Word Usage Graph Datasets in Multiple Languages</a>. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing.</p>
DWUG DE: Diachronic Word Usage Graphs for German
<p>This data collection contains diachronic Word Usage Graphs (WUGs) for German. Find a description of the data format, code to process the data and further datasets on the <a href="https://www.ims.uni-stuttgart.de/data/wugs">WUGsite</a>.</p> <p>See previous versions for additional testsets. Find a version of this dataset annotated with classical word sense definitions at <a href="https://zenodo.org/doi/10.5281/zenodo.8197552">DWUG DE Sense</a>.</p> <p>Please find more information on the provided data in the papers referenced below.</p> <h3>Reference</h3> <p>Dominik Schlechtweg, Nina Tahmasebi, Simon Hengchen, Haim Dubossarsky, Barbara McGillivray. 2021. <a href="https://aclanthology.org/2021.emnlp-main.567/">DWUG: A large Resource of Diachronic Word Usage Graphs in Four Languages</a>. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing.</p> <p>Dominik Schlechtweg, Pierluigi Cassotti, Bill Noble, David Alfter, Sabine Schulte im Walde, Nina Tahmasebi. <a href="https://aclanthology.org/2024.emnlp-main.796/">More DWUGs: Extending and Evaluating Word Usage Graph Datasets in Multiple Languages</a>. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing.</p>
DWUG SV: Diachronic Word Usage Graphs for Swedish
<p>This data collection contains diachronic Word Usage Graphs (WUGs) for Swedish. Find a description of the data format, code to process the data and further datasets on the <a href="https://www.ims.uni-stuttgart.de/data/wugs">WUGsite</a>.</p> <p>See previous versions for additional testsets.</p> <p>Please find more information on the provided data in the papers referenced below.</p> <h3>Reference</h3> <p>Dominik Schlechtweg, Nina Tahmasebi, Simon Hengchen, Haim Dubossarsky, Barbara McGillivray. 2021. <a href="https://aclanthology.org/2021.emnlp-main.567/">DWUG: A large Resource of Diachronic Word Usage Graphs in Four Languages</a>. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing.</p> <p>Dominik Schlechtweg, Pierluigi Cassotti, Bill Noble, David Alfter, Sabine Schulte im Walde, Nina Tahmasebi. More DWUGs: Extending and Evaluating Word Usage Graph Datasets in Multiple Languages. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing.</p>
Spanish Legal Domain Word & Sub-Word Embeddings
<p><strong>Spanish Legal Word and Sub-word Embeddings in FastText</strong></p> <p>These embeddings have been generated from the largest corpus (9GB) ever made from Spanish Legal resources till the date.</p> <p>More legal domain resources: https://github.com/PlanTL-GOB-ES/lm-legal-es</p> <p><strong>Citation</strong></p> <pre><code>@misc{gutierrezfandino2021legal, title={Spanish Legalese Language Model and Corpora}, author={Asier Gutiérrez-Fandiño and Jordi Armengol-Estapé and Aitor Gonzalez-Agirre and Marta Villegas}, year={2021}, eprint={2110.12201}, archivePrefix={arXiv}, primaryClass={cs.CL} }</code></pre> <p><strong>Copyright </strong></p> <p>Copyright (c) 2021 Secretaría de Estado de Digitalización e Inteligencia Artificial</p>
Spanish Skip-Gram Word Embeddings in FastText
<p>These Spanish word embeddings in FastText have been generated from the largest corpus ever made in Spanish till date. The corpus has more than 2TB of high-quality text, compiled from the different web crawlings done by the National Library of Spain from 2009 to 2019. </p> <p>These are the SKIP-GRAM embeddings, for the CBOW embeddings see: https://zenodo.org/record/5044988</p> <p><strong>Citation</strong></p> <pre><code>@article{gutierrezfandino2022, author = {Asier Gutiérrez-Fandiño and Jordi Armengol-Estapé and Marc Pàmies and Joan Llop-Palao and Joaquin Silveira-Ocampo and Casimiro Pio Carrino and Carme Armentano-Oller and Carlos Rodriguez-Penagos and Aitor Gonzalez-Agirre and Marta Villegas}, title = {MarIA: Spanish Language Models}, journal = {Procesamiento del Lenguaje Natural}, volume = {68}, number = {0}, year = {2022}, issn = {1989-7553}, url = {http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6405}, pages = {39--60} }</code></pre> <p><strong>Copyright </strong></p> <p>Copyright (c) 2021 Secretaría de Estado de Digitalización e Inteligencia Artificial</p>
Spanish CBOW Word Embeddings in FastText
<p>These Spanish word embeddings in FastText have been generated from the largest corpus ever made in Spanish till date. The corpus has more than 2TB of high-quality text, compiled from the different web crawlings done by the National Library of Spain from 2009 to 2019. </p> <p>These are the CBOW embeddings, for the SKIP-GRAM embeddings see: https://zenodo.org/record/5046525</p> <p><strong>Citation</strong></p> <pre><code>@article{gutierrezfandino2022, author = {Asier Gutiérrez-Fandiño and Jordi Armengol-Estapé and Marc Pàmies and Joan Llop-Palao and Joaquin Silveira-Ocampo and Casimiro Pio Carrino and Carme Armentano-Oller and Carlos Rodriguez-Penagos and Aitor Gonzalez-Agirre and Marta Villegas}, title = {MarIA: Spanish Language Models}, journal = {Procesamiento del Lenguaje Natural}, volume = {68}, number = {0}, year = {2022}, issn = {1989-7553}, url = {http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6405}, pages = {39--60} }</code></pre> <p><strong>Copyright </strong></p> <p>Copyright (c) 2021 Secretaría de Estado de Digitalización e Inteligencia Artificial</p>
Obtaining Better Static Word Embeddings Using Contextual Embedding Models
<p><strong>Obtaining Better Static Word Embeddings Using Contextual Embedding Models</strong></p> <p>This repository contains the dataset of pretrained word embeddings as well as datasets used to train them, released with the following <a href="https://arxiv.org/pdf/2106.04302.pdf">paper</a>.</p> <blockquote> <p>“Obtaining Better Static Word Embeddings Using Contextual Embedding Models” <em>ACL</em> (2021).</p> </blockquote> <p>The wikipedia datasets were preprocessed from the wikipedia dump downloaded from <a href="http://dumps.wikimedia.org">dumps.wikimedia.org</a> under Creative Commons Attribution-Share-Alike 3.0 License .</p> <p>If you found the provided resources useful, please cite the above paper. Here's a BibTeX entry you may use:</p> <blockquote> <p>@inproceedings{Gupta2021ObtainingPC,<br> title={Obtaining Better Static Word Embeddings Using Contextual Embedding Models},<br> author={Prakhar Gupta and Martin Jaggi},<br> booktitle={ACL},<br> year={2021}<br> }</p> </blockquote>
Sentiment Analysis and Cross-lingual Word Embeddings for Endangered Languages
<p>A sentiment analyzer and cross-lingual word embeddings for endangered languages (e.g., Erzya, Moksha, Skolt Sami, Komi-Zyrian).</p>
WikiMorph: Learning to Decompose Words into Morphological Structures
<p>WikiMorph is a JSON dataset that contains word breakdowns for English words. These word breakdowns primarily consist of morphological compounds (both from English and the word's etymology) along with each compound's associated definition. It also contains other fields that might be useful, such as syllables and parts-of-speech tags. The dataset contains entries for 355,782 unique words and 505,033 total entries. The data collection process for this dataset was described in the paper "<a href="https://link.springer.com/chapter/10.1007/978-3-030-78270-2_72">WikiMorph: Learning to Decompose Words into Morphological Structures</a>", with some additional updates after publication.</p> <p> </p> <pre><code class="language-json"> { "Word": "abduction", "PoS": "Noun", "Syllables": [ "ab", "duc", "tion" ], "Definition": "The act of abducing or abducting; a drawing apart; the movement which separates a limb or other part from the axis, or middle line, of the body.", "Morphemes": [ { "Affix": "abduct", "Language": "en", "PoS": "Verb", "Meaning": "To draw away, as a limb or other part, from the median axis of the body.", "Etymology Compounds": [ { "Affix": "ab", "Language": "la", "Decoded": "ab", "PoS": null, "Meaning": "away" }, { "Affix": "duco", "Language": "la", "Decoded": "duco", "PoS": null, "Meaning": "to lead" } ] }, { "Affix": "-ion", "Language": "en", "PoS": "Suffix", "Meaning": "an action or process, or the result of an action or process", "Etymology Compounds": [ { "Affix": "-iō", "Language": "la", "Decoded": "-io", "PoS": "Suffix", "Meaning": "Used to form abstract nouns from verbs." } ] } ] }</code></pre> <p> </p>
Pokémon Word Embeddings
<p>Word vector models for Pokémon (poke2vec). Word2Vec, FastText and Meta4Meaning models trained on a big Pokémon corpus.</p> <p>code.zip files has examples of how to load and use the models.</p> <p>Please cite the following paper if you use the resources:</p> <p>Hämäläinen, M., Alnajjar, K. & Partanen, N. (2021). <a href="https://researchportal.helsinki.fi/en/publications/nettikorpuksen-avulla-tuotettuja-sanavektorimalleja-pok%C3%A9monien-om">Nettikorpuksen avulla tuotettuja sanavektorimalleja Pokémonien ominaisuuksien kuvaamiseksi</a>. In Saarikivi, T. & Saarikivi, J. (eds.) <em>Turhan tiedon kirja — Tutkimuksista pois jätettyjä sivuja</em>. p. 199-214. SKS Kirjat.</p> <p><a href="https://www.researchgate.net/publication/354088508_How_Cute_is_Pikachu_Gathering_and_Ranking_Pokemon_Properties_from_Data_with_Pokemon_Word_Embeddings">English version of the paper</a></p>
DWUG LA: Diachronic Word Usage Graphs for Latin
<p>This data collection contains diachronic Word Usage Graphs (WUGs) for Latin. Find a description of the data format, code to process the data and further datasets on the <a href="https://www.ims.uni-stuttgart.de/data/wugs">WUGsite</a>.</p> <p>The annotation was coordinated by Barbara McGillivray, and done by Annie Burman, Daria Kondakova, Francesca Dell'Oro, Helena Bermudez Sabel, Hugo Burgess, Paola Marongiu, Rozalia Dobos and Tomaz Potocnik. The pre-annotation was coordinated and designed by Barbara McGillivray and done by Manuel Márquez Cruz.</p> <p>Please find more information on the provided data in the paper referenced below.</p>
Spanish CBOW Word Embeddings in Floret
<p><strong>Spanish CBOW Word Embeddings in Floret</strong></p> <p>The embeddings have been trained with the corpus from the National Library of Spain (<a href="http://www.bne.es/en/Inicio/index.html">Biblioteca Nacional de España</a> or BNE) using <a href="https://arxiv.org/pdf/1607.04606.pdf">floret</a> with the following hyperparameters:</p> <blockquote> <p>mode: str = "floret",<br> model: str = "cbow",<br> dim: int = 300,<br> mincount: int = 10,<br> minn: int = 5,<br> maxn: int = 6,<br> neg: int = 10,<br> hashcount: int = 2,<br> bucket: int = 50000,<br> thread: int = 128,</p> </blockquote> <p> </p> <p>Detailed information about the corpus can be found <a href="http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6405">here </a></p> <p>The processing took place on an HPC <a href="https://www.bsc.es/innovation-and-services/technical-information-cte-amd">node</a> equipped with an AMD EPYC 7742 (@ 2.250GHz) processor with 128 threads.</p> <p><strong>How to use</strong></p> <p>First initialize the spacy vectors from the floret table (.floret file):</p> <pre><code>spacy init vectors es floret_embeddings_bne_es.floret floret_embeddings_bne_es --mode floret</code></pre> <pre><code>import spacy # Load the floret vectors floret_embeddings = spacy.load("floret_embeddings_bne_es") # Get the embeddings of some words playa = floret_embeddings.vocab["playa"] frío = floret_embeddings.vocab["frío"] invierno = floret_embeddings.vocab["invierno"] verano = floret_embeddings.vocab["verano"] # Get some similarities print(frío.similarity(invierno)) print(frío.similarity(verano)) # frío should be more similar to invierno than verano. print(playa.similarity(invierno)) print(playa.similarity(verano)) # playa should be more similar to verano than invierno.</code></pre> <p><strong>Intended Uses and Limitations</strong></p> <p>At the time of submission, no measures have been taken to estimate the bias and toxicity embedded in the model. However, we are well aware that our models may be biased since the corpora have been collected using crawling techniques on multiple web sources. We intend to conduct research in these areas in the future, and if completed, this card will be updated.</p> <p><strong>Authors</strong></p> <p>The Text Mining Unit from Barcelona Supercomputing Center.</p> <p><strong>Contact Information</strong></p> <p>For further information, send an email to <a href="http://plantl-gob-es@bsc.es/">plantl-gob-es@bsc.es</a></p> <p><strong>Funding</strong></p> <p>This work was funded by the <a href="https://portal.mineco.gob.es/en-us/digitalizacionIA/Pages/sedia.aspx">Spanish State Secretariat for Digitalization and Artificial Intelligence (SEDIA)</a> within the framework of the Plan-TL.</p> <p><strong>Copyright </strong></p> <p>Copyright (c) 2021 Secretaría de Estado de Digitalización e Inteligencia Artificial</p>
Biomedical Spanish CBOW Word Embeddings in Floret
<p><strong>Biomedical Spanish CBOW Word Embeddings in Floret</strong></p> <p>The embeddings have been trained with a biomedical Spanish corpus using <a href="https://arxiv.org/pdf/1607.04606.pdf">floret</a> with the following hyperparameters:</p> <blockquote> <p>mode: str = "floret",<br> model: str = "cbow",<br> dim: int = 300,<br> mincount: int = 10,<br> minn: int = 5,<br> maxn: int = 6,<br> neg: int = 10,<br> hashcount: int = 2,<br> bucket: int = 50000,<br> thread: int = 128,</p> </blockquote> <p>The embeddings were trained on the concatenation of all corpora from the <strong>Spanish biomedical corpus</strong> that includes Spanish data from various sources for a total of 1.1B tokens across 2,5M documents.</p> <table> <thead> <tr> <th scope="col">Source</th> <th scope="col">No. tokens</th> </tr> </thead> <tbody> <tr> <td>Medical crawler</td> <td>903,558,136</td> </tr> <tr> <td>Clinical cases misc.</td> <td>102,855,267</td> </tr> <tr> <td>EHRs documents<strong>*</strong></td> <td>95,267,204</td> </tr> <tr> <td>Scielo</td> <td>60,007,289</td> </tr> <tr> <td>BARR2 Background</td> <td>24,516,442</td> </tr> <tr> <td>Wikipedia (Life Sciences)</td> <td>13,890,501</td> </tr> <tr> <td>Patents</td> <td>13,463,387</td> </tr> <tr> <td>EMEA</td> <td>5,377,448</td> </tr> <tr> <td>Mespen (MedlinePlus)</td> <td>4,166,077</td> </tr> <tr> <td>PubMed</td> <td>1,858,966</td> </tr> </tbody> </table> <p>More information about the corpus can be found here <a href="https://aclanthology.org/2022.bionlp-1.19/">https://aclanthology.org/2022.bionlp-1.19/</a> and here <a href="https://arxiv.org/abs/2109.07765">https://arxiv.org/abs/2109.07765</a></p> <p>The processing took place on an HPC <a href="https://www.bsc.es/innovation-and-services/technical-information-cte-amd">node</a> equipped with an AMD EPYC 7742 (@ 2.250GHz) processor with 128 threads.</p> <p><strong>How to use</strong></p> <p>First initialize the spacy vectors from the floret table (.floret file):</p> <pre><code class="language-bash">spacy init vectors es floret_embeddings_bio_es.floret floret_embeddings_bio_es --mode floret</code></pre> <pre><code>import spacy # Load the floret vectors floret_embeddings = spacy.load("floret_embeddings_bio_es") # Get the embeddings of some words diabetes = floret_embeddings.vocab["diabetes"] insulina = floret_embeddings.vocab["insulina"] radiografia = floret_embeddings.vocab["radiografia"] # Get some similarities print(diabetes.similarity(insulina)) print(diabetes.similarity(radiografia)) # diabetes should be more similar to insuline than radiografia </code></pre> <p> </p> <p><strong>Intended Uses and Limitations</strong></p> <p>At the time of submission, no measures have been taken to estimate the bias and toxicity embedded in the model. However, we are well aware that our models may be biased since the corpora have been collected using crawling techniques on multiple web sources. We intend to conduct research in these areas in the future, and if completed, this card will be updated.</p> <p><strong>Authors</strong></p> <p>The Text Mining Unit from Barcelona Supercomputing Center.</p> <p><strong>Contact Information</strong></p> <p>For further information, send an email to <a href="http://plantl-gob-es@bsc.es">plantl-gob-es@bsc.es</a></p> <p><strong>Funding</strong></p> <p>This work was funded by the <a href="https://portal.mineco.gob.es/en-us/digitalizacionIA/Pages/sedia.aspx">Spanish State Secretariat for Digitalization and Artificial Intelligence (SEDIA)</a> within the framework of the Plan-TL.</p> <p><strong>Copyright </strong></p> <p>Copyright (c) 2022 Secretaría de Estado de Digitalización e Inteligencia Artificial</p>
Catalan CBOW Word Embeddings in Floret
<p><strong>Embeddings with the Catalan Textual Corpus</strong></p> <p>The embeddings have been trained with a Catalan textual corpus of more than 34GB of data using <a href="https://arxiv.org/pdf/1607.04606.pdf">floret</a> with the following hyperparameters:</p> <blockquote> <p> mode: str = "floret",<br> model: str = "cbow",<br> dim: int = 300,<br> mincount: int = 10,<br> minn: int = 5,<br> maxn: int = 6,<br> neg: int = 10,<br> hashcount: int = 2,<br> bucket: int = 50000,<br> thread: int = 128,</p> </blockquote> <p>The Catalan Textual Corpus used to train this embeddings, is the extended version of the initial available corpora described in <a href="https://arxiv.org/pdf/2107.07903.pdf">Armengol-Estapé et al. (2021)</a>. This new version includes:</p> <table> <caption> </caption> <thead> <tr> <th scope="col">Corpus</th> <th scope="col">Size in GB</th> </tr> </thead> <tbody> <tr> <td>CaCrawlat</td> <td>13.00</td> </tr> <tr> <td>Wikipedia</td> <td>1.10</td> </tr> <tr> <td>DOGC</td> <td>0.78</td> </tr> <tr> <td>Catalan Open Subtitles</td> <td>0.02</td> </tr> <tr> <td>Catalan Oscar</td> <td>4.00</td> </tr> <tr> <td>CaWaC</td> <td>3.60</td> </tr> <tr> <td>Cat. General Crawling</td> <td>2.50</td> </tr> <tr> <td>Cat. Goverment Crawling</td> <td>0.24</td> </tr> <tr> <td>ACN</td> <td>0.42</td> </tr> <tr> <td>Padicat</td> <td>0.63</td> </tr> <tr> <td>RacoCatalà</td> <td>8.10</td> </tr> <tr> <td>NacióDigital</td> <td>0.42</td> </tr> <tr> <td>VilaWeb</td> <td>0.06</td> </tr> </tbody> </table> <p>From the new corpora, VilaWeb and NacióDigital come from digital newspapers, Padicat is composed of crawlings of the Biblioteca de Catalunya, and CaCrawlat comes from the Biblioteca Nacional de España (BNE).</p> <p>The processing took place on an HPC <a href="https://www.bsc.es/innovation-and-services/technical-information-cte-amd">node</a> equipped with an AMD EPYC 7742 (@ 2.250GHz) processor with 128 threads.</p> <p><strong>How to use</strong></p> <p>First initialize the spacy vectors from the floret table (.floret file):</p> <pre><code class="language-bash">spacy init vectors ca floret_embeddings_ca.floret floret_embeddings_ca --mode floret</code></pre> <pre><code class="language-python">import spacy # Load the floret vectors floret_embeddings = spacy.load("floret_embeddings_ca") # Get the embeddings of some words castanyes = floret_embeddings.vocab["castanyes"] flors = floret_embeddings.vocab["flors"] primavera = floret_embeddings.vocab["primavera"] tardor = floret_embeddings.vocab["tardor"] # Get some similarities print(flors.similarity(tardor)) print(flors.similarity(primavera)) # flors should be more similar to primavera than tardor. print(castanyes.similarity(primavera)) print(castanyes.similarity(tardor)) # castanyes should be more similar to tardor than primavera.</code></pre> <p><strong>Intended Uses and Limitations</strong></p> <p>At the time of submission, no measures have been taken to estimate the bias and toxicity embedded in the model. However, we are well aware that our models may be biased since the corpora have been collected using crawling techniques on multiple web sources. We intend to conduct research in these areas in the future, and if completed, this card will be updated.</p> <p><strong>Authors</strong></p> <p>The Text Mining Unit from Barcelona Supercomputing Center.</p> <p><strong>Contact Information</strong></p> <p>For further information, send an email to <a href="mailto:aina@bsc.es">aina@bsc.es</a>.</p> <p><strong>Funding</strong></p> <p>This work was funded by the <a href="https://politiquesdigitals.gencat.cat/ca/inici/index.html">Departament de la Vicepresidència i de Polítiques Digitals i Territori de la Generalitat de Catalunya</a> within the framework of <a href="https://politiquesdigitals.gencat.cat/ca/economia/catalonia-ai/aina">Projecte AINA</a>.</p> <p><strong>Copyright</strong></p> <p>Copyright (c) 2022 Text Mining Unit - Barcelona Supercomputing Center.</p>
MEG data during the presentation of Gabor patterns and word sets
<p><strong>Subjects</strong></p> <p>MEG was recorded in 7 healthy subjects (5 young right-handed and 1 left-handed and 1 elderly) in the waking state with their eyes open or closed, who were sitting in a comfortable chair. The experimental technique was approved by the ethical commission of the Institute of Higher Nervous Activity and Neurophysiology of RAS (protocol No. 5 dated 02.12.2020).</p> <p><strong>Equipment</strong></p> <p>MEG was recorded on a VectorView device (Elekta Neuromag Oy, Finland), which was placed inside a magnetically protected chamber made of multilayer permalloy (AK3b, Vacuumschmelze GmbH, Germany). Before MEG recording, the coordinates of anatomical reference points (left and right preauricular points and nasion) were determined, as well as indicator coils attached to the surface of the scalp of the subject in the upper part of the forehead and behind the auricles. These points were determined using a FASTRAK 3D digitizer (Polhemus, USA). Each subject had a virtual model of the brain and head obtained from an anatomical 3D MRI taken the day before (file: MRI_V1_7.zip).</p> <p><strong>Registration and pre-processing</strong></p> <p>The subject's head was covered by a helmet, which is part of a fiberglass Dewar vessel with an array of sensors immersed in liquid helium. The subject sat down in such a way that the surface of the head was as close as possible to the sensors. The magnetic signal was recorded from 102 triplets, each of which consisted of 1 magnetometer and 2 gradiometers at rest with eyes closed and upon presentation of visual and speech stimuli. Recording was performed with a sampling frequency of 1000 Hz in a bandwidth of 0.1–330 Hz and was processed by the MaxFilter program (Elekta Neuromag Oy, Finland), which eliminates artifacts (the tSSS method—spatio-temporal separation of signals). The signal levels were corrected in accordance with the data on the position of the subject's head in relation to the MEG sensors. The position of the head during the experiment was controlled using special inductors.</p> <p><strong>Visual and verbal stimuli</strong></p> <p>After recording the background MEG for 3 minutes with closed eyes, the subject opened his eyes on command and observed the fixation point on the projection screen. After 15 seconds, stimulation was started and the subject's responses were received in the form of pressing a button. In response to the 0 degrees and 90 degrees stimuli, the subject had to press the button with the index finger, and to the 45 degrees and 135 degrees inclined stimuli, the adjacent button with the middle finger. Stimuli lasting 100 ms were presented randomly every 3100±100 ms (intervals between stimuli varied randomly). In two series, 42 stimuli of each orientation were presented. Between the series, the subject rested for 2-3 minutes. Visual stimuli in the form of Gabor contrast gratings (1.9 cycles per angular degree) with dimensions of 5.25 angular degrees and an average brightness of 4 lux were projected onto a screen located at a distance of 95 cm from the subject's eyes using a Panasonic PT-stimulating projector D7700E-K, which is part of the MEG facility. Visual stimulus patterns were generated at http://www.cogsci.nl/pages/gabor-generator with edge parameters: Circular (sharp edge). The samples are contained in the GaborStim.zip file (the names of the sample files correspond to their name in the script file, but do not match their geometric meaning, see table below). The stimulator was programmed using the Presentation software (USA, Neurobehavioral Systems, Inc). Stimulation scripts are contained in the sce.zip file.</p> <p><strong>Table</strong></p> <p><em>Stimulus or response code Type of stimulus or response</em></p> <p>STI101_1 Fixation point<br> STI101_2 90 degrees<br> STI101_4 135 degrees</p> <p>STI101_8 0 degrees<br> STI101_16 45 degrees<br> STI101_32 First button (index finger)</p> <p>STI101_64 Second button (middle finger)</p> <p>After 2 series of visual stimuli, the subject closed his eyes and was presented with 3 series of speech stimuli for 2 minutes with a break of 1 minute. In each series, recordings of audio files of 8 separate adjectives of the Russian language were presented, which were repeated 5 times in a pseudo-random order. The series began with 3 words, which were not taken into account in further analysis. The subject listened to the words and had to press the button after he understood the meaning of the presented word. After pressing or no response, the next word followed in 2±1 s. The audio files are contained in the words101_343.zip file (the names correspond to the script file).</p> <p><strong>Data received</strong></p> <p>The records are contained in files with the name of the type V1m24r, where V1 is the number of the subject, m is the sex (m/f), 24 is the age, and r is the right-handed subject. This dataset can be easily loaded into the Brainstorm program. Spontaneous and evoked MEG can be used for source localization and reconstruction of traveling waves.</p> <p><strong>Acknowledgments</strong></p> <p>The reported study was funded by RFBR, project number 20-015-00475.</p>
Ancient Greek Fasttext Word Embeddings
<p>Word embeddings generated with Fasttext and 1 GB of Ancient Greek texts. These embeddings were produced for the study of social networks and social semantics in ancient Greece by the Diogenet project at the University of San Diego, California. </p>
ArabicSL-Net: A Benchmark Video Dataset for Arabic Words Sign Language
<p>The data was captured by mobile camera in four main organization namely Bank , Cafe , Hospital , and Train station. The ArabicSL-Net initially consists of 307 words recorded in approximately 30,000 videos. For each organization, we capture the most representative words that are used in those places. For Bank data, we have a total of 76 of words, while Cafe data contains 54 words. For Hospital, we collects videos for 102 words, and collects videos for 71 words in Train station.</p>
Supplementary material for "Using a parallel corpus to study patterns of word order variation: Determiners and quantifiers within the noun phrase in European languages"
<p>- output-{ciep,treebanks}-full.csv: frequency and entropy for all the categories, using four types of combinations of layers;<br> - plots.R: R script to draw plots from the output files;<br> - readReport-{CIEP+,treebanks}.R: R script to extract frequency and compute entropy from the report files (not included);<br> - ud-wordorder.py: Python script to extract word order pairs from conllu files and write them in report files.</p> <p>Unfortunately, I cannot include the report files, as CIEP+ is protected by copyright; the analysis can be however replicated with respect to the UD Treebanks.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.