Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

27

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

27 results for “word2vec”

Learn how ShareScore rates datasets ↗
zenodo28/100

Study of terminological subsystems of modern school textbooks in Russian with the help of word embedding models Word2Vec and neural networks

<p>The aim of the project is to analyse the inventory and functioning of scientific terms and special lexemes in textbooks for secondary schools of the Russian Federation with the help of modern methods of natural language processing and deep learning. The number of terms from different fields of knowledge that a pupil should learn during secondary school studies has never been evaluated. According to the preliminary evaluations made on the basis of the Model Basic Curriculum for General and Secondary Education in 2015 only the subject &quot;Russian language&quot; presupposes that a pupil finishing the 11th grade of secondary school should be able to understand, recognise and use about 1000 terms and terminological combinations. Thus, taking into account the number of school subjects, the total number of special vocabulary units studied in general education schools is measured in thousands. At the same time, the comparative characteristics of the inventory and functioning of terms in textbooks for different school subjects are not studied and remain unknown. The correlation between the terminological density of the text in school textbooks for different subjects and the place occupied by these subjects in the curriculum is not clear. The traditional way of compiling lists of scientific terms is simply by gleaning them from special texts and writing down manually. If this method is reliable in terms of intellectualisation of selection principles, it cannot be applied to large data sets and does not reflect either the frequency of use of terms, or the specificity of their syntagmatic connections, or the systemic relationship between terms. The current project is aimed at filling this gap by means of 1) creating a full-text corpus of school textbooks for 5&ndash;11 classes included in the Federal List compiled by the Ministry of Education, 2) automatic extraction, stratification, and mapping of terms with the help of distribution semantics algorithms, 3) creation and training of a deep neural network capable of predicting the subject, level of education and educational topic given a group of vector representations of terms as input. The results of the research can be of fundamental interest in the perspective of terminology science development and also have practical applications in the creation of different types of educational literature.</p> <p><em>Funding: The reported study was funded by RFBR, project number 19-29-14032</em></p>

opencc-by-4.0Oct 2020View details →
zenodo28/100

Study of terminological subsystems of modern school textbooks in Russian with the help of word embedding models Word2Vec and neural networks

<p>The reported study was funded by RFBR, project number 19-29-14032 mk.</p>

opencc-by-4.0Oct 2020View details →
zenodo28/100

word2vec

<p>Trained word2vec model</p>

opencc-by-4.0Nov 2022View details →
zenodo28/100

Study of terminological subsystems of modern school textbooks in Russian with the help of word embedding models Word2Vec and neural networks

<p>The reported study was funded by RFBR, project number 19-29-14032 mk.</p>

opencc-by-4.0Oct 2020View details →
zenodo20/100

VUDENC - python corpus for word2vec

<p>Python corpus for training a word2vec model, and one trained model.</p>

opencc-by-4.0Dec 2019View details →
zenodo12/100

Pre-trained word2vec models for ``Easy over Hard: A Case Study on Deep Learning''

<p>Since the whole stack overflow dump is so big, we can't easily handle well. Here, we provide 10 pre trained word2vec models with different seeds.</p> <p> </p> <p>More details about how to use it, please see paper </p>

restrictedMar 2017View details →
zenodo8/100

Pre-trained word2vec model for ``Easy over Hard: A Case Study on Deep Learning''

<p>Since the whole stack overflow dump is so big, we can't easily handle well. Here, we provide 10 pre trained word2vec models with different seeds using skip-gram algorithms</p> <p> </p> <p>More details about how to use it, please see paper </p>

restrictedMar 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record