Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
27
datasets available to search
ShareScore release 0.7.1
Dataset results
27 results for “word2vec”
Study of terminological subsystems of modern school textbooks in Russian with the help of word embedding models Word2Vec and neural networks
<p>The aim of the project is to analyse the inventory and functioning of scientific terms and special lexemes in textbooks for secondary schools of the Russian Federation with the help of modern methods of natural language processing and deep learning. The number of terms from different fields of knowledge that a pupil should learn during secondary school studies has never been evaluated. According to the preliminary evaluations made on the basis of the Model Basic Curriculum for General and Secondary Education in 2015 only the subject "Russian language" presupposes that a pupil finishing the 11th grade of secondary school should be able to understand, recognise and use about 1000 terms and terminological combinations. Thus, taking into account the number of school subjects, the total number of special vocabulary units studied in general education schools is measured in thousands. At the same time, the comparative characteristics of the inventory and functioning of terms in textbooks for different school subjects are not studied and remain unknown. The correlation between the terminological density of the text in school textbooks for different subjects and the place occupied by these subjects in the curriculum is not clear. The traditional way of compiling lists of scientific terms is simply by gleaning them from special texts and writing down manually. If this method is reliable in terms of intellectualisation of selection principles, it cannot be applied to large data sets and does not reflect either the frequency of use of terms, or the specificity of their syntagmatic connections, or the systemic relationship between terms. The current project is aimed at filling this gap by means of 1) creating a full-text corpus of school textbooks for 5–11 classes included in the Federal List compiled by the Ministry of Education, 2) automatic extraction, stratification, and mapping of terms with the help of distribution semantics algorithms, 3) creation and training of a deep neural network capable of predicting the subject, level of education and educational topic given a group of vector representations of terms as input. The results of the research can be of fundamental interest in the perspective of terminology science development and also have practical applications in the creation of different types of educational literature.</p> <p><em>Funding: The reported study was funded by RFBR, project number 19-29-14032</em></p>
Study of terminological subsystems of modern school textbooks in Russian with the help of word embedding models Word2Vec and neural networks
<p>The reported study was funded by RFBR, project number 19-29-14032 mk.</p>
word2vec
<p>Trained word2vec model</p>
Study of terminological subsystems of modern school textbooks in Russian with the help of word embedding models Word2Vec and neural networks
<p>The reported study was funded by RFBR, project number 19-29-14032 mk.</p>
VUDENC - python corpus for word2vec
<p>Python corpus for training a word2vec model, and one trained model.</p>
Pre-trained word2vec models for ``Easy over Hard: A Case Study on Deep Learning''
<p>Since the whole stack overflow dump is so big, we can't easily handle well. Here, we provide 10 pre trained word2vec models with different seeds.</p> <p> </p> <p>More details about how to use it, please see paper </p>
Pre-trained word2vec model for ``Easy over Hard: A Case Study on Deep Learning''
<p>Since the whole stack overflow dump is so big, we can't easily handle well. Here, we provide 10 pre trained word2vec models with different seeds using skip-gram algorithms</p> <p> </p> <p>More details about how to use it, please see paper </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.