Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
19
datasets available to search
ShareScore release 0.7.1
Dataset results
19 results for “Thesaurus”
HiMAT Thesaurus for Mining Research
<p>This datasets contains the Thesaurus for HiMAT-datasets including informations related to skos-concepts and the triples of the Thesaurus' RDF representation.</p>
Irish Traditional Music Instruments Thesaurus, Extended Version
<p>A Simple Knowledge Organisation System (SKOS) thesaurus. Incorporates extended instruments used in contemporary Irish traditional music. Developed for use at the Irish Traditional Music Archive. Contains Irish language and English terms.</p>
Irish Traditional Music Tune Types Thesaurus
<p>A Simple Knowledge Organisation System (SKOS) Thesaurus. Incorporates dance tunes and other tune types found in contemporary Irish traditional music. Developed for use at the Irish Traditional Music Archive. Contains Irish language and English terms.</p>
Irish Traditional Music Instruments Thesaurus, Core Version
<p>A Simple Knowledge Organisation System (SKOS) thesaurus. Incorporates core instruments used in contemporary Irish traditional music. Developed for use at the Irish Traditional Music Archive. Contains Irish language and English terms.</p>
The US LTER Thesaurus: Contents and Keyword Use Statistics in LTER Data Packages in 2006 and 2018
This dataset contains raw data and statistical summaries that reflect use of keywords in LTER Datasets in May 2018 and 2006. Specific summaries include: Number of uses and number sites by keyword (LTERVocabKeywordSummary.csv), Summary of keyword use by data package (LTERVocabDataPackageSummary.csv), Summary of Keyword Use by LTER Site in 2018(LTERVocabSiteSummary.csv), Summary of Keyword Use by LTER Site in 2006(KeyStats2006.csv). Raw data includes XML files containing the US LTER Thesaurus in Moodle format and the ResultSet containing the information for each dataset from the Environmental Data Initiative PASTA repository.
Crowdsourcing thesaurus
<p>Thesaurus used for text mining analysis on a corpus about digital libraries and crowdsourcing. It contains ideologies, taxonomies and motivations.</p> <p>Analysis were used for a French article : Andro, M. (2016). Bibliothèques numériques et crowdsourcing : analyses bibliométriques et text mining</p>
Russian Distributional Thesaurus (RDT): Word Embeddings
<p>This resource is a part of the Russian Distributional Thesaurus (RDT): see http://russe.nlpub.ru/downloads and http://nlpub.ru/RDT. </p> <p>This dataset contains a large scale word embeddings model for Russian trained using the SGNS model (Mikolov et al., 2013) on a 12.9 billion word collection of books in Russian. According to the results of our participation in the shared task on Russian semantic similarity (Panchenko et al., 2015), this approach scored in the top 5 among 105 submissions (Arefyev et al., 2015). Following our prior experiments (Arefyev et al., 2015) we have selected the following parameters for the model: minimal word frequency – 5, number of dimensions in a word vector – 500, three or five iterations of the learning algorithm over the input corpus, context window size of 1, 2, 3, 5, 7 and 10 words. Parameters of the model are listed below:</p> <ul> <li>Model: skip-gram</li> <li>Corpus: a 150Gb sample of the lib.rus.ec book collection.</li> <li>Context window size: 10 words</li> <li>Number of dimensions: 500</li> <li>Number of iterations: 3</li> <li>Minimal word frequency: 5</li> </ul> <p>References:</p> <ul> <li>Panchenko A., Ustalov D., Arefyev N., Paperno D., Konstantinova N., Loukachevitch N. and Biemann C. (2016): Human and Machine Judgements about Russian Semantic Relatedness. In Proceedings of the 5th Conference on Analysis of Images, Social Networks, and Texts (AIST'2016). Communications in Computer and Information Science (CCIS). Springer-Verlag Berlin Heidelberg</li> </ul> <ul> <li>Panchenko A., Loukachevitch N. V., Ustalov D., Paperno D., Meyer C. M., Konstantinova N. (2015): RUSSE: The First International Workshop on Russian Semantic Similarity. In Proceedings of the 21st International Conference on Computational Linguistics and Intellectual Technologies (Dialogue'2015). Moscow, Russia. RGGU</li> </ul> <ul> <li>Arefyev N., Panchenko A., Lukanin A., Lesota O., Romanov P. (2015): Evaluating Three Corpus-Based Semantic Similarity Systems for Russian. In Proceedings of the 21st International Conference on Computational Linguistics and Intellectual Technologies (Dialogue'2015). Moscow, Russia. RGGU</li> </ul>
The Earth Surface System Scientific Data Thesaurus
<p>The Earth Surface System Scientific Data Thesaurus</p>
Thesaurus for KeyWords Plus for MEJ-24 2000-2019
<p>This is a supplemental file for the STI2024 submission "Delineating the field of medical education over time: A case study on interdisciplinarity and interuniversity collaboration patterns, 2000-2019". This file can be used when generating the KeyWords Plus networks of the Web of Science data outputs for the MEJ-24 2000-2019.<br> </p>
Thesaurus iconographique Biblissima
<p>Ce dépôt met à disposition le thésaurus iconographique du <a href="https://portail.biblissima.fr">portail Biblissima</a> conformément au modèle <a href="https://www.w3.org/TR/skos-reference/">SKOS</a> (formats RDF/XML et Turtle). Ce thésaurus est constitué à partir des données des bases partenaires du projet <a href="https://projet.biblissima.fr">Biblissima</a>.</p> <p>Il est actuellement utilisé dans le portail Biblissima pour permettre la recherche et la navigation dans un corpus de plus de 306 000 notices d'enluminures et décors de manuscrits et imprimés anciens issues de deux bases iconographiques : Mandragore (BnF) et Initiale (IRHT-CNRS).</p> <p>Voir sur le portail Biblissima :<br> - la page <a href="https://portail.biblissima.fr/fr/ark:/43093/thb806db559f2abfe3bd6884def6909c7329f5a183">Thesaurus iconographique</a><br> - <a href="https://portail.biblissima.fr/fr/iconography">l'interface d'exploration et de visualisation de l'iconographie</a></p> <p>Les concepts de ce thésaurus sont stockés et gérés dans la plateforme des <a href="https://data.biblissima.fr">référentiels d'autorité de Biblissima</a> s'appuyant sur le logiciel Wikibase.</p> <p>Pour plus d'informations sur la constitution du référentiel des descripteurs iconographiques, voir <a href="https://data.biblissima.fr/w/R%C3%A9f%C3%A9rentiel_des_descripteurs_iconographiques">cette page de présentation</a>.</p>
The European Language Social Science Thesaurus (ELSST)
<p>The European Language Social Science Thesaurus (ELSST) is a broad-based, multilingual thesaurus for the social sciences. It is owned and published by the Consortium of European Social Science Data Archives (CESSDA) and its national Service Providers. The thesaurus consists of over 3,400 concepts and covers the core social science disciplines: politics, sociology, economics, education, law, crime, demography, health, employment, information and communication technology, and environmental science.</p> <p>ELSST is used for data discovery within CESSDA and facilitates access to data resources across Europe, independent of domain, resource, language or vocabulary.</p> <p><strong><em>Recommended Citation</em></strong>: CESSDA and Service Providers (2024) The European Language Social Science Thesaurus (ELSST) (Version 5), <a href="https://elsst.cessda.eu">https://elsst.cessda.eu</a>. DOI: 10.5281/zenodo.13843400</p>
FIGURE 2. A in ZooNom: an online thesaurus for alleviating ambiguity in the terminology of zoological nomenclature
FIGURE 2. A screenshot of ZooNom taken on 10/11/2021. The concept shown is "doxisonym" (URL: https://www.loterre.fr/ skosmos/FM8/en/page/-PD692RFQ-4). The "note" field contains the etymology, and the "scope note" field presents the Code's equivalent (objective synonym) and definition.
Thesaurus for Sustainable Development Goals
Open the record for dataset details and reuse information.
EuroVoc and thesauri from EU institutions and agencies Interoperability and perspectives for collaborative thesaurus management
<p>The EuroVoc thesaurus maintained by the EU Publications Office is available in 24 languages and is focused on the main areas of activity of the EU. Like EuroVoc, a number of multilingual specialised thesauri or controlled vocabularies are also maintained and disseminated independently with their own thesaurus management systems, inside the EU institutions and agencies. Thesaurus Alignment carried out in the Publications Office has shown that a number of concepts and terms are shared and duplicated amongst the EU thesauri. </p><p>With the objective of enhanced collaboration between EU multilingual thesauri, the Publications Office proposes to offer a collaborative environment for thesaurus maintenance, dissemination and alignment for EU institutions and agencies. The collaboration will enable mutual enrichment of our thesauri in terms of coverage, languages and the generation of Linked Data. Additionally, the project will help with reducing the efforts and costs of thesaurus maintenance, translation, hosting and development.</p>
CONSTRUCTION TECHNOLOGY AND SEMANTIC AREAS OF THE UZBEK-ENGLISH THESAURUS ON PHARMACY TERMS
Open the record for dataset details and reuse information.
Thesaurus Linguae Sericae (TLS) as interlinked Markdown files
Open the record for dataset details and reuse information.
Sino-Tibetan Etymological Dictionary and Thesaurus Database Software
Open the record for dataset details and reuse information.
FIGURE 1. A in ZooNom: an online thesaurus for alleviating ambiguity in the terminology of zoological nomenclature
FIGURE 1. A simplified visualization of the structure of the thesaurus, with the example of the concept "Onomatophore". The color gets darker for every entity contained ("narrower") in another.
NASA Thesaurus
The NASA Thesaurus contains the authorized NASA subject terms used to index and retrieve materials in the NASA Technical Reports Server (NTRS) and the NTRS Registered (Formerly NA&SD). The scope of this controlled vocabulary includes not only aerospace engineering, but all supporting areas of engineering and physics, the natural space sciences (astronomy, astrophysics, planetary science), Earth sciences, and the biological sciences. The NASA Thesaurus contains over 18,400 subject terms, 4,300 definitions, and more than 4,500 USE cross references.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.