Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
132
datasets available to search
ShareScore release 0.7.1
Dataset results
132 results for “vocabularies”
Vanuatu Basic Vocabulary List
<p>The Vanuatu Basic Vocabulary List was developed by and for linguists working on the Vanuatu Languages and Lifeways Project at the Max Planck Institute for the Science of Human History. It is used as a tool for eliciting basic vocabulary data of Vanuatu languages, and therefore both English and Bislama prompts are provided. The list of 215 concepts is based entirely on the Austronesian Basic Vocabulary Database (Greenhill et al 2008), with only minor additions in the pronominal domain that account for both inclusivity and exclusivity, as well as dual and plural forms.</p> <p>References:</p> <p>Greenhill, S.J., Blust. R, & Gray, R.D. (2008). <a href="https://abvd.shh.mpg.de/publications/index.php?pub=Greenhill_et_al2008">The Austronesian Basic Vocabulary Database: From Bioinformatics to Lexomics</a>. <em><a href="http://www.la-press.com/article.php?article_id=1129">Evolutionary Bioinformatics</a></em>, 4:271-283.</p>
Wind energy taxonomies and restricted vocabularies
<p>This is the update version of the wind energy taxonomies and restricted vocabularies which has been implemented in DTU Data (https://data.dtu.dk/DTU_Wind_Energy) with the purpose of accurately describing published data sets and data collections. The work on updating and implementing the wind energy taxonomies and restricted vocabularies has been done as a part of an internally funded 'FAIR digitalization' project of DTU Wind Energy(see 10.5281/zenodo.1493874). The work on updating the taxonomies and restricted vocabularies is a continuation of the work previously done under the IRPWind Open Data initiative (see 10.5281/zenodo.1199489). </p>
Archaeology Vocabulary DEU-ENG
<p>This collection of German-English archaeology vocabulary was created in the course of translating archaeological, art historical, historical academic scholarship from German to English for a variety of German and Austrian archaeological institutes and museums. All words underwent extensive research to make sure the words are correct and in some cases references are also included. Not all words have a direct translations.</p> <p>Subjects covered include field archaeology, architecture, pottery, methods, fortification, and many more.</p> <p>The collections archaeology_terminology and term_list_archaeology were originally created in SDL Multiterm 2015 and 2017 and then combined to create archaeology_terminology_complete which was also cleaned in OpenRefine 3.1</p> <p>File formats uploaded here include xlsx and MultiTerm Termbase.</p>
Lexis and tradition: variation in the vocabulary of Sanskrit Mahāyāna literature - datasets
<p>Lexical datasets containing annotated concordances of words pertaining to the conceptual domains of language and conceptualisation in Buddhist Sanskrit Literature. The smaller dataset contains linguistic annotations, the larger only metadata. The concordances have been taken from the segmented Sanskrit corpus 10.5281/zenodo.3526665.</p> <p>These datasets have been created as part of the project 'Lexis and Tradition: variation in the vocabulary of Sanskrit Mahāyāna literature', funded by the British Academy through a Newton International Fellowship (NF161436) and hosted at the Department of Theology and Religious Studies at King's College London under the supervision of Prof. Henrietta Kate Crosby. </p> <p>Dr. Bruno Galasek-Hul and Luis Quiñones have assisted me with semantic annotations thanks to funding from the Mangalam Research Center.</p> <p>The repository also contains R scripts for text clustering on the basis of the lexical data provided.</p> <p>the annotated dataset can be interactively explored at:</p> <ol> <li> <a href="https://ligeialugli.shinyapps.io/VisualDictionaryOfBuddhistSanskrit/">https://ligeialugli.shinyapps.io/VisualDictionaryOfBuddhistSanskrit/</a></li> <li> <a href="https://ligeialugli.shinyapps.io/VisualDictionaryOfBuddhistSanskrit/">https://ligeialugli.shinyapps.io/VisualThesaurusOfBuddhistSanskrit/</a></li> </ol> <p>Updated versions 1.1-1.2 correct some mistakes in metadata and add a few new lemmata.</p> <p> </p>
Usage and Impact Vocabularies Peer Review Dataset
<p>This dataset reflects responses received via the Qualtrics data collection instrument provided to usage and impact vocabulary stakeholders to solicit peer review on a draft version of the glossary and crosswalk spreadsheet developed for the "EAGER: Secure Research Impact Metric Data Exchange: Data Supply Chain" project funded by The National Science Foundation (Award # 2335827).</p>
PDF4 - Semantics and Vocabularies Hackathon SWOT Visualization
<p>This SWOT visualization outlines the community's feedback regarding the overall polar RDM strengths, weaknesses, opportunities, and threats. This visualization is valuable to convey the overarching themes/ topics of interest, and allows for stakeholders to determine where resources should be allocated. The content for this visualization was compiled through community input during the 'Semantics and Vocabularies' hackathon at the 4th Polar Data Forum in September 2021. </p>
“Tear down that wall” – updating the vocabulary of phage and bacterial lytic proteins
Open the record for dataset details and reuse information.
Proto-Western Kho-Bwa: regionally relevant basic vocabulary elicitation list
<p>This upload contains a pdf and a word file of a 555-entry word list that was used to elicit data from eight Western Kho-Bwa varieties (Khispi, Duhumbi, Khoina, Khoitam, Jerigaon, Rahung, Rupa and Shergaon), Bangru, Brokpa and Tshangla in Western and Central Arunachal Pradesh between March 2012 and November 2018. </p> <p>The word list contains concepts that are also found in the most commonly used lists for vocabulary elicitation, but also contains regionally relevant concepts, for example, those related to flora and fauna, local livelihood practices and agricultural crops, and cultural and religious concepts. The word list is partially translated in Roman Hindi to ease elicitation with Hindi-speaking consultants, but certain concepts for which Hindi equivalents could not be found have been translated into Tshangla, the erstwhile lingua franca in some areas and a language still spoken by the older generation.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: bodttim (at) gmail (dot) com</p>
Test Data from a Study on Latin Vocabulary Acquisition with Beginners (Textbook)
<p>The dataset contains test results from an intervention study with beginners in the 5th grade of a high school in Berlin. In total, 103 students participated in four groups (= classes). The intervention materials and tests are published as well.</p> <p>A key question of the still ongoing research project is: How can vocabulary competence in a historical language such as Latin be acquired and deepened by using corpus-based, i.e. context-based, methods? This question is based on a broad understanding of vocabulary that refers back to theories of the mental lexicon.</p>
Bharathi - Linked Data Vocabulary for the Indian Context
<p>Bharathi is a collection of linked data vocabularies for the Indian context linking the metadata of terms and inter relationships across different entities within Government functions. In its first release, Bharathi contains information regarding Government Organisations at the Union Government and State Government level, the various Administrative Regions and its respective classifications, Sectors, Sub Sectors and common topics within the hierarchy of the Government functions</p>
NCAS Instrument Vocabulary
Controlled Vocabularies of NCAS instruments
Synthetic Out-Of-Vocabulary IAM dataset with Latent Diffusion Models
<p>SyntheticHTR: Handwritten Text Image Synthesis based on Latent Diffusion Models</p>
Vocabulary of AI Risks (VAIR)
Open the record for dataset details and reuse information.
Supplementary methods and data for: Dogs with a vocabulary of object label remember labels for at least two years
<p>Long-term memory of words has a crucial role in the developing abilities of young children to acquire language. In dogs, the ability to learn object labels is present in only a small group of uniquely Gifted Word Learner (GWL) dogs. The ability of these dogs to acquire large vocabularies consisting of hundreds of names of dog toys through naturally occurring interactions in human families presents them as a valid model for studying language-related cognitive mechanisms. As they are very rare, little is known about the mechanisms through which they acquire such large vocabularies. In the current study, we tested the ability of five GWL dogs to retrieve 12 labelled objects two years after the object-label mapping acquisition. The dogs proved to remember the labels of between 3-9 objects. The results shed light on the process by which GWL dogs acquire an exceptionally large vocabulary of object names. As memory plays a crucial role in language development, these dogs supply a unique opportunity to study label retention in a non-linguistic species.</p>
Controlled vocabularies and knowledge organisation for Digital Humanities - Part I
<p>The workshop "Controlled vocabularies and knowledge organisation for the digital humanities" aimed to exchange experiences and applications of controlled vocabularies for research and projects in the field of digital humanities, including research infrastructures, libraries and other cultural institutions linked to social sciences, arts and humanities (SSAH).</p> <p>This is part I, presented by: Helen Goulis (DARIAH Thesaurus Maintenance WG); José Moreiro González (University Carlos III of Madrid); Bruno Almeida (ROSSIO Infrastructure / NOVA CLUNL); Filipa Medeiros (Art Library and Archives of the Calouste Gulbenkian Foundation).</p> <p>This online workshop was organized by ROSSIO Infrastructure, Department of Linguistics of NOVA FCSH, NOVA CLUNL and Art Library and Archives of the Calouste Gulbenkian Foundation (FCG), and took place on the 12th of July 2021.</p>
Controlled vocabularies and knowledge organisation for Digital Humanities - Part II
<p>The workshop "Controlled vocabularies and knowledge organisation for the digital humanities" aimed to exchange experiences and applications of controlled vocabularies for research and projects in the field of digital humanities, including research infrastructures, libraries and other cultural institutions linked to social sciences, arts and humanities (SSAH).</p> <p>This is part II, presented by: Sébastien Durost (BIBRACTE EPCC), Guillaume Reich (MSHE Ledoux) & Jean Pierre Girard (UMR Archéorient); Susana Medina (FEUP); Ana Paula Figueiredo (Directorate-General Cultural Heritage of Portugal); Teresa Borges (Portuguese Film Archives).</p> <p>This online workshop was organized by ROSSIO Infrastructure, Department of Linguistics of NOVA FCSH, NOVA CLUNL and Art Library and Archives of the Calouste Gulbenkian Foundation (FCG), and took place on the 12th of July 2021.</p>
Ugric vocabulary (appendix to Grünthal et al. 2022: Drastic demographic events triggered the Uralic spread)
<p>Ugric cognates (Appendix to the paper Grünthal, R., Heyd, V., Holopainen, S., Janhunen, J., Khanina, O., Miestamo, M., Nichols, J., Saarikivi, J. & Sinnemäki, K. 2022: Drastic demographic events triggered the Uralic spread. – Diachronica. https://doi.org/10.1075/dia.20038.gru)</p> <p> </p> <p>Words found only in Hungarian and Khanty and/or Mansi.</p> <p> </p> <p>Further work on Ugric etymologies (with updates to the information on this table) will be published on https://sanat.csc.fi/wiki/Hungarian_Historical_Phonology</p>
ERA Vocabulary
<p>This is the human and machine readable Vocabulary/Ontology governed by the <a href="https://www.era.europa.eu/">European Union Agency for Railways</a>. It represents the concepts and relationships linked to the sectorial legal framework and the use cases under the Agency´s remit, as described in the <a href="http://data.europa.eu/eli/reg_impl/2019/777/oj">Commission Implementing Regulation (EU) 2019/777 of 16 May 2019 on the common specifications for the register of railway infrastructure and repealing Implementing Decision 2014/880/EU</a>.</p> <p>Currently, this vocabulary covers the European railway infrastructure and the vehicles authorized to operate over it. It is a semantic/browsable representation of the <a href="https://www.era.europa.eu/sites/default/files/registers/docs/rinf_application_guide_for_register_en.pdf">RINF application guide</a> and <a href="https://www.era.europa.eu/sites/default/files/registers/docs/iu-eratv_application_guide_for_register_2016-797_en.pdf">ERATV</a> application guides that were built by domain experts in the RINF and ERATV working parties.</p> <p>The vocabulary also includes the routebook concepts described in appendix D2 "Elements the infrastructure manager has to provide to the railway undertaking for the Route Book" as presented in the <a href="https://eur-lex.europa.eu/eli/reg_impl/2019/773/oj">Comission Implementing Regulation (EU) 2019/773 of 16 May 2019 on the technical specification for interoperability relating to the operation and traffic management subsystem of the rail system within the European Union and repealing Decision 2012/757/EU</a> and the appendix D3 "ERTMS trackside engineering information relevant to operation that the infrastructure manager shall provide to the railway undertaking".</p>
Data Privacy Vocabulary (DPV)
<h1>Data Privacy Vocabulary (DPV)</h1> <p>The <a href="https://w3id.org/dpv" rel="nofollow">Data Privacy Vocabulary (DPV)</a> provides an ontology (classes and properties) and taxonomies of concepts to represent information regarding how personal data is processed in the form of an ontology or a knowledge graph. For example, it provides taxonomies associated with:</p> <ul> <li>purposes of processing</li> <li>personal data categories involved</li> <li>processing operations</li> <li>technical and organisational measures or restrictions applied</li> <li>legal basis used to justify processing</li> <li>information about legal basis for processing</li> <li>rights as applicable</li> <li>risks as applicable</li> </ul> <p>The namespace for DPV terms is <code>http://w3id.org/dpv#</code> with suggested prefix <code>dpv</code>, and serialisations are provided in RDF/XML, Turtle, JSON-LD, and N3 formats. The default serialisations are defined using RDFS/SKOS semantics, with an <a href="https://w3id.org/dpv/dpv-owl" rel="nofollow">alternate serialisation</a> defined using OWL2 semantics.</p> <div> <h2>Extensions</h2> <a href="https://github.com/w3c/dpv#extensions"></a>These extensions provide additional concepts that extend the concepts and scope of the main DPV specification:</div> <ul> <li><a href="https://w3id.org/dpv/pd" rel="nofollow">Personal Data (PD)</a> provides a taxonomy of personal data categories</li> <li><a href="https://w3id.org/dpv/loc" rel="nofollow">Location (LOC)</a> provides a taxonomy of location concepts based on ISO 3166 (countries, regions)</li> <li><a href="https://w3id.org/dpv/tech" rel="nofollow">Technology (TECH)</a> provides a taxonomy of technology concepts</li> <li><a href="https://w3id.org/dpv/ai" rel="nofollow">AI</a> provides a taxonomy of AI concepts extending the TECH extension</li> <li><a href="https://w3id.org/dpv/justifications" rel="nofollow">Justifications</a> provides concepts for representing justifications i.e. why something must be done or could not be done</li> <li><a href="https://w3id.org/dpv/risk" rel="nofollow">Risk</a> provides concepts for risk assessment and management</li> </ul> <div> <h2>Extensions for Jurisdictions and Regulations</h2> <a href="https://github.com/w3c/dpv#extensions-for-jurisdictions-and-regulations"></a>The legal extensions provide concepts associated with specific jurisdictions and the laws, authorities, and treaties within them. The <a href="https://w3id.org/dpv/legal" rel="nofollow">Legal</a> page provides an overview of these. The jurisdictions are represented by using their ISO 3166-2 codes.</div> <ul> <li><a href="https://w3id.org/dpv/legal/eu" rel="nofollow">European Union (EU)</a> <ul> <li><a href="https://w3id.org/dpv/legal/eu/gdpr" rel="nofollow">GDPR</a></li> <li><a href="https://w3id.org/dpv/legal/eu/dga" rel="nofollow">DGA</a></li> <li><a href="https://w3id.org/dpv/legal/eu/aiact" rel="nofollow">AI Act</a></li> <li><a href="https://w3id.org/dpv/legal/eu/rights" rel="nofollow">Charter of Fundamental Rights</a></li> </ul> </li> <li><a href="https://w3id.org/dpv/legal/de" rel="nofollow">Germany (DE)</a></li> <li><a href="https://w3id.org/dpv/legal/ie" rel="nofollow">Ireland (IE)</a></li> <li><a href="https://w3id.org/dpv/legal/in" rel="nofollow">India (IN)</a></li> <li><a href="https://w3id.org/dpv/legal/gb" rel="nofollow">United Kingdom (GB)</a></li> <li><a href="https://w3id.org/dpv/legal/usa" rel="nofollow">United States of America (USA)</a></li> </ul> <div> <h2>Acknowledgements and Citation</h2> <ul> <li>For use of DPV from v2 onwards, <strong>Cite as:</strong> <a href="https://arxiv.org/abs/2404.13426" rel="nofollow">Data Privacy Vocabulary (DPV) -- Version 2</a> by Harshvardhan J. Pandit, Beatriz Esteves, Georg P. Krog, Paul Ryan, Delaram Golpayegani, Julian Flake <a href="https://arxiv.org/abs/2404.13426" rel="nofollow">https://arxiv.org/abs/2404.13426</a> (2024)</li> </ul> </div> <ul> <li>For use of DPV up to v1 and v1.1, <strong>Cite as:</strong> The peer-reviewed article “<a href="https://link.springer.com/chapter/10.1007%2F978-3-030-33246-4_44" rel="nofollow">Creating A Vocabulary for Data Privacy</a>” presents a historical overview of the DPVCG, and describes the methodology and structure of the DPV along with describing its creation. An open-access version can be accessed <a href="http://hdl.handle.net/2262/91581" rel="nofollow">here</a>, <a href="http://doras.dcu.ie/23801/" rel="nofollow">here</a>, and <a href="https://aic.ai.wu.ac.at/~polleres/publications/pand-etal-2019ODBASE.pdf" rel="nofollow">here</a>.</li> </ul>
Elicitation with Sammy Mbipite on 07/29/2024 on Vocabulary
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.