Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

132

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

132 results for “vocabularies”

Learn how ShareScore rates datasets ↗
zenodo44/100

Hachidaishu Classical Japanese Poetic Vocabulary Dataset

<h1>Hachidaishu classical Japanese poetic vocabulary dataset</h1> <h2>Hilofumi Yamamoto, Ph.D. (Institute of Science Tokyo)</h2> <h2>Bor Hodo&scaron;ček, D.Engineering (The University of Osaka)</h2> <h3>Data offset</h3> <p>Example: # 1 Kokinshu</p> <pre><code>01:000001:0001 A00 BG-01-1630-01-0100 02 年 年 とし 年 とし 01:000001:0001 A10 BG-01-1911-03-1800 02 年 年 とし 年 とし 01:000001:0002 A00 BG-08-0061-07-0100 61 の の の の の 01:000001:0003 A00 BG-01-1770-01-0300 02 内 内 うち 内 うち 01:000001:0004 A00 BG-08-0061-05-0100 61 に に に に に 01:000001:0005 A00 BG-01-1624-02-0100 02 春 春 はる 春 はる 01:000001:0006 A00 BG-08-0065-07-0100 65 は は は は は 01:000001:0007 A00 BG-02-1527-01-0102 47 き 来 く 来 き 01:000001:0008 A00 BG-03-1200-02-0900 74 に ぬ ぬ に に 01:000001:0008 A10 BG-09-0010-01-0101 74 に ぬ ぬ に に 01:000001:0008 A20 BG-09-0010-03-0200 74 に ぬ ぬ に に 01:000001:0009 A00 BG-09-0010-04-0300 74 けり けり けり けり けり 01:000001:0010 B00 BG-01-1950-14-0100 02 一とせ 一年 ひととせ 一年 ひととせ 01:000001:0010 C00 BG-01-1950-01-0300 19 一 一 いち 一 いち 01:000001:0010 C01 BG-01-1630-01-0100 02 年 年 とし 年 とし 01:000001:0011 A00 BG-08-0061-10-0100 61 を を を を を 01:000001:0012 A00 BG-01-1642-02-0100 02 こそ 去年 こぞ 去年 こぞ 01:000001:0013 A00 BG-08-0061-04-0100 61 と と と と と 01:000001:0014 A00 BG-08-0065-14-0100 65 や や や や や 01:000001:0015 A00 BG-02-3120-01-0100 47 いは 言ふ いふ 言は いは 01:000001:0016 A00 BG-03-3012-03-2600 74 ん む む む む 01:000001:0016 A10 BG-09-0010-02-0102 74 ん む む む む 01:000001:0017 B00 BG-01-1641-02-0100 02 ことし 今年 ことし 今年 ことし 01:000001:0017 C00 BG-03-1000-01-0100 57 この この この この この 01:000001:0017 C01 BG-01-1630-01-0100 02 年 年 とし 年 とし 01:000001:0018 A00 BG-08-0061-04-0100 61 と と と と と 01:000001:0019 A00 BG-08-0065-14-0100 65 や や や や や 01:000001:0020 A00 BG-02-3120-01-0100 47 いは 言ふ いふ 言は いは 01:000001:0021 A00 BG-03-3012-03-2600 74 ん む む む む 01:000001:0021 A10 BG-09-0010-02-0102 74 ん む む む む </code></pre> <h3>A line consists of 7 columns separated by spaces.</h3> <pre><code>01:000001:0007 A00 BG-02-1527-01-0102 47 き 来 く 来 き </code></pre> <ul> <li>1st column "01:000001:0007" consists of 3 fields: 1) anthology, 2) number of poem, and 3) serial ID of the token. The anthology ID indicates respectively: 01..Kokinshu, 02..Gosenshu, 03..Shuishu, 04..Goshuishu, 05..Kin'yoshu, 06..Shikashu, 07..Senzaishu, and 08..Shinkokinshu.</li> <li>2nd column indicates type of token: A type is a single token; B type is a compound token; C type is a breakdown of B type. A00 indicates a single token; A01 indicates a single token and has another meaning; B00 indicates a compound token; B01 indicates a compound token which has another meaning; C00 indicates the first element of the B00/B01.. breakdown; C01 indicates the second element of the B00/B01.. breakdown.</li> <li>3rd column "BG-02-1527-01-0102": classification ID based on semantic categories according to Bunruigoihyo (Yamazaki et al. 2014).</li> <li>4th column indicates a Chasen POS number.</li> <li>5th column indicates surface form: a form appears in literary works.</li> <li>6th column indicates lemma in kanji writing.</li> <li>7th column indicates lemma in kana writing.</li> <li>8th column indicates conjugated form in kanji writing form.</li> <li>9th column indicates conjugated form in kana writing form.</li> </ul> <h2>Reference</h2> <ol> <li> <p>Yamamoto, Hilofumi (2007) Thesaurus of Japanese Poetic Vocabulary Based on the Semantic Classifications Chart, The 13th Annual Symposium for Database of the Humanities, 1-8, The Association for Database of the Humanities, Osaka.</p> </li> <li> <p>Yamamoto, Hilofumi (2009) Thesaurus for the Hachidaishu (ca. 905-1205) with the classification codes based on semantic principles, Nihongo no Kenkyu / Studies in the Japanese Language, 46-52, Society for Japanese Linguistics, 5, 1, ISSN1349-5119.</p> </li> <li> <p>Yamamoto, Hilofumi (2021) Hachidaishu vocabulary dataset, Zenodo, version 1.0.1, <a href="https://doi.org/10.5281/zenodo.4744170">https://doi.org/10.5281/zenodo.4744170</a></p> </li> <li> <p>Yamazaki, Makoto and Kashino, Wakako and Uchiyama, Kiyoko and Sunaoka, Kazuko, and Tajima, Ikudo and Yamamoto, Hilofumi and Han, Yoo-Sik and Seol, Geun-Su (2014) Bunruigoihyo zouhokaiteiban" e no anoteishion: kihongi no kettei (in Japanese), Keiryo Kokugo gakkai dai 58 kai taikai yokoshu, pp. 7--12.</p> </li> <li> <p><a href="http://kotenseki.nijl.ac.jp/biblio/200007092">国文学研究資料館二十一代集</a></p> </li> <li> <p><a href="http://codh.rois.ac.jp/pmjt/book/200007092/">二十一代集 DOI: 10.20730/200007092</a>: ROIS-DS人文学オープンデータ共同利用センター 新日本古典籍総合データベース(200007093)</p> </li> </ol> <p>&lt;!-- @dataset{yamamoto_hilofumi_2021_4735848, author = {Yamamoto, Hilofumi}, title = {Hachidaishu vocabulary dataset}, month = may, year = 2021, publisher = {Zenodo}, version = {1.0.0}, doi = {10.5281/zenodo.4735848}, url = {https://doi.org/10.5281/zenodo.4735848} } @Article{yamagen2009ae, author = {Yamamoto, Hilofumi}, title = {Thesaurus for the Hachidaishu (ca.\,905--1205) with the classification codes based on semantic principles}, journal = {Nihongo no Kenkyu / {S}tudies in the Japanese Language}, pages = {46--52}, OPTpublisher = {Society for Japanese Linguistics}, year = {2009}, volume = {5}, number = {1}, OPTedition = {ISSN1349-5119}, OPTmonth = {}, OPTnote = {}, OPTannote = {}, OPTlocation = {}, OPTmemo = {} } @InCollection{yamagen2007de, author = {Yamamoto, Hilofumi}, title = {Thesaurus of Japanese Poetic Vocabulary Based on the Semantic Classifications Chart}, year = {2007}, booktitle = {The 13th Annual Symposium for Database of the Humanities}, pages = {1--8}, publisher = {The Association for Database of the Humanities}, address = {Osaka}, OPTedition = {}, OPTmonth = {2007.12}, OPTmemo = {} }</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

CLDF dataset derived from Greenhill et al.'s "Austronesian Basic Vocabulary Database" from 2022 focusing on Oceanic languages

<p>Cite the source of the dataset as:</p> <blockquote> <p>Greenhill, S.J., Blust. R, &amp; Gray, R.D. (2008). The Austronesian Basic Vocabulary Database: From Bioinformatics to Lexomics. Evolutionary Bioinformatics, 4:271-283.</p> </blockquote>

opencc-by-4.0May 2023View details →
zenodo40/100

Figure 5 in Masner, a new genus of Ceraphronidae (Hymenoptera, Ceraphronoidea) described using controlled vocabularies

Figure 5. Male antenna of Ceraphronoidea A–B Masner lubomirus Deans and Mikó, sp. n. C Aphanogmus sp. D Ceraphron sp. E Dendrocerus sp. F Megaspilus sp. Unlabelled arrows indicate antennal sensilla located on sensillar patch. Scale bars in micrometer.

opencc-by-4.0Sep 2009View details →
zenodo40/100

Figure 7. Ceraphronoidea A–B Metasoma, ventral view A Trichosteresis glabra B in Masner, a new genus of Ceraphronidae (Hymenoptera, Ceraphronoidea) described using controlled vocabularies

Figure 7. Ceraphronoidea A–B Metasoma, ventral view A Trichosteresis glabra B Lagynodes sp. C Masner lubomirus Deans and Mikó, sp. n., lateral view. Scale bars in micrometer.

opencc-by-4.0Sep 2009View details →
zenodo40/100

Figure 6. Ceraphronoidea A in Masner, a new genus of Ceraphronidae (Hymenoptera, Ceraphronoidea) described using controlled vocabularies

Figure 6. Ceraphronoidea A Trichosteresis glabra, head, anterior view B Aphanogmus sp., mesosoma, lateral view, anterior to the left C Aphanogmus sp., head, posterior view D Aphanogmus sp., head, anterior view E Megaspilus sp., male genitalia, ventral view F Conostigmus sp., apex of mesotibia, ventral view. Scale bars in micrometer.

opencc-by-4.0Sep 2009View details →
zenodo40/100

Figure 4 in Masner, a new genus of Ceraphronidae (Hymenoptera, Ceraphronoidea) described using controlled vocabularies

Figure 4. Male genitalia and Waterston's evaporatorium of Ceraphronidae A–B Masner lubomirus Deans and Mikó, sp. n., male genitalia A ventral view8 B lateral view9 C–F Waterston's evaporatorium C Masner lubomirus Deans and Mikó, sp. n., dorsal view D Aphanogmus sp. dorsal view E Masner lubomirus Deans and Mikó, sp. n. F Aphanogmus sp., anterior view. Scale bars in micrometer.

opencc-by-4.0Sep 2009View details →
zenodo40/100

Supplementary Data for "Streamlining Vocabulary Conversion to SKOS: A YAML-based Approach to Facilitate Participation in the Semantic Web"

<p>This dataset contains quality assessment results for 26 vocabularies. The assessment was conducted using the <a href="https://skos-play.sparna.fr/skos-testing-tool/">qSKOS vocabulary quality assessment tool</a>.</p> <p>The 26 assessed vocabularies were converted from their original formats into the Simple Knowledge Organization System (SKOS) data model using the approach described in our paper titled <a href="https://doi.org/10.1007/978-3-031-62362-2_9">"Streamlining Vocabulary Conversion to SKOS: A YAML-based Approach to Facilitate Participation in the Semantic Web"</a>, presented at the <a href="https://doi.org/10.1007/978-3-031-62362-2">24th International Conference on Web Engineering (ICWE 2024)</a>.</p> <p>The dataset contains a quality assessment for the following vocabularies:</p> <ol> <li>A Taxonomy of Evaluation Towards Standards</li> <li>Cross-Device Taxonomy</li> <li>What Makes a Data-driven Business Model? A Consolidated Taxonomy</li> <li>DDI Aggregation Method</li> <li>DDI Mode of Collection</li> <li>Building a New Taxonomy for Data Discretization Techniques</li> <li>Demopaedia</li> <li>Data Science Glossary</li> <li>A Taxonomy of Evaluation Approaches in Software Engineering</li> <li>Evaluation Thesaurus</li> <li>The Glossary of Human Computer Interaction</li> <li>Human-Factors Taxonomy</li> <li>A Taxonomy to Structure and Analyze Human&ndash;Robot Interaction</li> <li>A Taxonomy of Interaction for Instructional Multimedia</li> <li>A Taxonomy of Interrogation Methods</li> <li>Design Vocabulary for Human&ndash;IoT Systems Communication</li> <li>Understanding Movement and Interaction: An Ontology for Kinect-Based 3D Depth Sensors</li> <li>Thesaurus Mass Communication</li> <li>Mixed-Initiative Human-Robot Interaction: Definition, Taxonomy, and Survey</li> <li>A Taxonomy of Quality of Service and Quality of Experience of Multimodal Human-Machine Interaction</li> <li>A Human-Centered Taxonomy of Interaction Modalities and Devices</li> <li>A Taxonomy of Spatial Interaction Patterns and Techniques</li> <li>A Taxonomy of Social Errors in Human-Robot Interaction</li> <li>Taxonomy of Digital Research Activities in the Humanities</li> <li>Virtual Reality and the CAVE: Taxonomy, Interaction Challenges and Research Directions&nbsp;</li> <li>Cross-Device Interaction</li> </ol>

opencc-by-4.0Feb 2024View details →
zenodo40/100

Matters of Police Ordinances. SKOS vocabulary

<p>The file <code>polmat.ttl</code>contains a systematic taxonomy of more than 1.800 matters regulated by police ordinances (last revised: Oct 2021). It is a SKOS vocabulary encoded in ttl/"turtle" format.</p> <p>The taxonomy has originated in the project "<a href="https://www.lhlt.mpg.de/research-project/repertory-of-policeyordnungen" rel="nofollow">Repertory of Policeyordnungen</a>" of the (then) <a href="https://www.lhlt.mpg.de/en" rel="nofollow">Max Planck Institute for Legal History and Legal Theory</a>. Initially this system of categories has been established in German in 1996 by Karl H&auml;rter and Michael Stolleis.</p> <p>The taxonomy has been SKOS-ified in November 2020 by Annemieke Romein and Andreas Wagner.</p> <p>The taxonomy is published under CC0, but we will be happy if you give attribution when you use it.</p>

opencc-zeroMar 2021View details →
zenodo40/100

A standardised climate change hazard vocabulary for heritage

<p>This dataset (.xlsx) is a vocabulary of climate change hazards for heritage. Hazards are the potential occurrences of natural or physical events that may cause damage or loss. Previously there had been no definitive list of climate hazards for heritage that were directly connected to changing climatic processes. This project addresses this gap by linking the created hazards to the Climatic Impact-Drivers (CIDs) produced by the Intergovernmental Panel on Climate Change (IPCC). The vocabulary consists of 52 primary and key related hazards for heritage. It is international in its remit.</p> <p>The list will be published as a vocabulary on <a href="https://www.heritage-standards.org.uk/fish-vocabularies/" target="_blank" rel="noopener">the Forum on Information Standards in Heritage (FISH)</a> where it can be accessed in multiple formats <a href="https://heritagedata.org/live/schemes/38076.html">including linked data</a>. This .xlsx format places the hazards in relationship to each other and in their CID context. Candidate terms can be submitted to the research group Heritage Environmental Risk and Data Analytics <a href="mailto:herada@ucl.ac.uk" target="_blank" rel="noopener noreferrer">herada@ucl.ac.uk</a> (terms submitted to <a href="mailto:Terminologies@HistoricEngland.org.uk">Terminologies@HistoricEngland.org.uk</a> will be directed to the research group for approval). An accompanying <a href="https://historicengland.org.uk/research/results/reports/13-2024?search=13%2F2024&amp;searchType=research+report">Historic England Research Report</a> provides more information, including the methodology of the project (available in both English and Welsh).</p> <p>The authors are interested in hearing from users of the vocabulary, specifically those that link the hazards to observed impacts of climate change on parts of the historic environment. This dataset was produced as part of a funded 6-month project between Historic England and the UCL Institute for Sustainable Heritage on developing a standardised vocabulary of climate change hazards for the historic environment. The Welsh version of the dataset was translated in 2025 by Lingo Soar, in collaboration with Fforest Fawr UNESCO Geopark, and as part of the UK National Commission for UNESCO's Climate Change and UNESCO Heritage project.</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

A controlled vocabulary for research and innovation in the field of Circular Bioeconomy

<p>We live in a world of limited resources. Facing global challenges such as climate change and degradation of natural capital, in addition to the increasing rate of resource consumption, we are compelled to look for new ways of producing and consuming that respect the ecological limits of our planet. A model based on the Circular Bioeconomy (<em>CBE</em>) is key to successfully tackling the complexity of this paradigm shift and addressing these challenges: CBE is a relatively new and fast evolving concept, which is still in the conceptualisation phase. It stems from the concepts of &ldquo;<em>bioeconomy</em>&rdquo; and &ldquo;<em>circular economy</em>&rdquo;, which have become progressively interlinked in recent years.</p> <p>Although the wider community of stakeholders has not reached yet a consensus on a single definition for CBE, we try to go beyond this limitation, by proposing a controlled vocabulary of keywords related to the field, built by eliciting domain knowledge from both experts and policy-makers. This effort responds to an explorative project launched by the Generalitat de Catalunya (the regional government of Catalonia, Spain - specifically, the Ministry of the Economy and Finance, Secretary for Economic Affairs and European Funds) in order to identify CBE research, development and innovation activities. The work, coordinated by SIRIS Academic, was carried out by consulting the advice of experts in the domain (see Acknowledgements, below), and it was ultimately applied to inform regional strategies on bioeconomy and the Research and Innovation Strategy for Smart Specialisation.</p> <p>After a qualitative analysis of numerous examples, the keywords extracted have been classified into three categories:</p> <ul> <li> <p><strong>Unequivocal</strong>: keywords positively associated with the CBE domain (e.g. <em>bioplastic conversion, biorefinery, or biomass</em>),</p> </li> <li> <p><strong>Bioeconomy</strong>: keywords in the bioeconomy domain but not necessarily within the circular paradigm&nbsp; (e.g. <em>agriculture, algae, or organic waste</em>),</p> </li> <li> <p><strong>Technologies and processes</strong>: keywords concerning circular processes and technologies (e.g. <em>valorisation, or bioconversion</em>).</p> </li> </ul> <p>This distinction allows one to apply the vocabulary to link a given text to the CBE domain. Indeed, a text may be&nbsp; considered in the CBE perimeter if it mentions an unequivocal keyword, or if it contains a concept concerning bioeconomy co-appearing with a concept concerning circular technologies and processes (e.g. food waste with composting, or vegetation with biosynthesis).</p> <p>The circular bioeconomy vocabulary, in its current form, contains 393 keywords (49 unequivocal keywords, 229 bioeconomy keywords, and 115 concerning circular processes and technologies).&nbsp;<br> &nbsp;</p> <p><strong>Datasets Provided</strong></p> <p><br> In this publication, we provide:</p> <p>&nbsp;</p> <ul> <li><strong>A document outlining the context that lead to the creation of the vocabulary</strong>, the methodology used to build it and the strategy to apply it</li> <li><strong>The full vocabulary of terms as a csv file</strong>. For each keyword of the vocabulary, the property &ldquo;type&rdquo; specifies whether the term is either an &ldquo;Unequivocal&rdquo; CBE keyword, a wider &ldquo;Bioeconomy&rdquo; term or a specification of &ldquo;Technologies and processes&rdquo;.</li> <li>In addition to the controlled vocabulary, we also provide <strong>a labelled dataset to foster the development of new text mining initiatives</strong> on the results obtained. The dataset consists of the description of 2000 R&amp;D projects funded by the Horizon 2020 framework of the European Commission. Of these, 1000 have been associated with the field of CBE via the vocabulary and further manual validation, while further 1000 descriptions are unrelated to CBE. Each project entry has the following attributes: title, abstract, ecRef (project identifier in the EU information service, CORDIS), and label (where True are projects in the CBE domain)</li> </ul> <p><strong>Acknowledgements</strong></p> <p>The work described above, leading to the production of the datasets provided, was developed in collaboration with Generalitat de Catalunya officials, as well as CBE research and innovation experts and stakeholders that contributed essential input to the conceptual definition of the CBE, the review and improvement of the controlled vocabulary and the quality of the classification, as well as the the revision of the analytical results described in the section Applicability, in particular:</p> <ul> <li> <p>Tatiana Fern&aacute;ndez (Direcci&oacute; General de Promoci&oacute; Econ&ograve;mica, Compet&egrave;ncia i Regulaci&oacute;, de la Generalitat de Catalunya)</p> </li> <li> <p>Merc&egrave; Balcells i Ram&oacute;n Canela (Universitat de Lleida)</p> </li> <li> <p>Teresa Botargues (Diputaci&oacute; de Lleida)</p> </li> <li> <p>Sergio Pons&agrave; (Centre Tecnol&ograve;gic BETA)</p> </li> <li> <p>Ignacio Rodr&iacute;guez and Jaume Si&oacute; (Departament d&rsquo;Agricultura, Ramaderia, Pesca i Alimentaci&oacute;, de la Generalitat de Catalunya)</p> </li> </ul>

opencc-by-sa-4.0Oct 2021View details →
zenodo40/100

Multi-domain and Explainable Prediction of Changes in Web Vocabularies (code & data)

<p>This deposit contains supplementary code &amp; data for the paper &#39;Multi-domain and Explainable Prediction of Changes in Web Vocabularies&#39; (K-CAP 2021).</p> <p>Web vocabularies (WV) have become a fundamental tool for structuring Web data: over 10 million sites use structured data formats and ontologies to markup content. Maintaining these vocabularies and keeping up with their changes are manual tasks with very limited automated support, impacting both publishers and users. Existing work shows that machine learning can be used to reliably predict vocabulary changes, but on specific domains (e.g. biomedicine) and with limited explanations on the impact of changes (e.g. their type, frequency, etc.). In this paper, we describe a framework that uses various supervised learning models to learn and predict changes in versioned vocabularies, independent of their domain. Using well-established results in ontology evolution we extract domain-agnostic and human-interpretable features and explain their influence on change predictability. Applying our method on 139 WV from 9 different domains, we find that ontology structural and instance data, the number of versions, and the release frequency highly correlate with predictability of change. These results can pave the way towards integrating predictive models into knowledge engineering practices and methods.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Statistics of the Network of Linked Vocabularies

<p>Reusing terms in the <a href="https://lod-cloud.net/">Linked Open Data cloud</a> results in a <a href="https://sites.google.com/view/nelo-evolution">Network of Linked vOcabularies (NeLO)</a>, where the nodes are the vocabularies that use at least one term from some other vocabulary and thus depend on each other. These dependencies become a problem when vocabularies in the network change, e.g., when terms are deprecated or deleted. In these cases, all dependent vocabularies in the network need to be updated. To address this shortcoming, we compute the state of NeLO from the available versions of the vocabularies in a period of time of over 17 years.</p> <p>Specifically, we provide the following statistics for each vocabulary in each year from 2001 and 2018 in three different formats (RDF/XML, JSON-LD, and CSV):</p> <ul> <li>in-degree;</li> <li>out-degree;</li> <li>degree;</li> <li>eccentricity;</li> <li>closeness centrality;</li> <li>harmonic closeness centrality;</li> <li>betweenness centrality;</li> <li>Authority;</li> <li>Hub;</li> <li>PageRank.</li> </ul> <p>The publication in the reference contains also further information on the analyzed vocabularies and on the methodology.</p> <p>This dataset is provided for non-commercial use only. If you find this dataset useful in your work, please cite the publication in the reference.</p>

opencc-by-nc-sa-4.0May 2019View details →
zenodo40/100

Amended Pama-Nyungan colour vocabularies

<p>Terminology for basic colours in about 187 Pama-Nyungan languages.</p> <p>Amended data table, based on the Word table http://www.pnas.org/content/suppl/2016/11/10/1613666113.DCSupplemental/pnas.1613666113.st01.docx published with H&amp;B</p> <p>H&amp;B: Haynie, Hannah J &amp; Claire Bowern. 2016. Phylogenetic approach to the evolution of color term systems. PNAS 113(48), 3666–13671. http://www.pnas.org/content/113/48/13666<br> <br> amendments (additions in green, deletions in red) as detailed in Supporting information http://hdl.handle.net/1885/123084 to:<br> Nash, David. 2017. Loss of color terms not demonstrated. PNAS 114. 201714007. https://doi.org/10.1073/pnas.1714007114</p>

opencc-by-4.0Oct 2017View details →
zenodo40/100

Figure 2. Screen capture of the current application-INTELLIGENT AGENT FOR ACQUISITION OF THE MOTHER TONGUE VOCABULARY

<p>From this screen capture (figure 2), you can observe that, for example, the noun car (masina)<br> is in a great correspondence with the correct word car(masina), and in the same correspondence<br> with the incorrect word small (mic). But this is due to the fact that the system worked only with few<br> examples. After a training with much more examples, the system will increase the correspondence<br> between the concept/object car and the word car, and the correspondence between the object car and<br> the word little will remain smallest.</p>

opencc-by-4.0Jan 2010View details →
zenodo40/100

Figure 1. Associations between objects and words-Intelligent Agent for Acquisition of the Mother Tongue Vocabulary

<p>So, the learning process is very complex, and it is a probabilistic one. Thus, a second<br> example is required, and even more than that. Normally, the mother speaks naturally to her child,<br> she speaks with love and affection, she doesn&rsquo;t &ldquo;judge&rdquo; or &ldquo;program&rdquo; what to say to her child.<br> Therefore, another day she will tell her child, for example, &ldquo;My darling son, let&rsquo;s drink the milk<br> from the cup.&rdquo;</p>

opencc-by-4.0Jan 2010View details →
zenodo40/100

Figure 1. Scatter plot of ECA exam by average vocabulary in Grade Three

<p>The mean score of every learner on the vocabulary tests was calculated. Then, through the<br> statistical procedure of regression analysis, the data from the mean scores and the ECA exam scores<br> were analyzed to derive a model which could reliably predict the learners&rsquo; performance on the ECA<br> exam.<br> To check whether the two variables of the ECA exam and the average vocabulary were<br> suitable for linear regression, its scatter plot was examined. The resulting scatter plot (Figure 1)<br> seemed to be sufficient for linear regression.</p>

opencc-by-4.0Oct 2010View details →
zenodo40/100

OrphaCode as a non-standard with mappings to a standard vocabulary in OMOP

<p>The files contain the necessary tables for the introduction of OrphaCode as a non-standard with mappings to a standard vocabulary in OMOP. These include:</p> <ul> <li>concept_manual</li> <li>concept_relationship_manual</li> <li>concept_synonym_manual.</li> </ul> <p>This file has not yet been quality assured, but will be made available for validation.</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Dataset of Vocabulary in Uzbek Primary Education

<p>This dataset compiles words from two main sources: the "Explanatory Vocabulary of the Uzbek Language" (EDUL) and textbooks used across grades 1-4 in Uzbek primary schools (UPSC). The EDUL.txt file contains 29,190 words meticulously compiled by Urgench State University between 2019 and 2023. Additionally, the UPSC dataset includes 208,204 words extracted from primary school textbooks, sorted into separate files for each grade level. The dataset also identifies specific vocabulary words for each grade, supporting the enhancement of Uzbek language education and facilitating the development of natural language processing tools.</p> <ul> <li> <p>Grade 1 lemma vocabulary: 3,188 words (all new words)</p> </li> <li> <p>Grade 2 lemma vocabulary: 4,630 words (including 1,997 new words)</p> </li> <li> <p>Grade 3 lemma vocabulary: 5,700 words (including 1,578 new words)</p> </li> <li> <p>Grade 4 lemma vocabulary: 6,397 words (including 1,356 new words)</p> </li> </ul> <p>All files are conveniently packaged into a single ZIP archive for easy access and distribution.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Name Size Figure 4.2. Ratio of male to female participants-The Effects of CALL on Vocabulary Learning: A Case of Iranian Intermediate EFL Learners

<p>In the past, vocabulary teaching and learning were often given little priority in second<br> language programs but recently there has been a renewed interest in the nature of vocabulary and its<br> role in learning and teaching. Although most teachers might be aware of the importance of<br> technology, say, computer, rarely teachers use it for teaching vocabulary. Thus, the current study<br> aims at exploring the effects of CALL on vocabulary learning of Iranian EFL Learners. In this<br> study, 40 intermediate EFL learners, both male and female aged from 16 to 18 studying New<br> Interchange, book III, were chosen randomly from a language institute in Tehran. They were divided<br> into two twenty-member groups. The experimental group was given the VTS.S (a computer<br> program for teaching vocabularies), a computerized dictionary and provided with teacher efeedback.<br> The control group received no special software and vocabularies were taught using the<br> conventional ways with the help of a paper dictionary.</p>

opencc-by-4.0Oct 2012View details →
zenodo40/100

Figure 4.1. Distribution of scores for the Nelson test-The Effects of CALL on Vocabulary Learning: A Case of Iranian Intermediate EFL Learners

<p>In the past, vocabulary teaching and learning were often given little priority in second<br> language programs but recently there has been a renewed interest in the nature of vocabulary and its<br> role in learning and teaching. Although most teachers might be aware of the importance of<br> technology, say, computer, rarely teachers use it for teaching vocabulary. Thus, the current study<br> aims at exploring the effects of CALL on vocabulary learning of Iranian EFL Learners. In this<br> study, 40 intermediate EFL learners, both male and female aged from 16 to 18 studying New<br> Interchange, book III, were chosen randomly from a language institute in Tehran. They were divided<br> into two twenty-member groups. The experimental group was given the VTS.S (a computer<br> program for teaching vocabularies), a computerized dictionary and provided with teacher efeedback.<br> The control group received no special software and vocabularies were taught using the<br> conventional ways with the help of a paper dictionary. A vocabulary pre-test based on the tests<br> available in their teacher&#39;s guide was given to both groups. The aim of this test was to make sure<br> that the students were not familiar with the words in advance. By pre-test/post-test comparison<br> researchers found learners exposed to VTS.S teacher e-feedback plus the computerized dictionary<br> scored higher than the control group. Both high-stake and low-stake holders can avail from the<br> findings of the study.</p>

opencc-by-4.0Oct 2012View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record