Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6
datasets available to search
ShareScore release 0.9.0
Dataset results
6 results for “POS tagging”
A part-of-speech (POS) tagged corpus of Classical Tibetan
<p>This part-of-speech (POS) tagged corpus of Classical Tibetan was prepared in the course of the research project 'Tibetan in Digital Communication' (2012-2015) hosted at SOAS, University of London and funded by the UK's Arts and Humanities Research Council (grant code: AH/J00152X/1). For a description of the tag set see Garrett et al. 2014. and Garrett et al. 2015. This corpus includes the <em>Mdzaṅs blun</em> (9th century, canonical), the <em>Bu ston chos ḥbyuṅ</em> (13th century, ecclesiastical history), the <em>Mi la ras paḥi rnam thar</em> and <em>Mar paḥi rnam thar</em> (15th century, biography).</p>
Specialised POS Tagged Syriac Corpus for State Morphology
<h1>Overview</h1> <p>A total of twelve .TXT files each representing a Syriac text that has been transcribed and tagged for part-of-speech (POS). This corpus forms part of a PhD research project on the historical syntax of Aramaic (Syriac) at The Australian National University (2020—current) in Canberra, Australia. This research project is interested in noun state morphology, among other topics, which is reflected in the POS scheme for this corpus.</p> <h2>Method</h2> <p>A detailed summary of this methodology is provided in El-Khaissi (data paper in review with the <em>Journal of Open Data Humanities</em>).</p> <ul> <li>Transcriptions are sourced from <a href="https://syriaccorpus.org/" target="_blank" rel="noopener">Digital Syriac Corpus.</a></li> <li>POS tags are based on word matches using <a href="https://sedra.bethmardutho.org/about/openapi" target="_blank" rel="noopener">SEDRA IV API (v1.0.0).</a></li> <li>Selection of Syriac texts was optimised to minimise external influence on Syriac grammar and maximise full coverage of key periods of the Syriac language from 2nd—13th century AD.</li> </ul> <h2>POS Format & Abbreviations </h2> <p>POS tags in the text files follow the following format:</p> <blockquote> <p><syntax-category>-<state>_<syriac_word></p> </blockquote> <p>Thus, an underscore '_' marks the beginning of a tag sequence while tag values are separated by hyphen(s) '-'. For example (noting text directionality constraints):</p> <blockquote> <pre>ܒܘܪܟܬܐ_EMP-N</pre> </blockquote> <p>The following abbreviation lists the definition of all POS tags, which are based on the parameters available in <a href="https://sedra.bethmardutho.org/about/openapi" target="_blank" rel="noopener">SEDRA IV API (v1.0.0).</a></p> <table> <tbody> <tr> <td>Absolute state noun (indeterminate relic)</td> <td>ABS</td> </tr> <tr> <td>Emphatic state noun (new indeterminate)</td> <td>EMP</td> </tr> <tr> <td>Construct state noun (bound noun)</td> <td>CNS</td> </tr> <tr> <td>State not applicable</td> <td>X</td> </tr> <tr> <td>particle</td> <td>PTCL</td> </tr> <tr> <td>pronoun</td> <td>PRO</td> </tr> <tr> <td>preposition</td> <td>PREP</td> </tr> <tr> <td>verb</td> <td>V</td> </tr> <tr> <td>denominative</td> <td>DEN</td> </tr> <tr> <td>noun</td> <td>N</td> </tr> <tr> <td>numeral</td> <td>NUM</td> </tr> <tr> <td>substantive</td> <td>SBV</td> </tr> <tr> <td>adjective</td> <td>ADJ</td> </tr> <tr> <td>proper noun</td> <td>PN</td> </tr> <tr> <td>adverb</td> <td>ADV</td> </tr> <tr> <td>demonym</td> <td>DNM</td> </tr> <tr> <td>participle adjective</td> <td>PTCPADJ</td> </tr> <tr> <td>adverb</td> <td>ADV</td> </tr> <tr> <td>idiom</td> <td>IDM</td> </tr> <tr> <td>See Quality Control & Limitations below</td> <td>DUP</td> </tr> </tbody> </table> <h2>Quality Control & Limitations</h2> <p>On average per manuscript, the POS-tagging process achieved a 63.13% saturation of texts. The POS tagging process was based on an exact-match process, which does not take into account syntactic or semantic context. Syriac words which exhibit homonymy are thus tagged with the value 'DUP' and should be assessed manually based on its original context. Among all 297,981 words in the corpus with an available POS tag, approximately 73,188 (24.56%) of tags reflected some kind of homonymy involving a word with various semantic and/or syntactic interpretations.</p> <p>Since this dataset was created as part of a research project investigating noun state morphology, additional tags were created targetting various state values. Grammatical elements, like number and gender, were not required as part of this investigation and therefore excluded from the POS-tagging process.</p> <h2>Contact</h2> <p>For any questions, please contact Charbel El-Khaissi <Charbel.El-Khaissi@anu.edu.au>.</p>
An unambiguous POS tagging set
<p>This data set contains 1,123 short, POS-tagged sentences (extracted from an earlier data set; see https://zenodo.org/records/7694423), using the Universal tag set. The sentences can easily be POS tagged by a human tagger. However, standard POS taggers struggle with these sentences. The data file contains a header row that describes each column. The first column indicates the type of sentence (either a transcript of spoken text (0) or a sentence originating from written text (1)), the second column contains the actual sentence, with ground truth POS tags. The third column indicates the index of the mistagged token, and the remaining five columns show the tags assigned (of which at least one is a mistagging) of five different taggers.</p>
The Annotated Corpus of Classical Tibetan (ACTib), Part II - POS-tagged version, based on the BDRC digitised text collection, tagged with the Memory-Based Tagger from TiMBL
<p>This corpus is a part-of-speech tagged version of</p> <p>Wallman, Jeff, Rowinski, Zach, Ngawang Trinley, Tomlinson, Chris, & Keutzer, Kurt. (2017). Collection of Tibetan etexts compiled by the Buddhist Digital Resource Center [Data set]. Zenodo. http://doi.org/10.5281/zenodo.821218</p> <p>using the training data of</p> <p>Hill, Nathan W., & Garrett, Edward. (2017). A part-of-speech (POS) tagged corpus of Classical Tibetan [Data set]. Zenodo. http://doi.org/10.5281/zenodo.574878</p> <p>Please note that the files are not post-processed or manually corrected and that a small number of files in the KarmaDelek directory were still annotated, although the original xml-input was corrupted already.</p> <p> </p> <p>using the memory based tagger of</p> <p>https://languagemachines.github.io/mbt/</p>
The Annotated Corpus of Classical Tibetan (ACTib) - Version 2.0 (Segmented & POS-tagged)
<p>This corpus consisting of >185 million tokens is a segmented and part-of-speech tagged version of</p> <p>Wallman, Jeff, Rowinski, Zach, Ngawang Trinley, Tomlinson, Chris, & Keutzer, Kurt. (2017). Collection of Tibetan etexts compiled by the Buddhist Digital Resource Center [Data set]. Zenodo. http://doi.org/10.5281/zenodo.821218</p> <p>using the training data of</p> <p>Hill, Nathan W., & Garrett, Edward. (2017). A part-of-speech (POS) tagged corpus of Classical Tibetan [Data set]. Zenodo. http://doi.org/10.5281/zenodo.574878</p> <p>The code for segmenting and POS tagging any Tibetan file can be found on GitHub.</p> <p>This Version 2 of ACTib is based on the same XML files as ACTib Version 1 (http://doi.org/10.5281/zenodo.823707), but contains both segmented and POS-tagged files and is improved in a number of ways, although post-processing was still done automatically and no manual correction was involved. For details of this improved annotation method see:</p> <p>Meelen, Marieke, Roux, Élie & Hill, Nathan (forthcoming). 'Optimisation of the largest annotated Tibetan corpus combining rule-based, memory-based & deep-learning methods' in <em>TALLIP.</em></p>
Tokenized and POS-Tagged Khmer Data of the Asian Language Treebank Project
<p>* Introduction</p> <p>This is the Khmer ALT of the Asian Language Treebank (ALT) Corpus. English texts sampled from English Wikinews were available under a Creative Commons Attribution 2.5 License.</p> <p>Please refer to<br> http://www2.nict.go.jp/astrec-att/member/mutiyama/ALT/index.html<br> for an introduction of the ALT project.</p> <p>Khmer ALT has been developed by NICT and NIPTICT. The license of Khmer ALT is</p> <p>Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) License<br> https://creativecommons.org/licenses/by-nc-sa/4.0/</p> <p><br> * Contents</p> <p>- data_km.km-[tok|tag].nova : tokenized/POS-tagged Khmer sentences by the nova annotation system<br> # based on the following two guildelines<br> # http://www2.nict.go.jp/astrec-att/member/mutiyama/ALT/Khmer-annotation-guideline.pdf<br> # http://www2.nict.go.jp/astrec-att/member/mutiyama/ALT/Khmer-annotation-guideline-supplementary.pdf</p> <p><br> * Disclaimer</p> <p>[1] The content of the selected English Wikinews articles have been translated for this corpus. English texts sampled from English Wikinews were available under a Creative Commons Attribution 2.5 License. Users of the corpus are requested to take careful consideration when encountering any instances of defamation, discriminatory terms, or personal information that might be found within the corpus. Users of the corpus are advised to read Terms of Use in https://en.wikinews.org/wiki/Main_Page carefully to ensure proper usage.</p> <p>[2] NICT bears no responsibility for the contents of the corpus and the lexicon and assumes no liability for any direct or indirect damage or loss whatsoever that may be incurred as a result of using the corpus or the lexicon.</p> <p>[3] If any copyright infringement or other problems are found in the corpus or the lexicon, please contact us at alt-info[at]khn[dot]nict[dot]go[dot]jp. We will review the issue and undertake appropriate measures when needed.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.