Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

10

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

10 results for “tagged corpus”

Learn how ShareScore rates datasets ↗
zenodo52/100

S1000 corpus, large-scale tagging results and other supplementary files

<p>Data associated with the S1000 corpus</p><p>The tagger software for which the dictionary files in <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/tagger-organisms-dictionary-S1000.tar.gz">tagger-organisms-dictionary-S1000.tar.gz </a>can be used with can be found here: <a href="https://github.com/larsjuhljensen/tagger">https://github.com/larsjuhljensen/tagger</a></p><p>The online version of the annotation documentation can be found here: <a href="https://katnastou.github.io/s1000-corpus-annotation-guidelines/">https://katnastou.github.io/s1000-corpus-annotation-guidelines/</a></p><p>The S1000 corpus split in training, development and test sets in BRAT format can be found in <a href="https://zenodo.org/api/records/10285825/files/S1000-corpus.tar.gz">S1000-corpus.tar.gz</a><a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000-corpus.tar.gz?versionId=ac7ce430-c265-49bb-8c8f-9b5f8e271cbe"> </a>and in CoNLL format here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/s1000-conll.tar.gz">s1000-conll.tar.gz</a></p><p>The tagging results of Jensenlab tagger for the S1000 test set are here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000-jensenlab-tagger.tar.gz?versionId=d8d9c9f5-ee3b-4738-aefa-a4a95475d25d">S1000-jensenlab-tagger.tar.gz</a></p><p>The result from the large scale run in entire PubMed and PMC Open Access articles for Jensenlab tagger is provided here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/Jensenlab_tagger_large_scale_matches_with_rank.tsv.gz?versionId=48825928-9fc9-423c-8a4c-4f8994e95805">Jensenlab_tagger_large_scale_matches_with_rank.tsv.gz</a></p><p>The model used for the large scale run of the transformer-based method is here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000_Transformer_based_tagger_large_scale_model.tar.gz?versionId=8e974f64-9abc-4449-a377-e3f97e91d612">S1000_Transformer_based_tagger_large_scale_model.tar.gz</a> and the results from the large scale tagging here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/Transformer_based_tagger_large_scale_matches_with_rank.tsv.zip?versionId=dc21a6ba-9763-4130-9f02-0341a885c692">Transformer_based_tagger_large_scale_matches_with_rank.tsv.zip</a></p>

opencc-by-4.0Sep 2022View details →
zenodo48/100

Corpus of Occitan Written Traditional Folktales Annotated with Part-Of-Speech (OWT-Tag)

<p>This resource contains 5 extracts of texts in Occitan which were manually annotated with lemmas and parts-of-speech, following the Grace standard. It was produced during the ExpressioNarration project, funded by a Marie Curie Individual Fellowship, in order to evaluate the performance of an Occitan Part-Of-Speech tagger, Talismane, to the specifities of the corpus of the project called Oral Occitan (OcOr), also available on https://zenodo.org/record/1451753#.W78FJWOYSpo.<br> Each extract contains around 1500 words. They are extracted from &#39;Contes et proverbes populaires recueillis en armagnac et Contes populaires recueillis en agenais&#39; de J.-F. Blad&eacute;, &#39;Coundes biarn&eacute;s, cou&eacute;ilhuts a&uuml;s pars&agrave;as mi&eacute;ytad&egrave;s dou p&eacute;ys d&eacute; Biarn&#39; de J.-V. Lalanne, &#39;Contes populaires du Languedoc&#39; de L. Lambert and &#39;Contes populaires recueillis dans la Grande-Lande&#39; de F. Arnaudin.<br> The annotation process is described in the following article available on https://www.openscience.fr/IMG/pdf/iste_modocv1n1_2.pdf.</p>

opencc-by-sa-4.0Oct 2018View details →
zenodo40/100

A part-of-speech (POS) tagged corpus of Classical Tibetan

<p>This part-of-speech (POS) tagged corpus of Classical Tibetan was prepared in the course of the research project &#39;Tibetan in Digital Communication&#39; (2012-2015) hosted at SOAS, University of London and funded by the UK&#39;s Arts and Humanities Research Council (grant code: AH/J00152X/1). For a description of the tag set see Garrett et al. 2014. and Garrett et al. 2015. This corpus includes the <em>Mdzaṅs blun</em> (9th century, canonical), the <em>Bu ston chos ḥbyuṅ</em> (13th century, ecclesiastical history), the <em>Mi la ras paḥi rnam thar</em> and <em>Mar paḥi rnam thar</em> (15th century, biography).</p>

opencc-by-4.0May 2017View details →
zenodo40/100

Specialised POS Tagged Syriac Corpus for State Morphology

<h1>Overview</h1> <p>A total of twelve .TXT files each representing a Syriac text that has been transcribed and tagged for part-of-speech (POS). This corpus forms part of a PhD research project on the historical syntax of Aramaic (Syriac) at The Australian National University (2020&mdash;current) in Canberra, Australia. This research project is interested in noun state morphology, among other topics, which is reflected in the POS scheme for this corpus.</p> <h2>Method</h2> <p>A detailed summary of this methodology is provided in El-Khaissi (data paper&nbsp;in review with the&nbsp;<em>Journal of Open Data Humanities</em>).</p> <ul> <li>Transcriptions are sourced from <a href="https://syriaccorpus.org/" target="_blank" rel="noopener">Digital Syriac Corpus.</a></li> <li>POS tags are based on word matches using&nbsp;<a href="https://sedra.bethmardutho.org/about/openapi" target="_blank" rel="noopener">SEDRA IV API (v1.0.0).</a></li> <li>Selection of Syriac texts was optimised to minimise external influence on Syriac grammar and maximise full coverage of key periods of the Syriac language from 2nd&mdash;13th century AD.</li> </ul> <h2>POS Format &amp; Abbreviations&nbsp;</h2> <p>POS tags in the text files follow the following format:</p> <blockquote> <p>&lt;syntax-category&gt;-&lt;state&gt;_&lt;syriac_word&gt;</p> </blockquote> <p>Thus, an underscore '_' marks the beginning of a tag sequence while tag values are separated by hyphen(s) '-'. For example (noting text directionality constraints):</p> <blockquote> <pre>ܒܘܪܟܬܐ_EMP-N</pre> </blockquote> <p>The following abbreviation lists the definition of all POS tags, which are based on the parameters available in&nbsp;<a href="https://sedra.bethmardutho.org/about/openapi" target="_blank" rel="noopener">SEDRA IV API (v1.0.0).</a></p> <table> <tbody> <tr> <td>Absolute state noun (indeterminate relic)</td> <td>ABS</td> </tr> <tr> <td>Emphatic state noun (new indeterminate)</td> <td>EMP</td> </tr> <tr> <td>Construct state noun (bound noun)</td> <td>CNS</td> </tr> <tr> <td>State not applicable</td> <td>X</td> </tr> <tr> <td>particle</td> <td>PTCL</td> </tr> <tr> <td>pronoun</td> <td>PRO</td> </tr> <tr> <td>preposition</td> <td>PREP</td> </tr> <tr> <td>verb</td> <td>V</td> </tr> <tr> <td>denominative</td> <td>DEN</td> </tr> <tr> <td>noun</td> <td>N</td> </tr> <tr> <td>numeral</td> <td>NUM</td> </tr> <tr> <td>substantive</td> <td>SBV</td> </tr> <tr> <td>adjective</td> <td>ADJ</td> </tr> <tr> <td>proper noun</td> <td>PN</td> </tr> <tr> <td>adverb</td> <td>ADV</td> </tr> <tr> <td>demonym</td> <td>DNM</td> </tr> <tr> <td>participle adjective</td> <td>PTCPADJ</td> </tr> <tr> <td>adverb</td> <td>ADV</td> </tr> <tr> <td>idiom</td> <td>IDM</td> </tr> <tr> <td>See Quality Control &amp; Limitations below</td> <td>DUP</td> </tr> </tbody> </table> <h2>Quality Control &amp; Limitations</h2> <p>On average per manuscript, the POS-tagging process achieved a 63.13% saturation of texts. The POS tagging process was based on an exact-match process, which does not take into account syntactic or semantic context. Syriac words which exhibit homonymy are thus tagged with the value 'DUP' and should be assessed manually based on its original context. Among all 297,981 words in the corpus with an available POS tag, approximately 73,188 (24.56%) of tags reflected some kind of homonymy involving a word with various semantic and/or syntactic interpretations.</p> <p>Since this dataset was created as part of a research project investigating noun state morphology, additional tags were created targetting various state values. Grammatical elements, like number and gender, were not required as part of this investigation and therefore excluded from the POS-tagging process.</p> <h2>Contact</h2> <p>For any questions, please contact Charbel El-Khaissi &lt;Charbel.El-Khaissi@anu.edu.au&gt;.</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

The Annotated Corpus of Classical Tibetan (ACTib), Part I - Segmented version, based on the BDRC digitised text collection, tagged with the Memory-Based Tagger from TiMBL.

<p>This corpus is a part-of-speech tagged version of</p> <p>Wallman, Jeff, Rowinski, Zach, Ngawang Trinley, Tomlinson, Chris, &amp; Keutzer, Kurt. (2017). Collection of Tibetan etexts compiled by the Buddhist Digital Resource Center [Data set]. Zenodo. http://doi.org/10.5281/zenodo.821218</p> <p>using the training data of</p> <p>Hill, Nathan W., &amp; Garrett, Edward. (2017). A part-of-speech (POS) tagged corpus of Classical Tibetan [Data set]. Zenodo. http://doi.org/10.5281/zenodo.574878</p> <p>using the memory based tagger of</p> <p>https://languagemachines.github.io/mbt/</p> <p>Please note that the files are not post-processed or manually corrected and that a small number of files in the KarmaDelek directory were still annotated, although the original xml-input was corrupted already.</p>

opencc-by-4.0Jul 2017View details →
zenodo36/100

The Annotated Corpus of Classical Tibetan (ACTib), Part II - POS-tagged version, based on the BDRC digitised text collection, tagged with the Memory-Based Tagger from TiMBL

<p>This corpus is a part-of-speech tagged version of</p> <p>Wallman, Jeff, Rowinski, Zach, Ngawang Trinley, Tomlinson, Chris, &amp; Keutzer, Kurt. (2017). Collection of Tibetan etexts compiled by the Buddhist Digital Resource Center [Data set]. Zenodo. http://doi.org/10.5281/zenodo.821218</p> <p>using the training data of</p> <p>Hill, Nathan W., &amp; Garrett, Edward. (2017). A part-of-speech (POS) tagged corpus of Classical Tibetan [Data set]. Zenodo. http://doi.org/10.5281/zenodo.574878</p> <p>Please note that the files are not post-processed or manually corrected and that a small number of files in the KarmaDelek directory were still annotated, although the original xml-input was corrupted already.</p> <p>&nbsp;</p> <p>using the memory based tagger of</p> <p>https://languagemachines.github.io/mbt/</p>

opencc-by-4.0Jul 2017View details →
zenodo36/100

Tagged Corpus of Early English Correspondence Extension Sampler (TCEECES)

<p>The <em>Tagged Corpus of Early English Correspondence Extension Sampler</em> (TCEECES) is the third public release from the 18th-century part of the <em>Corpora of Early English Correspondence</em> (CEEC-400).</p> <p>The TCEECES forms one part of the full CEECES. The other parts are the&nbsp;<a href="https://www.doi.org/10.5281/zenodo.4644243">CEECES part 1</a> (released 1 April 2021), and the&nbsp;<a href="https://www.doi.org/10.5281/zenodo.5887101">CEECES part 2</a> (released 11 April 2022).</p> <p>The TCEECES is an extract from the full&nbsp;<em>Tagged Corpus of Early English Correspondence Extension</em>&nbsp;(TCEECE), which remains unpublished.</p> <p>See the accompanying manual&nbsp;for more&nbsp;information on the TCEECES; see&nbsp;the manuals for&nbsp;the CEECES 1 and the CEECES 2 for&nbsp;more information on the CEECES. See <a href="https://varieng.helsinki.fi/CoRD/corpora/CEEC/">https://varieng.helsinki.fi/CoRD/corpora/CEEC/</a> for more on the CEEC-400.</p> <p>Citation:</p> <p>TCEECES&nbsp;= <em>Tagged Corpus of Early English Correspondence Extension Sampler</em>. Compiled by Terttu Nevalainen, Helena Raumolin-Brunberg, Samuli Kaislaniemi, Mikko Laitinen, Minna Nevala, Arja Nurmi, Minna Palander-Collin, Tanja S&auml;ily and Anni Sairio at the Department of Languages, University of Helsinki. Spelling standardised by Mikko Hakala, Minna Palander-Collin, Minna Nevala, Emanuela Costea, Anne Kingma and Anna-Lina Wallraff. Annotated by Lassi Saario and Tanja S&auml;ily. XML conversion and encoding by Lassi Saario. Helsinki: VARIENG, 2022.</p>

opencc-by-nc-nd-4.0Apr 2022View details →
zenodo36/100

A New Annotation Scheme for the Sejong Part-of-speech Tagged Corpus

<p>We produce Sejong-style morphological analysis and part-of-speech tagging results which have been the de facto&nbsp;standard for Korean language processing by using UDPipe (http://ufal.mff.cuni.cz/udpipe)&nbsp;</p> <p>&nbsp;</p> <p>udpipe --tokenize --tag sjmorph.model input &gt; output</p> <p>see&nbsp;https://github.com/jungyeul/sjmorph</p>

opencc-by-4.0May 2019View details →
zenodo28/100

The Annotated Corpus of Classical Tibetan (ACTib) - Version 2.0 (Segmented & POS-tagged)

<p>This corpus consisting of &gt;185 million tokens is a segmented and part-of-speech tagged version of</p> <p>Wallman, Jeff, Rowinski, Zach, Ngawang Trinley, Tomlinson, Chris, &amp; Keutzer, Kurt. (2017). Collection of Tibetan etexts compiled by the Buddhist Digital Resource Center [Data set]. Zenodo. http://doi.org/10.5281/zenodo.821218</p> <p>using the training data of</p> <p>Hill, Nathan W., &amp; Garrett, Edward. (2017). A part-of-speech (POS) tagged corpus of Classical Tibetan [Data set]. Zenodo. http://doi.org/10.5281/zenodo.574878</p> <p>The code for segmenting and POS tagging any Tibetan file can be found on GitHub.</p> <p>This Version 2 of ACTib is based on the same XML files as ACTib Version 1 (http://doi.org/10.5281/zenodo.823707), but contains both segmented and POS-tagged files and is improved in a number of ways, although post-processing was still done automatically and no manual correction was involved. For details of this improved annotation method see:</p> <p>Meelen, Marieke, Roux, &Eacute;lie &amp; Hill, Nathan (forthcoming). &#39;Optimisation of the largest annotated Tibetan corpus combining rule-based, memory-based &amp; deep-learning methods&#39; in <em>TALLIP.</em></p>

opencc-by-4.0May 2020View details →
zenodo28/100

THE PROBLEM OF SEMANTIC TAGGING OF PHILOSOPHICAL TERMS IN THE UZBEK LANGUAGE IN THE CORPUS

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record