Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1 result for “Named-entity recognition”

Learn how ShareScore rates datasets ↗
zenodo44/100

Named-Entity Recognition for Modern Tibetan Newspapers: Tagset, Guidelines and Training Data

<p>This dataset, tagset and guidelines were the output&nbsp;of a six-month incubator project on the feasibility of developing Named-Entity Recognition (NER) for modern Tibetan, primarily for use with contemporary Tibetan-language newspapers and media published inside the PRC.&nbsp;The project was carried out by the Mongolian and Inner Asian Studies Unit at Cambridge University&rsquo;s Department of Social Anthropology. It was funded by an incubator grant from Cambridge Language Sciences. The project title was&nbsp;&ldquo;Named-Entity Recognition in Tibetan and Mongolian Newspapers.&rdquo; The Project PI was Dr Hildegard Diemberger (Cambridge),&nbsp;the Coordinator and Lead Author was Dr Robert Barnett (SOAS), and Senior Advisers were Dr Nathan Hill (SOAS), Dr Marieke Meelen (Cambridge), and Dr Thomas White (Cambridge).&nbsp;<br> <br> Although some forms of NER and other NLP procedures have been developed within China for modern Tibetan (see Liu, Nuo <em>et al</em>, 2011), the data underlying those initiatives have not been made publicly available and their findings cannot be tested or reproduced. Significant work on developing NLP for Tibetan has been carried out outside China, but has focused largely on classical Tibetan and religious texts (see Hill &amp; Garrett, Edward, 2017).&nbsp;</p> <p>The Cambridge incubator project therefore produced a tagset, guidelines and training data for developing NER for modern Tibetan, with a focus on historical and political analysis of contemporary newspapers, media and other public documents in Tibetan. We compiled 3.11m syllables of data in Tibetan extracted from articles downloaded from Chinese-language news aggregator sites within China, primarily tibet.cpc.people.com.cn and tibet.people.com.cn. From this data, we selected texts containing 280,000 syllables in Tibetan, grouped in 26,000 utterances/sentences (available on request). Using Lighttag, an online annotation site, we developed a tagset for NER consisting of 17 tags (and one for wrong segmentation if using segmented data). We annotated approximately 186,000 syllables, leading to 9,884 annotations. Of these, after discounting flawed data, we produced training data containing c.6,700&nbsp;annotations.&nbsp; We carried out the secondary, manual review offline (for our method of converting Lighttag&nbsp;data for offline review, see the attached report &ldquo;Using Spreadsheets to Review Annotations Offline.pdf&rdquo;), and found an error rate of 3.6%. The final total of reviewed annotations was 6,624.&nbsp;</p> <p>The dataset, tagset, guidelines and reports were developed and documented by Robert Barnett, with assistance from Tsering Samdrup, Dr Hill and Dr Meelen. Primary annotation was by Tsering Samdrup, assisted by Dr Barnett.<br> <br> The datasets published here include:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</p> <ol> <li>The <strong>tagseet guidelines and annotation manual</strong>, including the 17-tag tagset, guidelines, and recommendations&nbsp;(&quot;NER for Modern Tibetan-tagset and guidelines.pdf&quot;).</li> <li>The <strong>tagged training data </strong>in .csv format (&quot;Tibetan NER Training Data-tagged, reviewed wth context-v10-UTF-8.csv&quot;) and .xls format (&quot;Tibetan NER Training Data-tagged with context-v10-UTF-8.xlsx&quot;). This includes&nbsp;6,624&nbsp;reveiwed annotations, arranged according to the Tibetan alphabet&nbsp;together&nbsp;with the tags and context (utterance) for each annotation.</li> <li>The <strong>raw annotation results </strong>downloaded&nbsp;from Lighttag as .json files&nbsp;(&quot;Raw Training Data for NER in Modern Tibetan -Jobs2-11-JSON.zip&quot;) and as .xls files&nbsp;(&quot;Training Data for NER in Modern Tibetan -Jobs2-11-XLS.zip&quot;). These include&nbsp;10 &quot;tasks&quot; or datasets of articles scraped from Tibetan-language websites within Tibet.&nbsp;&nbsp;&nbsp;&nbsp;</li> <li>A <strong>guide to preparing Lighttag annotation results for manual review offline </strong>(&ldquo;Using Spreadsheets to Review Annotations Offline.pdf&rdquo;).</li> </ol> <p>The project&#39;s findings regarding the status of NER and NLP for vertical Mongolian are available at DOI: 10.5281/zenodo.5103499.</p>

opencc-by-4.0Aug 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record