Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4
datasets available to search
ShareScore release 0.9.0
Dataset results
4 results for “Newar”
OCR model for Pracalit for Sanskrit and Newar MSS 16th to 19th C., Ground Truth
<p>Ground truth data (png and xml files) for a an OCR model. Will be continually updated.</p> <p>Originally trained on Transkribus with a PyLaia model created from ground truth data based on transcripts into Pracalit Unicode of four Nepalese manuscripts. The manuscripts used to create this model are Staatsbibliothek zu Berlin's Hitopadeśa (MIK I 4851) (mixed Newar and Sanskrit dating to 1561) and Vetālapañcaviṃśati (HS. Or. 6414) (Newar dating to 1675) as well as Cambridge Digital Library's Avalokiteśvaraguṇakāraṇḍavyūha (MS Add. 1322) (Sanskrit, 18th century) and the Royal Asiatic Society Online Collection's Madhyamasvayaṃbhūpurāṇa (RAS Hodgson MS 23) (Newar and Sanskrit dating to c. 1800).</p> <p>The training was done on 441 pages and validation on 242 pages.</p> <p>This model does not recognise spacing, except for large gaps (i.e. for pictures or string holes). Newar word divider markers may not be represented or may be transcribed as virama. In general, the model is made for MSS with scriptio continua and will transcribe into scriptio continua into Pracalit Unicode.</p> <p>Transcription was performed by Dr Alexander O'Neill (SOAS University of London). Transcription of the Vetālapañcaviṃśati (HS. Or. 6414) and Madhyamasvayaṃbhūpurāṇa (RAS Hodgson MS 23) was aided by unpublished materials provided by Dr Felix Otter (Philipps-Universität Marburg), as well as the published transcription in Shakya, Min Bahadur, and Shanta Harsha Bajracharya, eds. "Svayambhū Purāṇa." Lalitpur: Nagarjuna Institute of Exact Methods, 2001. The transcription of Avalokiteśvaraguṇakāraṇḍavyūha (MS Add. 1322) was aided by the transcription provided by the Digital Sanskrit Buddhist Canon Project based on Lokesh Chandra, "Guṇakāraṇḍavyūhasūtram," New Delhi: International Academy of Indian Culture, 1999.</p>
Diachronic Annotated Corpus of Newar
<p>This dataset contains segmented and part-of-speech-tagged files that comprise the ongoing Diachronic Annotated Corpus of Newar (DACON). Files are provided in .txt format.</p> <p>File names are explained as follows:</p> <p>cnew (Classical Newar)<br>century (e.g. 12)<br>short text name (and other information such as manuscript name (e.g., MSB) and line number completed for incomplete texts (e.g. 10000))<br>SEG or POS (segmented or part-of-speech-tagged</p> <p>e.g. cnew19-manicuda-10000_SEG.txt</p> <p>For full details, including text citations with discussion, please see: <br>O'Neill & Meelen, "The Diachronic Annotated Corpus of Newar: from Manuscript to Morphosyntax," <em>Cahiers de Linguistique Asie Orientale</em> (2024).</p> <p><a href="../doi/10.5281/zenodo.13117922">Annotation Manual Part I (Preprocessing)</a></p> <p><a href="../doi/10.5281/zenodo.13117961">Annotation Manual Part II (Segmentation and POS Tagging)</a></p> <p><a href="https://github.com/lothelanor/newarcorpora">Tools for the Diachronic Annotated Corpus of Newar (DACON)</a></p>
Wordlist of Dolakha Newar
<p>This is the word list in the appendix of Genetti's Newari Grammar.</p>
Lalitpur Newar 2022
<p>Dataset of recordings, plus metadata, for the study of the Lalitpur Newar dialect. Collected in 2022 with the consent of participants for its use for research purposes.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.