Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

46

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

46 results for “Text Collection”

Learn how ShareScore rates datasets ↗
zenodo36/100

The Annotated Corpus of Classical Tibetan (ACTib), Part II - POS-tagged version, based on the BDRC digitised text collection, tagged with the Memory-Based Tagger from TiMBL

<p>This corpus is a part-of-speech tagged version of</p> <p>Wallman, Jeff, Rowinski, Zach, Ngawang Trinley, Tomlinson, Chris, &amp; Keutzer, Kurt. (2017). Collection of Tibetan etexts compiled by the Buddhist Digital Resource Center [Data set]. Zenodo. http://doi.org/10.5281/zenodo.821218</p> <p>using the training data of</p> <p>Hill, Nathan W., &amp; Garrett, Edward. (2017). A part-of-speech (POS) tagged corpus of Classical Tibetan [Data set]. Zenodo. http://doi.org/10.5281/zenodo.574878</p> <p>Please note that the files are not post-processed or manually corrected and that a small number of files in the KarmaDelek directory were still annotated, although the original xml-input was corrupted already.</p> <p>&nbsp;</p> <p>using the memory based tagger of</p> <p>https://languagemachines.github.io/mbt/</p>

opencc-by-4.0Jul 2017View details →
zenodo36/100

Text-fig. 12. Lectotype specimen of Potentilla×hybrida WALLROTH. in Wallroth´S Collection Of Vascular Plants In The Herbarium Of The National Museum, Prague

Text-fig. 12. Lectotype specimen of Potentilla×hybrida WALLROTH.

opencc-by-4.0Sep 2008View details →
zenodo36/100

Text-fig. 13. Lectotype specimen of Senecio germanicus WALLROTH. in Wallroth´S Collection Of Vascular Plants In The Herbarium Of The National Museum, Prague

Text-fig. 13. Lectotype specimen of Senecio germanicus WALLROTH.

opencc-by-4.0Sep 2008View details →
zenodo36/100

Text-fig. 11. Lectotype specimen Orobanche rubens WALLROTH. in Wallroth´S Collection Of Vascular Plants In The Herbarium Of The National Museum, Prague

Text-fig. 11. Lectotype specimen Orobanche rubens WALLROTH.

opencc-by-4.0Sep 2008View details →
zenodo36/100

Text-fig. 8. Lectotype specimen of Malva neglecta WALLROTH. in Wallroth´S Collection Of Vascular Plants In The Herbarium Of The National Museum, Prague

Text-fig. 8. Lectotype specimen of Malva neglecta WALLROTH.

opencc-by-4.0Sep 2008View details →
zenodo36/100

Text-fig. 4. Lectotype specimen of Camelina sylvestris WALLROTH. in Wallroth´S Collection Of Vascular Plants In The Herbarium Of The National Museum, Prague

Text-fig. 4. Lectotype specimen of Camelina sylvestris WALLROTH.

opencc-by-4.0Sep 2008View details →
zenodo36/100

Text-fig. 1. Wallroth´s original handwriten page dealing with Valeriana collina WALLROTH. in Wallroth´S Collection Of Vascular Plants In The Herbarium Of The National Museum, Prague

Text-fig. 1. Wallroth´s original handwriten page dealing with Valeriana collina WALLROTH.

opencc-by-4.0Sep 2008View details →
zenodo32/100

A collection of Wa texts and a Wa wordlist (Bible orthography) for use in NLP

<p>This is a collection of Wa texts and a word list for use in NLP, following the orthography prevalent in Burma.</p>

openother-pdDec 2020View details →
zenodo28/100

A Korean raw text collection for creating a language model

<p><strong>A very large Korean raw text collection for creating a language model&nbsp;</strong></p> <p>&nbsp;</p> <p>We collected&nbsp;a very large monolingual dataset for Korean, which contains <strong>over 9.6M sentences and 130.6M eojeols</strong>, to create a language model: Korean Wikipedida (https://dumps.wikimedia.org/kowiki/20201101/, 5.3M sentences and 71.8M eojeols, respectively), the Sejong morphologically analyzed corpus (3.0M and 40.0M), and articles from <em>The Hankyoreh</em>&nbsp;daily newspaper during 2016 (1.2M and 18.6M).&nbsp;</p> <p>&nbsp;</p> <p>We preprocessed&nbsp;raw text into morpheme-segmented text &nbsp;using the POS tagging system (<a href="https://www.aclweb.org/anthology/W19-4022/">park-tyers:2019:LAW</a>).&nbsp;We also attached the POS label to the morpheme-segmented lexicon, and explicitly include a + symbol for consecutive morphemes.&nbsp;</p> <blockquote> <p>시인/NNG 윤동주/NNP +,/SP 이준익/NNP 감독/NNG 영화/NNG +로/JKB 부활/NNG</p> </blockquote> <p>&nbsp;</p> <p>See&nbsp;https://github.com/jungyeul/sjmorph for the&nbsp;POS tagging system described in&nbsp;<a href="https://www.aclweb.org/anthology/W19-4022/">park-tyers:2019:LAW</a>.&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo28/100

A collection of Wa texts and a Wa wordlist (PRC orthography) for use in NLP

<p>This is a collection of Wa texts for use in NLP, following the orthography prevalent in China.</p>

openother-pdDec 2020View details →
zenodo28/100

A collection of Drenjong texts for use in NLP

<p>This is a collection of Drejong (Sikkimese) texts for use in NLP.</p>

openother-pdDec 2020View details →
zenodo28/100

ELTeC-NIF: European Literary Text Collection LLOD

<p>The &nbsp;European Literary Text Collection - ELTeC is transformed into Linguistic Linked Open Data text corpora using the NLP Interchange Format (NIF). Namely, the ELTEC corpus subset, which consists of 1000 novels from the period 1840-1920 for 10 European languages, served as the basis for this edition. From each novel, not more than 1000 sentences were used. The annotated version of the novels, in the so-called TEI level-2 format, was transformed into NIF, an RDF/OWL-based format that aims to achieve interoperability between NLP tools, language resources, and annotations. &nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo28/100

Figureȱ 2.ȱ Graphȱ plottingȱ skullȱ max.ȱ widthȱ againstȱ totalȱ lengthȱ (billȱ tipȱ toȱ rearȱ ofȱ skull)ȱ forȱ skinsȱ ofȱ threeȱ species of Campephilus woodpecker in the NHMUK collection. Also shown are analogous measurements fromȱ theȱ unidentifiedȱ NHMUKȱ skeleton,ȱ forȱ whichȱ aȱ correctionȱ factorȱ upwardsȱ ofȱ 17.28%ȱ hasȱ beenȱ madeȱ to total skull length, to account for its missing rhamphotheca (see text for explanation), thereby making its measurements directly comparable with the others. in The conundrum of an overlooked skeleton referable to Imperial Woodpecker Campephilus imperialis in the collection of the Natural History Museum at Tring

Figureȱ 2.ȱ Graphȱ plottingȱ skullȱ max.ȱ widthȱ againstȱ totalȱ lengthȱ (billȱ tipȱ toȱ rearȱ ofȱ skull)ȱ forȱ skinsȱ ofȱ threeȱ species of Campephilus woodpecker in the NHMUK collection. Also shown are analogous measurements fromȱ theȱ unidentifiedȱ NHMUKȱ skeleton,ȱ forȱ whichȱ aȱ correctionȱ factorȱ upwardsȱ ofȱ 17.28%ȱ hasȱ beenȱ madeȱ to total skull length, to account for its missing rhamphotheca (see text for explanation), thereby making its measurements directly comparable with the others.

opencc-by-4.0Mar 2021View details →
zenodo28/100

Text-fig. 1. The geographic position of the localities mentioned in the text. A – position within the Czech Republic. B – Detailed map of the area. Subsilesian Unit: 1 – Kelč, 2 – Špičky, 3 – Horní Těšice; Silesian Unit: 4 – Loučka, 5 – Osíčko, 6 – Rožnov pod Radhoštěm; Ždánice Unit: 7 – Bohuslavice, 8 – Jestřabice, 9 – Kožušice, 10 – Křepice, 11 – Litenčice, 12 – Mouchnice, 13 – Nikolčice, 14 – Nítkovice, 15 – Nosislav, 16 – Židlochovice. The distribution of the units according to Čtyřoký and Stráník (1995). C – map of the Osíčko vicinity with distribution of the Silesian Unit sediments (gray spots) and collecting points (P1 and P2). The distribution of the Silesian Unit sediments according to Stráník (1999). in An Annotated List Of The Oligocene Fish Fauna From The Osíčko Locality (Menilitic Fm.; Moravia, The Czech Republic)

Text-fig. 1. The geographic position of the localities mentioned in the text. A – position within the Czech Republic. B – Detailed map of the area. Subsilesian Unit: 1 – Kelč, 2 – Špičky, 3 – Horní Těšice; Silesian Unit: 4 – Loučka, 5 – Osíčko, 6 – Rožnov pod Radhoštěm; Ždánice Unit: 7 – Bohuslavice, 8 – Jestřabice, 9 – Kožušice, 10 – Křepice, 11 – Litenčice, 12 – Mouchnice, 13 – Nikolčice, 14 – Nítkovice, 15 – Nosislav, 16 – Židlochovice. The distribution of the units according to Čtyřoký and Stráník (1995). C – map of the Osíčko vicinity with distribution of the Silesian Unit sediments (gray spots) and collecting points (P1 and P2). The distribution of the Silesian Unit sediments according to Stráník (1999).

opencc-by-4.0Dec 2013View details →
zenodo28/100

Text-fig. 2. Vertebral profiles of Auxis thazard. VL, central length; VH, central height; VW, central width. in On The Morphology Of The Vertebral Column Of The Frigate Tuna, Auxis Thazard (Lacepedea, 1800) (Family: Scombridae) Collected From The Sea Of Oman

Text-fig. 2. Vertebral profiles of Auxis thazard. VL, central length; VH, central height; VW, central width.

opencc-by-4.0Sep 2013View details →
zenodo24/100

Is there a (proto-)Lucianic stratum in the text of 1 Kings of the Old Latin manuscript La115? (Collected cases)

<p>Cases for the article &ldquo;Is there a (proto-)Lucianic stratum in the text of 1 Kings of the Old Latin manuscript La<sup>115</sup>?&rdquo; In Kristin De Troyer (ed.), <em>On Hexaplaric and Lucianic Readings and Recensions </em>(DSI; forthcoming).</p>

opencc-by-4.0Nov 2020View details →
zenodo24/100

Text-fig. 1. Vertebral column of Auxis thazard showing different regions in On The Morphology Of The Vertebral Column Of The Frigate Tuna, Auxis Thazard (Lacepedea, 1800) (Family: Scombridae) Collected From The Sea Of Oman

Text-fig. 1. Vertebral column of Auxis thazard showing different regions

opencc-by-4.0Sep 2013View details →
zenodo24/100

Text-fig. 9. Lectotype specimen of Monotropa hypophegea WALLROTH. in Wallroth´S Collection Of Vascular Plants In The Herbarium Of The National Museum, Prague

Text-fig. 9. Lectotype specimen of Monotropa hypophegea WALLROTH.

opencc-by-4.0Sep 2008View details →
zenodo24/100

Text-fig. 10. Lectotype specimen Montia arvensis WALLROTH. in Wallroth´S Collection Of Vascular Plants In The Herbarium Of The National Museum, Prague

Text-fig. 10. Lectotype specimen Montia arvensis WALLROTH.

opencc-by-4.0Sep 2008View details →
zenodo24/100

Text-fig. 5. Lectotype specimen of Hieracium lactucella WALLROTH. in Wallroth´S Collection Of Vascular Plants In The Herbarium Of The National Museum, Prague

Text-fig. 5. Lectotype specimen of Hieracium lactucella WALLROTH.

opencc-by-4.0Sep 2008View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record