Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,305

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,305 results for “text”

Learn how ShareScore rates datasets ↗
zenodo44/100

Duhumbi Procedural Texts - Transcribed, parsed, glossed, translated text files

<p>This data set contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the following recordings that belong to the following publication:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <ul> <li>[CHUK210512I1] /&nbsp;PHPT / Hunting for porcupine&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</li> <li>[CHUK230512A1A] /&nbsp;SBDC / Making fermented soybean</li> <li>[CHUK220413A1] /&nbsp;CTTT / Catching frogs&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</li> <li>[CHUK220413B1] /&nbsp;CWTT / Collecting beeswax&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</li> <li>[CHUK220413C1] /&nbsp;CHTT / Collecting hornets</li> <li>[CHUK240413A1] /&nbsp;SNAP / Collecting stinging nettle&nbsp;</li> <li>[CHUK240413B1] /&nbsp;LCYT / Leather craft&nbsp; &nbsp; &nbsp;&nbsp;</li> </ul> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file &ldquo;Settings&rdquo;, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>For the metadata of the sound files in this data set, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for&nbsp;commercial purposes&nbsp;<em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration &amp; payment for access, or sites that rely on advertisement (including YouTube)&nbsp;</em>is&nbsp;<strong>not</strong>&nbsp;permitted without&nbsp;<strong>specific written consent</strong>&nbsp;from the speakers and their community, obtained through the collector&nbsp;of the material. By downloading this material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit properly and license your new creations under the identical terms. License Deed on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p>

opencc-by-4.0Aug 2018View details →
zenodo44/100

Duhumbi Discussions - Transcribed, parsed, glossed, translated text files

<p>This dataset contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the following recordings that belong to the following publication:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <ul> <li>[CHUKxxxx13A6] / LEL / Local elections (not included on speaker&#39;s request)</li> <li>[CHUK300412J2] / LGT / Planning a trip to Lagam</li> <li>[CHUK260413A2A] / NNK / Nicknames</li> </ul> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file &ldquo;Settings&rdquo;, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>For the metadata of the sound files in this data set, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for&nbsp;commercial purposes&nbsp;<em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration &amp; payment for access, or sites that rely on advertisement (including YouTube)&nbsp;</em>is&nbsp;<strong>not</strong>&nbsp;permitted without&nbsp;<strong>specific written consent</strong>&nbsp;from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p>

opencc-by-4.0Aug 2018View details →
zenodo44/100

Biodiversity Hotspots Map (no text)

<p>This map displays the global <a href="https://zenodo.org/record/3261807#.X8_HgNhKg2x">Biodiversity Hotspots 2016.1 dataset</a>. The colors assigned to the hotspots are only used to distinguish adjacent hotspots and have no other meaning. The biodiversity hotspots represent terrestrial biodiversity only. The offshore lines are a cartographic device to group and highlight islands that are part of the same hotspots (e.g. Polynesia-Micronesia, Indo-Burma).&nbsp;The background image is from Natural Earth. This version is without labels; a version with the hotspots labelled in English is available:&nbsp;10.5281/zenodo.4311850</p> <p>There are currently&nbsp;<a href="https://www.cepf.net/node/1996">36 recognized biodiversity hotspots</a>. These are Earth&rsquo;s most biologically rich&mdash;yet threatened&mdash;terrestrial regions.</p> <p>To qualify as a biodiversity hotspot, an area must meet two strict criteria:</p> <ul> <li>Contain at least 1,500 species of vascular plants found nowhere else on Earth (known as &quot;endemic&quot; species).</li> <li>Have lost at least 70 percent of its primary&nbsp;native vegetation.</li> </ul> <p>Many of the biodiversity hotspots exceed the two criteria. For example, both the Sundaland Hotspot in Southeast Asia and the Tropical Andes Hotspot in South America have about&nbsp;<strong>15,000</strong>&nbsp;endemic plant species. The loss of vegetation in some hotspots has reached a startling&nbsp;<strong>95</strong>&nbsp;percent.</p>

opencc-by-4.0Jun 2016View details →
zenodo44/100

Connecting U.S. Supreme Court Case Information and Opinion Authorship (SCDB) to Full Case Text Data (CAP), 1791-2011

<p>This dataset was constructed to connect the rich metadata created by the Supreme Court Database (SCDB) to the Caselaw Access Project (CAP) full-text court opinion data. Since the SCDB includes only substantive opinions, it is necessarily a subset of the full range of opinions available through CAP.</p> <p>There are two parts to this data: the map connecting each SCDB ID to its corresponding CAP case number, and a more advanced (but error-prone) version in which the authorship of each opinion text identified for the case in CAP is attributed to the Justice who wrote it. Each of these data products have been hand-corrected to the best of this author&#39;s ability.</p> <p><strong>SCDB-CAP map</strong></p> <p>The SCDB-&gt;CAP map began as a relatively straightforward automated matching process, based on the US Reports citation for each case as expressed in both SCDB and CAP. Slightly over 80% of SCDB entries found a single CAP data match this way. From there, the data was entirely hand-corrected, with non-matches or duplicate matches individually investigated and manually corrected.</p> <p>Some SCDB entries simply could not be matched to an appropriate CAP text. Initially, the entirety of US Reports volume 44 was missing, but with the help of CAP staff, the volume was located as having been filed in the New York jurisdiction rather that the United States jurisdiction. The case numbers were then added to the map, but until the volume is relocated to the United States jurisdiction, it may be necessary to also incorporate the New York jurisdiction in full text analysis so that the cases from volume 44 can be searched. 108 more missing cases are from US Reports volume 131, which was a &quot;catch up&quot; volume published in the 19th century. These catch-up cases, many heard by the Supreme Court decades prior, were numbered with lowercase roman numerals instead of the ordinary&nbsp; numbers, which is almost certainly why CAP&#39;s software dismissed the catch-up section as prefatory material. Many of the rest of the errors seem largely to be examples where the SCDB project recognized a separate court action that CAP did not. Perhaps most of these seem to have been later rehearings for a case previously decided, which in the 19th century particularly were commonly reported out at the end of the first decision text. While SCDB sometimes gave these subsequent but related actions a separate SCDB entry, CAP seems to have largely incorporated them as part of the text of the main case. Additionally, there were a few that simply could not be found, despite a careful look through each database as well as the original US Reports and sometimes adjacent volumes. Finally, the cases were only matched up through the 2011 court term. After the 2011 term, the mismatches between CAP and SCDB were extensive and frequently seemed impossible to resolve.</p> <p>Even so, with the manual correction, the overall error rate is low. Of 28,304 cases, only 191 do not have a match, and of those, 108 are contained within the vol. 131 &quot;catch up&quot; volume. Since most of the rest are extremely short subsequent actions that were separately noted by SCDB, the effect of these non-matched cases would seem to be small in most cases.</p> <p>The typical use case would be that the researcher would generate some kind of results based on searching in the CAP full text, then could use the CAP ID to look up the SCDB ID in the map. With the SCDB ID, of course, the rich metadata from the SCDB can then be connected to each result as needed.</p> <p><strong>Opinion authorship</strong></p> <p>Being able to use the rich metadata of SCDB in conjunction with a case&#39;s full text is exciting, but it immediately prompts a further question -- what if the texts could be attributed directly to the Justices who authored them? SCDB produces its data in two forms; one is &quot;case centered,&quot; where each record represents one case, and the other is &quot;justice centered,&quot; in which each record is the vote of one Justice in one case. CAP, in turn, breaks the total text of the case into distinct opinions, and tries to attribute those opinions to their authors by scraping a string of text from the raw input. Therefore, the challenge was to connect these two sources at the opinion level.</p> <p>Connecting the opinions, like connecting the cases, involved an initial match by machines, followed by manual correction and revision. In this case, the scope of the manual effort was much larger than that posed by the case-level connection, and more errors were noted in both SCDB and CAP.</p> <p>The matching process involved a number of steps. First a list of opinions was generated from the CAP data, then matched to SCDB using the SCDB-CAP connector data described above. (Thus, a case without a CAP match in the SCDB-CAP data will not appear in the opinion author data either.) CAP opinions were numbered in the order they were encountered in each CAP case JSON object, and these numbers are used to distinguish the opinions.</p> <p>Next, a round of automatic matching was performed. If there was only one opinion, and only one author listed in the SCDB data, then the majority opinion author (as listed in SCDB) was safely assumed to be the author. If there was no author listed in SCDB, &quot;percuriam&quot; was recorded as the author in this data. If there were exactly two opinions and two authors, the process was also straightforward, as the SCDB-identified majority opinion author was assigned to opinion 1, and the remaining author assigned opinion 2.</p> <p>Subsequently, cases with more than two opinions were processed. A potential match (i.e. a &quot;guess&quot;) for each opinion in a given case was created by listing each Justice identified by SCDB as having written an opinion in the case. These guesses were then parsed using a semi-automatic procedure with Levenshtein distance fuzzy name matching. With sufficiently conservative parameters, a successful fuzzy match meant that the non-successful guesses for that opinion could be deleted. These sorted guesses were then reviewed manually. Particular care was also taken for any opinion that contained authored opinions by Justices who had similar names (for example, Clark and Black differ by only a single letter). These sorts of cases, as well as instances of co-authorship, were identified and fixed manually.</p> <p>Those opinions whose authorship could not be matched then were fixed by hand. These included some where the CAP author strings were more complicated than SCDB&#39;s strict interpretation; others where the OCR in CAP which contained the Justice name was especially bad; and a number of others where &quot;Mr. Chief Justice&quot; couldn&#39;t be directly matched with an author name by the machine. After this light manual correction, almost 500 opinions with substantial errors remained to be individually investigated in depth, by examining the CAP record, the SCDB record, and images of the US Reports for that case. For these last tough customers, errors in the source data were commonly the cause of matching problems. Typically these were of three kinds: examples where CAP should have split the text but didn&#39;t (e.g. 2 opinions together in one opinion entry in CAP); examples where SCDB either did not identify or mis-identified an author (such as attributing it to Swayne when it was written by Miller); and examples of non-valid opinions (such as where CAP mistakenly split the opinion too early, leaving an opinion fragment).</p> <p>For these errors, a system of codes was created in the author field to signal the error type so that researchers can be suitably cautious. The error code is always at the beginning of the field and is followed by a comma and the names of each author, separated by a comma with no space to facilitate parsing. Note also that co-authors are listed as comma-separated names in this same field with no error code. Researchers will probably want to disaggregate this field to create duplicate records with each individual author for most purposes. The justice number field also contains information about all justices authoring the opinion but the error codes have been omitted here.</p> <ul> <li>!C -- error: multiple opinion texts combined (i.e. CAP splitting error)</li> <li>!X -- error: unattributed or misattributed opinion (not listed in SCDB as writer)</li> <li>!D -- error: extra opinion that should be deleted, i.e. not a valid opinion</li> <li>!W -- error: listed as Writers by SCDB, but should be co-authors</li> </ul> <p>&nbsp;</p> <p><strong>Data file structure</strong></p> <p>&quot;scdb_cap-051820.tsv&quot; is a Tab-separated data file containing 5 columns: SCDB ID, CAP ID, US Reports citation, case date, and case name (the latter three from the SCDB data).</p> <p>&quot;scdb-cap-opinion-authorship_051920.tsv&quot; is a Tab-separated data file containing seven columns: SCDB ID, CAP ID, US Reports citation, case name, opinion number in the case, opinion author, and SCDB justice ID. See above for caveats about disaggregating and error codes in fields six and seven.</p> <p><strong>Errors</strong></p> <p>It is likely that errors remain in this data, and it is also hoped that some of the errors beyond the author&#39;s immediate control might be fixed in the upstream data so that they can be corrected here. Authors would be grateful for error reports, and also reports of errors fixed, if any.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Transkribus - Handwritten Text Recognition for Premodern Documents (SIMS 2020 Lightning Talk)

<p>Transkribus is a platform for text recognition and can be used via the Transkribus Expert Software (available after registration: transkribus.eu). Through Transkribus different tools for document analysis and text recognition can be directly applied. The intro demonstrates very briefly how Transkribus can help with regards to premodern documents especially since a variety of pre-trained models are already available: for Latin (prints and handwriting), for early modern vernaculars in French, Dutch, English, and German. For more information go to transkribus.eu.</p> <p>Presented as a Schoenberg Symposium 2020 Lightning Talk</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

WOLOF TTS(Text To Speech) Data

<p>This contains&nbsp;a WOLOF Text To Speech(TTS) dataset, it contains recordings from two natif Wolof actos (a male and female voice).<br> Each actor recored more than 20 000 sentences.<br> The notebook accompanying the dataset contains a brief analysis of the dataset and the code creating the appropriate train/validation and test set.<br> The file [male, female]train, [male, female]validation and [male, female]test are also present to extract the corespondig audios inside the data-commonvoice.zip</p> <p>The text dataset come from news website, Wikipedia and self curated text. We made sure with the help of our Wolof expert that the text dataset cover the different phonemes in the Wolof language.</p>

opencc-by-4.0Feb 2021View details →
zenodo44/100

AmphibiaWeb: AmphibiaWeb text w/traits based on Pensoft Annotator

from <p></p>http://amphibiaweb.org/. AmphibiaWeb is an online system enabling anyone with a Web browser to search and retrieve information relating to amphibian biology and conservation. This site was inspired by the global declines of amphibians, the study of which has been hindered by the lack of multidisplinary studies and a lack of coordination in monitoring, in field studies, and in lab studies. One of its major goals is to encourage a shared vision for the study of global amphibian declines and the conservation of remaining amphibians.<p></p><p></p>https://annotator.pensoft.net/about

opencc-by-4.0Aug 2024View details →
zenodo44/100

Text Generation using N-gram and GPT (metrics: ROUGE, BLEU and BERTScore)

<p>This publication presents a set of spreadsheets listing the user stories generated using N-gram and GPT models with metrics ROUGE, BLEU and BERTScore calculated. Each spreadsheet refers to one corpus of user stories processed.</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

The Curated Courier: Digital Text Corpora from the UNESCO Courier (1948–2020)

<p>Founded in 1948 as the official magazine of the United Nations Educational, Scientific and Cultural Organization, <i>The UNESCO Courier</i> represents an extraordinary resource for research on global themes in the humanities. The complete <a href="https://en.unesco.org/courier/archives">archive of the magazine</a> is available in PDF form through UNESCO.&nbsp;These files make it possible for users anywhere to read individual issues, but it does not allow for full-text searching, much less any of the computational text analysis methods that have recently made important advances in humanities research.</p><p>The Curated Courier 1.0 is a package of digital text corpora, text analysis tools, and supplementary materials that makes the complete archive of <i>The UNESCO Courier</i> from 1948 to 2020 machine-readable, accessible, and reusable for digital text analysis.&nbsp;</p><p>Here on Zenodo we publish two <i>Courier</i> corpora. The first corpus (curated_courier_article_corpus) consists of the texts of all articles published in the English-language edition of <i>The UNESCO Courier</i> between 1948 and 2020. For this corpus we have extracted and reconstructed&nbsp;the complete&nbsp;text of all articles, for example by pulling together non-contiguous pages where necessary and by removing non-article text&nbsp;(masthead, photo captions, letters to the editor, and so on). We have linked each article&nbsp;to a comprehensive curated metadata index, included in the download (document_index.csv).</p><p>The second corpus&nbsp;(curated_issues) compiles&nbsp;the complete text of all <i>Courier</i> issues (English-language edition), 1948-2020. To prepare this corpus we extracted text from <a href="https://en.unesco.org/courier/archives">the PDFs that UNESCO has made available</a>, used multiple modes of OCR, and rendered each issue as a simple text file. Our test of the OCR quality finds an average&nbsp;error rate of 0.7 %, which should be considered good quality.</p><p>Working data from the process can be found in our <a href="https://github.com/inidun/tagged_courier">GitHub repository "tagged Courier."</a> The products, text analysis tools, and additional documentation are in the <a href="https://github.com/inidun/curated_courier">repository "Curated Courier."</a></p><p>The text of <i>The UNESCO Courier</i> is&nbsp;<a href="https://courier.unesco.org/en/about">available in Open Access</a> under the Attribution-ShareAlike 3.0 IGO (CC-BY-SA 3.0 IGO) license, in the context of <a href="https://en.unesco.org/open-access/">UNESCO's open access publications policy</a>. This dataset is published under the most recent version of the same license: Attribution-ShareAlike 4.0 International (<a href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">CC BY-SA 4.0 Deed</a>).</p><p>These datasets&nbsp;was developed as part of the research project "International Ideas at UNESCO: Digital Approaches to Global Conceptual History" (INIDUN), led by Benjamin G. Martin at Uppsala University and funded by a grant from the Swedish Research Council (Vetenskapsrådet), 2020-2024. For more information, see: <a href="https://inidun.github.io">https://inidun.github.io</a>, as well as the<a href="https://github.com/inidun"> project repository on GitHub</a>, which includes documentation and files related to the curating process.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Accelerating Digital Skills for Music Researchers - Processing Text-Based Corpora for Musical Discourse Analysis - Episode 5

<p>Dataset containing four .xlsx and .csv files for the exercises in Episode 5 of the&nbsp;<a href="https://acceleratingdigitalskills.github.io/Processing-Text-Based-Corpora/">Processing Text-Based Corpora for Musical Discourse Analysis</a>&nbsp;lesson of the&nbsp;<a href="https://acceleratingdigitalskills.org/">Accelerating Digital Skills for Music Researchers</a>&nbsp;project. The original data was collected from&nbsp;<a href="https://boomkat.com/">Boomkat.com</a> with permission.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Glossed Hittite Texts with German Translation for Machine Learning

<p>This dataset contains 7,099 processed Hittite texts from 143 CTH numbers, sourced with permission from the <a href="https://www.hethport.uni-wuerzburg.de/HPM/index.php" target="_blank" rel="noopener"><strong>Hethitologie Portal Mainz (HPM)</strong></a>, which is the comprehensive resource of Hittite texts and culture (modern Turkey, c. 1,650 - 1,200 BCE). These texts have been converted from XML format to a tabular structure for computational and linguistic analysis, as well as machine learning applications, while attempting to preserve the full complexity of the source material.</p> <p><strong>Key Features:</strong></p> <ul> <li><strong>txtid</strong>: Identifier for the specific text (e.g., "IBoT 1.30+"), referencing entries in the HPM Konkordanz (S. Ko&scaron;ak, hethiter.net/: hetkonk (2.plus)).</li> <li><strong>lnr</strong>: Surface, column (if extant), and line number within the text.</li> <li><strong>cth_number</strong>: <em>Catalogue des textes hittites </em>(CTH) classification, organizing texts by genre and content (S. Ko&scaron;ak &ndash; G.G.W. M&uuml;ller &ndash; S. G&ouml;rke &ndash; Ch.W. Steitler, hethiter.net/: CTH (2022-10-26)).</li> <li><strong>word</strong>: Orthographic form of the text as presented in the HPM, not in cuneiform but in a standard Latin script representation.</li> <li><strong>translit</strong>: Detailed transliteration, preserving all nuances such as diacritics, broken parts, and editorial markings to reflect the original condition on the text.</li> <li><strong>gloss</strong>: Linguistic glosses, including grammatical, morphological, and semantic information.</li> <li><strong>trans_de</strong>: German translation of the word or phrase.</li> </ul>

opencc-by-4.0Dec 2024View details →
zenodo44/100

Patristisches Textarchiv. Ein Open Access-Archiv antiker christlicher Texte

The "Patristic Text Archive" offers anyone interested a collection of texts and translations of Christian texts from antiquity (i.e. "Patristic" is conceived in a very broad sense).

opencc-zeroOct 2022View details →
zenodo44/100

Antisemitism on Twitter: A Dataset for Machine Learning and Text Analytics

<h1><strong><span><span>Dataset from the Institute for the Study of Contemporary Antisemitism (ISCA) at Indiana University:&nbsp; </span></span></strong></h1> <p>&nbsp;</p> <div> <div> <p><span><span>The </span><span>Social Media</span><span> &amp; Hate research lab at the Institute for the Study of Contemporary Antisemitism compiled this dataset using an annotation portal (Jikeli, Soemer, and Karali 2024), which was used to label tweets as either antisemitic or non-antisemitic, among other labels. Note that annotation was done on live data, including images and context, such as threads. All data was annotated by two experts, and all discrepancies were discussed</span><span> (Jikeli et al. 2023)</span><span>.</span></span><span>&nbsp;</span></p> </div> </div> <p><br><strong>Content: </strong></p> <p><span><span>This dataset&nbsp;</span><span>contains</span> <span>1</span><span>1</span><span>311</span><span> tweets </span><span>covering</span><span>&nbsp;a wide range of topics common in conversations about Jews, Israel, and antisemitism between January 2019 and </span><span>April 2023</span><span>.&nbsp;</span><span>The dataset consists of random samples of relevant keywords during this </span><span>time period</span><span>.</span><span>&nbsp;1,</span><span>953</span><span> tweets (1</span><span>7</span><span>%) </span><span>are antisemitic </span><span>according to </span><span>the IHRA definition of antisemitism.</span><span>&nbsp;&nbsp;</span></span><span>&nbsp;</span></p> <div> <p><span><span>The distribution of tweets by year is as follows:</span><span>&nbsp;1499 (</span><span>13</span><span>%) from 2019, 371</span><span>2</span><span> (</span><span>33</span><span>%) from 2020, </span><span>2591</span><span> (2</span><span>3</span><span>%) from 2021</span><span>, 2644 from 2022 </span><span>(23%)</span> <span>and 865 </span><span>(8%)</span> <span>f</span><span>rom 2023</span><span>. </span><span>6365</span><span> (</span><span>56</span><span>%) </span><span>contain</span><span> the keyword "Jews,"</span><span> 4134 </span><span>(</span><span>3</span><span>7</span><span>%) include "Israel," 529 (</span><span>5</span><span>%) feature the derogatory term "</span><span>ZioNazi</span><span>*," and 283 (</span><span>3</span><span>%) use the slur "K---s." Some tweets may </span><span>contain</span><span> multiple keywords.&nbsp;</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>725</span><span> out of the </span><span>6365</span><span> tweets with the keyword "Jews" (11%) and </span><span>664</span><span> out of the </span><span>4134</span><span> tweets with the keyword "Israel" (1</span><span>6</span><span>%) were classified as antisemitic. 97 out of the 283 tweets using the antisemitic slur "K---s" (34%) are antisemitic.</span> <span>Interestingly, many tweets featuring the slur "K---s" actually </span><span>call out</span><span> its u</span><span>s</span><span>e.</span><span> In contrast, </span><span>the majority of</span><span> tweets </span><span>using</span><span> the derogatory term "</span><span>ZioNazi</span><span>*" are antisemitic, with 467 out of 529 (88%) being classified as such.&nbsp;</span></span><span>&nbsp;</span></p> </div> <p>&nbsp;</p> <p><strong>File Description:&nbsp;</strong></p> <div> <div> <p><span><span>The dataset is provided in a csv file format, with each row </span><span>representing</span><span> a single message, including replies, quotes, and retweets. The file </span><span>contains</span><span> the following columns:&nbsp;</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&nbsp;</span></span><span><span>&nbsp;</span><br></span><span><span>&lsquo;ID&rsquo;:</span> <span>Represents</span><span> the tweet ID.&nbsp;</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&lsquo;Username&rsquo;: </span><span>Represents</span><span> the username </span><span>that posted </span><span>the tweet</span><span>.&nbsp;&nbsp;</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&lsquo;Text&rsquo;: </span><span>Represents</span><span> the full text of the tweet (not pre-processed).</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&lsquo;</span><span>CreateDate</span><span>&rsquo;: </span><span>Represents</span><span> the date </span><span>on which </span><span>the tweet was created</span><span>.&nbsp;&nbsp;</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&lsquo;Biased&rsquo;: </span><span>Represents</span><span> the label given by our annotations as to whether the tweet </span><span>is antisemitic or no</span><span>t</span><span>.</span></span><span>&nbsp;</span></p> </div> <div> <p><span><span>&lsquo;Keyword&rsquo;: </span><span>Represents</span><span> the keyword that was used in the query. The keyword can be in the text, including </span><span>hashtags, </span><span>mentioned </span><span>users</span><span>, or the username</span><span> itself.</span><span>&nbsp;&nbsp;</span></span><span>&nbsp;</span></p> </div> </div> <p>&nbsp;</p> <p>Licences&nbsp;</p> <p>Data is published under the terms of the "Creative Commons Attribution 4.0 International" licence (https://creativecommons.org/licenses/by/4.0)&nbsp;</p> <p>&nbsp;</p> <p>Acknowledgements&nbsp;</p> <p>We are grateful for the support of Indiana University&rsquo;s Observatory on Social Media (OSoMe) (Davis et al. 2016) and the contributions and annotations of all team members in our Social Media &amp; Hate Research Lab at Indiana University&rsquo;s Institute for the Study of Contemporary Antisemitism, especially Grace Bland, Elisha S. Breton, Kathryn Cooper, Robin Forstenh&auml;usler, Sophie von M&aacute;ri&aacute;ssy, Mabel Poindexter, Jenna Solomon, Clara Schilling, and Victor Tschiskale.&nbsp;</p> <p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services &amp; Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Publication text: code, data, and new measures

<p>This Zenodo page describes data collection, processing, and different open access data files related to the text of scientific publications from OpenAlex. If you use the code or data, please cite the following paper:&nbsp;</p> <p>Sam Arts, Nicola Melluso, Reinhilde Veugelers; Beyond Citations: Measuring Novel Scientific Ideas and their Impact in Publication Text. <em>The Review of Economics and Statistics</em> 2025; 1&ndash;33 doi: <a href="https://doi.org/10.1162/rest_a_01561" target="_blank" rel="noopener">https://doi.org/10.1162/rest_a_01561</a></p> <p>&nbsp;</p>

openo-uda-1.0Oct 2024View details →
zenodo44/100

Adverse Drug Reaction (ADR) Text Dataset

<p>This repository contains text data and code related to the identification and clustering of Adverse Drug Reactions (ADR) using Sentence-BERT (S-BERT) embeddings and the SS-DBSCAN clustering algorithm. The dataset includes both labeled and unlabeled patient reports extracted from the publicly available MIMIC-III database.</p> <p>The labeled data has been manually annotated to distinguish between ADR and non-ADR cases. The unlabeled dataset is used for unsupervised clustering experiments, particularly to assess high-dimensional data clustering performance.</p> <p>New in This Version:<br>- Added Jupyter Notebook: `mimic-5k_PCA_tSNE_clustering.ipynb`<br>- Included detailed `README_ADR_Clustering_Task.txt` with step-by-step instructions to reproduce clustering results<br>- Explained how to scale experiments from 1,000 to full dataset size</p>

openmit-licenseOct 2024View details →
zenodo44/100

CWID-hi: A Dataset for Complex Word Identification in Hindi Text

<p>This dataset was created by conducting a human intelligence test, wherein native and non-native Hindi speakers annotated words they could not understand in Hindi text. They were then asked to rank the complexity of these words along with their synonyms. A word that received an average rank of &lt;=3 (out of 5) is labeled 1 and the word that received an average rank of &gt;3 is labeled 0. 1 indicates complex and 0 indicates simple.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Webis-MS-MARCO-Anchor-Texts-22

<p>The Webis MS MARCO Anchor Text 2022 dataset enriches Version 1 and 2 of the document collection of <a href="https://microsoft.github.io/msmarco/">MS MARCO</a> with anchor text extracted from six <a href="https://commoncrawl.org/">Common Crawl</a> snapshots. The six Common Crawl snapshots cover the years 2016 to 2021 (between 1.7-3.4 billion documents each). We sampled&nbsp; 1,000 anchor texts for documents with more than 1,000 anchor texts at random and all anchor texts for documents with less than 1,000 anchor texts (this sampling yields that all anchor text is included for 94% of the documents in Version 1 and 97% of documents for Version 2). Overall, the MS MARCO Anchor Text 2022 dataset enriches 1,703,834 documents for Version 1 and 4,821,244 documents for Version 2 with anchor text.</p> <p>Cleaned versions of the MS MARCO Anchor Text 2022 dataset are available in <a href="https://github.com/allenai/ir_datasets/issues/154">ir_datasets</a>, <a href="https://zenodo.org/record/5883456">Zenodo</a> and <a href="https://huggingface.co/datasets/webis/ms-marco-anchor-text">Hugging Face</a>. The raw dataset with additional information and all metadata for the extracted anchor texts (roughly 100GB) is available on <a href="https://huggingface.co/datasets/webis/ms-marco-anchor-text/tree/main/ms-marco-v1/anchor-text">Hugging Face</a> and <a href="https://files.webis.de/data-in-progress/ecir22-anchor-text/anchor-text-samples/">files.webis.de</a>.</p> <p>The details of the construction of the Webis MS MARCO Anchor Text 2022 dataset are described in the <a href="https://webis.de/publications.html#froebe_2022a">associated paper</a>. If you use this dataset, please cite<br> <code>@InProceedings{froebe:2022a,<br> &nbsp; address =&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; {Berlin Heidelberg New York},<br> &nbsp; author =&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; {Maik Fr{\&quot;o}be and Sebastian G{\&quot;u}nther and Maximilian Probst and Martin Potthast and Matthias Hagen},<br> &nbsp; booktitle =&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; {Advances in Information Retrieval. 44th European Conference on IR Research (ECIR 2022)},<br> &nbsp; editor =&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; {Matthias Hagen and Suzan Verberne and Craig Macdonald and Christin Seifert and Krisztian Balog and Kjetil N{\o}rv\r{a}g and Vinay Setty},<br> &nbsp; month =&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; apr,<br> &nbsp; publisher =&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; {Springer},<br> &nbsp; series =&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; {Lecture Notes in Computer Science},<br> &nbsp; site =&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; {Stavanger, Norway},<br> &nbsp; title =&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; {{The Power of Anchor Text in the Neural Retrieval Era}},<br> &nbsp; year =&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 2022<br> }</code></p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Greek Text to Trajectories Sign Language Dataset

<p>Entails the 2D human pose trajectories of Greek Elementary Sign Language Dataset&nbsp;and Greek News&nbsp;Sign Language Dataset (31681 examples).</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Modulares Erzählen: Annotierte Texte der Sieben weisen Meister

<p>Die einzelnen XML-Dateien enthalten den Text verschiedener deutschsprachiger&nbsp;Versionen und Fassungen der sp&auml;tmittelalterlichen&nbsp;Erz&auml;hltradition <em>Die sieben weisen Meister.&nbsp;</em>Sie bilden die Grundlage der textstatistischen Untersuchungen in&nbsp;Nico Kunkel: Modulares Erz&auml;hlen. Serialit&auml;t und Mouvance in der Erz&auml;hltradition der <em>Sieben weisen Meister</em>. Berlin/Boston (in Vorbereitung). Die Texte, die auf Editionen und Arbeitstranskriptionen zur&uuml;ckgehen,&nbsp;sind nach narratologischen Kriterien in einzelne Erz&auml;hleinheiten (= Erz&auml;hlmodule) eingeteilt.&nbsp;Einige Dateien (<em>Allegatio</em>, Gie&szlig;ener/Br&uuml;nner/Colmarer&nbsp;Fs.) liegen in Form von abgeleiteten Textformaten vor, da die Texte auf Editionen j&uuml;ngeren Datums (1997/2001/2008/2008) zur&uuml;ckgehen. Sie enthalten jeweils nur das erste und letzte Wort eines Erz&auml;hlmoduls, w&auml;hrend alle anderen W&ouml;rter durch Platzhalter (&quot;blank&quot;) ersetzt wurden. Auf diese Weise lassen sich&nbsp;statistische Abfragen (Modull&auml;nge/-sequenz) weiter umsetzen, w&auml;hrend ein Genuss des literarischen Werks nicht mehr m&ouml;glich ist.&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>Textgrundlagen:</strong></p> <ul> <li><em>Allegatio</em> = Text nach Steinmetz, Ralf-Henning (1997). &bdquo;Der &sbquo;Libellus muliebri nequitia plenus&lsquo;. Eine ungedruckte lateinische Version der &#39;Sieben weisen Meister&#39; und ihre deutsche &Uuml;bersetzung aus dem 15. Jahrhundert&ldquo;. In: <em>Zeitschrift f&uuml;r deutsches Altertum und deutsche Literatur</em> 126. 297-446 (abgeleitetes Textformat).</li> <li>anV = Text nach &bdquo;Von den sieben meistern&ldquo;. In: <em>Altdeutsche Gedichte.</em> Hg. von Adalbert Keller T&uuml;bingen 1846, 15-241.</li> <li>Aventewr = Arbeitstranskription von Heidelberg, Universit&auml;tsbibl., cpg 101, 29r-39r.</li> <li>Br&uuml;nner Fs = Text nach<em> Sieben weise Meister. Eine bairische und eine els&auml;ssische Fassung der &bdquo;Historia septem sapientum&ldquo;.</em> Hg. von Detlef Roth. Berlin 2008 (abgeleitetes Textformat).</li> <li>B&uuml;hnenfs = Arbeitstranskription von Wild Sebastian: <em>Schoener Comedien vnd Tragedien zwoelff: Au&szlig; heiliger goettlicher schrifft vnd auch au&szlig; etlichen historien gezogen [&hellip;] Auffs new in Truck verfertigt durch Sebastian Wilden</em>. Augsburg 1566.</li> <li>Colmarer Fs = Text nach&nbsp;<em> Sieben weise Meister. Eine bairische und eine els&auml;ssische Fassung der &bdquo;Historia septem sapientum&ldquo;.</em> Hg. von Detlef Roth. Berlin 2008 (abgeleitetes Textformat).)</li> <li><em>DL</em> = Text nach B&uuml;hel, Hans von. <em>Dyocletianus Leben</em>. Hg. von Adalbert Keller. Quedlinburg/Leipzig 1846.</li> <li>Donaueschinger Fs =&nbsp; Arbeitstranskription von Karlsruhe, Landesbibl., Cod. Donaueschingen 145, 5ra-58ra.</li> <li>Gie&szlig;ener Fs = Text nach <em>Die Historia von den sieben weisen Meistern und dem Kaiser Diocletianus</em>. Hg. von Ralf-Henning Steinmetz. T&uuml;bingen 2001 (abgeleitetes Textformat).</li> <li>Heidelberger Fs = Arbeitstranskription von Universit&auml;tsbibl., Cpg 149, 1r-108r.</li> <li><em>Hystorij</em> = Transkription von <em>Die Hystorij von Diocleciano. In Abbildungen aus dem Codex 407 des Wiener Schottenstifts</em>. Hg. Ralf-Henning Steinmetz. G&ouml;ppingen 1999.</li> <li>Vulgatfs = Arbeitstranskription von <em>Die Sieben weisen Meister</em>. Hg. von G&uuml;nter Schmitz. Hildesheim 1974.</li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo44/100

A Large-scale Dataset of (Open Source) License Text Variants

<p>We introduce a large-scale dataset of the complete texts of free/open source software (FOSS) license variants. To assemble it we have collected from the Software Heritage archive&mdash;the largest publicly available archive of FOSS source code with accompanying development history&mdash;all versions of files whose names are commonly used to convey licensing terms to software users and developers.<br> The dataset consists of 6.5 million unique license files that can be used to conduct empirical studies on open source licensing, training of automated license classifiers, natural language processing (NLP) analyses of legal texts, as well as historical and phylogenetic studies on FOSS licensing.<br> Additional metadata about shipped license files are also provided, making the dataset ready to use in various contexts; they include: file length measures, detected MIME type, detected SPDX license (using ScanCode), example origin (e.g., GitHub repository), oldest public commit in which the license appeared.<br> The dataset is released as open data as an archive file containing all deduplicated license blobs, plus several portable CSV files for metadata, referencing blobs via cryptographic checksums.</p> <p>For more details see the included&nbsp;README file and companion paper:</p> <ul> <li>Stefano Zacchiroli.&nbsp;<a href="https://doi.org/10.1145/3524842.3528491"><em>A Large-scale Dataset of (Open Source) License Text Variants</em></a>. In proceedings of the&nbsp;<a href="https://conf.researchr.org/home/msr-2022">2022 Mining Software Repositories Conference (MSR 2022)</a>. 23-24 May 2022 Pittsburgh, Pennsylvania, United States. ACM 2022.</li> </ul> <p>If you use this dataset for research purposes, please acknowledge its use by citing the above paper.</p> <ul> </ul>

opencc-by-4.0Mar 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record