Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,305

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,305 results for “text”

Learn how ShareScore rates datasets ↗
zenodo52/100

Invasion Biology WikiProject Scientific Papers: Text Data Mining and LLM-based Information Extraction of Species, Locations, Habitats, and Ecosystems

<p>This dataset contains the abstract and full-text for publication DOIs from the Invasion Biology WikiProject (DOI:&nbsp;<a href="https://www.doi.org/10.5281/zenodo.12518036">10.5281/zenodo.12518036</a>). The data was retrieved using the <a href="https://ask.orkg.org/">ask.orkg.org</a> <a href="https://api.ask.orkg.org/docs#tag/Semantic-Neural-Search/operation/explore_documents_index_explore_get">API</a>. For the <a href="https://github.com/jd-coderepos/invasion-biology-IE/blob/main/scripts/ask-doi-list-fulltext-search.py">script</a> used to obtain the data, refer to the accompanying GitHub repository: <a href="https://github.com/jd-coderepos/invasion-biology-IE/" target="_blank" rel="noopener">https://github.com/jd-coderepos/invasion-biology-IE/</a>.</p> <p>The resulting CSV file includes the following fields: <code>"ASK ID"</code>, <code>"DOI"</code>, <code>"Title"</code>, <code>"Abstract"</code>, and <code>"Full-text"</code>.</p> <p>Of the 49,438 queried DOIs, the ASK database provided:</p> <ul> <li><strong>Total DOIs processed:</strong> 12,636</li> <li><strong>DOIs with neither abstract nor full-text:</strong> 36 (abstract token count was less than 10)</li> <li><strong>DOIs with abstracts but no full-text:</strong> 12,636</li> <li><strong>DOIs with both abstract and full-text:</strong> 2,834</li> </ul> <p>The second part of the dataset contains structured information extracted from the publications using the GPT-4o Large Language Model. This structured data is included in the zipped folder <code>structured-publications.zip</code>.</p> <p>The accompanying GitHub repository provides access to the code and scripts used at various stages of the information extraction (IE) process.</p> <p><strong>Theme of the Study:</strong><br>"Mining for Species, Locations, Habitats, and Ecosystems from Scientific Papers in Invasion Biology: A Large-Scale Exploratory Study with Large Language Models."</p>

opencc-by-4.0Oct 2024View details →
zenodo52/100

WageIndicator Collective Agreements Database Dataset with Full Texts and Selected Clauses

<p>Since 2012, the&nbsp;<a href="https://wageindicator.org/">WageIndicator Foundation</a>&nbsp;has maintained a&nbsp;<a href="https://wageindicator.org/cbadatabase">Collective Agreements Database</a>, where the texts of 1600 collective agreements (CBAs) from 61 countries and in 27 languages have been uploaded, coded and annotated. This database is a unique example at global level: collective agreements are documents containing conditions of employment that result from negotiations between independent unions and employers, and their content is often surrounded by an atmosphere of secrecy. Under the&nbsp;<a href="https://sshopencloud.eu/">SSHOC project</a>&nbsp;and with the support of the&nbsp;<a href="https://www.clarin.eu/">CLARIN Research Infrastructure</a>, the agreements have been manually and automatically annotated on several levels: for each agreement, the team answers a series of questions and selects the appropriate piece of text (clause) for each.&nbsp;</p> <p>One of the results of the collective agreements&#39; annotation process is the&nbsp;dataset which is available here and includes all the clauses selected for each variable (WageIndicator_CBADatabase_Selected_Clauses). The full collective agreements&#39; texts are stored in another dataset, also available here (WageIndicator_CBADatabase_Full_Texts_211019). A codebook is also included (210125-wageindicator-cba-codebook.pdf).</p>

opencc-by-4.0Dec 2020View details →
zenodo52/100

Examining the Capacity of Text Mining and Software Metrics in Vulnerability Prediction [dataset]

<p>This dataset contains the extension of a publicly available dataset that was published initially by Ferenc et al. in their paper:</p> <p><em>&ldquo;Ferenc, R.; Hegedus, P.; Gyimesi, P.; Antal, G.; B&aacute;n, D.; Gyim&oacute;thy, T. Challenging machine learning algorithms in predicting vulnerable javascript functions. 2019 IEEE/ACM 7th InternationalWorkshop on Realizing Artificial Intelligence Synergies in Software Engineering (RAISE). IEEE, 2019, pp. 8&ndash;14.&rdquo;</em></p> <p>The dataset contained software metrics for source code functions written in JavaScript (JS) programming language. Each function was labeled as vulnerable or clean. The authors gathered vulnerabilities from publicly available vulnerability databases.</p> <p>In our paper entitled: &ldquo;<strong>Examining the Capacity of Text Mining and Software Metrics in Vulnerability Prediction</strong>&rdquo; and cited as:</p> <p><em>&ldquo;Kalouptsoglou I, Siavvas M, Kehagias D, Chatzigeorgiou A, Ampatzoglou A. Examining the Capacity of Text Mining and Software Metrics in Vulnerability Prediction. Entropy. 2022; 24(5):651. <a href="https://doi.org/10.3390/e24050651">https://doi.org/10.3390/e24050651</a>&rdquo;</em></p> <p>, we presented an extended version of the dataset by extracting textual features for the labeled JS functions. In particular, we got the dataset provided by Ferenc et al. in CSV format and then we gathered all the GitHub URLs of the dataset&#39;s functions (i.e., methods). Using these URLs, we collected the source code of the corresponding JS files from GitHub. Subsequently, by utilizing the start and end line information for every function, we cut off the code of the functions. Each function was then tokenized to construct a list of tokens per function.</p> <p>To extract text features, we used a text mining technique called sequences of tokens. As a result, we created a repository with all methods&#39; source code, the token sequences of each method, and their labels. To boost the generalizability of type-specific tokens, all comments were eliminated, as well as all integers and strings, which were replaced with two unique IDs.</p> <p>The dataset contains 12,106 JavaScript functions, from which 1,493 are considered vulnerable.</p> <p>This dataset was created and utilized during the Vulnerability Prediction Task of the Horizon2020 IoTAC Project&nbsp;as training and evaluation data for the construction of vulnerability prediction models. The dataset is provided in the csv format. Each row of the csv file has the following parts:</p> <ul> <li>Label: Flag with values &lsquo;1&rsquo; for vulnerable and &lsquo;0&rsquo; for non-vulnerable methods</li> <li>Name: The name of the JavaScript method</li> <li>Longname: The longname of the JavaScript method</li> <li>Path: The path of the file of the method in the repository</li> <li>Full_repo_path: The GitHub URL of the file of the method</li> <li>TokenX: Each next row corresponds to each token included in the method</li> </ul>

opencc-by-4.0Sep 2023View details →
edi52/100

Como and Phalen lake ChatBot water quality and recreation text survey, 2022 and 2023

We collected data from visitors to two urban lakes in Saint Paul, Minnesota, using a conversational chatbot to assess visitor perception of current lake water quality, trends in water quality over time, and other questions relevant to park managers. Data were collected at Como Lake in 2022 and 2023, and at Lake Phalen in 2023. Signs were installed at three locations around each lake with high pedestrian traffic. Each sign had a hook question (“How many watercraft are on the lake right now? Text the number to XXX-XXX-XXXX”). Visitors who responded to this question initiated a series of optional follow-up questions, using a conversational chatbot run by software that automates the sending and receiving of text messages. Survey questions included asking respondents about their primary purpose for visiting the lake today, how often they have visited the lake in the past 12 months, and whether they perceive that water quality in the lake is improving, getting worse, or remaining about the same. Respondents were asked to provide their ZIP code, used to estimate distance traveled to the lake. Respondents also had the opportunity to opt-in to future data collection via phone or text. An AI language model was then used to process the information and parse and synthesize responses. Unique anonymous identifiers were used to key survey response data.

openCC (other)Jun 2024View details →
zenodo48/100

Corpus Creation for Sentiment Analysis in Code-Mixed Tamil-English Text

<p>Understanding the sentiment of a comment from a video or an image is an essential task in many applications. Sentiment analysis of a text can be useful for various decision-making processes. One such application is to analyse the popular sentiments of videos on social media based on viewer comments. However, comments from social media do not follow strict rules of grammar, and they contain mixing of more than one language, often written in non-native scripts. Non-availability of annotated code-mixed data for a low-resourced language like Tamil also adds difficulty to this problem. To overcome this, we created a gold standard Tamil-English code-switched, sentiment-annotated corpus containing 15,744 comment posts from YouTube. In this paper, we describe the process of creating the corpus and assigning polarities. We present inter-annotator agreement and show the results of sentiment analysis trained on this corpus as a benchmark.</p>

opencc-by-4.0May 2020View details →
zenodo48/100

Fedora and Debian software package dependency networks along with description text associated with nodes

<p>Fedora (version 28) and Debian (version 9.5) software package dependency networks along with description text associated with nodes. Also includes learned vectors by using PCTADW-* as in &quot;Kexuan Sun, Shudan Zhong, and Hong Xu. 2020. Learning Embeddings of Directed Networks with Text-Associated Nodes---with Application in Software Package Dependency Networks. 2020 BigGraphs Workshop at IEEE BigData 2020.&quot;</p>

openmit-licenseSep 2018View details →
zenodo48/100

Relation Extraction Dataset for Dutch Biographical Texts

<p>A manually annotated dataset with relations relevant for biographical texts. The texts are in Dutch and are originally available in the Biographical portal of the Netherlands (http://www.biografischportaal.nl/)</p>

opencc-by-4.0Dec 2023View details →
zenodo48/100

Named Entity Recognition Dataset for Dutch Biographical Texts

<p>A dataset for Named Entity Recognition for Dutch biographies. The original data is available in the Biographical portal of the Netherlands (http://www.biografischportaal.nl/). The annotations are for 6 types of entities: PERSON, LOCATION, ORGANIZATION, DATE, ARTWORK, MISC. Additionally, the CoNLL formatted files were manually checked for tokenization and sentence splitting.</p>

opencc-by-4.0Dec 2023View details →
zenodo48/100

Multi-faceted analyses of Poland's Bronze and Early Iron Age hoards: Fig.1. Location of hoards mentioned in the text: white dots represent locations of hoards examined in the Biography of Hoards project; black dots represent locations of hoards examined in other multi-faceted projects

<p>The set contains a figure, with data, on the location of the hoards included (described in the related paper).<br><br>The paper and data were prepared as part of a project funded by the National Science Centre, Poland: <em>A Biography of Late Bronze and Early Iron Ages Hoards. A Multi-Faceted Analysis of Metal Objects Related to Monumental Constructions in Poland</em> (UMO-2021/41/B/HS3/00038)</p>

opencc-zeroSep 2023View details →
zenodo48/100

Automated Literature Screening for Systematic Reviews: Dataset for Evaluation Against Human Title and Abstract and Full-Text Screening Decisions

<p>This Zenodo entry contains the supplementary material associated with the manuscript titled&nbsp;<em>Automated Literature Screening for Systematic Reviews: A 5-Tier Prompting Approach Meeting Cochrane&rsquo;s Sensitivity Requirement of Greater Than 0.99.</em> The paper will be presented at <a href="https://dbis.rwth-aachen.de/LLMs4MI2024/">LLMsMI 2024</a> in November 2024.</p> <p>A script is provided for replicating the executed experiments, along with a comprehensive evaluation file that reports all the experiment results. Provided data files represent an extension to the original datasets as provided by [1]. For associated systematic review manuscripts and eligibility criteria, please refer to [1] as well.&nbsp;</p> <p>[1] Guo, Eddie; Gupta, Mehul; Deng, Jiawen; Park, Ye-Jean; Paget, Mike; Naugler, Christopher (2023). "Automated Paper Screening for Clinical Reviews Using Large Language Models."&nbsp;<em>Mendeley Data</em>, V1, doi: 10.17632/np79tmhkh5.1. Accessed from: <a href="https://data.mendeley.com/datasets/np79tmhkh5/1" target="_new" rel="noopener">https://data.mendeley.com/datasets/np79tmhkh5/1</a>.</p>

opencc-by-4.0Dec 2023View details →
zenodo48/100

Supplementary Material for Embodied Emotions in Ancient Neo-Assyrian Texts Revealed by Bodily Mapping of Emotional Semantics

<p>This dataset accompanies the article "Embodied Emotions in Ancient Neo-Assyrian Texts Revealed by Bodily Mapping of Emotional Semantics" (Lahnakoski &amp; Bennett et al., submitted).&nbsp;</p> <p>It includes the Neo-Assyrian text corpus that is the basis for the word embeddings, a list of the Akkadian emotion and body words of interest for this study, and the scripts, toolboxes, and data used to generate the heat maps of the body.</p> <p>There is an additional folder containing the high resolution figures included in the article.</p> <p>A detailed ReadMe (README.txt) provides an overview of the folders.</p>

opencc-by-4.0May 2024View details →
zenodo48/100

Data and analysis script for "The (non)effect of personalization in climate texts on credibility of climate scientists: A case study on sustainable travel"

<p>Dataset and analysis script for the article "<strong>The (non)effect of personalization in climate texts on credibility of climate scientists</strong><strong>: A case study on sustainable travel</strong>", under review at Geoscience Communication (https://doi.org/10.5194/egusphere-2024-543)</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

A Monologue Narrative Text of the Itoman Dialect of Okinawan: Yukkanuhii in My Childhood

<p>This dataset provides a monologue narrative text of the Itoman dialect of the Okinawan language spoken by a male speaker in his 70s. The speaker recounts his childhood memories of the event Yukkanuhii, which is held on the fourth day of the fifth lunar month. The event involves races in small boats called Haaree. The dataset includes an audio file (.wav) and an annotated xml file (.eaf). Japanese translations, morphological analyses, and interlinear glosses are provided in an .eaf file.</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Data set for (binary) text classification, involving spoken utterances and written text

<p>This data set contains sentences belonging to either of two classes: Transcripts of spoken<br> (informal) text (Class 0), and written, formal text (Class 1). Sentences in Class 0 were<br> obtained from publicly available transcripts of radio shows (e.g. NPR),<br> whereas Sentences in Class 1 were obtained from Wikipedia.</p> <p>The data set is divided into&nbsp;three subsets: Training, validation, and test (specified&nbsp;by the file names).<br> Each set contains a large number of sentences, belonging to either of the two classes:</p> <p>In total, there are 13,640,458 sentences, of which 6,374,487 in Class 0 and 7,265,971.<br> The training set contains 9.743,188 sentences (of which 4,553,205 in Class 0 and 5,189,983 in Class 1),&nbsp;<br> the validation set contains 1,948,639 sentences (of which 910,641 in Class 0 and 1,037,998 in Class 1), and the&nbsp;<br> test set contains 1,948,631 sentences (of which 910,641 in Class0 and 1,037,990 in Class1).&nbsp;</p> <p>The data sets are in plain text format. Every row contains (i) the class label (0 or 1) and<br> (ii) the text of the sentence, separated from the class label by a tab character.</p> <p>Note that the&nbsp;sentences contain 5 tokens or more (including punctuation marks).&nbsp;&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo48/100

Archi text corpus

<p>Archi belongs to the Lezgic group of the Nakh-Daghestanian (North-East Caucasian) languages, being quite loosely related to the rest of the group. It has been long time surrounded by non-Lezgic languages and therefore has kept and/or acquired a number of peculiar features.</p> <p>Here is presented a sample of texts collected in the village of Archi in 2006 and 2007. In total, over 50 texts of various genres have been recorded, including stories, conversations, tales, legends and songs. Most of them were recorded in both video and audio.</p> <p>Two kinds of texts were recorded. First, some 30 previously published (1977) texts were re-recorded in video and audio, read by one of three speakers. (The original recordings do not exist anymore). Second, new texts have been collected, mostly dialogues and stories. Texts available in this version were all originally published in [Kibrik et al. 1977].</p> <p><em>Кибрик А. Е., Кодзасов С. В., Оловянникова И. П., Самедов Д. С.</em> Арчинский язык. Тексты и словари.&nbsp;&mdash; М.: МГУ, 1977.<br> [Kibrik, Aleksandr E.; Kodzasov, S. V.; Olovjannikova, I. P. &amp; Samedov, D. S. (1977). <em>Arčinskij jazyk. Teksiy i slovari</em>. Moscow: Izdatel&#39;stvo moskovskogo universiteta.]</p> <p>The project was generously supported by NSF grant #0553546 &laquo;Five languages of Eurasia&raquo; (PI under the Documenting Endangered Languages Program, and by RFBR grants №&nbsp;05-06-80351 &laquo;Minority languages and cultures: On the verge of extinction&raquo; and №&nbsp;08-06-00345 &laquo;Multimedia corpora for endangered languages&raquo;.</p>

opencc-by-4.0Dec 2007View details →
zenodo48/100

OpenITI: a Machine-Readable Corpus of Islamicate Texts

<p><strong>Co-PIs</strong>: Matthew Thomas Miller (University of Maryland, College Park), Maxim G. Romanov (University of Hamburg), Sarah Bowen Savant (Aga Khan University&mdash;ISMC, London).</p> <p><em>Open Islamicate Texts Initiative</em> (<strong>OpenITI</strong>, see <a href="https://openiti.org/">https://openiti.org/</a>)&nbsp;is a multi-institutional effort to construct the first machine-actionable scholarly corpus of premodern Islamicate texts. Led by researchers at the Aga Khan University, Institute for the Study of Muslim Civilisations (AKU-ISMC), University of Hamburg (UH), and the Roshan Institute for Persian Studies at the University of Maryland (College Park) and an interdisciplinary advisory board of leading digital humanists and Islamic, Persian, and Arabic studies scholars, <strong>OpenITI</strong> aims to provide the essential textual infrastructure in Arabic, Persian and other Islamicate languages for new forms of textual analysis and digital scholarship. In the process, OpenITI will enable new synergies between Digital Humanities and the inter-related Islamicate fields of Islamic, Persian, and Arabic Studies. In addition to support from the researchers&rsquo; home institutions, it is supported by funding from the <a href="https://erc.europa.eu/">European Research Council</a> under the European Union&rsquo;s Horizon 2020 research and innovation programme, awarded to the <a href="http://kitab-project.org/">KITAB</a> project (Grant Agreement No. 772989, PI Sarah Bowen Savant) and the <a href="https://www.qnl.qa/en">Qatar National Library</a>.</p> <p>Currently, <strong>OpenITI</strong> contains almost exclusively Arabic texts, which were first assembled into a corpus within the <strong>OpenArabic</strong> project, developed first at Tufts University (at <em>The Perseus Project</em>, 2013&ndash;2015) and then at Leipzig University (at the Alexander von Humboldt Chair for Digital Humanities, 2015&ndash;2017)&mdash;in both cases with the support and under the patronage of Prof. Gregory Crane. The much more limited number of Persian texts were compiled during 2015&ndash;2016 in the Persian Digital Library (PDL) pilot (see <a href="https://persdigumd.github.io/PDL/">Persian Digital Library by PersDigUMD</a>) at Roshan Institute for Persian Studies at the University of Maryland. These texts have not been made fully compatible with OpenITI mARkdown yet and will be made fully available in next releases.</p> <p>This release contains all digital versions of the same text that are available in the OpenITI corpus . <strong>We also release a <a href="https://doi.org/10.5281/zenodo.7764025">'primary' version of the corpus</a></strong> that contains a single digital version for each text in the corpus that is marked as 'PRI' in the corpus metadata and may be more convenient for some use cases.</p> <p><strong>Note on Release Numbering</strong>: Version <strong>2019.1.1</strong>&mdash;where <strong>2019</strong> is the year of the release, the first dotted number&mdash;<strong>.1</strong>&mdash;is the ordinal release number in 2019, and the second dotted number&mdash;<strong>.1</strong>&mdash;is the overall release number; the first dotted number will reset every year, while the second one will continue on increasing.</p> <p>For more details: <a href="https://github.com/OpenITI/RELEASE">https://github.com/OpenITI/RELEASE</a></p> <p><strong>Note: </strong>In case of any issues with unzipping the files on Windows using built-in utilities, please use free softwares, such as&nbsp;WinRAR and 7zip.</p> <p>&nbsp;</p>

opencc-by-nc-sa-4.0Oct 2023View details →
edi48/100

Supplemental materials of the Castaño-Sánchez et. al. (2023) article (Agricultural Systems) containing the IFSM model input parameters not included in the main text, and the Criollo ranches survey form

CONTEXT: The southwestern United States is experiencing an increasingly warmer and drier climate that is affecting cattle production systems of the region. Adaptation strategies are needed that will not compromise environmental quality or profitability. Options include the use of desert-adapted beef cattle biotypes, such as Rarámuri Criollo cattle, and crossbreds of Criollo with more traditional British breeds. Currently, most calves raised in the Southwest are grain finished, often with irrigated crops produced in the hydrologically-threatened Ogallala Aquifer region. A viable alternative may be grass finishing with the rainfed forage of the arid and semi-arid rangeland of the Southwest or in the temperate grasslands of the Northern Plains. OBJECTIVE: Compare the environmental impacts and production costs of grain-finishing in Texas and grass-finishing in the Northern plains and the Southwest with traditional Angus cattle vs. Criollo and Criollo x Angus cattle. METHODS: Nine supply chain strategies were simulated using the Integrated Farm System Model to compare farm-gate life cycle intensities of greenhouse gas emissions (carbon footprint), fossil energy footprint, nitrogen footprint, blue water footprint and production costs using representative (appropriate soils, climate, and management) ranch and feedlot operations. RESULTS AND CONCLUSIONS: For both finishing options (grass, grain), Criollo x Angus cattle had the best environmental (3%-27% lower), and production cost (4-23% lower) outcomes followed by pure Criollo and then Angus cattle. Crossbred production combined the lower feed supplementation requirements of Criollo cows with heavier final carcasses of offspring from Angus genetics. Crossbred cattle with grass finishing in the Southwest or Northern Plains outperformed on most environmental variables as well as production costs, mostly due to reduced external input requirements (primarily feed). A downside for grass-finished crossbreds was greater carbon fo

openCC (other)Aug 2023View details →
zenodo44/100

Data_text section_ 11βHSD1_11β-Hydroxysteroid dehydrogenases control access of 7β,27-dihydroxycholesterol to retinoid-related orphan receptor γ

<p>Data from kinetic characterization (Km and vmaxapp) described in text section 3.1 of 11&beta;-Hydroxysteroid dehydrogenases control access of 7&beta;,27-dihydroxycholesterol to retinoid-related orphan receptor &gamma;</p> <p>Dataset (doi:10.1194/jlr.M092908) contains values from kinetic characterization (Km and vmaxapp) described in text section&nbsp;(Kinetic values_11&beta;HSD1.PNG)&nbsp;corresponding to raw data obtained from LC-MS/MS analysis provided as three files in CSV format (31003A-179400_DATE_KB_27Oxysterol_4_6_1-3). All further experiment related information and subsequent data analysis provided as two meta-data-files (31003A-179400_DATE_KB_27Oxysterol_4_6_M_1-2) as TXT format and PDF format.</p>

opencc-by-4.0Jul 2019View details →
zenodo44/100

Duhumbi Personal Narratives - Transcribed, parsed, glossed, translated text files

<p>This data set contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the following recordings that belong to the following publication:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <ul> <li>[CHUK230512D1A] / CMT / The story of the former CM&rsquo;s death</li> <li>[CHUK230512C1A] / LHT / The history of Laphek village</li> <li>[CHUK230512B1] / THT / Hunting takin</li> <li>[CHUK260413A3A]/ ACK / Alcohol consumption</li> <li>[CHUK131014] / DTPK / Chasing the demons</li> </ul> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file &ldquo;Settings&rdquo;, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>For the metadata of the sound files in this data set, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for&nbsp;commercial purposes&nbsp;<strong><em>of any kind</em></strong><em>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration &amp; payment for access, or sites that rely on advertisement (including YouTube)&nbsp;</em>is&nbsp;<strong>not</strong>&nbsp;permitted without&nbsp;<strong>specific written consent</strong>&nbsp;from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p>

opencc-by-4.0Aug 2018View details →
zenodo44/100

Duhumbi Religious Texts and Song - Transcribed, parsed, glossed, translated text files

<p>This data set contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the following recordings that belong to the following publication:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <ul> <li>[CHUK110413A2A] / RELJ / Buddhist admonition</li> <li>&nbsp;[CHUK221212D2A] / JIK / Bonpo prediction text</li> <li>&nbsp;[CHUK260413A1] / MSK / Impromptu song</li> </ul> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file &ldquo;Settings&rdquo;, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>For the metadata of the sound files in this data set, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for&nbsp;commercial purposes&nbsp;<strong><em>of any kind</em></strong><em>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration &amp; payment for access, or sites that rely on advertisement (including YouTube)&nbsp;</em>is&nbsp;<strong>not</strong>&nbsp;permitted without&nbsp;<strong>specific written consent</strong>&nbsp;from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p>

opencc-by-4.0Aug 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record