Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6
datasets available to search
ShareScore release 0.9.0
Dataset results
6 results for “Narrative text”
A Monologue Narrative Text of the Itoman Dialect of Okinawan: Yukkanuhii in My Childhood
<p>This dataset provides a monologue narrative text of the Itoman dialect of the Okinawan language spoken by a male speaker in his 70s. The speaker recounts his childhood memories of the event Yukkanuhii, which is held on the fourth day of the fifth lunar month. The event involves races in small boats called Haaree. The dataset includes an audio file (.wav) and an annotated xml file (.eaf). Japanese translations, morphological analyses, and interlinear glosses are provided in an .eaf file.</p>
Duhumbi Personal Narratives - Transcribed, parsed, glossed, translated text files
<p>This data set contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the following recordings that belong to the following publication:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <ul> <li>[CHUK230512D1A] / CMT / The story of the former CM’s death</li> <li>[CHUK230512C1A] / LHT / The history of Laphek village</li> <li>[CHUK230512B1] / THT / Hunting takin</li> <li>[CHUK260413A3A]/ ACK / Alcohol consumption</li> <li>[CHUK131014] / DTPK / Chasing the demons</li> </ul> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file “Settings”, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>For the metadata of the sound files in this data set, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <strong><em>of any kind</em></strong><em>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p>
A Monologue Narrative Text of the Itoman Dialect of Okinawan: How to Worship in Yukkanuhii
<p>This paper provides a monologue narrative text of the Itoman dialect of the Okinawan language spoken by a male speaker in his 70s. He describes the way worship is conducted in Yukkanuhi. Japanese translations, morphological analyses, and interlinear glosses are provided. The dataset includes an audio file (.wav) and an annotated xml file (.eaf). Japanese translations, morphological analyses, and interlinear glosses are provided in an .eaf file.</p>
Graphing/Markup Data: TNA SC8 Petitionary Texts and Community Narrative
<p>Cleaned CSV files of data samples taken from the British National Archives database; TNA petition XML/TEI data model for the purpose of digitizing and cataloguing documents from the TNA repository.</p> <p><em>The final iteration of this project will be an interactive web-based visual representation of medieval communication networks using petitionary texts. I am currently researching methods to reposition these 13th century petitions as a form of communication while viewing them as an early form of social media. Typically these petitions were either grievances or requests made to the king seeking favor or assistance. By analyzing the data that is found in each of these petitions both through a database found at the National Archives in the United Kingdom, I can look at the smaller storylines within the larger system. I am also studying the objects themselves including the handwritten information which is how I plan on transcribing and editing the texts using TEI markup. The goal is to make this model transferable for use with other collections of historical texts and for digital humanities projects where the collections, including the digitized images, the data, and the relationships established can be curated for educational use.</em></p>
NapSS: Paragraph-level Medical Text Simplification via Narrative Prompting and Sentence-matching Summarization
<p>Accessing medical literature is difficult for laypeople as the content is written for specialists and contains medical jargon. Automated text simplification methods offer a potential means to address this issue. In this work, we propose a summarize-then-simplify two-stage strategy, which we call NapSS, identifying the relevant content to simplify while ensuring that the original narrative flow is preserved. In this approach, we first generate reference summaries via sentence matching between the original and the simplified abstracts. These summaries are then used to train an extractive summarizer, learning the most relevant content to be simplified. Then, to ensure the narrative consistency of the simplified text, we synthesize auxiliary narrative prompts combining key phrases derived from the syntactical analyses of the original text. Our model achieves results significantly better than the seq2seq baseline on an English medical corpus, yielding 3%~4% absolute improvements in terms of lexical similarity, and providing a further 1.1% improvement of SARI score when combined with the baseline. We also highlight shortcomings of existing evaluation methods, and introduce new metrics that take into account both lexical and high-level semantic similarity. A human evaluation conducted on a random sample of the test set further establishes the effectiveness of the proposed approach.</p>
Evaluation of Text or Animation-based Narrative and Non-narrative Interventions to Enhance Cancer Literacy: a Pilot Study (CLARO)
ClinicalTrials.gov study NCT06026384. IPD Sharing: NO. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.