Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
78
datasets available to search
ShareScore release 0.9.0
Dataset results
78 results for “Glosses”
Würzburg Old Irish Glosses
<p>This dataset contains the digital text of the Würzburg Glosses on the Pauline Epistles, dated to about the 8th century, as they appear in <em>Thesaurus Palaeohibernicus Vol I</em><em> </em>(1901), along with other data relating to that collection. The data is collected in the format of a JSON document.</p> <p>The digital text and other data was created by Adrian Doyle in 2018 with funding from the National University of Ireland Galway through the Digital Arts and Humanities Scholarship, and with the kind permission of the Dublin Institute for Advanced Studies. This digital repository was created for the specific purposes of the Cardamom (Comparative Deep Models for Minority and Historical Languages) research project hosted by the Data Science Institute in the National University of Ireland Galway. The data available here has been collected in a format intended to meet the specific requirements of that project.</p> <p>Relevant selections of the Latin text of the Pauline Epistles are supplied as they appear in <em>Thesaurus Palaeohibernicus</em>, along with the Irish and Latin text of the glosses, and an English translation of glosses where one is available in <em>Thesaurus Palaeohibernicus</em>. Footnotes which appear in <em>Thesaurus Palaeohibernicus</em> relating to the Latin text of the Epistles, the glosses, or to gloss translations are also included. The Irish and Latin text of the glosses is supplied in three formats:</p> <p>1. Plain text (HTML tags used to italicise text as per <em>Thesaurus Palaeohibernicus</em>, no footnote markers included)</p> <p>2. Plain text with footnote markers (HTML tags used to italicise text and superscript footnote markers as per <em>Thesaurus Palaeohibernicus</em>)</p> <p>3. Fully tagged text (customised tag-set used to identify various features of the text including instances of code-switching, scribal contractions and abbreviations, text supplied by the editors, footnote markers, and more)</p> <p>Metadata collected in the JSON document relating to the glosses includes the relevant epistle, the manuscript folio, the gloss number, the scribal hand responsible for a given gloss, the relevant page in <em>Thesaurus Palaeohibernicus</em>, the biblical reference to the line of Latin text being glossed, the Latin lemma associated with a given gloss and the position of the first character of this lemma in the Latin text.</p>
Duhumbi Personal Narratives - Transcribed, parsed, glossed, translated text files
<p>This data set contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the following recordings that belong to the following publication:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <ul> <li>[CHUK230512D1A] / CMT / The story of the former CM’s death</li> <li>[CHUK230512C1A] / LHT / The history of Laphek village</li> <li>[CHUK230512B1] / THT / Hunting takin</li> <li>[CHUK260413A3A]/ ACK / Alcohol consumption</li> <li>[CHUK131014] / DTPK / Chasing the demons</li> </ul> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file “Settings”, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>For the metadata of the sound files in this data set, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <strong><em>of any kind</em></strong><em>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p>
Duhumbi Religious Texts and Song - Transcribed, parsed, glossed, translated text files
<p>This data set contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the following recordings that belong to the following publication:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <ul> <li>[CHUK110413A2A] / RELJ / Buddhist admonition</li> <li> [CHUK221212D2A] / JIK / Bonpo prediction text</li> <li> [CHUK260413A1] / MSK / Impromptu song</li> </ul> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file “Settings”, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>For the metadata of the sound files in this data set, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <strong><em>of any kind</em></strong><em>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p>
Duhumbi Procedural Texts - Transcribed, parsed, glossed, translated text files
<p>This data set contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the following recordings that belong to the following publication:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <ul> <li>[CHUK210512I1] / PHPT / Hunting for porcupine </li> <li>[CHUK230512A1A] / SBDC / Making fermented soybean</li> <li>[CHUK220413A1] / CTTT / Catching frogs </li> <li>[CHUK220413B1] / CWTT / Collecting beeswax </li> <li>[CHUK220413C1] / CHTT / Collecting hornets</li> <li>[CHUK240413A1] / SNAP / Collecting stinging nettle </li> <li>[CHUK240413B1] / LCYT / Leather craft </li> </ul> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file “Settings”, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>For the metadata of the sound files in this data set, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collector of the material. By downloading this material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit properly and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p>
Duhumbi Discussions - Transcribed, parsed, glossed, translated text files
<p>This dataset contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the following recordings that belong to the following publication:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <ul> <li>[CHUKxxxx13A6] / LEL / Local elections (not included on speaker's request)</li> <li>[CHUK300412J2] / LGT / Planning a trip to Lagam</li> <li>[CHUK260413A2A] / NNK / Nicknames</li> </ul> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file “Settings”, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>For the metadata of the sound files in this data set, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p>
Glossed Hittite Texts with German Translation for Machine Learning
<p>This dataset contains 7,099 processed Hittite texts from 143 CTH numbers, sourced with permission from the <a href="https://www.hethport.uni-wuerzburg.de/HPM/index.php" target="_blank" rel="noopener"><strong>Hethitologie Portal Mainz (HPM)</strong></a>, which is the comprehensive resource of Hittite texts and culture (modern Turkey, c. 1,650 - 1,200 BCE). These texts have been converted from XML format to a tabular structure for computational and linguistic analysis, as well as machine learning applications, while attempting to preserve the full complexity of the source material.</p> <p><strong>Key Features:</strong></p> <ul> <li><strong>txtid</strong>: Identifier for the specific text (e.g., "IBoT 1.30+"), referencing entries in the HPM Konkordanz (S. Košak, hethiter.net/: hetkonk (2.plus)).</li> <li><strong>lnr</strong>: Surface, column (if extant), and line number within the text.</li> <li><strong>cth_number</strong>: <em>Catalogue des textes hittites </em>(CTH) classification, organizing texts by genre and content (S. Košak – G.G.W. Müller – S. Görke – Ch.W. Steitler, hethiter.net/: CTH (2022-10-26)).</li> <li><strong>word</strong>: Orthographic form of the text as presented in the HPM, not in cuneiform but in a standard Latin script representation.</li> <li><strong>translit</strong>: Detailed transliteration, preserving all nuances such as diacritics, broken parts, and editorial markings to reflect the original condition on the text.</li> <li><strong>gloss</strong>: Linguistic glosses, including grammatical, morphological, and semantic information.</li> <li><strong>trans_de</strong>: German translation of the word or phrase.</li> </ul>
Vernacular Parallel Glosses in the Gloss-ViBe Corpus
<p>For this dataset I have collected all those instances in which there is an Old Irish gloss in the Vienna Bede that has a parallel gloss in at least one different language – Latin or Old Breton/Welsh – in one of the other three manuscripts recorded in the Gloss-ViBe corpus (https://gams.uni-graz.at/context:glossvibe). This dataset is used in <a href="https://doi.org/10.12688/openreseurope.16006.1">https://doi.org/10.12688/openreseurope.16006.1</a>.</p>
The glosses to the first book of the Etymologiae of Isidore of Seville: raw data
<p>This excel file contains the raw data behind the digital scholarly edition of the glosses to the first book of the <em>Etymologiae</em> of Isidore of Seville published at: <a href="https://db.innovatingknowledge.nl/edition">https://db.innovatingknowledge.nl/edition</a></p>
Gloss estimation of chocolate sprinkles with Hyperspectral Imaging
<p>Gloss is an important characteristic in the quality evaluation in chocolate production. However, standard glossing<br> measuring devices face several challenges when measuring food products, in particular those with curved surfaces and<br> small, such as chocolate sprinkles. Therefore, gloss evaluation of chocolate sprinkles is typically done by human visual<br> assessment. In this respect, hyperspectral imaging (HSI), combining spectroscopy and imaging, has gained attention as a<br> non-destructive and non-contact real-time detection tool for food quality analysis and control. This technique adds an<br> extra dimension to traditional machine vision techniques by providing images at a larger number of more narrow<br> wavebands. This can potentially increase the discrimination power.<br> <br> The main task of this dataset is to classify between 5 classes of chocolate production stages. The labels are indicated on the file names. The labels are: EXTRUDER, GLUCOSE, STAGE1, STAGE2, STAGE3.</p> <p> The .zip file contains two folders named:</p> <p>- envi_sprinkles_dataset_11_04_22 (Train Set)</p> <p>- envi_sprinkles_test_dataset_11_04_22 (Test Set)</p> <p>In envi_sprinkles_dataset_11_04_22 the file name convention is BATCH_CLASSNAME_INDEX.hdr</p> <p>In envi_sprinkles_test_dataset_11_04_22 the file name convention is CLASSNAME_INDEX.hdr</p> <p>Note that in the train set the batch information is present in the file name, if two samples belongs to the same batch means that they were produced at the same time.</p> <p>The samples format is ENVI. Consist of a header file and ENVI binary data file with file extensions <code>.hdr</code> and <code>.raw</code>, respectively. The function writes the wavelength and metadata information to the ENVI header file and the data cube containing the hyperspectral images to the ENVI binary data file.</p> <p>The image shape is [272x512x16] , [HEIGHT, WIDTH, BANDS], the pixel format is uint16.</p> <p> </p> <p> </p>
Dazzled by shine: gloss as an antipredator strategy in fast moving prey
<p>Previous studies on stationary prey have found mixed results for the role of gloss in predator avoidance – some have found that gloss can act as warning colouration or improve camouflage, whereas others detected no survival benefit. An alternative untested hypothesis is that gloss could provide protection in the form of dynamic dazzle. Fast-moving animals that are glossy produce flashes of light that increase in frequency at higher speeds, which could make it harder for predators to track and accurately locate prey. We tested this hypothesis by presenting praying mantids with glossy or matte targets moving at slow and fast speeds. Mantids were less likely to strike glossy targets, independently of speed. Additionally, we found that compared to matte targets, mantids were less likely to track glossy targets and more likely to hit the target with one rather than both raptorial arms, but only when targets were moving fast. These results support the hypothesis that gloss may have a function as an antipredator strategy by reducing the ability of predators to track and accurately target fast-moving prey. </p>
Dazzled by shine: gloss as an antipredator strategy in fast moving prey
Open the record for dataset details and reuse information.
Duhumbi Stories - Transcribed, parsed, glossed, translated text files
<p>This data set contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the following recordings that belong to the following publications:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <p>These Duhumbi stories can also be found in the 'Duhumbi Storybook' (Monpasang Publications, ISBN 978-90-818610-1-4) published in autumn 2018. A separate Zenodo DOI also contains all the sound files (http://doi.org/10.5281/zenodo.1400495).</p> <p>The following list contains the sound file names, the shortcut code for the sentence names and the title of the story in Duhumbi, English and Hindi.</p> <ul> <li>[CHUK260413A4A] / OMAK / Dangpu budunbakaq tsawa / The origin of mankind / मानव जाती की उत्पत्ति together with [CHUK260413A4A] / MOD / Bengkhannaq dontha / The meaning of dreams / सपनों का मतलब</li> <li>[CHUK290412A8A] / CHLN / Duhum chakpaqkho tam – 1 / Settlement history of Duhum – 1/ दुहुम के निवास का इतिहास – 1</li> <li>[CHUK230512E1B] / CHT / Duhum chakpaqkho tam – 2 / Settlement history of Duhum – 2 / दुहुम मे निवास का इतिहास – 2</li> <li>[CHUK240314A1] / SPZP / Shawa Pema Zomba / शावा पेमा जोम्बा</li> <li>[CHUK230512F1A; CHUK230512F2A; CHUK230512F3A; CHUK230512F4A; CHUK230512F5A; CHUK230512F6A; CHUK230512F7A] and [CHUK230512G1A; CHUK230512G2A; CHUK230512G3A] / KDZ1 – KDZ10 / Khandro Drowa Zangmu / खांडरों द्रोवा जांग्मो</li> <li>[CHUK110614A1; CHUK110614B1; CHUK110614C1; CHUK110614D1; CHUK110614E1; CHUK110614F1; CHUK110614G1; CHUK110614H1; CHUK110614I1] / LGG1 – LGG9 / Ling Gesar Gepuwaq namthar / The life of King Ling Gesar / लिंग गेसर गेपु की जीवनी</li> <li>[CHUK110413A1A] / TNZY / Tshongpon Norbu Zangpo dangngaq waq uda / Tshongpon Norbu Zangpo and his son / त्शोंग्पोन नोरबू जांगपो और उसका बेटा</li> <li>[CHUK070115A1] / BMDC / Shadong dangngaq gomchen phawang / The macaque and the bat / बंदर और चमगादर</li> <li>[CHUK070115B] / BUDC / Pempelingngaq tam / The butterfly effect / तितली का असर</li> <li>[CHUK211015E1] / RSTT / Samtu dangngaq grongthang / The squirrel and the rat / गिलहरी और चूहा</li> <li>[CHUK130115E] / FBDC / Shalaqbaknyi men shikhennaq ama / The mother that fed the children with ‘men’ / माँ जिसने बच्चों को ‘मेन’ खिलाया</li> <li>[CHUK300115C1] / MTDC / Mani tam / मनी ताम</li> <li> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file “Settings”, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>For the metadata of the sound files in this dataset, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <strong><em>of any kind</em></strong><em>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p> </li> </ul>
Travelling Annotations: Network Analysis as a Tool to Study Glossing Networks in Carolingian Europe
<p>While traditionally, the transmission of medieval texts is studied by means of stemmatics, certian types of textuality are well-known as being particularly resistent to traditional methods. This is also the case with annotations. While annotations can behave text-like, it is more often the case that each individual gloss must be treated as an autonomous entity. In different manuscripts different combinations of glosses are combined so that two manuscripts may contain a very different assembly of glosses and look dissimilar, while being closely related. In such cases, network analysis proves handy as a mean to reveal connection between manuscripts and trace the patterns of transmission of particular annotations, while opening new ways of using this transmission as a proxy for studying the intellectual networks that participated in such an exchange. In this presentation, I will exemplify this approach on the corpus of early medieval annotations to the Etymologiae of Isidore of Seville, the most important medieval Latin encyclopaedia. More specifically, it can be presupposed that most of the glosses to this text came into being in the context of its use for teaching in Carolingian period (c. 750 – 900). Their transfer, thus, may be related to the circulation of schoolmasters, students, and books through the networks of Carolingian schools.</p>
• Koranic Science [IO Bijapur 285] ٱلْكَشَّاف, glosses on
<ul> <li>Commentary on al-Qurʼān القرآن الكريم</li> <li><strong>This manuscript is now IO Bijapur 285 in the India Office collections.</strong></li> <li><strong>[metadata:</strong><a href="https://de.wikipedia.org/wiki/Otto_Loth"> <strong>Otto Loth, </strong></a><strong><em><a href="http://doi.org/10.5281/zenodo.3923636">A Catalogue of the Arabic Manuscripts in the Library of the India Office</a></em>, (volume 1), no. 60 here with notations and hyperlinks]</strong>.</li> </ul> <p><a href="https://archive.org/details/in.gov.ignca.35737/page/11/mode/2up?q=349"><strong>60</strong></a></p> <p>B 285. Size 8<sup>1/2</sup> in. by 5<sup>1/4</sup> in.; foll. 217. Seventeen lines in a page.</p> <p>Glosses of <a href="http://worldcat.org/identities/lccn-n88286025/">SAIYID SHARÎF JURJÂNÎ</a> (‘Alî b. Muḥammad, d. A. H. 816 [1413 CE]) [otherwise <a href="https://www.wikidata.org/wiki/Q3063378">ʿAlī ibn Muḥammad ibn ʿAli al‐Ḥusaynī al‐Jurjānī]</a> on the <a href="https://doi.org/10.1515/9783110533408"><em>Kashshâf</em></a> [of <a href="http://worldcat.org/identities/lccn-n81048213/">Zamakhsharī</a>, ob. 1044], terminating at Sû. 2, 23. Cf. <a href="https://en.wikipedia.org/wiki/Kashf_al-Zunun">Ḥ. Kh.</a> v. 187.</p> <p>Clearly written. Dated Sunday, 4<sup>th</sup> Rajab, 939 [=30 January 1533 CE, but the day seems to corresponds to Thursday]. In good preservation; one defect after fol. 88.</p> <p><a href="https://en.wikipedia.org/wiki/Bijapur_Collection">Bîj. Libr</a>., A. H. 1003 [=1594-95 CE]. Cat. 221, i. 2. [ed note: = Ḥaqīm Ḥamid al-Dīn's Urdu catalogue prepared at the behest of <a href="http://worldcat.org/identities/lccn-n85120180/">H. B. Frere</a>] </p> <ul> <li>[<strong>ed note</strong>: Quick <a href="https://www.fihrist.org.uk/catalog/person_19930855">link</a> to other manuscripts of works by Jurjānī.]</li> <li>[<strong>ed note</strong>: Guard cover over the original binding dated 8 October 1925]</li> </ul>
'Proto-rivalry': how the binocular brain identifies gloss
<p>Dataset relative to the following publication:</p> <p>Muryy, A. A., Fleming, R. W., & Welchman, A. E. (2016). ‘Proto-rivalry’: how the binocular brain identifies gloss. <em>Proceedings of the Royal Society B, 283(1830)</em>: 20160383.</p> <p>Each folder contains *.mat files with the data relative to the respective experiment(s) and a matlab script for loading and plotting the data. The matlab scripts also include comments regarding the data files.</p>
The morphologically glossed Rigveda - The Zurich annotation corpus revised and extended. Hosted by VedaWeb - Online Research Platform for Old Indic Texts.
<p>This file contains morphological and lexicographic annotations for the Rigveda. It was created in the DFG-funded research project Vedaweb and used as source data for the linguistic research platform <a href="https://vedaweb.uni-koeln.de">vedaweb.uni-koeln.de</a>.</p> <p>Prof. Dr. Paul Widmer and Dr. Salvatore Scarlata from the "Institut für Vergleichende Sprachwissenschaft" (Universität Zürich) provided the VedaWeb project a Filemaker file that was later transformed in Cologne into an Excel file. This data contained a version of the Rigveda by Prof. Dr. A. Lubotsky ("Indo-European Linguistics", Leiden University) that had been morphosytactically annotated over the course of more than 10 years at the University of Zurich. It also contained for each token, if available, a reference to an entry in Grassmann's dictionary for the Rigveda.</p> <p> </p> <p><strong>Modifications made by Jakob Halfmann and Natalie Korobzow to the data in 2020:</strong></p> <p>Disambiguation of the relevant categories, if unspecified in Zurich data, according to the Grassmann dictionary (updates from 6th edition partially included up to page 274):</p> <ul> <li>case, gender and number for nouns, pronouns (columns G–I)</li> <li>number, person, mood, tense and voice for verbs (columns I–M) up to line 109216</li> <li>case, gender, number, tense and voice for participles (columns G–I, L–M) up to line 109216</li> <li>absolutives are marked as Abs. in columns N and V</li> <li>Inconsistencies between the original file from Zurich and the Grassmann dictionary as well as internal inconsistencies in Grassmann are noted in column AE, whenever they were noticed.</li> <li>Zurich data was overwritten by conflicting Grassmann data in columns G–M but retained elsewhere.</li> <li>Verb classes according to Whitney (1885) and Jamison (1983) for class 10 in column Y, differences in root spelling between Whitney and Grassmann are noted in column Z. All potential verb classes provided by Whitney are given for every occurrence of the root.</li> <li>Local particles and verbal forms containing them are marked as LP in column AF.</li> <li>Comparatives and superlatives are marked as such in column X and desideratives as Des. in column Y.</li> </ul> <p> </p> <p><strong>Modifications made by Anna Fischer (data transformation, technical realisation) to the data:</strong></p> <p>New structure of data table for linguistic annotations with new column titles:</p> <ul> <li>A - "VERS_NR": renamed column (from "belege::stelleMMSSSRR")</li> <li>B - "PADA_NR": renamed column (from "belege::pada")</li> <li>C - "PADA_TEXT_LUBOTSKY": renamed column (from "belege::lubotskypada")</li> <li>D - "TOKEN_NR_VERS": renamed column (from "belege::wortnummer rc")</li> <li>E - "TOKEN_NR_PADA": renamed column (from "belege::wortnummer pada")</li> <li>F - "FORM": renamed column (from "belege::form")</li> <li>G - "KASUS": renamed column (from "belege::kasus")</li> <li>H - "GENUS": renamed column (from "belege::genus")</li> <li>I - "NUMERUS": renamed column (from "belege::numerus")</li> <li>J - "PERSON": renamed column (from "belege::person")</li> <li>K - "TEMPUS": moved and renamed column (from L "belege::tempus")</li> <li>L - "PRAESENSKLASSE": created new column for present stem class for each form</li> <li>M - "LEMMA_PRAESENSKLASSEN": created column for present stem classes of respective lemma: Moved and renamed column (from Y "formen::zusätzliche merkmale verb"), moved values "Abs." and "Inf." to column P "INFINIT", moved values "Prek." and "si-Ipv." to column N "MOOD", moved value "Des." to column Q "ABGELEITETE_KONJUGATION": moved value "se-Form" to column W "WEITERE_WERTE"</li> <li>N - "MODUS": moved and renamed column (from K "belege::modus")</li> <li>O - "DIATHESE": moved and renamed column (from M "belege::diathese")</li> <li>P - "INFINIT": created new column for infinite forms "Abs.", "Inf.", "Ptz.", "ta-Ptz.", "na-Ptz."</li> <li>Q - "ABGELEITETE_KONJUGATION": created new column for secondary conjugation "Des.", "Int.", "Kaus."</li> <li>R - "GRADUS": created new column for degree: "Comp.", "Sup."</li> <li>S - "LOKALPARTIKEL": moved and renamed column (from AF "LP")</li> <li>T - "LEMMA_ZÜRICH": moved and renamed column (from AA "lemmata klassisch::lemma")</li> <li>U - "LEMMA_ZÜRICH_LEMMATYP": moved and renamed column (from AB "lemmata klassisch::lemmatyp")</li> <li>V - "LEMMA_ZÜRICH_BEDEUTUNG": moved and renamed column (from AC "lemmata klassisch::bedeutung")</li> <li>W - "WEITERE_WERTE": created new column for all miscellaneous values: e.g. "Hyperchar.", "n-haltig", "se-Form"</li> <li>X - "KOMMENTAR": created new column merging former columns Z "formen::HELPformbestimmung", AD "lemmata klassisch::HELPbedeutung" and AE "anmerkungen abweichungen"</li> </ul> <p>Columns that were removed due to redundant information:</p> <ul> <li>"formen::zusätzliche merkmale nomen": values "superlative" And "comparative" were renamed "sup." and "comp." and moved to new column R "GRADUS", all other values were moved to new column for miscellaneous W "WEITERE_WERTE"</li> <li>"belege::belegbestimmung summe simpel": values "Ptz.", "ta-Ptz." and "na-Ptz." were moved to new new column P "INFINIT"</li> <li>"belege::kasus bestof"</li> <li>"belege::genus bestof"</li> <li>"belege::numerus bestof"</li> <li>"belege::person bestof"</li> <li>"belege::modus bestof"</li> <li>"belege::tempus bestof"</li> <li>"belege::diathese bestof"</li> <li>"belege::belegbestimmung bestof summe sophistiziert"</li> </ul> <p> </p> <p><strong>Revisions and additions made by Antje Casaretto to the data in 2023:</strong></p> <ul> <li>F-T: - revision and correction (wherever necessary) of all annotations (books 1-7)</li> <li>G,H,I - disambiguation of case forms, reg. pronouns and nominal forms, if unspecified in Zurich data (books 1-7)</li> <li>L - disambiguation of present stem classes (book 7 and book 1 up to line 21050 vers 01.125.01)</li> <li>M - disambiguation of denominal verbs from primary verbs of the 10th class (books 1-10)</li> <li>N - disambiguation of precative and optative forms wherever possible (books 1-7)</li> <li>Q - new annotations for "Int." (intensives) and "Kaus." (causatives) (books 1-7)</li> </ul> <p> </p> <p dir="ltr"><strong>Revisions and additions made by Antje Casaretto to the data in 2024 with support in data modeling and automation by Anna Fischer:</strong></p> <ul> <li>F-T: revision and correction (wherever necessary) of all annotations (books 8-10)</li> <li>G,H,I: disambiguation of case and gender forms in nominal and pronominal forms, if unspecified in Zurich data (books 8-10)</li> <li>L, M: disambiguation of present stem classes (books 1-10)</li> <li>N: disambiguation of precative and optative forms wherever possible (books 8-10)</li> <li>P: new annotations for "Gdv." (gerundives)</li> <li>Q: new annotations for "Den." (denominatives) (books 1-10) and further annotations of “Kaus.” (causatives) and “Int.” (intensives) (books 8-10)</li> <li>T: revision of lemmatization (books 1-10)</li> <li>V: update of meanings according to revised lemmatization; minimal revision</li> <li>W: revised annotation of ending -se (“se-Form”) (books 1-10); no systematic revision</li> <li>X: no systematic revision</li> <li>A-U: general revision of formal inconsistencies and typing errors (book 1-10)</li> </ul> <p> </p> <p dir="ltr"><strong>Revisions made by Natalie Korobzov and Pascal Coenen to the data in 2024 with computational support by Anna Fischer:</strong></p> <ul> <li>Y - "LEMMA_GRASSMANN_ID": new column for references to Grassmann dictionary (books 1-10) and revision of Grassmann references</li> </ul>
Etruscan black-gloss bowl
"Etruscan black-gloss bowl from the Castello Banfi Collection Black-gloss bowl. Orange clay with black paint. Circular stamp on interior. Etruscan-Lazio production. Third quarter of the 3rd c. BCE. Montalcino, Poggio alle Mura (Castello Banfi), n. inv. 109. Bischeri 2022: n. cat. 144. Processed in Reality Capture from 289 images. GDH ID C6_109 This project was done under the authority of the Soprintendenza Archeologia, belle arti e paesaggio per le province di Siena Grosseto e Arezzo in collaboration with Global Digital Heritage. We thank Elizabeth Koenig, Hospitality Project Director at Castello Banfi and President of the Scientific Committe of Fonadazione Banfi, for coordinating this project and making the collection accessible. The catalog of the Banfi collection was created by Dr. Matteo Bischeri (2022)." Source: Objaverse 1.0 / Sketchfab
Grammar [IO Bijapur 27] ʻAbd al-Ghafūr Lārī's glosses on فوائد الضيائية
<ul> <li><strong>Grammar.</strong></li> <li><strong>This manuscript is now IO Bijapur 27 </strong><strong>in the India Office collections.</strong></li> <li><strong>[metadata:</strong><a href="https://de.wikipedia.org/wiki/Otto_Loth"> <strong>Otto Loth, </strong></a><strong><em><a href="http://doi.org/10.5281/zenodo.3923636">A Catalogue of the Arabic Manuscripts in the Library of the India Office</a></em>, (volume 1), no. 928 here with further notations and hyperlinks]</strong>.</li> </ul> <p><a href="https://archive.org/details/catalogueofarabi01greauoft/page/260/mode/2up"><strong>928</strong></a>.</p> <p>B 27. Size 6<sup>3/4</sup> in. by 5 in.; foll. 151. Seventeen lines in a page.</p> <p>Glosses on <em>Jâmî</em>’<em>s</em> Commentary [= <em>al-Fawāʼid al-ḍiyāʼīyah</em>, Jāmī's commentary on <a href="https://id.loc.gov/authorities/names/n82162164.html">Ibn Ḥājib</a>'s <a href="https://archive.org/details/ldpd_14649875_000/page/n9/mode/2up"><em>al-Kāfiyah</em></a>], by his pupil, '<a href="https://id.loc.gov/authorities/names/n78052447.html">ABD AL-GHAFÛR LÂRÎ</a> (d. A.H. 912). Cf. <a href="https://en.wikipedia.org/wiki/Kashf_al-Zunun">Ḥ. Kh. </a>v. 11, and <a href="https://doi.org/10.5281/zenodo.6827074">Cat. St. Petersb</a>. 232. This work was printed at Constantinople, A.H. 1253. Another edition, which includes a continuation of the work ( تكملة ) by ׳Abd al-ḥakîm (Siyâlkûtî?), was printed A.H. 1254 (place not named – Calcutta ?) , in small quarto , pp. 728.</p> <p>Begins:</p> <p> قوله الحمد مصدر المعلوم و اللام للجنس</p> <p>The glosses extend to the paragraph اسماء الافعال (= fol. 120<em>v</em>. in no. 921).</p> <p>To this is added: -</p> <p>Foll. 149<em>v</em>.-151. A Shî’ah Legend, illustrating the miraculous powers of ‘Alî. Begins:</p> <p>خبر من خزانة مولانا مفترض الطاعة على الخلق اجمعين امير المؤمنين عم حدثنا ابوعبدالله بن زكرياء عن ابى جوهر بن اسود عن محمد بن عبدالله السابغ (؟) يرفعه الى سلمان الفارسى رضه انه قال كنا جلوسا عند مولانا امير المؤمنين الخ.</p> <p>The last portion of it is written on the margin, from the end backwards.</p> <p>Clearly written. Of the tenth century.</p> <p><a href="https://www.wikidata.org/wiki/Q105569536">Bîj. Libr</a>., A.H. 992, from Khalîl Allah b. Faḍl Allah Ja’farî.</p> <p>Seals of the latter (A.H. 977), and of his father.</p> <p>Cat. 235, iii. 1.</p> <ul> <li> <p>[<strong>ed note:</strong> for an account of Jāmī, see Nicholas L. Heer (tr.): <em><a href="https://archive.org/details/thepreciouspearlaljamisaldurrahalfakhirah/page/n11/mode/2up">The precious pearl: Al-Jāmī's al-Durrah al fakhirah, together with his glosses and the Commentary of 'Abd al-Ghafūr al-Lārī</a></em>. ix, 237 pp. Albany, N.Y.: State University of New York Press, 1979.]</p> </li> </ul> <p> </p>
Miscellanies [IO Bijapur 353] Gloss on the شرح الوقایة with Gloss on انوار التنزيل
<ul> <li><strong>Miscellanies.</strong></li> <li><strong>This manuscript is now IO Bijapur 353 </strong><strong>in the India Office collections.</strong></li> <li><strong>[metadata:</strong><a href="https://de.wikipedia.org/wiki/Otto_Loth"> <strong>Otto Loth, </strong></a><strong><em><a href="http://doi.org/10.5281/zenodo.3923636">A Catalogue of the Arabic Manuscripts in the Library of the India Office</a></em>, (volume 1), no. 1030 here with further notations and hyperlinks]</strong>.</li> </ul> <p><strong>MISCELLANIES.</strong></p> <p><a href="https://archive.org/details/catalogueofarabi01greauoft/page/284/mode/2up?view=theater"><strong>1030</strong></a>.</p> <p>B353. Size 10 in. by 6 in.; foll. 254. Twenty-five lines in a page.</p> <p><strong>I. Foll. 1-99.</strong> The beginning and two other fragments of a Gloss on the شرح الوقایة (see no. <a href="https://doi.org/10.5281/zenodo.4625492">221</a>). The author is, according to the modern inscription, <a href="http://worldcat.org/identities/lccn-n89260321/">SHÂH WAJÎH AL-DÎN</a>.</p> <p>Begins:</p> <p>الحمد لله رب العالمین ...قوله سعد جده و الانجح (= وانجح) جده الجد بالفتح البخت و بالکسر الاجتهاد الخ</p> <p>Ends in the کتاب الغصب.</p> <p>The first fragment inelegantly, the others well written.</p> <p>Bound with this is:</p> <p><strong>II. Foll. 100-254.</strong> A fragment of a Gloss on <a href="https://www.worldcat.org/identities/lccn-n82085851/"><em>Bai</em></a><em><a href="https://www.worldcat.org/identities/lccn-n82085851/">ḍâwî</a>’s </em>Commentary on the Koran (see no. <a href="https://doi.org/10.5281/zenodo.4479894">70</a> <a href="https://www.worldcat.org/title/tafsir-al-baydawi-al-musamma-anwar-al-tanzil-wa-asrar-al-tawil/oclc/51865591">انوار التنزيل</a>), which is also ascribed to the aforesaid <a href="https://en.wikipedia.org/wiki/Wajihuddin_Alvi">SHÂH WAJÎH AL-DÎN</a>.</p> <p>It extends from Sû. 2 to Sû. 13, and is imperfect both at the beginning and end. The first words are:</p> <p>کیف تکفرون</p> <p>Written like the latter portion of no. I. Defects after foll. 113, 123, and 238.</p> <p>Much worm-eaten, but carefully mended.</p> <p>Cat. 227, viii. 3.</p> <p> </p>
GlossReader at LSCDiscovery: Train to Select a Proper Gloss in English -- Discover Lexical Semantic Change in Spanish
<pre>Precomputed vectors for the GlossReader system. LSCDiscovery Competition: https://codalab.lisn.upsaclay.fr/competitions/2243. </pre>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.