Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

63

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

63 results for “charters”

Learn how ShareScore rates datasets ↗
zenodo36/100

UBA000014812 - Groot placaat en charter-boek van Vriesland: Vierde deel

<p>Titel:&nbsp;<em> </em><em>Groot placaat en charter-boek van Vriesland: Vierde deel</em></p> <p>Publisher: Willem Coulon</p> <p>Place: Leeuwarden</p> <p>Year: 1782</p> <p>Used version:&nbsp;The copy we used for the transcriptions is held at the University Library of Amsterdam, digitised by KB National Library of the Netherlands.</p> <p>Link digitised version of the book: https://books.google.nl/books?id=9FZkAAAAcAAJ</p> <p>(Main) Language: Dutch.</p> <p>Province: Friesland.</p> <p>Font: Roman.</p> <p>Model used:&nbsp;Dutch_Romantype_Print.</p> <p>Version of Transkribus used:&nbsp;v.1.9.1.</p> <p>Other info:&nbsp;Abbyy FineReader v.11 has been used.</p> <p>Link model:&nbsp;For more information on the HTR-model used, please visit:&nbsp;<a href="https://lab.kb.nl/dataset/entangled-histories-ordinances-low-countries">https://lab.kb.nl/dataset/entangled-histories-ordinances-low-countries</a>.</p> <p>Transcription conventions:</p> <ul> <li> <p>The abbreviations have been written out into full words.</p> </li> <li> <p>The hyphens at the end of a line have been kept (when there).</p> </li> </ul> <p>If you are in need of the original scans of the documents, please contact&nbsp;<a href="mailto:xxxxxx@kb.nl">dataservices@kb.nl</a>.</p> <p>This transcription is part of the dataset created with the &lsquo;Entangled Histories&rsquo;-project.</p> <p>PI: dr. C.A. Romein;<br> Scientific Programmer: S.F. Veldhoen, MSc;<br> Project Manager: drs. M. de Gruijter.</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2020View details →
zenodo36/100

UBA000014810 - Groot placaat en charter-boek van Vriesland: Derde deel

<p>Titel:&nbsp;<em> </em><em>Groot placaat en charter-boek van Vriesland: Derde deel</em></p> <p>Publisher: Willem Coulon</p> <p>Place: Leeuwarden</p> <p>Year: 1778</p> <p>Used version:&nbsp;The copy we used for the transcriptions is held at the University Library of Amsterdam, digitised by KB National Library of the Netherlands.</p> <p>Link digitised version of the book: https://books.google.nl/books?id=iFVkAAAAcAAJ</p> <p>(Main) Language: Dutch.</p> <p>Province: Friesland.</p> <p>Font: Roman.</p> <p>Model used:&nbsp;Dutch_Romantype_Print.</p> <p>Version of Transkribus used:&nbsp;v.1.9.1.</p> <p>Other info:&nbsp;Abbyy FineReader v.11 has been used.</p> <p>Link model:&nbsp;For more information on the HTR-model used, please visit:&nbsp;<a href="https://lab.kb.nl/dataset/entangled-histories-ordinances-low-countries">https://lab.kb.nl/dataset/entangled-histories-ordinances-low-countries</a>.</p> <p>Transcription conventions:</p> <ul> <li> <p>The abbreviations have been written out into full words.</p> </li> <li> <p>The hyphens at the end of a line have been kept (when there).</p> </li> </ul> <p>If you are in need of the original scans of the documents, please contact&nbsp;<a href="mailto:xxxxxx@kb.nl">dataservices@kb.nl</a>.</p> <p>This transcription is part of the dataset created with the &lsquo;Entangled Histories&rsquo;-project.</p> <p>PI: dr. C.A. Romein;<br> Scientific Programmer: S.F. Veldhoen, MSc;<br> Project Manager: drs. M. de Gruijter.</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2020View details →
zenodo36/100

UBA000168299 - Groot placaat en charter-boek van Vriesland: Vyfde deel

<p>Titel:&nbsp;<em> </em><em>Groot placaat en charter-boek van Vriesland</em><em>:&nbsp;</em><em>Vyfde deel </em></p> <p>Publisher: Hermannus Post</p> <p>Place: Leeuwarden</p> <p>Year: 1793</p> <p>Used version:&nbsp;The copy we used for the transcriptions is held at Amsterdam University Library and digitised by the KB National Library of the Netherlands.</p> <p>Link digitised version of the book: <a href="https://books.google.nl/books?id=xsjkvBzXPHgC">https://books.google.nl/books?id=xsjkvBzXPHgC</a></p> <p>(Main) Language: Dutch.</p> <p>Province: Friesland</p> <p>Font: Roman.</p> <p>Model used:&nbsp;Dutch_Romantype_Print.</p> <p>Version of Transkribus used:&nbsp;v.1.9.1.</p> <p>Other info:&nbsp;Abbyy FineReader v.11 has been used.</p> <p>Link model:&nbsp;For more information on the HTR-model used, please visit:&nbsp;<a href="https://lab.kb.nl/dataset/entangled-histories-ordinances-low-countries">https://lab.kb.nl/dataset/entangled-histories-ordinances-low-countries</a>.</p> <p>Transcription conventions:</p> <ul> <li> <p>The abbreviations have been written out into full words.</p> </li> <li> <p>The hyphens at the end of a line have been kept (when there).</p> </li> </ul> <p>If you are in need of the original scans of the documents, please contact&nbsp;<a href="mailto:xxxxxx@kb.nl">dataservices@kb.nl</a>.</p> <p>This transcription is part of the dataset created with the &lsquo;Entangled Histories&rsquo;-project.</p> <p>PI: dr. C.A. Romein;<br> Scientific Programmer: S.F. Veldhoen, MSc;<br> Project Manager: drs. M. de Gruijter.</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2020View details →
zenodo36/100

Confidence Matrices of Automatically Recognized Charters of the Monastery of Einsiedeln

<p>The data set consists of data gained through automatic text recognition with HTR+ models.<br>The model is usable within the Transkribus environment: <a href="https://readcoop.eu/model/charter-scripts-german-latin-french/">readcoop.eu/model/charter-scripts-german-latin-french/</a>.</p><p>Each folder is a group of documents, depending on the call-number in the archives (<a href="https://www.klosterarchiv.ch/">https://www.klosterarchiv.ch/</a>). Each sub-folder is a document and each CSV is a line of recognized text in the form of a Confidence Matrix.</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Paharpur (Bangladesh). Charter of Budhagupta dated year 159 (side 1).

<p><a href="https://siddham.network/inscription/in00065/">IN00065</a> Paharpur (Bangladesh). Charter of Budhagupta dated year 159 (side 1).</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Paharpur (Bangladesh). Charter of Budhagupta dated year 159 (side 2).

<p><a href="https://siddham.network/inscription/in00065/">IN00065 </a>Paharpur (Bangladesh). Charter of Budhagupta dated year 159 (side 2).</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Multilingual named entity recognition for medieval charters. Datasets and models

<p>Annotated dataset for training named entities recognition models for medieval charters in Latin, French and Spanish.</p> <p>&nbsp;</p> <p>The original raw texts for all charters were collected from four charters collections</p> <p>- HOME-ALCAR corpus : <a href="https://zenodo.org/record/5600884">https://zenodo.org/record/5600884</a></p> <p>- CBMA : <a href="https://www.google.com/url?sa=t&amp;rct=j&amp;q=&amp;esrc=s&amp;source=web&amp;cd=&amp;ved=2ahUKEwisvNOa3qP3AhULyIUKHenpBoAQFnoECA0QAQ&amp;url=http%3A%2F%2Fwww.cbma-project.eu%2F&amp;usg=AOvVaw0blsASXzOKSNz_EkixwfJT">http://www.cbma-project.eu</a></p> <p>- Diplomata Belgica : <a href="https://www.diplomata-belgica.be">https://www.diplomata-belgica.be</a></p> <p>- CODEA corpus :<a href="http://https://corpuscodea.es/"> https://corpuscodea.es/</a></p> <p>&nbsp;</p> <p>We include (i) the annotated training datasets, (ii) the contextual and static embeddings trained on medieval multilingual texts and (iii) the named entity recognition models trained using two architectures: Bi-LSTM-CRF + stacked embeddings and fine-tuning on Bert-based models (mBert and RoBERTa)</p> <p>Codes, datasets and notebooks&nbsp;used to train models can&nbsp;be consulted in our&nbsp;gitlab repository:&nbsp;<a href="https://gitlab.com/magistermilitum/ner_medieval_multilingual">https://gitlab.com/magistermilitum/ner_medieval_multilingual</a></p> <p>Our best RoBERTa model is also available in the HuggingFace library:&nbsp;<a href="https://huggingface.co/magistermilitum/roberta-multilingual-medieval-ner">https://huggingface.co/magistermilitum/roberta-multilingual-medieval-ner</a></p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Text zones in images of medieval charters from Stiftsarchiv Seitenstetten in monasterium.net

<p>PAGE annotation (<a href="http://schema.primaresearch.org/PAGE/gts/pagecontent/2013-07-15">http://schema.primaresearch.org/PAGE/gts/pagecontent/2013-07-15</a>) of image regions containing texts in a set of scan medieval charters, scanned in the Stiftsarchiv Seitenstetten (<a href="http://monasterium.net/mom/AT-StiASei/SeitenstettenOSB/fond">http://monasterium.net/mom/AT-StiASei/SeitenstettenOSB/fond</a>). Images created by the International Center for Archival Research ICARus (<a href="http://icar-us.eu/">http://icar-us.eu/</a>) Annotations created with the support of the Austrian Science Fund, project number P 26.706 (<a href="https://illuminierte-urkunden.uni-graz.at/">https://illuminierte-urkunden.uni-graz.at/</a>). It does not contain any further segmentation for words or characters.</p>

opencc-by-nc-4.0Mar 2018View details →
zenodo36/100

IN01006 Rawan Charter of Narendra

<p>IN01006 Rawan Charter of Narendra of the Śarabhapurīya dynasty, <em>circa</em> 500 CE. Bhopal, State Museum.</p>

opencc-by-nc-nd-4.0Apr 2018View details →
zenodo36/100

Copper-plate charter of Maharaja Rudradasa of the Valkha dynasty. Bhopal, State Museum

<p>Copper-plate charter of Maharāja Rudradasa of the Valkhā dynasty. Bhopal, State Museum</p>

opencc-by-nc-nd-4.0Apr 2018View details →
zenodo36/100

Charter School Websites: A Physical Education and Physical Activity Content Analysis

<p>Excel file for data set associated with analysis of 520 California elementary charter school websites&#39; mentioning of physical education and physical activity opportunities.</p>

opencc-by-4.0Jun 2018View details →
zenodo32/100

IN00611 Dharasena II charter year 252 (XML)

<p>Dharasena II charter year 252 (EI XXXVII) XML</p>

opencc-by-nc-nd-4.0Nov 2018View details →
zenodo32/100

OB00613B Dharasena II charter year 252

<p>OB00613B Dharasena II charter year 252 (EI XXXVII)</p>

opencc-by-nc-nd-4.0Nov 2018View details →
zenodo32/100

OB00613A Dharasena II charter year 252

<p>OB00613A&nbsp; Dharasena II charter year 252 (EI XXXVII)</p>

opencc-by-nc-nd-4.0Nov 2018View details →
zenodo28/100

UBA000168296 - Groot placaat en charter-boek van Vriesland:Tweede deel .

<p>Titel:&nbsp;<em> </em><em>Groot placaat en charter-boek van Vriesland</em><em>:</em><em>Tweede</em><em> deel .</em></p> <p>Publisher: Willem Coulon</p> <p>Place: Leeuwarden</p> <p>Year: 1773</p> <p>Used version:&nbsp;The copy we used for the transcriptions is held at Amsterdam University Library and digitised by the KB National Library of the Netherlands.</p> <p>Link digitised version of the book: https://books.google.nl/books?vid=KBNL:UBA000168296&amp;redir_esc=y</p> <p>(Main) Language: Dutch.</p> <p>Province: Friesland</p> <p>Font: Roman.</p> <p>Model used:&nbsp;Dutch_Romantype_Print.</p> <p>Version of Transkribus used:&nbsp;v.1.9.1.</p> <p>Other info:&nbsp;Abbyy FineReader v.11 has been used.</p> <p>Link model:&nbsp;For more information on the HTR-model used, please visit:&nbsp;<a href="https://lab.kb.nl/dataset/entangled-histories-ordinances-low-countries">https://lab.kb.nl/dataset/entangled-histories-ordinances-low-countries</a>.</p> <p>Transcription conventions:</p> <ul> <li> <p>The abbreviations have been written out into full words.</p> </li> <li> <p>The hyphens at the end of a line have been kept (when there).</p> </li> </ul> <p>If you are in need of the original scans of the documents, please contact&nbsp;<a href="mailto:xxxxxx@kb.nl">dataservices@kb.nl</a>.</p> <p>This transcription is part of the dataset created with the &lsquo;Entangled Histories&rsquo;-project.</p> <p>PI: dr. C.A. Romein;<br> Scientific Programmer: S.F. Veldhoen, MSc;<br> Project Manager: drs. M. de Gruijter.</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2020View details →
zenodo28/100

UBA000168294 - Groot placaat en charter-boek van Vriesland: Eerste deel

<p>Titel:&nbsp;<em> </em><em>Groot placaat en charter-boek van Vriesland</em><em>:&nbsp;</em><em>Eerste deel .</em></p> <p>Publisher: Willem Coulon</p> <p>Place: Leeuwarden</p> <p>Year: 1768</p> <p>Used version:&nbsp;The copy we used for the transcriptions is held at Amsterdam University Library and digitised by the KB National Library of the Netherlands.</p> <p>Link digitised version of the book: https://books.google.nl/books?vid=KBNL:UBA000168294&amp;redir_esc=y</p> <p>(Main) Language: Dutch.</p> <p>Province: Leeuwarden</p> <p>Font: Roman.</p> <p>Model used:&nbsp;Dutch_Romantype_Print.</p> <p>Version of Transkribus used:&nbsp;v.1.9.1.</p> <p>Other info:&nbsp;Abbyy FineReader v.11 has been used.</p> <p>Link model:&nbsp;For more information on the HTR-model used, please visit:&nbsp;<a href="https://lab.kb.nl/dataset/entangled-histories-ordinances-low-countries">https://lab.kb.nl/dataset/entangled-histories-ordinances-low-countries</a>.</p> <p>Transcription conventions:</p> <ul> <li> <p>The abbreviations have been written out into full words.</p> </li> <li> <p>The hyphens at the end of a line have been kept (when there).</p> </li> </ul> <p>If you are in need of the original scans of the documents, please contact&nbsp;<a href="mailto:xxxxxx@kb.nl">dataservices@kb.nl</a>.</p> <p>This transcription is part of the dataset created with the &lsquo;Entangled Histories&rsquo;-project.</p> <p>PI: dr. C.A. Romein;<br> Scientific Programmer: S.F. Veldhoen, MSc;<br> Project Manager: drs. M. de Gruijter.</p>

opencc-by-4.0Jan 2020View details →
zenodo28/100

KBNLB410014860 - Groot placaat en charter-boek van Vriesland: Den 12 October 1686-28 Februarij 1705. Zesde deel

<p>Titel:&nbsp;<em> </em><em>Groot placaat en charter-boek van Vriesland: Den 12 October 1686-28 Februarij 1705. Zesde deel</em></p> <p>Publisher: Hermannus Post</p> <p>Place: Leeuwarden</p> <p>Year: 1795</p> <p>Used version:&nbsp;The copy we used for the transcriptions is held at the KB National Library of the Netherlands.</p> <p>Link digitised version of the book:&nbsp;https://books.google.nl/books?id=hv4XuQEACAAJ</p> <p>(Main) Language: Dutch.</p> <p>Province: Friesland.</p> <p>Font: Roman.</p> <p>Model used:&nbsp;Dutch_Romanprint_Print</p> <p>Version of Transkribus used:&nbsp;v.1.9.1.</p> <p>Other info:&nbsp;Abbyy FineReader v.11 has been used.</p> <p>Link model:&nbsp;For more information on the HTR-model used, please visit:&nbsp;<a href="https://lab.kb.nl/dataset/entangled-histories-ordinances-low-countries">https://lab.kb.nl/dataset/entangled-histories-ordinances-low-countries</a>.</p> <p>Transcription conventions:</p> <ul> <li> <p>The abbreviations have been written out into full words.</p> </li> <li> <p>The hyphens at the end of a line have been kept (when there).</p> </li> </ul> <p>If you are in need of the original scans of the documents, please contact&nbsp;<a href="mailto:xxxxxx@kb.nl">dataservices@kb.nl</a>.</p> <p>This transcription is part of the dataset created with the &lsquo;Entangled Histories&rsquo;-project.</p> <p>PI: dr. C.A. Romein;<br> Scientific Programmer: S.F. Veldhoen, MSc;<br> Project Manager: drs. M. de Gruijter.</p>

opencc-by-4.0Jan 2020View details →
zenodo28/100

A hybrid approach to the small unannotated corpus-based language comparison and its application to the Old East Slavic charters - Supplementary material 5 (Corpus-based language distance measurement results)

<h1>General description</h1> <p>These are the results of the experiments with the use of <a href="https://doi.org/10.5281/zenodo.11395683" target="_blank" rel="noopener">corpus-based language distance measurement package</a> on the material of <a href="https://doi.org/10.5281/zenodo.14057668" target="_blank" rel="noopener">Old East Slavic</a>, <a href="https://doi.org/10.5281/zenodo.14148179" target="_blank" rel="noopener">modern East Slavic</a>, and <a href="https://doi.org/10.5281/zenodo.14148561">modern standard Slavic</a> lects. There are 40 possible experiments for each data set, divided by the usage of:</p> <ul> <li>topic antimodelling heuristic,</li> <li>Soerensen-Dice coefficient-based normalisation,</li> <li>the presence of hybridisation of frequency-based metric for coinciding units and combined frequency-based metric and string similarity measure for non-coinciding units,</li> <li>hybridisation type,</li> <li>exact type of string similarity measure used for combination,</li> <li>alphabet entropy-based normalisation for vector-based string similarity measures.</li> </ul> <p>In addition, modern standard Slavic dataset undergoes experiments 4 times that differ by the share of its size, used for measurements (0.1, 0.3, 0.6 and 1).</p> <p>For further information on each of the experiment parameters, refer to the documentation of the package.</p> <h1>Data set structure</h1> <h2>Executive summary</h2> <p>Data set consists of 240 folders that represent information on the experiments and 1&nbsp;<code>.csv</code>-file that aggregates the resulting values into a single table.</p> <p>Folders with indices 1-20 and 121-140 contain experiments with 0.1 share of the modern standard Slavic dataset; the first sequence applies topic antimodelling heuristic, the second sequence does not employ it.</p> <p>Folders with indices 21-40 and 141-160 contain experiments with 0.3 share of the modern standard Slavic dataset; the first sequence applies topic antimodelling heuristic, the second sequence does not employ it.</p> <p>Folders with indices&nbsp; 41-60 and 161-180 contain experiments with 0.6 share of the modern standard Slavic dataset; the first sequence applies topic antimodelling heuristic, the second sequence does not employ it.</p> <p>Folders with indices&nbsp; 61-80 and 181-200 contain experiments with the full share of the modern standard Slavic dataset; the first sequence applies topic antimodelling heuristic, the second sequence does not employ it.</p> <p>Folders with indices&nbsp; 81-100 and 201-220 contain experiments with the full share of the modern East Slavic dataset; the first sequence applies topic antimodelling heuristic, the second sequence does not employ it.</p> <p>Folders with indices&nbsp; 101-120 and 221-240 contain experiments with the full share of the Old East Slavic dataset; the first sequence applies topic antimodelling heuristic, the second sequence does not employ it.</p> <h2>.csv-file</h2> <p>Named <code>aggregated_results.csv</code>, lies in the root of the dataset. Separator is <strong>comma</strong> (<strong>,</strong>). Contains&nbsp;<strong>13 columns</strong> and&nbsp;<strong>241 row</strong>. The first row is&nbsp;<strong>header</strong>, the other 240 rows contain description for each conducted experiment and its resulting values, according to the columns. The columns are the following (in <strong>rtl</strong> order):</p> <ul> <li><strong>X.</strong> (<em>int</em>) - experiment ID; column is used as index.</li> <li><strong>Material&nbsp;</strong>(<em>string</em>) - the data set used for language distance measurement. The possible values are: <ul> <li>Slavic standard - Croatian, Slovenian, Slovak standard lects.</li> <li>Modern East Slavic - Northern Russian lect Megra, Central Russian lect Belogornoje, and Northern Belarusian lect Zialionka.</li> <li>Old East Slavic - Novgorod, Polack and Smolensk parts of the Old East Slavic continuum.</li> </ul> </li> <li><strong>Gensim&nbsp;</strong>(<em>int</em>) - the binary numeric indicator (0 or 1) of using the heuristic of&nbsp;<em>topic antimodelling</em>, namely, cleaning the words that were defined as a topic words by <em>gensim</em> Latent Dirichlet Association implementation (Rehurek &amp; Sojka, 2010). The intention of using this heuristic is to remove the tokens that are characteristic for the genre of the texts presented in the corpus for the sake of increasing the presence of the tokens that are characteristic of the lects themselves.</li> <li><strong>Split</strong> (<em>float</em>) - the used share of the data set (from 0 to 1); required to check the influence of the data set size on the metric efficiency.</li> <li><strong>Hybridisation</strong> (<em>string</em>) - the indicator of implementation of the hybridisation between the frequency-based metric between the 3-shingles (character 3-grams) that coincide for the compared lect pair, and the combination of frequency-based metric and string similarity measure between the 3-shingles that do not coincide for the compared lect pair. The possible values are: <ul> <li>TRUE: the experiment utilises hybridisation</li> <li>FALSE: the experiment does not utilise hybridisation.</li> </ul> </li> <li><strong>Hybridisation_type</strong> (<em>string</em>) - the indicator of how the frequency metric between coinciding 3-shingles and the combined metric between non-coinciding 3-shingles undergo the hybridisation process. The values are:<br> <ul> <li>JOINED - the approach is to multiply the means of the two.</li> <li>ARRAY -&nbsp;the approach is to join all the values into a single list, and then to score the mean.</li> <li>NOT_USED - experiment does not employ hybridisation. (<strong>Hybridisation&nbsp;</strong>is&nbsp;FALSE).</li> </ul> </li> <li><strong>Soerensen_normalisation</strong> (<em>string</em>) - the indicator of whether the frequency-based metric value undergoes normalisation with the division by Soerensen-Dice coefficient (a measure of number of coincidences between two lists) (Soerensen, 1948), in order to compensate the skewing between the coinciding and non-coinciding 3-shingles of the lects. The values are: <ul> <li>NOT_USED - <strong>Hybridisation_type&nbsp;</strong>is ARRAY, so there are no values to use Soerensen-Dice coefficient on.</li> <li>TRUE - the frequency-based metric undergoes division by the Soerensen-Dice coefficient.</li> <li>FALSE -&nbsp;the algorithm does not apply the normalisation by the Soerensen-Dice coefficient.</li> </ul> </li> <li><strong>Alphabet_normalisation</strong>&nbsp;(<em>string</em>) - indicator of whether the algorithm applies normalisation with the alphabet entropy&nbsp; (Shannon, 1948), the measure of differences in the skewings of symbols distribution in the texts, between the given lects. The values are:<br> <ul> <li>NOT_USED&nbsp; - a heuristic may not be implemented; present either in the cases, when <strong>Hybridisation&nbsp;</strong>is&nbsp;FALSE, or when the next parameter, <strong>Auxiliary_metrics</strong> is not&nbsp;VDND or&nbsp;VWJDND.</li> <li>TRUE<strong>&nbsp;</strong>- the experiment employs the heuristic.</li> <li>FALSE - the experiment does not employ the heuristic.</li> </ul> </li> <li><strong>Auxiliary_metrics&nbsp;</strong>(<em>string</em>) - the string similarity measure, used for the combination with the frequency-based metric for non-coinciding 3-shingles between analysed lects. There are five possible values: <ul> <li>LDND (Levenshtein distance normalised between analysed 3-shingles) (Holman et al., 2008).</li> <li>WJWDND (weighted Jaro-Winkler distance normalised between analysed 3-shingles) (Gueddah et al., 2015).</li> <li>VDND&nbsp;(Euclidean distance between the sums of symbol vector values between 3-shingles).</li> <li>VWJDND&nbsp;(VDND multiplied by scoring Jaro (Jaro, 1989) distance between analysed 3-shingles).</li> </ul> </li> <li><strong>Outgroup.identification</strong> (<em>string</em>) - the indicator of whether the outgroup detected in the given experiment coincides with the lect that preliminary manual classification supposes to be the outgroup. There are two possible values:<br> <ul> <li>CORRECT - the detected outgroup coincides with the supposed one.</li> <li>INCORRECT - the detected outgroup does not coincide with the supposed one.</li> </ul> </li> <li><strong>Outer.distance.split&nbsp;</strong>(<em>float</em>) - the length of the outgroup branch.</li> <li><strong>Inner.distance.split</strong> (<em>float</em>) - the distance between the split between the outgroup and the ingroup, and the split between the two ingroup lects.</li> <li><strong>Split.difference</strong> (<em>float</em>) - the division of <strong>Outer.distance.split</strong> by <strong>Inner.distance.split</strong>.</li> </ul> <h2>Folders</h2> <p>Each folder contains 6 files, each named according to the used experiment setup:</p> <ul> <li>3&nbsp;<code>.csv</code>-files that contain unit-by-unit comparison between each pair of the analysed lects. Each&nbsp;<code>.csv</code>-file is&nbsp;<strong>semi-colon</strong>-separated, and has&nbsp;<strong>4&nbsp;</strong>columns, <strong>header row</strong>, and rows that describes each unit-to-unit comparison. The columns contain the following information (in <strong>rtl</strong> order):<br> <ul> <li><code>[Name of the first compared lect]</code> : unit (character 3-shingle, or just 3-shingle) of the [name of the first compared lect] that undergoes comparison with units of the [name of the second compared lect]; datatype: string.</li> <li><code>[Name of the second compared lect] </code>: unit (character 3-shingle, or just 3-shingle) of the [name of the second compared lect] that undergoes comparison with units of the [name of the first compared lect]; if units coincide, contains value <code>id.</code>; datatype: string.</li> <li><code>[Experiment setup]</code>: name of the metric, a combination of the [experiment setup](concatenated through <code><em>-</em></code> parameters) and its exact part, which compares the two units; datatype: string. The possible values are: <ul> <li><code>[experiment setup] - DistRank</code> - the frequency-based metric that compares identical units</li> <li><code>[experiment setup] - hybrid</code> - the string similarity measure for non-identical units, combined with the frequency-based metric&nbsp;</li> </ul> </li> <li><code>Distance</code>: value of the metric; datatype: float</li> </ul> </li> <li><code>.info</code>-file that contains data on branch lengths along with coincidence/non-coincidence of the detected outgroup with the manually defined one. The file is a&nbsp;<strong>tabular-separated plain text</strong> that always contains three values: coincidence (CORRECT)/non-coincidence (INCORRECT) of the yielded classification with the supposed one; outer distance split (the length of the outgroup branch; datatype: float) and inner distance split (the length of the ingroup branch before split of its lects; datatype: float).</li> <li><code>.newick</code> -file that contains the result of an experiment, the phylogenetic tree built by UPGMA classifier. One can read it with <a href="https://cran.r-project.org/web/packages/TreeTools/vignettes/load-trees.html">ape::read.tree</a> (R), or <a href="https://biopython.org/wiki/Phylo">Phylo.read</a> (Python).</li> <li><code>.png</code> -file that contains the phylogenetic tree visualisation.</li> </ul> <p>&nbsp;</p> <h1>How-to</h1> <p>For the analysis of the results, download and unpack the archive, and further&nbsp;refer to the <a href="https://doi.org/10.5281/zenodo.14169792" target="_blank" rel="noopener">companion R notebook</a>.&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo28/100

Barcelona Copper Plate Charter

<p>Barcelona Copper Plate Charter</p>

opencc-by-4.0Feb 2019View details →
zenodo28/100

Barcelona Copper Plate Charter

<p>Barcelona Copper Plate Charter, plate 2 recto</p>

opencc-by-4.0Feb 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record