Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

24

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

24 results for “digital editions”

Learn how ShareScore rates datasets ↗
zenodo44/100

Towards a generic processingand presentation ofTEI encoded digital editions

<p>This dataset is the basis for the talk given at the TEI Member&#39;s Meeting 2014, Evanston, IL</p> <p>The abstract of the paper submitted:</p> <p>The set of XSL stylesheets provided and maintained by the TEI is relatively cautious about processing and presentation of transcriptions and content of digital editions. Only very basic functions are implemented, such as to surround abbreviations with brackets or to process from the element choice in plain mode only those children that represent the &quot;critical&quot; reading. Processing in plain mode means that the elements will be treated as in-line elements.[1]</p> <p>On the other hand the encoding must have been done with a special purpose. A general rule of text encoding is that the editor may encode only those structures and semantic features that he wants to process in the end. The processing might include elaborated examination and analysis of the encoded text or more complex queries as well as a reproduction of visual properties of the original document or provide a (simplified) reading text. Thus the encoding will tell something about the functionalities of the text in processing and presentation.</p> <p>According to Patrick Sahle&#39;s &quot;Textrad&quot;[2], &quot;the&quot; text does not exist in a transcription but the encoded text usually represents multiple properties and serves multiple purposes. Whatever the editor might state in some introductory notes and the documentation of the edition which should contain some statements about the encoding used, in the end the encoded text will speak on its own, can be interpreted and will be processed as is.</p> <p>In succession of the modelling of the TEI, realised in the modules, the grouping of elements and of attributes, the semantics of certain elements might let the processor estimate about the foreseen presentation, processing, and use:<br> - The elements &lt;pb&gt;, &lt;lb&gt;, &lt;l&gt;, &lt;lg&gt;, etc. as well as attributes @rend, @rendition or @style represent visual aspects of the text, therefore these might have to be reproduced; users may be given a choice to either see a document-centred view which visualises these aspects or switch to an editorial view on the text which eliminates these properties.<br> - The same applies to the element &lt;choice&gt;: If this is used the editor must have had in mind the opportunity to change the views on the document respectively encoded text.<br> - Entities encoded as &lt;rs&gt;, &lt;name&gt;, &lt;persName&gt;, &lt;placeName&gt;, etc might be referenced, especially if they are accompanied by the related list elements such as &lt;listPerson&gt;, &lt;listPlace&gt;, etc. Additionally, one might assume that there will be norm data available which allows for links into the open.<br> - Bibliographic records (&lt;bibl&gt;, &lt;msDesc&gt;) will serve a similar purpose and will have to be referenced.<br> - Quotations like &lt;cit&gt;, &lt;foreign&gt;, &lt;q&gt;, &lt;quote&gt; etc. will have to be distinguished from the surrounding text.</p> <p>Concerning the overall structure of an (critical) edition one might expect up to three apparatuses: The critical apparatus, the commentary and maybe a bibliographical apparatus. How many of these are present in a given edition is up to the editor but the presence of certain elements and especially of the amount of certain elements will give anybody an idea of how many apparatuses are &quot;appropriate&quot;: If the encoding contains editorial elements like &lt;choice&gt;, &lt;abbr&gt;/&lt;expan&gt;, &lt;add&gt;/&lt;del&gt;, etc. the representation of this information in an apparatus will be inevitable. If a certain amount of bibliographic references point to biblical or classical texts the tradition of the publication of editions has provided a separate apparatus as well. Last, editorial notes will have to be distinguished from the former two categories.</p> <p>This paper will examine existing editions with statistical methods and by clustering the elements used it might be possible to assign the encoded text to one or more text types of Sahle&#39;s typology. Additionally, the paper shall foster the discussion about the presentation of an edited text according to the intended purpose of the encoding. On the basis of the typology and purpose of the edition it will be more likely that a generic presentation of any edited text is possible. Finally, with some statistical data about the editions some remarks about the interoperability of the TEI-encoded texts shall be possible.</p> <p>[1] e.g. https://github.com/TEIC/Stylesheets/blob/master/html/html_core.xsl</p> <p>[2] Patrick Sahle: Digitale Editionsformen, 3 vols. 2013, esp. vol. 3, p. 9ff.</p>

opencc-by-4.0Jul 2020View details →
zenodo44/100

An open, collaborative, and scholarly digital edition of Anṭūn al-Jumayyil's monthly journal "al-Zuhūr" (Cairo, 1910--1913)

<p>This release has been necessitated by the need for documenting changes and improvements that took place over the last two years:</p> <ol> <li>New sets of facsimiles have been added using IIIF</li> <li>The TEI Boilerplate has been updated to the latest version</li> <li>The mark-up of entities and their links to our authority files have been much improved</li> </ol>

opencc-by-sa-4.0Jun 2021View details →
zenodo44/100

The e-NDP project : collaborative digital edition of the Chapter registers of Notre-Dame of Paris (1326-1504). Ground-truth for handwriting text recognition (HTR) on late medieval manuscripts.

<p>The <a href="https://endp.hypotheses.org/">e-NDP project</a>, funded by the ANR, is led by the <a href="https://lamop.hypotheses.org/6870">LaMOP</a> (Julie Claustre and Darwin Smith).</p> <p>The project&#39;s partners are the Archives nationales, the&nbsp;Biblioth&egrave;que nationale de France (Department of Manuscripts, Biblioth&egrave;que de l&#39;Arsenal), the &Eacute;cole nationale des chartes and the Biblioth&egrave;que Mazarine.</p> <p>The e-NDP project aims at renewing our knowledge on <strong>Notre-Dame de Paris cathedral</strong> through the creation of a collaborative digital edition of the registers of its Chapter (1326-1504, <em>AN LL 105-128</em>), the community of 51 canons meeting three times a week on set days to take all administrative, financial and practical decisions pertaining to the cathedral, its estate and the society living in its cloister. This corpus has never been the object of a comprehensive study to understand the workings and history of this urban enclave and powerful community. The collaborative digital edition is based on a process of<strong> handwriting text recognition (HTR)</strong>, tested and supervised by scholars, researchers and engineers combining expertise in Medieval history, paleography, philology and digital humanities. The edition shall allow a better insight into the Chapter&rsquo;s administration, into its economical and political power within Paris, and the relationships it maintained with other institutions in the city.</p> <p>&nbsp;</p> <p><strong>Section 1 : The e-NDP ground-truth dataset for Handwriting text recognition.</strong></p> <p>The full e-NDP corpus kept today in the French National Archives and was entirely digitized and described in its&nbsp;<a href="https://www.siv.archives-nationales.culture.gouv.fr/siv/rechercheconsultation/consultation/ir/consultationIR.action?formCaller=GENERALISTE&amp;irId=FRAN_IR_059635">catalog</a>&nbsp;in 2022.</p> <p>The first major goal of the&nbsp;e-NDP projet is to propose a first automatic transcription of the 14k pages composing the 26 chapter registers. To achieve this goal representative samples from&nbsp;each one of the volumes were selected and transcribed in order to train a specialized HTR model able to propose a high quality automatic transcription. The collected ground-truth released on this repository currently has <strong>512 pages from the 26 registers</strong> of the cathedral chapter preserved in the National Archives (LL105 - LL128, <strong>1326-1504</strong>). The transcriptions were manually completed in <strong>two rounds</strong> by a group of 12 contributors, historians and paleographers, over the course of 2021-2022 using <a href="https://escriptorium.paris.inria.fr/">eScriptorium </a>as annotation environment.&nbsp;&nbsp;</p> <p>&nbsp;</p> <p><strong>Ground-truth features :</strong></p> <p><br> <em>Number of hands </em>: according to our estimates no fewer than 18&nbsp;main hands were involved in the writing of the registers during the medieval period.&nbsp;</p> <p><em>Language</em> : More than 98% of the content of the registers was written in Latin, the rest in French. The exact percentage is hard to estimate because the vernacular language is often used in formulae, notes and comments. It is rare to find entire pages or blocks written in French.&nbsp;</p> <p><em>Script family</em> : The registers were written using a Cursive script (ca. late XIIIe - XVIe).</p> <p><em>Documental typology</em> : The volumes containing the chapter conclusions were conceived to serve&nbsp;as memorial&nbsp;records, but above all as documents for regular use and consultation in the daily practice of administration and management. In diplomatics the notion of &quot;documentary manuscripts&quot; is used to describe this kind of sources&nbsp;also by opposition to books and litterary or&nbsp;normative&nbsp;manuscripts.</p> <table align="center"> <caption><strong>Ground truth statistics</strong></caption> <tbody> <tr> <th>Text units</th> <th>Count</th> </tr> <tr> <td>Pages</td> <td>512</td> </tr> <tr> <td>Annotated regions (see section 2)</td> <td>2448</td> </tr> <tr> <td>Lines of text</td> <td>34231</td> </tr> <tr> <td>Tokens</td> <td>205083</td> </tr> <tr> <td>Characters</td> <td>3320407</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Rules of transcription :</strong></p> <ul> <li>The abbreviations have been resolved, both those by suspension (<code>facimꝰ</code> ---&gt; <code>facimus</code>) and by contraction (<code>d&ntilde;i</code> --&gt; <code>domini</code>). Likewise, those using conventional signs (<code>⁊</code> --&gt; <code>et</code> ; <code>ꝓ</code> --&gt; <code>pro</code>) have been resolved.&nbsp;</li> <li>The named entities (names of persons, places and institutions) have been <code>capitalized</code>. The beginning of a block of text as well as the original capitals used by the notary are also capitalized.</li> <li>The consonantal <code>i</code> and <code>u</code> characters have been transcribed as <code>j</code> and <code>v</code> in both French and Latin.</li> <li>The punctuation marks used in the text: <code>.</code> and <code>/</code> have been transcribed, but the transcription has not been standardized with modern punctuation.</li> <li>Corrections and words that appear cancelled in the manuscript have been transcribed surrounded by the sign <code>$</code> at the beginning and at the end.</li> <li>More specific transcription rules can be found into the file <code>transcription_guidelines.pdf</code></li> </ul> <p>&nbsp;</p> <p><strong>Section 2. e-NDP Layout Segmentation.</strong></p> <p>Layout segmentation is a compulsory step before HTR recognition in order to distinguish sections and regions inside a document. This process intend to separate interdependant page zones to produce a recognition in a section-sequence order and not in a line-sequence order which mix textual and peri-textual content.</p> <p>The regions of 364&nbsp;pages (see <code>GT-layout_list</code>) of the e-NDP corpus were annotated using a 5 sections vocabulary (see <code>endp_layout_regions</code>) in order to describe&nbsp;the page distribution in all the 26 volumes :</p> <ol> <li><em>Block</em>&nbsp;: All the central text blocks, that normally corresponds to the main content called &quot;conclusions&quot; in registers.</li> <li><em>Liste</em>&nbsp;: List of names of the canons who were present during the meeting. Normally located before the <em>conclusions</em>.</li> <li><em>Entr&eacute;e</em>&nbsp;: Marginal notes or entries to inform about the content of <em>conclusions</em>.</li> <li><em>Date</em>&nbsp;: Paragraph contending the date. Normally at the head of a <em>conclusion</em>, but separate of the main body.</li> <li><em>Num&eacute;rotation</em>&nbsp;: Page numbers in roman or arabic. Usually appear in the top corners of the pages.</li> </ol> <table align="center"> <caption><strong>Layout GT statistics</strong></caption> <tbody> <tr> <th>Region</th> <th>Count</th> </tr> <tr> <td>block</td> <td>833</td> </tr> <tr> <td>liste</td> <td>431</td> </tr> <tr> <td>date</td> <td>448</td> </tr> <tr> <td>entr&eacute;e</td> <td>205</td> </tr> <tr> <td>num&eacute;rotation</td> <td>531</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Section 3. The e-NDP HTR modeling.</strong></p> <p>The e-NDP project has progressively trained several HTR models adapted to work on late medieval cursive in order to accelerate the production of ground truth. Currently the best model delivers an average&nbsp;<strong>CER (Character error ratio) of 9.7%</strong> in handwriting recognition on&nbsp;the 26 registers (see <code>endp_learning_curve</code>) and can serve as generalist model&nbsp;for other manuscripts of the same period and similar script family. These models and their training implementation details can be found in the project&#39;s github <a href="https://github.com/chartes/e-NDP_HTR">repository</a>.&nbsp;</p> <p>Additionally, the automatic HTR transcriptions of the 26 registers (14k pages, 4.5M tokens) enriched with lexical and semantical information has been the subject of a first <a href="https://nosketch-engine.lamop.fr/#dashboard?corpname=endp">online publication</a> using the NoSketch engine that allows advanced data mining based on the combination of data, metadata and NLP features.&nbsp;</p> <p>&nbsp;</p> <p><strong>Section 4. Dataset content.</strong></p> <p>This zip dataset contains :</p> <p>- <code>HTR_ground_truth</code> : Two folders containing the jpg / jpeg images and their curated transcriptions in PAGE XML format.</p> <p>- <code>images_docs</code> : 4 files illustrating the different phases of the project (list of GT for layout segmentation, layout ontologie, transcription guideline and HTR evaluation curves)</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Catalogue of tools for digital scholarly editing

<p>This dataset is <strong>a catalogue of software products and computer tools that can be used to create digital scholarly editions (DSE)</strong>. The catalogue includes both tools specifically designed for use in the field of digital philology, mostly created by scholars and researchers in the field, and general-purpose tools that have been widely adopted in this field.</p> <p>For each tool, the following information is presented:</p> <ul> <li> <p>&ldquo;<strong>Name</strong>&rdquo;: the name of the tool.</p> </li> <li> <p>&ldquo;<strong>Website</strong>&rdquo;: the address of the official website, repository, or wiki that the creators of the tool offer as a showcase to learn about the tool.</p> </li> <li> <p>&ldquo;<strong>Category</strong>&rdquo;: one of the three macro-categories identified during the creation of the catalogue: &ldquo;Visualization&rdquo;, tools intended exclusively for the visualization (or publication) of a DSE; &ldquo;Mixed&rdquo;, tools that can be used for both processing and publishing a DSE; and finally, &ldquo;Production&rdquo;, tools that can replace or complement the philologist in a particular phase of the editorial process. The distinction between the proposed categories is not always clear-cut; tools have generally been categorized based on their main functionalities and the type of output or result they allow to obtain. A tool is categorized as either for visualization or mixed when it allows the production of a result that can be defined as a digital scholarly edition.</p> </li> <li> <p>&ldquo;<strong>TaDiRAH activity</strong>&rdquo;: one or more of the activities listed in the <a href="https://vocabs.dariah.eu/tadirah/en/">TaDiRAH taxonomy</a> that best describe the main functionalities of the tools.</p> </li> <li> <p>&ldquo;<strong>Begin date</strong>&rdquo;: the date in which the first public version of the tool was released. In cases where the tool was produced experimentally without reaching an official first version, the date of creation of the respective source code repository or the date indicated by the creator as the creation date on the website or in a publication is considered. When only the year is available the date is set on January first of that year. Respectively, when only the year and month are available, the date is set on the first day of the month.</p> </li> <li> <p>&ldquo;<strong>End date</strong>&rdquo;: the date in which the development of the tool was discontinued. When only the year is available the date is set on January first of that year. Respectively, when only the year and month are available, the date is set on the first day of the month.</p> </li> <li> <p>&ldquo;<strong>Description</strong>&rdquo;: a more or less brief text used on the website, repository, or wiki of the tool to present it to potential users. It is a more &ldquo;commercial&rdquo; than technical text, as its purpose is to encourage the use of the tool, and it briefly describes the main features and characteristics of the tool, with references to partner organizations. The description is included because it is interesting to analyze the words with which the creator presents their tool.</p> </li> <li> <p>&ldquo;<strong>Institutional Partner(s)</strong>&rdquo;: organizations that financially contributed to the development of the tool. These can be universities, research centers, cultural institutions, public projects, and funds. Since many of these funds are time-limited, I chose to include this information as it helps understand how stable a tool is in terms of development, maintenance, and user support.</p> </li> <li> <p>&ldquo;<strong>Creator</strong>&rdquo;: the name of the creator or creators of the tool. In the case of commercial software, the name of the company that produced it is provided. Similarly, if the tool was produced by a research center, a collective, or another type of group with an official name, the group&rsquo;s name is provided. If the tool was created by one or more scholars, their respective names and affiliations are indicated.</p> </li> <li> <p>&ldquo;<strong>Input Format(s)</strong>&rdquo;: the various formats accepted by the tool as input. In some cases, it is not possible to obtain exact formats, as official sources indicate &ldquo;text&rdquo; or &ldquo;image&rdquo; generically. In these cases, the most commonly used formats, which are likely supported, are listed: TXT for text and JPEG for images.</p> </li> <li> <p>&ldquo;<strong>Output Format(s)</strong>&rdquo;: the output formats in which the tool allows exporting the DSE or other types of data. It should be noted that some tools, especially among mixed or visualization-only ones, do not provide data export.</p> </li> <li> <p>&ldquo;<strong>Technologies</strong>&rdquo;: programming languages, tools, libraries, and frameworks used to develop the tool. Some of the listed tools have been integrated into other tools (for example, OpenSeadragon and VisColl have been integrated into EVT).</p> </li> <li> <p>&ldquo;<strong>System requirements</strong>&rdquo;: indicates the nature of the tool, such as desktop software, a web-based service, a library, a web-based application, etc., and system requirements (e.g. compatible operating systems).</p> </li> <li> <p>&ldquo;<strong>Collaborative working</strong>&quot;: indicates whether the tool is designed for collaborative work, allowing multiple users to work simultaneously on the same materials.</p> </li> <li> <p>&ldquo;<strong>Open Source</strong>&rdquo;: indicates whether the source code of the tool is freely accessible.</p> </li> <li> <p>&ldquo;<strong>Repository</strong>&rdquo;: the address of the repository where the source code of the tool is stored.</p> </li> <li> <p>&ldquo;<strong>License</strong>&rdquo;: the license under which the tool is available. If the tool is available for a fee, it is indicated as &ldquo;for a fee&rdquo;. If the tool is free but the exact open-source license is not specified, the license is simply marked as &ldquo;Free&rdquo;.</p> </li> <li> <p>&ldquo;<strong>Current Version</strong>&rdquo;: the number of the latest publicly released version of the tool.</p> </li> <li> <p>&ldquo;<strong>Editions</strong>&rdquo;: the titles of DSEs that have been created using the tool. This field is very useful for potential user-editors as it allows them to see concretely how a DSE is presented thanks to the tool and/or what scientific results the tool enables. The names are hyperlinks to the respective official websites of the DSEs.</p> </li> <li> <p>&ldquo;<strong>Publications</strong>&rdquo;: bibliographic references to publications in which the creators present their tool or other scholars review or report their experience using the software with a concrete use case.</p> </li> </ul> <p>For many tools, it was not possible to provide all the above-listed information, either because it was not available in official sources (primarily websites and repositories) or because it was not possible to identify updated and reliable bibliographic sources. The omission of information does not imply that it cannot be obtained through further in-depth study or by contacting the creators themselves. Furthermore, the information may no longer be up-to-date or valid; for example, hyperlinks to websites may no longer be active.</p> <p>The catalogue is not exhaustive, as it mainly includes tools produced in Italy, Europe, and the United States. The main objective of the work is to provide an initial overview of tools for digital philology, albeit partial and incomplete, and to propose useful categories for analyzing and classifying these tools from the perspective of user-editors who wish to create an DSE. This review constitutes the starting point for a larger research project that I would like to undertake in the future, namely, creating a catalog of computer tools for digital philology.</p> <p>I would like to thank Professors Roberto Rosselli Del Turco and Lino Leonardi for providing references to several of the analyzed tools.</p> <p>The list is also <a href="https://airtable.com/appmd2Z1LsaYMaYEF/shrb1dVa8QsJ4P8Ga">available for consultation online</a>.</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

TEI-XML-Datenset der Tagebücher, Briefe, Dokumente, Forschungsbeiträge, Chronologieeinträge und Register der edition humboldt digital

<p>Das Datenset enth&auml;lt alle edierten Texte (Tageb&uuml;cher, Briefe und weitere Dokumente) sowie Paratexte (Forschungsbeitr&auml;ge, Eintr&auml;ge der Chronologie zu Alexander von Humboldts Leben, Register und Glossar) der Version 11 der <a href="https://edition-humboldt.de">edition humboldt digital</a>, die am 4. Juni 2025 erschienen ist. Das Datenset enth&auml;lt gegen&uuml;ber der HTML-Version technische Fehlerkorrekturen, daher wird es als Version 11.0.1 ver&ouml;ffentlicht.</p> <p>Die Editionsrichtlinien stehen auf <a href="https://edition-humboldt.de/richtlinien/index.html">edition-humboldt.de</a> zur Vef&uuml;gung. Das Datenmodell ist in drei verschiedene ODDs aufgeteilt (f&uuml;r edierte Texte, Registereintr&auml;ge und Forschungsbeitr&auml;ge). Dem Datenset liegen die drei RNG-Schemata bei, die ODD-Ursprungsdateien sind im GitHub-Repository <a href="https://github.com/telota/ediarum.AVHR.data-model/">ediarum.AVHR.data-model</a> zu finden. Beachten Sie bitte, dass es f&uuml;r das Pflanzenregister derzeit noch kein Schema gibt, da dieses aus dem Tagging automatisch erstellt wird.</p> <p>Weitere Hinweise zur digitalen Methodik finden sich in <a href="https://edition-humboldt.de/H0016212">Dumont 2024</a> und zum Editionsvorhaben im Allgemeinen in <a href="https://doi.org/10.25365/wdr-01-03-02">Kraft/Dumont 2020</a>.</p> <p>Dieses Datenset ist auch auf <a href="https://github.com/telota/edition-humboldt-digital">GitHub</a> zug&auml;nglich.</p>

opencc-by-sa-4.0Jul 2023View details →
zenodo40/100

An open, collaborative, and scholarly digital edition of Anastās Mārī al-Karmalī's monthly journal "Lughat al-ʿArab" (Baghdad, 1911--14)

<p>This repository has seen a lot of edits since the last release:</p> <p>changes</p> <ul> <li>moved to central TEI boilerplate</li> <li>switched interface to Arabic</li> </ul> <p>edits</p> <ul> <li>added mastheads to vol.s 1 and 2</li> <li>some validation of automated structural mark-up</li> <li>manual validation of mark-up against the facsimile: vol. 1</li> <li>removed faulty automated mark-up of dates and periodical titles and re-added them through much improved process</li> <li>mark-up of named entities</li> <li>linked entity names to authority files</li> </ul>

openother-openFeb 2022View details →
zenodo40/100

An open, collaborative, and scholarly digital edition of ʿAbd al-Qādir al-Iskandarānī's monthly journal "al-Ḥaqāʾiq" (Damascus, 1910--12)

<p>The last release, v0.9, wasn&#39;t caught by Zenodo for some unknown reason. Nothing has been changed since. The following release message has been copied from v0.9:</p> <p>As the mark-up of al-Ḥaqāʾiq has been complete for some time now and we are waiting for our scans to be processed at Halle for some three years now, we have decided to release a practically complete version. There have been a couple of maintenance edits over the last years to unify the set-up of editions across the entire project. Thus, the folder with TEI files was renamed tei/. Entity linked as also considerably improved, particularly for periodical titles.</p>

openother-openFeb 2022View details →
zenodo40/100

Data of Digital Scholarly Edition of the Diary of the Travel of Heinrich XI. Reuß-Greiz 1740-1742

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
zenodo40/100

Overview of life and death of digital scholarly editions

<p>The dataset is based on the online catalogue for Digital Scholarly Editions (<a href="http://www.digitale-edition.de/">http://www.digitale-edition.de/</a>) on 7/8/2019. Based on the links provided in the catalogue, we queried the Wayback Machine (https://archive.org/web/) for first seen and last seen version of the website and added this information to the dataset.&nbsp;</p>

opencc-by-4.0Oct 2019View details →
zenodo40/100

Interviews on the future of digital editing and publishing

<p>Transcriptions of 46 interviews with theorists and practitioners in the field of digital scholarly editing on topics relating to the future of digital scholarly editions</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Knowledge graphs for interoperable NFDI: Digital editions

<p>Pr&auml;sentation im Rahmen des Text+ FAIR February Meetup 2023: <em>I wie Interoperability, </em>15. Februar 2023.</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

A Hybrid Focus Group for the Evaluation of Digital Scholarly Editions of Literary Authors

<p>Data from focus group conducted for DiXiT - Digital Scholarly Editions Initial Training Network. </p> <p>We are studying the use of digital scholarly editions, by observing what users do in interaction with different editions. For this analysis, we chose three digital scholarly editions, designed tasks specific to these editions, confronted typical users with these tasks, and analysed the quality of use in relation to efficiency and effectiveness.</p> <p>The research leading to these results has received funding from the People Programme (Marie Curie Actions) of the European Union's Seventh Framework Programme FP7/2007-2013/ under REA grant agreement n° 317436.</p>

opencc-by-4.0Jun 2017View details →
zenodo36/100

Digital Scholarly Edition of the 1761 Library Catalogue of the Paderborn Capuchins

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2022View details →
zenodo36/100

A digital edition of the 389 Chorale Settings (Choralgesänge) by J.S. Bach

<p>This repository has been made possible by Gertim Alberda, who 'manually' digitalized all the 389 Bach 4-voice chorales, as published by the <a href="https://imslp.org/wiki/Special:ReverseLookup/348824">Breitkopf edition &#39;nr. 3765&#39;</a>. For this he meticulously transcribed all the notes and lyrics of that edition, by using the music notation editor MuseScore (MS). He checked it against both BGA and NBA (Bach-Gesellschaft Ausgabe and Neue Bach-Ausgabe), if there was any reasonable doubt for it, and even found and improved a couple of mistakes that way (&lt;10). Solely for performance reasons, he also applied many hidden (grey) extras in MS to make the scores and its phrasing sound as realistic as possible; hidden fermatas, tempo changes, breath pauses, phrasing, note-cutbacks, etc. This was all done to &#39;humanize&#39; the playback and make it sound like a real choir performance (without words), and given the technical limitations at that time (2016-2018). Also see the <a href="https://gertim-alberda.com/chorales/info.html">Info</a> tab on his website.</p> <p>The full corpus resides on Gertim Alberda's <a href="https://gertim-alberda.com/chorales">dedicated website</a>. The scores there have synchronized play back (synthesized and human performances) available and can be downloaded in various formats (mscz/xml/midi/mp3/pdf). <a href="https://gertim-alberda.com/chorales/BachChorales/B288.html">Here is an example</a> of his playback page. His intention is to have only human performances available (YT videos and/or mp3's) for all the scores, but that is still an ongoing process (see the <a href="https://gertim-alberda.com/chorales/changelog_bach_chorales.html">changelog</a> on his website). He has agreed to share his MS files under a <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">Creative Commons BY-NC-SA 4.0</a> license. The license prohibits the use of the scores for any commercial purpose, including the training of machine learning models for use in commercial products.</p> <h2>Contents</h2> <p>The &quot;mother of all datasets&quot; is available as two versions, each contained in a separate folder:</p> <ul> <li><code>original_complete</code>: Gertim Alberda&#39;s original MuseScore files converted to uncompressed MuseScore 3 format (<code>.mscx</code>)</li> <li><code>vocal_parts_only</code>: Alternative version of the dataset containing no instrumental parts.</li> </ul> <h2>Version history</h2> <p>See the <a href="https://github.com/johentsch/389_chorale_settings/releases">GitHub releases</a>.</p> <h2>Questions, Suggestions, Corrections, Bug Reports</h2> <p>Please <a href="https://github.com/johentsch/389_chorale_settings/issues">create an issue</a> and/or feel free to fork and submit pull requests.</p> <h2>License</h2> <p>Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">CC BY-NC-SA 4.0</a>).</p>

opencc-by-nc-sa-4.0May 2024View details →
zenodo32/100

Digitally edited images.

<p>Digitally edited images of male and female models.</p>

opencc-by-4.0Jul 2023View details →
zenodo20/100

Digitized, corrected and supplemented data on net zooplankton of the Far East seas and adjacent Pacific Ocean waters from five rare reference books published in Russian in limited editions

<p>Data on plankton collected by the Juday net in 1984-2013 in the Far Eastern seas and the north-western Pacific containing the species composition, occurrence and abundance of zooplankton in the surveyed area. The data is aggregated by species, developmental stages, size fractions, regions, vertical layers of water, light and dark time of day, four seasons of the year and perennial periods. It is accompanied shape-files with polygons of the regions by which data is summarized with information about surface areas and water volumes in each of them. The scope of application of this data is fundamental to the management of marine resources, aquaculture development, nature conservation, and assessment of the damage of various anthropogenic factors on nature.</p>

restrictedJan 2021View details →
nasa20/100

IHW COMET SPEC EDITED DIGITIZED IMAGE RECORD CROMMELIN V1.0

In preparation for the concerted international study of Comet Halley, the IHW conducted a trial run with observations of Comet Crommelin, largely during February and March of 1984.

restrictedus-pdMar 2025View details →
nasa20/100

IHW COMET SPEC EDITED DIGITALIZED IMAGE DATA RECORD GZ V1.0

In preparation for the concerted international study of Comet Halley, the IHW conducted a trial run with observations of Comet Crommelin, largely during February and March of 1984.

restrictedus-pdApr 2025View details →
nasa20/100

IHW COMET LSPN EDITED DIGITALIZED IMAGE DATA RECORD GZ V1.0

In preparation for the concerted international study of Comet Halley, the IHW conducted a trial run with observations of Comet Crommelin, largely during February and March of 1984.

restrictedus-pdMar 2025View details →
dryad16/100

Incipits for the Catalogue of the Works of Hector Berlioz, Second edition, digital

Open the record for dataset details and reuse information.

publicJul 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record