Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
266
datasets available to search
ShareScore release 0.9.0
Dataset results
266 results for “xml”
IN02067 Inscription of Bhimarjuna and Visnugupta at Yengahiti. Sanskrit XML file, draft epidoc edition
<p>IN02067 Inscription of Bhimarjuna and Visnugupta at Yengahiti. Sanskrit XML file (without metadata). Draft epidoc edition to be incorporated into 'Siddham' archive</p>
IN02080 Yengu Bahaltole Inscription. Sanskrit XML file, draft epidoc edition
<p>IN02080 Yengu Bahaltole Inscription. Sanskrit XML file (without metadata). Draft epidoc edition to be incorporated into 'Siddham' archive</p>
IN02069 Tebahal Stone Inscription. Sanskrit XML file, draft epidoc edition
<p>IN02069 Tebahal Stone Inscription. Sanskrit XML file (without metadata). Draft epidoc edition to be incorporated into 'Siddham' archive</p>
IN02081 Sanku Fragment Inscription (revise title). Sanskrit XML file, draft epidoc edition
<p>IN02081 Sanku Fragment Inscription (revise title). Sanskrit XML file (without metadata). Draft epidoc edition to be incorporated into 'Siddham' archive</p>
Inundation maps of Donana for 23 dates within the period 2015/12/19 to 2017/08/20 and their accompanying INSPIRE metadata XML files
<p>Satellite-derived inundation maps offer an efficient solution for monitoring the spatial and temporal variability of the hydrological cycle of wetlands. This task is important for taking mitigation actions against factors (e.g. climate change and human pressures) threatening wetlands' functions and services.</p> <p>Inundation maps within the period 2015/12/19 to 2017/08/20 were generated for Donana based on the methodology presented in "Kordelas, G.A.; Manakos, I.; Aragonés, D.; Díaz-Delgado, R.; Bustamante, J. Fast and Automatic Data-Driven Thresholding for Inundation Mapping with Sentinel-2 Data. <em>Remote Sens.</em> <strong>2018</strong>, <em>10</em>, 910.".</p> <p>Each inundation map is named as " 'Date'_inundation_map_Donana_S2.tif ", and contains the following classes: Inundated Class, Non-inundated Class. In this map, Inundated and Non-inundated Classes are denoted with 0 and 1, respectively. 'Date' is in the form YYYY_MM_DD.</p>
Inundation maps of Danube Delta for 10 dates within the period 2016/10/05 to 2017/08/01 and their accompanying INSPIRE metadata XML files
<p>Satellite-derived inundation maps offer an efficient solution for monitoring the spatial and temporal variability of the hydrological cycle of wetlands. This task is important for taking mitigation actions against factors (e.g. climate change and human pressures) threatening wetlands' functions and services.</p> <p>Inundation maps within the period 2016/10/05 to 2017/08/01 were generated for Danube Delta based on the methodology presented in "Kordelas, G.A.; Manakos, I.; Aragonés, D.; Díaz-Delgado, R.; Bustamante, J. Fast and Automatic Data-Driven Thresholding for Inundation Mapping with Sentinel-2 Data. <em>Remote Sens.</em> <strong>2018</strong>, <em>10</em>, 910.".</p> <p>Each inundation map is named as " 'Date'_inundation_map_Danube_Delta_S2.tif ", and contains the following classes: Inundated Class, Non-inundated Class. In this map, Inundated and Non-inundated Classes are denoted with 0 and 1, respectively. The regions, which are manually denoted as affected by clouds, are denoted with 2. 'Date' is in the form YYYY_MM_DD.</p>
Inundation maps of Camargue for 47 dates within the period 2016/02/09 to 2018/06/19 and their accompanying INSPIRE metadata XML files
<p>Satellite-derived inundation maps offer an efficient solution for monitoring the spatial and temporal variability of the hydrological cycle of wetlands. This task is important for taking mitigation actions against factors (e.g. climate change and human pressures) threatening wetlands' functions and services.</p> <p>Inundation maps within the period 2016/02/09 to 2018/06/19 were generated for Camargue based on the methodology presented in "Kordelas, G.A.; Manakos, I.; Aragonés, D.; Díaz-Delgado, R.; Bustamante, J. Fast and Automatic Data-Driven Thresholding for Inundation Mapping with Sentinel-2 Data. <em>Remote Sens.</em> <strong>2018</strong>, <em>10</em>, 910.".</p> <p>Each inundation map is named as " 'Date'_inundation_map_Camargue_S2.tif ", and contains the following classes: Inundated Class, Non-inundated Class. In this map, Inundated and Non-inundated Classes are denoted with 0 and 1, respectively. 'Date' is in the form YYYY_MM_DD.</p>
AquaMaps: AquaMaps XML resource
from <p></p>http://www.aquamaps.org/. AquaMaps are computer-generated predictions of natural occurrence of marine species, based on the environmental tolerance of a given species with respect to depth, salinity, temperature, primary productivity, and its association with sea ice or coastal areas. These __environmental envelopes__ are matched against an authority file which contains respective information for the Oceans of the World. Independent knowledge such as distribution by FAO areas or bounding boxes are used to avoid mapping species in areas that contain suitable habitat, but are not occupied by the species. Maps show the color-coded likelihood of a species to occur in a half-degree cell, with about 50 km side length near the equator. Experts are able to review, modify and approve maps.<p></p>from EOL v2 database
AnAge: AnAge text (XML resource)
AnAge is a database of longevity and ageing in animals. It features quantitative life history data for over 4,000 species, including extensive longevity records, body masses at different developmental stages, reproductive data, and physiological traits related to metabolism. In addition to quantitative data, AnAge also features comments and observations related to ageing or relevant to the life history of individual taxa. AnAge features a manually-curated collection of animal longevity records. Moreover, AnAge has extensive life-history traits such as adult body weight, gestation or incubation time, age at sexual maturity and other reproductive data. Lastly, observations on physiological or pathological changes with age in animals are (where available) featured. AnAge focuses primarily on chordates. At the time of writing, AnAge features 4,122 entries, including 1,331 mammals, 1,098 birds, 539 reptiles, 169 amphibians, 962 fishes and 28 non-chordates. Our focus is on accuracy and quality, however, not quantity, and only species for which we have confidence in the data are featured. Numerous experts have contributed information to AnAge and helped us meet quality standards. Professor Steven Austad, a world-renowned expert in mammalian ageing at the Barshop Institute in San Antonio, is AnAge__s expert mammalogist and curator.<p></p>
AskNature: AskNature XML
Open the record for dataset details and reuse information.
TEI-XML Zürcher Regierungsratsbeschlüsse 1803-1887
<p><strong>Projekt TKR</strong></p> <p>Zwischen 2003–2016 wurden die Protokollbände des Regierungsrats und des Kantonsrats als Worddokumente seriell transkribiert und anschliessend in PDF und später als OGD-Datensätze in TEI-XML konvertiert.<br>Die Dateien werden unter Berücksichtigung der gesetzlichen Schutzfristen laufend (max. 80 Jahre) publiziert.</p> <p>Siehe auch: <a href="https://archives-quickaccess.ch/search/stazh/rrb">https://archives-quickaccess.ch/search/stazh/rrb</a></p> <p><strong>Inhalt</strong></p> <p>Dieses Datenset beinhaltet die <strong>handschriftlich verfassten Regierungsratsbeschlüsse des Kantons Zürich von 1803 bis 1887</strong>. Ab dem 1. Juli 1887 wurden die Regierungsratsbeschlüsse gedruckt. </p> <p>Die Beschlüsse des Regierungsrates gehören zu den zentralen Aktenserien des Kantons Zürich. <br>In ihnen spiegelt sich ein äusserst breites Spektrum an Themen, da der Regierungsrat nicht nur für grosse politische Entscheide, sondern oft auch für alltägliche Belange zuständig war. <br>Neben Themen wie Auswanderung oder Aufnahme von politischen Flüchtlingen, Bau von Eisenbahnen und Strassen oder Regulierung der stark ansteigenden Industrie beschäftigen den Regierungsrat stets auch Tagesgeschäfte wie Konzessionsgesuche für Tavernen oder Wasserkraftanlagen, Steuerrekurse, Einbürgerungen oder die Aufnahme von Kantonsfremden in kantonseigene Spitäler.<br>Die Metadaten eines Beschlusses (TEI/teiHeader) bestehen unter anderem aus dem Kürzel des/der Transkriptors/in, dem Transkriptionsdatum, der Signatur, dem Publikationsdatum, dem Titel des Beschlusses, dem Link zur Verzeichnung im Archivkatalog und dem Beschlussdatum.<br>Klassen und Bände bilden die hierarchische Ordnerstruktur.<br>Eine Klasse stellt jeweils den Zeitraum zwischen zwei politischen Umbrüchen dar: <br>Das Protokoll setzt 1803 mit der Bildung des modernen Kantons Zürich ein, die erste Klasse schliesst mit der Annahme der neuen restaurativen Verfassung im Juni 1814. <br>Die dritte Klasse setzt mit der liberalen Verfassung von 1831 ein; der Bruch ist auch daran erkennbar, dass der «Kleine Rat» von nun an «Regierungsrat» genannt wird. <br>Mit dem konservativen «Züriputsch» im September 1839 beginnt die vierte Klasse, welche 1849 mit einer gross angelegten Verwaltungsreform und der Integration des Kantons Zürich in den neuen schweizerischen Bundesstaat schliesst. <br>Mit der demokratischen Verfassung von 1869 setzt die sechste und letzte Klasse ein, welche nicht durch ein politisches Ereignis, sondern durch die Umstellung auf das gedruckte Protokoll endet. <br>Dies entspricht auch der Struktur im Archivverzeichnis.</p> <p><strong>Technische Erschliessung</strong></p> <p>Die Konvertierung von Worddokumenten in TEI-konforme XML-Dateien geschah mittels eines Python-Scripts, welches mittels Mustererkennung die einzelnen Datenelemente voneinander abgrenzte. </p> <p>Validierung:</p> <p>Alle Dateien sind gemäss TEI-Schema valide. Nicht valide xml-Dateien wurden manuell verbessert.</p> <p>XSL-Stylesheet:</p> <p>Das Stylesheet befindet sich im Ordner Ressourcen ("\Ressourcen\Stylesheet.xsl") und ist mit relativem Pfad eingebunden in den XML-Dateien (Zeile 2: <?xml-stylesheet type="text/xsl" href="../../Ressourcen/Stylesheet.xsl"?>). <br>Das Stylesheet ermöglicht eine Browseransicht, welche sich der originalen Transkription in Word annähert.</p> <p>Ab Version 3.0 wurden im ganzen Datensatz Eigennamen mittels maschinellem Lernen ausgezeichnet (vgl. <a href="https://github.com/machinelearningZH/named-entity-recognition_staatsarchiv">Github-Repository</a> zum Projekt). <em>Bitte beachten: </em>Die Auszeichnung wurde nicht manuell nachkontrolliert und kann Fehler enthalten!</p> <p><strong>Urheberrecht</strong></p> <p>Die Daten stehen unter einer Creative Commons CC-BY-SA 4.0 Lizenz.</p> <p>Herausgeber: Staatsarchiv des Kantons Zürich</p> <p>Technische Erschliessung: Rebekka Plüss, rebekka.pluess@zh.ch </p> <p>Projektleiter TKR: Luzi Schutz</p>
morethanbooks/XML-TEI-Bible: XML-TEI Bible: Entities and communication (66 Books)
<p>This release contains the biblical text in XML-TEI (66 books). The encoded text is in Spanish, but the codification (elements, attributes, values, ids) is in English. It makes explicit following information:</p> <ul> <li>Books, chapters, pericopes and verses.</li> <li>References to peoples, places, times, groups and books, using ids.</li> <li>Direct speech, including who is communicating, to whom and how (written, oral, prayer...).</li> </ul>
Ancient Greek Literature for Advanced Data Processing: A Text Fabric Representation of Open Access Texts in TEI XML
<p>This data set contains a full conversion of Greek texts available in the Perseus Digital Library and the Open Greek and Latin Project to the Text Fabric data format. The main advantage of the Text Fabric datatype over the original TEI XML format is that it utilizes a strict separation of text and annotation in a flat data structure. At the same time, it permits multiple distinct formats of the same text as well as an unlimited depth of (embedded) annotations. Because of its flat data structure, it facilitates easy and clean procedures to analyze, transform, and enrich the available data. Many of these processes are very difficult to conduct while departing from the hierarchically organized XML tree representation.</p>
XML_corpus
<p>All texts are from TextGrid licenced under CC-BY 3.0 (https://creativecommons.org/licenses/by/3.0/de/) and put together as a corpus by Dr. Katrin Dennerlein (http://www.germanistik.uni-wuerzburg.de/lehrstuehle/computerphilologie/mitarbeiter/dennerlein/).</p> <p>The corpus is mentioned in The Schiller-Kleist Uncertainty Principle.</p>
A word2vec model file built from the French Wikipedia XML Dump using gensim.
<p>A word2vec model file built from the French Wikipedia XML dump using gensim. The data published here includes three model files (you need all three of them in the same folder) as well as the Python script used to build the model (for documentation). The Wikipedia dump was downloaded on October 7, 2016 from https://dumps.wikimedia.org/. Before building the model, plain text was extracted from the dump. The size of that dataset is about 500 million words or 3.6 GB of plain text. The principal parameters for building the model were the following: no lemmatization was performed, tokenization was done using the "\W" regular expression (any non-word character splits tokens), and the model was built with 500 dimensions.</p>
TEI-XML-Datenset der Tagebücher, Briefe, Dokumente, Forschungsbeiträge, Chronologieeinträge und Register der edition humboldt digital
<p>Das Datenset enthält alle edierten Texte (Tagebücher, Briefe und weitere Dokumente) sowie Paratexte (Forschungsbeiträge, Einträge der Chronologie zu Alexander von Humboldts Leben, Register und Glossar) der Version 11 der <a href="https://edition-humboldt.de">edition humboldt digital</a>, die am 4. Juni 2025 erschienen ist. Das Datenset enthält gegenüber der HTML-Version technische Fehlerkorrekturen, daher wird es als Version 11.0.1 veröffentlicht.</p> <p>Die Editionsrichtlinien stehen auf <a href="https://edition-humboldt.de/richtlinien/index.html">edition-humboldt.de</a> zur Vefügung. Das Datenmodell ist in drei verschiedene ODDs aufgeteilt (für edierte Texte, Registereinträge und Forschungsbeiträge). Dem Datenset liegen die drei RNG-Schemata bei, die ODD-Ursprungsdateien sind im GitHub-Repository <a href="https://github.com/telota/ediarum.AVHR.data-model/">ediarum.AVHR.data-model</a> zu finden. Beachten Sie bitte, dass es für das Pflanzenregister derzeit noch kein Schema gibt, da dieses aus dem Tagging automatisch erstellt wird.</p> <p>Weitere Hinweise zur digitalen Methodik finden sich in <a href="https://edition-humboldt.de/H0016212">Dumont 2024</a> und zum Editionsvorhaben im Allgemeinen in <a href="https://doi.org/10.25365/wdr-01-03-02">Kraft/Dumont 2020</a>.</p> <p>Dieses Datenset ist auch auf <a href="https://github.com/telota/edition-humboldt-digital">GitHub</a> zugänglich.</p>
XML-Schema (Gemein-Nachrichten)
<p>Mit den <em><strong>Gemein-Nachrichten</strong></em> stellt das Unitätsarchiv Herrnhut der weltweiten Evangelischen Brüder-Unität - Herrnhuter Brüdergemeine (Unitas Fratrum / Moravian Church) das älteste und umfangreichste Mitteilungsblatt der Brüdergemeine digital zur Verfügung. Es enthält Berichte aus Gemeinden sowie dem Missions- und Diasporawerk der Brüdergemeine sowie Reden und Lebensläufe. Die <em>Gemein-Nachrichten</em> wurden ab 1765 in Fortsetzung des <em>Jüngerhaus-Diariums</em> (1747-1764) ausschließlich handschriftlich vervielfältigt. In Druck gingen 1817 und 1818 die <em>Beyträge aus der Brüder-Gemeine</em> und zwischen 1819 und 1894 die <em>Nachrichten aus der Brüder-Gemeine</em>. Das Nachrichtenblatt fand in den <em>Mitteilungen aus der Brüder-Gemeine zur Förderung christlicher Gemeinschaft</em> ab 1895 bis 1941 seine Fortsetzung.</p> <p>Mit dem hier vorliegenden <strong>XML-Schema</strong> werden erschlossene Transkripte mit standardisierten Metadaten angereichert.</p>
IN01055 Halsi Grant of Ravivarman (5 plates). Sanskrit XML file
<p><a href="https://siddham.network/inscription/in01055/">IN01055</a> Halsi Grant of Ravivarman (5 plates). Sanskrit XML file (without metadata).</p>
IN01048 Banavasi Inscription of Mrgesavarman. Sanskrit XML file
<p>IN01048 Banavāsi Inscription of Mṛgeśavarman. Sanskrit XML file (without metadata).</p>
IN01062 Sivalli Grant of Krsnavarman II, Year 7. Sanskrit XML file
<p>IN01062 Śivaḷḷi Grant of Kṛṣṇavarman II, Year 7. Sanskrit XML file (without metadata).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.