Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,085
datasets available to search
ShareScore release 0.9.0
Dataset results
1,085 results for “Documentation”
The LSST Dark Energy Science Collaboration (DESC) Science Requirements Document v1 Released Data Products
<p>This tarball includes software and data products associated with the DESC Science Requirements Document (SRD) v1. See the "Executive Summary and User Guide" in the enclosed PDF of the DESC SRD for instructions on how to use and cite those products. The DESC SRD is described on <a href="https://arxiv.org/abs/1809.01669">arXiv</a> as follows:</p> <p>The Large Synoptic Survey Telescope (LSST) Dark Energy Science Collaboration (DESC) will use five cosmological probes: galaxy clusters, large scale structure, supernovae, strong lensing, and weak lensing. The Science Requirements Document (SRD) quantifies the expected dark energy constraining power of these probes individually and together, with conservative assumptions about analysis methodology and follow-up observational resources based on our current understanding and the expected evolution within the field in the coming years. We then define requirements on analysis pipelines that will enable us to achieve our goal of carrying out a dark energy analysis consistent with the Dark Energy Task Force definition of a Stage IV dark energy experiment.</p>
Se repérer : les différents types de documents
<p>Ce sketchnote complète une section de cours d'une enseignante consacré à l'identification et la citation des sources documentaires.</p> <p>Les définitions sont en partie issues de https://bib.umontreal.ca/guides/types-documents</p>
Ontology files in .owl, rdf and xml format demonstrating the texts, documents and works (=entities) ontology
<p>See various writings by Robinson concerning this ontology (e.g. , <a href="https://wiki.usask.ca/pages/viewpage.action?pageId=1324745355">Creating and Implementing an Ontology of Documents and Texts (ADHO 2018)</a>.</p> <p>Note revision of 10/21: removal of parts of document, work, text. dc:hasPart and dc:isPartOf make this redundant.</p>
Stavronikita Monastery Greek handwritten document Collection no.79
<p>It comprises manuscripts made of paper, written in the 16th century and its dimensions are 220X165 mm. The manuscript is embellished with epititles and red initials. Tachygraphical symbols and abbreviations are encountered in the manuscript as well. The dataset of XΦ79 consists of 803 lines of text containing 4389 words (2069 unique words) that are distributed over 40 scanned handwritten text pages.<br> For each page, a PageXML is provided containing the following ground-truth:</p> <p>1) Text region polygon coordinates<br> 2) Text line polygon coordinates with the corresponding transcription text<br> 3) Word polygon coordinated with the corresponding transcription text</p>
Stavronikita Monastery Greek handwritten document Collection no.114
<p>It comprises manuscripts made of paper, written at the end of the 15th century and its dimensions are 218X150 mm. In various pages, we find red initials and epititles which enrich the manuscript’s decoration. <br> The dataset of ΧΦ114 consists of 1051 lines of text containing 5467 (2877 unique words) words that are distributed over 44 scanned handwritten text pages. <br> For each page, a PageXML is provided containing the following ground-truth:</p> <p>1) Text region polygon coordinates<br> 2) Text line polygon coordinates with the corresponding transcription text<br> 3) Word polygon coordinated with the corresponding transcription text </p>
Stavronikita Monastery Greek handwritten document Collection no.53
<p>The collection is one of the oldest Stavronikita Monastery on Mount Athos. It is a parchment, four-gospel manuscript which has been written between 1301 and 1350. It comprises 54 pages with dimensions that are approximately 250x185 mm. The script is elegant minuscule and the use of majuscule letters is rare. Tachygraphical symbols and abbreviations are encountered in the manuscript as well. Furthermore, the manuscript is enriched with chrysography, elegant epititles and initials. The dataset of ΧΦ53 consists of 1038 lines of text containing 5592 words (2374 unique words) that are distributed over 54 scanned handwritten text pages.</p>
2008 Nura earthquake surface rupture slip vector documentation
<p>This online data holds information related to the surface rupture resulting from the 2008 Nura earthquake in south Kyrgyzstan. The primary dataset is a Google Earth KMZ file with GPS-locations where slip vector measurements were taken along the rupture. A downloadable ZIP file accompanies the KMZ, containing photographs linked to each data point. Both files should be stored in one folder for proper linkage. In addition, a text file is available with all measurements and associated information. Five videos obtained with the drone are available to illustrate the surface rupture zones and geological overview in the Nura settlement surroundings. The entire data was gathered in 2018. </p> <p>Raster-files of high-resolution digital surface models of the surface rupture can be found on opentopography <a href="https://doi.org/10.5069/G9ZW1J4C" target="_blank" rel="noreferrer noopener">https://doi.org/10.5069/G9ZW1J4C</a></p> <p>The data presented in this repository was initially disseminated in a dissertation by Magda Patyniak. This project is part of the CaTeNA-project within the Client II program of and funded by the Federal Ministry of Education and Research (BMBF; Sub-project grant 03G0878E to Manfred Strecker).</p>
Document Liveness Challenge (DLC-2021) - part 3 (cc)
<p>Dataset DLC-2021 consists of 1424 video clips captured in a wide range of real-world conditions and focused on ID document forensics tasks. Each clip was shot vertically and was at least 5 seconds long. Frames extracted at 10 frames per second and for the 50 first extracted frames document position is manually annotated.<br> The novelty of the dataset is that it contains shots from video with color laminated mock ID documents, color unlaminated copies, grayscale unlaminated copies, and screen recaptures of the documents. The proposed dataset complies with the GDPR because it contains images of synthetic IDs with generated owner photos and artificial personal information.</p> <p>Part 1 contains videos, frames and markup for “original” laminated documents from MIDV-2020 collection and unlaminated gray copies. <br> Part 2 contains videos, frames and markup for documents recaptured from device screen<br> Part 3 contains videos, frames and markup for unlaminated color copies.</p> <p><strong>Share and Cite</strong></p> <p><em>MDPI and ACS Style</em></p> <p>Polevoy, D.V.; Sigareva, I.V.; Ershova, D.M.; Arlazarov, V.V.; Nikolaev, D.P.; Ming, Z.; Luqman, M.M.; Burie, J.-C. Document Liveness Challenge Dataset (DLC-2021). <em>J. Imaging</em> <strong>2022</strong>, <em>8</em>, 181. https://doi.org/10.3390/jimaging8070181</p> <p><em>AMA Style</em></p> <p>Polevoy DV, Sigareva IV, Ershova DM, Arlazarov VV, Nikolaev DP, Ming Z, Luqman MM, Burie J-C. Document Liveness Challenge Dataset (DLC-2021). <em>Journal of Imaging</em>. 2022; 8(7):181. https://doi.org/10.3390/jimaging8070181</p> <p><em>Chicago/Turabian Style</em></p> <p>Polevoy, Dmitry V., Irina V. Sigareva, Daria M. Ershova, Vladimir V. Arlazarov, Dmitry P. Nikolaev, Zuheng Ming, Muhammad M. Luqman, and Jean-Christophe Burie. 2022. "Document Liveness Challenge Dataset (DLC-2021)" <em>Journal of Imaging</em> 8, no. 7: 181. https://doi.org/10.3390/jimaging8070181</p>
Document Liveness Challenge (DLC-2021) - part 2 (re)
<p>Dataset DLC-2021 consists of 1424 video clips captured in a wide range of real-world conditions and focused on ID document forensics tasks. Each clip was shot vertically and was at least 5 seconds long. Frames extracted at 10 frames per second and for the 50 first extracted frames document position is manually annotated.<br> The novelty of the dataset is that it contains shots from video with color laminated mock ID documents, color unlaminated copies, grayscale unlaminated copies, and screen recaptures of the documents. The proposed dataset complies with the GDPR because it contains images of synthetic IDs with generated owner photos and artificial personal information.</p> <p>Part 1 contains videos, frames and markup for “original” laminated documents from MIDV-2020 collection and unlaminated gray copies. <br> Part 2 contains videos, frames and markup for documents recaptured from device screen<br> Part 3 contains videos, frames and markup for unlaminated color copies.</p> <p><strong>Share and Cite</strong></p> <p><em>MDPI and ACS Style</em></p> <p>Polevoy, D.V.; Sigareva, I.V.; Ershova, D.M.; Arlazarov, V.V.; Nikolaev, D.P.; Ming, Z.; Luqman, M.M.; Burie, J.-C. Document Liveness Challenge Dataset (DLC-2021). <em>J. Imaging</em> <strong>2022</strong>, <em>8</em>, 181. https://doi.org/10.3390/jimaging8070181</p> <p><em>AMA Style</em></p> <p>Polevoy DV, Sigareva IV, Ershova DM, Arlazarov VV, Nikolaev DP, Ming Z, Luqman MM, Burie J-C. Document Liveness Challenge Dataset (DLC-2021). <em>Journal of Imaging</em>. 2022; 8(7):181. https://doi.org/10.3390/jimaging8070181</p> <p><em>Chicago/Turabian Style</em></p> <p>Polevoy, Dmitry V., Irina V. Sigareva, Daria M. Ershova, Vladimir V. Arlazarov, Dmitry P. Nikolaev, Zuheng Ming, Muhammad M. Luqman, and Jean-Christophe Burie. 2022. "Document Liveness Challenge Dataset (DLC-2021)" <em>Journal of Imaging</em> 8, no. 7: 181. https://doi.org/10.3390/jimaging8070181</p>
9+10+8 Immersive Music Production Audio and Documentation Archive.GEIDAI.WH
<p>This repository contains several resources related to research on the effect of floor-level loudspeakers on 3D audio reproduction, undertaken by Will Howie, Toru Kamekawa, Miki Morinaga, and Atsushi Marui at Tokyo University of the Arts, November 2021 - November 2023. Please follow the guidelines of usage found in the READ ME file. "Audio" contains 29ch (9+10+8) interleaved audio files of short excerpts of five immersive recordings of musical sound scenes. "Documentation" contains text-based, diagrammatical, and photographic documentation of the recording sessions that yielded these audio excerpts.</p>
OSH Automated Documentation - Overview Video
<p>A video that shows the problem and the approach of OSH Automated Documentation. It gives an overview of the functionality of the software.</p>
OSH Automated Documentation - Creating an Assembly Manual for a Vise Video
<p>A video that shows how an assembly manual of a simple vise is generated (semi) automatically from a CAD specification and a textual description.</p>
UDAPDR Document and Question Datasets
<p>Question and document datasets for <a href="https://arxiv.org/abs/2303.00807">UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers</a></p>
Dataset od: "Towards a Taxonomy of Roxygen Documentation in R Packages"
<p>Replication package for the paper titled "Towards a Taxonomy of Roxygen Documentation in R Packages"</p>
China's new type rural social insurance pension (NRSIP): transposition central government pilot guidelines into local documents
<p>The list provides an overview of how the pilot guidelines issued by the central government have been transposed into local documents (provincial level).</p> <p>Matthias Stepan has compiled the data and information in the period 2009-2013. It has been part of his dissertation project at VU University Amsterdam.</p>
Spectra belonging to XRF instrument report. Identification of ink components through XRF analysis of Azzolino documents.
<p>Accumulated spectra. Details given in supporting information </p> <p><a href="https://journals.plos.org/plosone/article/file?type=supplementary&id=10.1371/journal.pone.0283539.s002">S2 File. </a>XRF instrument report.</p> <p>Identification of ink components through XRF analysis of Azzolino documents.</p> <p><a href="https://doi.org/10.1371/journal.pone.0283539.s002">https://doi.org/10.1371/journal.pone.0283539.s002</a></p> <p>(DOCX)</p> <p>Belonging to publication </p> <p>Lagerqvist Alidoost A, Hacke M, Winther T, Sandström T (2023) A closer look at the Azzolino collection. PLOS ONE 18(4): e0283539. <a href="https://doi.org/10.1371/journal.pone.0283539">https://doi.org/10.1371/journal.pone.0283539</a></p>
XRF maps and line scans project files belonging to XRF instrument report. Identification of ink components through XRF analysis of Azzolino documents.
<p>RTX Project files. Details given in supporting information </p> <p><a href="https://journals.plos.org/plosone/article/file?type=supplementary&id=10.1371/journal.pone.0283539.s002">S2 File. </a>XRF instrument report.</p> <p>Identification of ink components through XRF analysis of Azzolino documents.</p> <p><a href="https://doi.org/10.1371/journal.pone.0283539.s002">https://doi.org/10.1371/journal.pone.0283539.s002</a></p> <p>(DOCX)</p> <p>Belonging to publication </p> <p>Lagerqvist Alidoost A, Hacke M, Winther T, Sandström T (2023) A closer look at the Azzolino collection. PLOS ONE 18(4): e0283539. <a href="https://doi.org/10.1371/journal.pone.0283539">https://doi.org/10.1371/journal.pone.0283539</a></p>
Dataset for Paper: A System for Processing and Recognition of Greek Byzantine and Post-Byzantine Documents
<p>Dataset for the paper: "A System for Processing and Recognition of Greek Byzantine and Post-Byzantine Documents", P. Kaddas, K. Palaiologos, B. Gatos, V. Katsouros, K. Christopoulou, 17th International Conference on Document Analysis and Recognition (ICDAR), San Jose, California, USA</p> <p>The dataset consists of 57 pages from the third edition of the Greek New Testament published by Robert Estienne (1503–1559), who was appointed “Royal Typographer” by the King of France François I (1494–1547). Robert Estienne produced this edition in 1550 using the grecs du roi typeface, produced by Claude Garamont on the basis of the Greek minuscule style of the calligrapher Angelos Vergikios (1505–1569) from Crete, who active copying Greek manuscripts in Venice and France. The dataset consists of 2045 cropped text line images in .png format with their corresponding OCR in .txt format, where 1431 used for training, 204 for validation and 410 for test. Initial images acquired from: https://bibles-online.net/flippingbook/1550/</p>
Lack of funds for consent document translation impedes inclusive enrollment
<p>Data on all consent events for patients who participated in clinical trials at UCLA from January 2013 to December 2018.</p>
Documenting Reef-Fish Diversity in the Revillagigedo Archipelago, Pacific Mexico, November 2022
<p>Documenting Reef-Fish Diversity in the Revillagigedo Archipelago, Pacific Mexico, November 2022</p> <p>By Allison Morgan Estape, Citizen Scientist</p> <p>150 Nautilus Drive Islamorada, Florida 33036, USA. Email: <a href="mailto:allison.carlos@me.com">allison.carlos@me.com</a>. </p> <p>Photographic Website: <a href="https://carlosestape.photoshelter.com/gallery-list">https://carlosestape.photoshelter.com/gallery-list</a>.</p> <p>Between November 20 and December 2, 2022, a group of seven professional ichthyologists from Mexico, Panama and the USA, together with 11 SCUBA-diving photographers experienced in taking diagnostic images of reef-fishes, conducted a scientific expedition aboard the diving-support vessel Quino El Guardian to the Revillagigedo Archipelago, a Mexican Marine Protected Area that lies 250 miles south-southwest of the tip of Baja California. The objectives of the expedition were (1) to photograph all observed species of reef-fishes in their natural habitat at each of the archipelago’s four islands to provide permanent documentation of such species occurrences; and (2) to collect specimens for museums in Mexico and the United States at which they would be used for morphological and genetic assessments of endemism in the archipelagos’ reef-fish fauna and the general relationships of that fauna to the reef-fish fauna of Mexico and other parts of the Tropical Eastern Pacific. This video of a power-point show provides a visual record of the expedition and summarizes the results by island and the expedition’s overall findings to date.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.