Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,085
datasets available to search
ShareScore release 0.9.0
Dataset results
1,085 results for “Documentation”
Fig. 1 in A new fossil from the London Clay documents the convergent origin of a "mousebird-like" tarsometatarsus in an early Eocene near-passerine bird
Fig. 1. The bones preserved in the holotype of the morsoravid bird Sororavis solitarius gen. et sp. nov. (NMS.Z.2021.40.75), from the lower Eocene London Clay of Walton-on-the-Naze, UK. A1, tip of upper beak in dorsal view; A2, fragments of mandible; A3, A4, left coracoid in dorsal (A3) and ventral (A4) views; A5, A6, right coracoid in dorsal (A5) and ventral (A6) views; A7, partial furcula; A8, A9, cranial portion of sternum in ventral (A8) and lateral (A9) views; A10, A11, partial right humerus in cranial (A10) and caudal (A11) views; A12‒A15, proximal (A12, A13) and distal (A14, A15) portions of left humerus in caudal (A12, A14) and cranial (A13, A15) views; A16, proximal end of right ulna in cranioventral view; A17, A18, partial left tibiotarsus in caudal (A17) and cranial (A18) views; A19‒A24, right tarsometatarsus in dorsal (A19), medial (A20), plantar (A21), lateral (A22), proximal (A23), and distal (A24) views; A25, A26, proximal end of left tarsometatarsus in plantar (A25) and dorsolateral (A26) views; A27, first phalanx of third toe in dorsal and plantar view; A28, second to fourth phalanges of fourth toe in different views (plantar, dorsal, and lateral, respectively).
Experimental package for "Live Software Documentation of Design Pattern Instances"
Experimental package containing the materials and data for an empirical study conducted with the DesignPatterDoc plugin for IntelliJ IDEA.
SDADDS-Guelma : A Multi-purpose Dataset for Synthetic Degraded Arabic Documents
<h1><strong>SDADDS-Guelma : A Multi-purpose Dataset for Synthetic Degraded Arabic Documents </strong></h1> <h2><strong>Description:</strong></h2> <p>This is a partial release of the SDADDS-Guelma dataset.</p> <p>SDADDS-Guelma (Synthetic Degraded Arabic Document DataSet of the University of Guelma) is a database of synthetic noisy or degraded Arabic document images. It was created by Dr. Abderrahmane Kefali and his team to support research on preprocessing, analysis, and recognition of degraded Arabic documents, where having a large set of images for training and testing is essential. This dataset is made publicly available to researchers in the field of document analysis and recognition, with the hope that it will be useful and contribute to their research endeavors.</p> <p>In this first release of the dataset, 84 handwritten images and 120 printed images have been used, along with 25 images of historical backgrounds, forming a total of 26316 synthetic images of degraded Arabic documents along with their corresponding ground-truth files.</p> <p>This release is separated into two parts to facilitate upload and use: one for the handwritten documents and the second for the printed documents.</p> <h2><strong>Composition of the dataset:</strong></h2> <p>Each of the parts of the SDADDS-Guelma dataset is organized into directories as follows:</p> <ul> <li>TXT_Files: Contains texts in UTF-8 format. </li> <li>IMG: Contains images of printed and handwritten Arabic text constructed from the text files. </li> <li>Bin_IMG: Contains binary images corresponding to the original images. </li> <li>BG_IMG: Contains images of empty old document backgrounds used for the generation of synthetic historical document images. </li> <li>GT_Files: Contains XML annotation files corresponding to the text images.</li> <li>Degraded_IMG: This directory contains synthetically generated degraded images, separated into sub-directories based on noise types such as Local_Noise, Show_through, Rotation, Curvature, Comb_IMG, etc.</li> </ul> <h2><strong>Ground-truth information:</strong></h2> <p>Ground truth information is essential for a document dataset, as it annotates documents and represents their essential characteristics. Our dataset is designed to be a large-scale and multipurpose dataset. As such, our methodology ensures that ground truth information is provided at three levels: text level (character codes), pixel level (binary and cleaned image), and document physical structure and other annotation information level.</p> <ul> <li>Textual Ground Truth: these are identical to the original texts. </li> <li>Pixel-level ground truth: presented in the form of binary images.</li> <li>Ground truth at the document structure level: the structure of each document image, alongside the textual transcription of the words and PAWs, is recorded in a corresponding XML annotation file. The XML format utilized resembles that employed in similar works with adjustments made according to the specific characteristics of Arabic texts, including the presence of PAWs. </li> </ul> <p>Consequently, each original text image in our dataset is associated to an XML file detailing the entire ground truth and associated metadata. </p> <h3><em><strong>Structure of XML file:</strong></em></h3> <p>Each XML annotation file contains metadata about the document image and text content within the image, including the language, number of lines, and font attributes. It also provides detailed information about each text line, word, and Part of Arabic Words (PAWs), including their bounding boxes and textual transcriptions.</p> <p>Thus, each ground truth file takes the following form:</p> <pre><code><DOCUMENT imageName="PR1Kufi_bin.png" height="2631" width="1860" nbTextLines="8" language="Arabic" fontName="Kufi" fontSize="34"> <TEXTLINE id="0" nbWords="3" boundingBox="215,481,355,1379"> <WORD id="0" nbPAWs="2" boundingBox="217,1065,341,1379" transcription="خصائص"> <PAW id="0" nbCCs="2" boundingBox="217,1206,322,1379" transcription="خصا"> <CC id="0" nbPixels="4110" pixels="(217,1206,1206);(218,1206,1207);(219,1206,1208);(220,1206,1208);(221,1206,1211);(222,1206,1211);..."> </CC> <CC id="1" nbPixels="80" pixels="(263,1330,1336);(264,1329,1336);(265,1328,1337);(266,1328,1337);(267,1328,1337);(268,1328,1337);(269,1328,1337);...."> </CC> </PAW> .... </WORD> <WORD id="1" nbPAWs="2" boundingBox="215,817,338,1044" transcription="التفسير"> <PAW id="0" nbCCs="1" boundingBox="215,1030,322,1044" transcription="ا"> <CC id="0" nbPixels="1037" pixels="(215,1030,1030);(216,1030,1030);(217,1030,1031);..."></CC> .... </PAW> .... </WORD> </TEXTLINE> .... </DOCUMENT></code></pre> <h1><strong>Contact:</strong></h1> <p>Name: Dr. Abderrahmane Kefali<br>Affiliation: University of 8 May 1945-Guelma, Algeria<br>Email: kefali.abderrahmane@univ-guelma.dz</p>
Improving distribution models of sparsely-documented disease vectors by incorporating information on related species via joint modeling
<p>A necessary component of understanding vector-borne disease risk is the accurate characterization of the distributions of their vectors. Species distribution models have been successfully applied to data-rich species but may produce inaccurate results for sparsely-documented vectors. In light of global change, vectors that are currently not well-documented could become increasingly important, requiring tools to predict their distributions. One way to achieve this could be to leverage data on related species to inform the distribution of a<strong> </strong>sparsely-documented vector based on the assumption that the environmental niches of related species are not independent. Relatedly, there is a natural dependence of the spatial distribution of a disease on the spatial dependence of its vector. Here, we propose to exploit these correlations by fitting a hierarchical model jointly to data on multiple vector species and their associated human diseases to improve distribution models of sparsely-documented species. To demonstrate this approach, we evaluated the ability of twelve models—which differed in their pooling of data from multiple vector species and inclusion of disease data—to improve distribution estimates of sparsely-documented vectors. We assessed our models on two simulated data sets, which allowed us to generalize our results and examine their mechanisms. We found that when the focal species is sparsely documented, incorporating data on related vector species reduces uncertainty and improves accuracy by reducing overfitting. When data on vector species are already incorporated, disease data only marginally improve model performance. However, when data on other vectors are not available, disease data can improve model accuracy and reduce overfitting and uncertainty. We then assessed the approach on empirical data on ticks and tick-borne diseases in Florida and found that incorporating data on other vector species improved model performance. This study illustrates the value of exploiting correlated data via joint modeling to improve distribution models of data-limited species.</p>
1964 flamenca negra guitar from Faustino Conde - documentation and reverse engineered construction plans
<p>This data set documents a specific guitar made by Faustino Conde in Madrid in 1964.<br>(Comment: Felipe Conde identified the handwriting of Mariano Conde in the "Conde" signature on the label inside, the signature of his father. This does not necessarily mean that the guitar was built by Mariano, since experts believe, that it was the brother Faustino who built these kind of special guitars, as it has also been the case for the flamenca negra guitars).</p> <p>The guitar is a so-called flamenca negra, a flamenco guitar with rosewood used for rib and back plate.<br>(Errata: the guitar was sold as a flamenca negra, and experts in Granada also believed it is a flamenca negra. However, Felipe Conde now inspected the guitar and states that the instrument has been built as a classical guitar with only some changes in the setup that might lead to other conclusions. It is not clear whether this setup is original from Conde or whether the guitar has been changed in an aftermath.)</p> <p>The guitar has the exceptional character of showing four main resonances between the Helmholtz resonance (A0) and the first air mode (A1), while other guitars usually have one or two resonances in the same range.<br>This is observable across a wider range of old and contemporary guitars, documented in an archive.<br>Mores, R. 2021a. ‘Archive for the acoustical documentation of classical Spanish guitars, flamenco guitars and romantic guitars from private and public collections – bridge mobility’. <em>Zenodo</em>,<br><a href="https://doi.org/10.5281/zenodo.4604577" target="_blank" rel="noopener">doi: 10.5281/zenodo.4604577</a>.<br>This observation triggered a research project to understand this.</p> <p>The resulting analytical model reveals the delicate tuning of related parameters in the construction, published in the Journal of the Acoustical Society of America, JASA. <br>Mores, R. 2021. ‘Sound tuning in asymmetrically braced guitars’. <em>J. Acoust. Soc. Am.</em> 149(2), 1041–1057.<br><a href="https://doi.org/10.1121/10.0003378" target="_blank" rel="noopener">https://doi.org/10.1121/10.0003378</a>.<br>This model explains how to design multiple resonances into a guitar so that the fundamental tone is suported for every semitone played. The model matches with findings not only of this Conde guitar but explains tuning issues in general. There should be a translation into Spanish in due time.</p> <p>A brief talk (English) explains the main issues and demonstrates the congruence between the analysis and an mechanical model, build for demonstration purposes.<br>Mores, R. 2020a. ‘Tuning signature modes in guitars - a lesson by Faustino Conde’. <em>Zenodo</em>, <br><a href="https://doi.org/10.5281/zenodo.4624826" target="_blank" rel="noopener">doi: 10.5281/zenodo.4624826</a>.</p> <p>Supporting material (coded in MATLAB) allows researchers and guitar makers to explore the issue.<br>Mores, R. 2020. ‘Tuning asymmetrically braced guitars - analytical models and MATLAB code’, <em>Zenodo</em>, <br><a href="https://doi.org/10.5281/zenodo.4010596" target="_blank" rel="noopener">doi: 10.5281/zenodo.4010596</a>.</p> <p>This publication documents the construction plans of the Faustino Conde guitar.<br>Guitar makers asked for these plans to understand the principles of parameter tuning based on the construction. The documentation comprises:<br>1. construction plans (high resolution)<br>2. photos of the total instrument<br>3. photos of details inside </p> <p>***</p> <p>Comments on the guitar for those who consider to build this guitar model.<br>1. The guitar is original. Also the top and the mechanics. Experts in Granada, while inspecting the varnish, believed that the top plate might have been modified. However, Felipe Conde (*1959) states that the guitar is original. This also includes the fretboard. The fretboard height is declining towards the body and this caused questions by Granadian guitar makers. However, in the workshop of the Conde family several inspected guitars from the 50s through the 70s revealed a likewise decline of the height.<br>2. Felipe Conde inspected the guitar. He believes that the guitar is an experimental guitar and that it has been built as a customized guitar for a highly professional musician. But there are no records on this.<br>3. The setup of the guitar has been modified. The height of the bridge is lowered where the bone sits. This can be inspected in the construction plan. It is not known whether this was done by Conde himself or by someone else in an aftermath. Experts in Granada but also members of the Conde family state that the present setup is perfect.</p>
PSM-AP Comparative document analysis data: Priorities and challenges in the policy and digital strategies of ten PSM
<p>The document consists of a list of 61 key policy and strategy documents analysed as part of the comparative work conducted in WP1 of the project Public Service Media in the Age of Platforms (PSM-AP). It also contains a series of selected quotes supporting the three key areas prioritised by the policies and PSM digital strategies: People (reaching audiences), Personalisation (developing the video-on-demand portal), and Prominence (of PSM services and content). The data was collected and analysed in 2023, from documents concerning 10 PSM organisations in seven media markets: Belgium-Flanders (VRT), Belgium-Wallonia Brussels (RTBF), Canada (CBC/Radio-Canada), Denmark (DR, TV 2) Italy (RAI), Poland (TVP), and the UK (BBC, Channel 4, ITV). All quotes were translated to English by the authors.</p>
Documentation diagrams of the National Edition of Aldo Moro's works
<p>A series of Graffoo and Draw.io diagrams that graphically represent the conceptual modeling of the Edition's data.</p>
The Clarity Software Documentation Dataset
<p>This repository holds the Clarity Dataset which is a companion to the SANER'22 entitled "An Empirical Investigation into the Use of Image Captioning for Automated Software Documentation". The dataset consists of 45,998 captions 10,204 GUI screenshots and xml metadata files (akin to the "html" for stipulating GUIs) of Android applications. The NL captions were obtained from human labelers, underwent several quality control mechanisms, and contain both high- (screen-level) and low-(component) level descriptions of screen functionality. This dataset is meant as a new source of data to augment techniques for software documentation that can take advantage of the rich pixel-based information contained within screenshots.</p>
State of biodiversity documentation in the Philippines: Metadata gaps, taxonomic biases, and spatial biases in the DNA barcode data of animal and plant taxa in the context of species occurrence data
<p>These files can be categorized into three groups: (1) raw datasets obtained from public databases (i.e., GBIF, BOLD, and GenBank), (2) manually edited files needed for parsing and analysis, and (3) supplementary files for spatial analysis. All are used in the examination of gaps and biases present in Philippine biodiversity data, which can direct research on the taxa and spatial regions that need more sampling.</p>
Knowledge Graph: tyrolean mining documents 15th and 16th century
<p>The dataset contains a Knowledge Graph (.nq file) of two historical mining documents: “Verleihbuch der Rattenberger Bergrichter” ( Hs. 37, 1460-1463) and “Schwazer Berglehenbuch” (Hs. 1587, approx. 1515) stored by the Tyrolean Regional Archive, Innsbruck (Austria). The user of the KG may explore the montanistic network and relations between people, claims and mines in the late medieval Tyrol. The core regions concern the districts Schwaz and Kufstein (Tyrol, Austria).</p> <p>The ontology used to represent the claims is CIDOC CRM, an ISO certified ontology for Cultural Heritage documentation. Supported by the Karma tool the KG is generated as RDF (Resource Description Framework). The generated RDF data is imported into a Triplestore, in this case GraphDB, and then displayed visually. This puts the data from the early mining texts into a semantically structured context and makes the mutual relationships between people, places and mines visible.</p> <p>Both documents and the Knowledge Graph were processed and generated by the research team of the project “Text Mining Medieval Mining Texts”. The research project (2019-2022) was carried out at the university of Innsbruck and funded by go!digital next generation programme of the Austrian Academy of Sciences.</p> <p>Citeable Transcripts of the historical documents are online available:<br> Hs. 37 DOI: 10.5281/zenodo.6274562<br> Hs. 1587 DOI: 10.5281/zenodo.6274928</p>
Documentation and digital files in support of "Aftershock regions of Aleutian–Alaska megathrust earthquakes, 1938-2021" by Carl Tape and Anthony Lomax: Part A
<p>These files support a paper entitled "Aftershock regions of Aleutian–Alaska megathrust earthquakes, 1938–2021," by Carl Tape and Anthony Lomax, published in Journal of Geophysical Research Solid Earth. This collection contains Part A. A separate collection contains Parts, B, C, and D. This research was supported by the U.S. Geological Survey (USGS), Department of the Interior, under USGS award number G19AP00050.</p>
Data Cleaning, Translation & Split of the Dataset for the Automatic Classification of Documents for the Classification System for the Berliner Handreichungen zur Bibliotheks- und Informationswissenschaft
<ul> <li>Cleaned_Dataset.csv – The combined CSV files of all scraped documents from DABI, e-LiS, o-bib and Springer.</li> <li>Data_Cleaning.ipynb – The Jupyter Notebook with python code for the analysis and cleaning of the original dataset.</li> <li>ger_train.csv – The German training set as CSV file.</li> <li>ger_validation.csv – The German validation set as CSV file.</li> <li>en_test.csv – The English test set as CSV file.</li> <li>en_train.csv – The English training set as CSV file.</li> <li>en_validation.csv – The English validation set as CSV file.</li> <li>splitting.py – The python code for splitting a dataset into train, test and validation set.</li> <li>DataSetTrans_de.csv – The final German dataset as a CSV file.</li> <li>DataSetTrans_en.csv – The final English dataset as a CSV file.</li> <li>translation.py – The python code for translating the cleaned dataset.</li> </ul>
Fig. 2 in The first documented record of Chvalaea Papp & Földvári, 2002 (Diptera, Hybotidae, Ocydromiinae) from the Australasian Region: a new species and its possible relationship to other members of the genus
Fig. 2. Chvalaea australis sp. nov. A–D. Male terminalia, holotype (AMS). A. Ventral view. B. Dorsal view. C. Left lateral view. D. Right lateral view. E. Female terminalia, paratype (AMS), ventral view. Abbreviations: bac scl = bacilliform sclerite; cerc = cercus; epand = epandrium; hypd = hypandrium; hyprct = hypoproct; ph = phallus; st = sternite; subepand scl = subepandrial sclerite; sur = surstylus; tg = tergite.
Fig. 5 in The first documented record of Chvalaea Papp & Földvári, 2002 (Diptera, Hybotidae, Ocydromiinae) from the Australasian Region: a new species and its possible relationship to other members of the genus
Fig. 5. Living specimens of Chvalaea Papp & Földvári, 2002 from the Australasian Region. A–B. Specimens resting on tips of branches (Geeveston, Tasmania), provided by Tony Daley. C. Specimen resting on a flower of Bedfordia salicina D.C. (Wellington Park, Tasmania), provided by Keith Martin-Smith.
Fig. 3 in The first documented record of Chvalaea Papp & Földvári, 2002 (Diptera, Hybotidae, Ocydromiinae) from the Australasian Region: a new species and its possible relationship to other members of the genus
Fig. 3. Chvalaea australis sp. nov. Wing of male paratype (AMS). Abbreviations: bm = basal medial cell; br = basal radial cell; cua = anterior cubital cell; CuA+CuP = anterior branch of cubital vein + posterior branch of cubital vein; dm = discal medial cell; M1 = first branch of media; M4 = fourth branch of media; R1 = anterior branch of radius; R2+3= second branch of radius; R4+5= third branch of radius.
Fig. 1 in The first documented record of Chvalaea Papp & Földvári, 2002 (Diptera, Hybotidae, Ocydromiinae) from the Australasian Region: a new species and its possible relationship to other members of the genus
Fig. 1. Chvalaea australis sp. nov. A, C–F. ♂, holotype (AMS). A. Habitus, lateral view. B. ♀, paratype (AMS), habitus, lateral view. C. Frons, dorsal view. D. Head, lateral view. E. Hind leg, lateral view. F. Hind tarsus, lateral view.
Collection of videos documenting the silk weaving process (Silk pilot, Mingei)
<p>Documentation videos of the Weaving process from the Silk pilot of the Mingei project.</p>
PURE: a Dataset of Public Requirements Documents
<p>Please cite this dataset as <strong>Ferrari, A., Spagnolo, G. O., & Gnesi, S. (2017, September). PURE: A dataset of public requirements documents. In <em>2017 IEEE 25th International Requirements Engineering Conference (RE) </em>(pp. 502-505). IEEE.</strong></p> <p><a href="https://ieeexplore.ieee.org/abstract/document/8049173">https://ieeexplore.ieee.org/abstract/document/8049173</a></p> <p>This dataset presents PURE (PUblic REquirements dataset), a dataset of 79 publicly available natural language requirements documents collected from the Web. The dataset includes 34,268 sentences and can be used for natural language processing tasks that are typical in requirements engineering, such as model synthesis, abstraction identification and document structure assessment. It can be further annotated to work as a benchmark for other tasks, such as ambiguity detection, requirements categorisation and identification of equivalent re-quirements. In the associated paper, we present the dataset and we compare its language with generic English texts, showing the peculiarities of the requirements jargon, made of a restricted vocabulary of domain-specific acronyms and words, and long sentences. We also present the common XML format to which we have manually ported a subset of the documents, with the goal of facilitating replication of NLP experiments. The XML documents are also available for download.</p> <p>The paper associated to the dataset can be found here: </p> <p>https://ieeexplore.ieee.org/document/8049173/</p> <p>More info about the dataset is available here: </p> <p>http://nlreqdataset.isti.cnr.it</p> <p>Preprint of the paper available at ResearchGate:</p> <p>https://goo.gl/HxJD7X</p> <p>The dataset includes:</p> <p>- all the documents in PDF format</p> <p>- a subset of 19 documents in XML format</p> <p>- the .xsd schema of the XML files</p> <p>The dataset has been created by gathering data from web sources and we are not aware of license agreements or intellectual property rights on the requirements. The curator took utmost diligence in minimizing the risks of copyright infringement by using non-recent data that is less likely to be critical, by sampling a subset of the original requirements collection, and by qualitatively analyzing the requirements. In case of copyright infringement, please contact the dataset curator (Alessio Ferrari, alessio.ferrari@cnr.it, alessio.ferrari@ucd.ie) to discuss the possibility of removal of that dataset [see <a href="https://support.zenodo.org/help/en-gb/13-policies/140-what-is-your-take-down-procedure" target="_blank" rel="noopener">Zenodo's policies</a>].</p>
Mastic_video_documentations_Mingei
<p>Documentation material from the Mastic pilot of the Mingei project</p>
Fig. 6 in Documenting tenebrionid diversity: progress on Blaps Fabricius (Coleoptera, Tenebrionidae, Tenebrioninae, Blaptini) systematics, with the description of Fve new species
Fig. 6. Bayesian maximum consensus tree resulting from the analysis of the combined molecular and morphological dataset carried out with MrBayes. Support of nodes is indicated using empty circles (for PP> 0.50) and flled circles for PP> 0.95). Species groups for Blaps belonging to section I are highlighted using transparent coloured frames. Paraphyletic species groups are highlighted using bracketed names; orange arrows are also used to highlight the placement of the fve new species. For illustrative purposes, the habitus of some adults are fgured on the right side of the fgure; on the left upper side, the habitus of the fve new species are also presented (all photographs were taken by L. Soldati).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.