Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,085
datasets available to search
ShareScore release 0.9.0
Dataset results
1,085 results for “Documentation”
Fig. 1 in Short communication First documented observation of differential dorsoventral coat colouration in wild boar Sus scrofa (Artyodactyla: Suidae) in Italy
Fig. 1 - The juvenile wild boar object of this note showing the differential dorsoventral colouration pattern (right) next another wild-type individual (left). Additional footage available at: https://youtu.be/gTc0BFSE9kA. / Il giovane esemplare di cinghiale oggetto di questa nota in cui è visibile il pattern cromatico a demarcazione dorsoventrale (a destra) accanto a un altro individuo con la tipica colorazione marrone uniforme (a sinistra). È anche disponibile un filmato aggiuntivo: https://youtu.be/ gTc0BFSE9kA. (Photo and video: / Foto e video: Francesco Gallozzi).
Fig. 2 in Recent documentation of the tropical bed bug (Hemiptera: Cimicidae) in Florida since the common bed bug resurgence
Fig. 2. Auto-Montage photograph of a dissected pronotum from Cimex lectularius (A) and Cimex hemipterus (B).
Fig. 1 in Recent documentation of the tropical bed bug (Hemiptera: Cimicidae) in Florida since the common bed bug resurgence
Fig. 1. Auto-Montage photograph of an adult Cimex hemipterus male collected from Brevard County on the right and an adult Cimex lectularius male on the lef. Arrows are pointing to the lateral pronotum margin on both species.
Рис. 2. Чайка КумΛиена, зарегистрированная 2 марта 2020 г. в северо-восточной части Охотского моря Fig. 2. Kumlien's gull recorded on 2 March 2020 in the northeast Sea of Okhotsk in The first documented record of the Kumlien's gull Larus glaucoides kumlieni Brewster, 1883 in Russia
Рис. 2. Чайка КумΛиена, зарегистрированная 2 марта 2020 г. в северо-восточной части Охотского моря Fig. 2. Kumlien's gull recorded on 2 March 2020 in the northeast Sea of Okhotsk
Рис. 3. Оценка окраски меΛанином первостепенных маховых P6–P10 чайки КумΛиена (баΛΛы по критерию ИнгоΛфссона) Fig. 3. Primary Pattern Score of Kumlien's gull assessed using the Ingolfsson criteria (1970) in The first documented record of the Kumlien's gull Larus glaucoides kumlieni Brewster, 1883 in Russia
Рис. 3. Оценка окраски меΛанином первостепенных маховых P6–P10 чайки КумΛиена (баΛΛы по критерию ИнгоΛфссона) Fig. 3. Primary Pattern Score of Kumlien's gull assessed using the Ingolfsson criteria (1970)
Fig. 10 in First record of Gracixalus quangi Rowley, Dau, Nguyen, Cao & Nguyen, 2011, from Hoa Binh Province, Vietnam, including the first documentation of advanced larval stages and an extended tadpole description
Fig. 10. Drawing of the oral field of a larva of Gracixalus quangi from Hoa Binh Province (IEBR 4333b) in development stage 35: natural view on top, schematic drawing below). Drawing by C. Niggemann.
Fig. 9 in First record of Gracixalus quangi Rowley, Dau, Nguyen, Cao & Nguyen, 2011, from Hoa Binh Province, Vietnam, including the first documentation of advanced larval stages and an extended tadpole description
Fig. 9. Drawing of a larva of Gracixalus quangi from Hoa Binh Province (IEBR 4333b) in development stage 35 in dorsal and lateral view. Drawing by C. Niggemann.
Fig. 8 in First record of Gracixalus quangi Rowley, Dau, Nguyen, Cao & Nguyen, 2011, from Hoa Binh Province, Vietnam, including the first documentation of advanced larval stages and an extended tadpole description
Fig. 8. Advanced developmental stages of Gracixalus quangi from Hoa Binh Province (development stages indicated). Photo credit T. Ziegler and T.D. Tran.
Fig. 6 in First record of Gracixalus quangi Rowley, Dau, Nguyen, Cao & Nguyen, 2011, from Hoa Binh Province, Vietnam, including the first documentation of advanced larval stages and an extended tadpole description
Fig. 6. Early developmental stages of Gracixalus quangi from Hoa Binh Province (development stages indicated). Photo credit C.T. Pham.
Fig. 3 in First record of Gracixalus quangi Rowley, Dau, Nguyen, Cao & Nguyen, 2011, from Hoa Binh Province, Vietnam, including the first documentation of advanced larval stages and an extended tadpole description
Fig. 3. Adult couple of Gracixalus quangi from Hoa Binh Province (male above, female below). Photo credit C.T. Pham.
Fig. 11 in First record of Gracixalus quangi Rowley, Dau, Nguyen, Cao & Nguyen, 2011, from Hoa Binh Province, Vietnam, including the first documentation of advanced larval stages and an extended tadpole description
Fig. 11. Larva of Gracixalus quangi from Hoa Binh Province (IEBR 4333b) in development stage 35 in life (photo taken 4 May 2014). Photo credit T. Ziegler.
EcoregionsTreeFinder – a global dataset documenting observations of 48,129 tree species in 828 terrestrial ecoregions
<p>Check this article for a description of the methods used to develop the EcoregionsTreeFinder. Together with the citation for this Zenodo archive, it is the suggested citation for the database.</p> <p>Kindt, R. and Pedercini, F. (2025), EcoregionsTreeFinder—A Global Dataset Documenting the Abundance of Observations of >45,000 Tree Species in 828 Terrestrial Ecoregions. Global Ecol Biogeogr, 34: e70064. <a href="https://doi.org/10.1111/geb.70064">https://doi.org/10.1111/geb.70064</a></p> <p>Use this shinyapp to filter native tree species for a particular ecoregion or to see ecoregions where a species is expected to be native: <a href="https://patspo.shinyapps.io/EcoregionsTreeFinder/" target="_blank" rel="noopener">https://patspo.shinyapps.io/EcoregionsTreeFinder/</a></p> <p> </p> <p>The database was created from observation records filtered from: GBIF.org (16 March 2021) GBIF Occurrence Download <a href="https://doi.org/10.15468/dl.77gcvq" target="_blank" rel="noopener">https://doi.org/10.15468/dl.77gcvq</a></p> <p> </p> <p><strong>Funding </strong></p> <p>Development of the EcoregionsTreeFinder was supported by the <strong>Bezos Earth Fund</strong> via the Quality Tree Seed for Africa project, by <strong>Norway's International Climate and Forest Initiative</strong> via the Provision of Adequate Tree Seed Portfolio in Ethiopia (PATSPO) project, by the <strong>Darwin Initiative</strong> via project DAREX001 of Developing a Global Biodiversity Standard certification for tree-planting and restoration, by the <strong>Green Climate Fund</strong> via the Readiness proposal Burkina Faso and TREPA projects, and by the <strong>International Climate Initiative</strong> via the Right Tree for the Right Place and Right Purpose (RTRPRP) project.</p> <p> </p>
C. elegans data sample for Pergola documentation
<p>C. elegans data sample for Pergola documentation (<a href="http://cbcrg.github.io/pergola/quick_start.html">http://cbcrg.github.io/pergola/quick_start.html</a>). The sample data set consists in two folders: One named "worm_speeds" containing a CSV file for each of the tracked worms. From the several measures that can be found in the individual files, in this example we will use mid-body speed. The "mapping" folder contains the "worm_speed2pergola.txt", which sets the mappings between the information represented in the worm_speed files and the pergola ontology.</p>
Mouse data sample for Pergola documentation - Shiny visualization
<p>Data sample of feeding and drinking behavior recorded during three weeks of C57BL6/J male mice. The data correspond to 2 groups (9 control mice and 8 high-fat diet mice). Each animal was tracked individually on Phecomp cages for 9 weeks. During the first experimental week all animals were given <em>ad libitum</em> access to a standard chow (habituation phase). After this first week, control mice continued with the same diet regime while high-fat mice were exclusively given <em>ad libitum</em> access to a high-fat chow. Data was used originally in this publication <a href="http://onlinelibrary.wiley.com/doi/10.1111/adb.12595/abstract">10.1111/adb.12595.</a> The recordings were processed using Pergola to BED and BedGraph file formats.</p> <p>The data set consist in:</p> <p>- a exp_info.txt file setting mouse membership to the control or the HF mice.</p> <p>- a files folder containing BED and BedGraph files of mouse feeding behavior.</p>
A Corpus of Online Drug Usage Guideline Documents Annotated with Type of Advice
<p><strong>Introduction: </strong>The goal of this dataset is to aid NLP research on recognizing safety critical information from drug usage guideline or patient handout data. This dataset contains annotated advice statements from 90 online DUG documents that corresponds to 90 drugs or medications that are used in the prescriptions of patients suffering from one or more chronic diseases. The advice statements are annotated in eight safety-critical categories: activity or lifestyle related, disease or symptom related, drug administration related, exercise related, food or beverage related, other drug related, pregnancy related, and temporal. </p> <p><strong>Data Collection: </strong>The data was collected from <a href="https://www.medscape.com">MedScape</a>. It is one of the most widely used reference for health care providers. At first, 34 real anonymized prescriptions of patients suffering from one or more chronic diseases are collected. These prescriptions contains 165 drugs that are used to treat chronic diseases. Then, MedScape was crawled to collect the drug user guideline (DUG) / patient handout for these 165 drugs. But, MedScape does not have DUG document for all drugs. We found DUG document for 90 drugs in MedScape. </p> <p><strong>Data Annotation tool: </strong>The data annotation tool is developed to ease the annotation process. It allows the user to select a DUG document and select a position from the document in terms of line number. It stores the user log from the annotator and loads the most recent position from the log when the application is launched. It supports annotating multiple files for the same drug, as often there are multiple overlapping sources of drug usage guidelines for a single drug. Often DUG documents contain formatted text. This tool aids annotation of the formatted text as well. The annotation tool is also available upon request.<strong> </strong></p> <p><strong>Annotated Data Description: </strong>The annotated data contains the annotation tag(s) of each advice extracted from the 90 online DUG documents. It also contains the phrases or topics in the advice statement that triggers the annotation tag, such as, activity, exercise, medication name, food or beverage name, disease name, pregnancy condition (gestational, postpartum). Sometimes disease names are not directly mentioned rather mentioned as a condition (e.g., stomach bleeding, alcohol abuse) or state of a parameter (e.g., low blood sugar, low blood pressure). The annotated data is formatted as following:<br> drug name, drug number, line number of the first sentence of the advice in the DUG document, advice Text, advice tag(s), medication, food, activity, exercise, and disease names mentioned in the advice. </p> <p><br> <strong>Unannotated Data Description:</strong><br> The unannotated data contains the raw DUG document for 90 drugs. It also contains the drug interaction information for the 165 drugs. The drug interaction information is categorized in 4 classes, contraindicated, serious, monitor closely, and minor. This information can be utilized to automatically detect potential interaction and effect of interaction among multiple drugs. </p> <p><strong>Citation: </strong>If you use this dataset in your work, please cite the following reference in any publication:</p> <p>@inproceedings{preum2018DUG,<br> title={A Corpus of Drug Usage Guidelines Annotated with Type of Advice},<br> author={Sarah Masud Preum, Md. Rizwan Parvez, Kai-Wei Chang, and John A. Stankovic},<br> booktitle={ Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)},<br> publisher = {European Language Resources Association (ELRA)},<br> year={2018}<br> }</p>
Research data, sources and documents for thesis on Exploring Complexity Metrics for Artifact-Centric Business Process Models
<p>Research data, sources and documents for thesis on Exploring Complexity Metrics for Artifact-Centric Business Process Models This repository contains the supplemental material for the <a href="https://pqdtopen.proquest.com/pubnum/10759956.html">thesis "Exploring Complexity Metrics for Artifact-Centric Business Process Models" by Marin, Mike A., Ph.D., University of South Africa (South Africa), 2017.</a></p>
FN-RE: A Corpus of Requirements Documents Enriched with Semantic Frame Annotations
<p>FN-RE is a human-labelled dataset using FrameNet scheme. The dataset is distributed and can be viewed using a web-index page. For further details about the annotation procedures, please refer to the annotation guidelines included in the folder.</p>
ScriptNet: ICDAR2017 Competition on Historical Document Writer Identification (Historical-WI)
<p>This dataset contains the test set for the ICDAR2017 Competition on Historical Document Writer Identification (Historical-WI).</p> <p>The dataset used in this competition consists of 3600 handwritten pages originating from 13th to 20th century. It contains manuscripts from 720 different writers where each writer contributed five pages.</p> <p>Competition Website: https://scriptnet.iit.demokritos.gr/competitions/6/</p> <p>Changes August 1st, 2018: uploaded trainings set in color and binarized</p> <p>if you use the dataset please cite the ICDAR2017 Competition on Historical Document Writer Identification (Historical-WI) paper: https://doi.org/10.1109/ICDAR.2017.225</p> <p> </p>
Oficio de Hipotecas de Girona. A dataset of Spanish notarial deeds (18th Century) for Handwritten Text Recognition and Layout Analysis of historical documents.
<p>This dataset is a subset of 596 documents from the <em>Registre d'Hipoteques de Girona</em> of 1769 collection, guarded by the <a href="http://xac.gencat.cat/ca/llista_arxius_comarcals/girones/"><em>Arxiu Històric de Girona</em></a>. This collection, is composed by hundreds of thousands of notarial deeds from the XVIII-XIX century (1768-1862). Sales, redemption of censuses, inheritance and matrimonial chapters are among the most common documentary typologies in the collection.</p> <p>This dataset is composed of more than 23700 text lines written by a single hand, covering more that 50 different topics (documentary typologies) and a vocabulary of more than 2400 different words. The documents are transcribed using the so-called diplomatic criteria. Additionally, transcripts were tagged with <br> extra enriching/complementary information (e.g. expansion of the abbreviations, hyphen marks, etc.). Along with the transcripts the layout of the document is detected and recorded. Pages have been labeled using six different layout regions.</p> <p>The images along with their respective ground-truth was compiled in PAGE compliant XML format<br> by the <a href="http://www2.udg.edu/tabid/11296/Default.aspx"><em>Centre de Recerca d'Història Rural</em></a> and the HTR group of the <a href="https://www.prhlt.upv.es">Pattern Recognition and Human Language Technologies Research Center</a>.</p>
GT4HistOCR: Ground Truth for training OCR engines on historical documents in German Fraktur and Early Modern Latin
<p><strong>GT4HistOCR</strong> contains ground truth for research in Optical Character Recognition (OCR) technology applied to historical printings in German Fraktur and Early Modern Latin.</p> <p>The ground truth comes in pairs of images of single printed lines as they appear in book pages (*.png) and their corresponding diplomatic transcriptions (*.gt.txt), which are UTF-8 strings preserving the character forms (glyphs) as much as possible within the UNICODE standard. These pairs of line images and their transcriptions can be directly used to train recognition models with, e.g., the open source OCR engines <em>OCRopy</em> or <em>Tesseract</em>. A total of 313,173 ground truth lines are provided.</p> <p><strong>Please note that the subcorpora making up this collection used different transcription guidelines, so it is a bad idea to train a recognition model on the total collection! Rather train individual models for each subcorpus.</strong> Fur further information about the subcorpora, please see the README file and the accompanying publication.</p> <p>If these data are useful for you, please cite the accompanying publication:</p> <pre>@article{<a href="http://springmann.net/publications.html#springmann2018gt4hist">springmann2018gt4hist</a>, author = {Uwe Springmann and Christian Reul and Stefanie Dipper and Johannes Baiter}, title = {{Ground Truth for training {OCR} engines on historical documents in German Fraktur and Early Modern Latin}}, journal = {J. Lang. Technol. Comput. Linguistics}, volume = {33}, number = {1}, pages = {97--114}, year = {2018}, url = {https://jlcl.org/content/2-allissues/1-heft1-2018/jlcl_2018-1_5.pdf} }</pre>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.