Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,523
datasets available to search
ShareScore release 0.7.1
Dataset results
7,523 results for “Annotation”
Fig. 4 in An annotated checklist of the herpetofauna of the Sibiloi National Park in northern Kenya based on field surveys
Fig. 4. Reptile species recorded during the surveys: (A) Crocodylus niloticus; (B) Agama lionotus; (C) Agama rueppelli; (D) Holodactylus africanus; (E) Hemidactylus angulatus; (F) Hemidactylus barbierii; (G) Hemidactylus lanzai; (H) Hemidactylus ruspolii; (I) Homopholis fasciata; (J) Lygodactylus somalicus; (K) Stenodactylus sthenodactylus; (L) Heliobolus spekii; (M) Latastia longicaudata; (N) Philochortus rudolfensis; (O) Chalcides bottegi; (P) Mochlus sundevallii; (Q) Trachylepis striata; (R) Varanus albigularis; (S) Eryx colubrinus; (T) Platyceps brevis; (U) Psammophis cf. tanganicus; (V) Psammophis punctulatus; (W) Rhamphiophis rostratus; (X) Naja pallida; (Y) Bitis arietans; and (Z) Echis pyramidum.
Fig. 3 in An annotated checklist of the herpetofauna of the Sibiloi National Park in northern Kenya based on field surveys
Fig. 3. Amphibian species recorded during the surveys: (A) Poyntonophrynus lughensis; (B) Sclerophrys xeros; (C) Sclerophrys turkanae; (D) Ptychadena nilotica; (E) Ptychadena cf. schillukorum; and (F) Tomopterna wambensis.
Fig. 1 in An annotated checklist of the herpetofauna of the Sibiloi National Park in northern Kenya based on field surveys
Fig. 1. Location of Sibiloi National Park (UNEP-WCMC and IUCN 2022) in Kenya and the main study sites: IL (Ilkemere), KA (Karare), KF (Koobi Fora), LO (Lomosia), AB (Alia Bay), and TBI (Turkana Research Institute). The inset map shows the African continent, and the black square indicates the location of the enlarged map.
Fig. 22 in An annotated checklist of Dolichopodidae (Diptera) species from Iran, with new records and a bibliography
Fig. 22. Natural Resources & Agricultural Research Center, Bampour, Sistan & Balouchestan Prov., 30 March ‒ 2 May 2017. Photograph by courtesy of F. Basavand Рис. 22. Центр прироΑных ресурсов и сеΛьскохозяйственных иссΛеΑований, Бампур, пров. Систан и БеΛуΑжистан, 30 марта – 2 мая 2017 г. Фотография Ф. БасаванΑа
Fig. 19 in An annotated checklist of Dolichopodidae (Diptera) species from Iran, with new records and a bibliography
Fig. 19. Platanus and Populus tremula alley in the Sasan Park, Tehran. Tree trunks are the habitat of Medetera diadema, M. truncorum and M. veles flies. Photograph by I. Grichanov, 11 October 2022 Рис. 19. АΛΛея с пΛатаном и осиной в парке Сасан, Тегеран. СтвоΛы Αеревьев явΛяются местом обитания имаго Medetera diadema, M. truncorum и M. veles. Фото И. Гричанова, 11 октября 2022 г.
Figs. 6‒12 in An annotated checklist of Dolichopodidae (Diptera) species from Iran, with new records and a bibliography
Figs. 6‒12. Asyndetus transversalis (Becker) (6‒8), Campsicnemus pilitarsis Negrobov et Zlobin (9‒11), Dolichopus efflatouni (Parent) (12). Habitus (6, 9, 12); apex of abdomen, left (7) and right (8) lateral; fore and mid tarsi, lateral (10); mid tarsus, dorsal view (11) Рис. 6‒12. Asyndetus transversalis (Becker) (6‒8), Campsicnemus pilitarsis Negrobov et Zlobin (9‒11), Dolichopus efflatouni (Parent) (12). Габитус (6, 9, 12); вершина брюшка сΛева (7) и справа (8) сбоку; переΑние и среΑние Λапки, сбоку (10); среΑняя Λапка, виΑ сверху (11)
Figs. 13‒18 in An annotated checklist of Dolichopodidae (Diptera) species from Iran, with new records and a bibliography
Figs. 13‒18. Dolichopus diadema Haliday (13), Emiratomyia arabica Naglis (14), Medetera mixta Negrobov (15‒16), Teuchophorus monacanthus Loew (17‒18). Habitus (13, 14, 15, 17); hypopygium, left lateral view (16); hind femur and tibia, anterior view (18) Рис. 13–18. Dolichopus diadema Haliday (13), Emiratomyia arabica Naglis (14), Medetera mixta Negrobov (15–16), Teuchophorus monacanthus Loew (17–18). Габитус (13, 14, 15, 17); гипопигий, виΑ сΛева (16); заΑние беΑра и гоΛени, виΑ спереΑи (18)
Figs. 1‒5 in An annotated checklist of Dolichopodidae (Diptera) species from Iran, with new records and a bibliography
Figs. 1‒5. Arabshamshevia ajbanensis Naglis (1‒2), Asyndetus albifacies Parent (3‒4), Asyndetus separatus (Becker) (5). Habitus (1, 3, 5); apex of abdomen (2, 4) Рис. 1‒5. Arabshamshevia ajbanensis Naglis (1‒2), Asyndetus albifacies Parent (3‒4), Asyndetus separatus (Becker) (5). Габитус (1, 3, 5); вершина брюшка (2, 4)
Fig. 23 in An annotated checklist of Dolichopodidae (Diptera) species from Iran, with new records and a bibliography
Fig. 23. Karkheh National Park, Shoush, Khuzestan Prov. Photograph by E. Gilasian Рис. 23. НационаΛьный парк Кархе, Шуш, пров. Хузестан. Фотография Э. ГиΛасяна
Fig. 20 in An annotated checklist of Dolichopodidae (Diptera) species from Iran, with new records and a bibliography
Fig. 20. Rashakan Research Station for Lake Urmia National Park, West Azerbaijan Prov., 25 June 2015. Photograph by courtesy of M. Parchami-Araghi Рис. 20. Рашаканская научная станция национаΛьного парка «Озеро Урмия», провинция ЗапаΑный АзербайΑжан, 25 июня 2015 г. Фотография М. Парчами-Араги
Supplemental Results for Assembly, Annotation, and Analysis from HiFi reads of Gulf Toadfish Genome and Transcriptome fOpsBet2.1
<p>This repository contains gzipped tarballs of the results of the various assembly, annotation, and analysis steps performed during the assembly of the fOpsBet2.1 genome assembly for Opsanus beta at the University of Miami Rosenstiel School of Marine, Atmospheric, and Earth Science for the McDonald Toadfish Lab. These results are too numerous to include as supplemental data for a journal publication and so are available here for review. In this repository you will find results for:</p><p>Scripts:</p><p>-all bash and LSF scheduler job scripts used as part of the analysis, both exploratory and final. </p><p>QC:</p><p>-GenomeScope2 estimation of genome metrics from HiFi Reads</p><p>-QUAST genome statistics for each assembly step</p><p>-BUSCO completeness assessments for each assembly step </p><p>-inspector logs for polishing of initial assembly</p><p>-logs from Kraken2 contaminant screen</p><p>Assembly and Scaffolding:</p><p>-ntLINKS logs and intermediates for initial scaffolding</p><p>-ragtag logs and metrics for super-scaffolding to the ThaAma1.1 T. amazonica reference assembly</p><p>-mitoHIFI results for mitogenome assembly from HiFi reads, primary assembly, and purged alternate assembly</p><p>Annotation:</p><p>-PASA directory with full input and output for SQLite PASA assembly of transcriptome for gene predictors</p><p>-Results folder for Funannotate::annotate for gene models, annotations, and CDS/mRNA/protein fastas</p><p>-InterProSCan5 results for protein annotation used as input into Funannotate</p><p>-ghostKOALA KEGG assignment results for predicted proteins from funannotate results</p><p>Repetitive Elements:</p><p>-tidk telomere repeat analysis results</p><p>-TRAH satellite DNA analysis with subsequent analysis with HiCAT and StainedGlass</p><p>-RepeatModeler results for de novo TE prediction</p><p>-repclassifier results for TE curation</p><p>Comparative Analysis:</p><p>-OrthoFinder ortholog search for O. beta to several other vertebrates</p><p>-CAFE5 gene family expansion and contraction of Orthogroups from OrthoFinder results</p><p> </p>
Diachronic Annotated Corpus of Newar
<p>This dataset contains segmented and part-of-speech-tagged files that comprise the ongoing Diachronic Annotated Corpus of Newar (DACON). Files are provided in .txt format.</p> <p>File names are explained as follows:</p> <p>cnew (Classical Newar)<br>century (e.g. 12)<br>short text name (and other information such as manuscript name (e.g., MSB) and line number completed for incomplete texts (e.g. 10000))<br>SEG or POS (segmented or part-of-speech-tagged</p> <p>e.g. cnew19-manicuda-10000_SEG.txt</p> <p>For full details, including text citations with discussion, please see: <br>O'Neill & Meelen, "The Diachronic Annotated Corpus of Newar: from Manuscript to Morphosyntax," <em>Cahiers de Linguistique Asie Orientale</em> (2024).</p> <p><a href="../doi/10.5281/zenodo.13117922">Annotation Manual Part I (Preprocessing)</a></p> <p><a href="../doi/10.5281/zenodo.13117961">Annotation Manual Part II (Segmentation and POS Tagging)</a></p> <p><a href="https://github.com/lothelanor/newarcorpora">Tools for the Diachronic Annotated Corpus of Newar (DACON)</a></p>
AQL queries and benchmark results from PhD thesis "ANNIS: A graph-based query system for deeply annotated text corpora"
<p>These are the queries, the benchmark results and the evaluation scripts of the thesis "ANNIS: A graph-based query system for deeply annotated text corpora" (Thomas Krause 2018, Humboldt-Universität zu Berlin)</p> <p><strong>diss_2018-01-12_v0.5.0.csv </strong><br> Results of all configurations of executed benchmarks for graphANNIS and also the baseline times of relANNIS.</p> <p><strong>queries.zip</strong><br> Contains folders for each corpus containing all queries used for the benchmark. Each file-name begins with the ID of the query. The extension denotes the type, which can be one of the following:</p> <ul> <li><em>".</em>aql" contains the original AQL (ANNIS query language) query which was collected</li> <li>".json" is JSON representation of the parsed AQL query</li> <li>".count" is the number of matches a query should have</li> <li>".time" is the average time in milliseconds that was needed to execute the query in relANNIS on the benchmark system</li> <li>".corpora" contains the name of the corpus the query belongs to (should be only one corpus and the same as the folder name in the selection of queries in this data set)</li> <li>".relplan" contains the PostgreSQL plan for the query</li> <li>".graphplan" contains the graphANNIS plan for the query</li> </ul> <p><strong>evaluation-scripts.py/evaluation-scripts.ipynb</strong><br> Python scripts to perform the evaluation and generate the images. This are both a Python-file and the original notebook file that can be used with the Jupyter Notebook application.</p> <p><strong>relannis_benchmark_scripts.zip </strong><br> The files in this zip-file can be used to execute the benchmarks in the relANNIS system by piping the into the "annis.sh" command line tool of relANNIS</p>
A Corpus of Online Drug Usage Guideline Documents Annotated with Type of Advice
<p><strong>Introduction: </strong>The goal of this dataset is to aid NLP research on recognizing safety critical information from drug usage guideline or patient handout data. This dataset contains annotated advice statements from 90 online DUG documents that corresponds to 90 drugs or medications that are used in the prescriptions of patients suffering from one or more chronic diseases. The advice statements are annotated in eight safety-critical categories: activity or lifestyle related, disease or symptom related, drug administration related, exercise related, food or beverage related, other drug related, pregnancy related, and temporal. </p> <p><strong>Data Collection: </strong>The data was collected from <a href="https://www.medscape.com">MedScape</a>. It is one of the most widely used reference for health care providers. At first, 34 real anonymized prescriptions of patients suffering from one or more chronic diseases are collected. These prescriptions contains 165 drugs that are used to treat chronic diseases. Then, MedScape was crawled to collect the drug user guideline (DUG) / patient handout for these 165 drugs. But, MedScape does not have DUG document for all drugs. We found DUG document for 90 drugs in MedScape. </p> <p><strong>Data Annotation tool: </strong>The data annotation tool is developed to ease the annotation process. It allows the user to select a DUG document and select a position from the document in terms of line number. It stores the user log from the annotator and loads the most recent position from the log when the application is launched. It supports annotating multiple files for the same drug, as often there are multiple overlapping sources of drug usage guidelines for a single drug. Often DUG documents contain formatted text. This tool aids annotation of the formatted text as well. The annotation tool is also available upon request.<strong> </strong></p> <p><strong>Annotated Data Description: </strong>The annotated data contains the annotation tag(s) of each advice extracted from the 90 online DUG documents. It also contains the phrases or topics in the advice statement that triggers the annotation tag, such as, activity, exercise, medication name, food or beverage name, disease name, pregnancy condition (gestational, postpartum). Sometimes disease names are not directly mentioned rather mentioned as a condition (e.g., stomach bleeding, alcohol abuse) or state of a parameter (e.g., low blood sugar, low blood pressure). The annotated data is formatted as following:<br> drug name, drug number, line number of the first sentence of the advice in the DUG document, advice Text, advice tag(s), medication, food, activity, exercise, and disease names mentioned in the advice. </p> <p><br> <strong>Unannotated Data Description:</strong><br> The unannotated data contains the raw DUG document for 90 drugs. It also contains the drug interaction information for the 165 drugs. The drug interaction information is categorized in 4 classes, contraindicated, serious, monitor closely, and minor. This information can be utilized to automatically detect potential interaction and effect of interaction among multiple drugs. </p> <p><strong>Citation: </strong>If you use this dataset in your work, please cite the following reference in any publication:</p> <p>@inproceedings{preum2018DUG,<br> title={A Corpus of Drug Usage Guidelines Annotated with Type of Advice},<br> author={Sarah Masud Preum, Md. Rizwan Parvez, Kai-Wei Chang, and John A. Stankovic},<br> booktitle={ Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)},<br> publisher = {European Language Resources Association (ELRA)},<br> year={2018}<br> }</p>
Character mentions in the German novel "Corpus Delicti" by Juli Zeh and annotations
<p>This file contains all character mentions in the German novel "Corpus Delicti" by Juli Zeh. The annotation was conducted by members of the research group hermA (www.herma.uni-hamburg.de). The file includes the following columns:</p> <ul> <li>id</li> <li>token: tokens of the character mention</li> <li>entity_nr: the (arbitrary) entity number. All mentions of one character share the same entity number.</li> <li>token_nr: position of the mention in the text in tokens</li> <li>sentence_nr: position of the mention in the text in sentences</li> <li>chapter: number of the chapter the mention occurs in (49 chapters in total)</li> <li>direct_speech: True if the mention occurs between quotation marks</li> <li>entity_name: a mapping of the entity number to the most frequent proper name used for the entity (if available)</li> <li>form: grammatical form of the mention, derived from an automatic part-of-speech tagging (NE = proper name; NP = noun phrase; PPER = personal pronoun; PPOSAT = possessive pronoun)</li> </ul> <p> </p> <p> </p>
FN-RE: A Corpus of Requirements Documents Enriched with Semantic Frame Annotations
<p>FN-RE is a human-labelled dataset using FrameNet scheme. The dataset is distributed and can be viewed using a web-index page. For further details about the annotation procedures, please refer to the annotation guidelines included in the folder.</p>
Annotated dataset for sub-shot segmentation evaluation
<p>This dataset was created for evaluating the performance of the developed motion-driven approach for fine-grained temporal segmentation of user-generated videos, that is reported in [1].</p> <p>Given the fact that user-generated videos are most commonly captured without interruption using a single camera, thus being single-shot videos, their fine-grained temporal segmentation aims to identify visually coherent parts (called sub-shots) that correspond to individual actions taking place during the video recording, such as left camera panning, camera zoom-out or tracking of a moving object.</p> <p>Based on the above, the created dataset contains the ground-truth sub-shot segmentation for 33 single-shot videos. This ground-truth segmentation was created by human annotation of the sub-shot boundaries for each video, where each boundary indicates the end of a visually and temporally contiguous activity of the video recording device and the start of the next one (e.g. a downward camera tilting that is followed by a camera zoom-in). Overall, our dataset contains 674 sub-shot transitions.</p> <p>The set of videos can be divided in three parts:</p> <ol> <li><em>Own Videos</em>: 15 single-shot videos of total duration 6 minutes, which contain clearly defined fragments that correspond to several video recording activities;</li> <li><em>Amateur Videos</em>: 5 single-shot amateur videos of total duration 17 minutes, found on the YouTube platform;</li> <li><em>Movie Excerpts</em>: 13 single-shot parts of known movies of total duration 46 minutes which represent professional video content.</li> </ol> <p>The videos of the first part were recorded in the external spaces of our (CERTH-ITI) facilities using an iPhone 5 smart-phone. These videos and their ground-truth are freely provided through the "Provided files" section below.</p> <p>The videos of the second and third part can be found on the YouTube platform (links in the "Video collection" section below). For these videos we only provide the ground-truth data again through the "Provided files" section below.</p> <p>For each video of the dataset we provide:</p> <ul> <li>The video ID;</li> <li>The video filename for "Own videos" and the YouTube video name and URL for "Amateur videos" and "Movie Excerpts" (accessed on 12/11/2017);</li> <li>The video frame-rate that directly affects the ground-truth annotation; the videos of the second and third part should be downloaded at the corresponding frame-rates (default values when using the KeepVid web application - <a href="https://keepvid.com/">https://keepvid.com/</a>) to match the created annotations.</li> </ul> <p>The above information is included in the provided "video_collection_info.pdf" file.</p> <p>The needed files for using the created dataset are in the provided compressed "dataset.zip" file. After unpacking the compressed file a structure of directories will be generated. In this structure:</p> <ol> <li>The "videos/own_videos" directory contains: a) the video files of the first part of the dataset, and b) the "list.txt" file which lists the video id, the filename and the frame-rate for each video of the "Own videos" collection, in a tab-separated format.</li> <li>The "videos/amateur_videos" directory contains the "list.txt" file which lists the video id, the filename and the frame-rate for each video of the "Amateur videos" collection, in a tab-separated format.</li> <li>The "videos/movie_excerpts" directory contains the "list.txt" file which lists the video id, the filename and the frame-rate for each video of the "Movie excerpts" collection, in a tab-separated format.</li> <li>The "ground_truth" directory contains 33 txt files (one per video) where each file is named after the ID of the corresponding video and includes the annotation of the sub-shots in the video; each line of this file corresponds to the starting frame of a sub-shot of the video (using a zero-based index of frames).</li> <li>The "evaluation" directory contains a simple Matlab script, called "eval_segm.m", that was used for evaluating the results of the sub-shot segmentation analysis.</li> <li>A readme file that contains the documentation of the dataset (i.e. the information that is available in this webpage).</li> </ol> <p>CERTH-ITI holds the copyright of the “<em>Own videos</em>” part of the dataset. These videos can be freely used for academic/research purposes only and their public playback is not allowed. CERTH-ITI does not own the copyright of the “<em>Amateur Videos</em>” and “<em>Movie Excerpts</em>” parts of the dataset. The use of these videos must abide by the YouTube copyright policy (<a href="https://www.youtube.com/yt/about/copyright/">https://www.youtube.com/yt/about/copyright/</a>). Any users of these videos must accept full responsibility for the use of these parts of the dataset, including but not limited to the use of any copies of copyrighted videos that they may create from the dataset.</p> <p>If you use this dataset, please cite the following scientific work:</p> <p>[1] K. Apostolidis, E. Apostolidis, V. Mezaris, <em>"A motion-driven approach for fine-grained temporal segmentation of user-generated videos"</em>, Proc. 24th Int. Conf. on Multimedia Modeling (MMM2018), Bangkok, Thailand, Feb. 2018.</p> <p> </p> <p> </p>
eggNOG Mapper annotations of Mouse, Dog and Pig gut gene catalogs
<p><a href="https://github.com/jhcepas/eggnog-mapper">eggNOG-mapper</a> annotations of <a href="https://doi.org/10.1038/nbt.3353">mouse</a>, <a href="https://doi.org/10.1186/s40168-018-0450-3">dog</a> and <a href="https://doi.org/10.1038/nmicrobiol.2016.161">pig</a> gut, and <a href="https://doi.org/10.1038/nbt.2942">IGC</a> gene catalogs.</p>
Partially automatically annotated corpus to predict gestural cues in Embodied Conversational Agents
<p>#Structure of the corpus</p> <p>This corpus has been built using speeches of Spanish politicians freely available <a href="http://www.congreso.es/portal/page/portal/Congreso/Congreso/Intervenciones">here</a> along with their transcriptions.</p> <p>Each transcription has been analyzed in terms of:</p> <ul> <li>Surface Syntactic Structure*</li> <li>Deep Syntactic Structure*</li> <li>Morphology (Part of Speech)*</li> <li>Communicative Structure</li> </ul> <p>Gestures (beat vs. no gesture tags) have been annotated using the videos.</p> <p>*All those features have been automatically retrieved using the parser freely available in https://github.com/TalnUPF/miis. The other features have been annotated manually.</p> <p>#Concerns about the corpus</p> <p>This corpus has been mostly annotated manually. Annotation agreement has not been computed.</p> <p>Moreover, it is small. In order to extract reliable correlations from it, it should be extended.</p>
Scalasca analysis report of the ASCI Sweep3D benchmark on 294,912 processes in virtual-node mode on IBM Blue Gene/P with manually annotated iterations
<p>A Cube3 performance analysis report written by the Scalasca parallel analyzer of a measurement of the ASCI benchmark Sweep3D, executed in virtual-node mode on 294,912 processes of the IBM Blue Gene/P system JUQUEEN, operated by Forschungszentrum Jülich GmbH. The measurement includes system topology information and manual annotations of the twelve iterations.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.