Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7,523

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

7,523 results for “Annotation”

Learn how ShareScore rates datasets ↗
zenodo40/100

Worldwide Fraxinus Genome Assemblies, Annotations, and Gene Families v0.2

<p>The worldwide&nbsp;<em>Fraxinus</em>&nbsp;genome project was conducted to assess the pathogenic resistance of 34 Ash tree species to Ash Dieback and Emerald Ash Borer. The project is led by Dr. Richard Buggs. v0.1 genomes are available at ashgenome.org and on ENA. As part of Josiah Seaman&#39;s PhD thesis, he improved the assembly of 13 genomes included here as v0.2. New de novo annotations, gene families, and all associated files are included for future studies and reproducibility.&nbsp;</p> <p>Annotations are done with GeMoMa using F. excelsior as a reference&nbsp;(Keilwagen et al. 2016). Gene families are defined as genes originating from a single copy at the last common ancestor with Solanum. Orthofinder outputs reconciled gene trees, aligned CDS, and gene families (Emms and Kelly 2015; Tekaia 2016). Species tree was calibrated based on fossil evidence using r8s,&nbsp;RAxML across&nbsp;25,182,399 sites (SpeciesTreeAlignment.fa). More methods details can be found in the full Chapter two of Josiah Seaman&#39;s PhD thesis (2021).</p> <p>I&#39;d be happy to talk with you if you&#39;d like any additional information or help visualizing your genomic data. You can find the tools used to browse this data at&nbsp;https://fluentdna.com/ and&nbsp;http://graphgenome.org/ Contact me at josiah@newline.us</p>

opencc-by-nd-3.0Dec 2020View details →
zenodo40/100

CORAL: A corpus of ontological requirements annotated with Lexico-Syntactic Patterns

<p>In this work&nbsp;we present CORAL (Corpus of Ontological Requirements Annotated with Lexico-syntactic patterns), an openly available corpus of 834 ontological requirements annotated and 29 lexico-syntactic patterns, from which 12 are proposed in this work. CORAL is openly available in three different open formats, namely, HTML, CSV and RDF.</p>

opencc-by-4.0Dec 2018View details →
zenodo40/100

A studyforrest extension, an annotation of spoken language in the German dubbed movie ``Forrest Gump'' and its audio-description (annotation)

<p>This dataset contains the annotation of speech spoken in the research cut (Hanke et al. 2014; Hanke et al., 2016) of the movie &quot;Forrest Gump&quot; (Zemeckis, 1994) and its audio-description that was broadcast as an additional audio track (Koop et al., 2009) for visually impaired listeners on Swiss public television. The corresponding paper is hosted on github (https://github.com/psychoinformatics-de/studyforrest-paper-speechannotation) and published in f1000research (https://doi.org/10.12688/f1000research.27621.1).</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Classical Syriac annotated lexemes

<p>31,972 lexemes in Classical Syriac and their inflectional forms annotated according to Sylak-Glassman (2016).</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Dataset - research methodology annotation (Information Science)

<p>Datasets used for developing text mining methods for extracting research methods reported in Information Science journal articles.&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Annotated MRI and ultrasound volume images of the prostate

<p><strong>Introduction</strong></p> <p>The <em>Surgical Planning Laboratory (SPL) </em>and the <em>National Center for Image Guided Therapy (NCIGT) </em>are making this dataset available as a resource to aid in the development of algorithms and tools for deformable registration,&nbsp;segmentation and analysis of prostate magnetic resonance imaging (MRI) and ultrasound&nbsp;(US) images. &nbsp;</p> <p><strong>Description</strong></p> <p>This dataset contains anonymized images of the human prostate (N=3 patients) collected during two sessions for each patient:</p> <ol> <li>MRI&nbsp;examination of the prostate for the purposes of disease staging.</li> <li>US&nbsp;examination of the prostate for the purposes of volumetric examination in preparation to the brachytherapy implant.</li> </ol> <p>These are three-dimensional (multi-slice) scalar images.</p> <p>Image files are stored using NRRD file format (files with .nrrd extension), see details at http://teem.sourceforge.net/nrrd/format.html. Each image file includes a code for the case number (internal numbering at the research site) and the modality (US or MR).</p> <p>Image annotations were prepared by Dr. Fedorov (no professional training in radiology)&nbsp;and Dr. Tuncali (10+ in prostate imaging interpretation). Annotations include</p> <ol> <li>Manual contouring (segmentation) of the whole prostate gland, performed in 3D Slicer software. These segmentation images are coded in the same fashion as the image files, and saved in NRRD format, with &quot;-label&quot; suffix.</li> <li>Manually placed points (fiducials) corresponding to the location of urethra entry into the prostate at base (coded as UB), verumontanum (VM), urethra entry into the prostate at apex (UA), as well as centroids of cysts and calcifications. UB, UA and VM locations are annotated both in MR and US for all cases, while cysts and calcifications are annotated when applicable. Fiducial points are stored in comma-separated CSV-style format adopted by 3D Slicer software&nbsp;(.fcsv file extension). There is one row per point in these files, encoding the location of the point in RAS coordinate space relative to the image data, and the name of the point.</li> </ol> <p><strong>Viewing the collection</strong></p> <p>We tested visualization of images, segmentations and fiducials in 3D Slicer software, and thus recommend 3D Slicer as the platform for visualization. 3D Slicer is a free open source platform (see http://slicer.org), with the pre-compiled binaries available for all major operating systems. You can download 3D Slicer at http://download.slicer.org.</p> <p><strong>Acknowledgments </strong></p> <p>Preparation of this data collection was made possible thanks to the&nbsp;funding from the National Institutes of Health (NIH) through grants R01 CA111288 and P41 RR019703.</p> <p>If you use this dataset in a publication, please cite the following manuscript. You can also learn more about this dataset from the publication below.</p> <p>Fedorov, A., Khallaghi, S., Antonio S&aacute;nchez, C., Lasso, A., Fels, S., Tuncali, K., Sugar, E. N., Kapur, T., Zhang, C., Wells, W., Nguyen, P. L., Abolmaesumi, P. &amp; Tempany, C. Open-source image registration for MRI&ndash;TRUS fusion-guided prostate interventions.&nbsp;<em>Int J CARS</em>&nbsp;<strong>10,</strong>&nbsp;925&ndash;934 (2015). https://pubmed.ncbi.nlm.nih.gov/25847666/</p> <p><strong>Contact</strong></p> <p>Andrey Fedorov, fedorov@bwh.harvard.edu</p>

opencc-by-nc-sa-4.0Mar 2015View details →
zenodo40/100

Annotated T2-weighted MR images of the Lower Spine

<p><strong>Annotated T2-weighted MR images of the Lower Spine</strong></p> <p>Chengwen Chu, Daniel Belavy, Gabriele Armbrecht, Martin Bansmann, Dieter Felsenberg, and Guoyan Zheng&nbsp;</p> <p><strong>Introduction</strong><br /> The Institute for Surgical Technology and Biomechanics, University of Bern, Switzerland, Charit&eacute; - University Medicine Berlin, Centre of Muscle and Bone Research, Free University &amp; Humboldt-University Berlin, Germany,&nbsp;Centre for Physical Activity and Nutrition Research, School of Exercise and Nutrition Sciences, Deakin University Burwood Campus, Australia and Institut f&uuml;r Diagnostische und Interventionelle Radiologie, Krankenhaus Porz Am Rhein gGmbH, K&ouml;ln, Germany, are making this dataset available as a resource in the development of algorithms and tools for spinal image analysis.</p> <p><strong>Description</strong><br /> The database consists of T2-weighted turbo spin echo MR spine images of 23 anonymized patients, each containing at least 7 vertebral bodies (VBs) of the lower spine (T11 &ndash; L5). For each vertebral body, reference manual segmentation is provided in the form of a binary mask. All images and binary masks are stored in the Neuroimaging Informatics Technology Initiative (NIFTI) file format, see details at http://nifti.nimh.nih.gov/. Image files are stored as &quot;Img_xx.nii&quot; while the associated annotation files are stored as &quot;Img_xx_Labels.nii&quot;, where &quot;xx&quot; is the internal case number for the patient.&nbsp;</p> <p>Image annotations were prepared by Mr. Chengwen Chu (no professional training in radiology).&nbsp;</p> <p><strong>Acknowledgements</strong></p> <ul> <li>The acquisition of original images was supported by the&nbsp;Grant 14431/02/NL/SH2 from the European Space Agency,&nbsp; grant 50WB0720 from the German Aerospace Center (DLR) and the Charit&eacute; Universit&auml;tsmedizin Berlin.</li> <li>Preparation of this data collection was made possible thanks to the funding from the Swiss National Science Foundation (SNSF) through project: 205321 157207/1.</li> </ul> <p><strong>Reference</strong><br /> C. Chu, D. Belavy, W. Yu, G. Armbrecht, M. Bansmann, D. Felsenberg, and G. Zheng, &ldquo;Fully Automatic Localization and Segmentation of 3D Vertebral Bodies from CT/MR Images via A Learning-based Method&rdquo;, <strong>PLoS One</strong>.&nbsp;2015 Nov 23;10(11):e0143327. doi: 10.1371/journal.pone.0143327. eCollection 2015.</p>

opencc-zeroJul 2015View details →
zenodo40/100

PDEStrIAn: A phosphodiesterase structure and ligand interaction annotated database as a tool for structure-based drug design

<p>A systematic analysis is presented of the 220 phosphodiesterase (PDE) catalytic domain crystal structures present in the Protein Data Bank (PDB) with a focus on PDE-ligand interactions. The consistent structural alignment of 57 PDE ligand binding site residues enables the systematic analysis of PDE-ligand Interaction FingerPrints (IFPs), the identification of subtype-specific PDE-ligand interaction features, and the classification of ligands according to their binding modes. We illustrate how systematic mining of this phosphodiesterase structure and ligand interaction annotated (PDEStrIAn) database provides new insights into how conserved and selective PDE interaction hot spots can accommodate the large diversity of chemical scaffolds in PDE ligands. A substructure analysis of the co-crystalized PDE ligands in combination with those in the ChEMBL database provides a toolbox for scaffold hopping and ligand design. These analyses lead to an improved understanding of the structural requirements of PDE binding that will be useful in future drug discovery studies.</p>

opencc-zeroFeb 2016View details →
zenodo40/100

COMBAT TB Tuberculosis genome annotation database

<p>A Neo4j (version 2.3.3) format graph database containing annotation related to the M. tuberculosis H37Rv genome, created as part of the COMBAT TB project at the South African National Bioinformatics Institute.</p>

opencc-by-4.0May 2016View details →
zenodo40/100

PyPI Packages Annotated

<p>PyPI Package data, as a directed graph, in .gexf format. Directed edges are software dependencies. Package metadata includes information about release dates and number of downloads.</p>

openmit-licenseJul 2016View details →
zenodo40/100

FIGURE 2. Durckheimia lochi n in Description of Durckheimia lochi n. sp., with an annotated checklist of Australian Pinnotheridae (Crustacea: Decapoda: Brachyura)

FIGURE 2. Durckheimia lochi n. sp., female holotype, AM P 24478, in bivalve mollusc host, Ctenoides ales (Finlay, 1927) (AM C 103731). Photo: I. Loch.

opencc-zeroDec 2003View details →
zenodo40/100

FIGURE 4 in An annotated check-list of lophogastrids (Crustacea: Lophogastrida) from the seas of the Iberian Peninsula

FIGURE 4. Similarity of subregions (nMDS / clustering) based on presence / absence of lophogastrid species in the 6 main geographic areas of the Iberian Seas. The analysis included data for the Mediterranean coasts of Italy (* Wittmann &amp; Ariani 2010) and Celtic, South Icelandic, South Greenlandic and West Atlantic regions (** Petryashov 2009).

opencc-zeroDec 2016View details →
zenodo40/100

FIGURE 1 in An annotated check-list of lophogastrids (Crustacea: Lophogastrida) from the seas of the Iberian Peninsula

FIGURE 1. Map of the Iberian Peninsula and nearby waters showing the different areas considered to characterize the spatial distribution of lophogastrid species.

opencc-zeroDec 2016View details →
zenodo40/100

FIGURE 2 in An annotated check-list of lophogastrids (Crustacea: Lophogastrida) from the seas of the Iberian Peninsula

FIGURE 2. Cumulative number of the first record of species (N spp) and published reports of lophogastrids (N ms) over time in the Iberian Seas.

opencc-zeroDec 2016View details →
zenodo40/100

FIGURE 5 in An annotated check-list of lophogastrids (Crustacea: Lophogastrida) from the seas of the Iberian Peninsula

FIGURE 5. Similarity of species distributions (nMDS / clustering) based on presence / absence of lophogastrid species in the 6 main geographic areas of the Iberian Seas (this study), the Mediterranean coasts of Italy (Wittmann &amp; Ariani 2010) and Celtic, South Icelandic, South Greenlandic and North European Atlantic regions (Petryashov 2009).

opencc-zeroDec 2016View details →
zenodo40/100

Jingju a cappella singing syllable boundary and duration annotation dataset

<p>This dataset is a collection of syllable boundary annotations and syllable duration annotations of a cappella singing performed by jingju (京剧, Beijing opera) professional and amateur singers. This dataset was used as the experimental dataset in the following work:</p> <blockquote> <p>Rong Gong, Nicolas Obin, Georgi Dzhambazov and Xavier Serra, &ldquo;Score-Informed syllable segmentation for jingju a cappella singing voice with Mel-frequency intensity profiles,&quot; in<em>&nbsp;Folk Music Analysis workshop (FMA) 2017, M&aacute;laga, Spain</em></p> </blockquote> <p><strong>Audio Content</strong></p> <p>The audio files are the a cappella singing arias recordings, which are stereo or mono, sampled at 44.1 kHz, and stored as wav files. They can be found at this link http://doi.org/10.5281/zenodo.344932</p> <p>The wav files are recorded by two institutes: those file names ending with &lsquo;qm&rsquo; are recorded by C4DM Queen Mary University of London; others file names ending with &lsquo;upf&rsquo; or &lsquo;lon&rsquo; are recorded by MTG-UPF. If you use the dataset in your work, please cite the following publication.</p> <blockquote> <p>D. A. A. Black, M. Li, and M. Tian, &ldquo;Automatic Identification of&nbsp;Emotional Cues in Chinese Opera Singing,&rdquo; in&nbsp;<em>13th Int. Conf. on Music&nbsp;</em><em>Perception and Cognition</em>&nbsp;(ICMPC-2014), 2014, pp. 250&ndash;255.</p> </blockquote> <p><strong>Annotations</strong></p> <p>The syllable boundary annotation is in Textgrid format (Praat). The annotation is done in both phrase-level and syllable-level. The syllable duration annotation is in cvs format. Please consult Readme text in both folders for further details. The parsing code of the annotation files is provided in &lsquo;pycode&rsquo; folder.&nbsp;</p> <p><strong>Availability of the Dataset</strong></p> <p>The annotations and codes in this dataset are licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.</p> <p><strong>Contact</strong></p> <p>If you have any questions or comments about the dataset, please feel free to write to us.</p> <p>Rong Gong: rong&lt;dot&gt;gong&lt;at&gt;upf&lt;dot&gt;edu</p> <p>Rafael Caro Repetto: rafael&lt;dot&gt;caro&lt;at&gt;upf&lt;dot&gt;edu</p>

opencc-by-nc-4.0Mar 2017View details →
zenodo40/100

Dataset for "Bacterial genome annotation" and "AMR gene detection" workflows

<p>This dataset is associated with the workflows "Bacterial genome annotation" and "AMR gene detection in an assembled bacterial genome".</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

TweetC19SR-Eng - Manually annotated dataset of English language COVID-19 tweets containing self-reports of symptoms

<p><strong>In this work, we release two expert curated, manually annotated datasets of COVID-19 self-reported symptoms. The first dataset contains tweets in English and the second contains tweets in Spanish, both containing around 36,500 tweets in total. These datasets were used for the Sixth and Seventh Workshop on Social Media Mining For Health (2021 and 2022)</strong></p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

TweetC19SR-Spa - Manually annotated dataset of Spanish language COVID-19 tweets containing self-reports of symptoms

<p><strong>In this work, we release two expert curated, manually annotated datasets of COVID-19 self-reported symptoms. The first dataset contains tweets in English and the second contains tweets in Spanish, both containing around 36,500 tweets in total. These datasets were used for the Sixth and Seventh Workshop on Social Media Mining For Health (2021 and 2022)</strong></p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Annotated Corpus for the Alsatian Dialects

<p>This corpus contains a collection of texts in the Alsatian dialects which were manually annotated with parts-of-speech, lemmas, translations into French and location entities.</p><p>The corpus was produced in the context of the RESTAURE project, funded by the French ANR. The current version of the corpus contains 21 documents and 12,907 syntactic words. The annotation process is detailed in the following article: <a href="http://hal.archives-ouvertes.fr/hal-01704806">http://hal.archives-ouvertes.fr/hal-01704806</a></p><p><strong>Information about version 3</strong></p><p>Version 3 corrects some minor errors in the CONLL-U files: wrong token indexes after multiword tokens and missing _ in glosses. In addition, all files are concatenated into a single CONLL-U file.</p><p><strong>Information about version 2</strong></p><p>Version 2 contains the same annotated documents as version 1, but some errors have been corrected and the annotated corpus is provided in the <a href="http://universaldependencies.org/format.html">CoNLL-U format</a></p><p>The untokenised and unannotated versions of the documents are found in the "txt" folder. The annotated versions of the documents are found in the "ud" folder (<a href="http://universaldependencies.org/format.html">CoNLL-U format</a>).</p><p>In addition to the form, the lemma and the part-of-speech additional information is also provided:</p><ul><li>translation of the lemma into French (Gloss field)</li><li>annotation of location names (NamedType field)</li></ul>

opencc-by-sa-4.0Feb 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record