Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,523
datasets available to search
ShareScore release 0.7.1
Dataset results
7,523 results for “Annotation”
FIGURE 52 in The marine ichthyofauna of Lebanon: an annotated checklist, history, biogeography, and conservation status
FIGURE 52. Trachipterus trachypterus (Gmelin 1789), Sarafand, 7 May 2018, photograph: unknown (posted on local social media).
FIGURE 31 in The marine ichthyofauna of Lebanon: an annotated checklist, history, biogeography, and conservation status
FIGURE 31. Ophisurus serpens (Linnaeus 1758), Beirut, February 2007, Crocetta & Bariche in Dailianis et al. (2016).
FIGURE 22 in The marine ichthyofauna of Lebanon: an annotated checklist, history, biogeography, and conservation status
FIGURE 22 Glaucostegus cemiculus (Geoffroy St. Hilaire 1817), Beirut, 17 October 2007, AUBM (CH0032).
FIGURE 11 in The marine ichthyofauna of Lebanon: an annotated checklist, history, biogeography, and conservation status
FIGURE 11. Odontaspis ferox (Risso 1810), Tripoli, 10 June 2014, photograph: unknown (posted on local social media).
FIGURE 8 in The marine ichthyofauna of Lebanon: an annotated checklist, history, biogeography, and conservation status
FIGURE 8. Alopias vulpinus (Bonnaterre 1788), Naqoura, 2 June 2018, photograph: unknown (posted on local social media).
FIGURE 2 in The marine ichthyofauna of Lebanon: an annotated checklist, history, biogeography, and conservation status
FIGURE 2. Carcharhinus falciformis (Müller & Henle 1839), Saida, 15 September 2017, photograph: unknown (posted on local social media).
FIGURE 3 in The marine ichthyofauna of Lebanon: an annotated checklist, history, biogeography, and conservation status
FIGURE 3. Carcharhinus limbatus (Valenciennes 1839), Naqoura, October 2011, photograph: Ziad Samaha.
Figure 1 in A new and an unrecorded species of the family Psychidae (Lepidoptera) from Korea, with an annotated catalogue
Figure 1. Proutia maculatella Saigusa et Sugimoto, male. (A) adult; (B) close-up of right wing; (C) head, frontal view; (D) ditto, lateral view; (E) scales, upperside of forewing; (F) ditto, upperside of hindwing; (G) close-up of cucullus; (H) anellus, dorsal view; (I) close-up of saccus; (J) left valva; (K) dorsum, dorsal view; (L) genitalia, lateral view; (M) ditto, ventral view. Scale bar 0.5 mm.
Petunia parodii S7 annotations
<p>Petunia axillaris v4 annotations (https://zenodo.org/record/3922967#.X1IhSIZS-EJ) gmapped onto the Petunia axillaris subs. parodii assembly (https://www.ncbi.nlm.nih.gov/assembly/GCA_013625405.1) described in "Indentification of transcription factors controlling floral morphology in wild Petunia species with contrasting pollination syndroms" https://onlinelibrary.wiley.com/doi/abs/10.1111/tpj.14962</p>
MIDAS hand-annotated news articles
<p>This dataset was produced in 2020 from the data collected throughout 2019 for the development of the MIDAS project (http://www.midasproject.eu/)</p> <p>The data is distributed throughout 5 topics:</p> <p>- EUS: Childhood Obesity (UC Basque Country)<br> - FIN: Mental Health (UC Finland)<br> - IRE: Diabetes (UC Ireland)<br> - NIR: Children in Care (UC Northern Ireland)<br> - INF: Infectious Diseases including Coronavirus (UC Influenzanet)</p> <p>The available data comes in 3 kinds and file formats:</p> <p>TXT - the source of news including ID, title and body of text<br> CSV - the hand annotation of the news articles in TXT with 5 to 10 MeSH headings<br> JSON - the input file for the evaluation of the classifier, including the title, news article body and MeSH heading IDs (available from https://www.ncbi.nlm.nih.gov/mesh/)</p> <p>The CSV files with name starting in "f1_", "pr_", "re_" are the results of the F1/Precision/Recall evaluation for each of the cases.</p> <p>## AUTHORS</p> <p>Joao Pita Costa, Anthony Staines, Jarmo Pääkkönen, Jenni Konttila, Joseba Bidaurrazaga, Oihana Belar, Christine Henderson</p> <p>## ACKNOWLEDGMENTS</p> <p>This work was supported by the European Commission H2020 project MIDAS (G.A. nr. 727721). </p> <p><br> ## LICENSE</p> <p>This dataset is licensed over Creative Commons.</p>
FIG. 2 in Annotated checklist of bats (Mammalia: Chiroptera) of Mount Cameroon, southwestern Cameroon
FIG. 2. — Habitat sampled for bats on Mount Cameroon: A, slow flowing streams; B, cultivated farmland; C, fallow farmland; D, beside fruiting trees; E, cleared farmland; F, understory of primary forest; G, ecotone forest/ alpine grassland; H, cave; I, waterhole. Photos: © Aaron Manga Mongombe
FIG. 1 in Annotated checklist of bats (Mammalia: Chiroptera) of Mount Cameroon, southwestern Cameroon
FIG. 1. — Map of Cameroon, showing localities listed in the text (See Appendix 1 for names of localities).
PharmaCoNER corpus: gold standard annotations of Pharmacological Substances, Compounds and proteins in Spanish clinical case reports
<p><strong>Intro:</strong></p><p>The PharmaCoNER corpus (divided into train, dev and test) is a Gold Standard manually annotated dataset used for the the PharmaCoNER shared task posed at BIONLP-ST (at EMNLP). In addition, we include here the PharmaCoNER background set. It contains the train, development and test sets of the two subtasks (subtask-1 and subtask-2) with Gold Standard annotations. In addition, it contains the documents of the background set, without annotations.</p><p>The PharmaCoNER corpus consists of:</p><ul><li>Manually classified clinical case sections derived from Open access Spanish medical publications, named the Spanish Clinical Case Corpus (SPACCC).</li><li>It was manually selected by a practicing oncologist and revised by a clinical documentalist to assure that records were relevant/representative and resembled structure and content relevant to process clinical records.</li><li>The final corpus: 1000 clinical cases 16,504 sentences</li><li>The corpus contains a total of 396,988 words, with an average of 396.2 words per clinical case.</li><li>It covers a range of medical disciplines including oncology, urology, cardiology, pneumology or infections diseases, etc.</li><li>The corpus has been annotated at the mention level by experts in medicinal chemistry and pharmacology following a granular annotation scheme covering four mention types:<ul><li><i>Entity type 1 (NORMALIZABLES)</i>: mentions of chemicals that can be manually normalized to a unique concept identifier (primarily SNOMED-CT).</li><li><i>Entity type 2 (NO_NORMALIZABLES)</i>: mentions of chemicals that could not be normalized manually to a unique concept identifier.</li><li><i>Entity type 3 (PROTEINAS)</i>: mentions of proteins/genes following an adaptation of the BioCreative GPRO track annotation guidelines (includes peptides, peptide hormones & antibodies).</li><li><i>Entity type 4 (UNCLEAR )</i>: cases of general substance class mentions of clinical relevance, including certain pharmaceutical formulations, general treatments, chemotherapy programs, and vaccines.</li><li>Mentions class "<i>UNCLEAR</i>" (not evaluated for the PharmaCoNER track)</li></ul></li></ul><p> </p><p> </p><p><strong>Please, cite: </strong></p><p>A. G. Agirre, M. Marimon, A. Intxaurrondo, O. Rabal, M. Villegas, M. Krallinger, Pharmaconer: Pharmacological substances, compounds and proteins named entity recognition track, in: Proceedings of The 5th Workshop on BioNLP Open Shared Tasks, 2019, pp. 1–10.</p><p> </p><p><strong>Annotation quality</strong></p><p>Inter-annotator agreement: 93% for annotation, 73% for mapping.</p><p>For more information, see the <a href="https://paperswithcode.com/paper/pharmaconer-pharmacological-substances">paper</a>.</p><p> </p><p><strong>Format</strong></p><p>For subtask 1 annotations are distributed in <i>Brat</i> format. (More info at Brat webpage https://brat.nlplab.org/standoff.html)</p><p>For subtask-2, codes are associated with each document are given in a <i>TSV</i> file with the following columns: </p><blockquote><p>filename code</p></blockquote><p> </p><p><strong>Shared task goal:</strong></p><p>In the two subtasks, the goal is to predict the annotations of the test files (either the ANN files or the TSV with the codes) given only the plain text files. </p><p> </p><p><strong>Resources:</strong></p><ul><li><a href="https://temu.bsc.es/pharmaconer/"><strong>Web</strong></a></li><li><a href="https://www.aclweb.org/anthology/D19-5701.pdf"><strong>Citation</strong></a><strong>: </strong>A. G. Agirre, M. Marimon, A. Intxaurrondo, O. Rabal, M. Villegas, M. Krallinger, Pharmaconer: Pharmacological substances, compounds and proteins named entity recognition track, in: Proceedings of The 5th Workshop on BioNLP Open Shared Tasks, 2019, pp. 1–10.</li><li><a href="https://doi.org/10.5281/zenodo.4271908"><strong>Silver Standard corpus</strong></a></li><li><a href="https://doi.org/10.5281/zenodo.3763276"><strong>Annotation guidelines</strong></a></li><li><a href="https://github.com/TeMU-BSC/PharmaCoNER-Tagger"><strong>PharmaCoNER tagger</strong></a></li><li><a href="https://www.youtube.com/watch?v=B3ZzJl5OMkY"><strong>Youtube video(general setting)</strong></a></li><li><a href="https://www.slideshare.net/MartinKrallinger/pharmaconer-pharmacological-substances-compounds-and-proteins-named-entity-recognition-track-at-bionlpost-workshop-november-4-skycity-rm-2-hong-kong-emnlp2019"><strong>Slides PharmacoNER overview talk at BIONLP-ST / EMNLP </strong></a></li></ul><p>For further information, please visit <a href="https://temu.bsc.es/pharmaconer/">https://temu.bsc.es/pharmaconer/</a> or email us at encargo-pln-life@bsc.es</p><p>Copyright (c) 2018 Secretaría de Estado para el Avance Digital (SEAD)</p><p> </p><p><strong>License</strong></p><p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p><p> </p><p><strong>Contact</strong></p><p>If you have any questions or suggestions, please contact us at:</p><p><br>- Martin Krallinger (<krallinger [dot] martin [at] gmail [dot] com>)</p><p><strong>Additional resources and corpora</strong></p><p>If you are interested in PharmaCoNER, you might want to check out these corpora and resources:</p><ul><li><a href="https://zenodo.org/records/7614764">DisTEMIST</a> (Corpus of disease mentions and normalization to SNOMED CT, same document collection)</li><li><a href="https://zenodo.org/records/8224056">MedProcNER </a>(Corpus of clinical procedure mentions and normalization to SNOMED CT, same document collection)</li><li><a href="https://zenodo.org/records/8413866">SympTEMIST</a> (Corpus of symptoms, signs and findings mentions and normalization, same document collection)</li><li><a href="https://zenodo.org/records/7116201">MEDDOPROF</a> (Corpus of mentions of professions, occupations and working status and normalization, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/8403498">MEDDOPLACE</a> (Corpus of mentions of place-related entity mentions, including departments, nationalities or patient movements etc.. and normalization, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/4279323">MEDDOCAN</a> (Corpus of mentions of Personal Health Identifiers (PHI), modified synthetic verions of the document collection)</li><li><a href="https://zenodo.org/records/3978041">CANTEMIST</a> (Corpus of cancer tumor morphology mentions and normalization, different document collection)</li><li><a href="https://zenodo.org/records/3837305">CodiESp</a> (Corpus of clinical case reportes with assigned clinical codes from ICD10, Spanish version, same document collection)</li><li><a href="https://zenodo.org/records/7684093">LivingNER</a> (Corpus of mentions of species, including human/family members, pathogens, food, etc.. and normalization to NCBI Taxonomy, different document collection with some overlapping documents)</li><li><a href="https://zenodo.org/records/2560344">SPACCC-POS</a> (Corpus of clinical case reports in Spanish annotated with POS-tags, same document collection)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-TOKEN</a> (Corpus of clinical case reports in Spanish annotated with token-tags (word mention boundaries), same document collection)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-SPLIT</a> (Corpus of clinical case reports in Spanish annotated with sentence boundary-tags, same document collection)</li><li><a href="https://zenodo.org/records/5602914">MESINESP-2</a> (Corpus of manually indexed records with DeCS /MeSH terms comprising scientific literature abstracts, clinical trials, and patent abstracts, different document collection)</li></ul>
MEDDOCAN corpus: gold standard annotations for Medical Document Anonymization on Spanish clinical case reports
<p><strong>Intro:</strong></p> <p>Meddocan shared task dataset (divided in train, dev and test). In addition, we include here the Meddocan background set.</p> <p>It contains the training, development and test sets of the Meddocan shared task with Gold Standard annotations.</p> <p>In addition, it contains the documents of the background set, without annotations.</p> <p> </p> <p><strong>Annotation quality</strong></p> <p>Inter-annotator agreement: 98% </p> <p>For more information, see the <a href="http://ceur-ws.org/Vol-2421/MEDDOCAN_overview.pdf">paper</a>. </p> <p> </p> <p><strong>Format:</strong></p> <p>Annotations are distributed in Brat format. See <a href="https://brat.nlplab.org/standoff.html">Brat webpage</a> for more information.</p> <p>In addition, annotations are also distributed in XML format (based on i2b2 XML format).</p> <p>In the <a href="https://temu.bsc.es/meddocan/index.php/resources/">Meddocan webpage</a>, there is a script to convert between MEDDOCAN-Brat, MEDDOCAN-XML, and i2b2 formats.</p> <p> </p> <p><strong>Shared task goal:</strong></p> <p>In the three subtasks, the goal will be to predict the annotations given only the plain text files. </p> <p> </p> <p><strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/meddocan/">Web</a></strong></li> <li><strong>Citation: </strong>Montserrat Marimon et al. “Automatic De-identification of Medical Texts in Spanish: the MEDDOCAN Track, Corpus, Guidelines, Methods and Evaluation of Results.” In: IberLEF@ SEPLN. 2019, pp. 618–638.</li> <li><strong>Silver Standard corpus</strong></li> <li><a href="https://doi.org/10.5281/zenodo.4279337"><strong>Annotation guidelines</strong></a></li> </ul> <p> </p> <p>For further information, please visit <a href="https://temu.bsc.es/meddocan/">https://temu.bsc.es/meddocan/</a> or email us at encargo-pln-life@bsc.es</p> <p>Copyright (c) 2019 Secretaría de Estado para el Avance Digital (SEAD)</p>
Petunia exserta v 1.1.3 annotations
<p>Structural annotations of Petunia exserta 1.1.3 genome</p>
UnityMol Annotations and Measurements Walkthrough video
<p>This video provides more detailed supportive information about using Unitymol.</p> <p>In this brief annotated video we demo the measurement capacities of UnityMol (distance, angle and dihedral angle) as well as its free-form hand-drawn annotation possibility.</p> <p> </p>
FIG. 4 in Annotated checklist of the endemic Tetrapoda species of Iran
FIG. 4. — Tropiocolotes naybandensis Krause, Ahmadzadeh, Moazeni, Wagner & Wilms, 2013. Photo by A. Gholamifard.
FIG. 6 in Annotated checklist of the springtails (Hexapoda: Collembola) of the Collo massif, northeastern Algeria
FIG. 6. — Deutonura zana Deharveng, Zoughailech, Hamra-Kroua & Porco, 2015. Body size: 1 to 1.5 mm.
FIG. 7 in Annotated checklist of the endemic Tetrapoda species of Iran
FIG. 7. — Eumeces persicus Faizi, Rastegar-Pouyani, Rastegar-Pouyani, Nazarov, Heidari, Zangi, Orlova & Poyarkov, 2017. Photo by H. Faizi
Training data for 'Genome annotation with Maker' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial for genome annotation with Maker.</p> <p>It is based on data used in <a href="http://weatherby.genetics.utah.edu/MAKER/wiki/index.php/MAKER_Tutorial_for_WGS_Assembly_and_Annotation_Winter_School_2018">another Maker tutorial</a>.</p> <p>The full genome was <a href="https://www.ncbi.nlm.nih.gov/genome/?term=Schizosaccharomyces%20pombe[Organism]&cmd=DetailsSearch">downloaded from NCBI</a>, and mitochondria sequence removed from it for simplicity.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.