Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,523
datasets available to search
ShareScore release 0.7.1
Dataset results
7,523 results for “Annotation”
FIG. 5 in Non-marine mammals of Togo (West Africa): an annotated checklist
FIG. 5. — Two skulls from Fazao Malkafassa: Tragelaphus eurycerus (Ogilby, 1837) (below) and Tragelaphus gratus P. L. Sclater, 1880 (above). Photograph: Gabriel Hoinsoudé Segniagbeto.
FIG. 14 in Non-marine mammals of Togo (West Africa): an annotated checklist
FIG. 14. — Epomophorus gambianus (Ogilby, 1835) from southern Togo. Photograph: Gabriel Hoinsoudé Segniagbeto.
FIG. 27 in Non-marine mammals of Togo (West Africa): an annotated checklist
FIG. 27. — The apparently extinct Leimacomys büttneri Matschie, 1893: mature forest habitat in the Adéle area (from where the species was collected), and type specimens (body and skull) currently stored at the ZMB. Photograph: Luca Luiselli (A); Jan Decher (B, C, D).
FIG. 4. — A in Non-marine mammals of Togo (West Africa): an annotated checklist
FIG. 4. — A, Philantomba walteri Colyn, Hulselman, Sonet, Ooudé, De Winter, Natta, Tamas Nagy & Verheyen, 2010 male; B, Sylvicapra grimmia (Gray, 1843) from Fazao-Malkafassa. Photograph: Gabriel Hoinsoudé Segniagbeto.
FIG. 21 in Non-marine mammals of Togo (West Africa): an annotated checklist
FIG. 21. — Erythrocebus patas (Schreber, 1775) from Fazao-Malkafassa. Photograph: Gabriel Hoinsoudé Segniagbeto.
FIG. 9. — Crossarchus obscurus F. G. Cuvier, 1825 in Non-marine mammals of Togo (West Africa): an annotated checklist
FIG. 9. — Crossarchus obscurus F. G. Cuvier, 1825 from South-western Togo. Photograph: Gabriel Hoinsoudé Segniagbeto.
FIG. 28 in Non-marine mammals of Togo (West Africa): an annotated checklist
FIG. 28. — Squirrels are regularly traded as traditional medicine items throughout the country. In this case, Xerus erythropus (E. Geoffroy, 1803) traded in the Maritime Region, southern Togo. Photograph: Fabio Petrozzi.
FIG. 3 in Annotated catalogue of brachyuran type specimens (Crustacea, Decapoda, Brachyura) deposited in the Muséum national d'Histoire naturelle, Paris. Part I. Podotremata
FIG. 3. — First page of the handwritten catalogue of Crustacea of Latreille, referenced A 53 1, dated 1814.
Example variant files and corresponding annotations for GnomAD v3.1.1 on a subset of chromosome 22
<p>Example variant files and corresponding hg38 annotations for `chr22:15518158-20127355`.</p> <p>Sources:</p> <ul> <li><a href="http://dx.doi.org/10.1093/nar/gky955">Gencode v34 (hg38)</a></li> <li><a href="https://doi.org/10.1038/s41586-020-2308-7">GnomAD v3.1.1</a></li> </ul>
Annotation results of two human retina organoids
<p>Cell annotation was based on i) marker genes which were extracted from literature and expert knowledge, or (ii) via a transfer learning tool CaSTLe (Lieberman Y, Rokach L, Shay T, 2018) using the Cowan <em>et al</em>. (2020) organoid reference data set.</p> <p>Data preprocessing of both scRNA-seq data sets was done in scanpy (Wolf, F., Angerer, P. & Theis, 2018), and are included as h5ad-files.</p> <p>The annotation results are included in the .csv files</p>
OCTA image dataset with label annotation for quality assessment
<p>This dataset is publish by the research "<em>A Deep Learning-based Quality Assessment and Segmentation System with a Large-scale Benchmark Dataset for Optical Coherence Tomographic Angiography Image</em>"</p> <p>Detail:</p> <p>OCTA image dataset with label annotation for quality assessment. sOCTA-3x3-10k: 10,480 3 × 3 mm<sup>2</sup> superficial vascular layer OCTA (sOCTA) images divided into three classes; sOCTA-6x6-14k: 14,042 6 × 6 mm<sup>2</sup> sOCTA images divided into three classes. </p> <p>GitHub: <a href="https://github.com/shanzha09/COIPS">https://github.com/shanzha09/COIPS</a></p> <p>These datasets are public available, if you use the dataset or our system in your research, please <strong>cite</strong> our paper: <em><code>A Deep Learning-based Quality Assessment and Segmentation System with a Large-scale Benchmark Dataset for Optical Coherence Tomographic Angiography Image</code></em>.</p> <p>arXiv:<a href="https://arxiv.org/abs/2107.10476v1">https://arxiv.org/abs/2107.10476v1</a></p>
Gene and repeat annotation for snowy owl (Bubo scandiacus) and selected species
<p>Here we provide the gene and repeat annotation for snowy owl (<em>Bubo scandiacus</em>), in addition to gene and repeat annotation done for some species this was compared to. It is unfortunately currently not possible to upload repeat annotation tracks to an international nucleotide sequence database such as ENA. While uploading the gene annotation is possible, some of the cross references to different databases in the functional annotation are removed. Further, the names of the entries in the publicly available genome assemblies on ENA have different names than what is found in the annotation tracks here, so we also provide the FASTA files for the snowy owl assemblies (bBubSca1.1.hap1.fasta.gz and bBubSca1.1.hap2.fasta.gz). Ideally, all this should have been available via ENA.</p> <p>We annotated the snowy owl genome assemblies, in addition to downy woodpecker (<em>Dryobates pubescens</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCA_014839835.1">GCA_014839835.1</a>), Northern Carmine bee-eater (<em>Merops nubicus</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCA_009819595.1">GCA_009819595.1</a>), Northern goshawk (<em>Accipiter gentilis</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCA_929443795.2">GCA_929443795.2</a>) and barn owl (<em>Tyto alba</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCF_018691265.1#/st">GCF_018691265.1</a>), since no genome annotation was publicly available for these species. We used a pre-release version of the EBP-Nor genome annotation pipeline (<a href="https://github.com/ebp-nor/GenomeAnnotation">https://github.com/ebp-nor/GenomeAnnotation</a>). First, AGAT (https://zenodo.org/record/7255559) agat_sp_keep_longest_isoform.pl and agat_sp_extract_sequences.pl were used on the GRCg7b (GCA_016699485.1) chicken genome assembly and annotation to generate one protein (the longest isoform) per gene. Miniprot (Li, 2023) was used to align the proteins to the curated assemblies. UniProtKB/Swiss-Prot (Consortium et al., 2022) release 2022_03 in addition to the vertebrata part of OrthoDB v11 (Kuznetsov et al., 2022) were also aligned separately to the assemblies. Red (Girgis, 2015) was run via redmask (<a href="https://github.com/nextgenusfs/redmask">https://github.com/nextgenusfs/redmask</a>) on the snowy owl assemblies to mask repetitive areas (we used the soft-masked genome assemblies available at NCBI for the other species). GALBA (Brůna et al., 2023; Buchfink et al., 2015; Hoff and Stanke, 2018; Li, 2023; Stanke et al., 2006) was run with the chicken proteins using the miniprot mode on the masked assemblies. The funannotate-runEVM.py script from Funannotate was used to run EvidenceModeler (Haas et al., 2008) on the alignments of chicken proteins, UniProtKB/Swiss-Prot proteins, vertebrata proteins and the predicted genes from GALBA. The resulting predicted proteins were compared to the protein repeats that Funannotate distributes using DIAMOND blastp and the predicted genes were filtered based on this comparison using AGAT. The filtered proteins were compared to the UniProtKB/Swiss-Prot release 2022_03 using DIAMOND (Buchfink et al., 2015) blastp to find gene names and InterProScan was used to discover functional domains. AGATs agat_sp_manage_functional_annotation.pl was used to attach the gene names and functional annotations to the predicted genes. EMBLmyGFF3 (Norling et al., 2018) was used to combine the fasta files and GFF3 files into a EMBL format for submission to ENA. These files end in gff.gz (the ones ending in fa.out.gff.gz are repeat annotations), proteins.fa.gz and mrna.fa.gz. </p> <p>All species in this study downy woodpecker (<em>Dryobates pubescens</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCA_014839835.1">GCA_014839835.1</a>), Northern Carmine bee-eater (<em>Merops nubicus</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCA_009819595.1">GCA_009819595.1</a>), Northern goshawk (<em>Accipiter gentilis</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCA_929443795.2">GCA_929443795.2</a>), barn owl (T<em>yto alba</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCF_018691265.1#/st">GCF_018691265.1</a>), chicken (<em>Gallus gallus</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCF_016699485.2/">GCF_016699485.2</a>), zebra finch (<em>Taeniopygia guttat</em>a; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCA_003957565.4">GCA_003957565.4</a>) and California condor (<em>Gymnogyps californianus</em>; <a href="https://www.ncbi.nlm.nih.gov/assembly/GCF_018139145.2">GCF_018139145.2</a>) in addition to hap1 of snowy owl was repeat masked with a bird-specific library from <a href="https://www.pnas.org/doi/abs/10.1073/pnas.1616702114">https://www.pnas.org/doi/abs/10.1073/pnas.1616702114</a>, provided by Alexander Suh. These files are named such as MerNubi.fa.out.gff.gz, MerNubi.fa.masked.gz and MerNubi.fna.cat.gz. </p> <p>We have also included the species specific repeat library as generated by RepeatModeler running on hap1 of snowy owl. This is called bBubSca1.1.hap1.repeatlibrary.fa.gz, with the files bBubSca1.1.hap1.fasta.masked.gz, bBubSca1.1.hap1.fasta.out.gff.gz, <span>bBubSca1.1.hap1.divsum.gz </span>and bBubSca1.1.hap1.fasta.cat.gz resulting from running RepeatMasker one hap1 using that library.</p> <p>From the Genespace analyses we have included all files including OrthoFinder results. This is found in the file genespace.tgz.</p>
Supplementary Data for AGouTI - flexible Annotation of Genomic and Transcriptomic Intervals
<p>The data allows to replicate the use-case scenario described in the manuscript "AGouTI - flexible Annotation of Genomic and Transcriptomic Intervals" by Jan G. Kosiński and Marek Żywicki.</p>
A crowdsourced dataset of aerial images with annotated solar photovoltaic arrays and installation metadata
<p><strong>Summary</strong></p> <p>Photovoltaic (PV) energy generation plays a crucial role in the energy transition. Small-scale, residential PV installations are deployed at an unprecedented pace, and their safe integration into the grid necessitates up-to-date, high-quality information. Overhead imagery is increasingly used to improve the knowledge of residential PV installations with machine learning models capable of automatically mapping these installations. However, these models cannot be reliably transferred from one region or imagery source to another without incurring a decrease in accuracy. To address this issue, known as distribution shift, and foster the development of PV array mapping pipelines, we propose a dataset containing aerial images, segmentation masks, and installation metadata. We provide installation metadata for more than 28000 installations. We provide ground truth segmentation masks for 13000 installations, including 7000 with annotations for two different image providers. Finally, we provide installation metadata that matches the annotation for more than 8000 installations. Dataset applications include end-to-end PV registry construction, robust PV installations mapping, and analysis of crowdsourced datasets.</p> <p>This dataset contains the complete records associated with the article "A crowdsourced dataset of aerial images of solar panels, their segmentation masks, and characteristics", published in Scientific data. The article is accessible here : <a href="https://www.nature.com/articles/s41597-023-01951-4">https://www.nature.com/articles/s41597-023-01951-4</a> These complete records consist of:</p> <ol> <li>The complete training dataset containing RGB overhead imagery, segmentation masks and metadata of PV installations (folder <strong>bdappv</strong>),</li> <li>The raw crowdsourcing data, and the postprocessed data for replication and validation (folder <strong>data</strong>).</li> </ol> <p><strong>Data records</strong></p> <p>Folders are organized as follows:</p> <ul> <li><strong>bdappv/</strong> Root data folder <ul> <li><strong>google / ign:</strong> One folder for each campaign <ul> <li><strong>img/</strong>: Folder containing all the images presented to the users. This folder contains 28807 images for Google and 17325 images for IGN.</li> <li><strong>mask/</strong>: Folder containing all segmentations masks generated from the polygon annotations of the users. This folder contains 13303 masks for Google and 7686 masks for IGN.</li> </ul> </li> <li><em>metadata.csv</em> The <code>.csv</code> file with the installations' metadata.</li> </ul> </li> </ul> <p> </p> <ul> <li><strong>data/ </strong>Root data folder <ul> <li><strong>raw/</strong> Folder containing the raw crowdsourcing data and raw metadata; <ul> <li><em>input-google.json</em>: <code>.json </code>input data data containing all information on images and raw annotators’ contributions for both phases (clicks and polygons) during the first annotation campaign;</li> <li><em>input-ign.json</em>:<em> </em><code>.json </code>input data containing all information on images and raw annotators’ contributions for both phases (clicks and polygons) during the second annotation campaign;</li> <li><em>raw-metadata.json</em>: <code>.json </code>output containing the PV systems’ metadata extracted from the BDPV database before filtering. It can be used to replicate the association between the installations and the segmentation masks, as done in the notebook metadata.</li> </ul> </li> <li><strong>replication/</strong> Folder containing the compiled data used to generate the segmentation masks; <ul> <li><strong>campaign-google/campaign-ign</strong>: One folder for each campaign <ul> <li><em>click-analysis.json</em>: <code>.json </code>output on the click analysis, compiling raw input into a few best-guess locations for the PV arrays. This dataset enables the replication of our annotations,</li> <li><em>polygon-analysis.json</em>: <code>.json </code>output of polygon analysis, compiling raw input into a best-guess polygon for the PV arrays.</li> </ul> </li> </ul> </li> <li><strong>validation/</strong> Folder containing the compiled data used for technical validation. <ul> <li><strong>campaign-google/campaign-ign</strong>: One folder for each campaign <ul> <li><em>click-analysis-thres=1.0.json</em>: <code>.json </code>output of the click analysis with a lowered threshold to analyze the effect of the threshold on image classification, as done in the notebook annotation;</li> <li><em>polygon-analysis-thres=1.0.json</em>: <code>.json </code>output of polygon analysis, with a lowered threshold to analyze the effect of the threshold on polygon annotation, as done in the notebook annotations.</li> </ul> </li> <li><em>metadata.csv</em>: the <code>.csv </code>file of filtered installations' metadata.</li> </ul> </li> </ul> </li> </ul> <p><strong>License</strong></p> <p>We extracted the thumbnails contained in the <strong>google/img/</strong> folder using Google Earth Engine API and we generated the thumbnails contained in the <strong>ign/img</strong><strong>/</strong> folder from high resolution tiles downloaded from the online IGN portal accessible here: <a href="https://geoservices.ign.fr/bdortho">https://geoservices.ign.fr/bdortho</a>. Images provided by Google are subjet to Google's terms and conditions. Images provided by the IGN are subject to an open license 2.0.</p> <p>Access the terms and conditions of Google images at this URL: <a href="https://www.google.com/intl/en/help/legalnotices_maps/">https://www.google.com/intl/en/help/legalnotices_maps/</a></p> <p>Access the terms and conditions of IGN images at this URL: <a href="https://www.etalab.gouv.fr/wp-content/uploads/2018/11/open-licence.pdf">https://www.etalab.gouv.fr/wp-content/uploads/2018/11/open-licence.pdf</a></p>
Annotation-based Modeling of Non-functional Requirements and Analysis Results in Domain-driven Design
<p>This repo contains all supplementary data sets that we have created and used throughout this thesis. In particular, it contains<br> - expert interview material (elicitation): consent form and question catalogue for requirements elicitation<br> - expert interview material (evaluation): consent form and task description for expert evaluation<br> - Diagrams related to our modeling concept and Dqualizer<br> - Screenshots of the Domain Story Modeler with our Modeling Concept</p>
Lathyrus sativus LS007 genome assembly and annotation Rbp1.0
<p>Genome assembly of grass pea (<em>Lathyrus sativus</em> L.) genotype LS007, assembled from PromethION nanopore data and polished using Illumina HiSeq PE data. The assembly was annotated using the mikado-minos pipeline developed by the Earlham Institute. Also included is a separate annotation track for repeat sequences produced using the DANTE pipleline. </p> <p> </p> <p>For any questions regarding this dataset, contact peter.emmrich@jic.ac.uk</p> <p> </p> <p>Note: ctg14433 has been manually corrected based on sequenced amplicon data. Files have been updated accordingly.</p> <p> </p> <p><strong>Assembly files:</strong></p> <p>Lsativus_LS007_Rbp1.0.7z - compressed complete assembly without scaffolding. The annotation refers to this assembly</p> <p>Rbp_9 largest HiC scaffolds.7z - compressed fasta file of the largest 9 scaffolds following HiC scaffolding</p> <p>Lsat_LS007_Rbp_chloroplast.fasta - fasta file of the complete LS007 chloroplast genome</p> <p>Lsat_LS007_Rbp_mitochondrion.fasta - fasta file of the complete LS007 mitochondrial genome</p> <p> </p> <p><strong>Annotation tracks:</strong></p> <p>LATSA3860_EIv1.0.annotation.gff3</p> <p>DANTE_transposable_element_protein_domains.gff3</p> <p>Full_length_LTR_retrotransposons.gff3</p> <p>Repeat_annotation_classI_classII_satellites.gff3</p> <p> </p> <p><strong>Annotation FASTA files:</strong></p> <p>LATSA3860_EIv1.0.annotation.gff3.cds.fasta</p> <p>LATSA3860_EIv1.0.annotation.gff3.cdna.fasta</p> <p>LATSA3860_EIv1.0.annotation.gff3.pep.fasta</p> <p> </p> <p><strong>Summaries and statistics:</strong></p> <p>LATSA3860_EIv1.0.annotation.gff3.final_table.tsv</p> <p>LATSA3860_EIv1.0.annotation.gff3.mikado_stats.txt</p> <p>LATSA3860_EIv1.0.annotation.gff3.biotype_conf.summary</p> <p>LATSA3860_EIv1.0.annotation.gff3.final_table.tsv</p> <p>LATSA3860_EIv1.0.annotation.gff3.pep.fasta.functional_annotation.tsv</p> <p>NOT_UPDATED_LATSA3860_EIv1.0.annotation.gff3.metrics.tsv *</p> <p>Blobtools_passed_contigs.txt - list of all contigs of the assembly that pass the BlobTools filter (Streptophyta, 20-100x coverage, >50 kbp) </p> <p> </p> <p>*this file has not been updated to reflect the correction to ctg14433</p>
Brassica napus, rapa, oleracea pangenome assemblies and annotations
<p>All pangenome assemblies and annotations.</p> <p> </p> <p>Rapa_Oleracea_Napus_v1.4_Pangenomes.zip - the annotations in gff3 and fasta, once in the standard EVM naming format, once in the Brassica Consortium standard.</p> <p>Brassica_napus_rapa_oleracea_pangenome_assemblies.zip - the assemblies in fasta. The main references + unassigned pangenome contigs.</p>
Reuse of Model Transformations for Propagating Variability Annotations in Annotative Software Product Lines - Evaluation Data
<p>This package contains all data that was produced for and used in the doctoral thesis for evaluating commutativity of propagating annotations in model-driven product lines.<br> This includes the implementation that conducts the evaluation, the measured results, and the input subjects.</p>
FIG. 60 in An annotated checklist of the tree species of French Guiana, including vernacular nomenclature
FIG. 60. — Vochysiaceae: A, Qualea amapaensis Balslev & S.A.Mori (M.-F. Prévost & D. Sabatier 2755); B, Qualea tricolor Benoist (D. Sabatier 6341); C, Qualea moriboomiorum Marc.-Berti; D, Vochysia densiflora Spruce ex Warm.; E, Vochysia sabatieri Marc.-Berti (D. Sabatier & M.-F. Prévost 4850). © D. Sabatier/IRD.
FIG. 59 in An annotated checklist of the tree species of French Guiana, including vernacular nomenclature
FIG. 59. — Verbenaceae:A, Citharexylum macrophyllum Poir. (M.-F. Prévost 1404). Violaceae: B, C, Leonia glycycarpa Ruiz & Pav. (D. Sabatier 3502); D, Paypayrola hulkiana Pulle (M.-F. Prévost et al. 4587); E, Rinorea falcata (Mart. ex Eichler) Kuntze (D. Sabatier 5573). A, C, © M.-F. Prévost/IRD; B, D, E, © D. Sabatier/IRD.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.