Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,523
datasets available to search
ShareScore release 0.7.1
Dataset results
7,523 results for “Annotation”
Figure 9 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters
Figure 9. Distribution of the members of Gammaridae (Gammarus partly, species in alphabetical order) in Turkish inland waters.
Figure 6 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters
Figure 6. Distribution of the members of Gammaridae (Gammarus partly, species in alphabetical order) in Turkish inland waters.
Figure 7 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters
Figure 7. Distribution of the members of Gammaridae (Gammarus partly, species in alphabetical order) in Turkish inland waters.
Figure 17 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters
Figure 17. Distribution of the members of Leptocheliidae, Tanaellidae and Tanaididae in Turkish inland waters.
Figure 10 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters
Figure 10. Distribution of the members of Gammaridae (Gammarus partly, species in alphabetical order) in Turkish inland waters.
Figure 2 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters
Figure 2. Distribution of the members of Stenothoidae, Dexaminidae, Aoridae, Bogidiellidae, and Caprellidae in Turkish inland waters.
Figure 5 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters
Figure 5. Distribution of the members of Gammaridae (Amathillina, Dikerogammarus and Echinogammarus) in Turkish inland waters.
Figure 21 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters
Figure 21. Distribution of the members of Dromiidae, Paguridae, Varunidae, Pandalidae, and Penaeidae in Turkish inland waters.
Figure 15 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters
Figure 15. Distribution of the members of Anthuridae, Cirolanidae, Ligiidae, Tylidae, Sphaeromatidae and Idoteidae in Turkish inland waters.
Genome annotation file for a draft genome assembly for Nucella lapillus
<p><span>A male specimen of wild <em>Nucella lapillus</em>, measuring approximately 1.5–3 cm in length, was collected <span>from a rocky shore (mid-upper shore) at low tide from near the quay at <span>Portnahaven, Isle of Islay, Argyll and Bute, Scotland (National Grid Reference NR 16614 51966)</span> on <span>26th June 2023. Genomic DNA was extracted from the non-shell tissue of the specimen, and sequenced using PacBio HIFI and Oxford Nanopore technolgies (ONT) long read sequencing platforms. </span></span>The genome assembly was derived from 40.6 Gb of PacBio HiFi reads (read <span>N50, 11291; N90, 9246</span>), and 61.1 Gb of ONT data (read <span>N50, </span>3643<span>; N90, </span>1546). </span>Annotation of protein-coding genes in the cleaned and masked genome assembly of <em>Nucella lapillus</em> was performed using GALBA v1.0.11, an automated pipeline that uses proteins from a closely related species to assist in the training of gene prediction using AUGUSTUS. Proteins from <em>Rapana venosa</em> were provided for this purpose, and the miniprot option was used to perform the protein-to-genome alignments. Functional annotation of predicted protein-coding genes was performed using eggNOG-mapper v2.1.12, and additionally annotated with best hit BLAST results (v2.16.0) against the proteomes of the following marine gastropod species: <em>Rapana venosa</em>, <em>Littorina. saxatilis</em>, <em>Pomocea canaliculata, Stramonita haemastoma</em> and <em>Haliotis rufescens</em>.</p>
Real-World Signed Graphs Annotated for Whole Graph Classification
<p><strong>Description: </strong>this corpus was designed as an experimental benchmark for a task of signed graph classification. It is composed of three datasets derived from external sources and adapted to our needs:</p> <ul> <li><strong>SpaceOrigin Conversations [1]: </strong>set of conversational graphs, each one associated to a situation of verbal abuse vs. normal situation. These conversations model interactions happening in chatrooms hosted by an MMORPG/ The graphs were originally unsigned: we attributed signed to the edges based on the polarity of the exchanged messages. </li> <li><strong>Correlation Clustering Instances [2]: </strong>set of graph generated randomly as instances of the Correlation Clustering problem, which consists in partitioning signed graphs. These graphs are not associated in any class in the original paper. We proposed a class based on certain features of the space of optimal solutions explored in [2].</li> <li><strong>European Parliament Roll-Calls [3]: </strong>vote networks extracted from the activity of French Members of the European Parliament. The original data does not have any class associated to the networks: we proposed one based on the number of political factions identified in each network in [3]. </li> </ul> <p>These data were used in [4] in order to train and assess various representation learning methods. The authors proposed Signed Graph2vec, a signed variant of Graph2vec; WSGCN, a whole-graph variant of Signed Graph Convolutional Networks (SGCN), and use an aggregated version of Signed Network Embeddings (SiNE) as a baseline. The article provides more information regarding the properties of the datasets, and how they were constituted.</p> <p><strong>Software: </strong>the software used to train the representation learning methods and classifiers is publicly available online: <a href="https://github.com/CompNet/SWGE">SWGE</a>.</p> <p><strong>References:</strong></p> <ol> <li>Papegnies, É.; Labatut, V.; Dufour, R. & Linarès, G. Conversational Networks for Automatic Online Moderation. <em>IEEE Transactions on Computational Social Systems, </em>2019<em>, </em>6:38-55. DOI: <a href="http://doi.org/10.1109/TCSS.2018.2887240">10.1109/TCSS.2018.2887240</a> ⟨<a href="https://hal.science/hal-01999546">hal-01999546</a>⟩</li> <li>Arınık, N.; Figueiredo, R. & Labatut, V. Multiplicity and Diversity: Analyzing the Optimal Solution Space of the Correlation Clustering Problem on Complete Signed Graphs. <em>Journal of Complex Networks, </em>2020<em>, </em>8(6):cnaa025. DOI: <a href="http://doi.org/10.1093/comnet/cnaa025">10.1093/comnet/cnaa025</a> ⟨<a href="https://hal.science/hal-02994011">hal-02994011</a>⟩</li> <li>Arınık, N.; Figueiredo, R. & Labatut, V. Multiple partitioning of multiplex signed networks: Application to European parliament votes. <em>Social Networks, </em>2020<em>, </em>60:83-102. DOI: <a href="http://doi.org/10.1016/j.socnet.2019.02.001">10.1016/j.socnet.2019.02.001</a> ⟨<a href="https://hal.science/hal-02082574">hal-02082574</a>⟩</li> <li>Cécillon, N.; Labatut, V.; Dufour, R. & Arınık, N. Whole-Graph Representation Learning For the Classification of Signed Networks. <em>IEEE Access</em>, 2024, 12:151303-151316. DOI: <a href="https://dx.doi.org/10.1109/ACCESS.2024.3472474">10.1109/ACCESS.2024.3472474</a> <a href="https://hal.archives-ouvertes.fr/hal-04712854" rel="nofollow">⟨hal-04712854⟩</a></li> </ol> <p><strong>Funding: </strong>part of this work was funded by a grant from the <em>Provence-Alpes-Côte-d'Azur</em> region (PACA, France) and the <em>Nectar de Code</em> company.</p> <p><strong>Citation: </strong>If you use this data or the associated source code, please cite article [4]:</p> <p><code>@Article{Cecillon2024,</code><br><code> author = {Cécillon, Noé and Labatut, Vincent and Dufour, Richard and Arınık, Nejat},</code><br><code> title = {Whole-Graph Representation Learning For the Classification of Signed Networks},</code><br><code> journal = {IEEE Access},</code><br><code> year = {2024},</code><br><code> volume = {12},</code><br><code> pages = {151303-151316},</code><br><code> doi = {10.1109/ACCESS.2024.3472474},</code><br><code>}</code></p>
Trypanosoma cruzi DM28c_2018 annotation with UTR
<p>UTR regions were annotated to the T. cruzi genome 2018 (https://tritrypdb.org/common/downloads/release-68/TcruziDm28c2018/gff/data/). We used custom script for the 5' UTR and peaks2UTR - https://academic.oup.com/bioinformatics/article/39/3/btad112/7067741 for the process. </p>
Data folders for the Annotation part of OBRWR
<p>Phosphoproteomics experiments were conducted in MDCK cells (canis lupus).</p> <p>Some proteomics protein identification assays were performed in mus musculus.</p> <p>But interaction data (both Direct and Kinase-Phosphatase/Substrate) is mainly available in homo sapiens.</p> <p>Therefore protein ids were transferred from species of origine to homo sapiens.</p> <p>Using PANTHER's Quest for Orthologs dataset, uniprot proteom files and blast.</p>
Silicodata: An Annotated Benchmark CXR Dataset for Silicosis Detection
Open the record for dataset details and reuse information.
Mouse Annotated Dataset
<p>Contains 5 minutes clips from a 2 day recording of 10 mice (5*SWISS + 5*C57BL/6). For each video there is a .csv file that contains the manual annotation of every frame. The data was acquired with EthoProfile, a 10 cage rack developed as part of a research project, and used to train and test the DeepEthoProfile annotation software.</p> <p> </p> <p>Version update:</p> <p>Additionally, files ending with "_map.csv" have been processed from the original data files to contain exactly 8 annotations.</p> <p>Files ending with "_map_new.csv" are a reviewed version with more consistent annotaions.</p> <p>Files ending with "_map_new_v2.csv" are a reviewed version with an additional behavior category, 'None', for the frames where the mouse is turned away from the camera but not resting.</p>
Syzygium jambos genome assembly and annotation
<p><em>De novo</em> genome assembly and annotation of <em>Syzygium jambos</em>.</p> <p>The following files are available:</p> <ul> <li>sjam.fa.gz: reference genome sequence in fasta format</li> <li>sjam.gff3.gz: gene annotation in GFF3 format</li> <li>sjam.gtf.gz: gene annotation in GTF format</li> <li>sjam.tx.fa.gz: transcript sequences in fasta format</li> <li>sjam.cds.fa.gz: coding sequences in fasta format</li> <li>sjam.prot.fa.gz: protein sequences in fasta format</li> <li>sjam.tsv.gz: gene functional annotation in TSV format</li> <li>sjam.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>
Nicotiana tomentosiformis genome assembly and annotation
<p><em>De novo</em> genome assembly and annotation of <em>Nicotiana tomentosiformis</em>.</p> <p>The following files are available:</p> <ul> <li>ntom.fa.gz: reference genome sequence in fasta format</li> <li>ntom.gff3.gz: gene annotation in GFF3 format</li> <li>ntom.gtf.gz: gene annotation in GTF format</li> <li>ntom.tx.fa.gz: transcript sequences in fasta format</li> <li>ntom.cds.fa.gz: coding sequences in fasta format</li> <li>ntom.prot.fa.gz: protein sequences in fasta format</li> <li>ntom.tsv.gz: gene functional annotation in TSV format</li> <li>ntom.rt.fa.gz: retrotransposon sequences in fasta format</li> <li>ntom.rt.gff3.gz: retrotransposon annotation on GFF3 format</li> <li>ntom.rt.tsv.gz: retrotransposon annotation in TSV format</li> <li>ntom.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>
Syzygium malaccense genome assembly and annotation
<p><em>De novo</em> genome assembly and annotation of <em>Syzygium malaccense</em>.</p> <p>The following files are available:</p> <ul> <li>smal.fa.gz: reference genome sequence in fasta format</li> <li>smal.gff3.gz: gene annotation in GFF3 format</li> <li>smal.gtf.gz: gene annotation in GTF format</li> <li>smal.tx.fa.gz: transcript sequences in fasta format</li> <li>smal.cds.fa.gz: coding sequences in fasta format</li> <li>smal.prot.fa.gz: protein sequences in fasta format</li> <li>smal.tsv.gz: gene functional annotation in TSV format</li> <li>smal.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>
Syzygium aqueum genome assembly and annotation
<p><em>De novo </em>genome assembly and annotation of <em>Syzygium aqueum</em>.</p> <p>The following files are available:</p> <ul> <li>saqu.fa.gz: reference genome sequence in fasta format</li> <li>saqu.gff3.gz: gene annotation in GFF3 format</li> <li>saqu.gtf.gz: gene annotation in GTF format</li> <li>saqu.tx.fa.gz: transcript sequences in fasta format</li> <li>saqu.cds.fa.gz: coding sequences in fasta format</li> <li>saqu.prot.fa.gz: protein sequences in fasta format</li> <li>saqu.tsv.gz: gene functional annotation in TSV format</li> <li>saqu.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>
Nicotiana tabacum genome assembly and annotation
<p><em>De novo</em> genome assembly and annotation of <em>Nicotiana tabacum</em>.</p> <p>The following files are available:</p> <ul> <li>ntab.fa.gz: reference genome sequence in fasta format</li> <li>ntab.gff3.gz: gene annotation in GFF3 format</li> <li>ntab.gtf.gz: gene annotation in GTF format</li> <li>ntab.tx.fa.gz: transcript sequences in fasta format</li> <li>ntab.cds.fa.gz: coding sequences in fasta format</li> <li>ntab.prot.fa.gz: protein sequences in fasta format</li> <li>ntab.tsv.gz: gene functional annotation in TSV format</li> <li>ntab.rt.fa.gz: retrotransposon sequences in fasta format</li> <li>ntab.rt.gff3.gz: retrotransposon annotation on GFF3 format</li> <li>ntab.rt.tsv.gz: retrotransposon annotation in TSV format</li> <li>ntab.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.