Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7,523

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

7,523 results for “Annotation”

Learn how ShareScore rates datasets ↗
zenodo40/100

Figure 9 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters

Figure 9. Distribution of the members of Gammaridae (Gammarus partly, species in alphabetical order) in Turkish inland waters.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Figure 6 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters

Figure 6. Distribution of the members of Gammaridae (Gammarus partly, species in alphabetical order) in Turkish inland waters.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Figure 7 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters

Figure 7. Distribution of the members of Gammaridae (Gammarus partly, species in alphabetical order) in Turkish inland waters.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Figure 17 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters

Figure 17. Distribution of the members of Leptocheliidae, Tanaellidae and Tanaididae in Turkish inland waters.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Figure 10 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters

Figure 10. Distribution of the members of Gammaridae (Gammarus partly, species in alphabetical order) in Turkish inland waters.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Figure 2 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters

Figure 2. Distribution of the members of Stenothoidae, Dexaminidae, Aoridae, Bogidiellidae, and Caprellidae in Turkish inland waters.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Figure 5 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters

Figure 5. Distribution of the members of Gammaridae (Amathillina, Dikerogammarus and Echinogammarus) in Turkish inland waters.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Figure 21 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters

Figure 21. Distribution of the members of Dromiidae, Paguridae, Varunidae, Pandalidae, and Penaeidae in Turkish inland waters.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Figure 15 in An updated and annotated checklist of the Malacostraca (Crustacea) species inhabited Turkish inland waters

Figure 15. Distribution of the members of Anthuridae, Cirolanidae, Ligiidae, Tylidae, Sphaeromatidae and Idoteidae in Turkish inland waters.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Genome annotation file for a draft genome assembly for Nucella lapillus

<p><span>A male specimen of wild <em>Nucella lapillus</em>, measuring approximately 1.5&ndash;3 cm in length, was collected <span>from a rocky shore (mid-upper shore) at low tide from near the quay at <span>Portnahaven, Isle of Islay, Argyll and Bute, Scotland (National Grid Reference NR 16614 51966)</span> on <span>26th June 2023. Genomic DNA was extracted from the non-shell tissue of the specimen, and sequenced using PacBio HIFI and Oxford Nanopore technolgies (ONT) long read sequencing platforms. </span></span>The genome assembly was derived from 40.6 Gb of PacBio HiFi reads (read <span>N50, 11291; N90, 9246</span>), and 61.1 Gb of ONT data (read <span>N50, </span>3643<span>; N90, </span>1546).&nbsp; </span>Annotation of protein-coding genes in the cleaned and masked genome assembly of <em>Nucella lapillus</em> was performed using GALBA v1.0.11, an automated pipeline that uses proteins from a closely related species to assist in the training of gene prediction using AUGUSTUS. Proteins from <em>Rapana venosa</em> were provided for this purpose, and the miniprot option was used to perform the protein-to-genome alignments. Functional annotation of predicted protein-coding genes was performed using eggNOG-mapper&nbsp;v2.1.12, and additionally annotated with best hit BLAST results (v2.16.0) against the proteomes of the following marine gastropod species: <em>Rapana venosa</em>, <em>Littorina. saxatilis</em>,&nbsp; <em>Pomocea canaliculata, Stramonita haemastoma</em> and <em>Haliotis rufescens</em>.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Real-World Signed Graphs Annotated for Whole Graph Classification

<p><strong>Description: </strong>this corpus was designed as an experimental benchmark for a task of signed graph classification. It is composed of three datasets derived from external sources and adapted to our needs:</p> <ul> <li><strong>SpaceOrigin Conversations [1]: </strong>set of conversational graphs, each one associated to a situation of verbal abuse vs. normal situation. These conversations model interactions happening in chatrooms hosted by an MMORPG/ The graphs were originally unsigned: we attributed signed to the edges based on the polarity of the exchanged messages.&nbsp;&nbsp;</li> <li><strong>Correlation Clustering Instances [2]: </strong>set of graph generated randomly as instances of the Correlation Clustering problem, which consists in partitioning signed graphs. These graphs are not associated in any class in the original paper. We proposed a class based on certain features of the space of optimal solutions explored in [2].</li> <li><strong>European Parliament Roll-Calls [3]: </strong>vote networks extracted from the activity of French Members of the European Parliament. The original data does not have any class associated to the networks: we proposed one based on the number of political factions identified in each network in [3].&nbsp;</li> </ul> <p>These data were used in [4] in order to train and assess various representation learning methods. The authors proposed Signed Graph2vec, a signed variant of Graph2vec; WSGCN, a whole-graph variant of Signed Graph Convolutional Networks (SGCN), and use an aggregated version of Signed Network Embeddings (SiNE) as a baseline. The article provides more information regarding the properties of the datasets, and how they were constituted.</p> <p><strong>Software: </strong>the software used to train the representation learning methods and classifiers is publicly available online: <a href="https://github.com/CompNet/SWGE">SWGE</a>.</p> <p><strong>References:</strong></p> <ol> <li>Papegnies, &Eacute;.; Labatut, V.; Dufour, R. &amp; Linar&egrave;s, G. Conversational Networks for Automatic Online Moderation. <em>IEEE Transactions on Computational Social Systems, </em>2019<em>, </em>6:38-55. DOI: <a href="http://doi.org/10.1109/TCSS.2018.2887240">10.1109/TCSS.2018.2887240</a> ⟨<a href="https://hal.science/hal-01999546">hal-01999546</a>⟩</li> <li>Arınık, N.; Figueiredo, R. &amp; Labatut, V. Multiplicity and Diversity: Analyzing the Optimal Solution Space of the Correlation Clustering Problem on Complete Signed Graphs. <em>Journal of Complex Networks, </em>2020<em>, </em>8(6):cnaa025. DOI: <a href="http://doi.org/10.1093/comnet/cnaa025">10.1093/comnet/cnaa025</a> ⟨<a href="https://hal.science/hal-02994011">hal-02994011</a>⟩</li> <li>Arınık, N.; Figueiredo, R. &amp; Labatut, V. Multiple partitioning of multiplex signed networks: Application to European parliament votes. <em>Social Networks, </em>2020<em>, </em>60:83-102. DOI: <a href="http://doi.org/10.1016/j.socnet.2019.02.001">10.1016/j.socnet.2019.02.001</a> ⟨<a href="https://hal.science/hal-02082574">hal-02082574</a>⟩</li> <li>C&eacute;cillon, N.; Labatut, V.; Dufour, R. &amp; Arınık, N. Whole-Graph Representation Learning For the Classification of Signed Networks. <em>IEEE Access</em>, 2024, 12:151303-151316. DOI:&nbsp;<a href="https://dx.doi.org/10.1109/ACCESS.2024.3472474">10.1109/ACCESS.2024.3472474</a>&nbsp;<a href="https://hal.archives-ouvertes.fr/hal-04712854" rel="nofollow">⟨hal-04712854⟩</a></li> </ol> <p><strong>Funding: </strong>part of this work was funded by a grant from the <em>Provence-Alpes-C&ocirc;te-d'Azur</em>&nbsp;region (PACA, France) and the&nbsp;<em>Nectar de Code</em> company.</p> <p><strong>Citation: </strong>If you use this data or the associated source code, please cite article [4]:</p> <p><code>@Article{Cecillon2024,</code><br><code>&nbsp; author &nbsp; &nbsp;= {C&eacute;cillon, No&eacute; and Labatut, Vincent and Dufour, Richard and Arınık, Nejat},</code><br><code>&nbsp; title &nbsp; &nbsp; = {Whole-Graph Representation Learning For the Classification of Signed Networks},</code><br><code>&nbsp; journal &nbsp; = {IEEE Access},</code><br><code>&nbsp; year &nbsp; &nbsp; &nbsp;= {2024},</code><br><code>&nbsp; volume&nbsp; &nbsp; = {12},</code><br><code>&nbsp; pages&nbsp; &nbsp; &nbsp;= {151303-151316},</code><br><code>&nbsp; doi &nbsp; &nbsp; &nbsp; = {10.1109/ACCESS.2024.3472474},</code><br><code>}</code></p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Trypanosoma cruzi DM28c_2018 annotation with UTR

<p>UTR regions were annotated to the T. cruzi genome 2018 (https://tritrypdb.org/common/downloads/release-68/TcruziDm28c2018/gff/data/). We used custom script for the 5' UTR and peaks2UTR - https://academic.oup.com/bioinformatics/article/39/3/btad112/7067741 for the process.&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Data folders for the Annotation part of OBRWR

<p>Phosphoproteomics experiments were conducted in MDCK cells (canis lupus).</p> <p>Some proteomics protein identification assays were performed in mus musculus.</p> <p>But interaction data (both Direct and Kinase-Phosphatase/Substrate) is mainly available in homo sapiens.</p> <p>Therefore protein ids were transferred from species of origine to homo sapiens.</p> <p>Using PANTHER's Quest for Orthologs dataset, uniprot proteom files and blast.</p>

opencc-by-4.0Dec 2024View details →
Figshare40/100

Silicodata: An Annotated Benchmark CXR Dataset for Silicosis Detection

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Mouse Annotated Dataset

<p>Contains 5 minutes clips from a 2 day recording of 10 mice (5*SWISS + 5*C57BL/6). For each video there is a .csv file that contains the manual annotation of every frame. The data was acquired with EthoProfile, a 10 cage rack developed as part of a research project, and used to train and test the DeepEthoProfile annotation software.</p> <p>&nbsp;</p> <p>Version update:</p> <p>Additionally, files ending with "_map.csv" have been processed from the original data files to contain exactly 8 annotations.</p> <p>Files ending with "_map_new.csv" are a reviewed version with more consistent annotaions.</p> <p>Files ending with "_map_new_v2.csv" are a reviewed version with an additional behavior category, 'None', for the frames where the mouse is turned away from the camera but not resting.</p>

openmit-licenseDec 2022View details →
zenodo40/100

Syzygium jambos genome assembly and annotation

<p><em>De novo</em> genome assembly and annotation of <em>Syzygium jambos</em>.</p> <p>The following files are available:</p> <ul> <li>sjam.fa.gz: reference genome sequence in fasta format</li> <li>sjam.gff3.gz: gene annotation in GFF3 format</li> <li>sjam.gtf.gz: gene annotation in GTF format</li> <li>sjam.tx.fa.gz: transcript sequences in fasta format</li> <li>sjam.cds.fa.gz: coding sequences in fasta format</li> <li>sjam.prot.fa.gz: protein sequences in fasta format</li> <li>sjam.tsv.gz: gene functional annotation in TSV format</li> <li>sjam.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Nicotiana tomentosiformis genome assembly and annotation

<p><em>De novo</em> genome assembly and annotation of <em>Nicotiana tomentosiformis</em>.</p> <p>The following files are available:</p> <ul> <li>ntom.fa.gz: reference genome sequence in fasta format</li> <li>ntom.gff3.gz: gene annotation in GFF3 format</li> <li>ntom.gtf.gz: gene annotation in GTF format</li> <li>ntom.tx.fa.gz: transcript sequences in fasta format</li> <li>ntom.cds.fa.gz: coding sequences in fasta format</li> <li>ntom.prot.fa.gz: protein sequences in fasta format</li> <li>ntom.tsv.gz: gene functional annotation in TSV format</li> <li>ntom.rt.fa.gz: retrotransposon sequences in fasta format</li> <li>ntom.rt.gff3.gz: retrotransposon annotation on GFF3 format</li> <li>ntom.rt.tsv.gz: retrotransposon annotation in TSV format</li> <li>ntom.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Syzygium malaccense genome assembly and annotation

<p><em>De novo</em> genome assembly and annotation of <em>Syzygium malaccense</em>.</p> <p>The following files are available:</p> <ul> <li>smal.fa.gz: reference genome sequence in fasta format</li> <li>smal.gff3.gz: gene annotation in GFF3 format</li> <li>smal.gtf.gz: gene annotation in GTF format</li> <li>smal.tx.fa.gz: transcript sequences in fasta format</li> <li>smal.cds.fa.gz: coding sequences in fasta format</li> <li>smal.prot.fa.gz: protein sequences in fasta format</li> <li>smal.tsv.gz: gene functional annotation in TSV format</li> <li>smal.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Syzygium aqueum genome assembly and annotation

<p><em>De novo </em>genome assembly and annotation of <em>Syzygium aqueum</em>.</p> <p>The following files are available:</p> <ul> <li>saqu.fa.gz: reference genome sequence in fasta format</li> <li>saqu.gff3.gz: gene annotation in GFF3 format</li> <li>saqu.gtf.gz: gene annotation in GTF format</li> <li>saqu.tx.fa.gz: transcript sequences in fasta format</li> <li>saqu.cds.fa.gz: coding sequences in fasta format</li> <li>saqu.prot.fa.gz: protein sequences in fasta format</li> <li>saqu.tsv.gz: gene functional annotation in TSV format</li> <li>saqu.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Nicotiana tabacum genome assembly and annotation

<p><em>De novo</em> genome assembly and annotation of <em>Nicotiana tabacum</em>.</p> <p>The following files are available:</p> <ul> <li>ntab.fa.gz: reference genome sequence in fasta format</li> <li>ntab.gff3.gz: gene annotation in GFF3 format</li> <li>ntab.gtf.gz: gene annotation in GTF format</li> <li>ntab.tx.fa.gz: transcript sequences in fasta format</li> <li>ntab.cds.fa.gz: coding sequences in fasta format</li> <li>ntab.prot.fa.gz: protein sequences in fasta format</li> <li>ntab.tsv.gz: gene functional annotation in TSV format</li> <li>ntab.rt.fa.gz: retrotransposon sequences in fasta format</li> <li>ntab.rt.gff3.gz: retrotransposon annotation on GFF3 format</li> <li>ntab.rt.tsv.gz: retrotransposon annotation in TSV format</li> <li>ntab.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>

opencc-by-4.0Aug 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record