Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

302

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

302 results for “protein sequence”

Learn how ShareScore rates datasets ↗
zenodo32/100

Fig. 3 in Analysis of major sperm proteins in two nematode species from two classes, Enoplus brevis (Enoplea, Enoplida) and Panagrellus redivivus (Chromadorea, Rhabditida), reveals similar localization, but less homology of protein sequences than expected for Nematoda phylum

Fig. 3 Schematic representation of P. redivivus spermatozoa based on transmission electron microscopy. a Morphology of immature and mature spermatozoa. Immature spermatozoon is an unpolarized cell with nucleus devoid of nuclear envelope, mitochondria, and membranous organelles. Mature spermatozoon in female reproductive system is a bipolar cell with anterior pseudopodium and posterior main cell body containing chromatin, mitochondria, and membranous organelles that attached to cell membrane and open to the exterior via pores. Reproduced from Zograf (2014) with the permission from copyright holder (Russian Journal of Nematology). b Chain of conjugated mature spermatozoa in female reproductive system. Abbreviations: N, nucleus; mt, mitochondria; mo, membranous organelles; ch, nuclear chromatin; ps, pseudopodium; mcb, mail cell body

opennotspecifiedSep 2021View details →
zenodo32/100

Fig. 1 in Analysis of major sperm proteins in two nematode species from two classes, Enoplus brevis (Enoplea, Enoplida) and Panagrellus redivivus (Chromadorea, Rhabditida), reveals similar localization, but less homology of protein sequences than expected for Nematoda phylum

Fig. 1 Phylogeny of nematodes and MSP-based sperm motility. Phylogenetic relationships within phylum Nematoda derived primarily from SSU rDNA sequence data are given according to De Ley and Blaxter (2002). Suborders of the order Rhabditida, in which representatives highly homologous MSPs are found at DNA, RNA, or protein levels, are marked by underlining. Taxa whose species used in this study are marked with asterisks. Orders Trefusi- ida, Isolaimida, Dioctophyma- tida, Muspiceida, Marimermith- ida, and Desmoscolecida are not shown in this tree

opennotspecifiedSep 2021View details →
zenodo32/100

Fig. 8 in Analysis of major sperm proteins in two nematode species from two classes, Enoplus brevis (Enoplea, Enoplida) and Panagrellus redivivus (Chromadorea, Rhabditida), reveals similar localization, but less homology of protein sequences than expected for Nematoda phylum

Fig. 8 Putative MSPs those are most similar to peptide antigen. a P. redivivus MSPs aligned with peptide antigen. Protein sequences (Pan_g61.t1, Pan_g6018.t1, Pan_g6424.t1, Pan_g9068.t1, Pan_ g19433.t1, and Pan_g21178.t1) were found by Blast using peptide

opennotspecifiedSep 2021View details →
zenodo32/100

Fig. 4 in Analysis of major sperm proteins in two nematode species from two classes, Enoplus brevis (Enoplea, Enoplida) and Panagrellus redivivus (Chromadorea, Rhabditida), reveals similar localization, but less homology of protein sequences than expected for Nematoda phylum

Fig. 4 Immunolocalization of MSP in P. redivivus sperm. a Immature spermatozoa extracted from male. MSP localizes in granules. In some cells, MSP has strongest signals in the periphery (arrowheads) (scale bar 10 µm). b Chain of mature spermatozoa extracted from female.

opennotspecifiedSep 2021View details →
zenodo32/100

Plasmid sequences for: Recombinant venom proteins in insect seminal fluid reduces female lifespan

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
dryad32/100

Data from: Evaluating statistical multiple sequence alignment in comparison to other alignment methods on protein data sets

The estimation of multiple sequence alignments of protein sequences is a basic step in many bioinformatics pipelines, including protein structure prediction, protein family identification, and phylogeny estimation. Statistical co-estimation of alignments and trees under stochastic models of sequence evolution has long been considered the most rigorous technique for estimating alignments and trees, but little is known about the accuracy of such methods on biological benchmarks. We report the results of an extensive study evaluating the most popular protein alignment methods as well as the statistical co-estimation method BAli-Phy on 1192 protein data sets from established benchmarks as well as on 120 simulated data sets. Our study (which used more than 230 CPU years for the BAli-Phy analyses alone) shows that BAli-Phy has better precision and recall (with respect to the true alignments) than the other alignment methods on the simulated data sets, but has consistently lower recall on the biological benchmarks (with respect to the reference alignments) than many of the other methods. In other words, we find that BAli-Phy systematically under-aligns when operating on biological sequence data, but shows no sign of this on simulated data. There are several potential causes for this change in performance, including model misspecification, errors in the reference alignments, and conflicts between structural alignment and evolutionary alignments, and future research is needed to determine the most likely explanation. We conclude with a discussion of the potential ramifications for each of these possibilities.

opencc-zeroDec 2017View details →
zenodo32/100

In-house sequence database of S.pn recombinant protein vaccine target-related genes

<p>In-house sequence database of S.pn recombinant protein vaccine target-related genes.</p>

opencc-by-4.0Jul 2021View details →
zenodo32/100

Sequences used for building HMMER profiles to search for SNAREs proteins

<p>This file contains the four sequence datasets that were used to build HMMER search profiles in order to find all SNAREs protein present in the predicted proteome of Oopsacas minuta: a human SNARE profile, a Qa-SNAREs profile, a Qb-SNAREs profile, a Qc-SNAREs profile.</p>

opencc-by-4.0Jan 2023View details →
zenodo32/100

Case studies from: Sequence assignment validation in protein crystal structure models with checkMySequence

<p>Case studies from&nbsp;&quot;Sequence assignment validation in protein crystal structure models with checkMySequence&quot;</p>

opencc-by-4.0Feb 2023View details →
dryad32/100

Proteome database of 36 million proteins from 4,351 species, including marine microbial sequences

<p>A fasta-formatted database of 36,866,870 predicted proteins representing 4,351 unique species from 117 phyla.</p>

opencc-zeroFeb 2023View details →
zenodo32/100

Protein sequences for 2031 Saccharomyces cerevisiae genome assemblies

<p>Protein sequences for 2031&nbsp;Saccharomyces cerevisiae genome assemblies</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

Optimal control derived sensitivity-enhanced CA-CO mixing sequences for MAS solid-state NMR. Applications in sequential protein backbone assignments.

<p>Raw data pulse sequences and shapes for publication</p> <p><br> ## SEQUENCES ##<br> ./sequences_renamed<br> Pulse programs introduiced in this work. Previous pulse programs can be obtained from https://doi.org/10.5281/zenodo.7016441 or https://optimal-nmr.net/experiments.html</p> <p>## SHAPES ##<br> ./shapes_renamed<br> TROP shaped pulses for homonuclear 13C-13C homonuclear mixing discussed in this work. Heteronuclear shaped pulses can be obtained from https://doi.org/10.5281/zenodo.7016441 or https://optimal-nmr.net/sequences.html</p> <p>## RAW DATA ##<br> to decrease storage demands 3D-processed spectra were deleted and can be recovered using TopSpin command: ftnd 0<br> TopSpin NUS licence is required for processing of the NUS-sampled data. Transformed data can be obtained from authors on request.</p> <p># U-13C,15N,2H,1HN-SH3 sample at 55 kHz MAS<br> ./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/1<br> 1H saturation recovery</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/2<br> 1H hard pulse calibraton</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/3<br> 15N hard pulse calibration</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/4<br> conventional hNH with rampCP optimalization</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/5<br> sensitivity-enhahced se-hNH with TROP optimalization</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/11<br> 2D conventional hNH with rampCP&nbsp;</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/12<br> 2D sensitivtiy-enhanced se-hNH with TROP&nbsp;</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/13<br> CO hard pulse and hCO CP calibration</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/14<br> CA hard pulse and hCA CP calibration</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/27<br> CACO homoTROP power optimalization in sensitivity-enhanced se-hCACOHN experiment</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/28<br> CACO INEPT delay optimalization in &nbsp;conventional hcoCAcoHN experiment</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/31 to 38<br> comparison of hCANH experiment efficiency using combination of coherence transfer methods (rampCP, tmSPICE and TROP) for CAN and NN transfer (see experiment tiles)</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/41 to 45<br> comparison of hCONH experiment efficiency using combination of coherence transfer methods (rampCP, tmSPICE and TROP) for CON and NN transfer (see experiment tiles)</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/57 and 58<br> optimalization of selective CA and CO 90 and 180 pulse</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/59<br> COCA INEPT delay optimalization in &nbsp;conventional hCOcaHN experiment</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/60<br> COCA homoTROP power optimalization in sensitivity-enhanced se-hCOCAHN experiment</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/69<br> conventional 3D hCANH experiment using tmSPICE CAN transfer; with water suppression after first CP, 15% non-uniform sampled</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/70<br> sensitivity-enhanced 3D se-hCANH experiment using TROP transfer; with water suppression after first CP, 15% non-uniform sampled</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/72<br> conventional 3D hCONH experiment using tmSPICE CON transfer; with water suppression after first CP, 15% non-uniform sampled</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/73<br> sensitivity-enhanced 3D se-hCONH experiment using TROP transfer; with water suppression after first CP, 15% non-uniform sampled</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/74<br> sensitivity-enhanced 3D se-hCAcoNH experiment using TROP transfer and homoTROP; with water suppression after first CP, 15% non-uniform sampled</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/75<br> conventional se-hcoCAcoNH experiment using tmSPICE CN transfer and INEPT &lsquo;out-and-back&rsquo; coCAco transfer; with water suppression before NH transfer, 15% non-uniform sampled</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/76<br> sensitivity-enhanced 3D se-hCOcaNH experiment using TROP transfer and homoTROP; with water suppression after first CP, 15% non-uniform sampled</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/77<br> conventional 3D hCOcaNH experiment using tmSPICE CN transfer and INEPT &lsquo;complete forward&rsquo; COca; with water suppression after first CP, 15% non-uniform sampled</p> <p><br> # U-13C,15N glycine at 16.5 kHz MAS<br> ./MAS_COCa_seTesting/62<br> sensitivity-enhanced se-hCACO experiment with homoTROP pulse without diagonal-phase control</p> <p>./MAS_COCa_seTesting/63<br> sensitivity-enhanced se-hCOCA experiment with homoTROP pulse without diagonal-phase control</p> <p><br> # U-13C,15N-fMLF at 20 kHz MAS<br> ./jb.20211102_fMLF_3.2/2<br> 13C direct excitation spectra</p> <p>./jb.20211102_fMLF_3.2/3<br> 13C and 1H hard-pulse calibration and HC CP optimalization</p> <p>./jb.20211102_fMLF_3.2/4<br> 15N hard-pulse calibration and HC CP optimalization</p> <p>./jb.20211102_fMLF_3.2/5<br> NCA rampCP optimalization</p> <p>./jb.20211102_fMLF_3.2/6<br> NCO rampCP optimalization</p> <p>./jb.20211102_fMLF_3.2/79<br> sensitivity-enhanced se-hNCACO with TROP pulses</p> <p>./jb.20211102_fMLF_3.2/83<br> sensitivity-enhanced se-hNCOCA with TROP pulses</p> <p>./jb.20211102_fMLF_3.2/87<br> conventional hNCACO with rampCP and DREAM mixing</p> <p>./jb.20211102_fMLF_3.2/87<br> conventional hNCOCA with rampCP and DREAM mixing</p>

opencc-by-4.0Feb 2023View details →
zenodo32/100

EPSAPG: A Pipeline Combining MMseqs2 and PSI-BLAST to Quickly Generate Extensive Protein Sequence Alignment Profiles

<p>This repository contains all data, queries, and search results used in the analysis of</p> <ul> <li>Arab, Issar. &ldquo;<strong>EPSAPG: A Pipeline Combining MMseqs2 and PSI-BLAST to Quickly Generate</strong><br><strong>Extensive Protein Sequence Alignment Profiles</strong>.&rdquo; IEEE/ACM 10th International Conference on<br>Big Data Computing, Applications and Technologies (BDCAT &rsquo;23), December 4&ndash;7, 2023,<br>Taormina, Messina, Italy. <a href="https://doi.org/10.1145/3632366.3632384">doi.org/10.1145/3632366.3632384</a></li> </ul>

opencc-by-4.0Jun 2023View details →
zenodo32/100

Enhancing Protein Sequence Annotation in Viral Genomics Using Large Language Models and Soft Alignments.

<p>List of 200 most abundant VOG descriptions.</p>

opencc-by-4.0Jul 2023View details →
zenodo32/100

Protein sequences matching TIGRFAM models

<p>A set of 411 large and 3,001 small TIGRFAM&nbsp;protein families originally used for benchmarking sequence clustering programs. Large families contain at least 20,000 sequences with a genus label, whereas small contain fewer than 20,000. Sequences were&nbsp;downloaded from <a href="https://www.ncbi.nlm.nih.gov/genome/annotation_prok/tigrfams/">NCBI</a>&nbsp;and renamed by their original name (accession number)&nbsp;followed by their semi-colon separated NCBI taxonomy of the originating organism. For example, the first sequence in TIGRFAM_large/TIGRFAM00005.fas.gz is named:</p><blockquote><p>WP_000005837.1 RluA family pseudouridine synthase, partial [Bacillus anthracis];TIGR00005(group);cellular organisms(no rank);Bacteria(superkingdom);Terrabacteria group(clade);Firmicutes(phylum);Bacilli(class);Bacillales(order);Bacillaceae(family);Bacillus(genus);Bacillus cereus group(species group);Bacillus anthracis(species)</p></blockquote>

opencc-by-4.0Oct 2023View details →
dryad32/100

Data from: A hypervariable mitochondrial protein coding sequence associated with geographical origin in a cosmopolitan bloom-forming alga, Heterosigma akashiwo

Open the record for dataset details and reuse information.

publicMar 2017View details →
dryad32/100

Proteome database of 36 million proteins from 4,351 species, including marine microbial sequences

Open the record for dataset details and reuse information.

publicFeb 2023View details →
dryad32/100

Data from: Evaluating statistical multiple sequence alignment in comparison to other alignment methods on protein data sets

Open the record for dataset details and reuse information.

publicOct 2018View details →
dryad32/100

Data from: The evolution of heat shock protein sequences, cis-regulatory elements, and expression profiles in the eusocial Hymenoptera

Open the record for dataset details and reuse information.

publicFeb 2016View details →
dryad32/100

FASTA file of sequences of identified proteins in Anastrepha ludens reproductive tissues

Open the record for dataset details and reuse information.

publicJun 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record