Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
302
datasets available to search
ShareScore release 0.9.0
Dataset results
302 results for “protein sequence”
Fig. 3 in Analysis of major sperm proteins in two nematode species from two classes, Enoplus brevis (Enoplea, Enoplida) and Panagrellus redivivus (Chromadorea, Rhabditida), reveals similar localization, but less homology of protein sequences than expected for Nematoda phylum
Fig. 3 Schematic representation of P. redivivus spermatozoa based on transmission electron microscopy. a Morphology of immature and mature spermatozoa. Immature spermatozoon is an unpolarized cell with nucleus devoid of nuclear envelope, mitochondria, and membranous organelles. Mature spermatozoon in female reproductive system is a bipolar cell with anterior pseudopodium and posterior main cell body containing chromatin, mitochondria, and membranous organelles that attached to cell membrane and open to the exterior via pores. Reproduced from Zograf (2014) with the permission from copyright holder (Russian Journal of Nematology). b Chain of conjugated mature spermatozoa in female reproductive system. Abbreviations: N, nucleus; mt, mitochondria; mo, membranous organelles; ch, nuclear chromatin; ps, pseudopodium; mcb, mail cell body
Fig. 1 in Analysis of major sperm proteins in two nematode species from two classes, Enoplus brevis (Enoplea, Enoplida) and Panagrellus redivivus (Chromadorea, Rhabditida), reveals similar localization, but less homology of protein sequences than expected for Nematoda phylum
Fig. 1 Phylogeny of nematodes and MSP-based sperm motility. Phylogenetic relationships within phylum Nematoda derived primarily from SSU rDNA sequence data are given according to De Ley and Blaxter (2002). Suborders of the order Rhabditida, in which representatives highly homologous MSPs are found at DNA, RNA, or protein levels, are marked by underlining. Taxa whose species used in this study are marked with asterisks. Orders Trefusi- ida, Isolaimida, Dioctophyma- tida, Muspiceida, Marimermith- ida, and Desmoscolecida are not shown in this tree
Fig. 8 in Analysis of major sperm proteins in two nematode species from two classes, Enoplus brevis (Enoplea, Enoplida) and Panagrellus redivivus (Chromadorea, Rhabditida), reveals similar localization, but less homology of protein sequences than expected for Nematoda phylum
Fig. 8 Putative MSPs those are most similar to peptide antigen. a P. redivivus MSPs aligned with peptide antigen. Protein sequences (Pan_g61.t1, Pan_g6018.t1, Pan_g6424.t1, Pan_g9068.t1, Pan_ g19433.t1, and Pan_g21178.t1) were found by Blast using peptide
Fig. 4 in Analysis of major sperm proteins in two nematode species from two classes, Enoplus brevis (Enoplea, Enoplida) and Panagrellus redivivus (Chromadorea, Rhabditida), reveals similar localization, but less homology of protein sequences than expected for Nematoda phylum
Fig. 4 Immunolocalization of MSP in P. redivivus sperm. a Immature spermatozoa extracted from male. MSP localizes in granules. In some cells, MSP has strongest signals in the periphery (arrowheads) (scale bar 10 µm). b Chain of mature spermatozoa extracted from female.
Plasmid sequences for: Recombinant venom proteins in insect seminal fluid reduces female lifespan
Open the record for dataset details and reuse information.
Data from: Evaluating statistical multiple sequence alignment in comparison to other alignment methods on protein data sets
The estimation of multiple sequence alignments of protein sequences is a basic step in many bioinformatics pipelines, including protein structure prediction, protein family identification, and phylogeny estimation. Statistical co-estimation of alignments and trees under stochastic models of sequence evolution has long been considered the most rigorous technique for estimating alignments and trees, but little is known about the accuracy of such methods on biological benchmarks. We report the results of an extensive study evaluating the most popular protein alignment methods as well as the statistical co-estimation method BAli-Phy on 1192 protein data sets from established benchmarks as well as on 120 simulated data sets. Our study (which used more than 230 CPU years for the BAli-Phy analyses alone) shows that BAli-Phy has better precision and recall (with respect to the true alignments) than the other alignment methods on the simulated data sets, but has consistently lower recall on the biological benchmarks (with respect to the reference alignments) than many of the other methods. In other words, we find that BAli-Phy systematically under-aligns when operating on biological sequence data, but shows no sign of this on simulated data. There are several potential causes for this change in performance, including model misspecification, errors in the reference alignments, and conflicts between structural alignment and evolutionary alignments, and future research is needed to determine the most likely explanation. We conclude with a discussion of the potential ramifications for each of these possibilities.
In-house sequence database of S.pn recombinant protein vaccine target-related genes
<p>In-house sequence database of S.pn recombinant protein vaccine target-related genes.</p>
Sequences used for building HMMER profiles to search for SNAREs proteins
<p>This file contains the four sequence datasets that were used to build HMMER search profiles in order to find all SNAREs protein present in the predicted proteome of Oopsacas minuta: a human SNARE profile, a Qa-SNAREs profile, a Qb-SNAREs profile, a Qc-SNAREs profile.</p>
Case studies from: Sequence assignment validation in protein crystal structure models with checkMySequence
<p>Case studies from "Sequence assignment validation in protein crystal structure models with checkMySequence"</p>
Proteome database of 36 million proteins from 4,351 species, including marine microbial sequences
<p>A fasta-formatted database of 36,866,870 predicted proteins representing 4,351 unique species from 117 phyla.</p>
Protein sequences for 2031 Saccharomyces cerevisiae genome assemblies
<p>Protein sequences for 2031 Saccharomyces cerevisiae genome assemblies</p>
Optimal control derived sensitivity-enhanced CA-CO mixing sequences for MAS solid-state NMR. Applications in sequential protein backbone assignments.
<p>Raw data pulse sequences and shapes for publication</p> <p><br> ## SEQUENCES ##<br> ./sequences_renamed<br> Pulse programs introduiced in this work. Previous pulse programs can be obtained from https://doi.org/10.5281/zenodo.7016441 or https://optimal-nmr.net/experiments.html</p> <p>## SHAPES ##<br> ./shapes_renamed<br> TROP shaped pulses for homonuclear 13C-13C homonuclear mixing discussed in this work. Heteronuclear shaped pulses can be obtained from https://doi.org/10.5281/zenodo.7016441 or https://optimal-nmr.net/sequences.html</p> <p>## RAW DATA ##<br> to decrease storage demands 3D-processed spectra were deleted and can be recovered using TopSpin command: ftnd 0<br> TopSpin NUS licence is required for processing of the NUS-sampled data. Transformed data can be obtained from authors on request.</p> <p># U-13C,15N,2H,1HN-SH3 sample at 55 kHz MAS<br> ./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/1<br> 1H saturation recovery</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/2<br> 1H hard pulse calibraton</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/3<br> 15N hard pulse calibration</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/4<br> conventional hNH with rampCP optimalization</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/5<br> sensitivity-enhahced se-hNH with TROP optimalization</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/11<br> 2D conventional hNH with rampCP </p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/12<br> 2D sensitivtiy-enhanced se-hNH with TROP </p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/13<br> CO hard pulse and hCO CP calibration</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/14<br> CA hard pulse and hCA CP calibration</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/27<br> CACO homoTROP power optimalization in sensitivity-enhanced se-hCACOHN experiment</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/28<br> CACO INEPT delay optimalization in conventional hcoCAcoHN experiment</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/31 to 38<br> comparison of hCANH experiment efficiency using combination of coherence transfer methods (rampCP, tmSPICE and TROP) for CAN and NN transfer (see experiment tiles)</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/41 to 45<br> comparison of hCONH experiment efficiency using combination of coherence transfer methods (rampCP, tmSPICE and TROP) for CON and NN transfer (see experiment tiles)</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/57 and 58<br> optimalization of selective CA and CO 90 and 180 pulse</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/59<br> COCA INEPT delay optimalization in conventional hCOcaHN experiment</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/60<br> COCA homoTROP power optimalization in sensitivity-enhanced se-hCOCAHN experiment</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/69<br> conventional 3D hCANH experiment using tmSPICE CAN transfer; with water suppression after first CP, 15% non-uniform sampled</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/70<br> sensitivity-enhanced 3D se-hCANH experiment using TROP transfer; with water suppression after first CP, 15% non-uniform sampled</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/72<br> conventional 3D hCONH experiment using tmSPICE CON transfer; with water suppression after first CP, 15% non-uniform sampled</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/73<br> sensitivity-enhanced 3D se-hCONH experiment using TROP transfer; with water suppression after first CP, 15% non-uniform sampled</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/74<br> sensitivity-enhanced 3D se-hCAcoNH experiment using TROP transfer and homoTROP; with water suppression after first CP, 15% non-uniform sampled</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/75<br> conventional se-hcoCAcoNH experiment using tmSPICE CN transfer and INEPT ‘out-and-back’ coCAco transfer; with water suppression before NH transfer, 15% non-uniform sampled</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/76<br> sensitivity-enhanced 3D se-hCOcaNH experiment using TROP transfer and homoTROP; with water suppression after first CP, 15% non-uniform sampled</p> <p>./JB.1p3mm.20221122.sh3.2H13C15N100pcbackexch/77<br> conventional 3D hCOcaNH experiment using tmSPICE CN transfer and INEPT ‘complete forward’ COca; with water suppression after first CP, 15% non-uniform sampled</p> <p><br> # U-13C,15N glycine at 16.5 kHz MAS<br> ./MAS_COCa_seTesting/62<br> sensitivity-enhanced se-hCACO experiment with homoTROP pulse without diagonal-phase control</p> <p>./MAS_COCa_seTesting/63<br> sensitivity-enhanced se-hCOCA experiment with homoTROP pulse without diagonal-phase control</p> <p><br> # U-13C,15N-fMLF at 20 kHz MAS<br> ./jb.20211102_fMLF_3.2/2<br> 13C direct excitation spectra</p> <p>./jb.20211102_fMLF_3.2/3<br> 13C and 1H hard-pulse calibration and HC CP optimalization</p> <p>./jb.20211102_fMLF_3.2/4<br> 15N hard-pulse calibration and HC CP optimalization</p> <p>./jb.20211102_fMLF_3.2/5<br> NCA rampCP optimalization</p> <p>./jb.20211102_fMLF_3.2/6<br> NCO rampCP optimalization</p> <p>./jb.20211102_fMLF_3.2/79<br> sensitivity-enhanced se-hNCACO with TROP pulses</p> <p>./jb.20211102_fMLF_3.2/83<br> sensitivity-enhanced se-hNCOCA with TROP pulses</p> <p>./jb.20211102_fMLF_3.2/87<br> conventional hNCACO with rampCP and DREAM mixing</p> <p>./jb.20211102_fMLF_3.2/87<br> conventional hNCOCA with rampCP and DREAM mixing</p>
EPSAPG: A Pipeline Combining MMseqs2 and PSI-BLAST to Quickly Generate Extensive Protein Sequence Alignment Profiles
<p>This repository contains all data, queries, and search results used in the analysis of</p> <ul> <li>Arab, Issar. “<strong>EPSAPG: A Pipeline Combining MMseqs2 and PSI-BLAST to Quickly Generate</strong><br><strong>Extensive Protein Sequence Alignment Profiles</strong>.” IEEE/ACM 10th International Conference on<br>Big Data Computing, Applications and Technologies (BDCAT ’23), December 4–7, 2023,<br>Taormina, Messina, Italy. <a href="https://doi.org/10.1145/3632366.3632384">doi.org/10.1145/3632366.3632384</a></li> </ul>
Enhancing Protein Sequence Annotation in Viral Genomics Using Large Language Models and Soft Alignments.
<p>List of 200 most abundant VOG descriptions.</p>
Protein sequences matching TIGRFAM models
<p>A set of 411 large and 3,001 small TIGRFAM protein families originally used for benchmarking sequence clustering programs. Large families contain at least 20,000 sequences with a genus label, whereas small contain fewer than 20,000. Sequences were downloaded from <a href="https://www.ncbi.nlm.nih.gov/genome/annotation_prok/tigrfams/">NCBI</a> and renamed by their original name (accession number) followed by their semi-colon separated NCBI taxonomy of the originating organism. For example, the first sequence in TIGRFAM_large/TIGRFAM00005.fas.gz is named:</p><blockquote><p>WP_000005837.1 RluA family pseudouridine synthase, partial [Bacillus anthracis];TIGR00005(group);cellular organisms(no rank);Bacteria(superkingdom);Terrabacteria group(clade);Firmicutes(phylum);Bacilli(class);Bacillales(order);Bacillaceae(family);Bacillus(genus);Bacillus cereus group(species group);Bacillus anthracis(species)</p></blockquote>
Data from: A hypervariable mitochondrial protein coding sequence associated with geographical origin in a cosmopolitan bloom-forming alga, Heterosigma akashiwo
Open the record for dataset details and reuse information.
Proteome database of 36 million proteins from 4,351 species, including marine microbial sequences
Open the record for dataset details and reuse information.
Data from: Evaluating statistical multiple sequence alignment in comparison to other alignment methods on protein data sets
Open the record for dataset details and reuse information.
Data from: The evolution of heat shock protein sequences, cis-regulatory elements, and expression profiles in the eusocial Hymenoptera
Open the record for dataset details and reuse information.
FASTA file of sequences of identified proteins in Anastrepha ludens reproductive tissues
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.