Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
302
datasets available to search
ShareScore release 0.9.0
Dataset results
302 results for “protein sequence”
Gene and protein sequence features augment HLA class I ligand predictions
<p>Dataset and analyses supporting the manuscript "Gene and protein sequence features augment HLA class I ligand predictions".</p> <p>The "peptides" files contain the mass-spec detected peptides obtained from HLA ligandomics performed on the indicated tumor lines. </p> <p>The "protein data" files contain the RNAseq data (TPM) and Ribosome profiling data (ribosome occupancy) per protein, for each tumor line. </p> <p>The "source data" zip archive contains the source data underlying the figures of the manuscript.</p> <p>The "HLA ligandome analyses" zip archive contains the R scripts used for all data analysis in the manuscript, including all data and output files. These analyses can also be found at https://github.com/kasbress/HLA_Ligandome_Analyses/</p> <p> </p> <p> </p>
Mobility-assisted psuedo-MS3 sequencing of protein ions
<p>The sequencing of intact proteins within a mass spectrometer has many benefits but is frequently limited by the fact that tandem mass spectrometry (MS/MS) techniques often generate poor sequence coverages when applied to protein ions. To overcome this limitation exotic MS/MS techniques that rely on lasers and radical chemistry have been developed. These techniques generate high sequence coverages, but they require specialized instrumentation, create products through multiple dissociation mechanisms, and often require long acquisition times. Recently, we demonstrated that protein ions can be dissociated in a trapped ion mobility spectrometry (TIMS) device prior to mobility separation in a commercial timsTOF. All generated product ions were distributed throughout the mobility dimension and this separation enabled deconvolution of complex tandem mass spectra and could enable facile pseudo-MS<sup>3</sup> interrogation of generated product ions with the downstream quadrupole and collision cell. A second activation step improves sequence coverage because the most labile bonds have been depleted during the first dissociation and subsequent dissociation events are more evenly distributed throughout the product ion backbone. In this work, we explore the potential of this mobility-assisted pseudo-MS<sup>3</sup> (MAP) method on a commercial timsTOF and timsTOF Pro 2. We demonstrate that while MAP only generates 92% of the sequence coverage of the most effective MS/MS technique, it accomplished this feat in 1.5 mins and could be facilely integrated with liquid chromatographic separations.</p>
Multiple sequence alignment of USP Zf-UBD proteins
<p>Using Molsoft's ICM-Pro, a multiple sequence alignment of USP Zf-UBDs was done against HDAC6 Zf-UBD. </p>
Bridging the Chromosome-Centric and Biology and Disease Human Proteome Projects: Accessible and automated tools for interpreting biological and pathological impact of protein sequence variants detected via proteogenomics
<p>Bridging the Chromosome-Centric and Biology and Disease Human Proteome Projects: Accessible and automated tools for interpreting biological and pathological impact of protein sequence variants detected via proteogenomics</p>
Amino acid sequences of the proteins predicted from the whole genome of hilsa shad (Tenualosa ilisha) of the Bay of Bengal
<p>Gene prediction was performed by AUGUSTUS (Stanke et al., 2006) from the whole genome sequence of <em>T. ilisha</em> (<a href="https://www.ncbi.nlm.nih.gov/bioproject/400122">PRJNA400122</a>). The data contain amino acid sequences of 37,450 predicted protein coding genes.</p>
Data file with manuscript titled 'A Structurally Validated Sequence Alignment of 497 Human Protein Kinase Domains'
<p>The files used in different analysis reported in the manuscript titled - 'A Structurally-Validated Multiple Sequence Alignment of 497 Human Protein Kinase Domains' are shared at two locations. Following is a brief description of these files.</p> <p>Location - https://github.com/DunbrackLab/Kinases<br> 1. HMM profile files - HMM files for each of the nine groups computed separately labeled as Groupname.hmm, like AGC.hmm<br> 2. HMM profile file - HMM file computed from the full alignment including all the sequences - Human-PK.hmm<br> 3. Score files - HMM scores of each kinase sequence against all the groupwise HMMs both for iteration1 (HMM-iter1-scores-tables.txt) and iteration2 (HMM-iter1-scores-tables.txt)<br> 4. Jalview session file - Kinase alignment with sequences colored by secondary structure information from PDB file if the structure is known; or predicted secondary structure if the experimental structure is not known. The file could be opened in Jalview - kinases-PDB-SSPred.jvp</p> <p>Location - https://zenodo.org/record/3445533<br> 1. The file contains list of residue pairs aligned in pairwise structural alignments of 272 human protein kinases which were used as a benchmark in the study. The alignments were created by FATCAT and optimized by SE program.</p>
Treponema pallidum Cytoplasmic filament protein gene partial sequence
<p>A partial sequence of the Cytoplasmic filament protein gene (cfpA) of treponema pallidum from non human primates</p>
Phylogenetic tree of 1819 YfaL protein sequences identified in 2053 E. coli genomes.
<p><strong><span>Phylogenetic tree of 1819 YfaL protein sequences identified in 2053 <em>E. coli</em> genomes. </span></strong><span>The purple circles on the branches represent bootstrap values > 0.8. The different strain’s phylogroups are displayed outside of the tree.</span></p>
Data accompanying "In silico analysis of the profilaggrin sequence indicates alterations in the stability, degradation route, and intracellular protein fate in filaggrin null mutation carriers" article.
<p>This research was supported by the National Science Centre, Poland, grant PRELUDIUM number 2021/41/N/NZ1/03473 to NS, National Science Centre, Poland, grant SONATA BIS number 2019/34/E/NZ6/00354 to DG-O, as well as POIR.04.04.00-00-21FA/16–00 grant, carried out within the First TEAM programme of the Foundation for Polish Science co-financed by the European Union under the European Regional Development Fund (awarded to DG-O). WP was supported by the National Science Centre, Poland, grant SONATA-BIS number 2021/42/E/NZ1/00190. SB is supported by a Wellcome Trust Senior Research Fellowship (220875/Z/20/Z).</p>
Supplementary Information for Phylogenetic analyses of ray-finned fishes (Actinopterygii) using collagen type I protein sequences
<p>Ray-finned fishes (Actinopterygii) are the largest and most diverse group of vertebrates, comprising over half of all living vertebrate species. Phylogenetic relationships between ray-finned fishes have historically pivoted on the study of morphology, which has notoriously failed to resolve higher-order relationships, such as within the percomorphs. More recently, comprehensive genomic analyses have provided further resolution of actinopterygian phylogeny, including higher-order relationships. Such analyses are rightfully regarded as the 'gold standard' for phylogenetics. However, DNA retrieval requires modern or well-preserved tissue and is less likely to be preserved in archaeological or fossil specimens. In contrast some proteins, such as collagen, are phylogenetically informative and can survive into deep time. Here, we test the utility of collagen type I amino acid sequences for phylogenetic estimation of ray-finned fishes. We estimate topology using Bayesian approaches and compare the congruence of our estimated trees with published genomic phylogenies. Furthermore, we apply a Bayesian molecular clock approach and compare estimated divergence dates with previously published genomic clock analyses. Our collagen-derived trees exhibit 77% of node positions as congruent with recent genomic-derived trees, with the majority of discrepancies occurring in higher-order node positions, almost exclusively within the Percomorpha. Our molecular clock trees present divergence times that are fairly comparable with genomic-based phylogenetic analyses. We estimate the mean node age of Actinopteri at ~293 million years (Ma), the base of Teleostei at ~211 Ma and the radiation of percomorphs beginning at ~141 Ma (~350 Ma, ~250–283 Ma and ~120–133 Ma in genomic trees, respectively). Finally, we show that the average rate of collagen (I) sequence evolution is 0.9 amino acid substitutions for every million years of divergence, with the α3 (I) sequence evolving the fastest, followed by the α2 (I) chain. This is the quickest rate known for any vertebrate group. We demonstrate that phylogenetic analyses using collagen type I amino acid sequences generate tangible signals for actinopterygians that are highly congruent with recent genomic-level studies. However, there is limited congruence within percomorphs, perhaps due to clade-specific functional constraints acting upon collagen sequences. Our results provide important insights for future phylogenetic analyses incorporating extinct actinopterygian species via collagen (I) sequencing.</p>
Input Data for "Protein Function Prediction for newly sequenced organisms"
<p>The input sequence files in FASTA format and the detailed list of all organisms excluded when testing each specific bacterium.</p>
TIGRFAM protein sequences named by taxonomy
<p>A set of 411 TIGRFAM protein families originally used for benchmarking sequence clustering programs. Sequences were downloaded from <a href="https://www.ncbi.nlm.nih.gov/genome/annotation_prok/tigrfams/">NCBI</a> and renamed by their original name (accession number) followed by their semi-colon separated NCBI taxonomy. For example, the first sequence in TIGRFAM00005.fas.gz is named:</p> <blockquote> <p>WP_000005837.1 RluA family pseudouridine synthase, partial [Bacillus anthracis];TIGR00005(group);cellular organisms(no rank);Bacteria(superkingdom);Terrabacteria group(clade);Firmicutes(phylum);Bacilli(class);Bacillales(order);Bacillaceae(family);Bacillus(genus);Bacillus cereus group(species group);Bacillus anthracis(species)</p> </blockquote>
Fasta format protein sequences from assembled kyphosid fish gut metagenomes
<p>Predicted proteins sequences from kyposid fish gut metagenomic samples F5, F6, F7, and F8, obtained as described in the following study:</p> <p>Podell S, Oliver A, Kelly LW, Sparagon W, Plominsky, A, Nelson RS, Laurens LML, Augyte, S, Sims NA, Nelson CE, Allen EE. Herbivorous fish microbiome adaptations to sulfated dietary polysaccharides (2023)<br> manuscript submitted.</p>
Artificial intelligence method to design and fold alpha-helical structural proteins from the primary amino acid sequence
<p>Dataset for paper: Z. Qin, L. Wu, H. Sun, S. Huo, T. Ma, E. Lim, P.-Y. Chen, B. Marelli, M.J. Buehler, Artificial intelligence method to design and fold alpha-helical structural proteins from the primary amino acid sequence, Extreme Mechanics Letters, Vol. 36, 100652, 2020. <a href="https://doi.org/10.1016/j.eml.2020.100652">https://doi.org/10.1016/j.eml.2020.100652</a>.</p> <p>Code: https://github.com/lamm-mit/MNNN/ </p>
Dataset - Functional annotation for protein sequences
<p>These are the input and output files for the functional annotation of protein sequences workflow tests.</p>
extHomFam v37.0: structural benchmark for protein multiple sequence alignments
<p>extHomFam v37.0 was constructed by combining Homstrad reference alignments (2 December 2023 release) with Pfam 37.0 (UniProt release) families containing at least 200 sequences. Homstrad entries with less than 3 reference sequences and those pointing to dead Pfam families were discarded.</p> <p> </p>
Protein sequences of LbNoxA, LbNoxB and LbNoxR in L. bicolor and other different species
Open the record for dataset details and reuse information.
Long-read-based draft genome sequence of Indian black gram IPU-94-1 ‘Uttara’: Insights into disease resistance and seed storage protein genes
Open the record for dataset details and reuse information.
Mobility-assisted psuedo-MS3 sequencing of protein ions
Open the record for dataset details and reuse information.
Amino acid sequences of RWP-RK domain containing proteins used for the construction of phylogenetic tree shown in Fig. 1
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.