Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

302

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

302 results for “protein sequence”

Learn how ShareScore rates datasets ↗
dryad32/100

ProtASR2: Ancestral Reconstruction of Protein Sequences accounting for Folding Stability

Open the record for dataset details and reuse information.

publicDec 2019View details →
dryad32/100

Phylogenetic analysis of 218 protein sequences that bear sequence similarity to F. prausnitzii FAAH

Open the record for dataset details and reuse information.

publicOct 2024View details →
zenodo28/100

Evaluation SIHUMI dataset: A sectioning and database enrichment approach for improved peptide spectrum matching in large, genome-guided protein sequence databases

<p>Dataset for&nbsp;evaluation of&nbsp;database sectioning method for generating an enriched database for mass-spectrometry-based proteomic approaches&nbsp;using large databases. Our evaluation demonstrates that this method helps to&nbsp;increase the sensitivity of PSMs while maintaining acceptable FDR statistics. This dataset was acquired from the protein samples containing proteins from eight microorganisms (<em>Anaerostipes caccae, Bacteroides thetaiotaomicron, Bifidobacterium longum, Blautia producta, Clostridium butyricum, Clostridium ramosum, Escherichia coli, Lactobacillus plantarum</em>) that were grown in a bioreactor. The MS-data for this dataset was acquired by Dr. Hettich&#39;s group at Oak Ridge National Laboratory, and it was made available by Dr. Nico Jehmlich from&nbsp;Helmholtz Center for Environmental Research through&nbsp;the 3rd International Metaproteome Symposium (<a href="https://www.ufz.de/index.php?en=44639">https://www.ufz.de/index.php?en=44639</a>).&nbsp;</p>

opencc-by-4.0Apr 2020View details →
dryad28/100

Data from: A branch-heterogeneous model of protein evolution for efficient inference of ancestral sequences

Most models of nucleotide or amino acid substitution used in phylogenetic studies assume that the evolutionary process has been homogeneous across lineages and that composition of nucleotides or amino acids has remained the same throughout the tree. These oversimplified assumptions are refuted by the observation that compositional variability characterizes extant biological sequences. Branch-heterogeneous models of protein evolution that account for compositional variability have been developed, but are not yet in common use because of the large number of parameters required, leading to high computational costs and potential overparameterization. Here, we present a new branch-nonhomogeneous and nonstationary model of protein evolution that captures more accurately the high complexity of sequence evolution. This model, henceforth called Correspondence and likelihood analysis (COaLA), makes use of a correspondence analysis to reduce the number of parameters to be optimized through maximum likelihood, focusing on most of the compositional variation observed in the data. The model was thoroughly tested on both simulated and biological data sets to show its high performance in terms of data fitting and CPU time. COaLA efficiently estimates ancestral amino acid frequencies and sequences, making it relevant for studies aiming at reconstructing and resurrecting ancestral amino acid sequences. Finally, we applied COaLA on a concatenate of universal amino acid sequences to confirm previous results obtained with a nonhomogeneous Bayesian model regarding the early pattern of adaptation to optimal growth temperature, supporting the mesophilic nature of the Last Universal Common Ancestor.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Sequence specificity despite intrinsic disorder: how a disease-associated Val/Met polymorphism rearranges tertiary interactions in a long disordered protein

The role of electrostatic interactions and mutations that change charge states in intrinsically disordered proteins (IDPs) is well-established, but many disease-associated mutations in IDPs are charge-neutral. The Val66Met single nucleotide polymorphism (SNP) in precursor brain-derived neurotrophic factor (BDNF) is one of the earliest SNPs to be associated with neuropsychiatric disorders, and the underlying molecular mechanism is unknown. Here we report on over 250 μ s of fully-atomistic, explicit solvent, temperature replica exchange molecular dynamics (MD) simulations of the 91 residue BDNF prodomain, for both the V66 and M66 sequence. The simulations were able to correctly reproduce the location of both local and non-local secondary changes due to the Val66Met mutation when compared with NMR spectroscopy. We find that the change in local structure is mediated via entropic and sequence specific effects. We developed a hierarchical sequence-based framework for analysis and conceptualization, which first identifies "blobs" of 5-15 residues representing local globular regions or linkers. We use this framework within a novel test for enrichment of higher-order (tertiary) structure in disordered proteins; the size and shape of each blob is extracted from MD simulation of the real protein (RP), and used to parameterize a self-avoiding heterogenous polymer (SAHP). The SAHP version of the BDNF prodomain suggested a protein segmented into three regions, with a central long, highly disordered polyampholyte linker separating two globular regions. This effective segmentation was also observed in full simulations of the RP, but the Val66Met substitution significantly increased interactions across the linker, as well as the number of participating residues. The Val66Met substitution replaces β -bridging between Val66 and Val94 (on either side of the linker) with specific side-chain interactions between Met66 and Met95.The protein backbone in the vicinity of Met95 is then free to form β -bridges with residues 31-41 near the N-terminus, which condenses the protein. A significant role for Met/Met interactions is consistent with previously-observed non-local effects of the Val66Met SNP, as well as established interactions between the Met66 sequence and a Met-rich receptor that initiates neuronal growth cone retraction.

opencc-zeroOct 2019View details →
dryad28/100

Data from: Successful recovery of nuclear protein-coding genes from small insects in museums using illumina sequencing

In this paper we explore high-throughput Illumina sequencing of nuclear protein-coding, ribosomal, and mitochondrial genes in small, dried insects stored in natural history collections. We sequenced one tenebrionid beetle and 12 carabid beetles ranging in size from 3.7 to 9.7 mm in length that have been stored in various museums for 4 to 84 years. Although we chose a number of old, small specimens for which we expected low sequence recovery, we successfully recovered at least some low-copy nuclear protein-coding genes from all specimens. For example, in one 56-year-old beetle, 4.4 mm in length, our de novo assembly recovered about 63% of approximately 41,900 nucleotides in a target suite of 67 nuclear protein-coding gene fragments, and 70% using a reference-based assembly. Even in the least successfully sequenced carabid specimen, reference-based assembly yielded fragments that were at least 50% of the target length for 34 of 67 nuclear protein-coding gene fragments. Exploration of alternative references for reference-based assembly revealed few signs of bias created by the reference. For all specimens we recovered almost complete copies of ribosomal and mitochondrial genes. We verified the general accuracy of the sequences through comparisons with sequences obtained from PCR and Sanger sequencing, including of conspecific, fresh specimens, and through phylogenetic analysis that tested the placement of sequences in predicted regions. A few possible inaccuracies in the sequences were detected, but these rarely affected the phylogenetic placement of the samples. Although our sample sizes are low, an exploratory regression study suggests that the dominant factor in predicting success at recovering nuclear protein-coding genes is a high number of Illumina reads, with success at PCR of COI and killing by immersion in ethanol being secondary factors; in analyses of only high-read samples, the primary significant explanatory variable was body length, with small beetles being more successfully sequenced.

opencc-zeroDec 2015View details →
zenodo28/100

Learning the sequence code for protein abundance in human immune cells

<p><span>Protein abundance is defined by </span><span>transcriptional, post-transcriptional and post-translational regulatory mechanisms. Understanding the code for gene expression could inform novel therapies. Here, we developed a machine learning pipeline, termed SONAR, to decipher the endogenous sequence code that defines the abundance of protein in human cells. SONAR predicts up to 63% of protein abundance independently of promoter or enhancer information. Our analysis reveals a strong - yet dynamic - </span><span>cell-type specific sequence code<span>. The deep knowledge of SONAR provides a map of biologically active sequence features (SFs), which we leveraged to </span></span><span>manipulate protein expression and tailored to a specific cell-type. </span><span>Beyond providing fundamental insights in gene expression regulation, our study offers novel means to improve therapeutic and biotechnology applications.</span></p> <p><span>Datasets of the protein models (with gamma=1) are included in this repository. Please refer to GitHub for more details on how the models were generated.&nbsp;</span></p>

openmit-licenseApr 2024View details →
zenodo28/100

Appendix tables for Gaia: A Context-Aware Sequence Search and Discovery Tool for Microbial Proteins

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo28/100

A joint embedding of protein sequence and structure enables robust variant effect predictions

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo28/100

Sequence data and structural data utilized in the study and analysis of grain protein function prediction.

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
dryad28/100

Data from: Protein structure determination using metagenome sequence data

Despite decades of work by structural biologists, there are still ~5200 protein families with unknown structure outside the range of comparative modeling. We show that Rosetta structure prediction guided by residue-residue contacts inferred from evolutionary information can accurately model proteins that belong to large families and that metagenome sequence data more than triple the number of protein families with sufficient sequences for accurate modeling. We then integrate metagenome data, contact-based structure matching, and Rosetta structure calculations to generate models for 614 protein families with currently unknown structures; 206 are membrane proteins and 137 have folds not represented in the Protein Data Bank. This approach provides the representative models for large protein families originally envisioned as the goal of the Protein Structure Initiative at a fraction of the cost.

opencc-zeroDec 2016View details →
dryad28/100

FASTA sequences of 6 different proteins of SARS-CoV-2

<p>The acute respiratory disease induced by the severe acute respiratory syndrome-coronavirus-2 (SARS-CoV-2) has become a global epidemic in just less than a year by the first half of 2020. The subsequent efficient human-to-human transmission of this virus eventually affected millions of people worldwide. The virulence of the SARS-CoV-2 is mostly regulated by its proteins but very little is known about the protein structures and functionalities. Therefore, the main purpose of this study is to learn more about these proteins through bioinformatics approaches. In this study, ORF10, ORF7b, ORF7a, ORF6, membrane glycoprotein, and envelope protein have been selected from a Bangladeshi Corona-virus strain G039392 and a number of bioinformatics tools and strategies were implemented for multiple sequence alignment and phylogeny analysis with 9 different variants, predicting hydropathicity, amino acid compositions, protein-binding propensity, protein disorders, 2D and 3D protein modeling.</p>

opencc-zeroJul 2021View details →
zenodo28/100

294-fold mini all-α protein library encoded by 7,350 amino-acid sequences (project "flood of fold" )

<p>This repository includes&nbsp;3 compressed archive files for the &ldquo;flood-of-fold&rdquo; mini-protein library project.&nbsp;</p> <ol> <li>FloodOfFolds_294_backbone_models.tar.gz includes 294 mini all-&alpha; backbone models showing distinct topologies (folds). They are poly-VAL models.&nbsp;</li> <li>FloodOfFolds_7350_designs_MODEL.tar.gz includes 7,350 design protein models in the pdb format&nbsp;for the 294 mini-protein library. 25 amino-acid sequences were designed for each backbone model (25 x 294 = 7,350).&nbsp;</li> <li>FloodOfFolds_7350_designs_FASTA.tar.gz includes 7,350 fasta files derived from the pdb files in FloodOfFolds_7350_designs_MODEL.tar.gz.<br> &nbsp;</li> </ol> <p>See also here for results of folding simulations:&nbsp;https://zenodo.org/record/5526849#.YWRhFBBBw1I</p> <p>Acknowledgement: K.S. and S.M.&nbsp;would like to deeply thank Koga laboratory at Institute for Molecular Science providing computational resources.&nbsp;Most of the computations for model building and folding simulations were performed using the facilities at the Research Center for Computational Science, Okazaki, Japan.</p>

opencc-by-4.0Sep 2021View details →
zenodo28/100

Fasta format protein sequences from 391 ruminant gut MAGs

<p>Predicted proteins sequences MAGs previously published as nucleotide sequences from the following study:</p> <p>Glendinning L, Genc B, Wallace RJ, Watson M. Metagenomic analysis of the cow, sheep, reindeer and red deer rumen. Sci Rep. 2021;11(1):1990; doi: 10.1038/s41598-021-81668-9.</p> <p>Methods for protein prediction are described in the following study:</p> <p>Podell S, Oliver A, Kelly LW, Sparagon W, Plominsky, A, Nelson RS, Laurens LML, Augyte,&nbsp;S,&nbsp;Sims NA, Nelson CE, Allen EE. Herbivorous fish microbiome adaptations to sulfated dietary polysaccharides (2023)<br> manuscript submitted.</p>

opencc-by-4.0Feb 2023View details →
dryad28/100

Supplementary materials sequences of soluble chemical communication proteins in Anagrus nilaparvatae

<p class="MsoNormal"><a name="_Hlk77276479"></a><em><span>Anagrus nilaparvatae</span></em><span> is an important egg parasitoid wasp of pests such as the rice planthopper. Based on the powerful olfactory system of sensing chemical information in nature, <em>A. nilaparvatae</em> shows </span><span>complicated</span><span> life activities and behaviors, such as feeding, mating and hosting. <a name="_Hlk77277184"></a>We constructed a full-length transcriptome library and used this to identify the characteristics of soluble chemical communication proteins. We analyzed the sequence characteristics of soluble chemical communication protein genes and identified eight genes: <a name="_Hlk86678428"></a><em>AnilOBP2</em>, <em>AnilOBP9</em>, <em>AnilOBP23</em>, <em>AnilOBP56</em>, <em>AnilOBP83</em>, <em>AnilCSP5</em>, <em>AnilCSP6</em> and <em>AnilNPC2</em>. After sequence alignment and conserved domain prediction, the eight proteins encoded by the eight genes above were found to be consistent with the typical characteristics of odorant-binding proteins (OBPs), chemosensory proteins (CSPs) and Niemann-pick type C2 proteins (NPC2s) in other insects. Phylogenetic tree analysis showed that the eight genes share low homology with other species of Hymenoptera.</span><span> </span><span>This study provides the first data for <em>A. nilaparvatae</em>'s molecular characteristics of soluble chemical communication proteins, as well as an opportunity for understanding how <em>A. nilaparvatae</em> behaviors are mediated via soluble chemical communication proteins.</span></p>

opencc-zeroFeb 2023View details →
dryad28/100

Data from: Protein structure determination using metagenome sequence data

Open the record for dataset details and reuse information.

publicJan 2018View details →
dryad28/100

Data from: Sequence co-evolution gives 3D contacts and structures of protein complexes

Open the record for dataset details and reuse information.

publicOct 2014View details →
dryad28/100

Data from: A branch-heterogeneous model of protein evolution for efficient inference of ancestral sequences

Open the record for dataset details and reuse information.

publicMar 2013View details →
dryad28/100

FASTA sequences of 6 different proteins of SARS-CoV-2

Open the record for dataset details and reuse information.

publicJul 2021View details →
dryad28/100

Data from: Generalized bootstrap supports for phylogenetic analyses of protein sequences incorporating alignment uncertainty

Open the record for dataset details and reuse information.

publicDec 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record