Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
276
datasets available to search
ShareScore release 0.9.0
Dataset results
276 results for “protein domains”
Phase separation of SARS-CoV-2 nucleocapsid protein with TDP-43 is dependant on C-terminus domains
Open the record for dataset details and reuse information.
DPAM Domain Classification of Human Proteins against ECOD Reference
<p>Domain definitions of AlphaFold classifications of the human proteome (v1) from the AlphaFold Database. Also included are classifications of <em>Danio rerio</em>, <em>Mus musculus</em>, <em>Pan paniscus</em>, <em>Drosophila melanogaster</em>, <em>Caenorhabditis elegans </em>used for comparative analysis to human. See README file for descriptions of file formats.</p>
On the redundancy of the Josephin domain (JD) containing proteins and implications on the spinocerebellar ataxia type 3 (SCA3)
<p>Docking of the PPI reported in Figure 2, 3A, 3B, 4, 5, and 6 of the manuscript ‘On the redundancy of the Josephin domain (JD) containing proteins and implications on the spinocerebellar ataxia type 3 (SCA3).’</p>
Asymmetric evolution of protein domains in the leucine-rich repeat receptor-like kinase (LRR-RLK) family of plant developmental coordinators
<p><span>The coding sequences of developmental genes are expected to be conserved over deep time, with cis-regulatory change driving the modulation of gene function. In contrast, proteins with roles in defense are expected to evolve rapidly, in molecular arms races with pathogens. However, some gene families include both developmental and defense genes. In these families, do the tempo and mode of evolution differ between developmental and defense genes, despite shared ancestry and structure? The leucine-rich repeat receptor-like kinase (LRR-RLKs) protein family includes many members with roles in plant development and defense, thus providing an ideal system for answering this question. LRR-RLKs are receptors that traverse plasma membranes. LRR domains bind extracellular ligands, RLK domains initiate intracellular signaling cascades in response to ligand binding. In LRR-RLKs with roles in defense, LRR domains evolve faster than RLK domains. To determine whether this asymmetry extends to developmental LRR-RLKs, we assessed evolutionary rates and tested for selection acting on eleven clades of LRR-RLK proteins, using deeply sampled protein trees. To assess functional evolution, we performed heterologous complementation assays using <em>Arabidopsis thaliana</em> (arabidopsis) LRR-RLK mutants. We found that the LRR domains of developmental LRR-RLK proteins evolved faster than their cognate RLK domains. LRR-RLKs with roles in development and defense had strikingly similar patterns of molecular evolution. Heterologous transformation experiments revealed that the evolution of developmental LRR-RLKs likely involves multiple mechanisms, including changes to cis-regulation, coding sequence evolution, and escape from adaptive conflict. Our results indicate similar evolutionary pressures acting on developmental and defense signaling proteins, despite divergent organismal functions. In addition, deep understanding of the molecular evolution of developmental receptors can help guide targeted genome engineering in agriculture.</span></p>
Folding pathway of a discontinuous two-domain protein_1
<p>This data set contains all the raw data collected for the preparation of the manuscript "Folding pathway of a discontinuous two-domain protein".</p> <p>Raw data for the main figures Figure 1B-C, Figure 2A, C-D, Figure 4B-C, Figure 5B, C-D and for the supplementary figures Figure S1A-B, Figure S3C, E, G, Figure S5, Figure S7A-B, Figure S8A, D, Figure S9, and Figure S15A-C are deposited. Data is given figure-wise in folders and figure panels in sub-folders. The recurrent raw data used for multiple figures is mentioned for respective figures. Document explaining in detail about the figures and respective data set is also provided.</p> <p>Data is given as measured single-molecule TCSPC data as well as the background measurements as buffer and IRF as dpbs measurements in each respective folder.</p> <p>Due to size, the repository is uploaded in two parts. This part I has all the data for above figures except Figure S3, Figure S8, Figure S9 and Figure S15 for which associated data are deposited in part II repository with Zenodo DOI https://doi.org/10.5281/zenodo.8136592.</p>
Strategy of selection and optimization of single domain antibodies targeting the PHF6 linear peptide within the Tau intrinsically disordered protein
<p>Dataset pertaining to Strategy of selection and optimization of single domain antibodies targeting the PHF6 linear peptide within the Tau intrinsically disordered protein</p>
Fig. 4 in Phylogeny and domain architecture of plant ribosome inactivating proteins
Fig. 4. - Circular gene tree of RIPs. Our dataset comprised a curated selection from all the proteins available within NCBI's Conserved Domain Database containing a RIP domain. The phylogenetic tree was constructed using the maximum-likelihood method with the amino acid substitution model WAG + R9, ultrafast bootstrapping approximation (UFBoot) with 1000 iterations and the SH-like approximate likelihood ratio test with 1000 iterations. The legend titled 'Order (outer ring)' outlines the coloured dots around the perimeter of the tree and represents the phylogenetic order from which each protein sequence originated. Proteins without a coloured dot belong to orders with less than ten proteins. The legend titled 'RIP domains (inner lines)' defines colours in the branches of the tree, each representing a RIP group, based on presence/absence of a signal peptide and domains listed in both NCBI's Conserved Domain and Protein databases. (For interpretation of the references to colour in this figure legend, the reader is referred to the Web version of this article.)
Fig. 2 in Phylogeny and domain architecture of plant ribosome inactivating proteins
Fig. 2. - The number and type of RIPs within each plant order. Our dataset comprised a curated selection from all the proteins available within NCBI's Conserved Domain Database containing a RIP domain. The groups were sorted based on presence/absence of a signal peptide and domains listed in both NCBI's Conserved Domain and Protein databases. Colours represent different RIP groups. (For interpretation of the references to colour in this figure legend, the reader is referred to the Web version of this article.)
Fig. 6 in Phylogeny and domain architecture of plant ribosome inactivating proteins
Fig. 6. Tile plot of the most conserved amino acids within RIP domains of each protein group. The Y axis depicts the amino acid and corresponding position within the pokeweed antiviral protein (protein databank: 1QCI), the X axis depicts the proteins groups. For the two pink-highlighted amino acids on the Y axis, different amino acids were present in 70% of sequences at that position but were not present in the amino acid sequence of the crystal structure. The solid blue cells represent amino acids conserved at least 70% within each RIP group and with shared identity to the reference sequence 1QCI. Patterned blue cells represent amino acids conserved at least 70% within RIP groups but without identity to 1QCI. Any amino acid positions with 70% consensus in less than three protein groups were collapsed and shaded black.
Fig. 3 in Phylogeny and domain architecture of plant ribosome inactivating proteins
Fig. 3. - The number of RIPs within each plant species. Our dataset comprised a curated selection from all the proteins available within NCBI's Conserved Domain Database containing a RIP domain. The groups were sorted based on presence/absence of a signal peptide and domains listed in both NCBI's Conserved Domain and Protein databases. Colours represent different RIP groups. (For interpretation of the references to colour in this figure legend, the reader is referred to the Web version of this article.)
Fig. 1 in Phylogeny and domain architecture of plant ribosome inactivating proteins
Fig. 1. - Physicochemical characteristics of different RIP groups. Values were calculated from the amino acid coding sequence of each protein in our dataset with the R package 'peptides' and presented on boxplot format. Our dataset comprised a curated selection from all the proteins available within NCBI's Conserved Domain Database containing a RIP domain. The groups were sorted based on presence/absence of a signal peptide and domains listed in both NCBI's Conserved Domain and Protein databases. (A) molecular weight prediction; (B) theoretical net charge prediction; (C) Boman potential protein interaction index prediction; (D) aliphatic index prediction.
Fig. 5. - The most highly conserved RIP amino acids. Our dataset comprised a in Phylogeny and domain architecture of plant ribosome inactivating proteins
Fig. 5. - The most highly conserved RIP amino acids. Our dataset comprised a curated selection from all the proteins available within NCBI's Conserved Domain Database containing a RIP domain. Colours indicate amino acids conserved in at least 70% of sequences. For the two pink-highlighted proteins, the black bolded amino acids were present in 70% of sequences at that position but were not present in the amino acid sequence of the crystal structure. (A) Sequence alignment. RIP domain consensus: the consensus sequence generated from the multiple sequence alignment in Jalview excluding gaps; 1QCI: the amino acid sequence of pokeweed antiviral protein (protein databank: 1QCI). The third line denotes the similarity in the two sequences as determined by Clustal Omega. (B) The crystal structure of 1QCI visualized in UCSF ChimeraX in surface representation; (C) mesh representation; and (D) cartoon representation. (For interpretation of the references to colour in this figure legend, the reader is referred to the Web version of this article.)
Understanding structural and functional diversity of ATP-PPases using protein domains and functional families in CATH database
<p>The dataset of AF2-predicted HUP domains with overall pLDDT > 90, culled at 90% identity.</p>
KN046 (a Humanized PD-L1/CTLA4 Bispecific Single Domain Fc Fusion Protein Antibody) in Subjects With Thymic Carcinoma
ClinicalTrials.gov study NCT04469725. IPD Sharing: NO. Countries: 1. Publications: 1.
Study IL- 35 Level and Fibronectin Type Ⅲ Domain Containing Protein 5 \ Irsin (FNDC5 rs3480) Gene Variations in Chronic Hepatitis B.
ClinicalTrials.gov study NCT06023745. IPD Sharing: UNDECIDED. Countries: 1. Publications: 3.
Asymmetric evolution of protein domains in the leucine-rich repeat receptor-like kinase (LRR-RLK) family of plant developmental coordinators
Open the record for dataset details and reuse information.
Dataset for: "BACPHLIP: Predicting bacteriophage lifestyle from conserved protein domains"
<p>This is the dataset used in the manuscript titled "BACPHLIP: Predicting bacteriophage lifestyle from conserved protein domains". The dataset is necessary to run the code that can be found at <a href="https://github.com/adamhockenberry/dca-weighting">https://github.com/adamhockenberry/bacphlip-model-dev</a> and outside of the context of the code this data will likely not be super well-annotated or helpful. The code within that repository, however, should provide sufficient information about the structure and usage of this dataset. </p>
Data from: The basic keratin 10-binding domain of the virulence-associated pneumococcal serine-rich protein PsrP adopts a novel MSCRAMM fold
Streptococcus pneumoniae is a major human pathogen, and a leading cause of disease and death worldwide. Pneumococcal invasive disease is triggered by initial asymptomatic colonization of the human upper respiratory tract. The pneumococcal serine-rich repeat protein (PsrP) is a lung-specific virulence factor whose functional binding region (BR) binds to keratin-10 (KRT10) and promotes pneumococcal biofilm formation through self-oligomerization. We present the crystal structure of the KRT10-binding domain of PsrP (BR187–385) determined to 2.0 Å resolution. BR187–385 adopts a novel variant of the DEv-IgG fold, typical for microbial surface components recognizing adhesive matrix molecules adhesins, despite very low sequence identity. An extended β-sheet on one side of the compressed, two-sided barrel presents a basic groove that possibly binds to the acidic helical rod domain of KRT10. Our study also demonstrates the importance of the other side of the barrel, formed by extensive well-ordered loops and stabilized by short β-strands, for interaction with KRT10.
Evaluating the use of paralogous protein domains to increase data availability for missense variant classification [dataset]
Open the record for dataset details and reuse information.
Domain scanning results for a selected set of high-quality-annotation protein isoforms produced by human transcription factor genes
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.