Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
170
datasets available to search
ShareScore release 0.9.0
Dataset results
170 results for “Protein Family”
Asymmetric evolution of protein domains in the leucine-rich repeat receptor-like kinase (LRR-RLK) family of plant developmental coordinators
<p><span>The coding sequences of developmental genes are expected to be conserved over deep time, with cis-regulatory change driving the modulation of gene function. In contrast, proteins with roles in defense are expected to evolve rapidly, in molecular arms races with pathogens. However, some gene families include both developmental and defense genes. In these families, do the tempo and mode of evolution differ between developmental and defense genes, despite shared ancestry and structure? The leucine-rich repeat receptor-like kinase (LRR-RLKs) protein family includes many members with roles in plant development and defense, thus providing an ideal system for answering this question. LRR-RLKs are receptors that traverse plasma membranes. LRR domains bind extracellular ligands, RLK domains initiate intracellular signaling cascades in response to ligand binding. In LRR-RLKs with roles in defense, LRR domains evolve faster than RLK domains. To determine whether this asymmetry extends to developmental LRR-RLKs, we assessed evolutionary rates and tested for selection acting on eleven clades of LRR-RLK proteins, using deeply sampled protein trees. To assess functional evolution, we performed heterologous complementation assays using <em>Arabidopsis thaliana</em> (arabidopsis) LRR-RLK mutants. We found that the LRR domains of developmental LRR-RLK proteins evolved faster than their cognate RLK domains. LRR-RLKs with roles in development and defense had strikingly similar patterns of molecular evolution. Heterologous transformation experiments revealed that the evolution of developmental LRR-RLKs likely involves multiple mechanisms, including changes to cis-regulation, coding sequence evolution, and escape from adaptive conflict. Our results indicate similar evolutionary pressures acting on developmental and defense signaling proteins, despite divergent organismal functions. In addition, deep understanding of the molecular evolution of developmental receptors can help guide targeted genome engineering in agriculture.</span></p>
Understanding structural and functional diversity of ATP-PPases using protein domains and functional families in CATH database
<p>The dataset of AF2-predicted HUP domains with overall pLDDT > 90, culled at 90% identity.</p>
Effects of Soy Protein on Cholesterol Levels in Children Affected With Familial Hypercholesterolemia
ClinicalTrials.gov study NCT03563547. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Experimental evolution of ancestrally reconstructed BCL2 family proteins
Open the record for dataset details and reuse information.
Asymmetric evolution of protein domains in the leucine-rich repeat receptor-like kinase (LRR-RLK) family of plant developmental coordinators
Open the record for dataset details and reuse information.
Supplementary tables S5, S7, S9, S10, original protein models fasta files used for alignments, aligned and manually curated protein modes files used for phylogenies (PHYLIP format), and phylogenetic trees of plant cell wall decomposition gene families from 44 basidiomycete genomes (.tre files)
Open the record for dataset details and reuse information.
Annotation and analysis of the secondary structure elements in the Cytochrome P450 protein family
<p>We collected all currently available structures for proteins in the Cytochrome P450 family and annotated their secondary structure elements using SecStrAnnotator software (https://webchem.ncbr.muni.cz/Wiki/SecStrAnnotator). We used 2nnjA as the template domain for the annotation. Based on these annotations, we analysed the occurrence, length distribution, amino acid sequence, and presence of structural irregularities (β-bulges, 3<sub>10</sub>-helices, π-helices) of each secondary structure element class. We also statistically compared the bacterial vs eukaryotic structures. For the secondary structure element classes with sufficient sequence conservation, the most conserved residue is annotated as the reference residue.</p> <p>Main files:</p> <ul> <li><strong>set_ALL.json</strong> - Set-ALL: list of 1012 protein domains belonging to the Cytochrome P450 family (CATH accession 1.10.630.10 + Pfam accession PF00067, accessed on 7 July 2020, one domain per PDB entry)</li> <li><strong>set_NR.json</strong> - Set-NR: non-redundant list of 183 domains (one domain per UniProt ID)</li> <li><strong>domain_lists_table.tsv</strong> - Overview of Set-ALL and Set-NR and separation into subsets Set-NR-Bact (bacterial), Set-NR-Euka (eukaryotic), Set-NR-Arch (archaeal), Set-NR-Viru (viral)</li> <li><strong>structures/template_2NNJ-template.sses.json</strong> - Manually prepared annotation template (domain 2nnjA)</li> <li><strong>structures/template_2NNJ.cif</strong> - Structure of the template domain (2nnjA)</li> <li><strong>annotations_with_reference_residues_ALL.json, annotations_with_reference_residues_ALL.tsv</strong> - Annotation of secondary structure elements for Set-ALL</li> <li><strong>annotations_with_reference_residues_NR.json, annotations_with_reference_residues_NR.tsv</strong> - Annotation of secondary structure elements for Set-NR</li> <li><strong>aligments_NR</strong> - Multiple sequence alignments for each SSE class (Set-NR)</li> <li><strong>logos_NR</strong> - Sequence logos for each SSE class (Set-NR)</li> <li><strong>plots</strong> - Plots of SSE occurrence, length distribution, contained helix types and beta-bulge occurrence (Set-NR), some plots show the comparison between Set-NR-Bact and Set-NR-Euka</li> <li><strong>statistical_tests.ods</strong> - Comparison of SSE occurrence between Set-NR-Bact and Set-NR-Euka by the test of equal proportions and the Fisher test, comparision of the SSE length by the Kolmogorov-Smirnov test and the two-sample Wilcoxon test</li> </ul>
Data from: Evolution of the leucine-rich repeat receptor-like protein kinase gene family: Ancestral copy number and functional divergence of BAM1 and BAM2 in Brassicaceae
Gene duplication allows for functional divergence and innovation that provide selective advantages. However, in flowering plants genetic studies have revealed that single-gene mutations affecting one of two or more closely related paralogs often fail to cause detectable morphological defects, suggesting functional redundancy. Flowering plants have hundreds of genes encoding leucine-rich repeat receptor-like protein kinases (LRR-RLKs), several of which play important roles in anther development but little is known about their evolutionary history and possible functional divergence. We investigated the evolutionary relationship of the LRR-RLK gene family by phylogenetic analysis and found that these closely related paralogs resulted from multiple duplication events, such as the one resulting in BAM1 and BAM2. We further used qRT-PCR to verify gene expression changes in immature anthers from the bam1/bam2 single and double mutants compared with wild type providing strong evidence that the BAM1 and BAM2 genes have evolved different functions, with differential effects on anther gene expression. Moreover, careful examination of anther development in bam1 and bam2 single mutants revealed previously unrecognized extra cell division in tapetum cell layers. Thus our results from phylogenetic, molecular and morphological analyses uncover sequence and functional differences between paralogs whose single mutants lack obvious fertility defects, effectively revealing functional divergence of duplicate genes.
Data from: Distribution and bioinformatic analysis of the cerato-platanin protein family in Dikarya
The cerato-platanin family is a group of small cysteine-rich fungal proteins new to science. They usually are abundantly secreted extracellularly and are involved in fungus-host interactions. With the advance of available fungal genome sequences, we performed a genomewide study of the distribution of this family in fungi and analyzed the common characteristics of the protein sequences. A total of 55 fungal genomes, including 27 from Ascomycota and 28 from Basidiomycota, were used. A total of 130 cerato-platanin homolog protein sequences were obtained and analyzed. Our results showed that cerato-platanin homologs existed in both Ascomycota and Basidiomycota but were lost in early branches of jelly fungi as well as in some groups with yeast or yeast-like forms in their life cycle. Homolog numbers varied considerably between Ascomycota and Basidiomycota. Phylogenetic analysis suggested that the ancestor of the Dikarya possessed multiple copies of cerato-platanins, which sorted differently in Ascomycota and Basidiomycota, and that this gene family might have expanded in the Basidiomycota. Almost all homologs contained signal peptide sequences, and the length of mature proteins were mainly 105-134 amino acids. Four cysteines involved in forming two disulfide bridges and signature sequences (CSD or CSN) were highly conserved in most homologs. These results indicated a higher diversity of the cerato-platanin family in Basidiomycota than Ascomycota.
Annotation of Sox family proteins of the snail Lymnaea stagnalis (Mollusca: Pulmonata) using transcriptomic data
<p>This annotation of Sox family proteins of Lymnaea stagnalis snail was performed using transcriptome data from the article:</p> <p>Seppälä, O., Walser, JC., Cereghetti, T. et al. Transcriptome profiling of Lymnaea stagnalis (Gastropoda) for ecoimmunological research. BMC Genomics 22, 144 (2021). https://doi.org/10.1186/s12864-021-07428-1</p> <p>Sox2, Sox14, Sox10 and Sox5/6 were annotated using tBLASTn searches against transcriptome from the article upper. </p>
Data from: The divergence and positive selection of the plant-specific BURP-containing protein family
Open the record for dataset details and reuse information.
Data from: Evolution of the leucine-rich repeat receptor-like protein kinase gene family: Ancestral copy number and functional divergence of BAM1 and BAM2 in Brassicaceae
Open the record for dataset details and reuse information.
Data from: Distribution and bioinformatic analysis of the cerato-platanin protein family in Dikarya
Open the record for dataset details and reuse information.
The methyl binding domain 3 (MBD3) family proteins promote Tet2 enzymatic activity for mediating 5mC conversion [microarray]
GEO Series GSE74794. Homo sapiens. 3 samples. Type: Expression profiling by array.
Methylation-dependent and -independent genomic targeting principles of the MBD protein family
GEO Series GSE39610. Mus musculus. 30 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Methylation profiling by high throughput sequencing.
Snf2 Family Protein Fft3 Suppresses Nucleosome Turnover to Promote Epigenetic Inheritance of Heterochromatin and Proper Replication of the Genome [HU + BrdU IP]
GEO Series GSE93273. Schizosaccharomyces pombe. 2 samples. Type: Genome variation profiling by genome tiling array.
Structural and functional investigation of the DHH/DHHA1 family proteins in Deinococcus radiodurans
GEO Series GSE244345. Deinococcus radiodurans. 4 samples. Type: Expression profiling by high throughput sequencing.
An anciently diverged family of RNA binding proteins maintain correct splicing of ultra-long exons through cryptic splice site repression [iCLIP-seq]
GEO Series GSE233497. Homo sapiens. 3 samples. Type: Other.
Snf2 Family Protein Fft3 Suppresses Nucleosome Turnover to Promote Epigenetic Inheritance of Heterochromatin and Proper Replication of the Genome [H3K9me2 ChIP]
GEO Series GSE88902. Schizosaccharomyces pombe. 7 samples. Type: Genome binding/occupancy profiling by genome tiling array.
A frameshift variant in the specificity protein 1 triggers superactivation of Sp1-mediated transcription in familial bone marrow failure
GEO Series GSE152262. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.