Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
170
datasets available to search
ShareScore release 0.9.0
Dataset results
170 results for “Protein Family”
Alpha-Galactosaminidase family GH114 protein from Fusarium solani: X-ray diffraction images
<p>This submission includes h5-files with diffraction images recorded using the Dectris EIGER X 16M detector at the DIAMOND beamline I04. The model of the crystal structure and associated information can be found in the Protein Data Bank entry 9EP6. The model has P 31 2 1 symmetry and three molecules per asymmetric unit. This is a case of crystal pathology – partial disorder. There is electron density for the fourth molecule which could be modelled with occupancy 1/2 and would overlap with a symmetry-related molecule.</p>
PFAM Protein Families Dataset for Machine Learning
<p>A cleaned dataset of protein sequences and protein families for classification. The dataset is exported from PFAM as of June 2023 and curated to achieve the following characteristics:</p> <ul> <li>only protein families included with >=100 sequences</li> <li>families with >2000 sequences are truncated and only represented by 2000 sequences (chosen randomly)</li> <li>only proteins with sequence lengths between 100 and 1000</li> <li>amino acid sequences are form PDB; chains are concatenated only if not similar</li> </ul> <p>The dataset is not balanced, numbers of sequences per family in PFAM and in in dataset are:</p> <pre><code>families: 62, sequences: 46872 total (in PFAM) -> included (in dataset) Number in family ALLERGEN: 122 -> 122 Number in family APOPTOSIS: 381 -> 381 Number in family BIOSYNTHETIC PROTEIN: 346 -> 346 Number in family BIOTIN BINDING PROTEIN: 165 -> 165 Number in family BLOOD CLOTTING: 138 -> 138 Number in family CALCIUM BINDING PROTEIN: 135 -> 135 Number in family CELL ADHESION: 1116 -> 1116 Number in family CELL CYCLE: 511 -> 511 Number in family CHAPERONE: 964 -> 964 Number in family CONTRACTILE PROTEIN: 158 -> 158 Number in family CYTOKINE: 191 -> 191 Number in family DE NOVO PROTEIN: 253 -> 253 Number in family DNA BINDING PROTEIN: 1008 -> 1008 Number in family ELECTRON TRANSPORT: 841 -> 841 Number in family FLUORESCENT PROTEIN: 348 -> 348 Number in family GENE REGULATION: 607 -> 607 Number in family HORMONE: 272 -> 272 Number in family HORMONE GROWTH FACTOR: 159 -> 159 Number in family HORMONE RECEPTOR: 121 -> 121 Number in family HYDROLASE: 19551 -> 2000 Number in family HYDROLASE ANTIBIOTIC: 120 -> 120 Number in family HYDROLASE HYDROLASE INHIBITOR: 2890 -> 2000 Number in family HYDROLASE INHIBITOR: 315 -> 315 Number in family IMMUNE SYSTEM: 3333 -> 2000 Number in family IMMUNOGLOBULIN: 155 -> 155 Number in family ISOMERASE: 2457 -> 2000 Number in family ISOMERASE ISOMERASE INHIBITOR: 139 -> 139 Number in family LECTIN: 139 -> 139 Number in family LIGASE: 1780 -> 1780 Number in family LIGASE LIGASE INHIBITOR: 163 -> 163 Number in family LIPID BINDING PROTEIN: 421 -> 421 Number in family LIPID TRANSPORT: 115 -> 115 Number in family LUMINESCENT PROTEIN: 221 -> 221 Number in family LYASE: 4150 -> 2000 Number in family LYASE LYASE INHIBITOR: 298 -> 298 Number in family MEMBRANE PROTEIN: 1338 -> 1338 Number in family METAL BINDING PROTEIN: 951 -> 951 Number in family METAL TRANSPORT: 409 -> 409 Number in family MOTOR PROTEIN: 195 -> 195 Number in family OXIDOREDUCTASE: 11531 -> 2000 Number in family OXIDOREDUCTASE OXIDOREDUCTASE INHIBITOR: 766 -> 766 Number in family OXYGEN STORAGE: 127 -> 127 Number in family OXYGEN STORAGE TRANSPORT: 260 -> 260 Number in family OXYGEN TRANSPORT: 414 -> 414 Number in family PHOTOSYNTHESIS: 173 -> 173 Number in family PLANT PROTEIN: 255 -> 255 Number in family PROTEIN BINDING: 1613 -> 1613 Number in family PROTEIN TRANSPORT: 693 -> 693 Number in family RECEPTOR: 108 -> 108 Number in family REPLICATION: 161 -> 161 Number in family RNA BINDING PROTEIN: 546 -> 546 Number in family SIGNALING PROTEIN: 2312 -> 2000 Number in family STRUCTURAL PROTEIN: 869 -> 869 Number in family SUGAR BINDING PROTEIN: 1250 -> 1250 Number in family TOXIN: 546 -> 546 Number in family TRANSCRIPTION REGULATION: 3283 -> 2000 Number in family TRANSFERASE: 14724 -> 2000 Number in family TRANSFERASE INHIBITOR: 126 -> 126 Number in family TRANSFERASE TRANSFERASE INHIBITOR: 2465 -> 2000 Number in family TRANSLATION: 370 -> 370 Number in family TRANSPORT PROTEIN: 2782 -> 2000 Number in family VIRAL PROTEIN: 2150 -> 2000</code></pre> <p>Files:</p> <ul> <li>families.csv: list of protein families with frequencies</li> <li>pfam_46872x62.csv: full dataset with amino acid sequences as string (one-letter code)</li> <li>pfam-trn-xy.csv: training dataset with amino acid sequences as tokens (1..25) and padded to a common length of 1000 with padding token 0:</li> </ul> <pre><code> Amino acid | Token | Description -------------------------------- C | 1 | Cysteine S | 2 | Serine T | 3 | Threonine A | 4 | Alanine G | 5 | Glycine P | 6 | Proline D | 7 | Aspartic acid E | 8 | Glutamic acid Q | 9 | Glutamine N | 10 | Asparagine H | 11 | Histidine R | 12 | Arginine K | 13 | Lysine M | 14 | Methionine I | 15 | Isoleucine L | 16 | Leucine V | 17 | Valine W | 18 | Tryptophan Y | 19 | Tyrosine F | 20 | Phenylalanine B | 21 | Aspartic acid or Asparagine Z | 22 | Glutamic acid or Glutamine J | 23 | Leucine or Isoleucine U | 24 | Selenocysteine X | 25 | Unknown amino acid . | 0 | padding token</code></pre> <p> </p> <ul> <li>pfam-trn-labels.csv: plain-text labels for training data</li> <li>pfam-tst-xy.csv</li> <li>pfam-tst-labels.csv: test data</li> <li>pfam-balanced-trn-xy.csv</li> <li>pfam-balanced-trn-labels.csv:</li> <li>pfam-balanced-tst-xy.csv</li> <li>pfam-balanced-tst-labels.csv: balanced datasets, created by oversampling.</li> </ul>
Alpha-Galactosaminidase family GH191 protein from Environmental sample (99.2% identity to Myxococcus fulvus enzyme): X-ray diffraction images
<p><span>This submission includes a zip archive of diffraction images recorded with the Dectris EIGER X 9M detector at the DIAMOND beamline I04-1. The model of the crystal structure and associated information can be found in the Protein Data Bank entry 9EP5. This is a case of crystal pathology – partial disorder. The model has C 2 2 21 symmetry and two molecules per asymmetric unit with occupancies 1 and 1/3. The molecule with partial occupancy overlaps with a symmetry related molecule.</span></p>
Datasets for "The Venturia inaequalis effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins "
<p>Datasets for preprint entitled "The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi"</p> <p><strong>1) ViAnnotation.gff3</strong><br> Gene annotation of <em>Venturia inaequalis</em> MNH120 (<a href="https://genome.jgi.doe.gov/Venin1/Venin1.home.html">https://genome.jgi.doe.gov/Venin1/Venin1.home.html</a>) generated as part of the study "The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi". </p> <p>Gene reannotation was performed to include genes that would have been missed in the previous annotation by Deng et al. (2017), especially those genes encoding putative effector proteins, which are difficult to predict. For this purpose, we used a three-step approach. In the first step, coding sequences (CDSs) from <em>V. inaequalis</em> isolate 05/172, which were predicted as part of a previous study by Passey et al. (2018) (<a href="https://journals.asm.org/doi/full/10.1128/MRA.01062-18">https://journals.asm.org/doi/full/10.1128/MRA.01062-18</a>), were downloaded from the National Center for Biotechnology Information (<a href="https://www.ncbi.nlm.nih.gov/nuccore/QFBF00000000.1/">https://www.ncbi.nlm.nih.gov/nuccore/QFBF00000000.1/</a>) and mapped to the MNH120 genome using GMAP v2021-02-22. In the second step, RNA-seq reads from one biological replicate representing each <em>in planta</em> time point of <em>Malus domestica</em> infection by <em>V. inaequalis </em>(12 hour post-inoculation [hpi], 24 hpi, 2 days post-inoculation [dpi], 3 dpi, 5 dpi, 7 dpi), as well as one time point representing growth of the fungus in culture, were mapped to the MNH120 genome using HISAT2 v2.2.1. Then, a genome-guided <em>de novo</em> transcriptome assembly was performed using Trinity v2.12.0 and likely CDSs were identified using Transdecoder v5.5.0 (<a href="https://github.com/TransDecoder/TransDecoder">https://github.com/TransDecoder/TransDecoder</a>) in conjunction with a minimum open frame (ORF) length of 50 amino acids. Finally, in the third step, all annotations were visualized in Geneious v9.05, together with the previous annotation from Deng et al. (2017), and a manual curation was performed to create a consensus prediction. Note: this reannotation was generated with the aim of identifying as many genes as possible, and as a result, it contains many spurious genes. </p> <p><strong>2) Protein_sequences_ViAnnotation.fasta</strong></p> <p><strong>3) ECs_Families_AlphaFold.zip</strong></p> <p>This dataset is made up of predicted protein tertiary structures representing the main member of each up-regulated <em>V. inaequalis</em> effector candidate family. Structures were predicted using Alphafold with the ColabFold server (<a href="https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n">https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n</a>). In cases where the effector candidate had less than 30 proteins with amino acid sequence similarity in the NCBI database, a custom multiple sequence alignment (MSA) was generated and used as input for AlphaFold2. Here, mature protein sequences were used.</p> <p><strong>4) singletons_AlphaFold_OpenSourceCASP14.zip</strong></p> <p>This dataset set is made up of predicted protein tertiary structures representing up-regulated<em> V. inaequalis</em> singleton effector candidates. Structures were predicted using AlphaFold (<a href="https://github.com/deepmind/alphafold">https://github.com/deepmind/alphafold</a>) open source code v2.0.1 and v2.1.0, with pre-set casp14, max_template_date: 2020-05-14. Mature protein sequences were used as input. </p> <p><strong>5) ECs_Avrs_phytopathogens_AlphaFold.zip</strong></p> <p>Predicted tertiary structures of avirulence (Avr) proteins or candidate Avr proteins from other fungal pathogens included in the "The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi" study. These structures were predicted using Alphafold with the ColabFold server (<a href="https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n">https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n</a>). Mature protein sequences were used as input. </p> <p>If you have any questions about the datasets, please contact us.<br> Mercedes Rocafort: <a href="mailto:m.rocafort.ferrer@massey.ac.nz">m.rocafort.ferrer@massey.ac.nz</a><br> Carl Mesarich: <a href="mailto:c.mesarich@massey.ac.nz">c.mesarich@massey.ac.nz</a></p>
Protein codes for tetrameric and dimeric 6-phosphogluconate dehydrogenases and their signature sequence relevant to the cofactor specificity in the 6PGDH family.
<p>Here we present the Uniprot or Genbank code for tetrameric and dimeric 6-phosphogluconate dehydrogenases collected from different organisms. Besides, we include the sequence of the <span class="math-tex">\(\beta2-\alpha2\)</span> motif regarding the cofactor specificity in the 6PGDH family. </p> <p>The sequence data allows generating a phylogenetic tree of the 6PGDH family. The enzymes cluster first by their oligomerization state and then by their sequence in the <span class="math-tex">\(\beta2-\alpha2\)</span> motif, as is shown in the annexed figure. </p>
Data generated for the publication: Keeping it in the family: Using protein family templates to rescue poor AlphaFold models unliked
<p>Data and manuscript of:</p> <p>Keeping it in the family: Using protein family templates to rescue low confidence AlphaFold2 models</p> <p>Francesco Costa1, Matthias Blum1 and Alex Bateman1</p> <ol> <li>European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton. CB10 1SD. UK</li> </ol> <ul> <li>results contains the workflow results;</li> <li>AF2_seed contains results of the comparison with AF2 run with multiple seeds;</li> </ul>
Our predicted Crinkler (CRN) family effector proteins and the corresponding GFF3 files across 128 Phytophthora isolates
<p>These files contain predicted Crinkler (CRN) family effector proteins and the corresponding GFF3 files across 128 Phytophthora isolates.</p>
Рис. 7. Варианты преΑсказанной Αоменной структуры моΛекуΛ аΑгезии гемоцитов моΛΛюсков Planorbarius corneus. УсΛовные обозначения и сокращения: 1–3 — β-интегрины, 4–5 — α-интегрины, 6–7 — сеΛектины, 8–11 — моΛекуΛы семейства САМ (сell adhesiom molecues), INB — субъеΑиницы β-интегрина, IntegrinBcyt — цитопΛазматический Αомен β-интегрина, CY — цистатинопоΑобный Αомен, Int alpha — Αомен α-интегрина, FN3 — Αомен фибронектина типа 3, CCP — Αомен контроΛя компΛемента Fig. 7. Variants of the predicted domain structure of adhesion molecules from hemocytes of Planorbarius corneus molluscs. Symbols and abbreviations: 1–3 — β-integrins, 4–5 — α–integrins, 6–7 — selectins, 8–11 — molecules of the СAM family (cell adhesion molecules), INB — β-integrin subunits, IntegrinBcyt — cytoplasmic domain of β-integrin, CY — cystatin-like domain, Int alpha — α-integrin domain, FN3 — fibronectin type 3 domain, CCP — complement control protein domain in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)
Рис. 7. Варианты преΑсказанной Αоменной структуры моΛекуΛ аΑгезии гемоцитов моΛΛюсков Planorbarius corneus. УсΛовные обозначения и сокращения: 1–3 — β-интегрины, 4–5 — α-интегрины, 6–7 — сеΛектины, 8–11 — моΛекуΛы семейства САМ (сell adhesiom molecues), INB — субъеΑиницы β-интегрина, IntegrinBcyt — цитопΛазматический Αомен β-интегрина, CY — цистатинопоΑобный Αомен, Int alpha — Αомен α-интегрина, FN3 — Αомен фибронектина типа 3, CCP — Αомен контроΛя компΛемента Fig. 7. Variants of the predicted domain structure of adhesion molecules from hemocytes of Planorbarius corneus molluscs. Symbols and abbreviations: 1–3 — β-integrins, 4–5 — α–integrins, 6–7 — selectins, 8–11 — molecules of the СAM family (cell adhesion molecules), INB — β-integrin subunits, IntegrinBcyt — cytoplasmic domain of β-integrin, CY — cystatin-like domain, Int alpha — α-integrin domain, FN3 — fibronectin type 3 domain, CCP — complement control protein domain
Data from: The structure of an ancient genotype-phenotype map shaped the functional evolution of a protein family
Open the record for dataset details and reuse information.
How can we biochemically validate protein function predictions with the Ras GTPase family? - Associated data
<p>This is the data that accompanies the pub "<a href="https://doi.org/10.57844/arcadia-74ad-345f">How can we biochemically validate ProteinCartography with the Ras GTPase family?</a>" It's part of a group of pubs focused on validating ProtienCartography that begins with "<a href="https://doi.org/10.57844/arcadia-cae9-96c4">A strategy to validate protein functions <em>in vitro</em></a><a href="https://doi.org/10.57844/arcadia-cae9-96c4">." </a></p> <p>For this repository, we ran ProteinCartography <a href="https://github.com/Arcadia-Science/ProteinCartography/releases/tag/v0.5.0">v0.5.0</a> using human HRas and KRas as our inputs for a single run (UniProt ID: <a href="https://www.uniprot.org/uniprotkb/P01112/entry">P01112</a> and <a href="https://www.uniprot.org/uniprotkb/P01116/entry">P01116</a>). We asked for 3,000 Foldseek hits and 7,000 BLAST hits for a total of 10,000 structures. The updated configuration file is in the zipped folder in this repository. Also included in the zipped folder are the inputs, structures of all hits, and all ProteinCartography results. </p> <p>Finally, we created a custom overlay for the protein map using this <a href="https://github.com/Arcadia-Science/2023-actin-embedding/blob/main/notebooks/3_plotting_overlays.ipynb">notebook</a> and the manually annotated TSV file in this repository, where we denoted which group of substrates a protein is predicted to act on based on its annotation from UniProt.</p>
How can we biochemically validate protein function predictions with the deoxycytidine kinase family? - Associated data
<p>This is the data that accompanies the pub "<a href="https://doi.org/10.57844/arcadia-1e5d-e272">How can we biochemically validate ProteinCartography with the deoxycytydine kinase family?</a>" It's part of a group of pubs focused on validating ProtienCartography that begins with "<a href="https://doi.org/10.57844/arcadia-cae9-96c4">A strategy to validate protein functions <em>in vitro</em></a><a href="https://doi.org/10.57844/arcadia-cae9-96c4">." </a></p> <p>For this repository, we ran ProteinCartography <a href="https://github.com/Arcadia-Science/ProteinCartography/releases/tag/v0.5.0">v0.5.0</a> on the deoxycytidine kinase (dCK) using human dCK as our input (UniProt ID: <a href="https://www.uniprot.org/uniprotkb/P27707/entry">P27707</a>). We asked for 3,000 Foldseek hits and 7,000 BLAST hits for a total of 10,000 structures. The updated configuration file is in the zipped folder in this repository. Also included in the zipped folder are the inputs, structures of all hits, and all ProteinCartography results. </p> <p>Finally, we created a custom overlay for the protein map using this <a href="https://github.com/Arcadia-Science/2023-actin-embedding/blob/main/notebooks/3_plotting_overlays.ipynb">notebook</a> and the manually annotated TSV file in this repository, where we denoted which group of substrates a protein is predicted to act on based on its annotation from UniProt.</p>
Tudor genes of Holozoa: Early evolution and within Metazoa diversification of a multifaceted protein family
<p>Early metazoan evolution was characterized by the expansion of many gene families involved in novel multicellularity-related functions, like the Tudor family. In eukaryotes, Tudor genes are numerous and heterogeneous, mostly associated with gene expression regulation. However, the family underwent a lineage-specific expansion in animals, with novel elements almost exclusively involved in the germline-specific regulation of retrotransposons through piRNAs (as spatiotemporal regulators of the key-element Piwi, another previously supposedly animal-specific gene). In the present analysis, we used online-available proteomes for a total of 25 major taxonomic groups to characterize the Tudor gene family at a holozoan-wide level, and we confirmed the apomorphic expansion of piRNA-related Tudor genes in animals. However, we could also interestingly observe the presence of elements of the piRNA pathway, both Tudor and Piwi genes, in some Ichthyosporea species, suggesting that some elements of the pathway were already present in the last common ancestor of Holozoa. Moreover, we observed an outstanding variability (34-fold) of Tudor gene number both between and within metazoan phyla, that could be associated with convergent genomic and phenotypic evolutions. Expansions were usually sided by whole genome duplications and/or life history traits such as parthenogenesis, possibly leading to the expansion of retrotransposon silencing pathways. Reductions were instead mostly associated with overall phenotypic and genomic simplifications, like almost all endoparasites of our dataset. Lastly, we phylogenetically tested a previously proposed model for the evolution of the three possible secondary structures of the Tudor domains and we could mostly (but not completely) confirm the model.</p>
Protein families for 2032 Saccharomyces cerevisiae genome assemblies
<p>The protein families obtained with different cluster cutoffs for different sets of genome assemblies as well as the marker genes are presented in text files. In each file, a row indicates a family and a column (separated by TAB) indicates a genome assembly. Protein IDs for multiple homologues from the same assembly are separated by '|'. The family ID is shown in the first column and the tag for each assembly is shown in the first row. Cluster cutoffs used are 50%, 60%, 70%, 80%, and 90%. Genome sets shown are all genomes (all), non-redundant (nr) genomes, medium-high-quality genomes (mhq), and high-quality genomes (hq).</p>
Tudor genes of Holozoa: Early evolution and within Metazoa diversification of a multifaceted protein family
Open the record for dataset details and reuse information.
Supplementary tables S5, S7, S9, S10, original protein models fasta files used for alignments, aligned and manually curated protein modes files used for phylogenies (PHYLIP format), and phylogenetic trees of plant cell wall decomposition gene families from 44 basidiomycete genomes (.tre files)
<p><span><span><span><span><span><span><span><span><span><span><span>Litter-decomposing Agaricales play key role in terrestrial carbon cycling, but little is known about their decomposition mechanisms. We assembled datasets of 42 gene families involved in plant-cell-wall decomposition from seven newly sequenced litter decomposers and 35 other Agaricomycotina members, mostly white-rot and brown-rot species. Using sequence similarity and phylogenetics, we split the families into phylogroups and compared their gene composition across nutritional strategies. Subsequently, we used Raman spectroscopy to examine the ability of litter decomposers, white-rot fungi, and brown-rot fungi to decompose crystalline cellulose. Both litter decomposers and white-rot fungi share the enzymatic cellulose decomposition, whereas brown-rot fungi possess a distinct mechanism that disrupts cellulose crystallinity. However, litter decomposers and white-rot fungi differ with respect to hemicellulose and lignin degradation phylogroups, suggesting adaptation of the former group to the litter environment. Litter decomposers show high phylogroup diversity, which is indicative of high functional versatility within the group, whereas a set of white-rot species shows adaptation to bulk-wood decomposition. In both groups, we detected species that have unique characteristics associated with hitherto unknown adaptations to diverse wood and litter substrates. Our results suggest that the terms white-rot fungi and litter decomposers mask a much larger functional diversity.</span></span></span></span></span></span></span></span></span></span></span></p>
Dataset for Ha and Aylward 'Automated classification of giant virus genomes using a random forest model built on trademark protein families'
<ul><li>Genome sets used for model training and testing</li><li>Custom Python script that generated fragmented genomes at random completeness levels</li></ul>
Protein aggregation and calcium dysregulation are the earliest hallmarks of familial Parkinson's disease in human midbrain dopaminergic neurons
<p>Mutations in the <em>SNCA</em> gene cause autosomal dominant Parkinson’s disease (PD), with loss of dopaminergic neurons in the substantia nigra, and aggregation of α-synuclein. The sequence of molecular events that proceed from an <em>SNCA</em> mutation during development, to end stage pathology is unknown. Utilising human induced pluripotent stem cells (hiPSCs), we resolved the temporal sequence of SNCA induced pathophysiological events in order to discover early, and likely causative, events. Our small molecule-based protocol generates highly enriched midbrain dopaminergic (mDA) neurons: molecular identity was confirmed using single-cell RNA sequencing and proteomics, and functional identity through dopamine synthesis, and measures of electrophysiological activity. At the earliest stage of differentiation, prior to maturation to mDA neurons, we demonstrate the initial formation of small β-sheet rich oligomeric aggregates, in <em>SNCA</em>-mutant cultures. Aggregation persists and progresses, ultimately resulting in the accumulation of phosphorylated aggregates. Impaired intracellular calcium signalling, increased basal calcium, and impairments in mitochondrial calcium handling occurred early at day 34-41 post differentiation. Once midbrain identity fully developed, at day 48-62 post differentiation, <em>SNCA</em>-mutant neurons exhibited mitochondrial dysfunction, oxidative stress, lysosomal swelling and increased autophagy. Ultimately these multiple cellular stresses lead to abnormal excitability, altered neuronal activity, and cell death. Our differentiation paradigm generates an efficient model for studying disease mechanisms in PD, and highlights that protein misfolding to generate intraneuronal oligomers is one of the earliest critical events driving disease in human neurons, rather than a late-stage hallmark of the disease.</p>
Supplementary Materials: High-Quality Genome of a novel species belonging to the family Thermosynechococcaceae from Namibia and characterization of its protein expression patterns at elevated temperatures
<p><strong>Supplementary Materials</strong> to the scientific research article "<strong>High-Quality Genome of a novel species belonging to the family<em> Thermosynechococcaceae </em>from Namibia and characterization of its protein expression patterns at elevated temperatures</strong>" in MicrobiologyOpen.</p> <p>Including raw data of: the BCG analysis with antismash, dbCAN3 analysis of carbohydrate active enzymes, eggNOG mapper and KEGG (blastKOALA) functional annotations, whole genome alignments with progressiveMauve, ortholog-based phylogenetics with OrthoFinder, RAST annotations of the novel <em>Thermosynechococcus</em> sp. Okahandja<em> </em>and the type strains <em>Thermosynechococcus lividus</em> PCC 6715-6717, the full length 16S rRNA sequence extracted from the genome, proteomics analyses, the novel <em>T. </em>sp. Okahandja<em> </em>annotated genome, additional STEM microscopy figures and CRISPR analyses.</p>
Fueling ab initio folding with oceanic metagenomics enables structure and function predictions of new protein families
<p>Code and protein sequence database to construct multiple sequence alignment from Tara Ocean data.</p>
HerpesFolds: A proteome-wide structural systems approach reveals insights into protein families and activities of all nine human herpesviruses
<p>These are the AlphaFold output files and ChimeraX sessions for "HerpesFolds: A proteome-wide structural systems approach reveals insights into protein families and activities of all nine human herpesviruses" by Timothy K. Soh, Sofia Ognibene, Saskia Sanders, Robin Schäper, Benedikt B. Kaufer, and Jens B. Bosse.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.