Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
302
datasets available to search
ShareScore release 0.9.0
Dataset results
302 results for “protein sequences”
Protein haplotype sequences obtained by ProHap from the Haplotype Reference Consortium Release 1.1 dataset
<p>Database of protein sequences obtained using ProHap (<a href="https://github.com/ProGenNo/ProHap">https://github.com/ProGenNo/ProHap</a>) on the data set of phased genotypes published by the Haplotype Reference Consortium, Release 1.1 (<a href="https://ega-archive.org/datasets/EGAD00001002729" target="_blank" rel="noopener">https://ega-archive.org/datasets/EGAD00001002729</a>). We used Ensembl v.110 for the mapping of coordinates between genes, exons, and transcripts.</p> <p>Release 1.1 of the HRC is provided aligned with the GRCh37 reference genome. We have performed a liftover to the GRCh38 reference using GeneBe (https://genebe.net/tools/liftover). Variants for which the reported alternative allele is considered as reference in GRCh38 were removed. A threshold of 1% minor allele frequency was applied to filter the remaining variants. After translation, a frequency threshold of 0.5% was applied to filter the resulting unique non-canonical sequences. The complete configuration file for the ProHap run is attached to this repository.</p> <p>This dataset contains one compressed directory, contains the following files:</p> <ul> <li>F1: The concatenated fasta file ready to be used with search engines, contains the following: <ul> <li>Protein haplotype sequences obtained by ProHap</li> <li>Reference proteome as per Ensembl v. 110</li> <li>Contaminant sequences from the cRAP project (<a href="https://www.thegpm.org/crap/">https://www.thegpm.org/crap/</a>)</li> <li>The file is provided in two formats - full and simplified. The simplified fasta contains only the artificial protein identifier and the matching gene name, and is optimised for compatibility with a wide range of tools. For annotation of peptides using the PeptideAnnotator, please provide the header (F1.2) in addition to the fasta file. </li> </ul> </li> <li>F2: Additional information about the haplotype sequences, to be used for mapping identified peptides to the original haplotypes</li> <li>F3: Translations of haplotype cDNA sequences, before merging with the reference proteome</li> </ul> <p>For further description of the files, please refer to <a href="https://github.com/ProGenNo/ProHap/wiki/Output-files">https://github.com/ProGenNo/ProHap/wiki/Output-files</a>.</p> <p>For the usage of these databases with search engines, and downstream anaylsis of identified peptides, please refer to the project's wiki page: <a href="https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches">https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches</a>.</p> <p>When using these databases in your publication, please cite: Vašíček, J., Kuznetsova, K.G., Skiadopoulou, D. <em>et al.</em> ProHap enables human proteomic database generation accounting for population diversity. <em>Nat Methods</em> (2024). <a href="https://doi.org/10.1038/s41592-024-02506-0">https://doi.org/10.1038/s41592-024-02506-0</a></p>
Overcoming Limitation of AlphaFold2 by Deep-mutational Scanning and Stability-Selection of Protein Sequences
<p>This repository contains the processed datasets and corresponding code used in our study. While AlphaFold2 revolutionizes protein structure prediction, its accuracy critically depends on evolutionary information from natural homologs—limiting applications for proteins with sparse sequence families. Here, we bypass this bottleneck by employing deep mutational scanning and stability-guided selection to generate artificial homologs. Fed into AlphaFold2, these synthetic sequences match the accuracy achieved on well-predicted proteins with rich natural homology, while providing highly accurate predictions for difficult targets—including orphan proteins previously deemed "unpredictable." Our approach achieves high accuracy (<3 Å RMSD for 5/8 and <2 Å RMSD for 7/8 targets after excluding intrinsically flexible regions). Thus, integrating simple, scalable molecular biology (mutagenesis/selection) with high-throughput sequencing can deliver the accuracy similar to but at a fraction of the cost and time of traditional experimental structure-determination methods. This hybrid framework could democratize high-resolution structural biology, opening avenues to determine structures of protein complexes, modified proteins, and condition-dependent conformations. </p>
Protein haplotype sequences obtained by ProHap from the Human Pangenome Reference Consotruim dataset
<p>Database of protein sequences obtained using ProHap (<a href="https://github.com/ProGenNo/ProHap">https://github.com/ProGenNo/ProHap</a>) on the data set of phased genotypes published by the Human Pangenome Reference Consotruim (HPRC), first release (<a href="https://github.com/human-pangenomics/hpp_pangenome_resources">https://github.com/human-pangenomics/hpp_pangenome_resources</a>), 44 samples. We used Ensembl v.110 for the mapping of coordinates between genes, exons, and transcripts.</p> <p>This repository contains one database created using all 43 samples of the HPRC release (the haplotypes of the sample NA21309 did not encode any non-canonical sequences), and then a database for each of the 43 samples separately. No filtering on allele frequency or haplotype frequency was applied in any of the databases. The complete configuration file for the ProHap run is attached to this repository.</p> <p>There is one compressed directory for each of the databases, containing the following files:</p> <ul> <li>F1: The concatenated fasta file ready to be used with search engines, contains the following: <ul> <li>Protein haplotype sequences obtained by ProHap</li> <li>Reference proteome as per Ensembl v. 110</li> <li>Contaminant sequences from the cRAP project (<a href="https://www.thegpm.org/crap/">https://www.thegpm.org/crap/</a>)</li> <li>For this dataset, only the simplified format is provided. The simplified fasta contains only the artificial protein identifier and the matching gene name, and is optimised for compatibility with a wide range of tools. For annotation of peptides using the PeptideAnnotator, please provide the header (F1.2) in addition to the fasta file. </li> </ul> </li> <li>F2: Additional information about the haplotype sequences, to be used for mapping identified peptides to the original haplotypes</li> <li>F3: Translations of haplotype cDNA sequences, before merging with the reference proteome</li> </ul> <p>For further description of the files, please refer to <a href="https://github.com/ProGenNo/ProHap/wiki/Output-files">https://github.com/ProGenNo/ProHap/wiki/Output-files</a>.</p> <p>For the usage of these databases with search engines, and downstream anaylsis of identified peptides, please refer to the project's wiki page: <a href="https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches">https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches</a>.</p> <p>When using these databases in your publication, please cite: Vašíček, J., Kuznetsova, K.G., Skiadopoulou, D. <em>et al.</em> ProHap enables human proteomic database generation accounting for population diversity. <em>Nat Methods</em> (2024). <a href="https://doi.org/10.1038/s41592-024-02506-0">https://doi.org/10.1038/s41592-024-02506-0</a></p>
Discoba protein sequences for protein structure predictions
<p>Comprehensive database of Discoba protein sequences, gathered for the purpose of improving protein structure predictions of Discoba species (including <em>Trypanosoma </em>and <em>Leishmania</em>) by AlphaFold and RoseTTAFold. Originally gathered for use with: https://github.com/zephyris/discoba_alphafold</p>
Supplementary Data Files for the paper "Intrinsically disordered compositional bias in proteins: Sequence traits, region clustering, and generation of hypothetical functional associations"
<div> <div> <div> <div> <p><strong>Supplementary data files relating to <a href="https://doi.org/10.1177/11779322241287485">https://doi.org/10.1177/11779322241287485. </a></strong></p> <p><strong><span>Suppl. File 1: Protein Family Clusters.</span></strong></p> <p><strong><span>Suppl. File 2: Cluster GO enrichments/depletions. </span></strong></p> <p><strong><span>Suppl. File 3: The raw ID-CBR data with annotations. </span></strong></p> <p><strong><span>Suppl. File 4: ­ID-CBR Cluster membership.</span></strong></p> <p><strong><span>Each file has an explanatory header. </span></strong></p> <p> </p> </div> </div> </div> </div>
Predicting Exon Criticality from Protein Sequence
<p>Exon ByPASS (predicting Exon-skipping Based in Protein amino acid SequenceS), predictions on test exons from Human and Mouse transcripts. The exons in the test set from the two genomes are those that are not predicted to be skippable in hg38 and mm10 annotation and are also exons that in-frame when skipped. The preprocessed data includes the ensemble transcript id and exon rank as well the amino acid sequence for the upstream, downstream, and exon of interest. Additionally, the table contains the output probability from the model in the last column. The input data is transformed data of the amino acid sequence that is the Exon ByPASS model can use as an input.</p>
PSSH2 - database of protein sequence-to-structure homologies (including Sars-CoV-2 structures)
<p><strong>Protein sequence and structure data</strong></p> <p>This data set contains data from Uniprot (in the files called protein_sequence, protein_synonyms, protein_names, organism_synonyms) and PDB (in the files called PDB and PDB_chain) as used by the <a href="https://github.com/ODonoghueLab/Aquaria">Aquaria web resource</a> at the time of download (2022-02-08).</p> <p> </p> <p><strong>The PSSH2 data set</strong><br> <br> PSSH2 is a database of protein sequence-to-structure homologies based on HHblits, an alignment method employing iterative comparisons of hidden Markov models (HMMs). To ensure the highest possible final alignment quality for matches in Aquaria using HHblits, we first calculate HMM profiles for each unique PDB sequence (PDB_full) and also for each unique Swiss-Prot sequence. We generated PSSH2 using HHblits to find similarities between HMMs from PDB and HMMs from UniProt sequences.</p> <p> </p> <p><strong>Calculating PSSH2</strong></p> <p>The Swissprot and PDB data was downloaded in November 2021.<br> Generating PSSH2: We used <a href="https://gwdu111.gwdg.de/~compbiol/uniclust/2021_03/UniRef30_2021_03.tar.gz">UniRef30_2021_03</a> (originally called UniRef30_2021_06) from HH-suite, a database of non-redundant UniProt sequence clusters in which the highest pairwise sequence identity between clusters was 30%. The HHblits code and the code for running the calculations was retrieved from git (https://github.com/soedinglab/hh-suite.git and https://github.com/aschafu/PSSH2.git respectively) at the respective time of calculation in the timeframe until December 2021. <br> </p> <p><strong>PDB based sequence-to-structure alignments</strong></p> <p>In addition to the PSSH2 data, new PDB structures were retrieved based on the primary accession of the proteins, by querying for all chains in all PDB entries with exact matches using the sequence cross references records given in PDB. Sequence-to-structure alignments were then created, again based on information provided in each PDB entry. These are contained in the PDBchain data.</p> <p>This data covers sequences and PDB structures in the timeframe until February 2022. </p> <p> </p> <p><strong>Evaluating PSSH2</strong></p> <p>The resulting alignment data was analysed using CATH domain assignments downloaded from /cath/releases/all-releases/v4_2_0/cath-classification-data/ to define correct hits and false hits: </p> <ul> <li>The set of query sequences is defined by the CATH non-redundant S40_overlap_60 dataset (ftp://orengoftp.biochem.ucl.ac.uk/cath/releases/all-releases/v4_2_0/non-redundant-data-sets/)</li> <li>The set of all expected hits are all pdb structures containing a domain with the same CATH code if contained in the set of processed sequences (-> all) or only if also contained in the set of non redundant sequences (-> nr40).</li> <li>The set of true positives is defined by sharing the same CATH code up to the level of homology ("CATH") or up to the level of topology ("CAT").</li> </ul> <p>The data was evaluated with respect to false discovery rate (FDR) and recall (true positive rate TPR) by cumulatively considering all hits with an E-value below the threshold ("C") or in bins with an E-value between the threshold and one tenth of the threshold ("B"). This evaluation was carried out for the data obtained in November 2021 (202111) as well as previous data from October 2020 (202010), February 2020 (202002) and September 2017 (201709). The results are collected in <a href="https://zenodo.org/api/files/445add84-fcf1-4dfe-b8a1-63dc55f378ee/PSSH%20CATH%20validation.csv?versionId=a5df7473-6efd-442b-b422-3944e9452003">PSSH CATH validation.csv</a>. </p> <p> </p> <p><strong>Known errors</strong></p> <p>Due to processing error, the profile of pdb structure 5fia A / B (sequence md5 052667679fc644184f40063c7602c9e1) is incomplete in the pdb_full hhblits database which led to further errors in generating sequence based alignments for sequences for 1vtm P (sequence md5 c844aff103449363cb8489c78c58ebf1) and 434t A / B (sequence md5 d67aa1c3a36492c719cb48b5e7ecc624).<br> <br> </p>
Annotation of Phytozome V12 protein plant sequences using the ragp pipeline for hydroxyproline-rich glycoprotein mining
<p>Hydroxyproline aware annotation of hydroxyproline-rich glycoprotein (HRGP) sequences was performed on sequence data from 62 plant proteomes obtained from Phytozome database (<a href="https://phytozome.jgi.doe.gov/pz/portal.html">https://phytozome.jgi.doe.gov/pz/portal.html</a>, version 12) using the ragp R package (<a href="https://github.com/missuse/ragp">https://github.com/missuse/ragp</a>, version 0.3.0.0001). </p> <p>In each archive a single comma separated value table (.csv) is present along with a README.txt file describing the contents of the corresponding .csv file. The archives are:</p> <p>- phytozome_V12.tar.gz - sequences from 62 plant proteomes (phytozome V12) with a total of 2797062 protein sequences.</p> <p>- phytozome_V12_phobius.tar.gz -<strong> </strong> Signal peptide prediction using Phobius (<a href="http://phobius.sbc.su.se/">http://phobius.sbc.su.se/</a>) on sequences present in phytozome_V12.tar.gz.</p> <p>- phytozome_V12_signalp.tar.gz -<strong> </strong> Signal peptide prediction using SignalP 4.1 (<a href="http://www.cbs.dtu.dk/services/SignalP-4.1/">http://www.cbs.dtu.dk/services/SignalP-4.1/</a>) on sequences present in phytozome_V12.tar.gz.</p> <p>- phytozome_V12_targetp.tar.gz - Signal peptide prediction using TargetP 1.1 (<a href="http://www.cbs.dtu.dk/services/TargetP/">http://www.cbs.dtu.dk/services/TargetP/</a>) on sequences present in phytozome_V12.tar.gz.</p> <p>- phytozome_V12_predict_hyp.tar.gz - Probability of proline hydroxylation for each proline from 266135 protein sequences which were predicted to be secreted by a majority vote (using Phobius, SignalP 4.1 and TargetP 1.1).</p> <p>- phytozome_V12_maab.tar.gz - Motif and amino acid bias (MAAB) classification of hydroxyproline-rich glycoproteins performed on 266135 protein sequences which were predicted to be secreted by a majority vote (using Phobius, SignalP 4.1 and TargetP 1.1). The number of predicted hydroxyprolines in each sequence is also indicated (based on predictions provided in phytozome_V12_predict_hyp.tar.gz).</p> <p>- phytozome_V12_scan_ag.tar.gz. - Hydroxyproline aware arabinogalactan motif scan performed on 266135 protein sequences which were predicted to be secreted by a majority vote (using Phobius, SignalP 4.1 and TargetP 1.1). Hydroxyproline predictions are provided in phytozome_V12_predict_hyp.tar.gz. </p> <p>- phytozome_V12_scan_ag_hmmscan.tar.gz - Detection of domains in a subset of protein sequences which were found to contain arabinogalactan motifs (a subset of phytozome_V12_scan_ag.tar.gz).</p> <p>The list of the 62 plant species is provided in phytozome_V12.tar.gz README.txt.</p> <p>For questions contact mdragicevic@ibiss.bg.ac.rs.</p> <p> </p> <p> </p>
High-quality large curated dataset of protein sequences (1.83 million) and their corresponding Position Specific Scoring Matrices
<p>As part of his master thesis at the Rostlab, which is located at the Technical University of Munich (TUM), Mr. Issar Arab developed the first language model that encodes evolutionary information of proteins explicitly. The pre-training involved the creation of a novel high-quality dataset of protein sequences (around 1.83 million proteins, or ~0.8 Billion amino acids) with their corresponding Position Specific Scoring Matrices (PSSMs). Those matrices reflect the relative frequency of each amino acid at each position in a protein and is derived from evolutionarily related proteins.</p> <p>Mr. Arab makes this work publicly available to help other researchers speed up their work to leverage AI to learn the representation of protein evolutionary information more explicitly. The set of sequences was derived by extracting all PSSMs from the <a href="https://predictprotein.org/">PredictProtein</a> (PP) cache, which were also part o the UniProt Reference Cluster with 50% sequence identity (uniref50 2019_12). The overlap between PP and uniref50 was further filtered to only include high-quality samples, e.g. only multiple sequence alignments with a certain number of aligned sequences were considered. The processing led to a training set of 1.83 Million sequences, a validation set of 879 instances, and a test set of 879 entries. The training data of proteins is reduced to 40% sequence identity, with respect to the validation/test sets, and contains sequences ranging between 18 and 9858 residues in length.</p> <p>Refer to the Jupyter notebook for a detailed description of the files' structure and a Python code snippet to correctly manipulate this data.</p> <p>To access the full original work, please visit the following link: <a href="https://mediatum.ub.tum.de/node?id=1579236">Manuscript</a> <br><br><strong>Note:</strong> The dataset was recently used to fine tune a protein sequence language model (<a href="https://github.com/issararab/PEvoLM">PEvoLM</a>). The work was presented at the CIBCB'23 conference. If you use PEvoLM or this dataset in your work, please cite the following publication:</p> <p>- Issar Arab, <strong>PEvoLM: Protein Sequence Evolutionary Information Language Model</strong>, <em>IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), Eindhoven, Netherlands</em>, (2023), pp. 1-8, doi:<a href="https://ieeexplore.ieee.org/document/10264890">10.1109/CIBCB56990.2023.10264890</a></p>
Protein codes for tetrameric and dimeric 6-phosphogluconate dehydrogenases and their signature sequence relevant to the cofactor specificity in the 6PGDH family.
<p>Here we present the Uniprot or Genbank code for tetrameric and dimeric 6-phosphogluconate dehydrogenases collected from different organisms. Besides, we include the sequence of the <span class="math-tex">\(\beta2-\alpha2\)</span> motif regarding the cofactor specificity in the 6PGDH family. </p> <p>The sequence data allows generating a phylogenetic tree of the 6PGDH family. The enzymes cluster first by their oligomerization state and then by their sequence in the <span class="math-tex">\(\beta2-\alpha2\)</span> motif, as is shown in the annexed figure. </p>
Data supporting "Slowest-first translation scheme: Structural asymmetry along protein sequences and co-translational folding"
<p>Contains data for a set of 16,200 non-redundant protein structures taken from the Protein Data Bank. Associated code can be found at https://github.com/jomimc/FoldAsymCode.</p>
RNAseq sequences of the study "Transactive response DNA-binding Protein (TARDBP/TDP-43) regulates early HIV-1 entry and infection" (1/2)
<p>Each pair of FASTQ files corresponds to a specific sample condition:</p> <table> <thead> <tr> <th scope="col">Condition</th> <th scope="col">Sample</th> <th scope="col">FASTQ name R1</th> <th scope="col">FASTQ name R2</th> </tr> </thead> <tbody> <tr> <td>Cneg</td> <td>RNASEQ-AVF1</td> <td>RNASEQ-AVF1_S1_R1_001.fastq.gz</td> <td>RNASEQ-AVF1_S1_R2_001.fastq.gz</td> </tr> <tr> <td>Flag-wt-TDP-43</td> <td>RNASEQ-AVF2</td> <td>RNASEQ-AVF2_S2_R1_001.fastq.gz</td> <td>RNASEQ-AVF2_S2_R2_001.fastq.gz</td> </tr> <tr> <td>Flag-NLS-mut-TDP-43</td> <td>RNASEQ-AVF3</td> <td>RNASEQ-AVF3_S3_R1_001.fastq.gz</td> <td>RNASEQ-AVF3_S3_R2_001.fastq.gz</td> </tr> <tr> <td>Cneg</td> <td>RNASEQ-AVF4</td> <td>RNASEQ-AVF4_S4_R1_001.fastq.gz</td> <td>RNASEQ-AVF4_S4_R2_001.fastq.gz</td> </tr> <tr> <td>Scramble</td> <td>RNASEQ-AVF5</td> <td>RNASEQ-AVF5_S5_R1_001.fastq.gz</td> <td>RNASEQ-AVF5_S5_R2_001.fastq.gz</td> </tr> <tr> <td>TDP-43 siRNA A</td> <td>RNASEQ-AVF6</td> <td>RNASEQ-AVF6_S6_R1_001.fastq.gz</td> <td>RNASEQ-AVF6_S6_R2_001.fastq.gz</td> </tr> <tr> <td>TDP-43 siRNA B</td> <td>RNASEQ-AVF7</td> <td>RNASEQ-AVF7_S7_R1_001.fastq.gz</td> <td>RNASEQ-AVF7_S7_R2_001.fastq.gz</td> </tr> <tr> <td>TDP-43 siRNA C</td> <td>RNASEQ-AVF8</td> <td>RNASEQ-AVF8_S8_R1_001.fastq.gz</td> <td>RNASEQ-AVF8_S8_R2_001.fastq.gz</td> </tr> </tbody> </table> <p> </p>
RNAseq sequences of the study "Transactive response DNA-binding Protein (TARDBP/TDP-43) regulates early HIV-1 entry and infection" (2/2)
<p>Each pair of FASTQ files corresponds to a specific sample condition:</p> <table> <thead> <tr> <th scope="col">Condition</th> <th scope="col">Sample</th> <th scope="col">FASTQ name R1</th> <th scope="col">FASTQ name R2</th> </tr> </thead> <tbody> <tr> <td>TDP-43 siRNA D</td> <td>RNASEQ-AVF9</td> <td>RNASEQ-AVF9_S1_R1_001.fastq.gz</td> <td>RNASEQ-AVF9_S1_R2_001.fastq.gz</td> </tr> <tr> <td>Cneg</td> <td>RNASEQ-AVF10</td> <td>RNASEQ-AVF10_S2_R1_001.fastq.gz</td> <td>RNASEQ-AVF10_S2_R2_001.fastq.gz</td> </tr> <tr> <td>Flag-wt-TDP-43</td> <td>RNASEQ-AVF11</td> <td>RNASEQ-AVF11_S3_R1_001.fastq.gz</td> <td>RNASEQ-AVF11_S3_R2_001.fastq.gz</td> </tr> <tr> <td>Flag-NLS-mut-TDP-43</td> <td>RNASEQ-AVF12</td> <td>RNASEQ-AVF12_S4_R1_001.fastq.gz</td> <td>RNASEQ-AVF12_S4_R2_001.fastq.gz</td> </tr> <tr> <td>Cneg</td> <td>RNASEQ-AVF13</td> <td>RNASEQ-AVF13_S5_R1_001.fastq.gz</td> <td>RNASEQ-AVF13_S5_R2_001.fastq.gz</td> </tr> <tr> <td>Scramble</td> <td>RNASEQ-AVF14</td> <td>RNASEQ-AVF14_S6_R1_001.fastq.gz</td> <td>RNASEQ-AVF14_S6_R2_001.fastq.gz</td> </tr> <tr> <td>Oligos B+C</td> <td>RNASEQ-AVF15</td> <td>RNASEQ-AVF15_S7_R1_001.fastq.gz</td> <td>RNASEQ-AVF15_S7_R2_001.fastq.gz</td> </tr> <tr> <td>Oligos A+B+C</td> <td>RNASEQ-AVF16</td> <td>RNASEQ-AVF16_S8_R1_001.fastq.gz</td> <td>RNASEQ-AVF16_S8_R2_001.fastq.gz</td> </tr> </tbody> </table> <p> </p>
Sequence-structure-function relationships in the microbial protein universe
<p>The Microbiome Immunity Project (MIP) dataset contains models predicted with both Rosetta and DMPFold (folder `dataset/`). It also contains DeepFRI function predictions for all models. </p> <p>The `metadata` folder contains additional data which may be useful for searching the MIP database (FASTA files, BLAST databases and useful scripts for structure/function search) as well as retrieving the sequence/structural annotations.</p> <p>The `intermediate_data` folder contains preprocessed output for reproducing many of the figures in our manuscript in conjunction with scripts and Juypter notebooks found in our git repository: https://github.com/microbiome-immunity-project/protein_universe .</p> <p>More information about the dataset and associated metadata is provided in the `README.md` file).</p> <p>We are also providing workflows to search the MIP database against a protein sequence or structure or function of interest (see `SEARCHING.md` for more details).</p>
Coronavirus Spike Protein Sequences from NCBI Virus
<p>For the study of taxonomic classification of coronaviruses across all genera (Alpha-, Beta-, Gamma-, and Deltacoronavirus), we can use spike protein sequences downloaded from the NCBI Virus website, \url{https://www.ncbi.nlm.nih.gov/labs/virus/vssi/}. The spike protein sequences used in this study were downloaded on November 21 and 27, 2021 using search terms such as ``spike'', ``S1 protein'', ``S2 protein'', and ``S protein'', in order to download as many spike protein sequences as possible.</p>
Training data for 'Functional annotation of protein sequences' tutorial (Galaxy Training Material)
<p>The data provided here are part of a Galaxy Training Network tutorial for functional annotation of protein sequences.</p>
Protein haplotype sequences obtained by ProHap from the 1000 Genomes Project data set
<p>Database of protein sequences obtained using ProHap (<a href="https://github.com/ProGenNo/ProHap">https://github.com/ProGenNo/ProHap</a>) on the data set of phased genotypes published by the 1000 Genomes Project, aligned with the GRCh38 genome build (<a href="https://www.internationalgenome.org/data-portal/data-collection/grch38">https://www.internationalgenome.org/data-portal/data-collection/grch38</a>). We used Ensembl v.110 for the mapping of coordinates between genes, exons, and transcripts. The complete configuration file for each ProHap run is attached to this repository.</p> <p>This data set contains six compressed directories, five representing the superpopulations included in the 1000 Genomes Project (<a href="https://catalog.coriell.org/1/NHGRI/Collections/1000-Genomes-Project-Collection/1000-Genomes-Project">https://catalog.coriell.org/1/NHGRI/Collections/1000-Genomes-Project-Collection/1000-Genomes-Project</a>), and one created using all the samples included in the 1000 Genomes data set:</p> <ul> <li>AFR - African</li> <li>AMR - American</li> <li>EUR - European</li> <li>SAS - South Asian</li> <li>EAS - East Asian</li> <li>ALL - all participants in the 1000 Genomes Project</li> </ul> <p>Each of the directories contains the following files:</p> <ul> <li>F1: The concatenated fasta file ready to be used with search engines, contains the following: <ul> <li>Protein haplotype sequences obtained by ProHap, using alleles with at least 1 % frequency within the selected population</li> <li>Reference proteome as per Ensembl v. 110</li> <li>Contaminant sequences from the cRAP project (<a href="https://www.thegpm.org/crap/">https://www.thegpm.org/crap/</a>)</li> <li>The file is provided in two formats - full and simplified. The simplified fasta contains only the artificial protein identifier and the matching gene name, and is optimised for compatibility with a wide range of tools. For annotation of peptides using the PeptideAnnotator, please provide the header (F1.2) in addition to the simplified fasta file. </li> </ul> </li> <li>F2: Additional information about the haplotype sequences, to be used for mapping identified peptides to the original haplotypes</li> <li>F3: Translations of haplotype cDNA sequences, before merging with the reference proteome</li> </ul> <p>For further description of the files, please refer to <a href="https://github.com/ProGenNo/ProHap/wiki/Output-files">https://github.com/ProGenNo/ProHap/wiki/Output-files</a>.</p> <p>For the usage of these databases with search engines, and downstream anaylsis of identified peptides, please refer to the project's wiki page: <a href="https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches">https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches</a>.</p> <p>When using these databases in your publication, please cite: Vašíček, J., Kuznetsova, K.G., Skiadopoulou, D. <em>et al.</em> ProHap enables human proteomic database generation accounting for population diversity. <em>Nat Methods</em> (2024). <a href="https://doi.org/10.1038/s41592-024-02506-0">https://doi.org/10.1038/s41592-024-02506-0</a></p>
Figure 2. PAM250 matrix for the encoded sequence-Secondary Structure Prediction of Protein using Resilient Back Propagation Learning Algorithm
<p>The PAM matrix (Dayhoff et al., 1978) describes the probability that original amino acid<br> will be replaced by another amino acid over a defined evolutionary interval. The unit of<br> evolutionary divergence is defined as the interval in which 1% of the amino acids have been<br> changed between two sequences. The work uses PAM250, which assumes the occurrence of 250-<br> point mutations per 100 amino acids.<br> So, for the given the protein sequence GIVEQCCASVCSLYQLENYCN, A will be replaced<br> by 1 -3 0 1 -3 -1 0 5 -2 -3 -4 -2 -3 -5 0 1 0 -7 -5 -1 as shown in Figure 2.</p>
Deep learning model for characterizing protein-RNA interactions from sequence at single-base resolution
<p> </p> <p><a href="https://zenodo.org/api/records/14021440/draft/files/encode_eclip.h5/content" target="_blank" rel="noopener noreferrer">encode_eclip.h5</a> - This file contains the training, validation, and test data for the Reformer model.</p> <p><a href="https://zenodo.org/api/records/14021440/draft/files/encode_eclip_bc.h5/content" target="_blank" rel="noopener noreferrer">encode_eclip_bc.h5</a> - This file contains the training, validation, and test data for the Reformer-BC model.</p> <p><a href="https://zenodo.org/api/records/14027315/draft/files/Reformer-code.zip/content" target="_blank" rel="noopener">Reformer-code.zip</a> - This file contains the training code of Reformer.</p>
AlphaFold2-Based Characterization of Apo and Holo Protein Structures and Conformational Ensembles Using Randomized Alanine Sequence Scanning Adaptation: Capturing Shared Signature Dynamics and Ligand-Induced Conformational Changes
<p>Proteins often exist in multiple conformational states, influenced by the binding of ligands or substrates. The study of these states, particularly the apo (unbound) and holo (ligand-bound) forms, is crucial for understanding protein function, dynamics, and interactions. In the current study, we use AlphaFold2 that combines<span> randomized</span> <span><span> </span>alanine<span> </span>sequence masking<span> </span>with shallow multiple sequence alignment<span> </span>subsampling to expand the conformational diversity of the predicted structural<span> </span>ensembles and<span> </span>capture conformational changes between apo and holo protein forms. Using several well-established datasets of<span> </span>structurally diverse apo-holo protein pairs, the proposed approach </span><span>enables<span> </span>robust predictions of apo and holo structures and conformational ensembles, while also displaying notably similar dynamics distributions. These observations are consistent with<span> </span>the view </span><span> </span>that the intrinsic dynamics of allosteric proteins is defined by the structural topology of the fold and favors conserved conformational motions driven by soft modes among orthologs. We also found<span> </span>a significant <span>correlation </span>between conformational flexibility and <span> </span>AlphaFold2 metric of statistical significance pLDDT for the apo-holo pairs in which ligand binding induced local moderate conformational changes. For apo-holo pairs exhibiting larger structural changes, this relationship<span> </span>becomes nonlinear, reflecting inability of AlphaFold2 confidence metrics to identify high energy functional conformations. Our findings support the notion that AlphaFold2 approaches can yield reasonable accuracy in predicting minor conformational adjustments between apo and holo states, especially for proteins with <span> </span>moderate localized changes upon ligand binding. However, for large, hinge-like domain movements, AF2 tends to predict the most stable domain orientation which is typically the apo form rather than the full range of functional conformations characteristic of the holo ensemble. These results indicate that modeling of multiple functional states of proteins may require more accurate detection of flexible region conformations and cannot solely rely on the pLDDT metric as the major determinant of the prediction accuracy in reproducing functional conformational ensembles.<span> </span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.