Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,744
datasets available to search
ShareScore release 0.9.0
Dataset results
1,744 results for “peptide”
Dataset for: Stereorandomization as a Method to Probe Peptide Bioactivity
<p>The upload contains additional primary data associated with the publication, including raw data in the original file format whenever possible.</p> <p>Data content: HRMS, HPLC-MS, CD, MD, TEM, Serum stability, Vesicle leakage assay, Cytotoxicity, Hemolysis.</p>
Species-specific proteotypic peptides for characterization of non-tuberculosis mycobacteria
<p>Non-tuberculous mycobacteria are opportunistic bacteria that closely resemble <i>Mycobacterium tuberculosis,</i> causing respiratory infections in humans. While genotyping through genome sequencing is found accurate in detecting mycobacterial species but struggles with distinguishing bacterial co-infection. These challenges lead to delayed therapeutic intervention, drug resistance, and disease complications. Lately, mass spectrometry-based (MALDI-TOF-MS) proteomics has routinely been used in diagnosing mycobacterial species in clinical samples. However, it suffers accurate species detection owing to extensive bacterial cultures and poor specificity in polymicrobial infections. In contrast, due to its sensitivity, LC-MS/MS based proteomics is widely employed for accurate bacterial proteome mapping. In this study, in-depth proteomics of 9 NTM species with proteome database searches of 26 datasets was used. In total 20 million peptide spectrum matches were identified aiding to ≥40% proteome coverage in 7 NTMs with highest in <i>M. abscessus</i>. Further, metaproteomic analysis and rescoring of peptides resulted in high-confidence species-specificity proteotypic peptides in <i>M. smegmatis</i> (2342), <i>M. vaccae</i> (960), <i>M. abscessus</i> (75), <i>M. avium</i> subsp. <i>paratuberculosis</i> (3) and <i>M. fortuitum</i> (1). Finally, database search results were converted to spectral library format for easier future usage in targeted proteomic workflows. This workflow in deriving species-specific peptides with high confidence can be extended in distinguishing closely related bacterial species for enhancing microbial diagnostics.</p>
Citation network data sets for 'Oxytocin – a social peptide? Deconstructing the evidence'
<p><strong>Introduction</strong></p> <p>This note describes the data sets used for all analyses contained in the manuscript 'Oxytocin - a social peptide?’<a href="#_ftn1">[1]</a> </p> <p><strong>Data Collection</strong></p> <p>The datasets described here were originally retrieved from Web of Science (WoS) Core Collection via the University of Edinburgh’s library subscription <a href="#_ftn2">[2]</a>. The aim of the original study for which these data were gathered was to survey peer-reviewed primary studies on oxytocin and social behaviour. To capture relevant papers, we used the following query:</p> <p><em>TI = (“oxytocin” OR “pitocin” OR “syntocinon”) AND TS = (“social*” OR “pro$social” OR “anti$social”)</em></p> <p>The final search was performed on the 13 September 2021. This returned a total of 2,747 records, of which 2,049 were classified by WoS as ‘articles’. Given our interest in primary studies <em>only</em> – articles reporting original data – we excluded all other document types. We further excluded all articles sub-classified as ‘book chapters’ or as ‘proceeding papers’ in order to limit our analysis to primary studies published in peer-reviewed academic journals. This reduced the set to 1,977 articles. All of these were published in the English language, and no further language refinements were unnecessary.</p> <p>All available metadata on these 1,977 articles was exported as plain text ‘flat’ format files in four batches, which we later merged together via Notepad++. Upon manually examination, we discovered examples of papers classified as ‘articles’ by WoS that were, in fact, reviews. To further filter our results, we searched all available PMIDs in PubMed (1,903 had associated PMIDs - ~96% of set). We then filtered results to identify all records classified as ‘review’, ‘systematic review’, or ‘meta-analysis’, identifying 75 records <a href="#_ftn3">[3]</a> (thus, ~4% of records classified by WoS were classified as reviews in PubMed). After examining a sample and agreeing with the PubMed classification, these were removed these from our dataset - leaving a total of 1,902 articles.</p> <p>From these data, we constructed two datasets via parsing out relevant reference data via the Sci2 Tool <a href="#_ftn4">[4]</a>. First, we constructed a ‘node-attribute-list’ by first linking unique reference strings (‘Cite Me As’ column in WoS data files) to unique identifiers, we then parsed into this dataset information on the identify of a paper, including the title of the article, all authors, journal publication, year of publication, total citations as recorded from WoS, and WoS accession number. Second, we constructed an ‘edge-list’ that records the citations from a <em>citing paper</em> in the ‘Source’ column and identifies the <em>cited paper</em> in the ‘Target’ column, using the unique identifies as described previously to link these data to the node-attribute-list.</p> <p>We then constructed a network in which papers are nodes, and citation links between nodes are directed edges between nodes. We used Gephi Version 0.9.2 <a href="#_ftn5">[5]</a> to manually clean these data by merging duplicate references that are caused by different reference formats or by referencing errors. To do this, we needed to retain both all retrieved records (1,902) as well as including <em>all</em> of their references to papers whether these were included in our original search or not. In total, this produced a network of 46,633 nodes (unique reference strings) and 112,520 edges (citation links). Thus, the average reference list size of these articles is ~59 references. The mean indegree (within network citations) is 2.4 (median is 1) for the entire network reflecting a great diversity in referencing choices among our 1,902 articles.</p> <p>After merging duplicates, we then restricted the network to include <em>only</em> articles fully retrieved (1,902), and retrained <em>only</em> those that were connected together by citations links in a large interconnected network (i.e. the largest component). In total, 1,892 (99.5%) of our initial set were connected together via citation links, meaning a total of ten papers were removed from the following analysis – and these were neither connected to the largest component, nor did they form connections with one another (i.e. these were ‘isolates’).</p> <p>This left us with a network of 1,892 nodes connected together by 26,019 edges. <strong><em>It is this network that is described by the ‘node-attribute-list’ and ‘edge-list’ provided here</em></strong>. This network has a mean in-degree of 13.76 (median in-degree of 4). By restricting our analysis in this way, we lose 44,741 unique references (96%) and 86,501 citations (77%) from the full network, but retain a set of articles tightly knitted together, all of which have been fully retrieved due to possessing certain terms related to oxytocin AND social behaviour in their title, abstract, or associated keywords.</p> <p>Before moving on, we calculated indegree for all nodes in this network – this counts the number of citations to a given paper from other papers within this network – and have included this in the <em>node-attribute-list</em>. We further clustered this network via modularity maximisation via the Leiden algorithm <a href="#_ftn6">[6]</a>. We set the algorithm to resolution 1, and allowed the algorithm to run over 100 iterations and 100 restarts. This gave <em>Q</em>=0.43 and identified seven clusters, which we describe in detail within the body of the paper. We have included cluster membership as an attribute in the node-attribute-list.</p> <p>For additional analysis, we also analysed the full reference list data to examine the most commonly cited references between 2016 and 2021 - the results of this are described in OTSOC_Cited_2016-2021.csv. This takes the reference lists of all retrieved papers within the network and examines their full reference lists (including references to other papers not contained within the network). These data were cleaned by matching DOIs and manual cleansing. </p> <p><strong>Data description</strong></p> <p>We include here two network datasets: (i) ‘OTSOC-node-attribute-list.csv’ consists of the attributes of 1,892 primary articles retrieved from WoS that include terms indicating a focus on oxytocin and social behaviour; (ii) ‘OTSOC-edge-list.csv’ records the citations between these papers. Together, these can be imported into a range of different software for network analysis; however, we have formatted these for ease of upload into Gephi 0.9.2. Finally, we include (iii) 'OTSOC_Cited_2016-2021' that lists all papers cited by >10 papers in the OTSOC network following any analysis of the bibliographies of retrieved papers. Below, we detail their contents:</p> <p><strong>1. ‘OTSOC-node-attribute-list.csv’</strong> is a comma-separate values file that contains all node attributes for the citation network (n=1,892) analysed in the paper. The columns refer to:</p> <p><em>Id</em>, the unique identifier</p> <p><em>Label</em>, the reference string of the paper to which the attributes in this row correspond. This is taken from the ‘Cite Me As’ column from the original WoS download. The reference string is in the following format: last name of first author, publication year, journal, volume, start page, and DOI (if available). </p> <p><em>Wos_id</em>, unique Web of Science (WoS) accession number. These can be used to query WoS to find further data on all papers via the ‘UT= ’ field tag.</p> <p><em>Title</em>, paper title.</p> <p><em>Authors</em>, all named authors.</p> <p><em>Journal, </em>journal of publication.</p> <p><em>Pub_year</em>, year of publication.</p> <p><em>Wos_citations</em>, total number of citations recorded by WoS Core Collection to a given paper as of 13 September 2021</p> <p><em>Indegree</em>, the number of within network citations to a given paper, calculated for the network shown in Figure 1 of the manuscript.</p> <p><em>Cluster</em>, provides the cluster membership number as discussed within the manuscript (Figure 1). This was established via modularity maximisation via the Leiden algorithm (Res 1; Q=0.43|7 clusters)</p> <p><strong>2. ‘OTSOC-edge -list.csv’</strong> is a comma-separated values file that contains all citation links between the 1,892 articles (n=26,019). The columns refer to:</p> <p><em>Source</em>, the unique identifier of the citing paper.</p> <p><em>Target, </em>the unique identifier of the cited paper.</p> <p><em>Type, </em>edges are ‘Directed’, and this column tells Gephi to regard all edges as such.</p> <p><em>Syr_date, </em>this contains the date of publication of the citing paper.</p> <p><em>Tyr_date, </em>this contains the date of publication of the cited paper.</p> <p><strong>3. 'OTSOC_Cited_2016-2021.csv'</strong> is a comma-separated values file that contain citations to all cited references that were cited by at least 10 of the retrieved papers within the OTSOC network published from 2016 onwards. The columns refer to: </p> <p><em>Reference, </em>the cited reference string extracted from the bibliographies of retrieved papers.</p> <p><em>Publication year, </em>the publication year of the cited reference.</p> <p><em>DOI</em>, the DOI of the cited reference. </p> <p><em>indegree_2016, </em>the total number of citations to a cited reference from papers published in 2016 and contained within the OTSOC network. </p> <p><em>indegree_2017, </em>the total number of citations to a cited reference from papers published in 2017 and contained within the OTSOC network. </p> <p><em>indegree_2018, </em>the total number of citations to a cited reference from papers published in 2018 and contained within the OTSOC network. </p> <p><em>indegree_2019, </em>the total number of citations to a cited reference from papers published in 2019 and contained within the OTSOC network. </p> <p><em>indegree_2020, </em>the total number of citations to a cited reference from papers published in 2020 and contained within the OTSOC network. </p> <p><em>indegree_2021, </em>the total number of citations to a cited reference from papers published in 2021 and contained within the OTSOC network. </p> <p><em>total indegree 2016-21</em>, the total number of citation to a cited reference from papers published between 2016-2021 and contained within the OTSOC network. </p> <p><strong>Software recommended for analysis</strong></p> <p>Gephi version 0.9.2 was used for the visualisations within the manuscript, and both files can be read and into Gephi without modification.</p> <p><strong>Notes</strong></p> <p><a href="#_ftnref1">[1]</a> Leng, G., Leng, R. I., Ludwig, M. (Submitted). Oxytocin – a social peptide? Deconstructing the evidence.</p> <p><a href="#_ftnref2">[2]</a> Edinburgh University’s subscription to Web of Science covers the following databases: (i) Science Citation Index Expanded, 1900-present; (ii) Social Sciences Citation Index, 1900-present; (iii) Arts & Humanities Citation Index, 1975-present; (iv) Conference Proceedings Citation Index- Science, 1990-present; (v) Conference Proceedings Citation Index- Social Science & Humanities, 1990-present; (vi) Book Citation Index– Science, 2005-present; (vii) Book Citation Index– Social Sciences & Humanities, 2005-present; (viii) Emerging Sources Citation Index, 2015-present.</p> <p><a href="#_ftnref3">[3]</a> For those interested, the following PMIDs were identified as ‘articles’ by WoS, but as ‘reviews’ by PubMed: ‘34502097’ ‘33400920’ ‘32060678’ ‘31925983’ ‘31734142’ ‘30496762’ ‘30253045’ ‘29660735’ ‘29518698’ ‘29065361’ ‘29048602’ ‘28867943’ ‘28586471’ ‘28301323’ ‘27974283’ ‘27626613’ ‘27603523’ ‘27603327’ ‘27513442’ ‘27273834’ ‘27071789’ ‘26940141’ ‘26932552’ ‘26895254’ ‘26869847’ ‘26788924’ ‘26581735’ ‘26548910’ ‘26317636’ ‘26121678’ ‘26094200’ ‘25997760’ ‘25631363’ ‘25526824’ ‘25446893’ ‘25153535’ ‘25092245’ ‘25086828’ ‘24946432’ ‘24637261’ ‘24588761’ ‘24508579’ ‘24486356’ ‘24462936’ ‘24239932’ ‘24239931’ ‘24231551’ ‘24216134’ ‘23955310’ ‘23856187’ ‘23686025’ ‘23589638’ ‘23575742’ ‘23469841’ ‘23055480’ ‘22981649’ ‘22406388’ ‘22373652’ ‘22141469’ ‘21960250’ ‘21881219’ ‘21802859’ ‘21714746’ ‘21618004’ ‘21150165’ ‘20435805’ ‘20173685’ ‘19840865’ ‘19546570’ ‘19309413’ ‘15288368’ ‘12359512’ ‘9401603’ ‘9213136’ ‘7630585’</p> <p><a href="#_ftnref4">[4]</a> Sci2 Team. (2009). Science of Science (Sci2) Tool. Indiana University and SciTech Strategies. Stable URL: <a href="https://sci2.cns.iu.edu">https://sci2.cns.iu.edu</a></p> <p><a href="#_ftnref5">[5]</a> Bastian, M., Heymann, S., & Jacomy, M. (2009). Gephi: an open source software for exploring and manipulating networks. International AAAI Conference on Weblogs and Social Media. Gephi is available via <a href="https://gephi.org/">https://gephi.org/</a></p> <p><a href="#_ftnref6">[6]</a> Traag, V. A., Waltman, L., & van Eck, N. J. (2019). From Louvain to Leiden: guaranteeing well-connected communities. Scientific reports, 9(1), 5233. <a href="https://doi.org/10.1038/s41598-019-41695-z">https://doi.org/10.1038/s41598-019-41695-z</a></p>
Dataset for: The antibacterial activity of peptide dendrimers and polymyxin B increases sharply above pH 7.4
<p>The upload contains additional primary data associated with the publication, including raw data in the original file format whenever possible.</p> <p>Data content: HRMS, HPLC-MS, pH titration, CD, MD, TEM</p> <p> </p>
Machine learning designs non-hemolytic antimicrobial peptides
<p>The upload contains additional primary data associated with the publication, including raw data in the original file format whenever possible.</p> <p>Data content: HRMS, HPLC-MS, CD, MD, TEM</p>
Research data supporting: "Self-assembly of cyclic peptide monolayers by hydrophobic supramolecular hinges"
<p>This repository contains the set of modelling data shown in the paper:<strong> "Self-assembly of cyclic peptide monolayers by hydrophobic supramolecular hinges"</strong>, published on Chemical Science (DOI: 10.1039/d3sc03930g)</p>
MD simulations of phosphorylated peptides (GGXGG)
<p>This repository contains MD simulations and associated analyses of a peptide series of short peptides including phosphorylated residues. It is one of six repositories that are associated to the following research article:</p> <blockquote> <p>Bickel,D., and Vranken,W. (2024) Effects of Phosphorylation on Protein Backbone Dynamics and Conformational Preferences. <em>J. Chem. Theory Comput</em>. https://doi.org/10.1021/acs.jctc.4c00206.</p> </blockquote> <p>The full list of the related repositories is given here:</p> <ol> <li>Pentapeptide simulations: <code>10.5281/zenodo.10517328</code></li> <li>Hexapeptides simulations: <code>10.5281/zenodo.10518872</code></li> <li>Heptapeptides simulations: <code>10.5281/zenodo.10518971</code></li> <li>Octapeptides simulations: <code>10.5281/zenodo.10518993</code></li> <li>Nonapeptides simulations: <code>10.5281/zenodo.10519033</code></li> </ol>
MD simulations of phosphorylated peptides (GGXGXGG)
<p>This repository contains MD simulations and associated analyses of a peptide series of short peptides including phosphorylated residues. It is one of five repositories that are associated to the following research article:</p> <blockquote> <p>Bickel,D., and Vranken,W. (2024) Effects of Phosphorylation on Protein Backbone Dynamics and Conformational Preferences. <em>J. Chem. Theory Comput</em>. https://doi.org/10.1021/acs.jctc.4c00206.</p> </blockquote> <p>The full list of the related repositories is given here:</p> <ol> <li>Pentapeptide simulations: <code>10.5281/zenodo.10517328</code></li> <li>Hexapeptides simulations: <code>10.5281/zenodo.10518872</code></li> <li>Heptapeptides simulations: <code>10.5281/zenodo.10518971</code></li> <li>Octapeptides simulations: <code>10.5281/zenodo.10518993</code></li> <li>Nonapeptides simulations: <code>10.5281/zenodo.10519033</code></li> </ol>
MD simulations of phosphorylated peptides (GGXXGG)
<p>This repository contains MD simulations and associated analyses of a peptide series of short peptides including phosphorylated residues. It is one of five repositories that are associated to the following research article:</p> <blockquote> <p>Bickel,D., and Vranken,W. (2024) Effects of Phosphorylation on Protein Backbone Dynamics and Conformational Preferences. <em>J. Chem. Theory Comput</em>. https://doi.org/10.1021/acs.jctc.4c00206.</p> </blockquote> <p>The full list of the related repositories is given here:</p> <ol> <li>Pentapeptide simulations: <code>10.5281/zenodo.10517328</code></li> <li>Hexapeptides simulations: <code>10.5281/zenodo.10518872</code></li> <li>Heptapeptides simulations: <code>10.5281/zenodo.10518971</code></li> <li>Octapeptides simulations: <code>10.5281/zenodo.10518993</code></li> <li>Nonapeptides simulations: <code>10.5281/zenodo.10519033</code></li> </ol>
MD simulations of phosphorylated peptides (GGXGGXGG)
<p>This repository contains MD simulations and associated analyses of a peptide series of short peptides including phosphorylated residues. It is one of five repositories that are associated to the following research article:</p> <blockquote> <p>Bickel,D., and Vranken,W. (2024) Effects of Phosphorylation on Protein Backbone Dynamics and Conformational Preferences. <em>J. Chem. Theory Comput</em>. https://doi.org/10.1021/acs.jctc.4c00206.</p> </blockquote> <p>The full list of the related repositories is given here:</p> <ol> <li>Pentapeptide simulations: <code>10.5281/zenodo.10517328</code></li> <li>Hexapeptides simulations: <code>10.5281/zenodo.10518872</code></li> <li>Heptapeptides simulations: <code>10.5281/zenodo.10518971</code></li> <li>Octapeptides simulations: <code>10.5281/zenodo.10518993</code></li> <li>Nonapeptides simulations: <code>10.5281/zenodo.10519033</code></li> </ol>
MD simulations of phosphorylated peptides (GGXGGGXGG)
<p>This repository contains MD simulations and associated analyses of a peptide series of short peptides including phosphorylated residues. It is one of five repositories that are associated to the following research article:</p> <blockquote> <p>Bickel,D., and Vranken,W. (2024) Effects of Phosphorylation on Protein Backbone Dynamics and Conformational Preferences. <em>J. Chem. Theory Comput</em>. https://doi.org/10.1021/acs.jctc.4c00206.</p> </blockquote> <p>The full list of the related repositories is given here:</p> <ol> <li>Pentapeptide simulations: <code>10.5281/zenodo.10517328</code></li> <li>Hexapeptides simulations: <code>10.5281/zenodo.10518872</code></li> <li>Heptapeptides simulations: <code>10.5281/zenodo.10518971</code></li> <li>Octapeptides simulations: <code>10.5281/zenodo.10518993</code></li> <li>Nonapeptides simulations: <code>10.5281/zenodo.10519033</code></li> </ol>
AMPSphere : the worldwide survey of prokaryotic antimicrobial peptides
<p><strong>AMPSphere v.2022-03: the worldwide survey of prokaryotic antimicrobial peptides</strong></p> <p> </p> <p><strong>INTRODUCTION</strong></p> <p>AMPSphere is a comprehensive catalog of antimicrobial peptides predicted using Macrel (DOI: <a href="https://peerj.com/articles/10555/">10.7717/peerj.10555</a>) from 63,410 public metagenomes, <a href="http://progenomes.embl.de/">ProGenomes v2.2 database</a> (82,400 high-quality microbial genomes), and c.a. 4k non-whitelisted microbial genomes from NCBI. Currently, AMPSphere is available as a web resource at <a href="https://ampsphere.big-data-biology.org/">https://ampsphere.big-data-biology.org/</a>.</p> <p> </p> <p><strong>GENERATION</strong></p> <p>Peptides were predicted using <a href="https://www.big-data-biology.org/software/macrel/">Macrel</a>. Singleton peptides were removed, except those with a direct hit to <a href="http://dramp.cpu-bioinfor.org/">DRAMP<br> database</a>. Redundant peptides were coded using a reduced alphabet and hierarchically clustered using CD-HIT (version 4.6) at 100%, 85%, and 75% of amino acid identity (and 90% of overlap of the shorter peptide). The obtained clusters were numbered by decreasing size (number of peptides). Each level of clustering was called a SPHERE. Redundant nucleotide sequences for the gene variants of different AMPs also were included in this version of AMPSphere.</p> <p> </p> <p><strong>STATISTICS</strong></p> <p>AMPSphere v.2022-03 contains 863,498 sequences (avg length: 36 amino acids, range 8-98). DRAMP database was used to find confirmed sequences with strict homology to reference. This approach showed that 2,488 peptides were previously confirmed in our dataset.</p> <p> </p> <p><strong>IDENTIFIERS</strong></p> <p>Peptides are named in the form <strong>>AMP10.XXX_XXX</strong> where <strong>XXX_XXX</strong> is a unique numerical identifier (starting at zero). Numbers were assigned in order of increasing number of copies. So that the lower the number, the greater number of copies of that peptide were present in the input data. Annotations were also provided as separated fields in the fasta file, containing their:</p> <p>- SPHERE families at level III (corresponding to hierarchically obtained clusters using 100-85-75% of identity with a minimum overlap of 90% of the shorter gene).</p> <p>Example of header:</p> <pre><code class="language-bash">>AMP10.000_000 | SPHERE-III.001_493 </code></pre> <p><br> <strong>VERSION DETAILS</strong></p> <p><strong>Version 2022-03</strong> includes:</p> <p>- quality assessment of documented AMPs,</p> <p>- metadata associated with the genes,</p> <p>- a better taxonomic identification of AMP sources using GTDB.</p> <p><em>WARNING: Due to a different procedure of AMP sorting, now some entries and families may have changed their accessions.</em></p> <p> </p> <p><strong>FILES WITHIN THIS VERSION</strong></p> <p><em>README.md</em><br> This file.</p> <p> </p> <p><em>AMPsphere_v.2022-03.fna.xz</em><br> Multi-fasta with AMPSphere gene sequences (nucleotide).</p> <p> </p> <p><em>AMPsphere_v.2022-03.faa.gz</em></p> <p>Multi-fasta with AMPSphere peptide sequences (amino acid).</p> <p> </p> <p><em>SPHERE_v.2022-03.levels_assessment.tsv.gz</em></p> <p>TSV table relating AMP name and the hierarchically obtained clusters per level. Columns:</p> <p>- AMP accession<br> - evaluation vs. representative<br> - SPHERE_fam level I<br> - SPHERE_fam level II<br> - SPHERE_fam level III</p> <p>Levels of each SPHERE family:</p> <p>I: contains clusters obtained with 100% of identity cut-off and 90% of overlap of the shorter sequence;</p> <p>II: contains clusters obtained with the unclustered sequences and the representatives from level I at 85% of identity and 90% of overlap of the sorter sequence;</p> <p>III: contains clusters obtained with the unclustered sequences and the representatives from level II at 75% of identity and 90% of overlap of the sorter sequence;</p> <p>`evaluation vs. representatives` shows the percent of identity the sequence has in an alignment against the cluster representative, and also the overlap in percent.</p> <p>Example:</p> <p> * -- This means: this sequence is a cluster representative.</p> <p>OR something like this:</p> <p> 77.50%,1:40:1:40 -- This means: alignment identity against the<br> representative of the cluster equals 77.5% and the<br> alignment start and end position for the query (1 and<br> 40, respectively), and target (1 and 40, respectively).</p> <p> </p> <p><em>AMPSphere_v.2022-03.quality_assessment.tsv.gz</em><br> TSV table containing the results of each quality test (by sequence). Columns:</p> <p>- AMP ID<br> - Antifam<br> - RNAcode<br> - Metaproteomes<br> - Metatranscriptomes<br> - Coordinates</p> <p>Results are one of 'Passed', 'Failed', or 'Not tested'.</p> <p><a href="https://www.ebi.ac.uk/research/bateman/software/antifam-tool-identify-spurious-proteins">Antifam </a>results show if the sequence matches ('Fail') or does not match ('Pass') to Antifams, a set of well-known spurious ORFs.</p> <p><a href="https://github.com/ViennaRNA/RNAcode">RNAcode</a> relies on gene diversity, therefore, families with less than 3 different gene sequences could not be tested and were marked as such.</p> <p>The direct match of 50% of our peptide to transcripts (in at least 2 different samples) or peptides from meta-omics studies sampled from different environments assigned the peptide as passing the metatranscriptomes and metaproteomes tests, respectively.</p> <p>Finally, the coordinates test check if the start of the small ORF happens with at least one stop codon upstream, this ensures that the gene is not a fragment from a larger protein.</p> <p> </p> <p><em>AMPSphere_v.2022-03.general_geneinfo.tsv.gz</em><br> TSV table relating AMP, gene name, the microbial source, sample, environment, and geographical location. Columns:</p> <p>- gmsc (gene code access) <br> - amp <br> - sample (biosample)<br> - source (microbial origin, GTDB taxonomy)<br> - specI (species cluster according to ProGenomes v.2 classification)<br> - is_metagenomic (False if comming from a high-quality microbial genome)<br> - geographic_location<br> - latitude<br> - longitude<br> - general_envo_name<br> - environment_material</p> <p> </p> <p><strong>CONTACT</strong></p> <p>You can contact us via our discussion group: <a href="https://groups.google.com/g/ampsphere-users">https://groups.google.com/g/ampsphere-users</a></p> <p>AMPsphere main developers:</p> <p>- <a href="mailto:celio@big-data-biology.com?subject=AMPSphere%20v.2022-03&body=Dear%20Celio%2C%20%0A%0ARegarding%20AMPSphere%20v.2022-03.">Célio Dias Santos Júnior</a><br> - <a href="mailto:yiqian@big-data-biology.org?subject=AMPSphere%20v2022-03&body=Dear%20Yiqian%2C%20%0A%0ARegarding%20AMPSphere%20v2022-03.%0A">Yiqian Duan</a><br> - <a href="mailto:hui@big-data-biology.org?subject=AMPSphere%20v2022-03&body=Dear%20Hui%2C%20%0A%0ARegarding%20AMPSphere%20v.2022-03.">Hui Chong</a><br> - <a href="mailto:luispedro@big-data-biology.com?subject=AMPSphere%20v.2022-03&body=Dear%20Luis%2C%20%0A%0ARegarding%20AMPSphere%20v.2022-03.">Luis Pedro Coelho</a></p> <p><br> <strong>COPYRIGHT NOTICE</strong></p> <p><em>AMPSphere v.2022-03 - the worldwide survey of prokaryotic antimicrobial peptides.</em></p> <p>This work is a joint effort of Big Data Biology group from the Institute of Science and Technology for Brain-Inspired Intelligence (ISTBI) - Fudan University, Shanghai, China, and the Structural and Computational Biology Unit<br> (Heidelberg) - European Molecular Biology Laboratory (EMBL).</p> <p>Copyright (C) 2019-2022 The Authors</p> <p> AMPSphere IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,<br> EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES<br> OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.<br> IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM,<br> DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR<br> OTHERWISE,ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE<br> USE OR OTHER DEALINGS IN THE SOFTWARE.</p> <p> This database is free; you can redistribute it and/or modify it<br> as you wish, under the terms of the CC BY 4.0 license.</p> <p> You are allowed to:</p> <p> Share — copy and redistribute the material in any medium or format</p> <p> Adapt — remix, transform, and build upon the material for any purpose,<br> even commercially.</p> <p> You may also obtain a copy of the CC BY 4.0 license here:<br> <br> https://creativecommons.org/licenses/by/4.0/</p> <p><br> <strong>REFERENCES CITED</strong></p> <p>- Macrel: Santos-Júnior CD, Pan S, Zhao X, Coelho LP. 2020. Macrel: antimicrobial peptide screening in genomes and metagenomes. PeerJ 8:e10555. https://doi.org/10.7717/peerj.10555</p> <p>- ProGenomes: Mende DR, Letunic I, Maistrenko OM et al. 2020. proGenomes2: an improved database for accurate and consistent habitat, taxonomic and functional annotations of prokaryotic genomes. Nucleic Acids Research 48(D1): D621–D625. https://doi.org/10.1093/nar/gkz1002</p> <p>- DRAMP: Kang X, Dong F, Shi C et al. 2019. DRAMP 2.0, an updated data repository of antimicrobial peptides. Sci Data 6, 148. https://doi.org/10.1038/s41597-019-0154-y</p> <p>- ANTIFAM: Eberhardt RY, Haft DH, Punta M, Martin M, O’Donovan C, BatemanA. 2012. AntiFam: a tool to help identify spurious ORFs in protein annotation. Database, Bas003.</p> <p>- RNAcode: Washietl S, Findeiss S, Müller SA, Kalkhof S, von Bergen M, Hofacker IL, Stadler PF, Goldman N. 2011. RNAcode: robust discrimination of coding and noncoding regions in comparative sequence data. RNA 17(4):578-94.<br> </p>
Trypanosoma Epitope Dataset: Valid Epitopes and Randomly Generated Peptides with Biochemical Metrics and AI-Generated Scores
<p>This dataset contains information about valid linear B-Cell epitopes from the Trypanosoma genus, as well as randomly generated peptides. It includes biochemical metrics generated by the EpiBuilder-1.0 tool and scores generated by the BepiPred-3.0 software. The data was originally collected from the IEDB and UniProtKB platforms and has been processed and enhanced with these informations for researchers interested in understanding the molecular interactions between Trypanosoma protozoans and the immune system.</p>
CREMP-CycPeptMPDB: Conformer-rotamer ensembles of macrocyclic peptides for machine learning with permeability annotations
<p>CREMP-CycPeptMPDB: A resource generated for the rapid development and evaluation of machine learning models for permeable macrocyclic peptides. CREMP-CycPeptMPDB contains 3,258 unique macrocyclic peptides and their high-quality structural ensembles generated using the Conformer-Rotamer Ensemble Sampling Tool (CREST). Altogether, this dataset contains nearly 8.7 million unique macrocycle geometries, each annotated with energies derived from semi-empirical tight-binding DFT calculations and with experimental membrane permeability measurements obtained from the <a href="http://cycpeptmpdb.com/" target="_blank" rel="noopener">CycPeptMPDB</a> database. We anticipate that this dataset will enable the development of machine learning models that can improve peptide design and optimization for novel therapeutics.</p> <p>This dataset complements the <a title="CREMP" href="../doi/10.5281/zenodo.7931444" target="_blank" rel="noopener">CREMP dataset</a>, which contains a larger selection of conformer ensembles for homodetic macrocyclic peptides.</p> <p>We provide the data in two available formats, either as Python pickle files, which provide quick read access with RDKit version 2022.09.5 or later, and as text-based SDF files with associated metadata in JSON format. Each file is named based on its amino acid sequence, with residues separated by periods, using standard one-letter codes with lowercase letters representing D-amino acids and "Me" prefixes representing <em>N</em>-methylated amino acids. The sequences are in no particular order, e.g., "C.R.E.M.P" and "R.E.M.P.C" correspond to the same peptide macrocycle. The filename extensions are ".pickle", ".sdf", and ".json".</p> <p>Each file in the “pickle” folder contains a Python dictionary with amino acid sequence, SMILES, CREST metadata, and a single RDKit molecule object containing all conformers. All files in the folder were compressed into a single “pickle.tar.gz” archive. In the “sdf_and_json” folder, each individual SDF file contains all conformers, each associated with its own JSON file that contains CREST metadata. Similarly, all are compressed into another single archive, “sdf_and_json.tar.bz2”. A single summary CSV file is also provided containing ”sequence”, “smiles”, “num_monomers”, “num_atoms”, “num_heavy_atoms”, along with the CREST metadata “totalconfs”, “uniqueconfs”, “lowestenergy”, “poplowestpct”, “temperature”, “ensembleenergy”, “ensembleentropy”, and “ensemblefreeenergy”. The number of unique conformers with different 3D structures is given by “uniqueconfs”, while “totalconfs” includes the number of rotamers in addition.</p> <p>The unzipped sizes of the archives are approximately 13 GB for "pickle.tar.gz" and 84 GB for "sdf_and_json.tar.bz2". If you encounter errors when trying to load the pickle files, please make sure your RDKit version is at least 2022.09.5. If that doesn't work, try other Python versions.</p>
vPro-MS peptide spectral library for the identification of human-pathogenic viruses by untargeted proteomics
<p>The viral proteomics workflow (vPro-MS) enables identification of human-pathogenic viruses from patient samples by untargeted proteomics. vPro-MS is based on an in-silico derived peptide library covering the human virome in <a href="https://www.uniprot.org/" rel="nofollow">UniProtKB</a> (331 viruses, 20,386 genomes, 121,977 peptides). vPro-MS is intended to identify human-pathogenic viruses from DiaNN (<a href="https://github.com/vdemichev/DiaNN">https://github.com/vdemichev/DiaNN</a>) outputs of either DIA or diaPASEF data. A scoring algorithm (vProID) assesses the confidence of virus identification and the results are finally summarized in a report table. </p> <p>The vPro Peptide Library folder contains 3 peptide FASTA files (Contaminants.fasta, Human.fasta, vPro.Virus.fasta), which were used to predict the spectral library (vPro-lib.predicted.speclib). Please note, that the additional commands “--cut” and “--duplicate-proteins” are needed to reprocess the prediction in DiaNN. This spectral library should be used to identify peptide sequences from samples of human origin using DiaNN. Furthermore, the folder contains the metadata file of the viral peptide sequences (vPro.Peptide.Library.txt) and a summary file of the virus taxonomy covered by the library (Taxonomy.Summary.txt). The metadata file is used by the vPro script to identify viruses from the DiaNN main report.</p>
Dataset from "Venomous Peptides: Molecular Origin of the Toxicity of Snake Venom PLA2‑like Peptides"
<p>Dataset from "Venomous Peptides: Molecular Origin of the Toxicity of Snake Venom PLA2‑like Peptides", containing the most relevant all-atom output trajectories and input files ran with GROMACS 2021:</p> <p>1) <strong>pure_membrane_systems.7z</strong> - pure bilayer systems (AA1, AA2, AA5), including equilibration, calcium insertion, and umbrella sampling simulations;</p> <p>2) <strong>single_peptide_systems.7z</strong> - single peptide-containing systems (AA3, AA4, AA6), including equilibration, calcium insertion, and umbrella sampling simulations;</p> <p>3) <strong>multiple_peptide_systems.7z</strong> - multiple peptide-containing systems (AA7, AA8), including equilibration, calcium insertion, and umbrella sampling simulations.</p> <p>We have included the input files (.mdp), system topology (.top and .itp), initial and final structure files (.gro), the index file (.ndx), and the portable binary run input files (.tpr). We have also included the output trajectories of systems AA3, AA4, AA6-8 in .xtc format, and spaced every 500 ps.</p> <p>System composition is given in Table Z1. More details can be found in the related publication.</p> <p><strong>Table Z1. Simulated systems' details, including name, composition (in number of lipid and peptide molecules), number of atoms composing the systems, simulation (sim.) time, and total umbrella sampling (US) time.</strong> </p> <table> <tbody> <tr> <td><strong>System</strong></td> <td><strong>POPC/POPS/Peptide</strong></td> <td><strong>no. atoms (a)</strong></td> <td><strong>sim. time (µs)</strong></td> <td><strong>US time (µs)</strong></td> </tr> <tr> <td><strong>AA1</strong></td> <td>128/0/0</td> <td>40,226</td> <td>0.3</td> <td>10.8</td> </tr> <tr> <td><strong>AA2</strong></td> <td>0/128/0</td> <td>39,458</td> <td>0.3</td> <td>10.8</td> </tr> <tr> <td><strong>AA3</strong></td> <td>128/0/1</td> <td>40,504</td> <td>0.5</td> <td>32.3</td> </tr> <tr> <td><strong>AA4</strong></td> <td>0/128/1</td> <td>39,724</td> <td>0.5</td> <td>32.3</td> </tr> <tr> <td><strong>AA5</strong></td> <td>96/32/0</td> <td>40,034</td> <td>1.0</td> <td>-</td> </tr> <tr> <td><strong>AA6</strong></td> <td>96/32/1</td> <td>40,300</td> <td>1.0</td> <td>-</td> </tr> <tr> <td><strong>AA7</strong></td> <td>96/32/5</td> <td>41,364</td> <td>1.0</td> <td>10.8</td> </tr> <tr> <td><strong>AA8</strong></td> <td>96/32/13</td> <td>55,128</td> <td>2.0</td> <td>10.8</td> </tr> </tbody> </table> <p>(a) for the US simulations, the total number of atoms was reduced in 1 because two sodium ions were substituted by a single calcium ion.</p>
OASis peptide database
<p>OASis human 9-mer peptide database, generated from 118 million human antibody sequences from the Observed Antibody Space database.</p> <p>Attached is a gzipped SQLite database containing two tables: "peptides" and "subjects".</p> <p>Links:</p> <ul> <li>BioPhi codebase and documentation: https://github.com/Merck/BioPhi</li> <li>Public BioPhi server: https://biophi.dichlab.org</li> <li>OAS Database: http://opig.stats.ox.ac.uk/webapps/oas/</li> </ul>
CREMP: Conformer-rotamer ensembles of macrocyclic peptides for machine learning
<p>CREMP: A resource generated for the rapid development and evaluation of machine learning models for macrocyclic peptides. CREMP contains 36,198 unique macrocyclic peptides and their high-quality structural ensembles generated using the Conformer-Rotamer Ensemble Sampling Tool (CREST). Altogether, this dataset contains nearly 31.3 million unique macrocycle geometries, each annotated with energies derived from semi-empirical tight-binding DFT calculations. We anticipate that this dataset will enable the development of machine learning models that can improve peptide design and optimization for novel therapeutics.</p> <p>We provide the data in two available formats, either as Python pickle files, which provide quick read access with RDKit version 2022.09.5 or later, and as text-based SDF files with associated metadata in JSON format. Each file is named based on its amino acid sequence, with residues separated by periods, using standard one-letter codes with lowercase letters representing D-amino acids and "Me" prefixes representing <em>N</em>-methylated amino acids. The sequences are in no particular order, e.g., "C.R.E.M.P" and "R.E.M.P.C" correspond to the same peptide macrocycle. The filename extensions are ".pickle", ".sdf", and ".json".</p> <p>Each file in the “pickle” folder contains a Python dictionary with amino acid sequence, SMILES, CREST metadata, and a single RDKit molecule object containing all conformers. All files in the folder were compressed into a single “pickle.tar.gz” archive. In the “sdf_and_json” folder, each individual SDF file contains all conformers, each associated with its own JSON file that contains CREST metadata. Similarly, all are compressed into another single archive, “sdf_and_json.tar.bz2”. A single summary CSV file is also provided containing ”sequence”, “smiles”, “num_monomers”, “num_atoms”, “num_heavy_atoms”, along with the CREST metadata “totalconfs”, “uniqueconfs”, “lowestenergy”, “poplowestpct”, “temperature”, “ensembleenergy”, “ensembleentropy”, and “ensemblefreeenergy”. The number of unique conformers with different 3D structures is given by “uniqueconfs”, while “totalconfs” includes the number of rotamers in addition.</p> <p>The unzipped sizes of the archives are approximately 32 GB for "pickle.tar.gz" and 210 GB for "sdf_and_json.tar.bz2". If you encounter errors when trying to load the pickle files, please make sure your RDKit version is at least 2022.09.5. If that doesn't work, try other Python versions.</p>
Molecular dynamics trajectories for "Reservoir-REMD facilitates kinetic rescue from metastable peptide conformations
<p>The molecular dynamics-generated ensemble dataset for cyclo-(cGHHQKLV), used in the manuscript "Reservoir-REMD facilitates kinetic rescue from metastable peptide conformations". The dataset consists of 14 + 6 =20 .dcd files, and one .pdb file for rendering.</p>
A versatile "Synthesis Tag" (SynTag) for the chemical synthesis of aggregating peptides and proteins
<p>Raw data for the project "A versatile "Synthesis Tag" (SynTag) for the chemical synthesis of aggregating peptides and proteins".</p><p>Manuscript and supporting information available on ChemRxiv: https://doi.org/10.26434/chemrxiv-2023-7mz2c-v2.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.