Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,744
datasets available to search
ShareScore release 0.9.0
Dataset results
1,744 results for “peptide”
Screen of A6,B7,1G4 TCRs against a library of HLA-A*02:01 MHC-I peptides from the human exome
<p>T2 cells expressing a library of off targets (derived from A6 and B7 binding motifs in Hausmann 1999) are co-cultured with A6, B7 or1G4 expressing T cells (from a non-A2 donor). Minigenes from surviving cells are amplified and sequenced.</p>
PepBench: Dataset for Protein-Binding Peptide Design
<p>Datasets and splits of protein-peptide complexes benchmark from <a href="https://arxiv.org/abs/2402.13555">PepGLAD</a>.</p> <p>V2 Updates:</p> <ol> <li>The size of ProtFrag augmentation dataset is 70498 instead of 70645. The latest index file has deleted duplicated entries.</li> <li>Clustering results for complexes in training/validation sets are uploaded in train_valid.</li> </ol>
Multilinear regression results for the retention of amino acids of hydrolyzed peptide fractions and their properties - 5%Prolastin
<p>The data set is a part of the PhD thesis of Nattawan Chorhirankul, Wageningen University, The Netherlands.</p>
Efficient improvement of the proliferation, differentiation, and anti-arthritic capacity of mesenchymal stem cells by simply culturing on the immobilized FGF2 derived peptide, 44-ERGVVSIKGV-53
<p>All data needed to evaluate the results presented in the paper: Soo Bin Lee, Ahmed Abdal Dayem, Sebastian Kmiecik, Kyung Min Lim, Dong Sik Seo, Hyeong-Taek Kim, Polash Kumar Biswas, Minjae Do, Deok-Ho Kim, Ssang-Goo Cho, Efficient improvement of the proliferation, differentiation, and anti-arthritic capacity of mesenchymal stem cells by simply culturing on the immobilized FGF2 derived peptide, 44-ERGVVSIKGV-53,<br>Journal of Advanced Research, Volume 62,<br>2024,<br>Pages 119-141,<br>ISSN 2090-1232,<br>https://doi.org/10.1016/j.jare.2023.09.041.</p>
Deep learning-driven fragment ion series classification enables highly precise and sensitive de novo peptide sequencing
<p>This Zenodo record contains the dataset and model weights for "Deep learning-driven fragment ion series classification enables highly precise and sensitive de novo peptide sequencing".</p> <p> </p> <p>This repository contains the following files:</p> <ul> <li> <p>For the human dataset by Wang et al.:</p> <ul> <li> <p>train_val_test_split.csv containing the mapping of the correct peptide by MaxQuant to either train, validation or test set</p> </li> <li> <p>psms_train_val_test.csv containing the mapping of correct PSMs (scan number, raw file and correct peptide by MaxQuant) to either train, validation or test set</p> </li> <li> <p>updated_spectralis_test_out.csv as before containing Spectralis-EA predictions and scores on test set, as well as initial peptides and scores by Casanovo and Novor and now containing also correct peptides by MaxQuant and Spectralis-scores on the combination of Casanovo and Novor sequences (column named spectralis_score_onlyRescoring)</p> </li> <li> <p>spectralis_test_out_heart_analysis.csv subset of 20220822_spectralis_test_out.csv containing only PSMs for the tissue heart with the computation of precision and recall values</p> </li> <li> <p>spectralis_test_out_pointnovo_deepnovo.csv containing predictions by DeepNovo and PointNovo with original scores and Spectralis-score, as well as correct peptides by MaxQuant</p> </li> </ul> </li> </ul> <p> </p> <ul> <li> <p>For the nine-species dataset by Tran et al.:</p> <ul> <li> <p>spectralis_ninespecies_out.csv containing spectrum identifiers, correct peptides by PEAKSDB, predicted peptides by the different de novo sequencing tools as well as original scores and Spectralis-scores for the different PSMs.</p> </li> </ul> </li> </ul>
Supplementary materials for "High-Throughput Discovery of Substrate Peptide Sequences for E3 Ubiquitin Ligases Using a cDNA Display Method."
<p>The next-generation sequencing (NGS) data of the 5th rounds' samples for LX9 library and p53deg library. The csv files contain DNA sequences read out, amino acid sequences and their read counts in descending order. </p>
3D ED dataset of co-crystal of GRGDS peptide with trifluoroacetic acid (TFA)
<p>EPU-D electron diffraction dataset of the GRGDS TFA co-crystal. The original diffraction data is in MRC format, with all metadata stored in the PETS2 pts2 file. JANA files for absolute structure determination and dynamical refinement are also included.</p>
Supplementary Material: Investigating Interaction Dynamics of an Enantioselective Peptide Catalyzed Acylation Reaction
<p>Supplementary material to the publication "Investigating Interaction Dynamics of an Enantioselective Peptide Catalyzed Acylation Reaction".</p> <p>The files contain the raw nuclear magnetic resonance (NMR) spectra and computed structures.</p> <p>A README-file with detailed information is provided.</p>
An integrative characterisation of proline cis and trans conformers in a disordered peptide
<p>These metadynamic molecular dynamics simulations, nuclear magnetic resonance spectroscopy (NMR), and small-angle X-ray scattering (SAXS) data support the findings of the manuscript entitled 'An integrative characterisation of proline cis and trans conformers in a disordered peptide' by Pettitt et al. DOI: <a href="https://doi.org/10.1016/j.bpj.2024.09.028" target="_blank" rel="noopener">10.1016/j.bpj.2024.09.028</a></p> <h1>Metadynamic molecular dynamics simulations</h1> <p>The <code>Metadynamic_simulations_Zenodo.tar.xz</code> directory contains data for an N-terminally acetylated disordered peptide. The peptide is the C-terminal region of open reading frame 6 (ORF6-CTR) from SARS-CoV-2. The data were produced by metadynamics simulations in 3 force fields: AMBER03WS (a03ws), AMBER99SB-disp (a99sb), and CHARMM36m (C36m). We used PLUMED version 2.7.1 and GROMCS 2021.2. All the data and PLUMED input files required to reproduce the simulation results are available on PLUMED-NEST. This data should be used with the code provided on GitHub at <a href="https://github.com/hansenlab-ucl/orf6-ctr_cis_trans_conformers/tree/master/Metadynamic_simulations"><code>https://github.com/hansenlab-ucl/orf6-ctr_cis_trans_conformers/tree/master/Metadynamic_simulations</code></a></p> <p>Once downloaded, this directory should be extracted using the following command:</p> <p>tar -xzvf Metadynamics_simulations_Zenodo.tar.xz</p> <p>The directory should be saved with the name <code>Metadynamic_simulations_Zenodo</code> and placed in the same directory as the GitHub <code>README_metadynamic_simulations.md</code> file.</p> <h2>This dataset contains:</h2> <ul> <li>{system}.pdb - atomic coordinate files for the NAc-ORF6-CTR in the a03ws, a99sb, and C36m force fields.</li> <li>{system}_traj.trr - single concatenanated trajectory files for the a03ws run 1, a03ws run 2, a99sb, and C36m simulations.</li> <li>{system}_weights_corr.dat - weights for each frame in _traj.trr for the a03ws run 1, a03ws run 2, a99sb and C36m simulations. Here, the weights of frames in which the peptide interacts with its periodic image have been set to zero. The cutoff was 0.95 nm for C36m, and 1.2 nm for the other force fields.</li> <li>{system}_weights_saxs_bme.dat - SAXS Bayesian/Maximum Entropy (BME) reweighted weights for each frame in {system}_traj.trr for the a03ws run 1, a03ws run 2, and C36m simulations. Here, the weights of frames in which the peptide interacts with its periodic image have been set to zero. The cutoff was 0.95 nm for C36m, and 1.2 nm for the other force fields.</li> <li>top_frames_{system}.npy - numpy array of frames index for the a03ws run 1, a03ws run 2, a99sb, and C36m simulations. Frames were selected based on weights_corr.dat.</li> <li>top_frames_{system}.trr - trajectory of frames for the a03ws run 1, a03ws run 2, a99sb, and C36m simulations. Frames were selected based on weights_corr.dat</li> <li>top_frames_{system}_rw_cis.npy - numpy array of cis-P57 frames index for the a03ws run 1, a03ws run 2, and C36m simulations. Frames were selected based on _weights_saxs_bme.dat</li> <li>top_frames_{system}_rw_trans.npy - numpy array of trans-P57 frames index for the a03ws run 1, a03ws run 2, and C36m simulations. Frames were selected based on _weights_saxs_bme.dat</li> <li>top_frames_{system}_rw_cis.trr - trajectory of cis-P57 frames for the a03ws run 1, a03ws run 2, and C36m simulations. Frames were selected based on _weights_saxs_bme.dat</li> <li>top_frames_{system}_rw_trans.trr - trajectory of trans-P57 frames for the a03ws run 1, a03ws run 2, and C36m simulations. Frames were selected based on _weights_saxs_bme.dat</li> <li>CS_COLVAR_{system} - experimental and CamShift predicted chemical shifts for all four systems (a03ws run1, run2, a99sb, and C36m) as well as the cis and trans sub-ensembles for a03ws run1, a03ws run2, and C36m</li> </ul> <h1>Nuclear magnetic resonance spectroscopy (NMR) data</h1> <p>The <code>NMR_Zenodo.tar.xz</code> directory contains data for the backbone and sidechain assignment experiments, backbone 15N relaxation rates experiments, and 15N diffusion experiments (.ft2 and .ft3 files). Unless specified otherwise, NMR spectra were collected on uniformly 15N-labelled and 13C, 15N-labelled ORF6-CTR at concentrations of 300 micromolar and on unlabelled NAc-ORF6-CTR at concentations of 400 micromolar. All peptides were prepared in 25 mM HEPES buffer at pH 6.9, 150 mM NaCl, containing 5% D2O, 1 mM sodium azide, and 1 mM EDTA. NMR data were recorded at 288.15 K (15 degrees celsius). Spectra were measured at a static magnetic field strength of 14. T (600 MHz), unless specified otherwise.</p> <p>Once downloaded, this directory should be extracted using the following command:</p> <p><code>tar -xzvf NMR_Zenodo.tar.xz</code></p> <p>Note: this data is not required to run the NMR analysis on GitHub at <a href="https://github.com/hansenlab-ucl/orf6-ctr_cis_trans_conformers/tree/master/NMR"><code>https://github.com/hansenlab-ucl/orf6-ctr_cis_trans_conformers/tree/master/NMR</code></a> but it can be used for you to repeat your own analysis.</p> <p>NMR chemical shifts have been deposited in the Biological Magnetic Resonance Data Bank (BMRB; <a href="../records/13748215/preview/www.bmrb.wisc.edu"><code>www.bmrb.wisc.edu</code></a>) under the following accession codes: <strong>52459</strong> for the ORF6CTR cis-P57 and trans-P57 configurations, and <strong>52460</strong> for the unlabelled NAc-ORF6CTR.</p> <h2>This dataset contains:</h2> <ul> <li>1H-1H-tocsy-nac-orf6-ctr.ft2 - 2D 1H-1H tocsy on the NAc-ORF6-CTR. Standard dipsi2esgpph Bruker pulse sequence, with spectral widths of 14 ppm (direct dimension) and 12 ppm (indirect dimension), and 1,024 x 512 complex points, respectively. 32 scans were recorded, with an inter-scan delay of 1.5 s.</li> <li>1H-1H-tocsy.ft2 - 2D 1H-1H total correlation spectroscopy (TOCSY) on the ORF6-CTR. Same details as above.</li> <li>1H-13C-hsqc-nac-orf6-ctr.ft2 - 2D 1H-13C heteronuclear single quantum coherence (HSQC) spectra on the NAc-ORF6-CTR. Standard hsqcgpph Bruker sequence was employed, with 192 scans, spectral widths of 14 ppm (1H) and 80 ppm for (13C), and 512 x 180 complex points, respectively. The recycle delay was set to 1.5 s, with carriers set to the position of the water peak (1H) and 35 ppm relative to TMS (13C).</li> <li>1H-15N.ft2 - 2D 1H-15N HSQC spectra on the 15N-labelled ORF6-CTR. Spectra were acquired using the hsqcetf3gpsi2 Bruker pulse sequence. Spectral widths were set to 16 ppm (1H) and 20.5 ppm (15N), with 1,536 x 128 complex points acquired, respectively. Carriers were set to the position of the water peak (1H) and 121.5 ppm (15N), with 4 scans and a recycle delay of 1 s.</li> <li>1H-15N-hsqc-310K.ft2 - 2D 1H-15N HSQC spectra on the 15N-labelled ORF6-CTR at 310.15 K (37 degrees celsius). Same details as above.</li> <li>15N-edited-dosy-950MHz.ft2 - Pseudo-3D diffusion pulsed field-gradient spin echo (PFGSE) 1H-15N-Diffusion Ordered Spectroscopy-HSQC (1H-15N-DOSY-HSQC) experiment with a bipolar gradient at a static magnetic field strength of 22.3 T (950 MHz) on the 15N-labelled ORF6-CTR. Six gradient experiments were acquired for each data set, with the gradient strengths augmented linearly through the acquisition from 0.9 to 39.5 G/cm and all other delays and pulses held constant. Gradient pulses (δ) were applied for 3 ms and a diffusion delay (Δ) of 200 ms. 80 scans were acquired per gradient experiment with 1,536 x 176 complex points (1H, 15N), using the same spectral widths and carriers as used before for the 2D 1H-15N HSQC experiment.</li> <li>15N-edited-noesy.ft3 - 15N-edited nuclear Overhauser effect spectroscopy HSQC (15N-NOESY-HSQC) on the ORF6-CTR. Spectra were acquired using the noesyhsqcfpf3gpsi3d Bruker pulse sequence. Spectra were recorded with 8 scans, 2,048 x 24 x 96 complex points (1H, 15N, 1H), and a spectral width of 16 x 20.5 x 16 ppm, respectively. Carriers were set to the water peak (1H) and 121.5 ppm (15N). A recycle delay of 1.5 s and a mixing time of 0.25 s were used</li> <li>15N-edited-tocsy.ft3 - 15N-TOCSY-HSQC on the ORF6-CTR. Spectra were acquired using a pulse sequence based on the original Bruker dipsihsqcf3gpsi3d with a flip-flop spectroscopy (FLOPSY)-16 scheme. Spectra were acquired with 16 scans, 1,024 x 32 x 96 complex points (1H, 15N, 1H), and a spectral width of 16 x 20.5 x 16 ppm, respectively. Carriers were set to the water peak (1H) and 121.5 ppm (15N), with a recycle delay of 1 s. A TOCSY mixing time of 0.1 s was used and an 8 kHz TOCSY spin lock was applied.</li> <li>ccconh.ft3 - CC(CO)NH spectra were acquired on the 13C, 15N-labelled ORF6-CTR. Spectra were acquired using a pulse sequence based on the original Bruker ccconhgp3d.2 with a FLOPSY-16 scheme. Spectra were recorded with 16 scans, 1,024 x 30 x 84 complex points (1H, 15N, 13C), and a spectral width of 14 x 20.5 x 71 ppm, respectively. Carriers were set to the water peak (1H), 121.5 ppm (15N) and 43 ppm (13C). A recycle delay of 1.5 s, along with a TOCSY mixing time of 18 ms were used.</li> <li>con.ft2 - 2D CON spectra were acquired on the 13C, 15N-labelled ORF6-CTR. Spectra were acquired using the c_con_iasq Bruker pulse sequence. Spectral widths were set to 40 ppm in both dimensions, with carriers set to 173 ppm relative to TMS (13C) and 121.5 ppm (15N), and 512 x 128 complex points, respectively. 32 scans and a recycle delay of 2 s were applied.</li> <li>hncacb.ft3 - The 3D HNCACB spectra were acquired on the 13C, 15N-labelled ORF6-CTR using the hncacbgp3d Bruker pulse sequence. Spectra were recorded with 16 scans, 1,024 x 30 x 72 complex points (1H, 15N, 13C), and a spectral width of 14 x 20.5 x 60.2 ppm, respectively. Carriers were set to the water peak (1H), 121.5 ppm (15N), and 43 ppm (13C), with a recycle delay of 1 s.</li> <li>hncaco.ft3 - The 3D HN(CA)CO spectra were acquired using the hncacogpwg3d Bruker pulse sequence on the 13C,15N-labelled ORF6-CTR. Spectra were recorded with 16 scans, 1,024 x 29 x 80 complex points (1H, 15N, 13C), and a spectral width of 14 x 20.5 x 7 ppm, respectively. Carriers were set to the position of the water peak (1H), 121.5 ppm (15N) and 173 ppm relative to TMS (13C), with a recycle delay of 1 s.</li> <li>hnco.ft3 - The 3D HNCO spectra were acquired using the hncogp3d Bruker pulse sequence on the 13C, 15N-labelled ORF6-CTR. The same details were used as for the hncaco.ft3.</li> <li>noe.ft2 - The {1H}-15N hetNOEs were recorded using a pseudo-3D experiment, with and without proton saturation on the 15N-labelled ORF6-CTR. Amide proton magnetisation saturation was accomplished by using a 5 s train of high-power 120° pulses applied at 5 ms intervals. The reference and saturated spectra were alternately recorded. To ensure complete recovery of the initial magnetisation at the start of each increment of the reference experiment, a long recycle delay of 15 s was applied. The spectral widths were 16 ppm (1H) and 20.5 ppm (15N), and 2,048 x 160 (1H, 15N) complex points were recorded at a static magnetic field strength of 14.1 T.</li> <li>noe-800MHz.ft2 - The same details as for noe.ft2, except it was measured at a static magnetic field strength of 18.8 T (800 MHz).</li> <li>r1.ft2 - Rates were measured using established proton-detected pulse sequences based on a gradient-selected, sensitivity-enhanced, refocused 15N sequences on the 15N-labelled ORF6-CTR. Spectra were recorded with 2,048 x 160 complex points and spectral widths as used before in the 2D 1H-15N HSQC experiment. Gradient pulses were used to suppress the water signal and a long recycle delay of 3 s was employed. N-H cross-correlated relaxation pathways were suppressed by hard 180˚ pulses every 20 ms during the relaxation delay. The R1 N-H planes were recorded with 8 relaxation delays ranging from 20 ms to 700 ms.</li> <li>r1-800MHz.ft2 - Same details as for r1.ft2, except the rates were measured at a static magnnetic field strength of 18.8 T (800 MHz).</li> <li>r1rho.ft2 - Similar details to r1.ft2, except cross-correlated relaxation was suppressed by hard 180˚ pulses during the 15N spin-lock. Magnetisation was explicitly aligned with the spin-lock field. N-H planes were recorded using 8 relaxation delays ranging from 2 ms to 140 ms, with a 15N spin-lock field strength of 2 kHz.</li> <li>r1rho_800MHz.ft2 - Same as r1rho.ft2, except the rates were measured at a static magenetic field strength of 18.8 T (800 MHz).</li> <li>rdd.ft2 - Exchange-free 15N transverse relaxation (Rdd) rates were measured using pulse schemes for the four 1H-15N relaxation rates R1ρ(2HzN’z), R1ρ(2H’zNz), R1ρ2(2H’zN’z), and R1(2HzNz) on the 15N-labelled ORF6-CTR. A total of 8 relaxation delays between 2 ms and 26 ms were used for all experiments (same delays used for each rate measurement). A 10 kHz 1H spin lock and a 2 kHz 15N spin lock were applied. Spectral widths were the same as those used before in {1H}-15N hetNOE experiments, with 1,536 x 640 (1H, 15N) complex points.</li> </ul> <h1>Small-angle X-ray scattering (SAXS) data</h1> <p>Experimental and predicted SAXS data for the NAc-ORF6-CTR. Once downloaded, this directory should be extracted using the following command:</p> <p><code>tar -xzvf SAXS_calculations_Zenodo.tar.xz</code></p> <h2>This dataset contains:</h2> <ul> <li>saxs_exp_nac_orf6_ctr.dat - experimental SAXS data collected on Instrument B21 at Diamond Light Source (Didcot, UK). Measurements were recorded on 800 micromolar unlabelled ORF6-CTR at 310.15 K (37 degrees celsius). Data sets of 26 frames with a frame exposure time of 1 s each were acquired.</li> <li>{system}_saxs_calc_nac_orf6_ctr.dat - calculated SAXS data using PEPSI-SAXS for the a03ws run 1, a03ws run 2, and C36m metadynamics simulations. There is one column per experimental average (total = 2570) and a row for each frame in the trajectory, which depends on the number of frames in the a03ws run 1, a03ws run 2, and C36m simulations.</li> </ul> <h2>Jupyter notebooks</h2> <p>All Jupyter notebooks can be accessed from GitHub. Open Jupyter notebooks using <code>jupyter lab</code> in the analysis environment.</p>
Data for 'NovoRank: Refinement for De Novo Peptide Sequencing Based on Spectral Clustering and Deep Learning'
<p>The uploaded files include two archives for <a href="https://pubs.acs.org/doi/10.1021/acs.jproteome.4c00300" target="_blank" rel="noopener">NovoRank: Refinement for De Novo Peptide Sequencing Based on Spectral Clustering and Deep Learning</a>. The '<em>mgf_data</em>' archive contains all MGF files used in the study, while the '<em>sample_data</em>' archive includes sequencing data, clustering data generated using <code>MSCluster</code>, and XCorr calculation data computed with <code>CometX</code>, all of which were used in the research.</p>
Dataset - Exploring Ligand Binding to Calcitonin Gene-Related Peptide Receptors
<p>Spervised molecular dynamics simulations (SuMD) of:</p> <p>- CGRP binding to CGRPR TMD</p> <p>- telcagepant binding to CGRPR ECD</p> <p>- telcagepant unbinding from CGRPR ECD</p> <p>(Water molecules, POPC and ions removed)</p>
Code and data sets for "DeepLC can predict retention times for peptides that carry as-yet unseen modifications"
<p>Code used to prepare the data sets, calibrate retention times, generate DeepLC models, make predictions, and generate the figures. See README.md for more information on how to use these files and reproduce the results reported in the manuscript titled "DeepLC can predict retention times for peptides that carry as-yet unseen modifications".</p>
Protein Identification by Nanopore Peptide Profiling
<p>This dataset belongs to “Protein Identification by Nanopore Peptide Profiling” and describes the raw data and analysis of tryptic digested peptides translocating through a mutant Fragaceatoxin C nanopore. A jupyter notebook describing the analysis and structure is added to this dataset.</p> <p> </p> <p><strong>Data description:</strong></p> <p><strong>Protein Identification by Nanopore Peptide Profiling.ipynb</strong></p> <p> Jupyter notebook contained data analysis of data contained in data_0.zip and data_1.zip (Python 3.7)</p> <p><strong>python_scripts.zip</strong></p> <p> Supplementary scripts belonging to “Protein Identification by Nanopore Peptide Profiling.ipynb”. See explanation of custom classes in the jupyter notebook.</p> <p><strong>data_0.zip</strong> - Folder containing raw electrophysiology data and result after analysis with “Protein Identification by Nanopore Peptide Profiling.ipynb”, with each folder containing the following:</p> <p> Alpha casein: Tryptic digest of alpha casein</p> <p> Beta casein: Tryptic digest of beta casein</p> <p> BSA: Tryptic digest of bovine serum albumin</p> <p> Control: Tryptic digest of water (no protein, control measurement)</p> <p> Cytochrome c: Tryptic digest of cytochrome c</p> <p> DHFR_His6: Tryptic digest of dihydropholate reductase (His6 tagged)</p> <p> EFP: Tryptic digest of elongation factor P</p> <p> HMW1Act: Tryptic digest of high molecular weight adhesin protein</p> <p><strong>data_1.zip</strong> - Folder containing raw electrophysiology data, comma-separated MS peptide masses, and result after analysis with “Protein Identification by Nanopore Peptide Profiling.ipynb”, with each folder containing the following:</p> <p> Lysozyme: Tryptic digest of lysozyme </p> <p> PAN: Tryptic digest of proteasome-activating nucleotidase</p> <p> TbpA_Y27A: Tryptic digest of periplasmic binding protein</p> <p> Trypsin: Tryptic digest of bovine trypsin</p> <p> Mass_spec: csv files containing measured ESI-MS peptides</p> <p> Lysozyme synthetic peptides: Synthetic peptides:</p> <p> Lys1: TPGSR</p> <p> Lys2alk: C(+57.02)ELAAAMK</p> <p> Lys3: HGLDNYR</p> <p> Lys4alk: WWC(+57.02)NDGR</p> <p> Lys5: GTDVQAWIR</p> <p> Lys6alk: GYSLGNWVC(+57.02)AAK</p> <p> Lys7: FESNFNTQATNR</p> <p>The structure of the data files is registered data_1.zip in <strong>'index.csv'</strong> (digested proteins) and <strong>'index</strong><strong>_peptides</strong><strong>.csv'</strong> (synthetic peptides) contained in the data folder. In this file, we describe the protein that was measured as well as the folder location and the expected baseline / standard deviation.<br> <br> <strong>Structure of <em>./data/index.csv</em></strong></p> <p><strong>Protein (string) | Folder (string) | Baseline (pA) (float) | Baseline Error (pA) (float)</strong></p> <p> </p> <p>In each <strong>Folder</strong>, there is another <strong>'index.csv'</strong>, explaining which files are with protein and which are without (blank).<br> <br> <strong>Structure of <em>./data/[protein]/[repeat]/index.csv</em></strong></p> <p><strong>blank (boolean) | fname (string)</strong></p> <p> </p> <p>Each folder in data_0.zip and data_1.zip contains a folder for each measure protein, which contains a folder for each repeat. The repeats contain raw axon binary files (.abf), each file contains measurement conditions as follows:</p> <p> [Date of measurement]_[Pore type]_[Buffer conditions]_[added analyte(s)]_[operator initials]</p> <p> <em>e.g</em>: 20200312_1M_KCl_50mM_Citricacid_50mM_BTP_pH_38_FraC_G13F_neg70mV_20ul_CytC_TrypsinGold_FL_0000</p> <p> Measured on 12-03-2020, in 1M KCl buffered with Citricacid (50 mM) adjusted using bis-tris-propane to pH 3.8, using Fragaceatoxin C mutant G13F at a negatively applied potential of 70 mV. 20 µL cytochrome c was added to the cis compartment.</p> <p>The total volume of the container used for all electrophysiology experiments was 400 µL, all samples were prepared at a 1 g/L concentration. A prefix “perf” before analyte description indicates that the chamber was flushed with approximately 2 mL fresh buffer prior to analysis. The buffer condition "BTP" means bis-tris-propane, which is used to titrate to the exact pH of 3.8.</p> <p>Each analysed folder contains<strong> results.pkl</strong> file, containing the analysis result as provided by “Protein Identification by Nanopore Peptide Profiling.ipynb” - see the jupyter notebook</p> <p>Each analysed folder contains <strong>results_analysis.xlsx</strong>, which contains sheets with excluded currents, standard deviations, dwell time and beta value for the pore without analyte added “Blank” and results from the analyte added in “Results”. Parameters used for fitting are contained in “Parameters”. The “Histograms” tab shows the raw data of the excluded current spectra.</p> <p> </p> <p><strong>mass_spec_peaks.zip</strong> – Folder containing mass spectrometry files as analysed by PEAKS Studio</p> <p>The folder contains an subfolder for each protein measured using electrospray ionisation mass spectrometry (ESI-MS).</p> <p> acasein: alpha casein protein</p> <p> b_casein: beta casein protein</p> <p> BSA: bovine serum albumin</p> <p> CytC: cytochrome C digested</p> <p> DHFR: dihydropholate reductase</p> <p> HMW1_Act: high molecular weight adhesin protein</p> <p> PAN: proteasome-activating nucleotidase</p> <p> ThBP: periplasmic thiamine binding protein</p> <p> Trypsin: bovine trypsin</p>
SGLT2-Inhibition reverts urinary peptide changes associated with severe COVID-19: an in-silico proof-of-principle of proteomics-based drug repurposing
<p>Severe COVID-19 is reflected by significant changes in urine peptides. Based on this observation, a clinical test predicting COVID-19 severity, CoV50, was developed and registered as in vitro diagnostic in Germany. We have hypothesized that molecular changes displayed by CoV50, likely reflective of endothelial damage, may be reversed by specific drugs. Such an impact by a drug could indicate potential benefits in the context of COVID-19. To test this hypothesis, urinary peptide data from patients without COVID-19 prior to and after drug treatment were collected from the human urinary proteome database. The drugs chosen were selected based on availability of sufficient number of participants in the dataset (n>20) and potential value of drug therapies in the treatment of COVID-19 based on reports in the literature. In these participants without COVID-19, spironolactone did not demonstrate a significant impact on CoV50 scoring. Empagliflozin treatment resulted in a significant change in CoV50 scoring, indicative of a potential therapeutic benefit. The study serves as a proof-of-principle for a drug repurposing approach based on human urinary peptide signatures. The results support the initiation of a randomised control trial testing a potential positive effect of empagliflozin for severe COVID-19, possibly via endothelial protective mechanisms.</p>
Supporting data for the manuscript "Nerpa: a tool for discovering biosynthetic gene clusters of nonribosomal peptides"
<p>Preprocessed structures of nonribosomal peptides [NRPs] and genomic sequences (reference and representative genomes, biosynthetic gene clusters [BGCs]) used in the benchmark experiments in the Nerpa paper.</p> <p><strong>Files description</strong></p> <ul> <li><strong>mibig_nrp_bacteria_preprocessed.tar.gz</strong> contains the preprocessed dataset of 194 bacterial NRP BGCs from the MIBiG database.</li> <li><strong>mibig_nrp_bacteria_summary.tsv</strong> contains metadata for the MIBiG-NRP dataset.</li> <li><strong>bacterial_ref_and_repr_genomes_20210604_preprocessed.tar.gz</strong> contains the preprocessed dataset of 13,399 reference and representative bacterial genomes from the NCBI RefSeq database (retrieved on 2021/06/04).</li> <li><strong>bacterial_ref_and_repr_genomes_20210604_summary.txt</strong> contains metadata for the RefSeq dataset.</li> <li><strong>pnrpdb_preprocessed.info</strong> contains the Nerpa-preprocessed pNRPdb database, a database of 8,368 known and putative NRP structures.</li> <li><strong>pnrpdb_summary.tsv</strong> contains the pNRPdb database metadata.<br> </li> </ul>
MALDI imaging of mouse kidney peptides - test dataset
<p>This imzML test file is concise but meaningful as a training data set in the Galaxy training network (https://galaxyproject.github.io/training-material/).</p> <p>One 6 µm thick section of formalin-fixed paraffin-embedded mouse kidney (from 6 month old, male, C57 black 6 mice) was mounted onto an indium-tin oxide (ITO) glass slide, deparaffinized and subjected to antigen retrieval in citric acid (pH 6, 100°C, 1h). Before and after antigen retrieval the sample was washed with 10mM ammonium bicarbonate buffer. After air drying, four 1 µl spots of Bombesin (0.01 mg/ml) were placed around the tissue to control digestion.<strong> </strong>Trypsin was sprayed onto the tissue with the iMatrix Sprayer and the sample was incubated for 2 h at 50°C in a humid chamber. Internal Calibrants (Angiotensin I, Substance P, [Glu]-Fibrinopeptide B, ACTH 18-39) were mixed with α-Cyano-4-hydroxycinnamic acid (CHCA) matrix and sprayed onto the sample.</p> <p>The sample was measured with the Applied Biosystems/MDS SCIEX 4800 MALDI TOF/TOF™ Analyzer in reflector positive ion mode and a spatial resolution of 150 µm. The acquired Analyze7.5 file was loaded into Cardinal and filtered to reduce file size and decrease analysis time: Filtering was done for m/z values between 1220 and 1625 as well as pixel that represent about half of the kidney and one Bombesin digestion control spot. The data was exported in the common data format imzML.</p> <p> </p>
Allele‐specific cis‐regulatory methylation of the gene for vasoactive intestinal peptide in white‐throated sparrows
<p>White-throated sparrows (<em>Zonotrichia albicollis</em>) offer a unique opportunity to connect genotype with behavioral phenotype. In this species, a rearrangement of the second chromosome is linked with territorial aggression; birds with a copy of this "supergene" rearrangement are more aggressive than those without it. The supergene has captured the gene <em>VIP</em>, which encodes vasoactive intestinal peptide, a neuromodulator that causes aggression in other songbirds. In white-throated sparrows, <em>VIP</em> expression is higher in the anterior hypothalamus of birds with the supergene than those without it, and expression of <em>VIP</em> in this region predicts the level of territorial aggression regardless of genotype. Here, we aimed to identify epigenetic mechanisms that could contribute to differential expression of <em>VIP</em> both in breeding adults, which exhibit morph differences in territorial aggression, and in nestlings, before territorial behavior develops. We extracted and bisulfite-converted DNA from samples of the hypothalamus in wild-caught adults and nestlings and used high-throughput sequencing to measure DNA methylation of a region upstream of the <em>VIP</em> start site. We found that the allele inside the supergene was less methylated than the alternative allele in both adults and nestlings. The differential methylation was attributed primarily to CpG sites that were shared between the alleles, not to polymorphic sites, which suggests that epigenetic regulation is occurring independently of the genetic differentiation within the supergene. This work represents an initial step toward understanding how epigenetic differentiation inside chromosomal inversions leads to the development of alternative behavioral phenotypes.</p>
Fluorescence complementation enables quantitative imaging of cell penetrating peptide-mediated protein delivery in plants including WUSCHEL transcription factor
<p>These are data related to the manuscript titled "Fluorescence complementation enables quantitative imaging of cell penetrating peptide-mediated protein delivery in plants including WUSCHEL transcription factor" whose preprint can be found here: https://doi.org/10.1101/2022.05.03.490515</p>
Data for HydrAMP - a deep generative model for antimicrobial peptide discovery
<ul> <li>data- training data for peptides < 25 AA (16.8 MB)</li> <li>models - checkpoints of HydrAMP, PepCVAE, and Basic models for every training epoch (466 MB)</li> <li>results - dumped generation results for every model. Required for running comparison notebooks (832 MB)</li> <li>wheels - custom TensorFlow packages (1 GB)</li> </ul> <p> </p>
Data for: Self-cleaving 2A peptides allow for expression of multiple genes in Dictyostelium discoideum
<p>The social amoeba <em>Dictyostelium discoideum</em> is a model for a wide range of biological processes including chemotaxis, cell-cell communication, phagocytosis, and development. Interrogating these processes with modern genetic tools often requires the expression of multiple transgenes. While it is possible to transfect multiple transcriptional units, the use of separate promoters and terminators for each gene leads to large plasmid sizes and possible interference between units. In many eukaryotic systems this challenge has been addressed through polycistronic expression mediated by 2A viral peptides, permitting efficient, co-regulated gene expression. Here, we screen the most commonly used 2A peptides, porcine teschovirus-1 2A (P2A), <em>Thosea asigna</em> virus 2A (T2A), equine rhinitis A virus 2A (E2A), and foot-and-mouth disease virus 2A (F2A), for activity in <em>D. discoideum</em> and find that all the screened 2A sequences are effective. However, combining the coding sequences of two proteins into a single transcript leads to notable strain-dependent decreases in expression level, suggesting additional factors regulate gene expression in <em>D. discoideum</em> that merit further investigation. Our results show that P2A is the optimal sequence for polycistronic expression in <em>D. discoideum</em>, opening up new possibilities for genetic engineering in this model system.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.