Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
77
datasets available to search
ShareScore release 0.9.0
Dataset results
77 results for “intrinsically disordered”
Datasets of sequences, alignments and structural models generated for the structural prediction of complexes mediated by intrinsically disordered regions.
<p>This repository contains input and ouput files used and generated for the scanning of intrinsically disordered region and the prediction of their binding sites to receptor proteins using the <a href="https://github.com/i2bc/SCAN_IDR">SCAN_IDR</a> pipeline with AlphaFold2-Multimer.</p><p>It contains two archives: </p><ol><li><a href="https://zenodo.org/api/records/10068949/draft/files/scanidr_data_repository_corr6J08.tar/content"><i><strong>scanidr_data_repository_corr6J08.tar</strong></i></a> dedicated to the analysis of a dataset of 42 protein complexes non redundant with the dataset used for AlphaFold2 training,</li><li><a href="https://zenodo.org/api/records/10068949/draft/files/923_elm_cases_repository.tar.gz/content"><i><strong>923_elm_cases_repository.tar.gz</strong></i></a> dedicated to the analysis of 923 complexes from the ELM database.</li></ol><p>These data can be used to rerun specific sections of the pipeline and scripts provided in: <a href="https://github.com/i2bc/SCAN_IDR">https://github.com/i2bc/SCAN_IDR</a></p><h4><strong>Dataset of 42 non redundant complexes</strong></h4><p>The first archive <a href="https://zenodo.org/api/records/10068949/draft/files/scanidr_data_repository_corr6J08.tar/content"><i><strong>scanidr_data_repository_corr6J08.tar</strong></i></a> contains 3 compressed directories and a README file detailing their contents :</p><ul><li>the initial raw sequence and alignment data for every chain -> DIRECTORY <strong>fasta_msa/</strong></li><li>the input and output data of every Alphafold run for every complex -> DIRECTORY <strong>af2_runs/</strong></li><li>the native reference structures -> DIRECTORY <strong>ref_capri_curated/</strong></li></ul><p>The protein-peptide complex cases have been assigned a distinct index number, from 1 to 42, consistent across the several directories of the archive. Their corresponding directories are labelled as <i><index>_<pdbcode></i>.</p><p><i>The models in this archive were generated using AlphaFold2-Multimer v2.2</i></p><h4><strong>Dataset of 923 complexes selected from the ELM database</strong></h4><p>The second archive <a href="https://zenodo.org/api/records/10068949/draft/files/923_elm_cases_repository.tar.gz/content"><i><strong>923_elm_cases_repository.tar.gz</strong></i></a> contains input and ouput files used and generated for the analysis of 923 Eukaryotic Linear Motifs (ELM) database entries.</p><p>Each ELM entry is indexed with specific integer id and is composed of a receptor and a ligand protein. </p><p>The archive contains a Table associating ELM indexes with the ELM entry information, 5 directories and a README file detailing their contents:</p><ul><li>the table describing ELM entries -> FILE <strong>Table_923ELM_uid_delimitations_info_for_archive.txt</strong></li><li>the initial raw sequence and multiple sequence alignment (MSA) data for every chain -> DIRECTORY <strong>fasta_msa/</strong></li><li>the concatenated MSA model for every ELM complex and protocol used -> DIRECTORY <strong>af2_elm_coali_inputs/</strong></li><li>the best model of every AF2 protocol for every complex according to the AF2 -> DIRECTORY <strong>af2_elm_models/</strong></li><li>the best model cut in the ligand part to select only the ELM motifs as used for the evaluation of the models -> DIRECTORY <strong>elm_cut_models/</strong></li><li>the reference structures used for the evaluation of the models -> DIRECTORY <strong>ref_capri_curated/</strong></li></ul><p><i>The models in this archive were generated using AlphaFold2-Multimer v2.3</i></p>
Supplementary Data Files for the paper "Intrinsically disordered compositional bias in proteins: Sequence traits, region clustering, and generation of hypothetical functional associations"
<div> <div> <div> <div> <p><strong>Supplementary data files relating to <a href="https://doi.org/10.1177/11779322241287485">https://doi.org/10.1177/11779322241287485. </a></strong></p> <p><strong><span>Suppl. File 1: Protein Family Clusters.</span></strong></p> <p><strong><span>Suppl. File 2: Cluster GO enrichments/depletions. </span></strong></p> <p><strong><span>Suppl. File 3: The raw ID-CBR data with annotations. </span></strong></p> <p><strong><span>Suppl. File 4: ­ID-CBR Cluster membership.</span></strong></p> <p><strong><span>Each file has an explanatory header. </span></strong></p> <p> </p> </div> </div> </div> </div>
Intrinsic disorder, phase separation and fibrillation by the Henipavirus V and W proteins | Talk- I PhasAGE International Conference
<p>The <strong>I PhasAGE international conference</strong> brought together members of the PhasAGE consortium as well as outstanding international speakers showcasing high impact achievements in the field of liquid-liquid phase separation in aging and late-onset diseases.</p> <p>For details on conference program please see: https://phasage.eu/phasage-conference-1/ </p>
PhasAGE Expert Seminar- Intrinsic Protein Disorder and Conditional Folding in AlphaFoldDB
<p>The PhasAGE <strong>Expert Seminars</strong> consist of a series of talks with speakers from PhasAGE partner’s institutions to promote a successful transfer of knowledge about PhasAGE topics – biomolecular phase separation, aging and age-related diseases.</p>
PhasAGE Training School 1 - Computational prediction of intrinsic disorder in proteins-DisProt - PRACTICAL
<p>The Training School 1 <strong>“Computational Methods to Study Protein Phase Separation”</strong> is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of <strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide <strong>an overview of the available computational resources</strong> to navigate this knowledge. Participants will have <strong>hands-on training</strong> in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>
PhasAGE Training School 1 - Computational prediction of intrinsic disorder in proteins-MobiDB - PRACTICAL
<p>The Training School 1 <strong>“Computational Methods to Study Protein Phase Separation”</strong> is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of <strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide <strong>an overview of the available computational resources</strong> to navigate this knowledge. Participants will have <strong>hands-on training</strong> in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>
Extreme dynamics of a small molecule in its bound state with an intrinsically disordered protein
<p>These data support the manuscript entitled "Extreme dynamics of a small molecule in its bound state with an intrinsically disordered protein" by Heller, Shukla, Figueiredo, and Hansen.</p><p>This data should be used with the code provided on GitHub at https://github.com/hansenlab-ucl/R2_IDP_small_mol. Once downloaded, this directory should be extracted using the following command:</p><p> tar -xzvf Data.tar.gz</p><p>The directory should be saved with the name 'Data' placed in the same directory as the GitHub README.md file.</p><p><strong>This dataset contains: </strong><br><i>Nuclear Magnetic Resonance (NMR) spectroscopy data files (.ft2 format) including: </i></p><p>* 1H 1D ligand-detected chemical shift titration of 5-fluoroindole (50 uM) with increasing concentrations of the protein, non-structural protein 5A, domains 2 and 3 (NS5A-D2D3), in 1H_1D_ft2_data/</p><p>* 1H pseudo-2D Diffusion Ordered SpectroscopY (DOSY) data of 5-fluoroindole (50 uM) with and without NS5A-D2D3 (75 uM) in 1H_DOSY_data/</p><p>* 1H-15N Heteronuclear Single Quantum Coherence (HSQC) measurements of NS5A-D2D3 (40 uM) in the absence and presence of 5-fluoroindole (160 and 320 uM) in 1H_15N_HSQC_ft2_and_metadata/</p><p>* 19F 1D ligand-detected chemical shift titration of 5-fluoroindole (50 uM) with increasing concentrations of NS5A-D2D3 in 19F_1D_ft2_data/</p><p>* 19F pseudo-2D ligand-detected longitudinal (spin-lattice, R1,eff) relaxation titration data of 5-fluoroindole (50 uM) with increasing concentrations of NS5A-D2D3 in 19F_R1eff_ft2_data/</p><p>* 19F pseudo-2D ligand-detected longitudinal (spin-spin, R2,eff) relaxation titration data of 5-fluoroindole (50 uM) with increasing concentrations of NS5A-D2D3 in 19F_R2eff_ft2_data/</p><p><i>Circular Dichroism (CD) data files (.txt format) including: </i></p><p>* CD measurements of NS5A-D2D3 at increasing concentrations in CD_data/no_molecule/</p><p>* CD measurements of NS5A-D2D3 with and without the small molecule, 5-fluoroindole CD_data/with_molecule/</p><p><i>Metadata </i></p><p>* Metadata from the Biological Magnetic Resonance Data Bank (https://bmrb.io/) used to determine scaling factors for the calculation of chemical shift perturbations in 1H_15N_HSQC_ft2_and_metadata/</p>
III PhasAGE International Conference - PED in 2024: improving the community deposition of structural ensembles for intrinsically disordered proteins - Lecture
<p>The III PhasAGE International Conference "Multiscale understanding of protein aggregation and biomolecular condensates in aging and disease" brought together members of the PhasAGE consortium as well as outstanding international speakers from multidisciplinary fields dedicated to unraveling the intricacies of protein aggregation and biomolecular condensates in the context of aging and disease. For details on the conference program please see https://phasage.eu/iii-phasage-international-conference/. </p>
Borg tandem repeats undergo rapid evolution and are under strong selection to create new intrinsically disordered regions in proteins
<p>This repository contains files that accompany the Schoelmerich <em>et al. </em>(2022) bioRxiv preprint.</p> <p>These files include</p> <p>- all Borg proteins used for protein family clustering (<a href="https://zenodo.org/api/files/f75689a0-d40c-44a6-b154-b74e3594fc04/all_Borg_proteins.fasta">all_Borg_proteins.fasta</a>)</p> <p>- 37 additional aaTR-proteins from manually curated Borg contigs (<a href="https://zenodo.org/api/files/f75689a0-d40c-44a6-b154-b74e3594fc04/37_Borg_aaTR-proteins.fasta?versionId=bfdfe7b6-e3b5-48b8-b7e6-a20304097e7d">37_Borg_aaTR-proteins.fasta</a>)</p> <p>- IQ-TREE of Borg DNA polymerases and reference sequences from doi: 10.1093/nar/gkaa760 (<a href="https://zenodo.org/api/files/f75689a0-d40c-44a6-b154-b74e3594fc04/DNAPolB_iqtree.treefile?versionId=c588a98c-822d-4a2f-a92d-7dcb3b1675f0">DNAPolB_iqtree.treefile</a>)</p> <p>- Borg Sm ribonucleoprotein sequences (<a href="https://zenodo.org/api/files/f75689a0-d40c-44a6-b154-b74e3594fc04/21_Borg_Sm_ribonucleoproteins.fasta">21_Borg_Sm_ribonucleoproteins.fasta</a>)</p> <p>- Borg MHC sequences (<a href="https://zenodo.org/api/files/f75689a0-d40c-44a6-b154-b74e3594fc04/14_Borg_MHC_proteins.fasta">14_Borg_MHC_proteins.fasta</a>)</p> <p>- AlphaFold2 predicted structural models of Borg Sm ribonucleoproteins and MHCs with aaTRs</p>
Molecular dynamics simulations of intrinsically disordered proteins p53TAD and Pup
<p>Intrinsically disordered proteins (IDPs) are highly dynamic systems that play an important role in cell signaling processes and their misfunction often causes human disease. Proper understanding of IDP function not only requires the realistic characterization of their three-dimensional conformational ensembles at atomic-level resolution but also of the time scales of interconversion between their conformational substates. Large sets of experimental data are often used in combination with molecular modeling to restrain or bias models to improve agreement with experiment. It is shown here for the N-terminal transactivation domain of p53 (p53TAD) and Pup how the latest advancements in molecular dynamics (MD) simulations methodology produces native conformational ensembles by combining replica exchange with series of microsecond MD simulations. They closely reproduce experimental data at the global conformational ensemble level, in terms of the distribution properties of the radius of gyration tensor, and at the local level, in terms of NMR properties including <sup>15</sup>N spin relaxation, without the need for reweighting. The IDP ensembles were analyzed by graph theory to identify dominant inter-residue contact clusters and characteristic amino-acid contact propensities. These findings indicate that modern MD force fields with residue-specific backbone potentials can produce highly realistic IDP ensembles sampling a hierarchy of nano- and picosecond time scales providing new insights into their biological function.</p>
Data for: Intrinsically Disordered Proteins form Condensates with Gradually Collapsing Conformations at the Interface
<h3>Data for: Intrinsically Disordered Proteins form Condensates with Gradually Collapsing Conformations at the Interface</h3> <p>We ran simulations for four different systems:</p> <ul> <li>WT: A1-LCD WT (N=137), wild-type (WT) sequence of the low-complexity domain (LCD) of the heterogeneous nuclear ribonucleoprotein A1 (hnRNPA1), with electrostatic interactions, at temperature T=260K</li> <li>WT_noEL_T260: A1-LCD WT (N=137), without electrostatic interactions, at temperature T=260K</li> <li>WT_noEL_T290: A1-LCD WT (N=137), without electrostatic interactions, at temperature T=290K</li> <li>HP: homopolymer consisting of prolines (N=137), at temperature T=550</li> </ul> <p>For every system, we ran five independent simulations over 5µs (1000 frames) and used the last 900 frames (4.5µs) for our analysis.</p> <p>This data repository consists of<br> (1) folders containing the data for every seperate run (*_i, i=1,2,3,4,5) in simulation units<br> (2) folders containing the averaged data of all five runs (*_AVG), converted to SI units<br> (3) a droplet folder, containing the data (square radius of gyration and asphericity) for the whole droplet (for all four systems, all five runs)<br>Units are also clarified in each file's header.</p> <p>The simulation units can be converted to SI units via:</p> <ul> <li>Distance: D = 0.45nm</li> <li>Mass: M = 57.05amu</li> <li>Energy: epsilon = 0.2 kcal/mol</li> </ul> <p> </p> <p>Details for (1) and (2):<br>Each folder (*_i, i=1,2,3,4,5, and *_AVG) contains the following subfolders and files:</p> <p><strong>Ree:</strong></p> <ul> <li>distribCos2_all.dat: distribution of cos^2(θ_{ee}) of the whole chains, where θ_{ee} is the angle between the polymer's center r_c and the chain’s end-to-end vector Ree [Fig. S3b, Fig. S6b, Fig. S9b, Fig. S12b]</li> <li>distribCos2_segment_i.dat: distribution of cos^2(θ_{ee,s}) of segment seg_i, where θ_{ee,s} is the angle between the segment's center r_{c,s} and the segment’s end-to-end vector R_{ee,s} [Fig. S4d, Fig. S7d, Fig. S10d, Fig. S13d]</li> <li>distribCos2_segments_all.dat: distribution of cos^2(θ_{ee,s}) of all segments seg_i, where θ_{ee,s} is the angle between the segment's center r_{c,s} and the segment’s end-to-end vector R_{ee,s} [Fig. S4d, Fig. S7d, Fig. S10d, Fig. S13d]</li> <li>distribMonomer_all.dat: distribution of the monomers [Fig. 1, Fig. S4c, Fig. S7c, Fig. S10c, Fig. S13c]</li> <li>distribPolymer_all.dat: distribution of the polymers (whole chains, binned via polymer center position) [Fig. 1, Fig. S4c, Fig. S7c, Fig. S10c, Fig. S13c]</li> <li>distribPolymerEndPos.dat: distribution of the polymer end positions (whole chains) [Fig. S4c, Fig. S7c, Fig. S10c, Fig. S13c]</li> <li>distribPolymerEndPos_segment_i.dat: distribution of the polymer segment end positions of seg_i</li> <li>distribPolymerEndPos_segment_all.dat: distribution of the polymer segment end positions of all segments</li> <li>distribPolymerRee2_all.dat: distribution of Ree^2 (whole chains), binned via polymer center position r_c</li> <li>distribPolymerRee_segment_i.dat: distribution of Ree^2 of segment seg_i, binned via segment center position r_{c,s}</li> <li>distribPolymerRee_segments_all.dat: distribution of Ree^2 of all segments, binned via segment center position r_{c,s}</li> <li>distribPolymerSegment_i.dat: distribution of polymer segment seg_i, binned via segment center position r_{c,s}</li> <li>distribPolymerSegments_all: distribution of all polymer segments, binned via segment center position r_{c,s}</li> </ul> <p><strong>Rg:</strong></p> <ul> <li>distribCos2_all.dat: distribution of cos^2(θ) of the whole chains, where θ is the angle between the polymer's center r_c and the eigenvector belonging to the largest eigenvalue of the chain’s gyration tensor [Fig. S3b, Fig. S6b, Fig. S9b, Fig. S12b]</li> <li>distribCos2_segment_i.dat: distribution of cos^2(θ_s) of segment seg_i, where θ_s is the angle between r_{c,s} and the eigenvector belonging to the largest eigenvalue of the segment’s gyration tensor [Fig. S4b, Fig. S7b, Fig. S10b, Fig. S13b]</li> <li>distribCos2_segments_all.dat: distribution of cos^2(θ_s) of all segments seg_i, where θ_s is the angle between r_{c,s} and the eigenvector belonging to the largest eigenvalue of the segment’s gyration tensor [Fig. S4b, Fig. S7b, Fig. S10b, Fig. S13b]</li> <li>distribMonomer_all.dat: distribution of the monomers [Fig. 1, Fig. S4c, Fig. S7c, Fig. S10c, Fig. S13c]</li> <li>distribPolymer_all.dat: distribution of the polymers (whole chains, binned via polymer center position) [Fig. 1, Fig. S4c, Fig. S7c, Fig. S10c, Fig. S13c]</li> <li>distribMonomerRg_all.dat: distribution of monomer weighted Rg^2 (whole chains), referred to as R_{g,mono}^2 (following Farag et. al) [Fig. 2, Fig. S3a, Fig. S6a, Fig. S9a, Fig. S12a]</li> <li>distribPolymerRg_all.dat: distribution of Rg^2 (whole chains), binned via polymer center position r_c [Fig. 2, Fig. S3a, Fig. S6a, Fig. S9a, Fig. S12a]</li> <li>distribPolymerRg_segment_i.dat: distribution of Rg^2 of segment seg_i, referred to as R_{g,s}^2, binned via segment center position r_{c,s} [Fig. S4a, Fig. S7a, Fig. S10a, Fig. S13a]</li> <li>distribPolymerSegment_i.dat: distribution of polymer segment seg_i, binned via segment center position r_{c,s}</li> <li>distribPolymerSegments_all: distribution of all segments, binned via segment center position r_{c,s}</li> </ul> <p><strong>resDist:</strong></p> <ul> <li>distribPolymerRee2_base_resDistance_s.dat: distribution of Ree2 of all chain segments of length s=|j-i|, binned according to the segment base position r_i [Fig. 3, Fig. S5, Fig. S8, Fig. S11, Fig. S14]</li> <li>distribPolymerRee2_center_resDistance_s.dat: distribution of Ree2 of all chain segments of length s=|j-i|, binned according to the segment center position r_{c,s} [Fig. 3, Fig. S5, Fig. S8, Fig. S11, Fig. S14]</li> <li>distribPolymerRg2_base_resDistance_s.dat: distribution of Rg2 of all chain segments of length s=|j-i|, binned according to the segment base position r_i [Fig. 3, Fig. S5, Fig. S8, Fig. S11, Fig. S14]</li> </ul> <p>distribPolymerRg2_center_resDistance_s.dat: distribution of Rg2 of all chain segments of length s=|j-i|, binned according to the segment center position r_{c,s} [Fig. 3, Fig. S5, Fig. S8, Fig. S11, Fig. S14] </p> <p> </p> <p>Details for (3):<br>The folder '<strong>droplet</strong>' contains four system folders (HP, WT, WT_noEL_T260, WT_noEL_T290). Each of those folders contains the following files:</p> <ul> <li>runX_cluster_Rg2_Rg2Normal_kappa2.dat: for every run X, one finds the time evolution (in simulation units, with 1e8 timesteps = 1µs) of the square radius of gyration Rg2 of the full droplet, its x-, y- and z-components, its three eigenvalues and the droplet asphericity A (referred to as kappa2 in the header) [Fig.S1c, Fig.S1d]</li> <li>AVG_cluster_Rg2_Rg2Normal_kappa2.dat: average of the parameters from the runX_cluster_Rg2_Rg2Normal_kappa2.dat files, over all five runs, using the last 900 snapshots (4.5µs) of every run [Fig. S1a, Fig. S1b]</li> <li>STD_cluster_Rg2_Rg2Normal_kappa2.dat: standard deviation of the parameters from the runX_cluster_Rg2_Rg2Normal_kappa2.dat files, over all five runs, using the last 900 snapshots (4.5µs) of every run [Fig. S1a, Fig. S1b]</li> </ul>
Phosphorylation regulated conformational diversity and topological dynamics of an intrinsically disordered nuclear receptor
<p>Molecular dynamics simulations of AF1c region of human glucocorticoid receptor and its phosphovariants as described in the below paper: </p> <p>Phosphorylation regulated conformational diversity and topological dynamics of an intrinsically disordered nuclear receptor</p> <p>Vasily Akulov, Alba Jiménez Panizo, Eva Estébanez-Perpiñá, John van Noort, Alireza Mashaghi</p> <p> </p> <p>The data related to this project has been deposited in two repositories. This repository contains the second part of the data; the first part can be found at DOI: 10.5281/zenodo.13820169</p>
Molecular dynamics simulations of intrinsically disordered proteins p53TAD and Pup
Open the record for dataset details and reuse information.
Data repository associated with 'A Functional Map of the Human Intrinsically Disordered Proteome'
<p><strong>ES_MAP.zip</strong></p> <ul> <li>a hierarchically clustered map of the human IDR-ome</li> <li>.cdt and .gtr files - outputs of Cluster3.0 software</li> <li>can be visualized using JavaTreeView (see Tutorial_ES.pdf)</li> </ul> <p><strong>TUTORIAL.zip</strong>, information on:</p> <ul> <li>visualization and analysis of the human IDR-ome map</li> <li>search for proteins of interest and exploratory analyses of clusters</li> <li>automatic export and analysis of exported clusters (code available at https://github.com/IPritisanac/ES_PW)</li> </ul> <p><strong>IDROME_SEQUENCES.zip</strong></p> <ul> <li>human proteome fasta file</li> <li>IDRome fasta file</li> <li>SPOT-Disorder v1.0 disorder boundaries <ul> <li>13 044 unique protein sequences with at least one IDR (>=30 amino acids)</li> <li>21 252 total unique human IDRs</li> </ul> </li> </ul> <p><strong>IDR_ALN.zip</strong></p> <ul> <li>alignments of IDR sequences across ENSEMBL orthologs</li> <li>19 459 IDR alignments</li> <li>UniProt ID and IDR boundaries for the human sequence are indicated in the name of the file</li> </ul> <p><strong>FAIDR_TSTATS.zip</strong></p> <ul> <li>hierarchical clustering of FAIDR t-statistics for 148 GO terms<br> <ul> <li>.cdt, .gtr files from Cluster3.0</li> <li>can be visualized using JavaTreeView</li> <li>reveals the most predictive molecular features for the top performing 148 models</li> </ul> </li> </ul> <p><strong>CLUSTERS_EXPLORE.zip</strong></p> <ul> <li>clusters obtained through exploratory analysis of the map provided in ES_MAP.zip</li> <li>93 exported clusters in .cdt file format</li> </ul> <p><strong>CLUSTERS_AUTO.zip</strong></p> <ul> <li>clusters extracted from the hierarchically clustered IDR-ome map at a range of distance thresholds (0.4 - 0.8) in .cdt file format</li> <li>distance refers to the uncentered correlation distance between vectors of Z-scores representing human IDRs</li> <li>clusters extracted at different distance thresholds are split into separate archives</li> <li>AUTO_GO_FEATS.xlsx - summary of GO-term overrepresentation and feature enrichment analyses; each distance threshold is in a separate sheet</li> </ul> <p><strong>FAIDR_HIGH_AUC_PPV_GO.zip</strong></p> <ul> <li>target files with annotations of 148 GO terms for which good quality FAIDR models could be obtained (AUC >= 0.7, PPV >= 0.4)</li> <li>file format: three columns; 1st: IDR ID (includes IDR boundaries); 2nd: protein UniProt ID; 3rd: annotation of the protein to a GO term (1 if known to be associated with the GO term, 0 if not)</li> </ul> <p> </p> <p> </p>
Intrinsically disordered Prosystemin discloses biologically active repeat motifs
<p>Dataset referred to "Intrinsically disordered Prosystemin discloses biologically active repeat motifs".</p> <p>Abstract: The in-depth studies over the years on the defence barriers by tomato plants have shown that the Systemin peptide controls the response to a wealth of environmental stress agents. This multifaceted stress reaction seems to be related to the intrinsic disorder of its precursor protein, Prosystemin (ProSys). Since latest findings show that ProSys has biological functions besides Systemin sequence, here we wanted to assess if this precursor includes peptide motifs able to trigger stress-related pathways. Candidate peptides were identified in silico and synthesized to test their capacity to trigger defence responses in tomato plants against different biotic stressors. Our results demonstrated that ProSys harbours several repeat motifs which triggered plant immune reactions against pathogens and pest insects. Three of these peptides were detected by mass spectrometry in plants expressing ProSys, demonstrating their effective presence in vivo. These experimental data shed light on unrecognized functions of ProSys, mediated by multiple biologically active sequences which may partly account for the capacity of ProSys to induce defense responses to different stress agents.</p>
alfa-synuclein spectra for the paper "Linear discriminant analysis reveals hidden patterns in NMR chemical shifts of intrinsically disordered proteins"
<p>The experimental data for the paper "Linear discriminant analysis<br> reveals hidden patterns in NMR chemical shifts of intrinsically<br> disordered proteins" - four spectra of an intrinsically disordered<br> protein alfa-synuclein:</p> <p>3D HNCO</p> <p>4D HabCab(CO)NH</p> <p>4D (H)N(CA)CONH</p> <p>4D HNCACO</p> <p>All spectra were acquired using non-uniform sampling (schedules included<br> as 'schedule.txt' files in a format of 'Kozminski' method in VnmrJ). 3D<br> spectrum (HNCO_nuFT.ucsf) was processed using multidimensional Fourier transform<br> (http://nmr.cent3.uw.edu.pl/software, program 'toastd'). Additionally, we provide a spectrum cleaned of NUS artifacts (HNCO_artifacts_cleaned.ucsf) using program handy (http://nmr.cent3.uw.edu.pl/software, program 'handy').</p> <p>4D spectra were processed using sparse multidimensional Fourier transform<br> (http://nmr.cent3.uw.edu.pl/software, program 'reduced'), based on a<br> HNCO peak list ('peak.list' file). Files with processing parameters are<br> included as 'parameters.txt' files.</p> <p>The resulting spectra are included as Sparky .ucsf files. The files for<br> 4D spectra are the collections of 2D cross-sections with the<br> cross-section number corresponding to the numbering of HNCO peaks in the<br> 'peak.list' file.</p> <p>The folder 'Application2_assignment_transfer' contains Sparky 'ucsf' and 'save' files of the two-dimensional NH projection of the HNCO spectrum with the experimental (HSQC_exp.list) and BMRB (HSQC_bmrb.list) peak lists read into the spectrum. The experimental peaks are marked with arbitrary numbers and BMRB peaks - with names of the preceding aa-residues. The correct assignment is shown in the file 'assignment.txt'.</p>
MBE Growth of GeSn devices with intrinsic disorder
<p>These files consist of experimental data on GeSn devices. There are MATLAB files that were used to plot and analyse data.</p>
MD simulation data: Intrinsic disorder is a conserved feature of hepatitis C virus E2 glycoprotein
<p><strong>Background</strong></p> <p>Equilibration, relaxation and production runs were performed on GPUs using the CUDA version of PMEMD in AMBER 16 and AMBER ff14SB force field. Minimisation steps were performed on a CPU using PMEMD in AMBER 16 and the AMBER ff14SB force field. The GLYCAM_06j-1 force field was used for the simulations of the glycosylated E2. All software is available from http://ambermd.org/. </p> <p><strong>Contents</strong></p> <p>There are three tarball (<strong>.tar.gz</strong>) files, one for each of the HCV strain investigated. </p> <p>The contents of each tarball is as follows:</p> <p>1. a source PDB (<strong>.pdb</strong>) file</p> <p>2. <strong>leap.scr</strong> - a script used to create the .prmtop and .inpcrd files</p> <p>3. Two AMBER parameter/topology (<strong>.prmtop</strong>) (one with hydrogen mass repartitioning) and an AMBER coordinate (<strong>.inpcrd</strong>) file </p> <p>4. Multiple control (<strong>.ctl</strong>) files numbered 1 to 10 that are used to minimize (<strong>min</strong> prefix), relax (<strong>rel</strong> prefix) and equilibrate (<strong>equ</strong> prefix) the model</p> <p>5. Executable <strong>do_md</strong> that performed all the minimisation, relaxation and equilibration steps</p> <p>6. File <strong>parmed.txt</strong> used to repartition the hydrogen atom mass (for the unglycosylated simulations)</p> <p>7. control file <strong>prod.ctl</strong> used for the production run </p> <p>8. Executable <strong>run_prod</strong> that was used to perform the production run</p> <p>9. Two control files (<strong>prod_short.ctl </strong>and <strong>prod_short_2.ctl</strong>) for the short runs used to de-correlate the simulation for the independent runs</p> <p>10. Executable <strong>run_short</strong> and <strong>run_short_2</strong> used to carry out the de-correlated production runs.</p> <p>10. Five AMBER trajectory (<strong>.nc</strong>) files for five independent MD simulations, numbered 1 to 5. <strong>Note: </strong>each of these files is over 2GB.</p>
Phosphorylation regulated conformational diversity and topological dynamics of an intrinsically disordered nuclear receptor
<p>Molecular dynamics simulations of AF1c region of human glucocorticoid receptor and its phosphovariants as described in the below paper: </p> <p><strong>Phosphorylation regulated conformational diversity and topological dynamics of an intrinsically disordered nuclear receptor</strong></p> <p>Vasily Akulov, Alba Jiménez Panizo, Eva Estébanez-Perpiñá, John van Noort, Alireza Mashaghi</p> <p> </p> <p>The data related to this project has been deposited in two repositories. This repository contains the first part of the data; the second part can be found at the DOI: 10.5281/zenodo.13822438</p>
PhasAGE Training School 1 - Phase separation and emergent functions of Intrinsically Disordered Proteins- Lecture
<p>The Training School 1 <strong>“Computational Methods to Study Protein Phase Separation”</strong> is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of <strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide <strong>an overview of the available computational resources</strong> to navigate this knowledge. Participants will have <strong>hands-on training</strong> in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.