Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
347
datasets available to search
ShareScore release 0.9.0
Dataset results
347 results for “structural proteins”
Functional characterization of 3D-protein structures informed by human genetic diversity - data
<p>Supplementary data for https://www.biorxiv.org/content/early/2017/08/29/182287</p>
Synchrotron diffraction images for the 0.72-Å crystal structure of perdeuterated human myelin protein P2
<p>3600 synchrotron X-ray diffraction images used to refine the structure of perdeuterated human myelin protein P2 at 0.72-Å resolution. Processing files are included. The data were collected on the P11 synchrotron beamline at PETRAIII/DESY, Hamburg.</p>
Protein structure file as input 2AC0.sce (YASARA scene)
<p>Protein structure file as input for predicting the effect of mutations on protein structure</p> <p>2AC0.sce, adapted structure<br> This is a part of a tetrameric complex of the transcription factor P53 bound to DNA. 3 of the 4 P53 structures have been removed for simplicity and visualized some nice features.<br> <br> 2AC0_Repaired.sce, minimized structure, saved as YASARA scene object</p>
Data for "Machine Learning Scoring Functions for Drug Discovery from Experimental and Computer-generated Protein-Ligand Structures: Towards Per-target Scoring Functions"
<p>Data used in "<em>Machine Learning Scoring Functions for Drug Discovery from Experimental and Computer-generated Protein-Ligand Structures: Towards Per-target Scoring Functions</em>"<br> by F. Pellicani, D. Dal Ben, A. Perali, S. Pilati</p> <p>If you use these data or the python script for your research or other activities, please cite the corresponding journal article.</p> <p> </p> <p>====================</p> <p>Uncompressing the zipped file <em>DataSFUnicam.zip</em> provies the following files and folders:</p> <p><br> <strong>DataSFUnicam/</strong></p> <p> </p> <p> ExperimentalDataPDBFiles/<br> <em>This folder contains 2408 .pdb files of experimental complex structures. The files are named with a univocal code corresponding to the protein-ligand complex.</em></p> <p> </p> <p> ExperimentalDataXLSXFile.xlsx<br> <em>This Excel file reports the experimental protein-ligand chemical information. In the sheet named “Foglio1”, the first column contains the univocal code of the protein-ligand complex, the second column contains the experimentally measured pK_d.</em></p> <p> </p> <p> SyntheticDataPDBFiles/<br> <em>This folder contains the .pdb files of the synthetic complex structures. The .pdb files are grouped in 17 folders according to just as many target proteins. The folders are named after the corresponding protein. Each folder contains the .pdb files for the best position of each protein-ligand pair according to the MOE docking score. The files are named with a univocal code.</em></p> <p> </p> <p> SyntheticDataXLSXFiles/<br> <em> The folder contains 17 Excel files with the chemical information of the synthetic protein-ligand complexes. The files are named after the corresponding target protein. In the sheet named “Foglio1” of each .xlsx file, the first column contains a univocal code of the protein-ligand complex in each conformation, the second column contains an auxiliary numerical code corresponding to the protein-ligand pair, the third column contains the experimentally measured pK_i, and the fourth column contains the docking score provided by the MOE software.</em></p> <p>====================</p> <p>USER GUIDE FOR THE PYTHON SCRIPT</p> <p>Download and uncompress the zipped file "<em>SFUnicam.zip</em>" with a command like "<em>unzip SFUnicam.zip</em>". </p> <p>The following file structure is created:</p> <p><em>SFUnicam/</em></p> <p> <em>ComplexToBePredictedFolder/4ey5_30.pdb <br> MaxAssMatrix.npy<br> my_model<br> devStndSynt.npy<br> mediaSynt.npy<br> UnicamSF13prot.py<br> README.txt</em><br> <br> The subfolder "<em>ComplexToBePredictedFolder/</em>" contains the example PDB file "<em>4ey5_30.pdb</em>".</p> <p>-) To execute the script "<em>UnicamSF13prot.py</em>", Python 3 should be installed with the following libraries and sublibraries:<br> <em>Keras:<br> Regularizers<br> Sequential (keras.models)<br> Conv1D, Dense, MaxPooling1D, GlobalMaxPooling1D, GlobalAveragePooling1D, AveragePooling1D (keras.layers)<br> Adam (keras.optimizers)<br> Numpy</em><br> <em>Tensorflow</em></p> <p>Operation:<br> -) Copy the .pdb file related to the protein-ligand complex whose affinity is to be predicted in the subfolder “<em>ComplexToBePredictedFolder/</em>”.<br> -) Make sure the following files are in the same folder where the python script is:<br> <em>MaxAssMatrix.npy<br> mediaSynt.npy<br> devStndSynt.npy<br> my_model</em><br> -) Run the code using Python 3 with a command like "<em>python3.x UnicamSF13prot.py</em>".<br> -) Enter the name of the protein-ligand PDB file whose affinity is to be predicted (excluding the extension ".pdb").<br> -) Read the predicted affinity from screen.<br> </p> <p> </p>
Deep Supervised and Convolutional Generative Stochastic Network for Protein Secondary Structure Prediction, 2014
<p>This contains the protein sequence and secondary structure dataset from <a href="https://proceedings.mlr.press/v32/zhou14.html"><strong>Deep Supervised and Convolutional Generative Stochastic Network for Protein Secondary Structure Prediction</strong></a><strong>, ICML, 2014</strong></p> <p>This dataset was originally hosted at http://www.princeton.edu/~jzthree/datasets/ICML2014/. Since the original URL is no longer available and the dataset is still used by many, I moved the dataset here.</p>
Structure and Function of Salivary Proteins
ClinicalTrials.gov study NCT00916682. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Discovery and Validation of Protein Structural Complexes in Circulating Biofluids As Novel Biomarkers for Early Diagnosis, Prognosis and Therapeutic Management of Patients Affected by Neurodegenerativ
ClinicalTrials.gov study NCT06803784. IPD Sharing: NO. Countries: 1. Publications: 0.
Study to Evaluate the Potential of Air Structuring Protein to Elicit Allergic Reactions in Mold Sensitized People
ClinicalTrials.gov study NCT01494194. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Effect of Protein Supplementation and a Structured Exercise Program on Muscle in Women After Bariatric Surgery.
ClinicalTrials.gov study NCT04771377. IPD Sharing: NO. Countries: 1. Publications: 0.
CXXC zinc finger protein 1 (Cfp1) controls cardiomyocyte maturation by modifying histone H3K4me3 of structural, metabolic, and contractile related genes
GEO Series GSE240852. Mus musculus. 6 samples. Type: Expression profiling by high throughput sequencing.
Brucella MucR acts as an H-NS-like protein to silence virulence genes and structure the nucleoid
GEO Series GSE234935. Brucella abortus 2308. 12 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Other.
Structural annotation of equine protein-coding genes determined by mRNA sequencing
GEO Series GSE21925. Equus caballus. 8 samples. Type: Expression profiling by high throughput sequencing.
Data from: Diversity of opisthokont septin proteins reveals structural constraints and conserved motifs
Open the record for dataset details and reuse information.
Data from: Modelling dynamics in protein crystal structures by ensemble refinement
Open the record for dataset details and reuse information.
RNA G-quadruplex secondary structure promotes alternative splicing via the RNA binding protein hnRNPF
GEO Series GSE107542. Homo sapiens. 4 samples. Type: Expression profiling by high throughput sequencing.
Cooperative engagement and subsequent selective displacement of SR proteins define the pre-mRNA 3D structural scaffold for early spliceosome assembly
GEO Series GSE188652. Human adenovirus 2. 22 samples. Type: Other.
Reciprocal regulation of cardiac chromatin by the chromatin structural proteins HMGB and CTCF
GEO Series GSE80453. Rattus norvegicus. 2 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Genome-wide DNA-binding profile of the Vibrio cholerae histone-like nucleoid structuring protein (H-NS)
GEO Series GSE64249. Vibrio cholerae. 2 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
RNA structure probing to characterize RNA-protein interations on a low abundance pre-mRNA in living cells
GEO Series GSE159719. Mus musculus. 13 samples. Type: Other.
Structural analysis of the lncRNA SChLAP1 reveals protein binding interfaces and a conformationally heterogenous retroviral insertion
GEO Series GSE243328. Homo sapiens. 36 samples. Type: Other.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.