Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
41
datasets available to search
ShareScore release 0.9.0
Dataset results
41 results for “transmembrane protein”
Prediction and Visualization of Human Transmembrane Proteins using AlphaFold and Protein Language Models
<p><strong>Description:</strong> <strong>TMvis</strong> ("TMvis496.tar.gz") is a dataset containing 496 3D-structures of predicted human transmembrane proteins (TMP) and their predicted membrane embedding. The method TMbed [1], based on the protein language model ProtT5 [2] predicted 4.967 TMP for the human proteome (20,375 proteins, UniProt [3] version April 2022; excluding TITIN_HUMAN due to length). For these proteins, we obtained AlphaFold [4] structures from AlphaFoldDB [5] with an average per-residue confidence score (pLDDT) of more than 90%. This resulted in the 496 proteins of TMvis, as can be found in "TMvis496.fasta". The membrane embedding was predicted using the methods ANVIL [6], PPM3 [7], and per-residue TMbed predictions. As the three methods are based on different approaches, we decided to publish results for all. The figure “TMvis_project_overview.png” provides a graphical overview for each step described above.</p> <p><strong>TMvis Folder Structure:</strong> TMvis is separated into “alpha” containing predicted alpha-helical TMPs, and “beta” containing predicted beta-barrel TMPs. Within these folders, each protein is assigned one folder, identifiable by the respective unique UniProt ID. Each protein folder consists of:<br> - “UniprotID.fasta” with UniProt ID, sequence, TMbed per-residue prediction<br> - “AF-UniprotID-F1-model_v2.pdb” with the AlphaFold structure<br> - “AF-UniprotID-F1-model_v2.cif” with the AlphaFold structure<br> - “AF-UniprotID-F1-model_v2_ANVIL.pdb” with predicted ANVIL membrane embedding<br> - “AF-UniprotID-F1-model_v2_ppm.pdb” predicted PPM3 membrane embedding</p> <p>TMvis <br> | <br> ├── alpha <br> │ │ <br> │ ├── A0A087X1C5 <br> │ │ ├── A0A087X1C5.fasta <br> │ │ ├── AF-A0A087X1C5-F1-model_v2.pdb <br> │ │ ├── AF-A0A087X1C5-F1-model_v2.cif <br> │ │ ├── AF-A0A087X1C5-F1-model_v2_ANVIL.pdb <br> │ │ └── AF-A0A087X1C5-F1-model_v2_ppm.PDB <br> │ └── ... <br> └── beta <br> └── P45880</p> <p><strong>TMvis visualization:</strong> The 3D-visualization of every protein in the dataset TMvis can be easily accessed using the Jupyter Notebook “TMvis.ipynb”. It contains detailed descriptions the different membrane prediction tools ANVIL, PPM3, and TMbed as well as the respective code. Additionally, it allows to visualize the per-residue confidence scores (pLDDT) of AlphaFold.</p> <p>——————————————————————————————————————————————————————————————————————————</p> <p><strong>References:</strong></p> <p>[1] TMbed - TMbed Bernhofer, Michael, and Burkhard Rost. 2022. “TMbed – Transmembrane Proteins Predicted through Language Model Embeddings.” bioRxiv.</p> <p>[2] ProtT5 - A. Elnaggar et al., "ProtTrans: Towards Cracking the Language of Lifes Code Through Self-Supervised Deep Learning and High Performance Computing," in IEEE Transactions on Pattern Analysis and Machine Intelligence, doi: 10.1109/TPAMI.2021.3095381.</p> <p>[3] UniProt - UniProt Consortium (2021). UniProt: the universal protein knowledgebase in 2021. Nucleic acids research, 49(D1), D480–D489.</p> <p>[4] AlphaFold - AlphaFold Jumper, John, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, et al. 2021. “Highly Accurate Protein Structure Prediction with AlphaFold.” Nature 596 (7873): 583–89.</p> <p>[5] Alphafold DB - Varadi, Mihaly, Stephen Anyango, Mandar Deshpande, Sreenath Nair, Cindy Natassia, Galabina Yordanova, David Yuan, et al. 2022. “AlphaFold Protein Structure Database: Massively Expanding the Structural Coverage of Protein-Sequence Space with High-Accuracy Models.” Nucleic Acids Research 50 (D1): D439–44.</p> <p>[6] ANVIL - ANVIL Postic, Guillaume, Yassine Ghouzam, Vincent Guiraud, and Jean-Christophe Gelly. 2016. “Membrane Positioning for High- and Low-Resolution Protein Structures through a Binary Classification Approach.” Protein Engineering, Design & Selection: PEDS 29 (3): 87–91.</p> <p>[7] PPM3 - PPM3 Lomize, Mikhail A., Irina D. Pogozheva, Hyeon Joo, Henry I. Mosberg, and Andrei L. Lomize. 2012. “OPM Database and PPM Web Server: Resources for Positioning of Proteins in Membranes.” Nucleic Acids Research 40 (Database issue): D370–76.</p> <p>——————————————————————————————————————————————————————————————————————————</p> <p><strong>License:</strong></p> <p>This work is licensed under a Creative Commons Attribution 4.0 International License (CC-BY 4.0).</p> <p> </p>
Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain in Pathogen recognition molecules from hemocytes of Planorbarius corneus molluscs (Planorbidae, Pulmonata)
Рис. 2. Варианты преΑсказанной Αоменной структуры патогенраспознающих моΛекуΛ гемоцитов моΛΛюсков Planorbarius corneus. a — фибриногенпоΑобные беΛки, b — гаΛектины, c — F-Λектины. УсΛовные обозначения и сокращения, зΑесь и ΑаΛее: горизонтаΛьные красные поΛоски — сигнаΛьный пептиΑ, горизонтаΛьные розовые — обΛасть низкой сΛожности, вертикаΛьные синие поΛоски — трансмембранная обΛасть, FBG — фибриногеновый Αомен, FTP — Αомен фукоΛектина, EGF — Αомен эпиΑермаΛьного фактора роста, EGF_CA — каΛьцийсвязывающий EGF-поΑобный Αомен, PAN_AP — APPLE-поΑобный Αомен, SCAN — обΛасть, богатая Λейцином, GLECT — гаΛактозосвязывающий Λектин, CLECT — Λектин C-типа, Gal-bind — гаΛактозиΑ–связывающий Λектин, ML — MD-2- поΑробный Αомен распознавания ΛипиΑов Fig. 2. Variants of the predicted domain structure of pattern recognition molecules from hemocytes of Planorbarius corneus molluscs. a — fibrinogen-related proteins, b — galectins, c — F-lectins. Symbols and abbreviations (here and further): horizontal red stripes — signal peptide, horizontal pink stripes — a low complexity region, vertical blue stripes — transmembrane region, FBG — fibrinogen-related domain, FTP — fucolectin domain, EGF — epidermal growth factor-like domain, EGF_CA — calcium-binding EGF-like domain, PAN_AP — APPLE-like domain, SCAN — leucine rich region, Apple — APPLE domain, GLECT — galactose-binding lectin, CLECT — C-type lectin, Gal-bind — galactoside-binding lectin, ML — MD-2-related lipid-recognition domain
Dataset for the Transmembrane protein 106B antibody screening study
<p><strong>This antibody characterization dataset is related to the F1000 research article openly available at F1000Research.</strong></p> <p><em>This project contains the following underlying data included in a study aiming at characterizing antibodies for the Transmembrane protein 106B protein, encoded by the TEM106b gene. The original study is also available on the Zenodo YCharOS community (<a href="https://doi.org/10.5281/zenodo.7459629">https://doi.org/10.5281/zenodo.7459629</a>).</em></p>
Scaling protein-water interactions in the Martini 3 coarse-grained force field to simulate transmembrane helix dimers in different lipid environments
<p>This dataset contains molecular dynamics (MD) trajectories used for preparation of the following manuscript: <br> "Scaling protein-water interactions in the Martini 3 coarse-grained force field to simulate transmembrane helix dimers in different lipid environments". </p>
Docking data for "The evolution of the SARS-CoV-2 spike protein for differential usage of the host transmembrane serine proteases entry pathway"
<p><br>The dataset includes predicted complexes of the SARS-CoV-2 Spike protein (specifically at the S2' cleavage site) with Hepsin and TMPRSS2 proteins. It contains data on three variants: Wuhan, Delta, and Omicron BA.1.</p> <p><strong>Compressed folders:</strong></p> <p>-357596-DeltaHepsin.tgz</p> <p>-357597-DeltaTMPRSS2.tgz</p> <p>-360039-WuhanHepsin.tgz</p> <p>-360042-TMPRSSWuhan.tgz</p> <p>-392981-TMPRSS-BA1_all.tgz</p> <p>-392982-Hepsin-BA-all.tgz</p> <p><strong>Each compressed folder contains the following:</strong></p> <p>-Initial structures in pdb format</p> <p>-Output complexes in pdb format</p> <p>-Clusters in pdb format</p> <p>-Protocols</p> <p>-Parameters</p> <p>-Scoring files</p> <p> </p> <p><strong>Protein-protein docking </strong><br>Molecular docking between the SARS-CoV-2 S protein of Wuhan, Delta (PDB: 7W92, [DOI: 10.1038/s41467-022-28528-w]), and BA.1 (PDB: 7XO5, [DOI: 10.1038/s41422-022-00672-4]) and the human proteases TMPRSS2 (PDB: 8HD8, [DOI: 10.1038/s41467-023-42527-5]) and Hepsin (PDB: 1Z8G, [DOI: 10.1042/BJ20041955]) was performed using the HADDOCK v2.5-2024.03 webserver ([DOI: 10.1021/ja026939x], [DOI: 10.1016/j.jmb.2015.09.014]). Missing loops in the protein structures were reconstructed using Modeller v10.5 ([DOI: 10.1006/jmbi.1993.1626]). Every heteroatom was removed from the reference structures. The relaxed atomistic coordinates for each S protein variant were derived via all-atom molecular dynamics (MD) simulations. These simulations were performed using AMBER22 with the FF19SB force fields and the pmemd.cuda module for enhanced performance ([DOI: 10.1021/acs.jcim.3c01153], [DOI: 10.1021/jz501780a], [DOI:10.1021/ct400314y]). For the Wuhan variant the S protein was retrieved from our previous modeling study [DOI: 10.1039/D0NR03969A] where for Delta and BA.1, ecah S protein was placed in a dodecahedral box, extending 20 Å beyond the solute in every cartesian direction, and solvated with the four-site OPC water model ([DOI: 10.1021/jz501780a]). The systems were neutralized with counterions, specifically one Cl− ion for the Delta variant and three Cl- ions for the BA.1 variant. To remove local clashes, a geometric optimization was performed using the steepest descent algorithm for 5000 cycles. The MD equilibration process consisted of several stages. First, temperature equilibration in the NVT ensemble was performed by gradually increasing the temperature through steps of 150, 200, 250, 300, and finally 310 K, each lasting 200 ps. During this phase, position restraints were applied to the heavy atoms of the proteins, with progressively decreasing spring constants of 5.0, 4.0, 3.0, and 1.0 kcal mol−1 Å−2, facilitating gradual relaxation. This was followed by a 1 ns equilibration at 310 K in the NPT ensemble without restraints. For production MD, the simulations were run in the NPT ensemble with periodic boundary conditions and Particle Mesh Ewald (PME) method ([DOI: 10.1063/5.0040966], [DOI: 10.1021/ct9001015]) using a grid spacing of 1.0 Å for long-range electrostatics. Non-bonded interactions were modeled with a Lennard-Jones potential using a 9Å cutoff. Temperature control was maintained using Langevin dynamics ([DOI: 10.1021/ct800573m]) with a collision frequency of 4.0 ps−1, and pressure control was managed by the Monte Carlo barostat ([DOI: 10.1016/j.cplett.2003.12.039]) with a 2.0 ps relaxation time at 1 bar. Bond constraints on hydrogen atoms were applied using the SHAKE algorithm ([DOI: 10.1016/0021-9991(77)90098-5]), and the hydrogen mass repartitioning scheme was applied via ParmEd ([DOI: 10.1371/journal.pcbi.1005659]), enabling a 4 fs integration time step ([DOI: 10.1021/ct5010406]). Each protein complex was simulated for a total of 20 ns. For the Wuhan variant, the 3D coordinates were retrieved from [DOI: 10.5281/zenodo.3817446].<br>The active interaction region on the spike protein was defined as the cleavage site (residues P809-R815). For TMPRSS2 and Hepsin, the active sites were defined based on their catalytic residues: H296, D345, D435, S441, S460, and G462 for TMPRSS2, and H203, D257, D347, A348, and S353 for Hepsin. These specific regions were selected to guide the docking process and maximize biologically relevant interactions. Docking clusters were analyzed by selecting those with the lowest interaction energies for further structural analysis. To evaluate binding accuracy, native contacts between the S protein and proteases were computed using the contact map analysis based on the OV+rCSU method ([DOI: 10.12693/APhysPolA.145.S9, 10.1021/acs.jctc.6b00986]), which allows for a precise identification of critical stabilizing interactions, both specific and non-specifics. High-frequency contacts, defined as those appearing in over 70% of the generated models, were highlighted as key determinants of protein-protein recognition, providing insight into the most stable and consistent interactions across docking configurations.</p>
Lipidomic datasets for: Transmembrane protein 135 regulates lipid homeostasis through its role in peroxisomal DHA metabolism
<p>Transmembrane protein 135 (TMEM135) is thought to participate in the cellular response to increased intracellular lipids yet no defined molecular function for TMEM135 in lipid metabolism has been identified. In this study, we performed a lipid analysis of tissues from <em>Tmem135</em> mutant mice and found striking reductions of docosahexaenoic acid (DHA) across all <em>Tmem135</em> mutant tissues, indicating a role of TMEM135 in the production of DHA. Since all enzymes required for DHA synthesis remain intact in <em>Tmem135</em> mutant mice, we hypothesized that TMEM135 is involved in the export of DHA from peroxisomes. The <em>Tmem135</em> mutation likely leads to the retention of DHA in peroxisomes, causing DHA to be degraded within peroxisomes by their beta-oxidation machinery. This may lead to generation or alteration of ligands required for the activation of peroxisome proliferator-activated receptor a (PPARa) signaling, which in turn could result in increased peroxisomal number and beta-oxidation enzymes observed in <em>Tmem135</em> mutant mice. We confirmed this effect of PPARa signaling by detecting decreased peroxisomes and their proteins upon genetic ablation of <em>Ppara</em> in <em>Tmem135</em> mutant mice. Using <em>Tmem135</em> mutant mice, we also validated the protective effect of increased peroxisomes and peroxisomal beta-oxidation on the metabolic disease phenotypes of leptin mutant mice which has been observed in previous studies. Thus, we conclude that TMEM135 has a role in lipid homeostasis through its function in peroxisomes.</p>
An amphipol-stabilized multi-pass transmembrane protein as an immunogen to generate mouse memory B cells against native VMAT2
Open the record for dataset details and reuse information.
High-throughput discovery of transmembrane helix dimers from human single-pass membrane proteins with TOXGREEN sort-seq
Open the record for dataset details and reuse information.
Lipidomic datasets for: Transmembrane protein 135 regulates lipid homeostasis through its role in peroxisomal DHA metabolism
Open the record for dataset details and reuse information.
Ribosomal stalk proteins RPLP1 and RPLP2 promote biogenesis of flaviviral and cellular multi-pass transmembrane proteins
<p>The ribosomal stalk proteins, RPLP1 and RPLP2 (RPLP1/2), which form the ancient ribosomal stalk, were discovered decades ago but their functions remain mysterious. We had previously shown that RPLP1/2 are exquisitely required for replication of dengue virus (DENV) and other mosquito-borne flaviviruses. Here, we show that RPLP1/2 function to relieve ribosome pausing within the DENV envelope coding sequence, leading to enhanced protein stability. We evaluated viral and cellular translation in RPLP1/2-depleted cells using ribosome profiling and found that ribosomes pause in the sequence coding for the N-terminus of the envelope protein, immediately downstream of sequences encoding two adjacent transmembrane domains (TMDs). We also find that RPLP1/2 depletion impacts a ribosome density for a small subset of cellular mRNAs. Importantly, the polarity of ribosomes on mRNAs encoding multiple TMDs was disproportionately affected by RPLP1/2 knockdown, implying a role for RPLP1/2 in multipass transmembrane protein biogenesis. These analyses of viral and host RNAs converge to implicate RPLP1/2 as functionally important for ribosomes to elongate through ORFs encoding multiple TMDs. We suggest that the effect of RPLP1/2 at TMD associated pauses is mediated by improving the efficiency of co-translational folding and subsequent protein stability.</p>
Supporting data for transmembrane domain self association simulations in "Recalibration of protein interactions in Martini 3"
<p>This repository contains the data of transmembrane helix self-association simulation from "Recalibration of protein interactions in Martini 3". Simulations were run with the Martini 3.0 force field, along with two modified versions of Martini 3.0 in which the well-depth, ε, in the Lennard-Jones potential between all protein and water beads was rescaled by a factor <em>λ</em><sub>PW</sub>, ε in the Lennard-Jones potential between all protein beads was rescaled by a factor <em>λ</em><sub>PP</sub>. The simulation files are kept in one single zip file, which contains trajectories for two protein EphA1 and ErbB1 systems with three versions of force fields. The trajectory files are in xtc format, and are accompanied by a structure in pdb format for system topology and a tpr file to start the simulation. In each version of force field for each protein, name of the files corresponds to that specific umbrella sampling window. Umbrella sampling windows ranges from 0.6 nm to 3.4 nm with a spacing of 0.2 nm. </p>
Validation of de novo designed water-soluble and transmembrane proteins by in silico folding and melting
<p>Here are all of the datasets generated and analysed during this study. </p> <p>Here is a breakdown of their content:</p> <ul> <li><strong>8_stranded_transmembrane_barrels.zip</strong> - raw data from Alphafold (3 and 48 recycles), ESMFold and raptor predictions of the 8 stranded TMBs. A file with all the sequences is also given</li> <li><strong>12_stranded_transmembrane_barrels.zip - </strong>raw data from the Alphafold and ESMfold predictions of the 12 stranded TMBs. A file with all the sequences is also given</li> <li><strong>water_soluble_barrels.zip</strong> - raw data from the Alphafold and ESMfold predictions of the water soluble beta barrels (designable and non-designable). A file with all the sequences is also given</li> <li><strong>all design models.zip</strong> - original design models for water-soluble (designable and non-designable), 8-stranded and 12-stranded TMBs</li> </ul> <p> </p> <ul> <li><strong>ESMfold_masking_exp.tar - </strong>this tar file contains all the ESMfold masking experiments performed to the water-soluble, 8 and 12-stranded transmembrane barrels. Inside there are zipped datasets for each masking experiment<br> </li> <li> <p><strong>ziped_raw_csv_files.zip - </strong>raw csv files with all the data necessary to analyse the figures </p> </li> <li> <p><strong>analysis_notebooks.zip </strong>- Jupyter notebooks used to analyse the output prediction data for all figures</p> </li> </ul> <p> </p> <p> </p>
Monitoring the binding and insertion of a single transmembrane protein by an insertase.
<p>Data underlying the figures in the publication “Monitoring the binding and insertion of a single transmembrane protein by an insertase”, published in <em>Nat. Commun., </em><em><strong>2021</strong></em><em>, 12, 7082.</em></p> <p><em><a href="https://doi.org/10.1038/s41467-021-27315-3">https://doi.org/10.1038/s41467-021-27315-3</a></em></p> <p>Table of contents:</p> <p><strong>1. Dataset 1</strong>: Experimental data (excel file) for the <em>Figures 1-5 and the Supplementary Figures 1, 5, 7-16.</em></p>
A Study of TACI(Transmembrane Activator and Calcium-modulator and Cyclophilin Ligand (CAML) Interactor)-Antibody Fusion Protein Injection (RC18) in Subjects With Systemic Myasthenia Gravis
ClinicalTrials.gov study NCT04302103. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Ribosomal stalk proteins RPLP1 and RPLP2 promote biogenesis of flaviviral and cellular multi-pass transmembrane proteins
Open the record for dataset details and reuse information.
Annealing simulations of a lipid membrane with different amounts of transmembrane proteins
<p>soon...</p>
Annealing simulations of a lipid membrane with different transmembrane proteins
<p>soon...</p>
Simulations of a lipid membrane with different amounts of transmembrane proteins at different temperatures
<p>soon...</p>
A Mutation in Transmembrane Protein 135 Impairs Lipid Metabolism in Mouse Eyecups
GEO Series GSE184160. Mus musculus. 15 samples. Type: Expression profiling by high throughput sequencing.
Epigenetically regulated Fibronectin leucine rich transmembrane protein 2 (FLRT2) shows tumor suppressor activity in breast cancer cells
GEO Series GSE85247. Homo sapiens. 4 samples. Type: Expression profiling by array.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.