Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

11,855

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

11,855 results for “proteins”

Learn how ShareScore rates datasets ↗
zenodo48/100

The evolution of genomic, transcriptomic, and single-cell protein markers of metastatic upper tract urothelial carcinoma

<p>The molecular characteristics of metastatic upper tract urothelial carcinoma (UTUC) are unknown. The genomic and transcriptomic differences between primary and metastatic UTUC is not well described either. We combined whole-exome sequencing, RNA-sequencing, and Imaging Mass Cytometry<sup>TM</sup>&nbsp;(IMC<sup>TM</sup>) of 44 tumor samples from 28 patients with high-grade primary and metastatic UTUC. IMC enables spatially resolved single-cell analyses to examine the evolution of cancer cell, immune cell, and stromal cell markers using mass cytometry with lanthanide metal-conjugated antibodies. We discovered that actionable genomic alterations are frequently discordant between primary and metastatic UTUC tumors in the same patient. In contrast, molecular subtype membership and immune depletion signature were stable across primary and matched metastatic UTUC. Molecular and immune subtypes were consistent between bulk RNA-sequencing and mass cytometry of protein markers from 340,798 single-cells. Molecular subtyping at the single cell level was highly conserved between primary and metastatic UTUC tumors within the same patient.</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

Associated data from: An end-to-end workflow to study newly synthesized mRNA following rapid protein depletion in Saccharomyces cerevisiae

<p>This dataset includes two custom BED files described in "An end-to-end workflow to study newly synthesized mRNA following rapid protein depletion in&nbsp;<em>Saccharomyces cerevisiae</em>" (Ridenour and Donczew, submitted), which were used to define counting windows for processing SLAM-seq data in SLAM-DUNK (version 0.4.3) [1]. The BED files contain all annotated open reading frames (ORFs) in the<em> Saccharomyces cerevisiae</em> genome or the <em>Schizosaccharomyces</em><em>&nbsp;pombe</em> genome and were created using BEDOPS (version 2.4.3) [2]. All ORFs were then extended 250 bp beyond their stop position to capture 3&prime; untranslated regions (UTRs) using SAMtools (version 1.14) [3] and BEDTools (version 2.30.0) [4]. The reference genome annotations for <em>S. cerevisiae</em> strain S288C (version R64-3-1, RefSeq Assembly GCF_000146045.2) and <em>S. pombe</em> strain 972h- (version ASM294v2, RefSeq Assembly GCF_000002945.1) were retrieved from the NCBI Datasets repository. The <em>S. cerevisiae </em>chromosome names were modified to reflect standard nomenclature (https://www.yeastgenome.org/).</p>

opencc-by-4.0Feb 2024View details →
zenodo48/100

Alpha-Galactosaminidase family GH114 protein from Fusarium solani: X-ray diffraction images

<p>This submission includes h5-files with diffraction images recorded using the Dectris EIGER X 16M detector at the DIAMOND beamline I04. The model of the crystal structure and associated information can be found in the Protein Data Bank entry 9EP6. The model has P 31 2 1 symmetry and three molecules per asymmetric unit. This is a case of crystal pathology &ndash; partial disorder. There is electron density for the fourth molecule which could be modelled with occupancy 1/2 and would overlap with a symmetry-related molecule.</p>

opencc-by-4.0Apr 2024View details →
zenodo48/100

Temperature-dependent fold-switching mechanism of the circadian clock protein KaiB

<p>Derived data accompanying publication of&nbsp;<em>Temperature-dependent fold-switching mechanism of the circadian clock protein KaiB</em> (Zhang et al., PNAS 2024).</p> <p>&nbsp;</p> <p>This dataset contains data for fold-switching of KaiB from simulations performed using the Upside coarse-grained model (Jumper et al. PLoS Comput. Bio 2017). Files contained include collective variables, kinetic quantities (committors), and initial structures used to seed unbiased simulations. These data should be sufficient recreate the analysis shown in the associated publicaion. Raw trajectory files have not been deposited due to their size; contact the author (Spencer Guo) to request.</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Inter-Chemical Correlation results for the study: HHEARx2017-1967 (Perfluoroalkyl and Polyfluroalkyl Substances (PFAS), Protein Biomarkers, Adiposity and Cardiometabolic Risk Factors in a 3-year Cohort of Low-Income Latino Children with Overweight and Obesity from the Stanford GOALS Randomized Controlled Trial)

Title: Perfluoroalkyl and Polyfluroalkyl Substances (PFAS), Protein Biomarkers, Adiposity and Cardiometabolic Risk Factors in a 3-year Cohort of Low-Income Latino Children with Overweight and Obesity from the Stanford GOALS Randomized Controlled Trial <br>Species: Homo sapiens <br>Number of samples: 1085 <br>Number of named analytes: 8 <br>Datasource url: https://hheardatacenter.mssm.edu/PublicFile/ViewPublicFile?projectid=36 <br>

opencc-zeroMay 2024View details →
zenodo48/100

LigPCDS: Labeled Dataset of X-ray Protein Ligand Images in 3D Point Cloud and Validated Deep Learning Models

<p>The difference electron density from X-ray protein crystallography was used to create the first dataset of labeled ligand images in 3D point clouds, named <strong>LigPCDS</strong>. The dataset contain 244,226 entries of free organic ligands containing 3D representations labeled with two major labeling approaches: SP-based and AtomSymbol-based.</p> <p>&nbsp;</p> <p>The data from free organic molecules (non-covalent ligands) was retrieved from the Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) in december 2019 with resolutions ranging from 1.5 to 2.2 &Aring;. The ligand images (blobs) were interpolated from their calculated difference electron density map in a 3D grid-like bounding box, around their atomic positions, and stored in point clouds. These ligand grid representations were further processed to retrive the final ligands representation in 3D point clouds using a mask of the shape of the ligand. A grid spacing of 0.5 &Aring; gave the best results. The density value of the grid points was used as feature. The labeling approach used the structure of the ligands to propose vocabularies of chemical classes based on the chemical atoms themselves and their cyclic substructures. These structure annotations were applied pointwise to the ligand 3D representations using an atomic sphere model. Four proposed vocabularies were validated by successfully training good performance deep learning models for the semantic segmentation of a stratified dataset from LigPCDS, using 78902 entries.</p> <p>The four validated deep learning models are: (i) the LigandRegion, composed by generic atoms of any type; (ii) the AtomCycle, composed by generic atoms outside cycles and generic cycles; (iii) the AtomC347CA56, composed by generic atoms outside cycles, not aromatic cycles of size 3 to 7 and aromatic cycles of size 5 and 6; and (iv) the AtomSymbolGroups, composed by the atoms symbols with groupings. The mean accuracy of these models in their cross-validation was between 49.7% <span lang="EN-GB">[-19.4,20.</span><span lang="EN-GB">2]</span> and 77.4% <span lang="EN-GB">[-11.7,12.1]</span> in terms of Intersection over Union (mIoU) metric and between 62.4% <span lang="EN-GB">[-18.8,19.</span><span lang="EN-GB">7]</span> and 87.0% <span lang="EN-GB">[-8.4,8.8]</span> in F1-score (mF1), confidence interval between squared brackets. The models i, ii and iii and the used labeled representations in 3D point cloud are contained in the SP-based record; and model iv and its used labeled representations are contained in the AtomSymbol-based record.</p> <p>The dataset and validated models may be used to tackle problems regarding known and unknown ligand building to drug discovery and fragment screening pipelines.&nbsp;</p> <p>The code used to create and validated the LigPCDS is available at the following repository: https://github.com/danielatrivella/np3_ligand</p> <p>This repository also contains the NP&sup3; Blob Label application for ligand building using the validated deep learning models from LigPCDS.</p>

opencc-by-4.0May 2023View details →
zenodo48/100

Ligand binding remodels protein side chain conformational heterogeneity

<p>While protein conformational heterogeneity plays an important role in many aspects of biological function, including ligand binding, its impact has been difficult to quantify. Macromolecular X-ray diffraction is commonly interpreted with a static structure, but it can provide information on both the anharmonic and harmonic contributions to conformational heterogeneity. Here, through multiconformer modeling of time- and space-averaged electron density, we measure conformational heterogeneity of 743 stringently matched pairs of crystallographic datasets that reflect unbound/apo and ligand-bound/holo states. When comparing the conformational heterogeneity of side chains, we observe that when binding site residues become more rigid upon ligand binding, distant residues tend to become more flexible, especially in non-solvent exposed regions. Among ligand properties, we observe increased protein flexibility as the number of hydrogen bonds decrease and relative hydrophobicity increases. Across a series of 13 inhibitor bound structures of CDK2, we find that conformational heterogeneity is correlated with inhibitor features and identify how conformational changes propagate differences in conformational heterogeneity away from the binding site. Collectively, our findings agree with models emerging from NMR studies suggesting that residual side chain entropy can modulate affinity and point to the need to integrate both static conformational changes and conformational heterogeneity in models of ligand binding.</p>

opencc-by-4.0Sep 2021View details →
zenodo48/100

Slimfield: Escherichia coli DNA repair proteins (RecA-mGFP and RecB-sfGFP)

<p>Imaging modality / instrument: <em>Brightfield</em> + <em>Slimfield</em></p> <p>Image format:<em> OME TIFF (16 bit) + MicroManager metadata files</em></p> <p>Microscope settings:</p> <p><em>488 nm triggered excitation; split detection, cropped to GFP or RFP/GFP (left/right) channels; 3 ms/frame laser exposure; Photometrics Prime95b CMOS</em></p> <p>Samples and acquisitions:</p> <p>Fluorescent fusions in live E.coli cells.&nbsp; MMC = mitomycin C</p> <table> <tbody> <tr> <td> <p>No. fields of view</p> </td> <td> <p>MMC-</p> </td> <td> <p>MMC+ (0.5 ug/ml 3h)</p> </td> </tr> <tr> <td> <p>RecA-mGFP</p> </td> <td> <p>7</p> </td> <td> <p>15</p> </td> </tr> <tr> <td> <p>RecB-sfGFP</p> </td> <td> <p>17</p> </td> <td> <p>21</p> </td> </tr> <tr> <td> <p>MG1655 control</p> </td> <td> <p>10</p> </td> <td> <p>10</p> </td> </tr> </tbody> </table> <p>Approx. size before/after compression: 21 GB / 7 GB</p>

opencc-by-4.0Jun 2022View details →
zenodo48/100

Begomovirus DNA-B Movement Protein and Nuclear Shuttle Protein Ref-Seq Datasets

<p>Multiple Sequence Alignment of proteins encoded on begomovirus DNA-B (movement protein and nuclear shuttle protein; n=131). One isolate per species based on ICTV reference list.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo48/100

Protein structure files for the paper "Multiplexed identification of RAS paralog imbalance as a driver of lung cancer growth" in Nature Cell Biology by Tang et al.

<p>This archive contains models of HRAS, KRAS, and NRAS homo- and heterodimers with various mutations discussed in the paper,&nbsp; &quot;Multiplexed identification of RAS paralog imbalance as a driver of lung cancer growth&quot; in Nature Cell Biology by Tang et al.<br> as well as crystallographic dimers of these proteins as identified by the ProtCAD database, http://dunbrack2.fccc.edu/ProtCAD/Results/PfamArchClusterInfo.aspx?GroupId=8 (cluster 5). Several of the models are shown in Supp. Figure 11b and the crystallographic dimers of RAS that provide evidence for the possible biological relevance of these models are shown in Supp. Figure 11a.</p> <p>The crystallographic dimers were identified by clustering all possible interfaces generated by symmetry operators in crystals of HRAS, KRAS, and NRAS as described in the paper: Xu, Q., Dunbrack, R.L. ProtCID: a data resource for structural information on protein interactions. <em>Nat Commun</em> <strong>11</strong>, 711 (2020). https://doi.org/10.1038/s41467-020-14301-4.</p> <p>The models were created by superposing monomers of HRAS, KRAS, or NRAS onto the alpha4-alpha5 dimer present in the crystal of PDB entry 3k8y. Mutations were made in PyMOL. The structures were relaxed with the FastRelax protocol and the Ref2015 scoring function in the program Rosetta, which uses the backbone-dependent rotamer library of Shapovalov and Dunbrack to repack side chains.</p> <p>The crystallographic dimers are contained in a zipped PyMOL session. The mmCIF format for all the structures is present in a zip file, Tang_et_al_crystallographic_and_modeled_RAS_dimer_ciffiles.zip. The PyMOL session and zip file contains 87 HRAS dimers, 14 KRAS dimers, and 1 NRAS dimer, all having the interface consisting of the alpha4 and alpha5 helices. The PyMOL session also contains the modeled structures. Only Mg ions and GTP/GNP/GDP ligands are shown. Others are present but hidden and may be displayed by PyMOL (&quot;show sticks, het&quot;).</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo48/100

Protein haplotype sequences obtained by ProHap from the Human Pangenome Reference Consotruim dataset

<p>Database of protein sequences obtained using ProHap (<a href="https://github.com/ProGenNo/ProHap">https://github.com/ProGenNo/ProHap</a>) on the data set of phased genotypes published by the Human Pangenome Reference Consotruim (HPRC), first release (<a href="https://github.com/human-pangenomics/hpp_pangenome_resources">https://github.com/human-pangenomics/hpp_pangenome_resources</a>), 44 samples. We used Ensembl v.110 for the mapping of coordinates between genes, exons, and transcripts.</p> <p>This repository contains one database created using all 43 samples of the HPRC release (the haplotypes of the sample NA21309 did not encode any non-canonical sequences), and then a database for each of the 43 samples separately. No filtering on allele frequency or haplotype frequency was applied in any of the databases. The complete configuration file for the ProHap run is attached to this repository.</p> <p>There is one compressed directory for each of the databases, containing the following files:</p> <ul> <li>F1: The concatenated fasta file ready to be used with search engines, contains the following: <ul> <li>Protein haplotype sequences obtained by ProHap</li> <li>Reference proteome as per Ensembl v. 110</li> <li>Contaminant sequences from the cRAP project (<a href="https://www.thegpm.org/crap/">https://www.thegpm.org/crap/</a>)</li> <li>For this dataset, only the simplified format is provided. The simplified fasta contains only the artificial protein identifier and the matching gene name, and is optimised for compatibility with a wide range of tools. For annotation of peptides using the PeptideAnnotator, please provide the header (F1.2) in addition to the fasta file.&nbsp;</li> </ul> </li> <li>F2: Additional information about the haplotype sequences, to be used for mapping identified peptides to the original haplotypes</li> <li>F3: Translations of haplotype cDNA sequences, before merging with the reference proteome</li> </ul> <p>For further description of the files, please refer to&nbsp;<a href="https://github.com/ProGenNo/ProHap/wiki/Output-files">https://github.com/ProGenNo/ProHap/wiki/Output-files</a>.</p> <p>For the usage of these databases with search engines, and downstream anaylsis of identified peptides, please refer to the project's wiki page: <a href="https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches">https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches</a>.</p> <p>When using these databases in your publication, please cite: Va&scaron;&iacute;ček, J., Kuznetsova, K.G., Skiadopoulou, D. <em>et al.</em> ProHap enables human proteomic database generation accounting for population diversity. <em>Nat Methods</em> (2024). <a href="https://doi.org/10.1038/s41592-024-02506-0">https://doi.org/10.1038/s41592-024-02506-0</a></p>

opencc-by-4.0Jul 2024View details →
zenodo48/100

A large data-set of CASP protein refinement simulations for machine-learning

<p>The uploaded trajectory data originates from our own laboratory&#39;s refinement method in CASP11 and CASP12 for which the reference crystal structure is available in the PDB. In total the trajectory data consists of&nbsp; 904 trajectories with 3419 ns cumulative simulation time and 1,709,704 snapshots with a delta t =2 ps from 42 different protein systems.</p> <p><strong>File Overview</strong></p> <ul> <li><strong>trajectory_data_pdbs.tar.gz :</strong> contains the PDB files of the different trajectories as well as the starting model and reference crystal structure for each target</li> <li><strong>casp_normalized_all_data_final.csv.gz :&nbsp; </strong>contains the trajectory features calculated for each snapshot from the trajectory PDBs</li> <li><strong>cv_folds.csv : </strong>contains the 7 fold cross-validation assignment used to assess the performance of the model<br> &nbsp;</li> </ul> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2018View details →
zenodo48/100

Homologous membrane protein structures (HOMEP) dataset version v2

<p><strong>Protein structures from the dataset of&nbsp;Homologous MEmbrane Protein structures (HOMEP)</strong> version v2 created in 2010, published in 2013. A more automated version of HOMEP v1:&nbsp;<a href="https://doi.org/10.5281/zenodo.2646534">10.5281/zenodo.2646534</a><br> &nbsp;</p> <p><strong>Table 1</strong> = List of protein databank&nbsp;structure entries<br> From Stamm et al, PLOS One 2013,&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pubmed/23469223">https://www.ncbi.nlm.nih.gov/pubmed/23469223</a>, Supplementary Table 1, with the following entries:<br> Family grouping, Protein databank identifier, Name, Source organism, Resolution (&Aring;)</p> <p>&nbsp;</p> <p><strong>Table 2</strong> = List of pairs of structures<br> From Stamm et al, PLOS One 2013,&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pubmed/23469223">https://www.ncbi.nlm.nih.gov/pubmed/23469223</a>, Supplementary Table 2, with the following entries:<br> Family grouping, PDB code for first structure, Chain ID from PDB1, PDB for second structure, Chain ID from PDB2, protein structural difference (PSD), % sequence identity</p> <p>&nbsp;</p> <p><strong>File S2 HOMEP2 Dataset.tar.gz</strong> = Protein databank format files (PDB) are attached in the Dataset tar zipped file,&nbsp;organized by family.&nbsp;From Stamm et al, PLOS One 2013,&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pubmed/23469223">https://www.ncbi.nlm.nih.gov/pubmed/23469223</a>, Supplementary dataset.</p>

openother-openMar 2013View details →
zenodo48/100

Datasets and Jupyter notebook for the structural analysis of protein-RNA interface evolution

<p>The present repository contains data and code related to our manuscript "Structural comparison of protein-RNA homologous interfaces reveals widespread overall conservation contrasted with versatility in polar contacts". In the manuscript, we analyze the evolution of protein-RNA interfaces by building a dataset of protein-RNA interologs (homologous interfaces) and exploring how interface contacts are conserved between homologous interfaces, as well as possible explanations for non-conserved contacts.</p> <p>This repository contains the following files:</p> <ul> <li>DataAnalysisNotebook.ipynb is a Jupyter notebook to reproduce contact conservation analysis and all figures from our manuscript, and to explore data</li> <li>env.yaml is an environment file in order to build a Conda/Mamba environment to run the Jupyter notebook&nbsp;</li> <li>2022-02-21-PDB.csv contains data from the PDB about 3D structures of complexes containing interacting protein and RNA chains (PDB structure identifier, chain identifiers, experimental technique and resolution)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.tsv contains more detailed information about interacting protein and RNA chains from these complexes (PDB and chain identifiers, protein and RNA size, interface size and number of contacts)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.txt.selectXE_2.50_p30_r10_pi5_ri5_rep_bc-100.out_RNAcl_0.99.tsv contains the same detailed information, restricted to the filtered dataset used as a starting point in our interolog search pipeline</li> <li>PDBinterfaceAlign.csv contains information about the structural alignment of pairs of protein-RNA interactions (structural alignment TM-scores, sequence identity and coverage)</li> <li>DataInterologsParam.tsv contains information about a pre-filtered set of 2587 potential interologs (including interface RMSD, sequence identity and coverage and interface size)</li> <li>DataInterologsContactsFixedSASA.tsv contains detailed information about conserved and non-conserved contacts in the final set of 2022 interologs (atomic contacts, apolar contacts, hydrogen bonds, salt bridges and stacking information for aminoacid-nucleotide pairs, as well as information about whether each belongs to the interface, secondary structures, and the aminoacid surface accessibility and evolutionary conservation metrics) - compared to version 1, the calculation of solvent accessibility was fixed for a number of interolog pairs</li> <li>DataCons.csv contains precomputed contact conservation metrics for each of the 2022 interolog pairs, for fast reproduction of manuscript figures</li> <li>DataInterologsContactsResampledMaintainStructSeqId.tsv, DataInterologsContactsShuffled.tsv and DataInterologsShuffled.tsv relate to baselines computed for contact conservation assessment</li> <li>clan.txt, clan_membership.txt, ecod.latest.domains.uniq.txt, rfam_interfaces_977.txt, DataGroupsECOD.tsv, DataGroupesRFAM.tsv, DataGroupsRFAMClan.tsv, DataInterfaceGroupsECOD.tsv and DataInterfaceGroupsRFAM.tsv relate to the ECOD (respectively Rfam) classification of protein domains (respectively RNA) in protein-RNA interfaces from our dataset</li> <li>ListeIntraHbonds.pkl and ListeIntraSaltBridges.pkl are pickle-format data files containing intra-molecular hydrogen bonds and salt bridges (respectively) that are used to analyse scenarii of compensation for non-conserved polar contacts.</li> </ul>

opencc-by-4.0May 2024View details →
zenodo48/100

LukProt - an animal evolution-centric eukaryotic protein database

<p>LukProt is the EukProt database with additional species added, mostly the undersampled animal and some holozoan taxa.&nbsp;The database is composed of sequences translated from annotated genomes, transcriptomes or ESTs. <strong>The main purposes of the database are to consolidate sequences from undersampled animal taxa</strong> and provide usable search tools. The publication associated with LukProt can be found here: <a href="https://doi.org/10.1093/gbe/evae231">https://doi.org/10.1093/gbe/evae231</a>.</p> <p>The current version of the database (v1.5.1) is based on <a href="https://doi.org/10.24072/pcjournal.173">EukProt v3</a>. The home of all public versions of LukProt is this page (Zenodo).</p> <p>Proteomes that are novel in LukProt are denoted as LPXXXXX and those coming from AniProtDB are called APXXXXX. The sequence IDs from EukProt are conserved in LukProt. This means that each sequence is assigned an ID in the following format:</p> <pre><code>(A/E/L)PXXXXX_Species_epithet_(strain)_PYYYYYY</code></pre> <p>where XXXXX is a number from 00001 to 99999 and YYYYYY is a number from 000001 to 999999. Each sequence is assigned a unique number YYYYYY, and each taxon XXXXXX. All the IDs are compatible with BLAST v5 "-parse_seqids" option and the database can be readily deployed, for example on a server running <a href="https://doi.org/10.1093/molbev/msz185">SequenceServer</a>. Within each of the source fasta files, the source sequence identifier was kept after a blank space, so that it can still be retrieved if needed.</p> <p>A publicly available BLAST server providing LukProt search is available at: <a title="LukProt BLAST server" href="https://lukprot.hirszfeld.pl/" target="_blank" rel="noopener">https://lukprot.hirszfeld.pl/</a>.</p> <p>Comparison of EukProt v2/v3, LukProt 1.4.1 and LukProt v1.5.1 in their main areas of difference:</p> <table> <tbody> <tr> <th>Taxogroup</th> <th>EukProt v2</th> <th>EukProt v3</th> <th>LukProt v1.4.1</th> <th>LukProt v1.5.1</th> </tr> <tr> <th> <p>Holozoa</p> <p>(excluding Metazoa)</p> </th> <td>31</td> <td>40</td> <td>39</td> <td>43</td> </tr> <tr> <th>Ctenophora</th> <td>2</td> <td>2</td> <td>35</td> <td>38</td> </tr> <tr> <th>Porifera</th> <td>4</td> <td>5</td> <td>30</td> <td>47</td> </tr> <tr> <th>Placozoa</th> <td>2</td> <td>2</td> <td>3</td> <td>6</td> </tr> <tr> <th>Cnidaria</th> <td>3</td> <td>5</td> <td>65</td> <td>88</td> </tr> <tr> <th>Bilateria</th> <td>51</td> <td>51</td> <td>94</td> <td>142</td> </tr> </tbody> </table> <p>Included with the database are:</p> <ul> <li>ready to use main database files: <ul> <li><em>LukProt_v1.5.1_single_species_FASTA.7z</em> &ndash; a FASTA file with the sequences - <a href="https://en.wikipedia.org/wiki/7z">7-zipped</a>, <strong>uncompressed size: 17.6 GB</strong><br> <ul> <li>to concatenate all into one file, run this in the parent directory: <code>for file in $(find . -type f -name "*.fasta"); do awk 'FNR==1{print ""}1' $file &gt;&gt; LukProt_v1.5.1.fa; done</code>. This will create single FASTA file with all the sequences in the parent directory. <code>awk</code> is used to insert a new line after every file because&nbsp;<code>cat</code> would sometimes merge the last sequence with the header of the first sequence.</li> </ul> </li> <li><em>LukProt_v1.5.1_full_BLAST_db.7z</em> &ndash; a preformatted, full BLAST database (NCBI BLAST database format version: v5, masked with segmasker), <strong>uncompressed size: 28.3 GB</strong></li> <li><em>LukProt_v1.5.1_taxogroup_BLAST_db.7z</em> &ndash; a collection of BLAST databases where each proteome is one taxogroup and is placed within the eukaryotic tree of life directory structure, <strong>uncompressed size: 26.3 GB</strong></li> <li><em>LukProt_v1.5.1_single_species_BLAST_db.7z</em> &ndash; a collection of BLAST databases where each proteome is one BLAST database and is placed within the eukaryotic tree of life directory structure, <strong>uncompressed size: 26.4 GB</strong></li> </ul> </li> <li>auxiliary database files: <ul> <li><em>LukProt_v1.5.1.cdhit70.7z</em> &ndash; the full database clustered at 70% identity using CD-HIT with the following command: <code>cd-hit -g 1 -d 0 -T 20 -M 90000 -c 0.7 -uL 0.2 -uS 0.9 -s 0.2</code>,&nbsp;<strong>uncompressed sizes: fasta file - 11 GB, clstr file - 2.5 GB</strong></li> <li><em>LukProt_IDs_mapped.txt.gz</em> &ndash; a text file mapping the LukProt IDs to the AniProtDB IDs and EukProt IDs that are different</li> <li><em>BUSCO_tables.ods</em> &ndash; a spreadsheet with full result tables generated by BUSCO analysis</li> <li><em>OMAmer_output.zip</em> &ndash; a folder with full results of OMAmer analyses (includes per-sequence taxonomy classification)</li> <li><em>OMArk_output.zip</em> &ndash; a folder with the results of all OMArk analyses</li> </ul> </li> <li>metadata: <ul> <li><em>README.md</em> &ndash; a README file describing the metadata</li> <li><strong><em>LukProt_metadata_sheet.ods</em> &ndash; main metadata file. A spreadsheet with information about each proteome (in an open .ods format, most compatible with <a href="https://www.libreoffice.org/">LibreOffice</a>)</strong></li> <li><em>LukProt_metadata_other.zip</em> &ndash; an archive with other metadata files, documented in the README. Contents include:<br> <ul> <li>the LukProt taxonomy in various formats</li> <li>supporting scripts for data manipulation and visualization</li> </ul> </li> <li>a recoloring script (modified by LFS, originally by Dr. Celine Petitjean). The script is in&nbsp;<a title="formatFigtree2" href="https://doi.org/10.5281/zenodo.10654583">public domain</a> and reuploaded here only for convenience.&nbsp;</li> <li>other files - see README</li> </ul> </li> <li><em>changelog.md</em> &ndash; database changelog</li> </ul> <p>Words of caution:</p> <ul> <li>The database has been synchronized to EukProt v3 in version v1.5.1. This means that identifiers were modified in comparison to LukProt v1.4.1. The convention is not expected to change any more in future updates.</li> <li>Many proteomes, especially those transcriptome-based, may contain contamination from different species. In addition, the translation algorithms often introduce errors (e.g. the transcript may not represent a full length protein). For this reason, to get accurate sequences from each organism, users are directed to source data and to the included OMAmer, OMArk and BUSCO data for details.</li> <li>The taxonomy is different to UniEuk/EukMap, but UniEuk data were integrated where possible.</li> <li>A few NCBI taxids are missing and will be added in due course.</li> <li>Proteomes from NCBI and UniProt will be updated to current versions.</li> <li>A number of proteomes present in some metadata, are unpublished and were held back.</li> <li>While the database contains metadata that present a particular phylogeny of animals, holozoans and other eukaryotes, no particular claims or hypotheses are made by the author(s). However, in the future efforts will be made to name clades officially, once they are more firmly established.</li> </ul> <p><strong>Please report any problems or suggestions to Lukasz Sobala: lukasz.sobala (at) hirszfeld.pl.</strong></p> <p>&nbsp;</p> <p>Acknowledgements:</p> <ul> <li> <p>Andrew E. Allen Lab for creating the original <a href="https://allenlab.ucsd.edu/data/" target="_blank" rel="noopener">PhyloDB</a>.</p> </li> <li> <p>Daniel Richter <em>et al.</em> for creating <a href="https://doi.org/10.6084/m9.figshare.12417881">EukProt</a> and keeping it updated.</p> </li> <li> <p>Members of <a href="https://multicellgenome.com/">the Multicellgenome Lab</a>, especially Michelle Leger (for donating her database), for the bioinformatics support and for doing great science.</p> </li> <li> <p>All the authors of the original data.</p> </li> <li> <p>National Science Centre of Poland for funding of the project 2020/36/C/NZ8/00081, "The role of glycosylation in the emergence of animal multicellularity", which enabled the creation of this database.</p> </li> </ul>

opencc-by-4.0Sep 2022View details →
zenodo48/100

Proteins required for stereocilia elongation during mammalian hair cell development ensure precise and steady heights during adult life

<p>This dataset contains all source data for Hartig <em>et al </em>2024, PNAS, including:</p> <p>Data files</p> <p>Raw images and TDT ABR/DPOAE files</p> <p>ROIS and raw measurements from quantifications in ImageJ</p> <p>R scripts for data visualization and statistics</p> <p>Reports of statistical analyses including diagnostic qq plots and distributions</p>

opencc-by-4.0Aug 2024View details →
zenodo48/100

Data for the publication: Recombinant silk protein condensates show widely different properties depending on the sample background

<p>This entry includes raw data for the publication "Recombinant silk protein condensates show widely different properties depending on the sample background". The original publication was published in: Journal of Materials Chemistry B, DOI: 10.1039/d4tb01422g</p> <p>The folder "Videos_Micropipette_Aspiration_Zenodo.zip" contains 9 TIF files, labeled Number1 - Number9. The numbering corresponds to the numbering of IMAC condensates studied with micropipette aspiration in the publication. Each TIF file is an image stack from a time series.</p> <p>The folders "Videos_IMAC_silk_with_BG_lysate_coalescence.zip", "Videos_HT_silk_coalescence.zip", and "Videos_IMAC_silk_coalescence.zip" all contain subfolders labeled with the purification method, the framerate of the videos and then consecutive numbering. Each of these folders contains the frames of the video as single TIF files.</p> <p>Please find more information in the read_me file uploaded.</p>

opencc-by-4.0Oct 2024View details →
zenodo48/100

Identification and characterization of the cell division protein MapZ of Streptococcus suis

<p>Supplementary data and code related to the manuscript &quot;Identification and characterization of the cell division protein MapZ of <em>Streptococcus suis</em>&quot;.</p>

opencc-by-4.0Aug 2021View details →
zenodo48/100

Data in: Aging power spectrum of membrane protein transport and other subordinated random walks

<p>Datasets generated in the report &quot;Aging power spectrum of membrane protein transport and other subordinated random walks&quot;. Included data are:</p> <p><strong>Numerical simulations&nbsp;</strong><br> RWdata1.mat: 10,000 realizations, subordinated random walk with Hurst exponent, <em>H</em>=0.3&nbsp;and <span class="math-tex">\(\alpha\)</span>=0.4.<br> RWdata3.mat:&nbsp;10,000 realizations, subordinated random walk with Hurst exponent, <em>H</em>=0.7&nbsp;and <span class="math-tex">\(\alpha\)</span>=0.4.<br> RWdata8.mat:&nbsp;5,000 realizations, subordinated random walk with Hurst exponent, <em>H</em>=0.75&nbsp;and <span class="math-tex">\(\alpha\)</span>=0.8.<br> RWdataCTRW.mat:&nbsp;10,000 realizations, continuous time random walk (CTRW),&nbsp;<span class="math-tex">\(\alpha\)</span>=0.7.</p> <p><strong>Spectra of&nbsp;simulations</strong><br> PSDdata1.mat: Power spectral density (PSD) of a subordinated random walk with Hurst exponent, <em>H</em>=0.3&nbsp;and <span class="math-tex">\(\alpha\)</span>=0.4. Five different realization times are used to compute the PDS: 2^8,&nbsp;2^10,&nbsp;2^12,&nbsp;2^14, and 2^16.<br> PSDdata3.mat:&nbsp;PSD of a subordinated random walk with Hurst exponent, <em>H</em>=0.7&nbsp;and <span class="math-tex">\(\alpha\)</span>=0.4. Five different realization times are used to compute the PDS: 2^8,&nbsp;2^10,&nbsp;2^12,&nbsp;2^14, and 2^16.<br> PSDdata8.mat: PSD of a&nbsp;subordinated random walk with Hurst exponent, <em>H</em>=0.75&nbsp;and <span class="math-tex">\(\alpha\)</span>=0.8.&nbsp;Four&nbsp;different realization times are used to compute the PDS: 2^15,&nbsp;2^16,&nbsp;2^17, and 2^18.<br> PSDs_CTRW.mat: PSD of a&nbsp;continuous-time random walk (CTRW),&nbsp;<span class="math-tex">\(\alpha\)</span>=0.7. Five different realization times are used to compute the PDS: 2^8,&nbsp;2^10,&nbsp;2^12,&nbsp;2^14, and 2^16.</p> <p><strong>Experimental data of Nav1.6 channels in the soma of hippocampal neurons</strong><br> NavMSDtimes.csv: ensemble-averaged (EA) MSD and time-averaged (TA) MSD. The TA-MSD is measured&nbsp;for three observation times, 64, 128, and 256 frames (3.2, 6.4, and 12.8 s).<br> NavPSD.csv: Power spectral density (PSD) measured for&nbsp;three observation times, 64, 128, and 256 frames.</p>

opencc-by-4.0Sep 2021View details →
zenodo48/100

Discoba protein sequences for protein structure predictions

<p>Comprehensive database of Discoba protein sequences, gathered for the purpose of improving protein structure predictions of Discoba species (including <em>Trypanosoma </em>and <em>Leishmania</em>) by AlphaFold and RoseTTAFold. Originally gathered for use with:&nbsp;https://github.com/zephyris/discoba_alphafold</p>

opencc-by-4.0Oct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record