Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
295
datasets available to search
ShareScore release 0.9.0
Dataset results
295 results for “structure prediction”
Cross-phyla protein annotation by structural prediction and alignment
<p><strong>Background:</strong> Protein annotation is a major goal in molecular biology, yet experimentally determined knowledge is typically limited to a few model organisms. In non-model species, the sequence-based prediction of gene orthology can be used to infer protein identity, however this approach loses predictive power at longer evolutionary distances. Here we propose a workflow for protein annotation using structural similarity, exploiting the fact that similar protein structures often reflect homology and are more conserved than protein sequences.</p> <p><strong>Results:</strong> We propose a workflow of openly available tools for the functional annotation of proteins via structural similarity (MorF: <strong>Mor</strong>pholog<strong>F</strong>inder) and use it to annotate the complete proteome of a sponge. Sponges are highly relevant for inferring the early history of animals, yet their proteomes remain sparsely annotated. MorF accurately predicts the functions of proteins with known homology in >90% cases, and annotates an additional 50% of the proteome beyond standard sequence-based methods. We uncover new functions for sponge cell types, including extensive FGF, TGF and Ephrin signalling in sponge epithelia, and redox metabolism and control in myopeptidocytes. Notably, we also annotate genes specific to the enigmatic sponge mesocytes, proposing they function to digest cell walls.</p> <p><strong>Conclusions:</strong> Our work demonstrates that structural similarity is a powerful approach that complements and extends sequence similarity searches to identify homologous proteins over long evolutionary distances. We anticipate this to be a powerful approach that boosts discovery in numerous -omics datasets, especially for non-model organisms.</p>
Chemical structures, Cell Painting and transcriptional profiles for compound bioactivity prediction.
<p>This is the related data, both input and produced for the paper <a href="https://doi.org/10.1101/2020.12.15.422887">"Predicting compound activity from phenotypic profiles and chemical structures"</a>.</p> <p>This data can be merged with <a href="https://github.com/CaicedoLab/2023_Moshkov_NatComm">paper's GitHub repository</a> for reproduction.</p> <p>Folders and files and are described below:</p> <pre><code>├── assay_data ├── assay_matrix_discrete_270_assays.csv Assay matrix with hits for assays (270) and compounds (16170). Note that this is the final file that we used to produce splits. ├── assay_metadata.csv Assay metadata ├── broad_ids.txt List of broad ids used in this study. That is an unfiltered list of compounds required by some analysis scripts. ├── smiles.txt Same as broad_ids.txt, but SMILES strings. ├── feature_data (for 16978 compounds, can be masked with ./misc/compounds16978to16170.npy) ├── cp.npz Classical chemical features ├── ge.npz Gene expression features ├── ge_scale.npz Gene expression scaled features ├── mo.npz Morphology features (not batch corrected) ├── mobc.npz Morphology features (batch corrected) ├── misc ├── compound_analysis.npz Compounds in the dataset identified as PAINS ├── compounds16978to16170.npy Used to filter features from the bigger set of compounds to the final one ├── fingerprints.npz Calculated fingerprints of compounds, those were then used to calculate similarity ├── similarity_fingerprints.npz Similarity matrix for compounds (16978) ├── population_normalized.csv.gz Well-level morphological profiles that were used for batch-correction ├── Table for PUMA Excel file with additional data and plots ├── predictions ├── scaffold_median(mean)_AUC.csv Aggregated median(mean) AUC scores over scaffold-based cross-validation splits. In the paper, median results were reported. ├── scaffold_median(mean)_EF.csv Aggregated median(mean) enrichment factor (EF) over scaffold-based cross-validation splits. In the paper, median results were reported. ├── toprank_chemical_cv{}_hitsnorm.csv Those files are needed to create enrichment plots and contain hit rate and top rank hit rate. ├── Each folder here stands for an experiment type, the number in the folder name is a number of the split. Inside each folder there are the following elements: ├── predictions Folder with predictions for each assay-compound pair for each modality ├── 2022_01_evaluation_all_data.csv File with AUC scores for each assay for the test set in the split ├── 2022_01_evaluation_all_data_EF.csv File with enrichment factor (EF) values for each assay for the test set in the split. Those files exist only for *chemical* folders. ├── assay_matrix_discrete_train(test)_old_scaff.csv Training and test subsets of data for the split. The first column contains broad_id. ├── assay_matrix_discrete_train(test)_old_scaff.csv Same, but SMILES strings in the first column. Those files are used as input to ChemProp! Experiments in this folder are the following: - chemical Scaffold-based 5-fold cross-validation splits, the main results in the paper are reported with this series of experiments. - chemical_bal Same splits as in chemical, but training were run with ChemProp built-in data balancing. - chemical_st Same splits as in chemical, but separate models were trained for each assay. - CV Random 5-fold cross-validation splits. - GE 5-fold cross-validation splits based on same-size clustering of gene expression features. - MOBC 5-fold cross-validation splits based on same-size clustering of batch-corrected morphology features. - random 10 random splits, ~80% of compounds in the training set and the rest in the test set. ├── splitting This folder contains numpy files which help to match compounds and features to create training and test sets for a split, which can be reused in the analysis notebook for data preparation. ├── scaffold_based_split.npz Splitting for scaffold-based splits. ├── random_split_{}.npz Random split indices of test set compounds (10 files). ├── cross_validation_indicies.npz Indices for random cross-validation splits ├── GE_clusters_size_constrained.npz Indicies of clusters of same-size clustering for gene-expression features. ├── MOBC_clusters_size_constrained.npz Indices of clusters of same-size clustering for batch-corrected morphology features.</code></pre> <p> </p>
Discoba protein sequences for protein structure predictions
<p>Comprehensive database of Discoba protein sequences, gathered for the purpose of improving protein structure predictions of Discoba species (including <em>Trypanosoma </em>and <em>Leishmania</em>) by AlphaFold and RoseTTAFold. Originally gathered for use with: https://github.com/zephyris/discoba_alphafold</p>
Datasets of sequences, alignments and structural models generated for the structural prediction of complexes mediated by intrinsically disordered regions.
<p>This repository contains input and ouput files used and generated for the scanning of intrinsically disordered region and the prediction of their binding sites to receptor proteins using the <a href="https://github.com/i2bc/SCAN_IDR">SCAN_IDR</a> pipeline with AlphaFold2-Multimer.</p><p>It contains two archives: </p><ol><li><a href="https://zenodo.org/api/records/10068949/draft/files/scanidr_data_repository_corr6J08.tar/content"><i><strong>scanidr_data_repository_corr6J08.tar</strong></i></a> dedicated to the analysis of a dataset of 42 protein complexes non redundant with the dataset used for AlphaFold2 training,</li><li><a href="https://zenodo.org/api/records/10068949/draft/files/923_elm_cases_repository.tar.gz/content"><i><strong>923_elm_cases_repository.tar.gz</strong></i></a> dedicated to the analysis of 923 complexes from the ELM database.</li></ol><p>These data can be used to rerun specific sections of the pipeline and scripts provided in: <a href="https://github.com/i2bc/SCAN_IDR">https://github.com/i2bc/SCAN_IDR</a></p><h4><strong>Dataset of 42 non redundant complexes</strong></h4><p>The first archive <a href="https://zenodo.org/api/records/10068949/draft/files/scanidr_data_repository_corr6J08.tar/content"><i><strong>scanidr_data_repository_corr6J08.tar</strong></i></a> contains 3 compressed directories and a README file detailing their contents :</p><ul><li>the initial raw sequence and alignment data for every chain -> DIRECTORY <strong>fasta_msa/</strong></li><li>the input and output data of every Alphafold run for every complex -> DIRECTORY <strong>af2_runs/</strong></li><li>the native reference structures -> DIRECTORY <strong>ref_capri_curated/</strong></li></ul><p>The protein-peptide complex cases have been assigned a distinct index number, from 1 to 42, consistent across the several directories of the archive. Their corresponding directories are labelled as <i><index>_<pdbcode></i>.</p><p><i>The models in this archive were generated using AlphaFold2-Multimer v2.2</i></p><h4><strong>Dataset of 923 complexes selected from the ELM database</strong></h4><p>The second archive <a href="https://zenodo.org/api/records/10068949/draft/files/923_elm_cases_repository.tar.gz/content"><i><strong>923_elm_cases_repository.tar.gz</strong></i></a> contains input and ouput files used and generated for the analysis of 923 Eukaryotic Linear Motifs (ELM) database entries.</p><p>Each ELM entry is indexed with specific integer id and is composed of a receptor and a ligand protein. </p><p>The archive contains a Table associating ELM indexes with the ELM entry information, 5 directories and a README file detailing their contents:</p><ul><li>the table describing ELM entries -> FILE <strong>Table_923ELM_uid_delimitations_info_for_archive.txt</strong></li><li>the initial raw sequence and multiple sequence alignment (MSA) data for every chain -> DIRECTORY <strong>fasta_msa/</strong></li><li>the concatenated MSA model for every ELM complex and protocol used -> DIRECTORY <strong>af2_elm_coali_inputs/</strong></li><li>the best model of every AF2 protocol for every complex according to the AF2 -> DIRECTORY <strong>af2_elm_models/</strong></li><li>the best model cut in the ligand part to select only the ELM motifs as used for the evaluation of the models -> DIRECTORY <strong>elm_cut_models/</strong></li><li>the reference structures used for the evaluation of the models -> DIRECTORY <strong>ref_capri_curated/</strong></li></ul><p><i>The models in this archive were generated using AlphaFold2-Multimer v2.3</i></p>
Robust Method for Property Prediction via Artificial Neural Networks: Incorporating Key Structural Features for Carbon Dioxide – Ionic Liquid Mixtures
<p>This Dataset comprises two sub-sets of information:</p> <ul> <li>Database and Results of the work present in the paper "Robust Method for Property Prediction via Artificial Neural Networks: Incorporating Key Structural Features for Carbon Dioxide – Ionic Liquid Mixtures" published in The Journal of Physical Chemistry B (https://doi.org/10.1021/acs.jpcb.4c04432).</li> <li>Sample of the code used, in order to reproduce any of the results presented above. This can be found in the previous version of this Dataset (v1.0 https://zenodo.org/records/11216901)</li> </ul> <p> </p> <p>Regarding the sample code, an example for all ANN Models used in this work is provided. This includes the three models used:</p> <ol> <li>One based only on Critical Properties of Ionic Liquids (CRT Model)</li> <li>One based only on Structural Properties of Ionic Liquids (STR Model)</li> <li>One combination of the previous models, taking into account both Critical and Structural Properties (COMB Model)</li> </ol> <p>In this manner, it is possible to observe the differences between the performance of the different models, either through statiscal analysis or using graphical representation. This allows for the benchmarking to be done in a more concise way.</p>
KiSSim: Predicting off-targets from structural similarities in the kinome
<p><strong>KiSSim: Predicting off-targets from structural similarities in the kinome</strong></p> <p><strong>Project description.</strong></p> <p>KiSSim (Kinase Structural Similarity) is a novel fingerprint designed specifically for kinase pockets, allowing for similarity studies across the structurally covered kinome. The kinase fingerprint is based on the <a href="https://klifs.net/">KLIFS</a> pocket alignment, which defines 85 pocket residues for all kinase structures. This enables a residue-by-residue comparison without a computationally expensive alignment step.</p> <p>The pocket fingerprint encodes each pocket residue’s spatial and physicochemical properties. The spatial properties describe the residue’s position in relation to the kinase pocket center and important kinase subpockets, i.e. the hinge region, the DFG region, and the front pocket. The physicochemical properties encompass for each residue its size and pharmacophoric features, solvent exposure, and side chain orientation.</p> <p>Some datasets are not part of the `kissim_app` GitHub repository due to their size but can be downloaded from here to the respective kissim_app folders.</p> <p><strong>Data.</strong></p> <ul> <li>`20210902_KLIFS_HUMAN.tar.gz` --- save in `kissim_app/data/external/structures`</li> <li>`complete_SiteAlign.txt.gz` --- save in `kissim_app/data/external/sitealign`</li> </ul> <p><strong>Results.</strong></p> <ul> <li>`results.tar.bz2`--- save as `kissim_app/results`</li> </ul> <p>These are the KiSSim results: fingerprints, feature/fingerprint distances, kinase matrices, and kinase trees for structures in all (`all`), DFG-in (`dfg_in`), and DFG-out (`dfg_out`) conformation. In the case of the DFG-in conformation, we also have KiSSim runs with fingerprint subsets based on only residues that interact with certain ligands in KLIFS IFPs: Erlotinib (`dfg_in_IRE`), Imatinib (`dfg_in_STI`), Bosutinib (`dfg_in_DB8`), and Dopamapimod (`dfg_in_B96`). The folder contains README with a detailed file list.</p> <p><strong>Usage.</strong></p> <p>This dataset can be used to run the notebooks available on <a href="https://github.com/volkamerlab/kissim_app">https://github.com/volkamerlab/kissim_app</a>.</p> <ol> <li>Clone the kissim_app repository.</li> <li>Download the files provided here.</li> <li>If applicable, extract the archive content to the folders as indicated above and run the notebooks.</li> </ol> <pre><code class="language-bash">cd /path/to/your/download tar -xvf results.tar.bz2 -C /path/to/kissim_app/ tar -xvf 20210902_KLIFS_HUMAN.tar.bz2 -C /path/to/kissim_app/data/external/structures/ # In case you want the raw SiteAlign data mv complete_SiteAlign.txt.gz /path/to/kissim_app/data/external/sitealign</code></pre> <p><strong>Citation.</strong></p> <p>These datasets are part of the KiSSim publication: TBA</p>
19th Century United States Newspaper images predicted as Photographs with labels for "human", "animal", "human-structure" and "landscape"
<p>The Dataset contains images derived from the Newspaper Navigator (news-navigator.labs.loc.gov/), a dataset of images drawn from the Library of Congress Chronicling America collection (<a href="https://chroniclingamerica.loc.gov/">chroniclingamerica.loc.gov/</a>). </p> <blockquote> <p>[The Newspaper Navigator dataset] consists of extracted visual content for 16,358,041 historic newspaper pages in <em>Chronicling America</em>. The visual content was identified using an object detection model trained on annotations of World War 1-era Chronicling America pages, including annotations made by volunteers as part of the <a href="https://labs.loc.gov/work/experiments/beyond-words/">Beyond Words</a> crowdsourcing project.</p> <p>source:<a href="https://news-navigator.labs.loc.gov/"> https://news-navigator.labs.loc.gov/</a></p> </blockquote> <p>One of these categories is 'photographs'. This dataset contains a sample of these images with additional labels indicating if the photograph has one or more of the following labels: "human", "animal", "human-structure" and "landscape"</p> <p>The data is organised as follows:</p> <ul> <li>The images themselves can be found in `images.zip`</li> <li>`newspaper-navigator-sample-metadata.csv` contains metadata about each image drawn from the Newspaper Navigator Dataset.</li> <li>`multi_label.csv` contains the labels for the images as a CSV file</li> <li>`annotations.csv` conains the labels for the images with additional metadata</li> </ul> <p>This dataset was created for use in an under-review Programming Historian tutorial (<a href="http://programminghistorian.github.io/ph-submissions/lessons/computer-vision-deep-learning-pt2">http://programminghistorian.github.io/ph-submissions/lessons/computer-vision-deep-learning-pt2</a>) The primary aim of the data was to provide a realistic example dataset for teaching computer vision for working with digitised heritage material. The data is shared here since it may be useful for others. <strong>This data documentation is a work in progress and will be updated when the Programming Historian tutorial is released publicly. </strong></p> <p>The metadata CSV file contains the following columns:</p> <p>- filepath<br> - pub_date<br> - page_seq_num<br> - edition_seq_num<br> - batch<br> - lccn<br> - box<br> - score<br> - ocr<br> - place_of_publication<br> - geographic_coverage<br> - name<br> - publisher<br> - url<br> - page_url<br> - month<br> - year<br> - iiif_url</p>
Datasets for "The Venturia inaequalis effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins "
<p>Datasets for preprint entitled "The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi"</p> <p><strong>1) ViAnnotation.gff3</strong><br> Gene annotation of <em>Venturia inaequalis</em> MNH120 (<a href="https://genome.jgi.doe.gov/Venin1/Venin1.home.html">https://genome.jgi.doe.gov/Venin1/Venin1.home.html</a>) generated as part of the study "The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi". </p> <p>Gene reannotation was performed to include genes that would have been missed in the previous annotation by Deng et al. (2017), especially those genes encoding putative effector proteins, which are difficult to predict. For this purpose, we used a three-step approach. In the first step, coding sequences (CDSs) from <em>V. inaequalis</em> isolate 05/172, which were predicted as part of a previous study by Passey et al. (2018) (<a href="https://journals.asm.org/doi/full/10.1128/MRA.01062-18">https://journals.asm.org/doi/full/10.1128/MRA.01062-18</a>), were downloaded from the National Center for Biotechnology Information (<a href="https://www.ncbi.nlm.nih.gov/nuccore/QFBF00000000.1/">https://www.ncbi.nlm.nih.gov/nuccore/QFBF00000000.1/</a>) and mapped to the MNH120 genome using GMAP v2021-02-22. In the second step, RNA-seq reads from one biological replicate representing each <em>in planta</em> time point of <em>Malus domestica</em> infection by <em>V. inaequalis </em>(12 hour post-inoculation [hpi], 24 hpi, 2 days post-inoculation [dpi], 3 dpi, 5 dpi, 7 dpi), as well as one time point representing growth of the fungus in culture, were mapped to the MNH120 genome using HISAT2 v2.2.1. Then, a genome-guided <em>de novo</em> transcriptome assembly was performed using Trinity v2.12.0 and likely CDSs were identified using Transdecoder v5.5.0 (<a href="https://github.com/TransDecoder/TransDecoder">https://github.com/TransDecoder/TransDecoder</a>) in conjunction with a minimum open frame (ORF) length of 50 amino acids. Finally, in the third step, all annotations were visualized in Geneious v9.05, together with the previous annotation from Deng et al. (2017), and a manual curation was performed to create a consensus prediction. Note: this reannotation was generated with the aim of identifying as many genes as possible, and as a result, it contains many spurious genes. </p> <p><strong>2) Protein_sequences_ViAnnotation.fasta</strong></p> <p><strong>3) ECs_Families_AlphaFold.zip</strong></p> <p>This dataset is made up of predicted protein tertiary structures representing the main member of each up-regulated <em>V. inaequalis</em> effector candidate family. Structures were predicted using Alphafold with the ColabFold server (<a href="https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n">https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n</a>). In cases where the effector candidate had less than 30 proteins with amino acid sequence similarity in the NCBI database, a custom multiple sequence alignment (MSA) was generated and used as input for AlphaFold2. Here, mature protein sequences were used.</p> <p><strong>4) singletons_AlphaFold_OpenSourceCASP14.zip</strong></p> <p>This dataset set is made up of predicted protein tertiary structures representing up-regulated<em> V. inaequalis</em> singleton effector candidates. Structures were predicted using AlphaFold (<a href="https://github.com/deepmind/alphafold">https://github.com/deepmind/alphafold</a>) open source code v2.0.1 and v2.1.0, with pre-set casp14, max_template_date: 2020-05-14. Mature protein sequences were used as input. </p> <p><strong>5) ECs_Avrs_phytopathogens_AlphaFold.zip</strong></p> <p>Predicted tertiary structures of avirulence (Avr) proteins or candidate Avr proteins from other fungal pathogens included in the "The <em>Venturia inaequalis</em> effector repertoire is expressed in waves, and is dominated by expanded families with predicted structural similarity to avirulence proteins from other fungi" study. These structures were predicted using Alphafold with the ColabFold server (<a href="https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n">https://colab.research.google.com/github/sokrypton/ColabFold/blob/main/beta/AlphaFold2_advanced.ipynb#scrollTo=rowN0bVYLe9n</a>). Mature protein sequences were used as input. </p> <p>If you have any questions about the datasets, please contact us.<br> Mercedes Rocafort: <a href="mailto:m.rocafort.ferrer@massey.ac.nz">m.rocafort.ferrer@massey.ac.nz</a><br> Carl Mesarich: <a href="mailto:c.mesarich@massey.ac.nz">c.mesarich@massey.ac.nz</a></p>
DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction
<p>This dataset contains replication data for the paper titled "DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction". The dataset consists of pickled Pandas DataFrame files, along with training, validation, and (for DB5-Plus) test filename lists for cross-validation, that can be used to develop and evaluate protein interface prediction models. This dataset also contains the externally generated residue-level PSAIA and HH-suite3 features for users' convenience (e.g. raw MSAs and profile HMMs for each protein complex). Our GitHub repository linked in the "Additional notes" metadata section below provides more details on how we parsed through these files to create our cross-validation datasets. The GitHub repository for DIPS-Plus also includes scripts that can be used to impute missing feature values and convert the final "raw" complexes into DGL-compatible graph objects. Since our final DGL graph representation for each complex uses PyTorch tensors in its construction of residue embeddings, the final representation of each complex can easily be adapted to fit the users' needs (e.g. feeding a complex's 2D residue feature tensors into a convolutional neural network).</p>
Research data supporting "Tin phosphide anodes for potassium-ion batteries: insights from crystal structure prediction"
<p>This dataset contains the output files of crystal structure prediction calculations (density-functional theory relaxations, bandstructures, phonon calculations, GIPAW-NMR calculations) on the ternary K-Sn-P phase diagram. All calculations were performed with the CASTEP DFT package (https://www.castep.org/) and the "matador" Python library (https://github.com/ml-evs/matador).</p> <p><strong>Contents:</strong></p> <ul> <li>"convergence_tests.zip": contains the results of convergence tests on the K-P system at two levels of accuracy "polish" and "searches" on the corresponding edge of the K-Sn-P ternary system</li> <li>"phonons.zip": contains CASTEP output files for phonon calculations on the predicted low-lying phases on the corresponding edge of the K-Sn-P phase diagram</li> <li>"polish.zip": contains CASTEP output files of relaxations on the corresponding edge of the K-Sn-P system at the "polish" level of accuracy using various different xc-functionals or external pressures.</li> <li>"searches.zip": contains ".res" files that provide the relaxed structure from each different crystal structure prediction method on the corresponding edge of the K-Sn-P system.</li> <li>"bulk_modulus.zip" contains CASTEP output files for calculation of E(V) curves for low-lying KP phases with different xc-functionals.</li> <li>"nmr.zip" contains CASTEP output files for GIPAW-NMR calculations of chemical shifts for low-lying K-Sn-P phases.</li> <li>"spectral.zip" contains CASTEP and OptaDOS output files for projected bandstructure and DOS calculations of low-lying K-Sn-P phases.</li> <li>"digests.zip" contains JSON representations of all the structures from polish and searches, broken down into K-P and K-Sn-P specific digests.</li> </ul>
Dataset and structure database for an ML model to predict diffusivity in ZIF variants
<p>This dataset accompanies the publication titled "Data Mining for Predicting Gas Diffusivity in Zeolitic-imidazolate Frameworks (ZIFs)" (DOI: <a href="https://doi.org/10.1039/D2TA02624D">https://doi.org/10.1039/D2TA02624D</a>)</p> <p><a href="https://zenodo.org/api/files/b80f6d07-3bf4-484c-97ac-5d579fb0cc27/ESI_2_dataset.xlsx?versionId=dc4525d0-1c5c-478c-9bef-a156587ad69b">ESI_2_dataset.xlsx</a>: Descriptors for all ZIFs of the publication and simulations output, in the form of diffusivities of gas molecules (He up to iso-butane), in all ZIFs.</p> <p>ZIF_database.zip: ZIP file containing all ZIFs prepared by the authors (as discussed in the publication), through various units replacements, in the SOD topology, in .pdb format.</p>
ESM Atlas v0 representative random sample of predicted protein structures
<p>A representative random sample of the ESM Atlas v0 dataset introduced in "Evolutionary-scale prediction of atomic level protein structure with a language model.".<br> All predictions can be accessed in the ESM Metagenomic Atlas (<a href="https://esmatlas.com/">https://esmatlas.com</a>) open science resource, released on 2022-11-01.<br> Sample size: 997,405.</p>
ESM Atlas v0 random sample of high confidence predicted protein structures
<p>A random sample out of the 225M high confidence predictions in the ESM Atlas v0 dataset introduced in "Evolutionary-scale prediction of atomic level protein structure with a language model.".<br> All predictions can be accessed in the ESM Metagenomic Atlas (<a href="https://esmatlas.com/">https://esmatlas.com</a>) open science resource, released on 2022-11-01.<br> High confidence is defined as mean pLDDT > 0.7 and pTM > 0.7 and corresponds to ∼36% of the total 617M proteins folded.<br> This is the random sample used for analysis in the paper as well as visualization on the <a href="http://esmatlas.com/">esmatlas.com</a> Explore page.<br> Sample size: 999,520 based on 999,996 unique randomly sampled IDs and 0.05% missing data in the processing pipeline.</p>
Structural Interaction Fingerprints and Machine Learning for predicting and explaining binding of small molecule ligands to RNA: a benchmark dataset
<p><b>Structural Interaction Fingerprints and Machine Learning for predicting and explaining binding of small molecule ligands to RNA: a benchmark dataset.</b></p><p>Ribonucleic acids (RNA) play crucial roles in living organisms as they are involved in key processes necessary for proper cell functioning. Some RNA molecules, such as bacterial ribosomes and precursor messenger RNA, are targets of small molecule drugs, while others, e.g., bacterial riboswitches or viral RNA motifs are considered as potential therapeutic targets. Thus, the continuous discovery of new functional RNA increases the demand for developing compounds targeting them and for methods for analyzing RNA—small molecule interactions. We recently developed fingeRNAt - a software for detecting non-covalent bonds formed within complexes of nucleic acids with different types of ligands. The program detects several non-covalent interactions, such as hydrogen and halogen bonds, ionic, Pi, inorganic ion- and water-mediated, lipophilic interactions, and encodes them as computational-friendly Structural Interaction Fingerprint (SIFt). Here we present the application of SIFts accompanied by machine learning methods for binding prediction of small molecules to RNA targets. We show that SIFt-based models outperform the classic, general-purpose scoring functions in virtual screening. We discuss the aid offered by Explainable Artificial Intelligence in the analysis of the binding prediction models, elucidating the decision-making process, and deciphering molecular recognition processes.</p>
Michael Deem's PCOD database of Predicted Zeolitic Structures
<p>These are the SLC 'good' sections of version 2.1 of <a href="https://mwdeem.org">Michael Deem</a>'s database of predicted zeolite structures. Version 1 was published in M. W. Deem, R. Pophale, P. A. Cheeseman, and D. J. Earl, ``Computational Discovery of New Zeolite-Like Materials,'' with cover image, <em>J. Phys. Chem. C</em> <strong>113</strong> (2009) 21353-21360. Version 2 was published in R. Pophale, P. A. Cheeseman, and M. W. Deem, "A Database of New Zeolite-Like Materials," <em>Phys. Chem. Chem. Phys.</em> <strong>13</strong> (2011) 12407-12412:</p> <p>"We here describe a database of computationally predicted zeolite-like materials. These crystals were discovered by a Monte Carlo search for zeolite-like materials. Positions of Si atoms as well as unit cell, space group, density, and number of crystallographically unique atoms were explored in the construction of this database. The database contains over 2.6 M unique structures. Roughly 15% of these are within +30 kJ mol<sup>−1</sup>Si of α-quartz, the band in which most of the known zeolites lie. These structures have topological, geometrical, and diffraction characteristics that are similar to those of known zeolites. The database is the result of refinement by two interatomic potentials that both satisfy the Pauli exclusion principle." <a href="https://doi.org/10.1039/C0CP02255A">https://doi.org/10.1039/C0CP02255A</a></p> <p>These files are also available at <a href="https://mwdeem.org/PCOD">https://mwdeem.org/PCOD</a></p> <p> </p>
DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction (Supplementary Data)
<p>This dataset contains supplementary replication data for the paper titled "DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction". In particular, it contains a new version of our `final_raw_dips.tar.gz` protein pair representations which now contain (1) residue-level annotations for intrinsic disorder regions (IDRs) as well as (2) a copy of each protein pair representation in the HDF5 file format for programming language-agnostic read capabilities. In addition, this record also contains (3) raw MSAs (in HDF5 file format) generated for each protein pair using Jackhmmer and AlphaFold's small version of the Big Fantastic Database (BFD). Lastly, this record contains (4) PDB metadata derived for each DIPS-Plus complex using Graphein's PDBManager API as well as (5) structure-based (i.e., FoldSeek-based) training and validation splits of the dataset's complexes in the form of respective text files containing the file paths of complexes assigned to each split.</p>
AlphaFold structures reported in "AlphaFold2 Can Predict Single-Mutation Effects"
<p>This contains AlphaFold predictions for X proteins that are found in the Protein Data Bank (PDB), that were used to evalluate AlphaFold's predictions of mutation effects. This includes one set of structures predicted by AlphaFold2.0, using default settings, and one structure for each of 5 models. This also includes structures predicted by the ColabFold version of AlphaFold (6 recycles, 5 models, no template, amber minimization, 4 repeats).</p><p>There are also additional predicted structures that are found in the PDB that were not analyzed in the paper.</p><p>There are AlphaFold predictions for three proteins (BFP / RFP, GFP, and PafA), covering either all (BFP/RFP, PafA) or a subset (GFP) of the sequences in three datasets of phenotype measurements from high-throughput experiments.</p><p>Results are separated into tar files based on whether DeepMind (AF2.0) or ColabFold implementation was used.</p><p>Folders under "ColabFold/PDB" are labelled according to a sequence ID, since multiple PDB structures can exist for a single sequence. These sequence IDs can be mapped back to PDB IDs using the information in "seq_id_pdb_id.json".</p><p>All PDB files have been compressed using Foldcomp (<a href="https://github.com/steineggerlab/foldcomp">https://github.com/steineggerlab/foldcomp</a>). Foldcomp is required to decompress the ".fcz" files in order to recover the ".pdb" files.</p>
Data for: Brain structural connectivity predicts brain functional complexity
<p>Data used in analyses for "Brain structural connectivity predicts brain functional complexity: DTI derived centrality accounts for variance in fractal properties of fMRI signal"</p>
LiDAR-derived forest structure data and predictions of the locations of old-growth forests for Central Finland.
<p><strong>INTRO</strong><br> This archive contains data and analysis code for the Biodiversity Map -project conducted by Open Knowledge Finland (http://fi.okfn.org/projects/biodiversity-map/)</p> <p><strong>LICENCE</strong><br> The files listed below are all released to the public domain under a CC0 public domain dedication (https://creativecommons.org/publicdomain/zero/1.0/)</p> <p><strong>FILE DESCRIPTIONS</strong></p> <p><em><strong>FILE 1:</strong></em> background.zip<br> Inside the archive is a comma-separated file "background.csv" containing LiDAR-derived forest structure variables for 2/3 of Central Finland. These were derived from 3 raster data sets describing forest canopy maximum height (mh), forest canopy cover (cc) and lidar return intensity (in). The rasters had resolutions of 6 metres, 6 metres and 2 metres, respectfully. An 18 m resolution grid was then used to aggregate the rasters into average, minimum and maximum values + standard deviations of the original variables. The original LiDAR data was made available by the National Land Survey of Finland.</p> <p><br> <em><strong>FILE 2:</strong></em> conservation.lambdas<br> This file contains fitted parameters for the maxent model. For more information, check maxent documentation at https://www.cs.princeton.edu/~schapire/maxent/</p> <p><strong><em>FILE 3:</em></strong> conserved_swd.csv<br> Forest structure variables at 18 meter resolution for old-growth conservation areas in Central Finland. A subset of background.csv. This file still has a header, the variables are the same as in background.csv</p> <p><em><strong>FILE 4:</strong></em> grass_create_forest_rasters_from_las.sh<br> A shell script used to convert LiDAR files to raster maps of forest structure with GRASS 7.</p> <p><em><strong>FILE 5:</strong></em> lidar_coverage.png<br> A map showing the extent of LiDAR data available for Central Finland when we did the analyses.</p> <p><em><strong>FILE 6:</strong></em> maxent_model_run_product.sh<br> A shell script used to fit the maximum entropy model to predict the locations of conservation-area-like forests in Central Finland.</p> <p><em><strong>FILE 7:</strong></em> projection_product.csv<br> The results of the maxent model in a comma separated file. The first row has the variable names: x,y,product_fit. x and y are coordinates in the CRS ETRS-TM35FIN (EPSG:3067). product_fit is "the probablility that this 18*18 meter grid cell is old-growth conservation area".</p> <p><em><strong>FILE 8:</strong></em> README<br> A file with a description of the dataset in human-readable form.</p> <p><strong>VALIDATION FILES</strong><br> The data in these files was collected to validate the results of the aforementioned maxent model. The data were collected in a hierarchical sampling scheme: six randomly determinded unintersecting 9 km * 9 km landscape windows were chosen for sampling. From each window, three samples were taken. One sample from conservation areas, one sample from the "best" 10 % of forests as determined by the maxent model excluding conservation areas and one random sample. Not all windows contained conservation areas, and not all areas were accessible (islands, for example). In addition a few areas were skipped due to time constraints.</p> <p>The sampled points are identified by their lanscape window (suuralue), their sample (otos) and their sample number (mittauspiste).</p> <p><em><strong>FILE 9:</strong></em> validation_felled.csv<br> A comma separated list of those points that were not measured because they were felled.</p> <p><em><strong>FILE 10:</strong></em> validation_gps_results_2016-09-07.csv<br> A list of gps coordinates for all the sample points. product_fit is the value of the geographically closest prediction from the maxent model described above.</p> <p><em><strong>FILE 11:</strong></em> validation_lying_deadwood_transects_2016-08-30.csv<br> A comma separated file with data from deadwood transects. From each validation point, three 30 m long transects were made with 120 degree angles between them, and all lying deadwood more than 2 cm in diameter were measured. For some validation points, there were geographical obstructions which prevented the full 90 m of transect being surveyed, this is also recorded in the data. Each row holds measurements from one lying trunk.<br> </p> <p><em><strong>FILE 12:</strong></em> validation_relascope_2016-08-30.csv<br> Relascope measurements from the validation points. Each row is measurements for one species from one validation point. Dead and alive trees are counted separately.<br> </p> <p><strong>MORE INFORMATION</strong></p> <p>For more in-depth descritions of the files, read the file named README.<br> For some auxilliary files and information, check our old hackathon repository on github: https://github.com/Koalha/bdm_hackathon</p>
Male song structure predicts offspring recruitment to the breeding population in a migratory bird
<p>Bird song is a classic example of a sexually selected trait, but much of the work relating individual song components to fitness has not accounted for song typically being composed of multiple, often-correlated components, necessitating a multivariate approach. We explored the role of sexual selection in shaping complex male song of house wrens (<em>Troglodytes aedon</em>) by simultaneously relating its multiple components to fitness using multivariate selection analysis, which is widely used in insect and anuran studies but not in birds. The analysis revealed significant variation in the form and strength of selection acting on song across different selection episodes, from nest-site defense to recruitment of offspring to the breeding population. Males that sang more song typically employed in close communication sired more offspring that were subsequently recruited to the breeding population than those that sang far-communication song. However, this relationship was not consistent across earlier selection episodes, as evidenced by non-linear selection acting on these song components in other contexts. Collectively, our results present a complex picture of multivariate selection on male song structure that would not be evident using univariate approaches and suggest possible trade-offs within and among song components at different points of the breeding season. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.