Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
573
datasets available to search
ShareScore release 0.9.0
Dataset results
573 results for “Structure analysis”
Dataset of "Sensitivity analysis in photodynamics: How the electronic structure controls cis-stilbene photodynamics?"
<p>The techniques of computational photodynamics are increasingly employed to unravel reaction mechanisms and interpret experiments. However, inaccuracies in nonadiabatic dynamics can lead to misinterpretations, particularly when calculated observables exhibit low sensitivity to the underlying dynamics. This issue is exemplified in the photochemistry of cis-stilbene, where similar experimental outcomes have been differently interpreted based on the electronic structures supporting nonadiabatic dynamics. This study examines the predictions of cis-stilbene photochemistry using trajectory surface hopping methods coupled with various electronic structures (OM3-MRCISD, SA2-CASSCF, XMS-SA2-CASPT2, and XMS-SA3-CASPT2) and assesses their ability to interpret experimental observations. Although the excited-state lifetimes show consistency, ranging from 360 fs to 295 fs, the reaction quantum yields vary significantly. The quantum yield for cyclization ranges from nearly zero to 35% while the photoisomerization channel can either exceed 50% or be entirely suppressed completely in the second case. Intriguingly, the calculated photoelectron signal is not strikingly different for different reaction scenarios, making the methods seemingly reliable when treated separately Furthermore, analyzing stationary points on the potential energy surface does not reliably predict simulation outcomes, nor does it aid in selecting a specific method before simulations. Therefore, we advocate for incorporating sensitivity analyses in the simulation protocol. While employing an ensemble of methods is impractical, nonadiabatic simulations with external bias present a resource-efficient approach to achieve this goal.</p>
Datasets and Jupyter notebook for the structural analysis of protein-RNA interface evolution
<p>The present repository contains data and code related to our manuscript "Structural comparison of protein-RNA homologous interfaces reveals widespread overall conservation contrasted with versatility in polar contacts". In the manuscript, we analyze the evolution of protein-RNA interfaces by building a dataset of protein-RNA interologs (homologous interfaces) and exploring how interface contacts are conserved between homologous interfaces, as well as possible explanations for non-conserved contacts.</p> <p>This repository contains the following files:</p> <ul> <li>DataAnalysisNotebook.ipynb is a Jupyter notebook to reproduce contact conservation analysis and all figures from our manuscript, and to explore data</li> <li>env.yaml is an environment file in order to build a Conda/Mamba environment to run the Jupyter notebook </li> <li>2022-02-21-PDB.csv contains data from the PDB about 3D structures of complexes containing interacting protein and RNA chains (PDB structure identifier, chain identifiers, experimental technique and resolution)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.tsv contains more detailed information about interacting protein and RNA chains from these complexes (PDB and chain identifiers, protein and RNA size, interface size and number of contacts)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.txt.selectXE_2.50_p30_r10_pi5_ri5_rep_bc-100.out_RNAcl_0.99.tsv contains the same detailed information, restricted to the filtered dataset used as a starting point in our interolog search pipeline</li> <li>PDBinterfaceAlign.csv contains information about the structural alignment of pairs of protein-RNA interactions (structural alignment TM-scores, sequence identity and coverage)</li> <li>DataInterologsParam.tsv contains information about a pre-filtered set of 2587 potential interologs (including interface RMSD, sequence identity and coverage and interface size)</li> <li>DataInterologsContactsFixedSASA.tsv contains detailed information about conserved and non-conserved contacts in the final set of 2022 interologs (atomic contacts, apolar contacts, hydrogen bonds, salt bridges and stacking information for aminoacid-nucleotide pairs, as well as information about whether each belongs to the interface, secondary structures, and the aminoacid surface accessibility and evolutionary conservation metrics) - compared to version 1, the calculation of solvent accessibility was fixed for a number of interolog pairs</li> <li>DataCons.csv contains precomputed contact conservation metrics for each of the 2022 interolog pairs, for fast reproduction of manuscript figures</li> <li>DataInterologsContactsResampledMaintainStructSeqId.tsv, DataInterologsContactsShuffled.tsv and DataInterologsShuffled.tsv relate to baselines computed for contact conservation assessment</li> <li>clan.txt, clan_membership.txt, ecod.latest.domains.uniq.txt, rfam_interfaces_977.txt, DataGroupsECOD.tsv, DataGroupesRFAM.tsv, DataGroupsRFAMClan.tsv, DataInterfaceGroupsECOD.tsv and DataInterfaceGroupsRFAM.tsv relate to the ECOD (respectively Rfam) classification of protein domains (respectively RNA) in protein-RNA interfaces from our dataset</li> <li>ListeIntraHbonds.pkl and ListeIntraSaltBridges.pkl are pickle-format data files containing intra-molecular hydrogen bonds and salt bridges (respectively) that are used to analyse scenarii of compensation for non-conserved polar contacts.</li> </ul>
Global comparative structural analysis of responses to protein phosphorylation
<p>This contains the structures and data used for the structural analysis presented in <em>Global comparative structural analysis of responses to protein phosphorylation</em> (Correa Marrero et al., https://doi.org/10.1101/2024.10.18.617420 ). To summarize:</p> <ul> <li>filtered_df.xlsx: dataset of paired phosphorylated structures and their non-phosphorylated counterparts. Each row contains one such pair.</li> <li>chains_by_protein.zip: each directory (named with a UniProt ID) contains the used structures that form the basis for the analysis. The structures are in PDB format, in a separate directory for each protein in the dataset. The exception is the annotation_per_psite directory, which contains annotation as a csv file for each phosphosite.</li> <li>extracted_domains.zip: contains structures of Pfam domains (extracted from the previous dataset) in PDB format. Each filename follows the format {PDB ID}_{Chain ID}_{Pfam domain ID}. The domain_coverage.csv file lists the domain coverage of the structure, as well as its length compared to the whole sequence and the whole structure it was extracted from. These are the structures used for the analysis shown in Fig. 1 f-h.</li> <li>extracted_pfam_domains.zip: contains structures of a broader set Pfam domain structures (the whole set of Pfam domains found to contain a phosphosite in filtered_df.csv) in PDB format. Each directory (named with the Pfam ID) contains the structures. merged_pfam_data.tsv contains metadata about the structures (structure quality, coverage of the domain structure, phosphosite location...). These are the structures used for the analysis shown in Fig. 2.</li> </ul>
Data set for the manuscript 'Robustness Analysis of Metasurfaces: Perfect Structures are not always the Best'
<p>In this data set, there are 1 PDF, 3 m-files, and 3 zip files.</p> <p>The manuscript (<strong><em>Readme document<em>.</em>pdf</em></strong>) contains three sections: Quasi-analytical model (<em><strong>analytical_model_EnergyConservation.m</strong></em>), post-processing full-wave simulations (<strong><em>post_processing_from_COMSOL.m</em></strong>), and post-processing experimental data (<strong><em>post_processing_from_experiment.m</em></strong>). Each Matlab code is explained in this manuscript. Corresponding raw data (<em><strong>COMSOL simulation data for reflective metallic metasurfaces.zip</strong>,<strong> COMSOL simulation data for transmitive dielectric metasurfaces.zip</strong>, </em>and <strong><em>experimental data.zip</em></strong>) are attached. One can move the required m-file into the folder and run the m-file directly. In the COMSOL simulation data.zip file, one can find two COMSOL files, which retain the settings for simulation and extracting the required data. </p> <p>The Matlab codes are implemented with version R2018b.</p> <p>The COMSOL files are created with version COMSOL Multiphysics 5.6.</p> <p> </p>
Tabular datasets for "In situ structural analysis reveals membrane shape transitions during autophagosome formation"
<p>Tabular source data for all plots in the manuscript "In situ structural analysis reveals membrane shape transitions during autophagosome formation". The article is available at https://doi.org/10.1101/2022.05.02.490291. The naming of the sheets in the .xlsx files corresponds to the figure number and panel.</p>
Crystallization process of organic-inorganic methylammonium lead bromide perovskite (MAPbBr3), GIXD analysis results: diffraction features and crystal structure
<p>Analysis result of an <em>in-situ</em> measurement of the crystallization process of organic-inorganic methylammonium lead bromide perovskite (MAPbBr3) on a glass substrate.</p> <p>This dataset contains the positions, sizes, and integrated intensities of extracted diffraction peaks with 0.1s time resolution.</p> <p>For crystal structure matching, the provided CIF file (CCDC 1446529) was used.</p>
An Automatic Neuroimaging Infrastructure For Synthesis and Analysis of Structural MRI Data
<p>We have designed, implemented and distributed a fully automatic neuroimaging infrastructure for the synthesis and analysis of structural magnetic resonance imaging (MRI) data. The framework provides a concrete environment for the quantitative validation of various methods for the analysis of brain asymmetries, for comparisons of methods and measures of brain shape asymmetry, and possibly for clarifying contradicting neuroimaging findings of brain lateralizations.</p> <p>See <a href="https://sites.google.com/site/brainmorphorg/home">https://sites.google.com/site/brainmorphorg/home </a></p> <p>and </p> <p>A. Pepe, I. Dinov, and J. Tohka . An Automatic Framework for Quantitative Validation of Voxel Based Morphometry Measures of Anatomical Brain Asymmetry. <a href="http://dx.doi.org/10.1016/j.neuroimage.2014.06.029">NeuroImage , 100: 444 - 459, 2014</a><a href="https://doi.org/10.1016/j.neuroimage.2014.06.029"> </a></p> <p>for more information. </p>
Project files provided as supporting information to the manuscript "A deep learning approach to the structural analysis of proteins"
<p><strong>README file to the project files provided as supporting information to the manuscript “A deep learning approach to the structural analysis of proteins”</strong></p> <p>Dec. 30, 2018</p> <p>Authors: Marco Giulini and Raffaello Potestio</p> <p>==================================</p> <p>The dataset contains the following files:</p> <p> </p> <p>- datasets.zip: archive containing five .csv files, namely:</p> <p> - decoys_cm.csv : all the data for 10728 protein decoys, training set</p> <p> - evaluation_cm.csv : all data for 146 proteins in the evaluation set</p> <p> - random_CG.csv : 1200 Coulomb matrices. 100 CG models for each protein with 120 amino acids</p> <p> - 1e5g_centered_sphere.csv : 100 CG models in which the central atoms in 1e5g are not removed</p> <p> - 1e5g_random_sphere.csv : 10 CG models for 10 different (random) locations for the sphere that includes atoms that have to be retained. 100 CG models in total</p> <p> </p> <p>- decoys_labels.lab containing the labels associated to the 10728 decoys present in the training set</p> <p>- evaluation_labels.lab containing the labels associated to the 146 pdb files in the evaluation set</p> <p>- random_CG_labels.lab containing the labels associated to the 6 proteins with 120 amino acids</p> <p>- network_development_training: a python script that performs cross validation and full training of the model</p> <p>- saved_networks.zip FOLDER containing 10 networks: the architecture is included in .json files while weight parameters are inside .hs files</p> <p> </p> <p>- pdb_files.zip FOLDER containing the PDB files that have been employed in the project, namely:</p> <p> - pdb_files_len100 : pdb files with 100 amino acids</p> <p> - pdb_files_len101-110 : pdb files with a number of amino acids between 101 and 110</p> <p> - decoys : decoys of length 100 extracted from the above folder: name syntax == PDBNAME_decoy_STARTRES_ENDRES.pdb</p> <p> EXAMPLE 6gsp.pdb will give rise to 6gsp_decoy_0_100.pdb , 6gsp_decoy_1_101.pdb , 6gsp_decoy_2_102.pdb , 6gsp_decoy_3_103.pdb , 6gsp_decoy_4_104.pdb</p> <p> - pdb_files_len100 : 6 pdb files with 120 amino acids</p> <p> </p>
R scripts for analyzing LiDAR data to assess forest canopy structure and perform Principal Component Analysis (PCA) on derived metrics
<p>This repository contains R scripts for analyzing LiDAR data to assess forest canopy structure and perform Principal Component Analysis (PCA) on spectral and LiDAR-derived metrics. The scripts cover LiDAR data processing, canopy height model (CHM) generation, calculation of forest canopy metrics, and PCA analysis.</p>
RMSD and Trp assays for "Structural analysis of a motor with increased mechanical output reveals new transitions in kinesin microtubule motility"
<p>This entry is for our manuscript, "Structural analysis of a motor with increased mechanical output reveals new transitions in kinesin microtubule motility" by Satoki Shibata*, Matthew Y. Wang*, Tsuyoshi Imasaki*, Hideki Shigematsu, Diego Ugarte La Torre, Yuanyuan Wei, Chacko Jobichen, Hajime Hagio, J. Sivaraman, Yuji Sugita, Sharyn A. Endow & Ryo Nitta.</p> <p>*Equal contribution</p> <p>Corresponding Authors: Tsuyoshi Imasaki*, Sharyn A. Endow, Ryo Nitta</p> <p>The deposited datasets are 1) beta-strand residue all-atom RMSD between kinesin-14 NcdY485K ADP and ADP + free Pi nucleotide states and 2) fluorimeter assays of NcdY485K intrinsic Trp fluorescence changes upon addition of free Pi . The data are presented in Fig. 6e, Fig. 7b,c and Supplementary Fig 10. Files are named for the figures in which the data are shown. A computer script for the RMSD analysis and methods for the fluorimeter assays and data analysis are included in the deposited files.</p>
Structural and Molecular Analysis of Adult Mouse Astrocytes and Vascular Connectivity in the Cortex and Hippocampus
<p>After image acquisition (0-RAW_CL230331_E2_serie1) and deconvolution (1-Deconvolved_CL230331_E2_serie1) using confocal microscopy and the SVI Huygens software,respectively, the image processing was conducted using Imaris, Fiji, and Matlab software. This process involved a sequence of manual operations (2-Imaris_surfaces_CL230331_E2_serie1) and custom Groovy scripts (5-Groovy scripts).</p> <p>The dataset analysis (3-Imaris_final_CL230331_E2_serie1_ims) allowed for a deeper investigation of morphological and molecular properties of adult mouse astrocytes (4-Image analysis_CL230331_E2_serie1) in two brain regions, the Isocortex and the Hippocampus, known to be interconnected to support multiple cognitive functions.</p>
Analysis of copy number variation in dogs implicates genomic structural variation in the development of anterior cruciate ligament rupture
<p>Anterior cruciate ligament (ACL) rupture is an important condition of the human knee. Second ruptures are common and societal costs are substantial. Canine cranial cruciate ligament (CCL) rupture closely models the human disease. CCL rupture is common in the Labrador Retriever (5.79% prevalence), ~100-fold more prevalent than in humans. Labrador Retriever CCL rupture is a polygenic complex disease, based on genome-wide association study (GWAS) of single nucleotide polymorphism (SNP) markers. Dissection of genetic variation in complex traits can be enhanced by studying structural variation, including copy number variants (CNVs). Dogs are an ideal model for CNV research because of reduced genetic variability within breeds and extensive phenotypic diversity across breeds. We studied the genetic etiology of CCL rupture by association analysis of CNV regions (CNVRs) using 110 case and 164 control Labrador Retrievers. CNVs were called from SNPs using three different programs (PennCNV, CNVPartition, and QuantiSNP). After quality control, CNV calls were combined to create CNVRs using ParseCNV and an association analysis was performed. We found no strong effect CNVRs but found 46 small effect (max(T) permutation P<0.05) CCL rupture associated CNVRs in 22 autosomes; 25 were deletions and 21 were duplications. Of the 46 CCL rupture associated CNVRs, we identified 39 unique regions. Thirty four were identified by a single calling algorithm, 3 were identified by two calling algorithms, and 2 were identified by all three algorithms. For 42 of the associated CNVRs, frequency in the population was <10% while 4 occurred at a frequency in the population ranging from 10-25%. Average CNVR length was 198,872bp and CNVRs covered 0.11 to 0.15% of the genome. All CNVRs were associated with case status. CNVRs did not overlap previous canine CCL rupture risk loci identified by GWAS. Associated CNVRs contained 152 annotated genes; 12 CNVRs did not have genes mapped to CanFam3.1. Using pathway analysis, a cluster of 19 homeobox domain transcript regulator genes was associated with CCL rupture (P=6.6E-13). This gene cluster influences cranial-caudal body pattern formation during embryonic limb development. Clustered genes were found in 3 CNVRs on chromosome 14 (HoxA), 28 (NKX6-2), and 36 (HoxD). When analysis was limited to deletion CNVRs, the association was strengthened (P=8.7E-16). This study suggests a component of the polygenic risk of CCL rupture in Labrador Retrievers is associated with small effect CNVs and may include aspects of stifle morphology regulated by homeobox domain transcript regulator genes.</p>
Data Set for the Journal Article "Automated Preparation of Nanoscopic Structures: Graph-Based Sequence Analysis, Mismatch Detection, and pH-Consistent Protonation with Uncertainty Estimates"
<p>This repository containes the data generated by ASAP and discussed in the journal article [Csizi, K.-S. and Reiher, M., 2023, arXiv:2307.16344], including Cartesian coordinates of training and test set molecules, and MD trajectories. </p>
Dataset of structural and energetic descriptors for optimized geometries of the pyrene dimer in the S1 state and Jupyter Notebbok used for the unsupervised clustering, analysis and visualization
<p>Geometrical and energy data extracted from a set of 188 optimized geometries for the pyrene dimer in the first excited singlet state at the TD-CAMB3LYP + D3BJ / 6-31G* / C-CPCM(Cyclohexane) level. </p> <p>Coordinates (in xyz format) of the 188 optimized geometries considered.</p> <p>Jupyter notebook used to perform unsupervised clustering and analysis of the available structures.</p>
Output of simulations for "Atmosphere Response to an Oceanic Sub-mesoscale SST Front: A Coherent Structure Analysis": Part 1
<p>This dataset contains the output of the simulations for the paper "Atmosphere Response to an Oceanic Sub-mesoscale SST Front: A Coherent Structure Analysis" doi: [TO BE COMPLETED]. This is part 1. It contains data for the S1 simulation.</p>
Output of simulations for "Atmosphere Response to an Oceanic Sub-mesoscale SST Front: A Coherent Structure Analysis": Part 2
<p>This dataset contains the output of the simulations for the paper "Atmosphere Response to an Oceanic Sub-mesoscale SST Front: A Coherent Structure Analysis" doi: [TO BE COMPLETED]. This is part 2. It contains the remaining of data for the S1 simulation, the data of the reference simulations RefC and RefW, and the data for the sensitivity analysis of the supplementary material.</p>
Data from: Using model analysis to unveil hidden patterns in tropical forest structures
<p>Data set of the article entitled: <strong>Using model analysis to unveil hidden patterns in tropical forest structures</strong></p> <p>This data set gives the following structural attributes for 133 forest plots at 9 sites in the tropics:</p> <ul> <li>tree density (ha<sup>-1</sup>)</li> <li>basal area (m<sup>2</sup> ha<sup>-1</sup>)</li> <li>mean diametere (cm)</li> <li>equivalent diameter (cm)</li> <li>density of trees in the dbh class 10-30 cm (ha<sup>-1</sup>)</li> <li>density of trees in the dbh class 30-60 cm (ha<sup>-1</sup>)</li> <li>density of trees with dbh ≥ 60 cm (ha<sup>-1</sup>)</li> <li>aboveground dry biomass (Mg ha<sup>-1</sup>)</li> <li>fraction of the biomass of trees with dbh ≥ 60 cm</li> <li>weighted mean wood density (g cm<sup>-3</sup>)</li> <li>density of trees in the dbh class 10-20 cm (ha<sup>-1</sup>)</li> <li>density of trees in the dbh class 20-30 cm (ha<sup>-1</sup>)</li> <li>density of trees in the dbh class 30-40 cm (ha<sup>-1</sup>)</li> <li>density of trees in the dbh class 40-50 cm (ha<sup>-1</sup>)</li> <li>density of trees in the dbh class 50-60 cm (ha<sup>-1</sup>)</li> <li>density of trees in the dbh class 60-70 cm (ha<sup>-1</sup>)</li> <li>density of trees in the dbh class 70-80 cm (ha<sup>-1</sup>)</li> <li>density of trees in the dbh class 80-90 cm (ha<sup>-1</sup>)</li> <li>density of trees in the dbh class 90-100 cm (ha<sup>-1</sup>)</li> <li>density of trees in the dbh class 100-110 cm (ha<sup>-1</sup>)</li> <li>density of trees in the dbh class 110-120 cm (ha<sup>-1</sup>)</li> <li>density of trees in the dbh class 120-130 cm (ha<sup>-1</sup>)</li> <li>density of trees with dbh ≥ 130 cm (ha<sup>-1</sup>)</li> </ul>
AutoDock and CB-Dock data for (NPA)6Zn3(H2O)2 in Synthesis, structural analysis, and docking studies with SARS-CoV-2 of a trinuclear zinc complex with N-phenylanthranilic acid ligands
<p>AutoDock 4.2 and CB-Dock data for (NPA)<sub>6</sub>Zn<sub>3</sub>(H<sub>2</sub>O)<sub>2</sub> with M<sup>pro</sup> from SARS-CoV-2 from PDB Id: 6LU7. </p>
NMR data for (NPA)6Zn3(H2O)2 in Synthesis, structural analysis, and docking studies with SARS-CoV-2 of a trinuclear zinc complex with N-phenylanthranilic acid ligands
<p><sup>1</sup>H, <sup>13</sup>C, COSY, HMBC, and HSQC NMR data in fid format for (NPA)<sub>6</sub>Zn<sub>3</sub>(H<sub>2</sub>O)<sub>2</sub> (NPA = 2-(phenylamino) benzoate) in DMSO-<em>d</em><sub>6.</sub></p>
Comparative analysis of surface sanitization protocols on the bacterial community structures in the hospital environment
<p>In this study, we used 16S rRNA gene sequencing approaches to characterize the bacterial microbiota on different surfaces of the hospital environment. The longitudinal data was then subjected to comprehensive comparisons between different sanitation strategies (disinfectants, detergents and probiotics) to measure their potential effect on the microbial community structures in the hospital environment.</p> <p>This archive contains results and data of the 16S rRNA amplicon sequencing performed on 1019 environmental and 271 patient DNA samples collected over the time course of 40 weeks in a newly opened ward in the neurological station at the Charité Hospital (Berlin). The files include a study information and sample metadata sheets, BIOM-tables and information about the taxonomy results and diversity metrics.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.