Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
21,281
datasets available to search
ShareScore release 0.9.0
Dataset results
21,281 results for “molecular”
Molecular dynamics of ALDH1/2
<p>Members of the aldehyde dehydrogenase (ALDHs) superfamily have unique characteristics at their substrate binding site, where they catalyze the NAD+-dependent oxidation of aldehydes to their respective carboxylic acids. Molecular dynamics simulations, performed with the GROMACS 2022 package, have investigated the dynamics of the substrate binding site, using ALDH1 (PDB ID: 6B5G) and ALDH2 (PDB ID: 5L13) crystallographic structures. The tetrameric structures were simulated for 100 ns in: (1) APO (without any ligand bounded) state; (2) NAD-bounded state; (3) subtrate-bounded state; and (4) NAD- and substrate-bounded state. A total of 101 frames were extracted at regular intervals of 100 ps from the molecular dynamics trajectory. The ligands were removed from the substrate binding site after production to ease analysis.</p>
Transcriptomic analyses of normal-appearing CNS white matter from multiple sclerosis donors reveal subtype-specific molecular signatures of disease (REVISED)
<p>Datasets of bulk RNA-sequencing of NAWM from MS donors + supplementary images of RNA deconvolution of cell trajectories</p>
An Ice Age JWST inventory of dense molecular cloud ices
<p>This dataset is the first observational data release for the JWST Early Release Science Ice Age program (#1309). More information about this program can be found at our team website (<a href="http://jwst-iceage.org/">http://jwst-iceage.org/</a>) and the STScI website (<a href="https://www.stsci.edu/jwst/science-execution/approved-programs/dd-ers/program-1309">https://www.stsci.edu/jwst/science-execution/approved-programs/dd-ers/program-1309</a>). The data is analyzed in the article by McClure et al. (2023), to be published on January 24th, 2023 by Nature Astronomy.</p> <p>The dataset consists of 5 files, each containing a spectrum of one of two background stars analyzed in that publication. The spectra are given as wavelength in microns (column 1), flux in milli-Janskys (column 2), and uncertainty in the flux (column 3). Some basic information about the JWST pipeline version is given in the header of each text file, but users should refer to the Methods section of McClure et al. (2023) for the full description of how the data were processed and extracted, as it extends beyond the basic pipelines. Information about how the data were observed is given in the APT file, accessible through a query in STScI's APT application (search by PID 1309) or here at STScI (<a href="https://www.stsci.edu/jwst/science-execution/program-information.html?id=1309">https://www.stsci.edu/jwst/science-execution/program-information.html?id=1309</a>).</p> <p>Three of the files correspond to the background star NIR38. These represent spectra of this star taken separately by JWST with the NIRCam WFSS, NIRSpec FS, and MIRI LRS FS instrument modes on JWST. The other two files correspond to the background star J110621 and are spectra taken separately by JWST using the NIRSpec FS and MIRI LRS FS instrument modes.</p> <p>It is necessary to cite the McClure et al. (2023) Nature Astronomy publication when making use of these data, to fully describe the data processing, as well as this Zenodo DOI.</p>
Alignments from "Caecilian genomes reveal molecular basis of adaptation and convergent evolution of limblessness in vertebrates"
<p>Compressed file containing the alignments at both nucleotide and amino acid level for the manuscript "Caecilian genomes reveal molecular basis of adaptation and convergent evolution of limblessness in vertebrates" </p>
Uncovering the spatial landscape of molecular interactions within the tumor microenvironment through latent spaces (Figures)
<p>High resolution figures related to the below manuscript:</p> <p>Atul Deshpande, Melanie Loth, et al., <a href="https://doi.org/10.1101/2022.06.02.490672">Uncovering the spatial landscape of molecular interactions within the tumor microenvironment through latent spaces</a>. <em>bioRxiv</em> 2022. doi:10.1101/2022.06.02.490672</p>
Research data supporting: "Unsupervised Data-Driven Reconstruction of Molecular Motifs in Simple to Complex Dynamic Micelles"
<p>This repository contains the set of data shown in the paper <strong>"Unsupervised Data-Driven Reconstruction of Molecular Motifs in Simple to Complex Dynamic Micelles"</strong>, published on The Journal of Physical Chemistry B (DOI:10.1021/acs.jpcb.2c08726).</p>
Bielefeld Molecular Organic Glasses (BIMOG) Database
<p>The Bielefeld Molecular Organic Glasses (BIMOG) Database is based on a compiled dataset of experimental glass transition temperatures (Tg). The BIMOG database is the basis for our machine learning model for predicting the glass transition temperature of molecular organic compounds. For this purpose, we extended the previously unpublished data set from <a href="https://dx.doi.org/10.1039/C1CP22617G"><strong>Koop et al. 2011</strong></a> with further data from the literature. All experimental data are listed with their respecitve source.</p> <p>Further information is provided here: <strong><a href="https://tgml.chemie.uni-bielefeld.de">https://tgml.chemie.uni-bielefeld.de</a></strong></p> <p>To expand the database, we welcome the submission of further experimental data from the community. To do so, please follow this <a href="https://tgml.chemie.uni-bielefeld.de/submit_data"><strong>link</strong></a>.</p>
Binding site plasticity regulation of the FimH catch-bond mechanism: Molecular Dynamics dataset
<p>Dataset of Molecular Dynamics simulations and analysis scripts used in the article "Binding site plasticity regulation of the FimH catch-bond mechanism" [<a href="https://doi.org/10.1016/j.bpj.2023.05.029">paper</a>][<a href="https://doi.org/10.1101/2022.11.15.516604">bioRxiv</a>].</p> <p>Contains:</p> <ul> <li>Replica Exchange with Solute Scaling (REST2) simulations of the FimH protein lectin domain in its two main allosteric states (Associated and Separated), in presence and absence of its synthetic ligand heptyl α-ᴅ-mannose (input files and trajectories of the unscaled replicas)</li> <li>Replica Exchange Umbrella Sampling (REUS) simulations of the liganted systems along a collective variable (CV) describing binding site opening (input files and trajectories)</li> <li>REUS simulations in presence of a pulling force on the protein-ligand complex.</li> </ul> <p>See the article for more details.</p>
Controlling the Supramolecular Polymerization of Squaraine Dyes by a Molecular Chaperone Analogue
<p>ABSTRACT: Molecular chaperones are proteins that assist in the (un)folding and (dis)assembly of other macromolecular structures toward their biologically functional state in a non-covalent manner. Transferring this concept from nature to artificial self-assembly processes, here, we show a new strategy to control supramolecular polymerization via a chaperone-like two-component system. A new kinetic trapping method was developed that enables efficient retardation of the spontaneous self-assembly of a squaraine dye monomer. The suppression of supramolecular polymerization could be regulated with a cofactor, which precisely initiates self-assembly. The presented system was investigated and characterized by ultraviolet–visible, Fourier transform infrared, and nuclear magnetic resonance spectroscopy, atomic force microscopy, isothermal titration calorimetry, and single-crystal X-ray diffraction. With these results, living supramolecular polymerization and block copolymer fabrication could be realized, demonstrating a new possibility for effective control over supramolecular polymerization processes.</p>
RattanID - a molecular identification toolkit for rattan palms
<p>This repository contains laboratory protocols, reference datasets, auxiliary files (target file and sequencing adapters) and example data for the RattanID molecular identification toolkit (https://github.com/BenKuhnhaeuser/RattanID). It also contains a dataset detailing rattan occurrence records at species level, rattan uses and extinction risk predictions, as well as distribution maps built based on the rattan occurrence records dataset.</p>
Benchmark Data for Attracting Cavities 2.0 Small-Molecular Docking Program
<p>This repository provides data from the following article:<br> <br> U.F. Roehrig, M. Goullieux, M. Bugnon, V. Zoete,<br> Attracting Cavities 2.0: Improving the Flexibility and Robustness for Small-Molecule Docking.<br> J. Chem. Inf. Modeling 2023<br> https://doi.org/10.1021/acs.jcim.3c00054<br> <br> </p>
Fine-Tuning of Colloidal Polymer Crystals by Molecular Simulation
<p>Data archive corresponding to the manuscript "Fine-Tuning of Colloidal Polymer Crystals by Molecular Simulation" by M. Herranz et al., Phys. Rev. E 107, 064605 (2023); DOI: 10.1103/PhysRevE.107.064605</p> <p>Please see README.txt for instructions on how to access and read the files from the crystallographic analysis based on the CCE norm descriptor.</p> <p>All snapshots have been generated and successively analyzed by the Simu-D software.</p>
Molecular adaptations in response to exercise training are associated with tissue-specific transcriptomic and epigenomic signatures
<p>Processed data associated with the manuscript DOI: <a href="https://doi.org/10.1016/j.xgen.2023.100421" target="_blank" rel="noopener">10.1016/j.xgen.2023.100421 </a></p> <p>Analysis code on GitHub: <a href="../doi/10.5281/zenodo.8253917" target="_blank" rel="noopener">10.5281/zenodo.8253917</a></p> <p> </p>
Molecular dynamics trajectories for "Reservoir-REMD facilitates kinetic rescue from metastable peptide conformations
<p>The molecular dynamics-generated ensemble dataset for cyclo-(cGHHQKLV), used in the manuscript "Reservoir-REMD facilitates kinetic rescue from metastable peptide conformations". The dataset consists of 14 + 6 =20 .dcd files, and one .pdb file for rendering.</p>
TUK-FFDat - Data scheme and data format for transferable force fields for molecular simulation
<p>Online repository to suplement the following publication:</p> <p>G. Kanagalingam, S. Schmitt, F. Fleckenstein, S. Stephan: Data scheme and data format for transferable force fields for molecular simulation, Scientific Data, accepted (2023).</p>
Three systems of molecular markers reveal genetic differences between varieties sabina and balkanensis in the Juniperus sabina L. range
<p>Genotypes of 94 Juniperus sabina samples from 14 populations at SNP (Jsabina_SNPs.txt) and SilicoDArT (Jsabina_SilicoDArTs.txt) loci investigated using the DArTseq technology developed by Diversity Array Technology Pty Ltd (DArT, Canberra, ACT, Australia)</p>
Dataset of Molecular Dynamics Simulations for the Upregulated Biomarker PSMB8: 3UNF and its G210V Mutant in Experimental Autoimmune Encephalomyelitis
<p>This dataset contains molecular dynamics simulations data generated using GROMACS for the upregulated biomarker 3UNF and its G210V mutant in the context of Experimental Autoimmune Encephalomyelitis (EAE). EAE is a widely studied animal model for multiple sclerosis, and investigating the behavior of biomarkers in this model is crucial for understanding disease progression and potential therapeutic interventions.</p> <p>The dataset includes trajectory files, coordinate files, and relevant parameters used in the simulations. These simulations provide valuable insights into the structural dynamics, conformational changes, and interactions of the PSMB8 biomarker 3UNF and its G210V mutant within the EAE system. The data offers researchers an opportunity to analyze and explore the behavior of these biomarkers at the atomic level, aiding in the identification of potential binding partners, functional sites, and mechanisms associated with disease progression.</p> <p>By sharing this dataset, we aim to contribute to the scientific community by providing a valuable resource for further analysis, validation, and comparison of the molecular behavior of the upregulated biomarker 3UNF and its G210V mutant in Experimental Autoimmune Encephalomyelitis.</p>
Molecular dynamics simulation data 1: Structure of the connexin-43 gap junction channel in a putative closed state
<p>Molecular dynamics data for the manuscript Qi C.*, Acosta-Gutierrez S.*, Lavriha P., Othman A., Lopez-Pigozzi D., Bayraktar E., Schuster D., Picotti P., Zamboni N., Bortolozzi M., Gervasio F.L., Korkhov V.M. Structure of the connexin-43 gap junction channel in a putative closed state. eLife (2023) <a href="https://doi.org/10.7554/eLife.87616.2">https://doi.org/10.7554/eLife.87616.2</a></p> <p>The dataset includes:</p> <p>1. The starting coordinates, topology, MD inputs</p> <p>2. Production run gromacs trajectories for the Cx43 gap junction channel</p>
FooDrugs database: A database with molecular and text information about food - drug interactions
<p>FooDrugs database is a development done by the Computational Biology Group at IMDEA Food Institute (Madrid, Spain), in the context of the Food Nutrition Security Cloud (FNS-Cloud) project. Food Nutrition Security Cloud (FNS-Cloud) has received funding from the European Union's Horizon 2020 Research and Innovation programme (H2020-EU.3.2.2.3. – A sustainable and competitive agri-food industry) under Grant Agreement No. 863059 – <a href="http://www.fns-cloud.eu">www.fns-cloud.eu</a> (See more details about FNS-Cloud below)</p> <p>FooDrugs stores information extracted from transcriptomics and text documents for foo-drug interactiosn and it is part of a demonstrator to be done in the FNS-Cloud project. The database was built using MySQL, an open source relational database management system. FooDrugs_V2 host information for a total of 161 transcriptomics GEO series with 585 conditions for food or bioactive compounds (see below changes in versions V3 and V4). Each condition is defined as a food/biocomponent per time point, per concentration, per cell line, primary culture or biopsy per study. FooDrugs includes information about a bipartite network with 510 nodes and their similarity scores (tau score; https://clue.io/connectopedia/connectivity_scores) related with possible drug interactions with drugs assayed in conectivity map (https://www.broadinstitute.org/connectivity-map-cmap). The information is stored in eight tables: </p> <ul> <li> <p>Table “study” : This table contains basic information about study identifiers from GEO, pubmed or platform, study type, title and abstract </p> </li> <li> <p>Table “sample”: This table contains basic information about the different experiments in a study, like the identifier of the sample, treatment, origin type, time point or concentration.</p> </li> <li> <p>Table “misc_study”: This table contains additional information about different attributes of the study.</p> </li> <li> <p>Table “misc_sample”: This table contains additional information about different attributes of the sample.</p> </li> <li> <p>Table “cmap”: This table contains information about 70895 nodes, compromising drugs, foods or bioactives, overexpressed and knockdown genes (see section 3.4). The information includes cell line, compound and perturbation type.</p> </li> <li> <p>Table “cmap_foodrugs”: This table contains information about the tau score (see section 3.4) that relates food with drugs or genes and the node identifier in the FooDrugs network.</p> </li> <li> <p>Table “topTable”: This table contains information about 150 over and underexpressed genes from each GEO study condition, used to calculate the tau score (see section 3.4). The information stored is the logarithmic fold change, average expression, t-statistic, p-value, adjusted p-value and if the gene is up or downregulated.</p> </li> <li> <p>Table “nodes”: This table stores the information about the identification of the sample and the node in the bipartite network connecting the tables “sample”, “cmap_foodrugs” and “topTable”.</p> </li> </ul> <p>In addition, FooDrugs_V2 database stores a total of 6422 food/drug interactions from 2849 text documents, obtained from three different sources: 2312 documents from PubMed, 285 from DrugBank, and 252 from drugs.com. These documents describe potential interactions between 1464 food/bioactive compounds and 3009 drugs (see below changes in versions V3 and V4). The information is stored in two tables:</p> <ul> <li> <p>Table “texts”: This table contains all the documents with its identifiers where interactions have been identified with strategy described in section 4. </p> </li> <li> <p>Table “TM_interactions”: This table contains information about interaction identifiers, the food and drug entities, and the start and the end positions of the context for the interaction in the document.</p> </li> </ul> <p> </p> <p>FNS-Cloud will overcome fragmentation problems by integrating existing FNS data, which is essential for high-end, pan-European FNS research, addressing FNS, diet, health, and consumer behaviours as well as on sustainable agriculture and the bio-economy. Current fragmented FNS resources not only result in knowledge gaps that inhibit public health and agricultural policy, and the food industry from developing effective solutions, making production sustainable and consumption healthier, but also do not enable exploitation of FNS knowledge for the benefit of European citizens.<br> FNS-Cloud will, through three Demonstrators; Agri-Food, Nutrition & Lifestyle and NCDs & the Microbiome to facilitate:<br> (1) Analyses of regional and country-specific differences in diet including nutrition, (epi)genetics, microbiota, consumer behaviours, culture and lifestyle and their effects on health (obesity, NCDs, ethnic and traditional foods), which are essential for public health and agri-food and health policies;<br> (2) Improved understanding agricultural differences within Europe and what these means in terms of creating a sustainable, resilient food systems for healthy diets; and<br> (3) Clear definitions of boundaries and how these affect the compositions of foods and consumer choices and, ultimately, personal and public health in the future.<br> Long-term sustainability of the FNS-Cloud will be based on Services that have the capacity to link with new resources and enable cross-talk amongst them; access to FNS-Cloud data will be open access, underpinned by FAIR principles (findable, accessible, interoperable and re-useable). FNS-Cloud will work closely with the proposed Food, Nutrition and Health Research Infrastructure (FNHRI) as well as METROFOOD-RI and other existing ESFRI RIs (e.g. ELIXIR, ECRIN) in which several FNS-Cloud Beneficiaries are involved directly. (https://cordis.europa.eu/project/id/863059)</p> <p><strong>***** changes between version FooDrugs_v2 and FooDrugs_V3 (31st January 2023) are:</strong></p> <ul> <li> <p>Increased the amount of text documents by 85.675 from PubMed and ClinicalTrials.gov, and the amount of Text Mining interactions by 168.826.</p> </li> <li> <p>Increased the amount of transcriptomic studies by 32 GEO series.</p> </li> <li> <p>Removed all rows in table <em>cmap_foodrugs</em> representing interactions with values of <em>tau</em>=0</p> </li> <li> <p>Removed 43 GEO series that after manually checking didn't correspond to food compounds.</p> </li> <li> <p>Added a new column to the table <em>texts</em>: <em>citation</em> to hold the citation of the text. </p> </li> <li> <p>Added these columns to the table <em>study</em>: <em>contributor</em> to contain the authors of the study, <em>publication_date</em> to store the date of publication of the study in GEO and <em>pubmed_id</em> to reference the publication associated with the study if any.</p> </li> <li> <p>Added a new column to <em>topTable </em>to hold the top 150 up-regulated and 150 down-regulated genes</p> </li> </ul> <p><strong>***** changes between version FooDrugs_v3 and FooDrugs_V4 (28th July 2023) are:</strong></p> <ul> <li> <p>Increased the amount of text documents by 439.338 from PubMed, ClinicalTrials.gov and DDI corpus (<a href="https://www.sciencedirect.com/science/article/pii/S1532046413001123">Herrero-Zazo et al.</a>), and the amount of Text Mining interactions by 1108429.</p> </li> </ul> <p> </p>
Molecular dynamics simulation based analysis of celecoxib-polymer interactions
<p>This dataset contains scripts and coordinate files for running and analysing molecular dynamics simulations to investigate celecoxib-polymer interactions in aqueous solution. Trajectory files (stripped of water and ions) are included.</p> <p>The associated study is described in:</p> <p>" Comparative analysis of drug-salt-polymer interactions by experiment and molecular simulation improves biopharmaceutical performance", Sumit Mukesh, Goutam Mukherjee, Ridhima Singh, Nathan Steenbuck, Carolina Demidova, Prachi Joshi, Abhay T. Sangamwar, Rebecca C. Wade, submitted.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.