Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11
datasets available to search
ShareScore release 0.9.0
Dataset results
11 results for “molecular mixture”
Predicted room temperature electrical conductivity of molecular mixtures
<p>In the associated manuscript, we propose the MolSets machine learning model for molecular mixture properties. Using the MolSets architecture, we train a model on a dataset curated by Bradford et al. (2023) to predict the room temperature (298 K) electrical conductivity of mixtures. Here, we report the model-predicted conductivities of all equal-weight binary mixtures among 28 types of small molecules, combined with 30 types of Li<sup>+</sup> salts (1 mol·kg<sup>-1</sup>), totaling 11,340 candidate lithium battery electrolytes. Note that the current model has a limitation of not taking salt solubility into account. This dataset is for demonstration purposes and should be used with caution.</p>
Molecular dynamics simulation of SpoIVFB:Pro-SigmaK complex (in POPE/POPG mixture)
<p>Simulation in 2:1 POPE:POPG mixture.</p> <p>Found here are all files needed to reproduce or visualize the results of molecular dynamics simulation of the SpoIVFB intramembrane protease bound to the transcription factor Pro-sigmaK. The protein complex was embedded in a POPE:POPG bilayer using CHARMM-GUI and simulated using OpenMM. The README file is a C-shell script that will run equilibration and 250ns of unrestrained simulation. </p> <p>Individual output (.out) and trajectory (.dcd) files are provided for each checkpoint of the simulation. A combined trajectory containing 250 ns of unrestrained simulation is also provided (combined_250ns_traj.dcd). Together with the step5_input.psf file, this combined dcd file can be used with common software such as VMD to visualize the molecular dynamics trajectory.</p>
Predicted room temperature electrical conductivity of molecular mixtures
Open the record for dataset details and reuse information.
Discovering molecular regulators of ageing using mixture models with RNA-sequencing data
<p>Identifying the molecular regulators that control ageing is challenging because the ageing process is influenced by a combination of genetic and environmental factors which makes it difficult to source the contribution of a single gene. Multiple studies have demonstrated that as humans age, increased gene expression heterogeneity results in the dysregulation of key regulators and pathways. Given the dynamic nature of gene expression, it is vital that this data be modelled by statistical approaches that can appropriately account for changes in variability to understand the contribution of heterogeneity during the aging process and properly identify its regulators. This study demonstrates the utility of using mixture models to model biological variability of gene expression occurring during ageing and how novel potential regulators of ageing can be identified.</p> <p>Our mixture modelling approach was applied to gene expression data from the Genotype-Tissue Expression (GTEx) cohort. For every gene, the expression profile was modelled using a mixture model across the cohort where the subset of donors corresponding to each mode was tested for a significant change in age group. The multi-tissue aspect of GTEx was leveraged to find ageing regulators based on this mixture model approach genes that were common across multiple tissues, suggesting that the regulation of ageing may also be controlled through a set of genes that have non-tissue-specific activity.</p> <p>Our approach identified well-documented ageing regulators <em>mTOR </em>and <em>RICTOR</em> and other potential ageing regulators such as <em>IL4</em> and <em>GPR4</em> which were detected only by our approach. Genes identified by edgeR, DESeq2 and the mixture model-based approach were enriched for similar biological pathways. This suggests that while the specific ageing regulators identified from our approach may be distinct, they generally belong in the same pathways as the genes identified by standard approaches. Overall, these results indicate that modelling gene expression variability using mixture models in conjunction with standard differential gene expression can help uncover new regulators that have a potential role for understanding human ageing.</p> <p>I</p>
Molecular dynamics simulations of hyaluronan octamer–tetrapeptide mixtures
<p><strong>To cite:</strong> Riopedre-Fernandez, M.; Biriukov, D.; Dračínský, M.; Martinez-Seara, H. Hyaluronan-arginine enhanced and dynamic interaction emerges from distinctive molecular signature due to electrostatics and side-chain specificity. Carbohydr. Polym. 2024, 325, 121568. DOI: <a href="https://doi.org/10.1016/j.carbpol.2023.121568">10.1016/j.carbpol.2023.121568</a></p> <p>MD simulations of hyaluronan octamer (HA8) with tetrapeptides. Simulation files and molecular topologies are provided.</p> <p>Tetrapeptides simulated: tetraarginine (R4), tetralysine (K4), tetraalanine (A4), tetraproline (P4), tetraglycine (G4), arginine-lysine-arginine-lysine (RKRK), and tetraglutamic acid (E4).</p> <p>Two force fields were compared: CHARMM (version charmm36-jul2020.ff.tgz from <a href="http://mackerell.umaryland.edu/charmm_ff.shtml#gromacs">http://mackerell.umaryland.edu/charmm_ff.shtml#gromacs</a>) and prosECCo75 (<a href="https://gitlab.com/sparkly/prosecco/prosECCo75">https://gitlab.com/sparkly/prosecco/prosECCo75</a>).</p> <p>Each simulation contained one hyaluronan polymer, one tetrapeptide, CHARMM-specific TIP3P water, and potassium counterions (standard "K" model in CHARMM and "K_s" model in prosECCo75) when necessary.</p> <p>We also additionally performed: (i) reference simulations with only hyaluronan octamer and potassium counterions; (ii) reference simulations with only R4 or K4 peptide and chloride counterions; (iii) simulations with two peptides (R4, K4, or G4) and two hyaluronan octamers.</p> <p>Simulations were done in Gromacs.</p> <p>Length - at least 1 µs, prolonged till 2 µs for prosECCo75 systems with R4, K4, or G4 peptides.</p> <p>Temperature - 300 K.</p> <p>For peer-reviewing process, extracted PDB configurations (a configuration every 10 ns of 2 μs simulations excluding the first 100 ns of equilibration) from selected systems were separately uploaded to <a href="http://doi.org/10.5281/zenodo.8423276">http://doi.org/10.5281/zenodo.8423276</a></p>
Molecular dynamics simulations of hyaluronan octamer–tetrapeptide mixtures (PDB configurations)
<p>Extracted PDB configurations from MD simulations of hyaluronan octamer (HA8) with R4 and K4 tetrapeptides. Simulated explicit water was removed for clarity.</p> <p>See <a href="https://doi.org/10.5281/zenodo.8028600">10.5281/zenodo.8028600</a> for more details about simulated systems.</p>
Molecular dynamics simulations of lipid bilayers containing POPC and POPS (various mixtures) with ECC-lipids force field, and Na+ (K+) counterions
<p>Classical molecular dynamics simulations of various mixtures of POPC:POPS lipid bilayers in water solution with only Na+ counterions (or with K+ counterions when noted with "_KCl" suffix).</p> <p>ECC-lipids force field parameters used for lipids, SPC/E water model and ECC-ions, all parameters available at <a href="https://github.com/jmelcr/ecc_lipids">https://github.com/jmelcr/ecc_lipids</a></p> <p>simulations performed with Gromacs 2018.0 (*.xtc files)</p> <p>simulation length 1000 ns = 1 microsecond</p> <p>temperature 298 K</p> <p>Gromacs simulation setting is in the file npt_lipid_bilayer.mdp</p>
Data from: The impact of the tree prior on molecular dating of data sets containing a mixture of inter- and intraspecies sampling
In Bayesian phylogenetic analyses of genetic data, prior probability distributions need to be specified for the model parameters, including the tree. When Bayesian methods are used for molecular dating, available tree priors include those designed for species-level data, such as the pure-birth and birth-death priors, and coalescent-based priors designed for population-level data. However, molecular dating methods are frequently applied to data sets that include multiple individuals across multiple species. Such data sets violate the assumptions of both the speciation and coalescent-based tree priors, making it unclear which should be chosen and whether this choice can affect the estimation of node times. To investigate this problem, we used a simulation approach to produce data sets with different proportions of within- and between-species sampling under the multispecies coalescent model. These data sets were then analysed under pure-birth, birth-death, constant-size coalescent, and skyline coalescent tree priors. We also explored the ability of Bayesian model testing to select the best-performing priors. We confirmed the applicability of our results to empirical data sets from cetaceans, phocids, and coregonid whitefish. Estimates of node times were generally robust to the choice of tree prior, but some combinations of tree priors and sampling schemes led to large differences in the age estimates. In particular, the pure-birth tree prior frequently led to inaccurate estimates for data sets containing a mixture of inter- and intraspecific sampling, whereas the birth-death and skyline coalescent priors produced stable results across all scenarios. Model testing provided an adequate means of rejecting inappropriate tree priors. Our results suggest that tree priors do not strongly affect Bayesian molecular dating results in most cases, even when severely misspecified. However, the choice of tree prior can be significant for the accuracy of dating results in the case of data sets with mixed inter- and intraspecies sampling.
Data from: The impact of the tree prior on molecular dating of data sets containing a mixture of inter- and intraspecies sampling
Open the record for dataset details and reuse information.
Characterization of POP mixture redistribution and identification of their molecular signature in xenografted fat mice
GEO Series GSE279353. Mus musculus. 54 samples. Type: Expression profiling by high throughput sequencing.
Transcriptomic Profiling of Plasma Extracellular Vesicles Enables Reliable Annotation of the Cancer-specific Transcriptome and Molecular Subtype (cell mixture)
GEO Series GSE255774. Homo sapiens. 31 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.