Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,878

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3,878 results for “Molecular data”

Learn how ShareScore rates datasets ↗
dryad36/100

Data from: Benchmarking ultra-high molecular weight DNA preservation methods for long-read and long-range sequencing

<p>Studies in vertebrate genomics require sampling from a broad range of tissue types, taxa, and localities. Recent advancements in long-read and long-range genome sequencing have made it possible to produce high-quality chromosome-level genome assemblies for almost any organism. However, adequate tissue preservation for the requisite ultra-high molecular weight DNA (uHMW DNA) remains a major challenge. Here we present a comparative study of preservation methods for field and laboratory tissue sampling, across vertebrate classes and different tissue types. We find that no single method is best for all cases. Instead, the optimal storage and extraction methods vary by taxa, by tissue, and by down-stream application. Therefore, we provide sample preservation guidelines that ensure sufficient DNA integrity and amount required for use with long-read and long-range sequencing technologies across vertebrates. Our best practices generate the uHMW DNA needed for the high-quality reference genomes for Phase 1 of the Vertebrate Genomes Project (VGP), whose ultimate mission is to generate chromosome-level reference genome assemblies of all ~70,000 extant vertebrate species.</p>

opencc-zeroApr 2022View details →
dryad36/100

Morphological and molecular data support the separate species status of Syringa fauriei in Korea

<p><span><em>Syringa fauriei </em>is an unresolved lilac taxon (Oleaceae) native to specific geographical areas and present in the rivers and valleys of Gangwon province in South Korea. It was first identified on Mt. Geumgang by Faurie in 1906 and recognized as an independent species. Nevertheless, <em>S. fauriei </em>has been controversial as a synonymous species for <em>S. reticulata </em>subsp.<em> amurensis </em>without clear evidence. Therefore, we conducted this study on the taxonomic position of <em>S. fauriei</em> by examining the morphological and molecular differences between <em>S. fauriei </em>and its related taxa, the <em>S. reticulata</em> complex. Phylogenetic relationships were also inferred using molecular markers of nuclear ribosomal DNA internal transcribed spacer (ITS1–5.8S–ITS2; ITS) and external transcribed spacer regions. Consequently, <em>S. fauriei </em>and its related taxa were classified into two distinct clades: <em>S. fauriei</em> and the <em>S. reticulata </em>complex. Morphological and molecular data revealed that <em>S. fauriei </em>was endemic to Korea. The species characteristics of <em>S. fauriei </em>were further described for taxonomic re-establishment.</span></p>

opencc-zeroMay 2022View details →
dryad36/100

Data for a preliminary molecular phylogeny of the family Hydroptilidae (Trichoptera): exploring the combination of targeted enrichment data and legacy Sanger sequence data

<p><span>The purpose of this study is to provide a proof-of-concept that the use of molecular data, particularly targeted enrichment data, and statistically supported methods of analysis can result in the construction of a stable phylogenetic framework for the microcaddisflies (Trichoptera: Hydroptilidae). Here, we use a combination of targeted enrichment data for ca. 300 nuclear protein-coding genes and legacy (Sanger-based) sequence data for the mitochondrial COI gene and partial sequence from the 28S rRNA gene.</span></p>

opencc-zeroJun 2022View details →
dryad36/100

Molecular data and analysis specifications for a study on Atriophallophorus parasites from New Zealand

<p>This dataset accompanies the manuscript "Phylogeography and Cryptic Species Structure of a Locally Adapted Parasite in New Zealand". In our study we have shown that multiple species of the trematode parasite from the <em>Atriophallophorus</em> genus co-exist within the same lakes. When focusing on the most common of these species, <em>Atriophallophorus winterbourni</em>, we found that the Southern Alps and Pleistocene glaciation are a likely explanation for its phylogeographic patterns.</p> <p>This dataset consists of three parts: (I) An Excel file containing SNP data, all sample information and the output of phylogeographic analysis with MASCOT<sup>1</sup>, (II) all sequence alignments, including NADH5, 28S, ITS2 and COI, as they are used for the manuscript and (III) input files for the various analyses as they were prepared for the analysis.</p> <p><sup>1</sup>MASCOT: marginal approximation of the structured coalescent as implemented in BEAST2</p>

opencc-zeroJun 2022View details →
zenodo36/100

Data for "Quantum-corrected thickness-dependent thermal conductivity in amorphous silicon predicted by machine learning molecular dynamics simulations"

<p>This is the data set for the preprint&nbsp;<a href="https://arxiv.org/abs/2206.07605">arXiv:2206.07605</a>&nbsp;[cond-mat.mtrl-sci], obtained by the GPUMD code.</p> <p>Here are 6 directories.<br> &nbsp;&nbsp; &nbsp;1). NEMD<br> &nbsp;&nbsp; &nbsp;2). NEPpotential<br> &nbsp;&nbsp; &nbsp;3). PDOS<br> &nbsp;&nbsp; &nbsp;4). kappa-quenchRate<br> &nbsp;&nbsp; &nbsp;5). kappa-size<br> &nbsp;&nbsp; &nbsp;6). kappa-temperature<br> &nbsp;&nbsp; &nbsp;<br> 1). NEMD directory contains calculations of ballistic conductance using NEMD method, where 6 independent cycles are run to average.</p> <p>2). NEPpotential directory is the trained NEP potential.</p> <p>3). PDOS directory contains phonon density of states of a-Si samples generated by the quench rate of 10^{11} K/s.</p> <p>4). kappa-quenchRate directory contains HNEMD calculations of a-Si samples which are prepared using melt-quench temperature protocols with the quench rates covering from 10^{11} to 5x10^{12} K/s. In each case, 3 independent cycles are run.</p> <p>5). kappa-size directory contains HNEMD calculations based on different supercells. 6 independent cycles are run.</p> <p>6). kappa-temperature directory contains HNEMD calculations of a-Si samples which are prepared for different targeted temperatures using slow quench rate of 10^{11} K/s.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
dryad36/100

Underlying data for 'Rapid molecular assays for the detection of the four dengue viruses in infected mosquitoes'

<p>The pantropic emergence of severe dengue disease can partly be attributed to the co-circulation of different dengue viruses (DENVs) in the same geographical location. Effective monitoring for circulation of each of the four DENVs is critical to inform disease mitigation strategies. In low resource settings, this can be effectively achieved by utilizing inexpensive, rapid, sensitive and specific assays to detect viruses in mosquito populations. In this study, we developed four rapid DENV tests with direct applicability for low-resource virus surveillance in mosquitoes. The test protocols utilize a novel sample preparation step, a single-temperature isothermal amplification, and a simple lateral flow detection. Analytical sensitivity testing demonstrated tests could detect down to 1,000 copies/µL of virus-specific DENV RNA, and analytical specificity testing indicated tests were highly specific for their respective virus, and did not detect closely related flaviviruses. All four DENV tests showed excellent diagnostic specificity and sensitivity when used for detection of both individually infected mosquitoes and infected mosquitoes in pools of uninfected mosquitoes. With individually infected mosquitoes, the rapid DENV-1, -2 and -3 tests showed 100% diagnostic sensitivity (95% CI = 69% to 100%, n=8 for DENV-1; n=10 for DENV 2,3) and the DENV-4 test showed 92% diagnostic sensitivity (CI: <span>62% to 100%, n=12</span>) along with 100% diagnostic specificity (CI: 48–100%) for all four tests. Testing infected mosquito pools, the rapid DENV-2, -3 and -4 tests showed 100% diagnostic sensitivity (95% CI = 69% to 100%, n=10) and the DENV-1 test showed 90% diagnostic sensitivity (<span>55.50% to 99.75%, n=10</span>) together with 100% diagnostic specificity (CI: 48–100%). Our tests reduce the operational time required to perform mosquito infection status surveillance testing from &gt; two hours to only 35 minutes, and have potential to improve accessibility of mosquito screening, improving monitoring and control strategies in low-income countries most affected by dengue outbreaks.</p>

opencc-zeroJun 2022View details →
dryad36/100

Data from: Combined analysis of extant Rhynchonellida (Brachiopoda) using morphological and molecular data

Independent molecular and morphological phylogenetic analyses have often produced discordant results for certain groups which, for fossil-rich groups, raises the possibility that morphological data might mislead in those groups for which we depend upon morphology the most. Rhynchonellide brachiopods, with more than 500 extinct genera but only 19 extant genera represented today, provide an opportunity to explore the factors that produce contentious phylogenetic signal across datasets, as previous phylogenetic hypotheses generated from molecular sequence data bear little agreement with those constructed using morphological characters. Using a revised matrix of 66 morphological characters, and published ribosomal DNA sequences, we performed a series of combined phylogenetic analyses to identify conflicting phylogenetic signals. We completed a series of parsimony-based and Bayesian analyses, varying the data used, the taxa included, and the models used in the Bayesian analyses. We also performed simulation-based sensitivity analyses to assess whether the small size of the morphological data partition relative to the molecular data influenced the results of the combined analyses. In order to compare and contrast a large number of phylogenetic analyses and their resulting summary trees, we developed a measure for the incongruence between two topologies, and simultaneously ignore any differences in phylogenetic resolution. Phylogenetic hypotheses generated using only morphological characters differed amongst each other, and with previous analyses, while molecular-only and combined Bayesian analyses produced extremely similar topologies. Characters historically associated with traditional classification in the Rhynchonellida have very low consistency indices on the topology preferred by the combined Bayesian analyses. Overall, this casts doubt on the use of morphological systematics to resolve relationships among the crown rhynchonellide brachiopods. However, expanding our dataset to a larger number of extinct taxa with intermediate morphologies is necessary to exclude the possibility that the morphology of extant taxa is not dominated by convergence along long branches.

opencc-zeroDec 2016View details →
zenodo36/100

Discovering molecular regulators of ageing using mixture models with RNA-sequencing data

<p>Identifying the molecular regulators that control ageing is challenging because the ageing process is influenced by a combination of genetic and environmental factors which makes it difficult to source the contribution of a single gene. Multiple studies have demonstrated that as humans age, increased gene expression heterogeneity results in the dysregulation of key regulators and pathways. Given the dynamic nature of gene expression, it is vital that this data be modelled by statistical approaches that can appropriately account for changes in variability to understand the contribution of heterogeneity during the aging process and properly identify its regulators. This study demonstrates the utility of using mixture models to model biological variability of gene expression occurring during ageing and how novel potential regulators of ageing can be identified.</p> <p>Our mixture modelling approach was applied to gene expression data from the Genotype-Tissue Expression (GTEx) cohort. For every gene, the expression profile was modelled using a mixture model across the cohort where the subset of donors corresponding to each mode was tested for a significant change in age group. The multi-tissue aspect of GTEx was leveraged to find ageing regulators based on this mixture model approach genes that were common across multiple tissues, suggesting that the regulation of ageing may also be controlled through a set of genes that have non-tissue-specific activity.</p> <p>Our approach identified well-documented ageing regulators <em>mTOR </em>and <em>RICTOR</em> and other potential ageing regulators such as <em>IL4</em> and <em>GPR4</em> which were detected only by our approach. Genes identified by edgeR, DESeq2 and the mixture model-based approach were enriched for similar biological pathways. This suggests that while the specific ageing regulators identified from our approach may be distinct, they generally belong in the same pathways as the genes identified by standard approaches. Overall, these results indicate that modelling gene expression variability using mixture models in conjunction with standard differential gene expression can help uncover new regulators that have a potential role for understanding human ageing.</p> <p>I</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Data for Functional Group Pair Distance Based Descriptor for Isomerisation in Porous Molecular Framework Materials

<p>This is a&nbsp;dataset of&nbsp;isomer structure files&nbsp;for&nbsp;pore topology: Tri2Di3, Tri4Di6, Tri4-2Di6,Tri6Di9, Tet2Di4, Tet3Di3, Tet4-4Di8, Tet5Di10, and Tet6Di12.&nbsp;</p> <p>All.tar.bz2 contains all pore topologies, the total disk space after unzipping the bundle is 1.8 Gb. The total disk space for pore Tet6Di12 alone is 1.5Gb.</p> <p>The base structure&nbsp;of all pore topologies are constructed using a metal node of radius ~5 (represented by a Zirconium atom) and a benzene linker, while the functional group is represented by a Nitrogen atom.</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Sintering of alumina nanoparticles: comparison of interatomic potentials, molecular dynamics simulations, and data analysis

<p>This is the dataset for the publication in MSMSE 2022 containing all plot scripts and data for reproducing all figures. The dataset is a snapshot of the repository https://gitlab.com/computational-materials-science/public/publication-data-and-code/2022_MSMSE_Roy_et_al_MD-sintering (SHA 7ad2f421deb055f3384c00ba29f2fb1acd0e78ea) that might contain additional/newer&nbsp;data and scripts.</p>

opencc-by-4.0Jul 2022View details →
dryad36/100

Phylogenetic position of Centroglossa and Dunstervillea (Ornithocephalus clade: Oncidiinae: Orchidaceae) based on molecular and morphological data

<p><span>Even though the monophyly of the <em>Ornithocephalus </em>clade (Oncidiinae) is currently well defined, the systematic positioning of <em>Centroglossa </em>and <em>Dunstervillea </em>remains obscure in the clade due to the absence in previous phylogenetic studies. <em>Centroglossa </em>has a very similar habit and is indistinguishable from <em>Zygostates</em>, whereas <em>Dunstervillea </em>has as its main characteristic the calcarate labellum, also found in <em>Centroglossa</em>. We clarify the systematic and phylogenetic positioning of Centroglossa and Dunstervillea in the <em>Ornithocephalus </em>clade (OC) through analysis of maximum likelihood, <em>Bayesian </em>inference, and maximum parsimony from molecular data (nrITS and <em>matK </em>cpDNA) and morphology. Our results indicate that Dunstervillea is phylogenetically close to <em>Eloyella</em>; both genera have a psigmoid habit, single-sided and flattened leaves, floral perianth with the same coloring, petals with entire margins, and a short rostellum. <em>Centroglossa </em>appears as a subclade within <em>Zygostates</em>. In addition to several homoplastic features, these two genera have the dorsal position of viscidium as a synapomorphy. The calcarate labellum, common to <em>Centroglossa </em>and <em>Dunstervillea</em>, originated more than once in the OC. Based on the phylogenetic results, we propose the nomenclatural changes to include <em>Dunstervillea </em>in <em>Eloyella </em>and Centroglossa in Zygostates. Lectotypes are indicated to <em>Centroglossa macroceras </em>and <em>C</em>. <em>glaziovii</em>.</span></p>

opencc-zeroAug 2022View details →
zenodo36/100

Raw NGS Data for "Deep mutational scanning and machine learning reveal structural and molecular rules governing allosteric hotspots in homologous proteins"

<p>This directory contains relevant fastq files used for deep sequencing analysis in the publication &ldquo;Deep mutational scanning and machine learning reveal structural and molecular rules governing allosteric hotspots in homologous proteins&rdquo;.&nbsp;</p> <p>Fastq files are provided for presorted, uninduced and induced populations from DMS experiments of&nbsp;four&nbsp;homologs (TtgR, TetR, RolR, and MphR). Three replicates were performed for each sample.</p> <p>Data analysis of this&nbsp;deep sequencing data was performed using custom scripts, which are described in the methods section of the publication.</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

BioExcel Use Case 1: collection of output data from molecular dynamics simulation

<p>The Use Case aims to address all the challenges related to antibody design through an integrative approach combining the core BioExcel software comprising of GROMACS, HADDOCK and PMX.</p> <p>The folder&nbsp; contains the GROMACS output files (xtc and pdb file). Molecular Dynamics simulations have been performed with GROMACS version 2020 and CHARMM36 force field. The input files and scripts of the final protocol are publicly available on BioExcel GitHub https://github.com/bioexcel/BioExcel-UseCase1.</p> <p>The Use Case 1 protocol was presented at the BioExcel Summer School on Biomolecular Simulation in 2021 (see <a href="https://doi.org/10.5281/zenodo.7009238">https://doi.org/10.5281/zenodo.7009238</a> or <a href="https://youtu.be/_TDKfKX4kwM">https://youtu.be/_TDKfKX4kwM</a>)</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo36/100

Data samples for Flow-matching -- efficient coarse-graining molecular dynamics without forces

<p>CG samples generated during the training and validation processes in the flow-matching project. Accompanying the preprint &quot;Flow-matching -- efficient coarse-graining molecular dynamics without forces&quot;: https://arxiv.org/abs/2203.11167. Detailed descriptions can be found in the preprint as well as the included README.</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Molecular dynamics simulations data of Caspase-3 enzyme with pentapeptide ligand DEVDG and its chiral mutant DEVdG having D-Asp at fourth position

<p>Amino acids in proteins are maintained in one specific L chiral form in the body. D-amino acids are not normally incorporated into proteins and their accumulation has been associated with several conditions including schizophrenia, amyotrophic lateral sclerosis, and other age-related disorders. However, the mechanisms by which the accumulation of D-amino-acids in proteins may lead to pathophysiological consequences remain poorly understood. In this work, we studied a model protease system, caspase-3 that specifically hydrolyses the 4&rsquo;&ndash;5&rsquo; peptide bond of the pentapeptide substrate DEVDG. Through extensive molecular dynamics simulations, free energy calculations and distance maps, we reveal that caspase-3 naturally rejects the pentapeptide containing D-Asp substrate, DEVdG and prevents catalytic activity by caspase. The importance of this chiral discriminating capacity is evident from chiral-selective in vivo experimental assays to detect caspase-bound D-Asp in Drosophila where altering the chiral balance created impaired caspase activity and impaired apoptosis, increased tumour formation, and premature death. The modelling data reveals the molecular level charge balancing that enforces the chiral recognition necessary to maintain homeostasis across the cell, tissue, and organ level.</p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

Data supplementing the article "Avoiding quantification bias in metabarcoding: application of a cell biovolume correction factor in diatom molecular biomonitoring" V. Vasselon, A. Bouchez, F. Rimet, S. Jacquet, R. Trobajo, M. Corniquel, K. Tapolczai, I. Domaizon submitted to Methods in Ecology and Evolution journal

<p>These data supplement the article &quot;Avoiding quantification bias in metabarcoding: application of a cell biovolume correction factor in diatom molecular biomonitoring&quot; V. Vasselon, A. Bouchez, F. Rimet, S. Jacquet, R. Trobajo, M. Corniquel, K. Tapolczai, I. Domaizon submitted to Methods in Ecology and Evolution journal</p> <p>The directory contains the following files:</p> <p>1<strong>5&nbsp;fastq files raw reads (5 mock communities, 3 replicates)</strong><strong>.rar </strong>- contains the 15&nbsp;fastq files provided by the sequencing platform with demultiplexed DNA reads (raw data prior any bioinformatics treatments).</p> <p><strong>15 fastq files information.xlsx</strong>&nbsp;:</p> <p>- contains the information relative to the 15 fastq files corresponding to the PGM raw data of the 5 mock communities (sequenced with 3 replicates), including:&nbsp;the ID of the fastq files, the mock community name,&nbsp;the replicate number, the final sample Id and the number of raw reads per fastq file.</p> <p>- contains the information of the proportion of the 8 diatoms species (%) used to create the 5 mock communities (estimated from microscopy).</p>

opencc-by-4.0Jun 2017View details →
zenodo36/100

Data from "Allostery and evolution: a molecular journey throught the structural and dynamical landscape of an enzyme super family."

<p>This data&nbsp;accompanies the paper&nbsp;entitled Allostery and evolution: a molecular journey throught the structural and dynamical landscape of an enzyme super family.</p> <p>The zip archive contains:&nbsp;</p> <p>1- Starting configurations of the proteins after equilibration in PDB format and trajectories of unrestrained molecular dynamics simulations with the positions of the proteins every 100 ps in XTC gromacs format are provided for all systems.&nbsp;</p> <p>2- The free energy profiles and histograms are provided for all umbrella sampling simulations and the scripts used to run it with gromacs.</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Data for "Water Sorption Controls Extreme Single-Crystal-to-Single Crystal Molecular Reorganization in Hydrogen Bonded Organic Frameworks"

<p>Paper DOI: <a href="https://doi.org/10.1002/chem.202201929">10.1002/chem.202201929</a></p> <p>Previously uploaded in 10.5281/zenodo.8432296 and&nbsp;<a href="https://github.com/andrewtarzia/citable_data" rel="noopener noreferrer">https://github.com/andrewtarzia/citable_data</a></p> <p>Each .out file is generated from zeo_runs_production.py, which includes the output from Zeo++ for the probe radius and sampling value in the file name.</p> <p>zeo_runs.py tests sampling values to check for convergence, those output files are not included here.</p> <p>Each python script includes the list of CIFs to run the analysis on. Only the CIFs shown in the manuscript are included here, as testing was done on a series to see the effect of symmetry, disorder and cell size.</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Molecular Dynamics (MD) Simulation Data for Dynamics Underlie the Drug Recognition Mechanism by the Efflux Transporter EmrE

<p>MD simulations on the proton bound (PDB 8UWU), deprotonated on E14A (PDB 8UWU), TPP Bound (PDB 8UWU) on our NMR derived structures.</p> <p>&nbsp;</p> <p>MD simulations on the proton bound (7MH6) and deprotonated on E14A (7MH6) on X-ray structures.&nbsp;</p> <p>&nbsp;</p> <p>Total raw simulation data would be too large for uploading to repositories.&nbsp;&nbsp;To reduce size of file, starting structure and tpr files are uploaded.&nbsp;&nbsp;Final structure at 2.5 &mu;s are also uploaded.&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Data supporting: "Interaction of MRI Contrast Agent [Gd(DOTA)]− with Lipid Membranes: A Molecular Dynamics Study"

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record