Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
409
datasets available to search
ShareScore release 0.7.1
Dataset results
409 results for “molecular genetics”
Linked collectors and determiners for: Endangered beauties: micro-CT cranial osteology, molecular genetics and external morphology reveal three new species of chameleons in the Calumma boettgeri complex (Squamata: Chamaeleonidae).
Natural history specimen data linked to collectors and determiners held within, "Endangered beauties: micro-CT cranial osteology, molecular genetics and external morphology reveal three new species of chameleons in the Calumma boettgeri complex (Squamata: Chamaeleonidae)". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/99b3b009-f006-4146-b335-cbfdc2fcd60b">https://bionomia.net/dataset/99b3b009-f006-4146-b335-cbfdc2fcd60b</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/99b3b009-f006-4146-b335-cbfdc2fcd60b">https://gbif.org/dataset/99b3b009-f006-4146-b335-cbfdc2fcd60b</a>. Formatted as a Frictionless Data package.
Merging metabolomics and genomics provides a catalog of genetic factors that infuence molecular phenotypes in pigs linking relevant metabolic pathways
<h3>Content</h3> <p>Metabolites included in the study. Summary statistics of metabolite levels for the Large White and Duroc pig populations are provided.</p>
Data from: Drift happens: molecular genetic diversity and differentiation among populations of jewelweed (Impatiens capensis Meerb.) reflect fragmentation of floodplain forests
Open the record for dataset details and reuse information.
Data from: The molecular biogeography of the Indo-Pacific: testing hypotheses with multispecies genetic patterns
Aim: To test hypothesized biogeographic partitions of the tropical Indo-Pacific Ocean with phylogeographic data from 56 taxa, and to evaluate the strength and nature of barriers emerging from this test. Location: The Indo-Pacific Ocean. Time Period: Pliocene through the Holocene. Major Taxa Studied: 56 marine species. Methods: We tested eight biogeographic hypotheses for partitioning of the Indo-Pacific using a novel modification to analysis of molecular variance. Putative barriers to gene flow emerging from this analysis were evaluated for pairwise ΦST, and these ΦST distributions were compared to distributions from randomized datasets and simple coalescent simulations of vicariance arising from the Last Glacial Maximum. We then weighed the relative contribution of distance vs. environmental or geographical barriers to pairwise ΦST with a distance-based redundancy analysis (dbRDA). Results: We observed a diversity of outcomes, although the majority of species fit a few broad biogeographic regions. Repeated coalescent simulation of a simple vicariance model yielded a wide distribution of pairwise ΦST that was very similar to empirical distributions observed across five putative barriers to gene flow. Three of these barriers had median ΦST that were significantly larger than random expectation. Only 21 of 52 species analyzed with dbRDA rejected the null model. Among these, 15 had overwater distance as a significant predictor of pairwise ΦST, while 11 were significant for geographical or environmental barriers other than distance. Main Conclusions: Although there is support for three previously described barriers, phylogeographic discordance in the Indo-Pacific oceans indicates incongruity between processes shaping the distributions of diversity at the species and population levels. Among the many possible causes of this incongruity, genetic drift provides the most compelling explanation: given massive effective population sizes of Indo-Pacific species, even hard vicariance for tens of thousands of years can yield ΦST values that range from 0 to nearly 0.5.
Figure 1 in Guidelines and quantitative standards to improve consistency in cetacean subspecies and species delimitation relying on molecular genetic data
Figure 1. Guidelines for studies of cetacean taxonomy based on genetic data.
Comparative transcriptomics reveals the molecular genetic basis of pigmentation loss in Sinocyclocheilus cavefishes
<p><span><span><span>Cave-dwelling animals evolve distinct troglomorphic traits, such as loss of eyes, skin pigmentation, and augmentation of senses following long-term adaptation to the perpetual darkness. However, the molecular genetic mechanisms underlying these phenotypic variations remain unclear. In this study, we conducted comparative histology and comparative transcriptomics study of the skin of eight <i>Sinocyclocheilus</i> species (Cypriniformes: Cyprinidae) that included surface and cave-dwelling species. We analyzed four surface and four cavefish species by using next-generation sequencing, and a total of 802,798,907 clean reads were generated and assembled into 505,495,009 transcripts, which contributed to 1,037,334 unigenes. Bioinformatics comparisons of four different surface-cave fish groups revealed between 10,629 and 6442 significantly differentially expressed unigenes. Further, tens of differentially expressed genes (DEGs) potentially related to skin pigmentation were identified. Most of these DEGs (including <i>GNAQ</i>, <i>PKA</i>, <i>NRAS</i>, and <i>p38</i>) are downregulated in cavefish species. They are involved in key signaling pathways of pigment synthesis, such as the melanogenesis, Wnt, and MAPK pathways. This trend of downregulation was confirmed through qPCR experiments. This study will deepen our understanding of the formation of troglomorphic traits in cavefishes.</span></span></span></p>
Molecular Genetic Analysis of SARS-CoV-2 Lineages in Armenia - additional data
<p>Sequencing of SARS-CoV-2 provides essential information on viral evolution, transmission, and epidemiology. In this study, we performed whole-genome sequencing of SARS-CoV-2 using nanopore and Illumina short-read sequencing to describe the circulation of the virus lineage in Armenia.</p> <p>This dataset contains Nextstrain configuration files, the auspice JSON file, BEAST output logs, and trees files, and resulting log and tree files as well as R scripts and data files used in phylogenetic and functional analyses. </p> <p> </p>
Collection data and molecular datasets for: Defining species-specific seed sourcing strategies for restoration: An example of how to use genetic data to inform seed collections for multiple co-occurring species
<p>Two files for each dataset are provided:</p> <p>Metadata files contain colelcting information for the samples in each molecular dataset as well as the group assignments (species, genetic neighbourhood and sites) used in analyses, saved as an excel spreadsheet.</p> <p>Molecular datasets containing samples and SNPs used in analyses. The data is formatted as a genlight object saved as an RData file that can be read into the R statisical environment and analysed using the 'dartR' package (Gruber et al. 2018).</p>
Genetic variation influencing DNA methylation provides new insights into the molecular pathways regulating genomic function - Selected Supplementary Tables
<p><strong>Selected Supplementary Tables - ST5, 7, 8 and 9</strong></p> <p><strong>Supplementary Table 5. Cosmopolitan results. </strong>Cosmopolitan SNP-CpG associations identified through genome-wide association amongst Europeans and South Asians.</p> <p><strong>Supplementary Table 7. Cross-tissue replication.</strong> Results for further testing of the 11,165,559 cosmopolitan SNP-CpG associations identified (by genome-wide association in blood), in 4 isolated white cell subsets (CD4+ lymphocytes, CD8+ lymphocytes, monocytes and neutrophils), in adipocytes isolated from subcutaneous adipose tissue or visceral adipose tissue, and in whole adipose tissue.</p> <p><strong>Supplementary Table 8. Conditional analysis.</strong> Results of conditional analysis to identify SNPs independently associated with each of the ~360K CpG sites tested. </p> <p><strong>Supplementary Table 9. Sentinel SNPs and CpGs.</strong> Results of R2 pruning and locus merging to identify discrete genetic and methylation loci that are associated, and their respective sentinel SNPs and sentinel CpG sites. </p> <p>Other files (e.g. annotation files and 'intermediate' processing files) referenced in our code are also provided. </p> <p> </p>
Supplementary data for "Testing the efficacy of different molecular tools for parasite conservation genetics: a case study using horsehair worms (Phylum Nematomorpha)"
<p>Supplementary data for "Testing the efficacy of different molecular tools for parasite conservation genetics: a case study using horsehair worms (Phylum Nematomorpha)"</p> <p>alignments: alignments used for BEAST ("bayes") and PopArt ("popart"). The "popart" folder also has a traits file per each species.</p> <p>bayesian_plots: TSVs ("tsv") and PDF files ("ogs") generated by BEAST. The "tsv" folder also has the scripts for plotting the results in R.</p> <p>easysfs: scripts, population file and results from the VCF to SFS conversione done by easySFS.</p> <p>fineRADstructure: fineRADstructure input files and output PDF plots ("plots") for <em>C. formosanus</em> ipyrad and Stacks ("stacks") data. </p> <p>logs: logs for ipyrad, ModelTest, PGDspider, PopArt ("popart") and Stacks ("stacks"). The "popart" folder also have the generated networks in a TXT file. The "stacks" folder also has ODS files for calculating the amount of loci per each M/n fixed value.</p> <p>snapclust: STR files used with R for snapclust. Scripts included.</p> <p>stairway_plot: input (blueprint files) and outputs for Stairway Plot 2 analyses. The <em>C. formosanus</em> folder ("chordodes") also has scripts for R plotting.</p> <p>vcfs: VCF and HDF5 files used in this study. Also scripts for filtering/converting data and plotting the PCA with ipyrad (activate python first!) for <em>C. formosanus</em>.</p> <p>"acutogordius" = <em>A. taiwanensis</em><br> "chordodes" = <em>C. formosanus</em><br> "gordius" = <em>G. chiashanus</em></p>
Prospective Evaluation of Immunological, Molecular-genetic, Image-based and Microbial Analyzes to Characterize Tumor Response and Control in Patients With Inoperable Stage III NSCLC Treated With Chemo
ClinicalTrials.gov study NCT05027165. IPD Sharing: NO. Countries: 1. Publications: 1.
Pleiotropic effects of trisomy and pharmacologic modulation on structural, functional, molecular, and genetic systems in a Down syndrome mouse model
Open the record for dataset details and reuse information.
Data from: The molecular biogeography of the Indo-Pacific: testing hypotheses with multispecies genetic patterns
Open the record for dataset details and reuse information.
Comparative transcriptomics reveals the molecular genetic basis of pigmentation loss in Sinocyclocheilus cavefishes
Open the record for dataset details and reuse information.
Data from: Molecular population genetics of the northern elephant seal Mirounga angustirostris
Open the record for dataset details and reuse information.
Data from: Small, but mitey: investigating the molecular genetic basis for mite domatia development and intraspecific variation in Vitis riparia using transcriptomics
Open the record for dataset details and reuse information.
Data from: Disease swamps molecular signatures of genetic-environmental associations to abiotic factors in Tasmanian devil (Sarcophilus harrisii) populations
Landscape genomics studies focus on identifying candidate genes under selection via spatial variation in abiotic environmental variables, but rarely by biotic factors such as disease. The Tasmanian devil (Sarcophilus harrisii) is found only on the environmentally heterogeneous island of Tasmania and is threatened with extinction by a nearly 100% fatal, transmissible cancer, devil facial tumor disease (DFTD). Devils persist in regions of long-term infection despite epidemiological model predictions of species' extinction, suggesting possible adaptation to DFTD. Here, we test the extent to which spatial variation and genetic diversity are associated with the abiotic environment and/or DFTD. We employ genetic-environment association analyses using a RAD-capture panel including 6,886 SNPs from 3,286 individuals sampled pre- and post-disease arrival. Pre-disease, we find significant correlations of allele frequencies with environmental variables, including 365 unique loci linked to 71 genes, suggesting local adaptation to abiotic environment. The majority of candidate loci detected pre-DFTD were not detected post disease arrival. Several post-DFTD candidate loci were associated with disease prevalence and were in linkage disequilibrium with genes involved in tumor suppression and immune response. Loss of apparent signal of abiotic local adaptation post-disease suggests swamping by the strong selection resulting from the rapid onset of DFTD.
Data from: Friends and Family: a software program for identification of unrelated individuals from molecular marker data. And from: Genetic diversity, relatedness and inbreeding of ranched and fragmented Cape buffalo populations in southern Africa
The identification of related and unrelated individuals from molecular marker data is often difficult, particularly when no pedigree information is available and the data set is large. High levels of relatedness or inbreeding can influence genotype frequencies and thus genetic marker evaluation, as well as the accurate inference of hidden genetic structure. Identification of related and unrelated individuals is also important in breeding programmes, to inform decisions about breeding pairs and translocations. We present Friends and Family, a Windows executable program with a graphical user interface that identifies unrelated individuals from a pairwise relatedness matrix or table generated in programs such as COANCESTRY and GenAlEx. Friends and Family outputs a list of samples that are all unrelated to each other, based on a user-defined relatedness cut-off value. This unrelated data set can be used in downstream analyses, such as marker evaluation or inference of genetic structure. The results can be compared to that of the full data set to determine the effect related individuals have on the analyses. We demonstrate one of the applications of the program: how the removal of related individuals altered the Hardy-Weinberg equilibrium test outcome for microsatellite markers in an empirical data set. Friends and Family can be obtained from https://github.com/DeondeJager/Friends-and-Family.
Data from: Multi-objective optimization for plant germplasm collection conservation of genetic resources based on molecular variability
Germplasm collections play a significant role among strategies for conservation of diversity. It is common to select a core collection to represent the genetic diversity of a germplasm collection, in order to minimize the cost of conservation, while ensuring the maximization of genetic variation. We aimed to solve two main problems: (1) to select a set of individuals, from an in situ data set, that is genetically complementary to an existing germplasm collection, and (2) to define a core collection for a germplasm collection. We proposed a new multi-objective optimization (MOO) approach based on principles of systematic conservation planning (SCP) incorporating heterozygosity information; therefore, optimization takes genotypic diversity and variability patterns into account as well. As a case study, we used Dipteryx alata microsatellite loci information from two sources, an ex situ germplasm collection located at the Agronomy School of the Federal University of Goiás (UFG-AS), and an in situ data set composed of 642 sampled individual trees. We were able to identify within a population of several individuals, the exact accessions/samples that should be chosen in order to preserve the species diversity. We found that material from nine in situ individual trees are enough to complement the UFG-AS germplasm collection as it is, and that it is possible to define a core collection of 20 individual trees representing all studied genetic diversity. Moreover, we defined a method (a protocol) to deal with large amounts of accessions in the context of MOO. The proposed approach can be used to help constructing collections with maximal allelic richness and can also be extended to the in situ conservation. As far as we know, this is the first time that principles of SCP and the MOO approach are applied to the problem of complementing a germplasm collection and of finding a core collection for a germplasm collection.
Data from: A molecular genetic time scale demonstrates Cretaceous origins and multiple diversification rate shifts within the order Galliformes (Aves)
The phylogeny of Galliformes (landfowl) has been studied extensively; however, the associated chronologies have been criticized recently due to misplaced or misidentified fossil calibrations. As a consequence, it is unclear whether any crown-group lineages arose in the Cretaceous and survived the Cretaceous–Paleogene (K–Pg; 65.5 Ma) mass extinction. Using Bayesian phylogenetic inference on an alignment spanning 14,539 bp of mitochondrial and nuclear DNA sequence data, four fossil calibrations, and a combination of uncorrelated lognormally distributed relaxed-clock and strict-clock models, we inferred a time-calibrated molecular phylogeny for 225 of the 291 extant Galliform taxa. These analyses suggest that crown Galliformes diversified in the Cretaceous and that three-stem lineages survived the K–Pg mass extinction. Ideally, characterizing the tempo and mode of diversification involves a taxonomically complete phylogenetic hypothesis. We used simple constraint structures to incorporate 66 data-deficient taxa and inferred the first taxon-complete phylogenetic hypothesis for the Galliformes. Diversification analyses conducted on 10,000 timetrees sampled from the posterior distribution of candidate trees show that the evolutionary history of the Galliformes is best explained by a rate-shift model including 1–3 clade-specific increases in diversification rate. We further show that the tempo and mode of diversification in the Galliformes conforms to a three-pulse model, with three-stem lineages arising in the Cretaceous and inter and intrafamilial diversification occurring after the K–Pg mass extinction, in the Paleocene–Eocene (65.5–33.9 Ma) or in association with the Eocene–Oligocene transition (33.9 Ma).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.