Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
247
datasets available to search
ShareScore release 0.9.0
Dataset results
247 results for “genetic integration”
Data from: Dispersal in a house sparrow metapopulation: an integrative case study of genetic assignment calibrated with ecological data and pedigree information
<p class="western">Dispersal has a crucial role determining eco-evolutionary dynamics through both gene flow and population size regulation. However, to study dispersal and its consequences, one must distinguish immigrants from residents. Dispersers can be identified using telemetry, capture-mark-recapture (CMR) methods, or genetic assignment methods. All of these methods have disadvantages, such as, high costs and substantial field efforts needed for telemetry and CMR surveys, and adequate genetic distance required in genetic assignment. In this study, we used genome-wide 200K Single Nucleotide Polymorphism data and two different genetic assignment approaches (GSI_SIM, Bayesian framework; BONE, network-based estimation) to identify the dispersers in a house sparrow (<i>Passer domesticus</i>) metapopulation sampled over 16 years. Our results showed higher assignment accuracy with BONE. Hence, we proceeded to diagnose potential sources of errors in the assignment results from the BONE method due to variation in levels of inter-population genetic differentiation, intra-population genetic variation and sample size. We show that assignment accuracy is high even at low levels of genetic differentiation and that it increases with the proportion of a population that has been sampled. Finally, we highlight that dispersal studies integrating both ecological and genetic data provide robust assessments of the dispersal patterns in natural populations.</p>
Integrative genetic analyses illuminates ALS heritability and identifies novel risk genes
<p>Amyotrophic lateral sclerosis (ALS), the major adult onset motor neuron disease, has substantial heritability, in part shared with fronto-temporal dementia (FTD). We show here that ALS heritability is enriched in splicing variants and in binding sites of 6 RNA binding proteins including TDP-43 and FUS. A discovery and replication transcriptome wide association study (TWAS) identified 6 loci associated with ALS, 3 in known ALS loci (<em>C9ORF72, SCFD1, SLC9A8</em>) and 3 novel loci including <em>NUP50 </em>encoding for the nucleopore basket protein NUP50 In our meta-analysis of TWAS cohorts, <em>NUP50 </em>common variant was associated with ALS and to decreased expression of <em>NUP50 </em>in the central nervous system. Independently, we further show association of rare variants in <em>NUP50</em> with ALS risk (<em>P</em> = 3.71.10<sup>-03</sup>; odds ratio = 3.29; 95%CI, 1.37 to 7.87) in a cohort of 9,390 ALS/FTD patients and 4,594 controls. Cells from one patient carrying a <em>NUP50 </em>frameshift mutation displayed a decreased levels of NUP50. Loss of NUP50 leads to neuronal death in cultured neurons, and motor defects in <em>Drosophila </em>and zebrafish models. Thus, our study identifies alterations in splicing in neurons as a critical pathogenic process in ALS, uncovers several new loci potentially contributing to ALS, and provides genetic evidence linking nuclear pore defects to ALS.</p>
Yield Prediction Through Integration of Genetic, Environment, and Management Data Through Deep Learning: Cleaned Data
<p>The included files and script are to allow for reconstruction of the data directory and cleaned data used in "Yield Prediction Through Integration of Genetic, Environment, and Management Data Through Deep Learning" ( https://doi.org/10.1101/2022.07.29.502051 ). Code used is available at 10.5281/zenodo.7401113 .</p> <table> <tbody> <tr> <th>Filename</th> <th>Description</th> </tr> <tr> <td>interim.tar.gz</td> <td>Contains site grouping dictonary</td> </tr> <tr> <td>processed.tar.gz</td> <td>Processed data</td> </tr> <tr> <td>raw.tar.gz</td> <td>Input data</td> </tr> <tr> <td>SetupInstructions.sh</td> <td>Bash script to prepare folders and unzipped data expected by code in 10.5281/zenodo.7401113</td> </tr> <tr> <td>SetupInstructions.txt</td> <td>Instructions for unzipping the data</td> </tr> <tr> <td>Train_Test_Split_Reference_Phenotypes.csv</td> <td>Reference spreadsheet to allow for easily exploring training and test set groupings</td> </tr> </tbody> </table> <ul> </ul> <p>This work was supported through funding from the USDA Agricultural Research Service, ARS project number 5070-21000-041-000-D. Raw data provided by the [Genomes to Field Initiative](https://www.genomes2fields.org/) and the [Daymet database](https://daymet.ornl.gov/).</p>
Figure 5 in Genetic integrity of the European grayling (Thymallus thymallus) populations within the Vienne River drainage basin after five decades of stockings
Figure 5 - Results of the Structure analysis (Q-values) for 8 loci (A) and 10 loci (B). Colours represent the different "k" units chosen using the approach of Evanno et al. (2005). Population names are shown on the x-axis, and Q-values (contribution or assignment from each k partition) are shown on the Y-axis. For 8 loci, k = 8, and for 10 loci k = 11.
Figure 2 in Genetic integrity of the European grayling (Thymallus thymallus) populations within the Vienne River drainage basin after five decades of stockings
Figure 2. - Close-up of grayling distribution and four sampled sites (stars) within the upper Vienne district in 2012 (relief background from IGN-Geoportail).
Figure 4 in Genetic integrity of the European grayling (Thymallus thymallus) populations within the Vienne River drainage basin after five decades of stockings
Figure 4. - Factorial Correspondence Analyses (FCA) of individuals based on the presence and absence of microsatellite alleles. A: Bi-variate plot of the first two FCA factors (F1 & F2), graphically partitioned by population. B: Bi-variate plot of the third and fourth FCA factors (F3 & F4) partitioned by population. The percentage inertia for each factor is as follows: FC1 = 4.44%; FC2 = 4.03%; FC3 = 3.46%; FC4 = 2.82%.
Figure 1 in Genetic integrity of the European grayling (Thymallus thymallus) populations within the Vienne River drainage basin after five decades of stockings
Figure 1. - Distribution of grayling in Europe (green) and its former distribution in France around 1900 (red). Sample sites (white stars = wild populations, black stars = hatcheries) are numbered following table II.
Figure 3 in Genetic integrity of the European grayling (Thymallus thymallus) populations within the Vienne River drainage basin after five decades of stockings
Figure 3. - Median Joining network of haplotypes found in this study augmented by additional published haplotypes provided for reference. All new haplotypes use the simple abbreviation "Ht" with a serial number reflecting the order in which they were first found in our data. The small black dots indicate substitutional steps whereas the small red dots represent multifurcating nodes. The following haplotypes and corresponding clade names stem from Weiss et al. (2002). "Danube drainage (Northern Alps)": Da1, Da2, Da4, Da11; "Mixed Central Europe": Rh1, Rh4, Rh6 all first reported from the Rhone basin; "Mixed Central Europe": At14, At15 both frequent and widespread in the Rhine basin; "Scandinavia": At6 reported from Finland; and AT1 first reported from the Loire River.
High-density genetic linkage mapping in Sitka spruce advances the integration of genomic resources in conifers
Open the record for dataset details and reuse information.
Data from: harnessing the power of regional baselines for broad-scale genetic stock identification: a multistage, integrated, and cost-effective approach
Open the record for dataset details and reuse information.
Data from: Dispersal in a house sparrow metapopulation: an integrative case study of genetic assignment calibrated with ecological data and pedigree information
Open the record for dataset details and reuse information.
Chromonomer: a tool set for repairing and enhancing assembled genomes through integration of genetic maps and conserved synteny
<p class="BodyAA">The pace of the sequencing and computational assembly of novel reference genomes is accelerating. Though DNA sequencing technologies and assembly software tools continue to improve, biological features of genomes such as repetitive sequence as well as molecular artifacts that often accompany sequencing library preparation can lead to fragmented or chimeric assemblies. If left uncorrected, defects like these trammel progress on understanding genome structure and function, or worse, positively mislead this research. Fortunately, integration of additional, independent streams of information, such as a marker-dense genetic map and conserved orthologous gene order from related taxa, can be used to scaffold together unlinked, disordered fragments and to restructure a reference genome where it is incorrectly joined. We present a tool set for automating these processes, one that additionally tracks any changes to the assembly and to the genetic map, and which allows the user to scrutinize these changes with the help of web-based, graphical visualizations. Chromonomer takes a user-defined reference genome, a map of genetic markers, and, optionally, conserved synteny information to construct an improved reference genome of chromosome models: a "chromonome". We demonstrate Chromonomer's performance on genome assemblies and genetic maps that have disparate characteristics and levels of quality.</p>
Data from: Integrating population genetics to define conservation units from the core to the edge of Rhinolophus ferrumequinum western range
The greater horseshoe bat (<i>Rhinolophus ferrumequinum</i>) is among the most widespread bat species in Europe but it has experienced severe declines, especially in Northern Europe. This species is listed Near Threatened in the European IUCN Red List of Threatened Animals and it is considered to be highly sensitive to human activities and particularly to habitat fragmentation. Therefore, understanding the population boundaries and demographic history of populations of this species is of primary importance to assess relevant conservation strategies. In this study, we used 17 microsatellite markers to assess the genetic diversity, the genetic structure and the demographic history of <i>R. ferrumequinum</i> colonies in the western part of its distribution. We identified one large population showing high levels of genetic diversity and large population size. Lower estimates were found in England and northern France. Analyses of clustering and isolation by distance suggested that the Channel and the Mediterranean seas could impede <i>R. ferrumequinum</i> gene flow. These results provide important information to improve the delineation of <i>R. ferrumequinum</i> management units. We suggest that a large management unit corresponding to the population ranging from Spanish Basque country to northern France must be considered. Particular attention should be given to mating territories as they seem to play a key role in maintaining the high levels of genetic mixing between colonies. Smaller management units corresponding to English and northern France colonies must also be implemented. These insular or peripheral colonies could be at higher risk of extinction in a near future.
Predicting amphibian intraspecific diversity with machine learning: Challenges and prospects for integrating traits, geography, and genetic data
<p>The growing availability of genetic datasets, in combination with machine learning frameworks, offer great potential to answer long-standing questions in ecology and evolution. One such question has intrigued population geneticists, biogeographers, and conservation biologists: What factors determine intraspecific genetic diversity? This question is challenging to answer because many factors may influence genetic variation, including life history traits, historical influences, and geography, and the relative importance of these factors varies across taxonomic and geographic scales. Furthermore, interpreting the influence of numerous, potentially correlated variables is difficult with traditional statistical approaches. To address these challenges, we analyzed repurposed data using machine learning and investigated predictors of genetic diversity, focusing on Nearctic amphibians as a case study. We aggregated species traits, range characteristics, and >42,000 genetic sequences for 299 species using open-access scripts and various databases. After identifying important predictors of nucleotide diversity with random forest regression, we conducted follow-up analyses to examine the roles of phylogenetic history, geography, and demographic processes on intraspecific diversity. Although life history traits were not important predictors for this dataset, we found significant phylogenetic signal in genetic diversity within amphibians. We also found that salamander species at northern latitudes contain lower genetic diversity. Data repurposing and machine learning provide valuable tools for detecting patterns with relevance for conservation, but concerted efforts are needed to compile meaningful datasets with greater utility for understanding global biodiversity.</p>
Weak genetic signal for phenotypic integration implicates developmental processes as major regulators of trait covariation
<p>Phenotypic integration is an important metric that describes the degree of covariation among traits in a population, and is hypothesized to arise due to selection for shared functional processes. Our ability to identify the genetic and/or developmental underpinnings of integration is marred by temporally overlapping cell-, tissue-, and structure-level processes that serve to continually 'overwrite' the structure of covariation among traits through ontogeny. Here we examine whether traits that are integrated at the phenotypic level, also exhibit a shared genetic basis (e.g., pleiotropy). We micro-CT scanned two hard tissue traits, and two soft tissue traits (mandible, pectoral girdle, atrium, and ventricle respectively) from an F<sub>5</sub> hybrid population of Lake Malawi cichlids, and used geometric morphometrics to extract 3D shape information from each trait. Given the large degree of asymmetric variation that may reflect developmental instability, we separated symmetric- from asymmetric-components of shape variation. We then performed quantitative trait loci (QTL) analysis to determine the degree of genetic overlap between shapes. While we found ubiquitous associations among traits at the phenotypic level, except for a handful of notable exceptions, our QTL analysis revealed few overlapping genetic regions. Taken together, this indicates developmental interactions can play a large role in determining the degree of phenotypic integration among traits, and likely obfuscate the genotype to phenotype map, limiting our ability to gain a comprehensive picture of the genetic contributors responsible for phenotypic divergence.</p>
Genetic admixture and evolutionary history of Han Chinese in the Shandong Peninsula inferred from integrative modern and ancient genomic resources
<p>The allele frequency data of 264 individuals from Shandong Province and supplementary table.</p>
Integrating top-down and bottom-up approaches to understand the genetic architecture of speciation across a monkeyflower hybrid zone
<p><span>Understanding the phenotypic and genetic architecture of reproductive isolation is a longstanding goal of speciation research. In several systems, large-effect loci contributing to barrier phenotypes have been characterized, but such causal connections are rarely known for more complex genetic architectures. In this study, we combine 'top-down' and 'bottom-up' approaches with demographic modeling toward an integrated understanding of speciation across a monkeyflower hybrid zone. Previous work suggests that pollinator visitation acts as a primary barrier to gene flow between two divergent red- and yellow-flowered ecotypes of <em>Mimulus</em> <em>aurantiacus</em>. Several candidate isolating traits and anonymous SNP loci under divergent selection have been identified, but their genomic positions remain unknown. Here, we report findings from demographic analyses that indicate this hybrid zone formed by secondary contact, but that subsequent gene flow was restricted by widespread barrier loci across the genome. Using a novel, geographic cline-based genome scan, we demonstrate that candidate barrier loci are broadly distributed across the genome, rather than mapping to one or a few 'islands of speciation.' Quantitative trait locus (QTL) mapping reveals that most floral traits are polygenic, with little evidence that QTL co-localize, indicating that most traits are genetically independent. Finally, we find little evidence that QTL and candidate barrier loci overlap, suggesting that some loci contribute to other forms of reproductive isolation. Our findings highlight the challenges of understanding the genetic architecture of reproductive isolation and reveal that barriers to gene flow aside from pollinator isolation may play an important role in this system.</span></p>
Literature compilation for: The rise of animal biotelemetry and genetics research data integration
<p><span>The advancement and availability of innovative animal biotelemetry and genomic technologies are improving our understanding of how the movements of individuals influence gene flow within and between populations and ultimately drive evolutionary and ecological processes. There is a growing body of work that is integrating what were once disparate fields of biology, and here we reviewed the published literature up until January 2023 (139 papers) to better understand the drivers of this research and how it is improving our knowledge of animal biology. The review showed that the predominant drivers for this research were: i) understanding how individual-based movements affect animal populations, ii) analyzing the relationship between genetic relatedness and social structuring, and iii) studying how the landscape affects the flow of genes, and how this is impacted by environmental change. However, there was a divergence between taxa as to the most prevalent research aim, and the methodologies applied. We also found that after 2010 there was an increase in studies that integrated the two data types using innovative statistical techniques instead of analyzing the data independently using traditional statistics from the respective fields. This new approach greatly improved our understanding of the link between </span><span>the individual, the population, and the environment and is being used to better conserve and manage species. We discuss the challenges and limitations, as well as the potential for growth and diversification of this research approach. The paper provides a guide for researchers who wish to consider applying these disparate disciplines and advancing the field.</span></p>
Integrative multi-ancestry genetic analysis of gene regulation in coronary arteries prioritizes disease risk loci
<p>All full-sample files contain results generated in coronary artery tissue from 138 American adults. Subset analyses utilized 80 individuals selected from the original 138. Scripts accompanying some of these data in downstream analyses can be viewed on our Github, which also contains a link to the current version of our accompanying manuscript: https://github.com/MillerLab-CPHG/CAD_QTL</p> <p>Full summary statistics for eQTL associations using mixQTL (https://github.com/hakyimlab/mixqtl/wiki) by chromosome are located in UVA_coronary_mixQTL_sumstats_by_chromosome.zip</p> <p>Full summary statistics for eQTL associations using mixQTL in the subset of 100% European-ancestry study sample members by chromosome are located in Hodonsky_mixQTL_Euro_sumstats.zip</p> <p>Full summary statistics for eQTL associations using mixQTL in the genetically diverse downsampled subset by chromosome are located in Hodonsky_mixQTL_downsample_sumstats.zip</p> <p>Full summary statistics for nominal pass for all genes identified as significant in the permutation pass using QTLtools (https://qtltools.github.io/qtltools/) adjusting for local ancestry by gene by chromosome are located in Local_ancestry_UVA_coronary_QTLtools_nominal_sumstats.zip</p> <p>Full summary statistics for sQTL associations with splice junctions using QTLtools by gene are located in sQTL_results_UVA_coronary_full_sumstats.zip</p>
Lynch Syndrome Integrative Epidemiology and Genetics
ClinicalTrials.gov study NCT06582914. IPD Sharing: NO. Countries: 1. Publications: 4.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.