Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,445

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,445 results for “Genetics: population”

Learn how ShareScore rates datasets ↗
edi44/100

State Water Project, Genetic Determination of Population of Origin 2011-2021

Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring

openCC (other)Dec 2021View details →
zenodo40/100

Geographical gradients of genetic diversity and differentiation among the southernmost marginal populations of Abies sachalinensis revealed by EST-SSR polymorphism

Research Highlights: We detected the longitudinal gradients of genetic diversity parameters, such as the number of alleles, effective number of alleles, heterozygosity, and inbreeding coefficient, and found that these might be attributable to climatic conditions, such as temperature and snow depth. Background and Objectives: Genetic diversity among local populations of a plant species at its distributional margin has long been of interest in ecological genetics. Populations at the distribution center grow well in favorable conditions, but those at the range margins are exposed to unfavorable environments, and the environmental conditions at establishment sites might reflect the genetic diversity of local populations. This is known as the central-marginal hypothesis in which marginal populations show lower genetic variation and higher differentiation than do central populations. In addition, genetic variation in a local population is influenced by phylogenetic constraints and the population history of selection under environmental constraints. In this study, we investigated this hypothesis in relation to Abies sachalinensis, a major conifer species in Hokkaido. Materials and methods: A total of 1,189 trees from 25 natural populations were analyzed using 19 EST-SSR loci. Results: The eastern populations; namely, those in the species distribution center, showed greater genetic diversity than did the western peripheral populations. Another important finding is that the southwestern marginal populations were highly differentiated from the other populations. Conclusions: These differences might be due to genetic drift in the small and isolated populations at the range margin. Therefore, our results indicated that the central-marginal hypothesis held true for the southernmost A. sachalinensis populations in Hokkaido.

opencc-zeroJan 2020View details →
zenodo40/100

Fig. 3 in Genetic diversity and population structure of Brycon nattereri (Characiformes: Bryconidae): a Neotropical fish under threat of extinction

Fig. 3. Haplotype network based on partial sequencing of the D-loop region (mtDNA) of 92 individuals of Brycon nattereri from the Laranjinha River. Circle sizes are pro- portional to haplotype frequency.

opencc-by-4.0Apr 2019View details →
zenodo40/100

Genetic variation in early fitness traits across European populations of silver birch (Betula pendula)

<p>Early life phenotypic data from three <em>Betula pendula</em> common garden experiments spread across the species latitudinal range in Europe.</p>

opencc-by-4.0May 2020View details →
dryad40/100

Data from: Genetic and environmental canalization are not correlated among altitudinally varying populations of Drosophila melanogaster

<p>Organisms are exposed to environmental and mutational effects influencing both mean and variance of phenotypes.  Potentially deleterious effects arising from this variation can be reduced by the evolution of buffering (canalizing) mechanisms, ultimately reducing phenotypic variability. There has been interest regarding the conditions enabling the evolution of canalization. Under some models, the circumstances under which genetic canalization evolves is limited, despite apparent empirical evidence for it. It has been argued that genetic canalization evolves as a correlated response to environmental canalization (congruence model). Yet, empirical evidence has not consistently supported predictions of a correlation between genetic and environmental canalization. In a recent study, a population of <em>Drosophila </em>adapted to high altitude showed evidence of genetic decanalization relative to those from low altitudes. Using strains derived from these populations, we tested if they varied for multiple aspects of environmental canalization We observed the expected differences in wing size, shape, cell (trichome) density and mutational defects between high- and low-altitude populations. However, we observed little evidence for a relationship between measures of environmental canalization with population or with defect frequency. Our results do not support the predicted association between genetic and environmental canalization.</p>

opencc-zeroJul 2020View details →
dryad40/100

Geographic patterns in morphometric and genetic variation for coyote populations with emphasis on southeastern coyotes

Prior to 1900, coyotes (Canis latrans) were restricted to the western and central regions of North America, but by the early 2000s coyotes became ubiquitous throughout the eastern United States. Information regarding morphological and genetic structure of coyote populations in the southeastern United States is limited, and where data exist, they are rarely compared to those from other regions of North America. We assessed geographic patterns in morphology and genetics of coyotes with special consideration of coyotes in the southeastern United States. Mean body mass of coyote populations increased along a west-to-east gradient, with southeastern coyotes being intermediate to western and northeastern coyotes. Similarly, principal component analysis of body mass and linear body measurements suggested that southeastern coyotes were intermediate to western and northeastern coyotes in body size but exhibited shorter tails and ears from other populations. Genetic analyses indicated that southeastern coyotes represented a distinct genetic cluster that differentiated strongly from western and northeastern coyotes. We postulate that southeastern coyotes experienced lower immigration from western populations than did northeastern coyotes, and over time, genetically diverged from both western and northeastern populations. Coyotes colonizing eastern North America experienced different selective pressures than did stable populations in the core range and we offer that the larger body size of eastern coyotes reflect an adaptation that improved dispersal capabilities of individuals in the expanding range.

opencc-zeroDec 2018View details →
dryad40/100

Data from: Genome wide assessment of genetic variation and population distinctiveness of the pig family in South Africa

<p>Genetic diversity is of great importance and a prerequisite for genetic improvement and conservation programs in pigs and other livestock populations. The present study provides a genome wide analysis of the genetic variability and population structure of pig populations from different production systems in South Africa relative to global populations. A total of 234 pigs sampled in South Africa and consisting of village (n = 91), commercial (n = 60), indigenous (n = 40), Asian (n = 5) and wild (n = 38) populations were genotyped using Porcine SNP60K BeadChip. In addition, 389 genotypes representing village and commercial pigs from America, Europe and Asia were accessed from a previous study and used to compare population clustering and relationships of South African pigs with global populations. Moderate heterozygosity levels, ranging from 0.204 for Warthogs to 0.371 for village pigs sampled from Capricorn municipality in Eastern Cape province of South Africa were observed. Principal Component Analysis of the South African pigs resulted in four distinct clusters of (i) Duroc; (ii) Vietnamese; (iii) Bush pig and Warthog and (iv) a cluster with the rest of the commercial (SA Large White and Landrace), village, Wild Boar and indigenous breeds of Koelbroek and Windsnyer. The clustering demonstrated alignment with genetic similarities, geographic location and production systems.  The PCA with the global populations also resulted in four clusters that where populated with (i) all the village populations, wild boars, SA indigenous and the large white and landraces; (ii) Durocs (iii) Chinese and Vietnamese pigs and (iv) Warthog and Bush pig. <i>K</i>= 10 (The number of population units) was the most probable ADMIXTURE based clustering, which grouped animals according to their populations with the exception of the village pigs that showed presence of admixture. AMOVA reported 19.92% – 98.62% of the genetic variation to be within populations. Sub structuring was observed between South African commercial populations as well as between Indigenous and commercial breeds. Population pairwise <i>F<sub>ST</sub></i>analysis showed genetic differentiation <i>(P &lt; 0.05)</i>between the village, commercial and wild populations. A per marker per population pairwise <i>F<sub>ST</sub></i>analysis revealed SNPs associated with QTLs for traits such as meat quality, cytoskeletal and muscle development, glucose metabolism processes and growth factors between both domestic populations as well as between wild and domestic breeds. Overall, the study provided a baseline understanding of porcine diversity and an important foundation for porcine genomics of South African populations.</p>

opencc-zeroJun 2020View details →
zenodo40/100

Figure 1 in Non-invasive genetic study and population monitoring of the brown bear (Ursus arctos) (Mammalia: Ursidae) in Kastoria region - Greece

Figure 1. The study area in Kastoria region and capture locations (red dots) for the 75 living bears.

opencc-by-4.0Jan 2014View details →
zenodo40/100

New genetic markers for Sapotaceae phylogenomics: more than 600 nuclear genes applicable from family to population levels

<p>Some tropical plant families, such as the Sapotaceae, have a complex taxonomy, which can be resolved using Next Generation Sequencing (NGS). For most groups however, methodological protocols are still missing. Here we identified 531 monocopy genes and 227 Short tandem repeats (STR) markers and tested them on Sapotaceae using target capture and NGS.&nbsp;The probes were designed using two genome skimming samples from<em>Capurodendron delphinense</em>&nbsp;and&nbsp;<em>Bemangidia lowryi</em>, both from the Tseboneae tribe, as well as the published&nbsp;<em>Manilkara zapota</em>&nbsp;transcriptome from the Sapotoideae tribe.&nbsp;We combined our probes with 261 additional ones previously published and designed for the entire angiosperm group. On a total of 792 low-copy genes, 638 showed no signs of paralogy and were used to build a phylogeny of the family with 231 individuals from all main lineages. A highly supported topology was obtained at high taxonomic ranks but also at the species level. This phylogeny revealed the existence of more than 20 putative new species. Single nucleotide polymorphisms (SNPs) extracted from the 638 genes were able to distinguish lineages within a species complex and to highlight geographical structuration. STR were recovered efficiently for the species used as reference (<em>C. delphinense</em>) but the recovery rate decreased dramatically with the phylogenetic distance to the focal species. All together, the new loci will help reaching a sound taxonomic understanding of the family Sapotaceae for which many circumscriptions and relationships are still debated, at the species, genus and tribe levels.</p>

opencc-by-4.0Jan 2021View details →
dryad40/100

Data from: Drift happens: molecular genetic diversity and differentiation among populations of jewelweed (Impatiens capensis Meerb.) reflect fragmentation of floodplain forests

Landscape features often shape patterns of gene flow and genetic differentiation in plant species. Populations that are small and isolated enough also become subject to genetic drift. We examined patterns of gene flow and differentiation among 12 floodplain populations of the selfing annual jewelweed (Impatiens capensis Meerb.) nested within four river systems and two major watersheds in Wisconsin, USA. Floodplain forests and marshes provide a model system for assessing the effects of habitat fragmentation within agricultural/urban landscapes and for testing whether rivers act to genetically connect dispersed populations. We generated a panel of 12,856 single nucleotide polymorphisms and assessed genetic diversity, differentiation, gene flow, and drift. Clustering methods revealed strong population genetic structure with limited admixture and highly differentiated populations (mean multilocus FST = 0.32, FST' = 0.33). No signals of isolation by geographic distance or environment emerged, but alleles may flow along rivers given that genetic differentiation increased with river distance. Differentiation also increased in populations with fewer private alleles (R2 = 0.51) and higher local inbreeding (R2 = 0.22). Populations varied greatly in levels of local inbreeding (FIS = 0.2 to 0.9) and FIS declined in smaller, more isolated populations. These results suggest that genetic drift dominates other forces in structuring these Impatiens populations. In rapidly changing environments, species must migrate or genetically adapt. Habitat fragmentation limits both processes, potentially compromising the ability of species to persist in fragmented landscapes.

opencc-zeroDec 2018View details →
dryad40/100

Data from: Different genetic structures revealed resident populations of a specialist parasitoid wasp in contrast to its migratory host

Genetic comparisons of parasitoids and their hosts are expected to reflect ecological and evolutionary processes that influence the interactions between species. The parasitoid wasp, Cotesia vestalis, and its host diamondback moth (DBM), Plutella xylostella, provide opportunities to test whether the specialist natural enemy migrates seasonally with its host or occurs as resident population. We genotyped 17 microsatellite loci and two mitochondrial genes for 158 female adults of C. vestalis collected from 12 geographical populations, as well as nine microsatellite loci for 127 DBM larvae from six separate sites. The samplings covered both the likely source (southern) and immigrant (northern) areas of DBM from China. Populations of C. vestalis fell into three groups, pointing to isolation in northwestern and southwestern China and strong genetic differentiation of these populations from others in central and eastern China. In contrast, DBM showed much weaker genetic differentiation and high rates of gene flow. TESS analysis identified the immigrant populations of DBM as showing admixture in northern China. Genetic disconnect between C. vestalis and its host suggests that the parasitoid did not migrate yearly with its host but likely consisted of resident populations in places where its host could not survive in winter.

opencc-zeroDec 2016View details →
dryad40/100

Data from: Sporadic genetic connectivity among small insular populations of the rare geoendemic plant Caulanthus amplexicaulis var. barbarae (Santa Barbara Jewelflower)

Globally, a small number of plants have adapted to terrestrial outcroppings of serpentine geology, which are characterized by soils with low levels of essential mineral nutrients (N, P, K, Ca, Mo) and toxic levels of heavy metals (Ni, Cr, Co). Paradoxically, many of these plants are restricted to this harsh environment. Caulanthus ampexlicaulis var. barbarae (Brassicaceae) is a rare annual plant that is strictly endemic to a small set of isolated serpentine outcrops in the coastal mountains of central California. The goals of the work presented here were to 1) determine the patterns of genetic connectivity among all known populations of Caulanthus ampexlicaulis var. barbarae, and 2) estimate contemporary effective population sizes (Ne), in order to inform ongoing genomic analyses of the evolutionary history of this taxon, and to provide a foundation upon which to model its future evolutionary potential and long-term viability in a changing environment. Eleven populations of this taxon were sampled, and population-genetic parameters were estimated using 11 nuclear microsatellite markers. Contemporary effective population sizes were estimated using multiple methods and found to be strikingly small (typically Ne &lt; 10). Further, our data showed that a substantial component of genetic connectivity of this taxon is not at equilibrium, and instead showed sporadic gene flow. Several lines of evidence indicate that gene flow between isolated populations is maintained through long-distance seed dispersal (e.g. &gt; 1 km), possibly via zoochory.

opencc-zeroDec 2019View details →
dryad40/100

Varying genetic imprints of roads and human density in North American mammal populations

<p>Road networks and human density are major factors contributing to habitat fragmentation and loss, isolation of wildlife populations and reduced genetic diversity. Terrestrial mammals are particularly sensitive to road networks and encroachment by human populations. However, there are limited assessments of the impacts of road networks and human density on population-specific nuclear genetic diversity, and it remains unclear how these impacts are modulated by life history traits. Using generalized linear mixed models and microsatellite data from 1444 North American terrestrial mammal populations we show that taxa with large home range sizes, dense populations, and large body sizes had reduced nuclear genetic diversity with increasing road impacts and human density, but the overall influence of life history traits was generally weak. Instead, we observed a high degree of genus-specific variation in genetic responses to road impacts and human density. Human density negatively affected allelic diversity or heterozygosity more than road networks (13 versus 5-7 of 25 assessed genera, respectively); increased road networks and human density also positively affected allelic diversity and heterozygosity in 15 and 6-9 genera, respectively. Large bodied, human-averse species were generally more negatively impacted than small, urban-adapted species. Genus-specific responses to habitat fragmentation by ongoing road development and human encroachment likely depend on the specific capability to (i) navigate roads as either barriers or movement corridors, and (ii) exploit resource-rich urban environments. The non-uniform genetic response to roads and human density highlights the need to implement efforts to mitigate the risk of vehicular collisions, while also facilitating gene flow between populations of particularly vulnerable taxa.</p>

opencc-zeroJun 2021View details →
zenodo40/100

Convergent geographic patterns between grizzly bear population genetic structure and Indigenous language groups in coastal British Columbia

<p>Microsatellite loci calls, sex, and mean centre detection per individual (GrizzlyMicroLociMeanXY.csv)&nbsp;and code associated with the paper: &quot;Convergent geographic patterns between grizzly bear population genetic structure and Indigenous language groups in coastal British Columbia&quot;. All code is from published R packages or GitHub repositories not created by the author. Code used is best described in these alternate resources.&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Raw Genotyping data from: Variation in recombination rate and its genetic determinism in sheep populations from combining multiple genomewide datasets

<p>Data supporting :</p> <p><strong>Variation in recombination rate and its genetic determinism in sheep populations from combining multiple genomewide datasets</strong></p> <p>Morgane Petit, Jean-Michel Astruc, Julien Sarry, Laurence Drouilhet, Stephane Fabre, Carole Moreno, Bertrand Servin</p> <p>http://doi.org/10.1534/genetics.117.300123</p> <p><strong>Abstract</strong></p> <p>Recombination is a complex biological process that results from a cascade of multiple events during meiosis. Understanding the genetic determinism of recombination can help to understand if and how these events are interacting. To tackle this question, we studied the patterns of recombination in sheep, using multiple approaches and datasets. We constructed male recombination maps in a dairy breed from the south of France (the Lacaune breed) at a fine scale by combining meiotic recombination rates from a large pedigree genotyped with a 50K SNP array and historical recombination rates from a sample of unrelated individuals genotyped with a 600K SNP array. This analysis revealed recombination patterns in sheep similar to other mammals but also genome regions that have likely been affected by directional and diversifying selection. We estimated the average recombination rate of Lacaune sheep at 1.5 cM/Mb, identified about 50,000 crossover hotspots on the genome and found a high correlation between historical and meiotic recombination rate estimates. A genome-wide association study revealed two major loci affecting inter-individual variation in recombination rate in Lacaune, including the <em>RNF212</em> and<em> HEI10</em> genes and possibly 2 other loci of smaller effects including &nbsp;the <em>KCNJ15</em> &nbsp;and <em>FSHR</em> genes. Finally, we compared our results to those obtained previously in a distantly related population of domestic sheep, the Soay. This comparison revealed that Soay and Lacaune males have a very similar distribution of recombination along the genome and that the two datasets can be combined to create more precise male meiotic recombination maps in sheep. Despite their similar recombination maps, we show that Soay and Lacaune males exhibit different heritabilities and QTL effects for inter-individual variation in genome-wide recombination rates.</p> <p>&nbsp;</p> <p>Data files are provided in Plink format ( https://www.cog-genomics.org/plink2 ).</p> <p>&nbsp;</p>

opencc-by-nc-4.0Feb 2017View details →
zenodo40/100

Mountain landscape connectivity and subspecies appurtenance shape genetic differentiation in natural plant populations of the snapdragon (Antirrhinum majus L.)

<p>This dataset provides the raw data for the population genetic analyses for the article: "Mountain landscape connectivity and subspecies appurtenance shape genetic differentiation in natural plant populations of the snapdragon (Antirrhinum majus L.)" by Benoit Pujol; Juliette Archambeau; Aurore Bontemps; Mylène Lascoste; Sara Marin; and Alexandre Meunier found in the journal "Botany Letters", Vol 164 pp. 111-119 (DOI: 10.1080/23818107.2017.1310056).</p> <p>Link to journal open access article: http://www.tandfonline.com/doi/pdf/10.1080/23818107.2017.1310056</p> <p>Link to Zenodo article reporsitory: https://zenodo.org/record/801169</p> <p>The datafile includes three data sheets:</p> <p>Data, which contains for each plant : the name of the population, the name of the sampled individual, the subspecies, the latitude of the population, the longitude of the population, the altitudinal elevation of the population in meters, and the microsatellite genotype of each plant. Genotype data is recorded by locus (two columns for the two alleles at one locus). Locus name is found as the title of the column. The record for each allele is its allele size.</p> <p>valleys 1 and valleys 2, which contains the association between populations and valleys following the two scenarios that we analyzed in the paper.</p> <p>Microsatelite loci were developed during previous work: see the following paper for more details: Debout, G., E. Lhuillier, P.-J. Malé, B. Pujol, and C. Thébaud. 2012. Development and characterization of 24 polymorphic microsatellite loci in two Antirrhinum majus subspecies (Plantaginaceae) using pyrosequencing technology. Conservation Genetics Resources 4:75-79.</p>

opencc-by-4.0Apr 2017View details →
zenodo40/100

Population-scale skeletal muscle single-nucleus multi-omic profiling reveals extensive context specific genetic regulation

<p>Data accompanying the manuscript "Population-scale skeletal muscle single-nucleus multi-omic profiling reveals extensive context specific genetic regulation".</p> <p>Note: For ATAC fragment files, e,caQTL full cis scan summary files, clustering objects, please see the CMDGA portal (https://cmdga.org/search/?searchTerm=stephen-parker%3AVarshney2024)<br>For raw data including fastq files, please see dbGaP repo phs001048.v3.p1</p> <p>Data in this repository includes:</p> <p>Filename: Description</p> <p>1. list of 8,666 genes for which exon-only counts were considered. See methods section "Adjusting RNA counts for overlapping gene annotations" in the manuscript.</p> <p>2. nucleus_sample_cluster_map.tsv: nucleus-sample-cluster map with other QC info.&nbsp;<br># index: nucleus identified syntax &lt;modality&gt;.&lt;batch&gt;.NM.&lt;10X channel&gt;.&lt;barcode&gt;&nbsp;<br># UMAP_1, UMAP_2: UMAP coordinates for visualization<br># modality: rna or atac<br># batch: processing batch identifier<br># hqaa_umi: high quality autosomal alignments (HQAA) for atac nuclei, unique molecular identifier (UMI) for tna&nbsp;<br># fraction_mitochondrial: fraction of reads mapping to the mitochondrial genome<br># cohort: sample cohort<br># tss_enrichment: TSS enrichment for atac nuclei<br># coarse_cluster_name: cluster name</p> <p>3. peaks.tar.gz: snATAC peak features including:<br># consensus-summits.bed: consensus summits along with the cell type that the summits was highest in.<br># narrow peaks in clusters<br># consensus summit feature (summit +- 150bp) identified in each cluster - these were used in GWAS enrichments.</p> <p>4. snrna-cell-type-specific-genes.tsv: Normalized expression scores for genes in each cell-type cluster</p> <p>5. eqtl_permute.tar.gz: Permutation scan eQTL in each cell-type cluster. Columns:&nbsp;<br># variant: syntax &lt;chrom&gt;:&lt;hg38 pos&gt;:&lt;ref&gt;:&lt;alt&gt;<br># effect_allele: effect allele (was the alt allele)<br># other_allele: non-effect allele<br># feature: gene name<br># featureCoordinates_tss: gene TSS<br># p-value: nominal p value<br># beta: slope/beta of the linear regression. Keyed on the alt allele<br># se: standard error of the slope<br># snp: SNP ID<br># strand: gene strand<br># n_variants_tested: number of variants tested for the gene<br># distance_var_pheno: distance of the variant with the gene TSS<br># n_effective_tests: number of effective tests<br># p_beta: beta distribution adjusted p value<br># qvalue: qvalue (Storey)</p> <p>6. caqtl_permute.tar.gz: # Permutation scan caQTL in each cell-type cluster. Columns:&nbsp;<br># variant: syntax &lt;chrom&gt;:&lt;hg38 pos&gt;:&lt;ref&gt;:&lt;alt&gt;<br># effect_allele: effect allele (was the alt allele)<br># other_allele: non-effect allele<br># feature: peak feature coordinates<br># p-value: nominal p value<br># beta: slope/beta of the linear regression. Keyed on the alt allele<br># se: standard error of the slope<br># snp: SNP ID<br># n_variants_tested: number of variants tested for the gene<br># distance_var_pheno: distance of the variant with the gene TSS<br># n_effective_tests: number of effective tests<br># p_beta: beta distribution adjusted p value<br># qvalue: qvalue (Storey)</p> <p>7. eqtl_credible_sets.tar.gz: # eQTL credible set. The file name denotes the egene and the signal hit id. Bed file columns:&nbsp;<br># 1: snp chromosome<br># 2: snp start<br># 3: snp end<br># 4: snp chrom_pos_ref_alt<br># 5: Bayes Factor&nbsp;<br># 6: PIP<br># 7: SNP rsid</p> <p>8. caqtl_credible_sets.tar.gz: # caqtl credible set. The file name denotes the capeak and the signal hit id. Bed file columns:&nbsp;<br># 1: snp chromosome<br># 2: snp start<br># 3: snp end<br># 4: snp chrom_pos_ref_alt<br># 5: Bayes Factor&nbsp;<br># 6: PIP<br># 7: SNP rsid</p> <p>9. cicero_all.tar.gz # Cicero coaccessibility results. Columns<br># Peak 1: Macs2 narrowpeak coordinate for peak 1<br># Peak 2: Macs2 narrowpeak coordinate for peak 2<br># coaccess: Cicero coaccessibility score</p> <p>10. cicero_gene_tss.tar.gz: Cicero coaccessibility results between peak and genes. Macs2 narrow peaks in the TSS+1kb upstream region are assigned that gene name. Columns<br># Cicero coaccessibility results between peak and genes. Macs2 narrow peaks in the TSS+1kb upstream region are assigned that gene name.Columns<br># Peak 1: Macs2 narrowpeak coordinate for peak 1<br># gene_name: Assigned gene<br># Peak 2: Macs2 narrowpeak coordinate for peak 2<br># coaccess: Cicero coaccessibility score<br>## &nbsp;Peak1 is the narrowpeak in the TSS region, peak2 is the distal peak</p> <p>11. mash.tar.gz Mashr results for e/caQTL - lfsr, posterior means and posterior SD for each tested eSNP-eGene, caSNP-caPeak pair.&nbsp;</p> <p>12. cellregmap.tar.gz: Cellregmap results for endothelial nucleus-level eQTL scans.<br>## Persistent genetic effect beta_g was calculated in a simple association model.&nbsp;<br>## An interaction model was fit to test for GxC effect. columns:<br># rho1, g2, e1, and eps2 are variance component measures outputs from CellRegMap corresponding to interaction, genetic, environment and residual variance components.&nbsp;<br># p_nominal: nominal p from cellRegMap<br># kind: model kind in CellRegMap - simple association or interaction<br># beta_g: &nbsp;Persistent genetic effect<br># gene_name: gene name for eQTL or peak feature name for caQTL<br># context: context used either factors (continuous) or subclusters (discrete)<br># snp: index snp for which model is fit. This is the most significant identified snp from our standard e,caQTL scans. chrom-hg38pos-rsid</p> <p><br>13. coloc-eqtl-caqtl.tsv: # Summary of eQTL-caQTL coloc in each cluster. Columns:<br># nsnps: Number of SNPs in the region<br># eqtl_hit: SNP with the highest Bayes factor in the SuSiE eQTL credible set<br># caqtl_hit: SNP with the highest Bayes factor in the SuSiE caQTL credible set<br># PP.H0.abf: Coloc posterior probability for no signal<br># PP.H1.abf: Coloc posterior probability for signal in dataset 1<br># PP.H2.abf: Coloc posterior probability for signal in dataset 2<br># PP.H3.abf: Coloc posterior probability for different signals in datasets 1 and 2<br># PP.H4.abf: Coloc posterior probability for shared signal in datasets 1 and 2<br># idx1: Index of the SuSiE credible set for dataset 1<br># idx2: Index of the SuSiE credible set for dataset 2<br># cluster: cluster name<br># egene: eGene name<br># capeak: caPeak coordinates</p> <p>14. cit-mrs-summary.tsv: &nbsp;Summary from CIT and MR Steiger directionality tests. Columns:<br># cluster: cluster name<br># egene: eGene name<br># capeak: caPeak coordinates<br># eqhit: SNP with the highest Bayes factor in the SuSiE eQTL credible set<br># cahit: SNP with the highest Bayes factor in the SuSiE caQTL credible set<br># p.cit_c_c-e: P value for CIT causal cahit-ca-to-e model<br># q.cit_c_c-e: q value for CIT causal cahit-ca-to-e model<br># p.cit_rc_c-e: P value for CIT reverse-causal eqhit-ca-to-e model&nbsp;<br># q.cit_rc_c-e: value for CIT reverse-causal eqhit-ca-to-e model&nbsp;<br># p.cit_c_e-c: P value for CIT causal eqhit-e-to-ca model<br># q.cit_c_e-c: q value for CIT causal eqhit-e-to-ca model<br># p.cit_rc_e-c: P value for CIT reverse-causal cahit-e-to-ca model&nbsp;<br># q.cit_rc_e-c: q value for CIT reverse-causal cahit-e-to-ca model&nbsp;<br># cit_direction: Direction inferred from CIT &nbsp;<br># correct_causal_direction--ca-to-e: MR Steiger directionality test - is ca-to-e direction correct?<br># correct_causal_direction--e-to-ca: MR Steiger directionality test - is e-to-ca direction correct?<br># sensitivity_ratio--ca-to-e: MR Steiger Sensitivity ratio for ca-to-e model&nbsp;<br># sensitivity_ratio--e-to-ca: &nbsp;MR Steiger Sensitivity ratio for e-to-ca model<br># steiger_test--ca-to-e: MR Steiger directionality test P value for ca-to-e model<br># steiger_test--e-to-ca: MR Steiger directionality test P value for e-to-ca model<br># steiger_q--ca-to-e: MR Steiger directionality test q value for ca-to-e model<br># steiger_q--e-to-ca: MR Steiger directionality test q value for e-to-ca model<br># mrs_direction: Direction inferred from MR Steiger<br># direction: Direction inferred requiring consistent results between CIT and MR Steiger directionality test</p> <p>15. coloc-gwas-eqtl.tsv and<br>16. coloc-gwas-caqtl.tsv # Summary of e/caQTL coloc with GWAS in each cluster. Columns:<br># nsnps: Number of SNPs in the region<br># gwas_hit: SNP with the highest bayes factor in the SuSiE GWAS credible set<br># eqtl_hit: SNP with the highest bayes factor in the SuSiE eQTL credible set<br># caqtl_hit: SNP with the highest bayes factor in the SuSiE caQTL credible set<br># PP.H0.abf: Coloc posterior probability for no signal<br># PP.H1.abf: Coloc posterior probability for signal in dataset 1<br># PP.H2.abf: Coloc posterior probability for signal in dataset 2<br># PP.H3.abf: Coloc posterior probability for different signal in datasets 1 and 2<br># PP.H4.abf: Coloc posterior probability for shared signal in datasets 1 and 2<br># idx1: Index of the SuSiE credible set for dataset 1<br># idx2: Index of the SuSiE credible set for dataset 2<br># cluster: cluster name<br># egene: eGene name<br># capeak: caPeak coordinates<br># p12min: Min prior p12 where the PP H4 &gt; 0.5. Lower this value, more robust is the colocalization<br># trait: GWAS trait name<br># gwas_locus: GWAS locus name for the coloc test - a 250kb left and right flanking genomic window on this SNP was considered for testing coloc between all pairs of GWAS/QTL signals identified in this region &nbsp;<br># traitname: Expanded GWAS trait name<br># variable_type: GWAS type&nbsp;<br># source: Source of GWAS - either UKBB or other study</p> <p>17. supplementary_tables.xlsx: Supplementary tables from the manuscript.<br>Information included in sheets:<br>1. "marker_genes": Marker genes known from literature used to annotate clusters<br>2. "n_nuclei": n pass-QC nuclei per modality-sample-cluster</p> <p>2. "snrna_GO_enrichment": GO term enrichment: matrix of cluster vs top 2 GO terms</p> <p>3. "qtl_scan_info": &nbsp;e/caQTL scan info<br>cluster: cluster<br>ntested_eqtl: N genes tested for eQTL<br>nsig_eqtl: N significant (5% FDR) eGenes<br>n_pheno_pcs_eqtl: N phenotype PCs considered for eQTL<br>ratio_eqtl: Ratio of N eGenes/N genes tested<br>nsig_caqtl: &nbsp;N peaks tested for caQTL<br>ntested_caqtl: N significant (5% FDR) caPeaks<br>n_pheno_pcs_caqtl: N phenotype PCs considered for caQTL<br>ratio_caqtl: Ratio of N caPeaks/N peaks tested<br>nsamples_eqtl: N samples for eQTL<br>nsamples_caqtl: N samples for caQTL</p> <p>4. "gwas_trait_list": GWAS trait info<br>trait: GWAS trait ID<br>traitname: GWAS trait description<br>variable_type: GWAS type. case/control (cc), continuous_irnt=continuous inverse-normal transformed<br>source: GWAS source<br>doi: GWAS study DOI</p> <p>5. "traits_in_ldsc_baseline" - list of annotations included in the baseline model for LDSC</p> <p>6. "gwas_enrichment_in_peaks" GWAS enrichment in cluster peaks (S-LDSC)</p> <p>7. "gwas_enrichment_in_qtl_peaks" GWAS enrichment in QTL peaks (fGWAS) # fGWAS results comparing GWAS enrichment in type 1 annotations<br>CI_lower_ln, estimate_ln, CI_upper_ln: natural log of lower confidence interval, estimate, and upper confidence interval<br>trait: trait id<br>traitname: trait name<br>annotation: annotation<br>sig: 1 if CIs don't overlap 0, otherwise 0</p> <p>8. t2d_gwas_caqtl_coloc and<br>9. t2d_gwas_eqtl_coloc:<br>Summary of e,caQTL coloc with T2D GWAS in each cluster, along with target gene nominations. Columns:<br>nsnps: Number of SNPs in the region<br>gwas_hit: SNP with the highest bayes factor in the SuSiE GWAS credible set<br>eqtl_hit: SNP with the highest bayes factor in the SuSiE eQTL credible set<br>caqtl_hit: SNP with the highest bayes factor in the SuSiE caQTL credible set<br>PP.H0.abf: Coloc posterior probability for no signal<br>PP.H1.abf: Coloc posterior probability for signal in dataset 1<br>PP.H2.abf: Coloc posterior probability for signal in dataset 2<br>PP.H3.abf: Coloc posterior probability for different signal in datasets 1 and 2<br>PP.H4.abf: Coloc posterior probability for shared signal in datasets 1 and 2<br>idx1: Index of the SuSiE credible set for dataset 1<br>idx2: Index of the SuSiE credible set for dataset 2<br>cluster: cluster name<br>egene: eGene name<br>capeak: caPeak coordinates<br>p12min: Min prior p12 where the PP H4 &gt; 0.5. Lower this value, more robust is the colocalization<br>trait: GWAS trait id<br>diamante_gwas_locus: GWAS signal from the DIAMANTE 2018 study. Some signals that our SuSiE runs identified were not present in the original study in which case this column is NA<br>traitname: Expanded GWAS trait name<br>capeak_in_tss: caPeak in TSS + 1kb upstream region of a gene<br>gene_target_standard_cicero: caPeak coaccessible with TSS peak of a gene considering nuclei from all samples for co-accessibility<br>gene_target_allelic_cicero: &nbsp;caPeak coaccessible with TSS peak of a gene considering nuclei from samples homozygous for the caSNP allele associated with increased accessibility<br>gwashit_nominal_egene: gwas_hit nominally associated with these genes nominated in the columns capeak_in_tss, gene_target_standard_cicero, and &nbsp;gene_target_allelic_cicero</p> <p>10. MPRA results for the C2CD4A locus</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Georeferenced data for the study Environmental suitability throughout the late Quaternary explains population genetic diversity

<p>Data filtered from GBIF (datasetKey: 50c9509d-22c7-4a22-a47d-8c48425ef4a7) &nbsp;Contains 150 records of the <i>Sciurus aberti </i>squirrel filtered in latitudinal windows of 5 degrees from 20 to 45 degrees N. &nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Germline CpG methylation signatures in the human population inferred from genetic polymorphism

<p>This repository contains data released accompanying the manuscript "Germline CpG methylation signatures in the human population inferred from genetic polymorphism".&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Figure 3 in Morphogeometric and genetic variations among North African populations of the Mediterranean killifish Aphanius fasciatus (Valenciennes, 1821) from different habitats

Figure 3. – Projection of Procrustes coordinates and scores of canonical variate on Aphanius fasciatus shape analysis (A) and comparative transformation grids (B). LM: Mellah lagoon, M: Mellah marsh, BZ: Lagoon of Bizerte, LA: Ayata Lake.

opencc-by-4.0Dec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record