Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

899

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

899 results for “allele”

Learn how ShareScore rates datasets ↗
dryad36/100

Template-specific optimization of NGS genotyping pipelines reveals allele-specific variation in MHC gene expression

<p>Using high-throughput sequencing for precise genotyping of multi-locus gene families, such as the Major Histocompatibility Complex (MHC), remains challenging, due to the complexity of the data and difficulties in distinguishing genuine from erroneous variants. Several dedicated genotyping pipelines for data from high-throughput sequencing, such as next-generation sequencing (NGS), have been developed to tackle the ensuing risk of artificially inflated diversity. Here, we thoroughly assess three such multi-locus genotyping pipelines for NGS data, the DOC method, AmpliSAS and ACACIA, using MHC class IIβ datasets of three-spined stickleback gDNA, cDNA, and "artificial" plasmid samples with known allelic diversity. We show that genotyping of gDNA and plasmid samples at optimal pipeline parameters was highly accurate and reproducible across methods. However, for cDNA data, gDNA-optimal parameter configuration yielded decreased overall genotyping precision and consistency between pipelines. Further adjustments of key clustering parameters were required tο account for higher error rates and larger variation in sequencing depth per allele, highlighting the importance of template-specific pipeline optimization for reliable genotyping of multi-locus gene families. Through accurate paired gDNA-cDNA typing and MHC-II haplotype inference, we show that MHC-II allele-specific expression levels correlate negatively with allele number across haplotypes. Lastly, sibship-assisted cDNA-typing of MHC-I revealed novel variants linked in haplotype blocks and a higher-than-previously-reported individual MHC-I allelic diversity. In conclusion, we provide novel genotyping protocols for the three-spined stickleback MHC-I and -II genes and evaluate the performance of popular NGS-genotyping pipelines. We also show that fine-tuned genotyping of paired gDNA-cDNA samples facilitates amplification bias-corrected MHC allele expression analysis.</p>

opencc-zeroJan 2024View details →
dryad36/100

Autoimmunity-associated allele of tyrosine phosphatase gene PTPN22 enhances anti-viral immunity

<p>The 1858C&gt;T allele of the tyrosine phosphatase <em>PTPN22</em> is present in 5-10% of the North American population and is strongly associated with numerous autoimmune diseases. Although research has been done to define how this allele potentiates autoimmunity, the influence <em>PTPN22</em> and its pro-autoimmune allele have in anti-viral immunity remains poorly defined. Here, we use single-cell RNA-sequencing and functional studies to interrogate the impact of this pro-autoimmune allele on anti-viral immunity during Lymphocytic Choriomeningitis Virus clone 13 (LCMV-cl13) infection. Mice homozygous for this allele (PEP-619WW) clear the LCMV-cl13 virus whereas wildtype (PEP-WT) mice cannot. This is associated with enhanced anti-viral CD4 T cell responses and a more immunostimulatory CD8a<sup>-</sup> cDC phenotype. Adoptive transfer studies demonstrated that PEP-619WW enhanced anti-viral CD4 T cell function through virus-specific CD4 T cell-intrinsic and extrinsic mechanisms. Taken together, our data show that the pro-autoimmune allele of <em>Ptpn22</em> drives a beneficial anti-viral immune response thereby preventing what is normally a chronic virus infection.</p>

opencc-zeroApr 2024View details →
zenodo36/100

A phased genome of the highly heterozygous 'Texas' almond uncovers patterns of allele-specific expression linked to heterozygous structural variants

<h2># Genomic datasets associated to the publication:&nbsp;</h2> <h3># Gene-ID conversion with previous genome version</h3> <p>Texasv3_vs_Texasv2_GeneID.txt -- gene ID conversion between Texasv3 and Texasv2 (https://www.rosaceae.org/analysis/295)</p> <p>pdulcis26_to_F1_liftoff_polished.gff3 -- Texasv2 gene annotation liftoff on Phase-1 assembly (Phase-1 coordinates)</p> <h3># Phase-1</h3> <p>Texas_F1_K80_chr.fasta &nbsp;-- genome asssembly, phase-1&nbsp;<br>Texas_F1_gene_models.gff3 -- &nbsp;phase-1 &nbsp;gene annotation (de novo annotation)<br>Functional_annotation_TexasF1.csv -- &nbsp;phase-1 gene functions &nbsp;<br>Texas_F1_ref_SV.vcf --- Structural variations relative to phase-0 (This file uses Phase-1 as reference)</p> <h3># Phase-0</h3> <p>Texas_F0_K80_chr.fasta -- genome asssembly, phase-0&nbsp;<br>Texas_F0_gene_models.gff3 -- &nbsp;phase-0 &nbsp;gene annotation (liftoff from Phase-1)<br>Functional_annotation_TexasF0.csv -- &nbsp;phase-0 &nbsp;gene functions &nbsp;<br>Texas_F0_ref_SV.vcf --- Structural variations relative to phase-1 (This file uses Phase-0 as reference)</p> <h3># Transposable element annotation</h3> <p>Texas_F0_HiConf_TE_v3.gff3 --- TE annotation in Phase-0<br>Texas_F1_HiConf_TE_v3.gff3 --- TE annotation in Phase-1<br>Texasv3_TElib.fa --- TE library of TexasV3 (non-redundant repeat consensuses taking the account the two genome phases)</p> <h3># Gene sequences in fasta</h3> <p>Transcript, CDS and protein sequences in fasta format for phase-0 (F0) and phase-1 (F1)&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Ancestral allele estimates for cattle using est-sfs software with the K2 model

<h1>Overview</h1> <p>The assignment of bovine ancestral alleles was based on a model comparison of alleles from cattle with alleles from outgroup species: Water Buffalo, Sheep, and White-Tailed Deer.&nbsp;</p> <p>The frequency of cattle alleles are determined using 79 representative individuals from 1000 Bull Genomics Project. We utilized multiple sequence alignments of 110 species (78 ruminants and 32 mammalian outgroup species), available from http://animal.omics.pro/code/index.php/RGD/loadByGet?address[]=RGD/Download/comSynDownload.php, to determine the alleles in Water Buffalo, Sheep, and White-Tailed Deer at each locus.</p> <p>We employed the est-sfs software with the K2 model to infer the probability (Pancs) of the major allele in cattle being ancestral. Alleles were determined to be ancestral if they were the major allele at a site with Pancs &gt; 0.8 or the minor allele at a site with Pancs &lt; 0.2.</p> <p>Please email bft990914@163.com for any queries.</p> <p>The columns of this dataframe are</p> <p>chrome: chromosome index.</p> <p>pos: physical location of SNV.</p> <p>cattle_ref: reference allele of cattle.</p> <p>cattle_alt: alternative allele of cattle.</p> <p>cattle_maj: &nbsp;major allele of cattle.</p> <p>water_buffalo: the sequence of water_buffalo.</p> <p>sheep: the sequence of sheep.</p> <p>white_tailed_deer: the sequence of white_tailed_deer.</p> <p>p_maj_anc: the probability of the major allele of cattle being ancestral.</p> <p>ancestral_allele: the inferred ancestral allele.</p>

opencc-by-4.0Dec 2024View details →
dryad36/100

Figure e-1.- Effect of APOE alleles on cognitive domain-specific composite measures.

<p>Supplemental Figure of our original article "Association of <em>APOE</em> Genotype with Heterogeneity of Cognitive Decline Rate in Alzheimer's Disease" published in <em>Neurology</em> showing model 1-based trajectories of cognitive domain-specific composite measures by <em>APOE</em> genotype, as well as the difference of the <em>APOE</em>e2 and <em>APOE</em>e4 groups with respect to the <em>APOE</em>e3/e3 reference group.</p>

opencc-zeroFeb 2022View details →
zenodo36/100

Evolution of allele frequencies in the cattle breed Asturiana de los Valles

<p>Genotype data for 153 animals from the Asturiana de los Valles bovine breed, with birth dates from 1980 to 2013. These genotypes were obtained from 3 different sources:&nbsp;</p> <p>- 88 sires were genotyped using the Illumina&rsquo;s BovineSNP50 v.2 chip; the resulting data are provided at plink ped/map format (asturiana_50K.tar.xz).</p> <p>- 50 animals (25 sires and 25 dams) were genotyped using the Illumina&rsquo;s Bovine High Density BeadChip 770K SNP; the resulting data are provided at plink ped/map format (asturiana_800K.tar.xz).</p> <p>- 15 sires were sequenced on a HiSeq 3,000; the resulting genotype calls are provided at vcf format (asturiana_WGS_SNP.vcf.gz).</p> <p>Boitard et al (2021) combined these 3 datasets in order to detect recent and historical selection signatures in this breed. The scripts used for this analysis can be found at https://github.com/sboitard/Asturiana_analysis.</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Computational models from: Allele-specific activation, enzyme kinetics, and inhibitor sensitivities of EGFR exon 19 deletion mutations in lung cancer

<p>Computational models,&nbsp;compressed molecular dynamics (MD) simulation trajectories, and sample input files&nbsp;for &quot;Allele-specific activation, enzyme kinetics, and inhibitor sensitivities of EGFR exon 19 deletion mutations in lung cancer&quot;. An early version of this manuscript is available as a preprint here:&nbsp;https://www.biorxiv.org/content/10.1101/2022.03.16.484661v1</p>

opencc-by-4.0May 2022View details →
zenodo36/100

ARF / YA16Sdb collection of curated 16S rRNA alleles

<p>A collection of 16s rRNA alleles filtered by read quality and annotation quality.</p> <p>This repository contains 376,934 full-length 16s rRNA alleles with validated (by majority-rules) taxonomic annotations. These are a subset of 727,361 full-length 16s rRNA alleles.</p> <p>The pipeline used to create this repo is also available at: github.com/jgolob/arf</p>

opencc-by-4.0Jul 2022View details →
dryad36/100

Revisiting the number of self‐incompatibility alleles in finite populations: From old models to new results

<p>Under gametophytic self-incompatibility (GSI), plants are heterozygous at the self-incompatibility locus (S-locus) and can only be fertilized by pollen with a different allele at that locus. The last century has seen a heated debate about the correct way of modeling the allele diversity in a GSI population that was never formally resolved. Starting from an individual-based model, we derive the deterministic dynamics as proposed by Fisher (1958), and compute the stationary S-allele frequency distribution. We find that the stationary distribution proposed by Wright (1964) is close to our theoretical prediction, in line with earlier numerical confirmation. Additionally, we approximate the invasion probability of a new S-allele, which scales inversely with the number of resident S-alleles. Lastly, we use the stationary allele frequency distribution to estimate the population size of a plant population from an empirically obtained allele frequency spectrum, which complements the existing estimator of the number of S-alleles. Our expression of the stationary distribution resolves the long-standing debate about the correct approximation of the number of S-alleles and paves the way to new statistical developments for the estimation of the plant population size based on S-allele frequencies.</p>

opencc-zeroAug 2022View details →
zenodo36/100

Variable allelic expression of imprinted genes at the Peg13, Trappc9, Ago2 cluster in single neural cells

<p>Fig&nbsp;1 5&#39; RACE - Sequence tracks and analysis of alternative Trappc9 transcriptional start sites</p> <p>Fig 2 and suppl fig 3-&nbsp;Brain and Kidney tissue or Neural stem cells isolated from the Hippocampus region of newborn mice generated from a C57BL/6 (female) and a Cast/EiJ (male) cross (and its reciprocal cross) was used to determine allelic bias expression. A SNP located within an exon of Peg13, Trappc9, Ago2, Chrac1 and Kcnk9 was identified and amplified via pyrosequencing PCR with the percentage of SNP identification used to determine allele expression percentages. Additionally, Pyrorun sequences of reverse transcribed RNA from Trappc9 expression in different tissues. A hybrid cross between C57BL/6 and JF1 mouse was used to generate hybrid pups that were used to determine allele specificity of Trappc9 expression in Kidney and brain tissues.</p> <p>Figs 3, 4 &amp; 5- Single neural stem cells were isolated from the Hippocampus of newborn mice generated from a hybrid cross. Some of these cells were differentiated In vitro and either the NSC or differentiated neurons&nbsp;were lysed and underwent a reverse transcription. The newly formed cDNA was used as a template&nbsp;to amplify expressed Peg13, Trappc9 or Ago2 transcripts which was then sent for Sanger sequencing. A SNP located within the exon was used to determine whether the transcript from that cell was generated from the maternal or paternal allele.</p> <p>Fig 6- Brain-specific regulatory elements were cloned into a pGL [Luc] vector containing a Trappc9 promoter. These newly generated plasmids were then transfected into either primary neuron or fibroblast cultures alongside a Renilla vector for normalization&nbsp;using&nbsp;Lipofectamine as a&nbsp;transfection reagent. After 48 hours the cells were lysed and analyzed using a Glomax illuminator to determine their impact on Luciferase expression compared to that of the pGL vector containing just the Trappc9 promoter. Additionally,&nbsp;Plasmid vectors that were used for transfection of primary neurons to determine the impact of brain-specific regulatory elements on transcription. Plasmids can be visualized using the free software pdraw.32 downloaded from http://acaclone.com/download/install.htm</p> <p>Suppl fig 2-&nbsp;Single Neural stem cells and in vitro differentiated neurons&nbsp;isolated from Mouse Hippocampus tissue underwent reverse transcription and a qPCR reaction intended to amplify cDNA&nbsp;of genes associated with specific neural cell types as a method of detecting cell fate and whether this had an impact on allele-specific expression and differential methylation.</p> <p>Suppl fig 4 &amp; 5-&nbsp;DNA isolated from Neural stem cells was bisulfite converted for downstream identification of methyl group presence. CpG islands located at or near the promoters of the Peg13,&nbsp;Trappc9, Ago2, Chrac1 and Kcnk9 genes were amplified and cloned into a TOPO vector. The cloned segments were Sanger sequenced and compared to the original non-bisulfite converted sequence using the Quantification for methylation analysis (QUMA) tool to determine which CG dinucleotides were methylated and which weren&#39;t. Pyroruns determining methylation frequency at the&nbsp;CpG islands of these genes can be found in a separate upload on Zenodo.</p>

opencc-by-4.0Aug 2022View details →
dryad36/100

PAM-altering SNP-based allele-specific CRISPR-Cas9 therapeutic strategies for Huntington's disease

<p>Huntington's disease (HD) is caused by an expanded CAG repeat in huntingtin (<em>HTT</em>). Since HD is dominant, and loss of <em>HTT </em>leads to neurological abnormalities, safe therapeutic strategies require selective inactivation of mutant <em>HTT</em>. Previously, we proposed a concept of CRISPR-Cas9 using mutant-specific PAM sites generated by SNPs to selectively inactivate mutant <em>HTT</em>. Aiming at revealing suitable targets for clinical development, we analyzed the largest HD genotype dataset to reveal target <strong>P</strong>AM-<strong>a</strong>ltering <strong>S</strong>NPs (PAS) and subsequently evaluated their allele specificities. The gRNAs based on the PAM sites generated by rs2857935, rs16843804, and rs16843836 showed high levels of allele specificity in patient-derived cells. Simultaneous use of two gRNAs based on rs2857935-rs16843804 or rs2857935-rs16843836 produced selective genomic deletions in mutant <em>HTT </em>and prevented the transcription of mutant <em>HTT </em>mRNA without impacting the expression of normal counterpart or re-integration of the excised fragment elsewhere in the genome. RNAseq and off-target analysis confirmed high levels of allele specificity and the lack of recurrent off-targeting. Approximately 60% of HD subjects are eligible for mutant-specific CRISPR-Cas9 strategies of targeting one of these 3 PAS in conjunction with one non-allele-specific site, supporting high applicability of PAS-based allele-specific CRISPR approaches in the HD patient population.</p>

opencc-zeroAug 2022View details →
dryad36/100

Complete allele-specific silencing of the gain-of-function mutation of Huntington's disease

<p>Dominant gain-of-function mechanism in Huntington's disease (HD) suggests selective inactivation of mutant <em>HTT</em> produces the biggest therapeutic benefit. Here, we developed a complete allele-specific CRISPR/Cas9 strategy to permanently silence mutant <em>HTT</em> through nonsense-mediated decay (NMD), capitalizing on an exonic PAM (protospacer adjacent motif)-Altering SNP (PAS). Comprehensive sequence/haplotype analysis identified PAS-generated NGG PAM sites on exons of common <em>HTT </em>haplotypes in HD patients, revealing a single clinically meaningful PAS-based mutant-specific NMD-CRISPR/Cas9 strategy. The alternative allele of rs363099 eliminates NGG PAM site on the most frequent normal <em>HTT </em>haplotype in HD, permitting mutant <em>HTT-</em>specific CRISPR/Cas9 therapeutics in ~20% of HD patients with European ancestry. Our rs363099-based CRISPR/Cas9 showed perfect allele specificity and good targeting efficiencies in cells derived from HD patients. Dramatically reduced mutant <em>HTT </em>mRNA and complete loss of mutant HTT protein indicate that our allele-specific CRISPR/Cas9 strategy completely inactivates mutant <em>HTT </em>through NMD. RNAseq analysis also supported high levels of on-target gene specificity because no other genes except <em>HTT </em>were altered in clonal lines developed through our NMD-CRISPR/Cas9 strategy. Together, our data demonstrating significant target population, selective inactivation of mutant <em>HTT</em>, good targeting efficiency, and lack of recurrent off-targeting establish its therapeutic value of novel rs363099-based mutant <em>HTT-</em>specific NMD-CRISPR/Cas9 strategy in HD.</p>

opencc-zeroAug 2022View details →
dryad36/100

Raw data from: Natural alleles at the Doa locus underpin evolutionary changes in Drosophila lifespan and fecundity

<p>Evolve and resequence' (E&amp;R) studies in <em>Drosophila melanogaster</em> have identified many candidate loci underlying the evolution of ageing and life history, but experiments that validate the effects of such candidates remain rare. In a recent E&amp;R study we have identified several alleles of the LAMMER kinase <em>Darkener of apricot</em> (<em>Doa</em>) as candidates for evolutionary changes in lifespan and fecundity. Here, we use two complementary approaches to confirm the functional role of <em>Doa</em> in life-history evolution. First, we used transgenic RNAi to study the effects of<em> Doa</em> at the whole-gene level. Ubiquitous silencing of expression in adult flies reduced both lifespan and fecundity, indicating pleiotropic effects. Second, to characterize segregating variation at <em>Doa</em>, we examined four candidate single nucleotide polymorphisms (SNPs;<em> Doa-1, -2, -3, -4</em>) using a genetic association approach. Three candidate SNPs had effects that were qualitatively consistent with expectations based on our E&amp;R study: <em>Doa</em>-2 pleiotropically affected both lifespan and late-life fecundity; <em>Doa</em>-1 affected lifespan (but not fecundity), and <em>Doa</em>-4 affected late-life fecundity (but not lifespan). Finally, the last candidate allele (<em>Doa</em>-3) also affected lifespan, but in the opposite direction than predicted.</p>

opencc-zeroOct 2022View details →
zenodo36/100

Data and code from "Temporal allele frequency changes in large-effect loci reveal potential fishing impacts on salmon life-history diversity" (Miettinen et al. 2024)

<p>This archive contains code and data files to perform analyses detailed in Miettinen et al. (2024, Evolutionary Applications, https://doi.org/10.1111/eva.13690).</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

1000 Genomes Autosomal Allele Frequency Differences Across Ancestries (build 38)

<p>This dataset contains chromosome-specific allele frequency differences across different genetic ancestry groups for 73,159,508 autosomal SNPs captured in 2,548 individuals from the 1000 Genomes Project (build 38), available <a href="http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1000_genomes_project/release/20190312_biallelic_SNV_and_INDEL/">here</a>.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2014View details →
zenodo36/100

Distinguishing mutations and null alleles from genotyping errors using mother progeny comparisons in Brazilian pine (Araucaria angustifolia)

The use of microsatellite markers provides a window into the evolutionary processes of a given species. As such, these markers are widely used in scientific and applied research and are praised for their practicality and ease of use, however, the unavoidable incidence of genotyping deviations has been broadly neglected in the literature. Therefore, the present study aimed to estimate the rate of null alleles, mutations and genotyping errors in microsatellite loci, using Araucaria angustifolia, a threatened species, as a case study. We estimated the rates of the different types of genotyping deviations using mother-progeny genotype comparison from 50 seed-trees and their respective progeny (seeds). A total of 2336 A. angustifolia samples were genotyped, and we found that the rate of null alleles was 0.045. From the 1972 mother-progeny comparisons, the overall genotype deviation rate was 1.58%, consisting of 145 inconsistences (mutations), 339 null alleles and 210 genotyping errors. In terms of seed numbers, 128 (6.5%) showed inconsistencies in at least one locus, 118 (6.0%) null alleles, and 321 (16.3%) genotyping errors. This is the first study to describe the inconsistences (mutations) between mother-progeny genotypes for A. angustifolia, and the outcome makes it clear that an understanding of these genotyping deviations must be considered in assessing the accuracy of inferences made based on population genetics analyses.

opencc-zeroSep 2019View details →
zenodo36/100

Presence of the APOE4 allele is associated with an increased risk of sepsis progression

<p><strong>Supplementary Table 1. The Hardy-Weinberg equilibrium assay for APOE genotypes in healthy controls, sepsis, septic shock and all sepsis patients</strong>.</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

Distinguishing mutations and null alleles from genotyping errors using mother progeny comparisons in Brazilian pine (Araucaria angustifolia)

the rate of null alleles, mutations and genotyping errors in microsatellite loci, using Araucaria angustifolia, a threatened species, as a case study. We estimated the rates of the different types of genotyping deviations using mother-progeny genotype comparison from 50 seed-trees and their respective progeny (seeds). A total of 2336 A. angustifolia samples were genotyped, and we found that the rate of null alleles was 0.045. From the 1972 mother-progeny comparisons, the overall genotype deviation rate was 1.58%, consisting of 145 inconsistences (mutations), 339 null alleles and 210 genotyping errors. In terms of seed numbers, 128 (6.5%) showed inconsistencies in at least one locus, 118 (6.0%) null alleles, and 321 (16.3%) genotyping errors. This is the first study to describe the inconsistences (mutations) between mother-progeny genotypes for A. angustifolia, and the outcome makes it clear that an understanding of these genotyping deviations must be considered in assessing the accuracy of inferences made based on population genetics analyses.

opencc-zeroOct 2019View details →
zenodo36/100

Supplementary Dataset for the paper "Population suppression with dominant female-lethal alleles is boosted by homing gene drive"

<p>Supplementary Dataset for the paper "Population suppression with dominant female-lethal alleles is boosted by homing gene drive"</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Datasets belonging to the publication "Allelic variants confer Arabidopsis adaptation to small regional environmental differences".

<p>Datasets belonging to the publication "Allelic variants confer Arabidopsis adaptation to small regional environmental differences".</p> <p>A short description of each file is found below. Samples in the Variant Call Format (VCF) files are named according to how they are stored in the National Center for Biotechnology Information (NCBI) BioSample database. Detailed descriptions of how each file was generated are found in the manuscript.</p> <p><strong>Dartmap_Dutch_1001G_nuclear.vcf.gz</strong></p> <p>2,712,612 bi-allelic SNPs and 353,974 bi-allelic small indels called in the DartMap panel and the Dutch 1001G accessions (collectively referred to as DartMap + 1001G panel for brevity) relative to the <em>Arabidopsis thaliana</em> Col-0 nuclear genome reference sequence.</p> <p><strong>Dartmap_Dutch_1001G_mitochondrial.vcf.gz</strong></p> <p>231 bi-allelic SNPs and 23 bi-allelic small indels called in the DartMap panel and the Dutch 1001G accessions (collectively referred to as DartMap + 1001G panel for brevity) relative to the <em>A. thaliana</em> Col-0 mitochondrial genome reference sequence.</p> <p><strong>Dartmap_Dutch_1001G_chloroplast.vcf.gz</strong></p> <p>400 bi-allelic SNPs and 79 bi-allelic small indels called in the DartMap panel and the Dutch 1001G accessions (collectively referred to as DartMap + 1001G panel for brevity) relative to the <em>A. thaliana</em> Col-0 chloroplast genome reference sequence.</p> <p><strong>Dartmap_CNVs.vcf.gz</strong></p> <p>29,155 copy number variants (CNVs) called in the DartMap panel relative to the <em>A. thaliana</em> Col-0 reference genome.</p> <p><strong>Dartmap_Dutch_1001G_neighbouring_countries_nuclear.vcf.gz</strong></p> <p>Filtered VCF file containing variants of the DartMap + 1001G panel and that of neighbouring countries relative to the <em>A. thaliana</em> Col-0 nuclear genome reference sequence. Used for population structure analysis in the manuscript.</p> <p><strong>Dartmap_Dutch_1001G_neighbouring_countries_mitochondrial.vcf.gz</strong></p> <p>Filtered VCF file containing variants of the DartMap + 1001G panel and that of neighbouring countries relative to the <em>A. thaliana</em> Col-0 mitochondrial genome reference sequence. Used for population structure analysis in the manuscript.</p> <p><strong>Dartmap_Dutch_1001G_neighbouring_countries_chloroplast.vcf.gz</strong></p> <p>Filtered VCF file containing variants of the DartMap + 1001G panel and that of neighbouring countries relative to the <em>A. thaliana</em> Col-0 chloroplast genome reference sequence. Used for population structure analysis in the manuscript.</p> <p><strong>Dartmap_GWAS_GEA.vcf.gz</strong></p> <p>Filtered VCF file containing all variants of the DartMap panel that were used for all genome-wide association (GWA) and genome-environment association (GEA) analyses.</p> <p><strong>DartMap_phenotypical_data.xlsx</strong></p> <p>All relevant phenotypical data of the DartMap panel.&nbsp;</p> <p><strong>Dartmap_sample_ids_NCBI_sample_names.csv</strong></p> <p>Table showing which numerical ID of each DartMap sample corresponds to which NCBI BioSample ID.&nbsp;</p>

opencc-by-4.0Aug 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record