Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,785
datasets available to search
ShareScore release 0.9.0
Dataset results
2,785 results for “Genotype”
HPAIV genotypes in Germany
<p>Highly pathogenic avian influenza viruses (HPAIV) of the H5 goose/Guangdong (gs/GD) lineage have repeatedly emerged in Germany since 2006. Rooted in the respective gs/GD lineages HPAIV in Germany have been genetically diversified into a plethora of clades and subclades and evolved into an assortment of sub- and genotypes. The image summarizes the identified genotypes in Germany.</p>
Data for: Intergenerational genotypic interactions drive collective behavioural cycles in a social insect
<p>Many social animals display collective activity cycles based on synchronous behavioural oscillations across group members. A classic example is the colony cycle of army ants, where thousands of individuals undergo stereotypical biphasic behavioural cycles of about one month. Cycle phases coincide with brood developmental stages, but the regulation of this cycle is otherwise poorly understood. Here, we probe the regulation of cycle duration through interactions between brood and workers in an experimentally amenable army ant relative, the clonal raider ant. We first establish that cycle length varies across clonal lineages using long-term monitoring data. We then investigate the putative sources and impacts of this variation in a cross-fostering experiment with four lineages combining developmental, morphological, and automated behavioural tracking analyses. We show that cycle length variation stems from variation in the duration of the larval developmental stage, and that this stage can be prolonged not only by the clonal lineage of brood (direct genetic effects), but also of the workers (indirect genetic effects). We find similar indirect effects of worker line on brood adult size and, conversely but more surprisingly, indirect genetic effects of the brood on worker behaviour (walking speed and time spent in the nest).</p>
43 longevity-associated SNPs genotyped in a Croatian sample of oldest-old individuals
<p>This dataset presents genotype data for 43 single nucleotide polymorphisms (SNPs) that have been genotyped in an anonymised sample of 314 oldest-old individuals (85+ years) from Croatia. The SNPs are located in or near candidate genes for longevity, and were selected from publicly available literature databases (PubMed and repositories specialized for human longevity such as https://genomics.senescence.info/longevity/, http://ageing-map.org/). They were selected based on their strong or repeatedly reported association with human longevity and involvement in various metabolic pathways. Genotyping was performed by Kompetitive Allele Specific PCR (KASP) on genomic DNA isolated from peripheral blood using the salting-out method. The dataset also contains recoding of the genotypes for each participant according to their association with longevity: a value of 2 was assigned to the homozygous genotype of longevity allele, a value of 1 to the heterozygous genotype, and a value of 0 to the homozygous genotype of an allele not associated with longevity in our sample. In cases where there were less than 10 of either homozygous genotypes, and in cases where a dominant or recessive coding gave a more significant result in further analyses, they were additionally recoded as binary variables with only the values 0 and 1, with heterozygote being added to the less common homozygote. This data was used to perform logistic regression analyses to create the best models for predicting survival to the ages of 90 and 95. Those models were then used to create genetic risk scores (here named genetic longevity scores, GLS) for predicting that phenotype, which are shown in this dataset as well. Information about the selected SNPs is also presented: rs code, nearest gene, chromosome position, and references for literature sources where association with longevity is reported; along with data that refers to the studied Croatian population: alleles (major/minor), minor allele frequencies (MAF), genotyping success rate, and HWE p-values. </p>
Multi-location trials and population-based genotyping reveal high diversity and adaptation to breeding environments in a large collection of red clover
<p>This dataset accompanies the article with the same title made available on bioRxiv <a href="https://doi.org/10.1101/2022.12.19.520744">https://doi.org/10.1101/2022.12.19.520744</a> </p>
Genome-wide SNP discovery in native American and Hungarian Robinia pseudoacacia genotypes using next-generation double-digest restriction-site-associated DNA sequencing (ddRAD-Seq)
<p>Initial filtered ddRADseq dataset with highly variable SNP markers from native American and Hungarian <em>Robinia pseudoacacia</em> L. individuals</p>
Data: More than 1000 genotypes are required to derive robust relationships between yield, yield stability and physiological parameters: a computational study on wheat crop
<p>APSIM-Wheat <strong>(</strong><a href="">www.apsim.info</a><strong>)</strong> was used to simulate a data set (for details, see Casadebaig<em> et al.</em>, 2016) with 9100 virtual genotypes (<em>N</em><sub>gen</sub>= 9100) grown under 9000 environments (<em>N</em><sub>env</sub>=9000). In short, virtual genotypes were created by varying the value of 90 independent physiological parameters in a range of ±20% from the reference cultivar <em>Hartog</em>. Environments in the dataset contain historical climate data of 125 years (1889-2013) in four locations (Emerald, Narrabri, Yanco and Merredin) in Australia, in combination with two CO<sub><sup>2</sup></sub> levels (380 and 555 ppm), three nitrogen levels (low: 50%, control: 100% and high fertilization: 100% plus 50 kg‧ha<sup>-1</sup>) and three sowing dates (early, control and late).</p>
Index for use with the genotyper KAGE
<p>Index for use with the genotyper KAGE. This index is for all variants > 0.1% allele frequency from the 1000 Genomes Project and is generated by this pipeline: https://github.com/ivargr/genotyping-benchmarking, commit id 54678efa7336129b26f58a36227f5bbe5c8b9599.</p>
Detecting frequency-dependent selection through the effects of genotype similarity on fitness components
<p>Frequency-dependent selection (FDS) is an evolutionary regime that can maintain or reduce polymorphisms. Despite the increasing availability of polymorphism data, few effective methods are available for estimating the gradient of FDS from the observed fitness components. We modeled the effects of genotype similarity on individual fitness to develop a selection gradient analysis of FDS. This modeling enabled us to estimate FDS by regressing fitness components on the genotype similarity among individuals. We detected known negative FDS on the visible polymorphism in a wild <em>Arabidopsis</em> and damselfly by applying this analysis to single-locus data. Further, we simulated genome-wide polymorphisms and fitness components to modify the single-locus analysis as a genome-wide association study (GWAS). The simulation showed that negative or positive FDS could be distinguished through the estimated effects of genotype similarity on simulated fitness. Moreover, we conducted the GWAS of the reproductive branch number in <em>Arabidopsis thaliana</em> and found that negative FDS was enriched among the top-associated polymorphisms of FDS. These results showed the potential applicability of the proposed method for FDS on both visible polymorphism and genome-wide polymorphisms. Overall, our study provides an effective method for selection gradient analysis to understand the maintenance or loss of polymorphism.</p>
Datset related to article "G507D MUTATION IN FUS GENE CAUSES FAMILIAL AMYOTROPHIC LATERAL SCLEROSIS WITH A SPECIFIC GENOTYPE-PHENOTYPE CORRELATION"
<p><strong>NGS_analysis performed at Fondazione Besta carried out as part of the study reported at title</strong></p>
A multi-lab experimental assessment reveals that replicability can be improved by using empirical estimates of genotype-by-lab interaction
<p>Raw data sets, curated data sets and R code underlying the paper "A multi-lab experimental assessment reveals that replicability can be improved by using empirical estimates of genotype-by-lab interaction" </p> <p>https://doi.org/10.1101/2021.12.05.471264</p>
Dataset related to article"G507D MUTATION IN FUS GENE CAUSES FAMILIAL AMYOTROPHIC LATERAL SCLEROSIS WITH A SPECIFIC GENOTYPE-PHENOTYPE CORRELATION"
<p>NGS Analyses of patients included in the study reported at title followed at Fondazione Besta</p>
Genotype and genetic diversity data for: Contrasts in riverscape patterns of intraspecific genetic variation in a diverse Neotropical fish community of high conservation value
<p><span>Spatial patterns in genetic variation compared across species provide information about the predictability of genetic diversity of natural populations and areas requiring conservation measures. Due to their remarkable fish diversity, rivers in Neotropical regions are ideal systems to confront theory with observations and would benefit greatly from such approaches given their increasing vulnerability to anthropogenic pressures. We used SNP data from 18 fish species with contrasting life-history traits, co-sampled across 12 sites in the Maroni – a major river system from the Guiana Shield – to compare patterns of intraspecific genetic variation and identify their underlying drivers. Analyses of covariance revealed a decrease in genetic diversity as distance from the river outlet increased for 5 of the 18 species, illustrating a pattern commonly observed in riverscapes for species with low-to-medium dispersal abilities. However, mean within-site genetic diversity was lowest in the two easternmost tributaries of the Upper Maroni and around an urbanized location downstream, indicating the need to address the potential influence of local pressures in these areas, such as goldmining or fishing. Finally, the relative influence of isolation by stream distance, isolation by discontinuous river flow and isolation by spatial heterogeneity in effective size on pairwise genetic differentiation varied across species. Species with similar dispersal and reproductive guilds did not necessarily display shared patterns of population structure. Increasing the knowledge of specific life history traits and ecological requirements of fish species in these remote areas should help further understand factors that influence their current patterns of genetic variation.</span></p>
Dataset for Multiplex-PCR detection and Nanopore-based genotyping of fish pathogens
<p>This is a revised zip file contains scripts, initial fastq files, assembled amplicon (public and from this study) as well as bioinformatics intermediate files used for this study.</p> <p>Changelog:</p> <p>1. Fixed a bug in the 02_consensus.sh to enable proper removal of amplicons with zero depth</p> <p>2. Added a script (06_unclassified_read.sh) to extract and annotate reads that previously could not align to the 4 reference gene segment. Now the previously unclassified reads will be re-align (raw fastq) back to the gene segments as well as an additional tilapia genome assembly to gauge amount of reads mapping to the host genome. Furthermore, any read that still fail to align with minimap2 was subsequently aligned using blastn (-word_size 15 -evalue 0.01) against the same sequences.</p> <p>File Structure and Descriptions</p> <p>├── 01_process.sh : primer trimming, length-based filtering, read alignment, alignment filtering (unique hit) and extraction of uniquely hit reads for consensus generation<br> ├── 02_consensus.sh : [need artic conda env] Generation of consensus based on uniquely-mapped reads and minimal read depth of 20x required to call a variant (or it will be masked)<br> ├── 03_cleanup.sh: General folder and intermediate file re-organization<br> ├── 04_filter.sh: [need quast conda env] statistic of consensus generated and filtering of consensus with one or more ambiguous base (N), not suitable for haplotype<br> ├── 05_cluster.sh: clustering of consensus based on 100% identity threshold to generate putative haplotype<br> ├── 06_unclassified_read.sh: Extraction and annotation of unclassified reads using lenient criteria and with host reference genome as added reference<br> ├── Amplicon_FastQ folder: uniquely mapped fastq files for consensus generation<br> ├── BAM: alignment files generated from minimap2 used as input for the artic pipeline to identify variants<br> ├── Cluster_Rep.txt: Consensus sequences that were chosen to represent each haplotype<br> ├── Consensus folder: consensus fasta files generated for each sample containing sequences for each specific pathogen<br> ├── Coverage folder: coverage and base-level read depth for each sample and each pathogen reference genes<br> ├── Filter: individual fasta sequences (only 1 sequence per file) for each pathogen and each sample without any ambiguous base for subsequent clustering analysis<br> ├── Full_Haplotype.fasta: all possible haplotype sequences generated for each pathogen<br> ├── Gap_Analysis.tsv: Table with percentage of gap (0-100%) for each consensus sequence generated (used for filtering)<br> ├── Haplotype folder: Intermediate file and sample-level haplotype used to infer final haplotype and generate haplotype summary<br> ├── Haplotype_summary.tsv: Table with sample ID and their respectively pathogen haplotype<br> ├── Minimap2_PAF: Intermediate alignment generated from minimap2 used to generate the count table<br> ├── FailMinimap2 folder: FastQ files that didn't align using minimap2. Will be subsequently aligned using blastN (more sensitive) against the same reference sequences as minimap2<br> ├── Host_4Pathogen.fasta: Fasta file containing the tilapia genome and 4 pathogen (primer binding site included)<br> ├── Original: fastq with original naming prior to renaming based on sampleID. a script (rename.sh) was included to show renaming scheme<br> ├── primer.fasta: Primer sequences used for identifying and trimming reads with flanking primer sequence<br> ├── primer.fasta.fai: the index file for primer.fasta<br> ├── PrimerTrim folder: Primer-trimmed reads<br> ├── quast_results: consensus statistics generated by quast<br> ├── RawCount.tsv: Count table generated that can used as a input to generate figure<br> ├── RawFastq folder: Raw reads that have been renamed to reflect sample information<br> ├── readme.md: The current readme file<br> ├── ref_full_latest.fasta: Reference sequence of (gene segments) 4 pathogens e.g. TilV, ISKNV, SAG (Streptococcus agalactiae), FNO (Francisella noatunensis subsp. orientalis)<br> ├── ref_full_latest.primer.fasta: Same as above but with their primer binding sequence trimmed similar to the processed reads<br> ├── ref_full_latest.primer.fasta.fai<br> ├── RenameHaplotype: Script to perform reorganization of cdhit output<br> ├── Seq.stat.tsv: Sequencing statistics<br> ├── Uniq_PAF: Minimap2 alignment file for raw reads that initially failed quality check (no primer present and/or less than 80% query coverage / not unique alignment)<br> ├── Unmap: Raw reads that initial failed quality check (no primer on both ends / less than 80% query coverage / not unique alignment) <br> └── VCF: VCF files from medaka variant calling used to generate the final consensus</p>
Data for: Single-gene resolution of diversity-driven overyielding in plant genotype mixtures
<p>In plant communities, diversity often increases productivity and functioning, but the specific underlying drivers are difficult to identify. Most ecological theories attribute positive diversity effects to complementary niches occupied by different species or genotypes. However, the specific nature of niche complementarity often remains unclear, including how it is expressed in terms of trait differences between plants. Here, we use a gene-centred approach to study positive diversity effects in mixtures of natural <em>Arabidopsis </em><em>thaliana</em> genotypes. Using two orthogonal genetic mapping approaches, we find that between-plant allelic differences at the <em>AtSUC8</em> locus are strongly associated with mixture overyielding. <em>AtSUC8</em> encodes a proton-sucrose symporter and is expressed in root tissues. Genetic variation in <em>AtSUC8</em> affects the biochemical activities of protein variants and natural variation at this locus is associated with different sensitivities of root growth to changes in substrate pH. We thus speculate that - in the particular case studied here - evolutionary divergence along an edaphic gradient resulted in the niche complementarity between genotypes that now drives overyielding in mixtures. Identifying such genes important for ecosystem functioning may ultimately allow linking ecological processes to evolutionary drivers, help identify traits underlying positive diversity effects, and facilitate the development of high-performing crop variety mixtures.</p>
Data for: Selective elimination of enterovirus genotypes by activated sludge and chlorination
<p>Raw data underlying the journal article "Selective elimination of enterovirus genotypes by activated sludge and chlorination" by Larivé et al., <em>Environmental Science: Water Research and Technology</em>, 2023 (doi: 10.1039/d3ew00050h)</p> <p>One CSV file for each of Figures 2-6 of the main manuscript $</p> <p>One CSV file for each of Figures S5, S6 and S7 of the Supplementary information. The data for Figures S2, S3 and S4 are summarized in a single CSV file.</p>
Phyloseq 16S v3v4 object accompanying paper: Host genotype affects endotoxin release in excreta of broilers at slaughter age
<p>This ready to load <strong>phyloseq</strong> R S4 object contains the ASV table, taxonomy table and sample metadata. This data was build using the DaDa2 (version 1.18.0) and phyloseq (version 1.34.0) R packages using the SILVA 138.1 nr99 database and raw MiSeq PE300 sequencing data deposited at NCBI-SRA under BioProject: PRJNA975731.</p> <p>F. Marcato, J.M.J. Rebel, S.K. Kar, I. Wouters, D. Schokker, A. Bossers, F. Harders, J.W. van Riel, M. Wolthuis-Fillerup and I.C. de Jong. </p> <p>The main goal of the current study was to investigate whether or not it was possible to modulate the fecal microbiome and thereby reducing endotoxin concentrations in the excreta of broiler chickens. An experiment was carried out with a 2 x 2 x 2 factorial arrangement including 3 factors, 1) genetic strain (fast-growing Ross 308 vs. slower-growing Hubbard JA757), 2) no vs. combined use of probiotics and prebiotics in the diet and drinking water, 3) early feeding at the hatchery vs. non early feeding. A total of 624 Ross 308 and 624 Hubbard JA757 day-old male broiler chickens were included until d 37 and d 51 of age, respectively. Broilers (26 chicks per pen) were housed in a total of 48 pens, and there were 6 replicate pens/treatment group. Pooled cloacal swabs (10 chickens per pen) for microbiome and endotoxin analyses were collected at a target body weight (BW) of 200 g, 1 kg and 2.5 kg.</p>
Linkage maps and genotype data of strawberry produced with skim-sequencing data
<p>The following set of files contain the results and scripts to produce those results, described in Chapter 5 of the PhD thesis of Alejandro Thérèse Navarro, entitled "How to map a million markers: linkage mapping of skim-sequencing data in strawberry". In this study, a large dataset of markers produced by whole genome resequecning of a strawberry (<em>Fragaria </em>x <em>ananassa</em>) biparental population are used to generate linkage maps. To that end the software <a href="https://github.com/Alethere/SmoothDescent">Smooth Descent</a> is used, since it is oriented to obtaining linkage maps in usin low quality (error-prone) genotype data. With this methodology we were able to produce a linkage map of 27 out of 28 chromosomes of strawberry which containing 1.85M markers in ~2400 unique genetic mpositions. We also compare this map with a linkage map produced using SNP array data and with the genome sequence assembly "Camarosa".</p>
Trans-acting genotypes drive mRNA expression affecting metabolic and thermal tolerance traits
<p>Evolutionary processes driving physiological trait variation depend on the underlying genomic mechanisms. Evolution of these mechanisms depends on whether traits are genetically complex (involving many genes) and how gene expression that impacts the traits is converted to phenotype. Yet, genomic mechanisms that impact physiological traits are diverse and context-dependent (e.g., vary by environment or among tissues), making them difficult to discern. Here we examine the relationships between genotype, mRNA expression, and physiological traits to discern the genetic complexity and whether the gene expression affecting the physiological traits is primarily cis or trans-acting. We use low-coverage whole genome sequencing and tissue-specific mRNA expression among individuals to identify polymorphisms directly associated with physiological traits and expressed quantitative trait loci (eQTL) driving variation in six temperature-specific physiological traits (standard metabolic rate, thermal tolerance, and four substrate-specific cardiac metabolic rates). Not surprisingly, there were few, only five, SNPs directly associated with physiological traits. Yet, by focusing on a select set of mRNAs belonging to co-expression modules that explain up to 82% of temperature specific (12°C or 28°C) metabolism and thermal tolerance, we identified hundreds of significant eQTL for mRNA whose expression affects physiological traits. Surprisingly, most eQTL (97.4% for heart and 96.7% for brain) of eQTL were trans-acting. This could be due to higher effect size or greater importance of trans versus cis-acting eQTLs for mRNAs that are central to co-expression modules. That is, we may have enhanced the identification of trans-acting factors by looking for SNPs associated with mRNAs in co-expression modules that are known to be correlated with the expression of 10s or 100s of other genes, and thus have identified eQTLs with widespread effects on broad gene expression patterns. Overall, these data indicate that the genomic mechanism driving physiological variation across environments is driven by trans-acting tissue-specific mRNA expression.</p>
Data from: Maximum mutational robustness in genotype-phenotype maps follows a self-similar blancmange-like curve
<div class="section abstract"> <p>Phenotype robustness, defined as the average mutational robustness of all the genotypes that map to a given phenotype, plays a key role in facilitating neutral exploration of novel phenotypic variation by an evolving population. By applying results from coding theory, we prove that the maximum phenotype robustness occurs when genotypes are organised as bricklayer's graphs, so called because they resemble the way in which a bricklayer would fill in a Hamming graph. The value of the maximal robustness is given by a fractal continuous everywhere but differentiable nowhere sums-of-digits function from number theory. Interestingly, genotype-phenotype (GP) maps for RNA secondary structure and the HP model for protein folding can exhibit phenotype robustness that exactly attains this upper bound. By exploiting properties of the sums-of-digits function, we prove a lower bound on the deviation of the maximum robustness of phenotypes with multiple neutral components from the bricklayer's graph bound, and show that RNA secondary structure phenotypes obey this bound. Finally, we show how robustness changes when phenotypes are coarse-grained and derive a formula and associated bounds for the transition probabilities between such phenotypes.</p> </div>
Data platform (genotyping data set) related to ERDF postdoctoral project No. 1.1.1.2/VIAA/4/20/718 "The role of vitamin D gene polymorphisms and its receptors in the modulation of intestinal inflammation in patients with relapsing and progressive forms of multiple sclerosis".
<p><strong>Data platform </strong><strong>(genotyping dataset)</strong> <strong>related to the ERDF postdoctoral project No. </strong><strong>1.1.1.2/VIAA/4/20/718</strong><strong> “</strong><strong>The role of vitamin D and its receptor gene polymorphisms in the modulation of intestinal inflammation in patients with relapsing and progressive forms of multiple sclerosis</strong><strong>”.</strong></p> <p><strong>About the project and gathered data:</strong></p> <p>The dataset contains genotyping data on 289 sex-balanced samples (approximately 60% women / 40% men)) were created at the the multiple sclerosis (MS) Clinic of the Latvian Maritime Medical Center (LMMC) in 2011 (disease duration of 1-51 years); the collection was updated within the framework of the ERDF MS project (2017-2020) and replenished during the ERDF postdoctoral project No. 1.1.1.2/VIAA/4/20/718 “The role of vitamin D and its receptor gene polymorphisms in the modulation of intestinal inflammation in patients with relapsing and progressive forms of multiple sclerosis” (2021-2023).</p> <p>For the <strong>Genotyping dataset </strong>relevant information for each patient from the MS disease cohort, referring to proteasomal gene genetic variations (microsatellites and SNPs): (HSMS006 <em>(PSMA6),</em> HSMS602 <em>(FAM177A1),</em> HSMS701 <em>(KIAA0391)</em>, HSMS702 <em>(KIAA0391)</em> HSMS801 <em>(KIAA0391)</em>, rs11543947<em>(PSMB5), </em>rs2277460 (mi110), rs1048990 (mi8)<em> (PSMA6),</em> rs1048990 (mi8)<em> (PSMA6),</em> rs2295826/rs2295827<em>(PSMC6),</em> rs2348071 <em>(PSMA3),</em> rs2071543, rs9357155 <em>(PSMB8),</em> rs17587<em>(PSMB9),</em> rs74421874 <em>(PSMD9); </em>rs9275596 from HLA region; vitamin D-related genes (VDR and GC) polymorphisms: rs2228570, rs1544410, rs7975232, rs731236 (<em>VDR</em>) and rs7041, rs4588 <em>(GC).</em></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.