Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
219
datasets available to search
ShareScore release 0.7.1
Dataset results
219 results for “genotyping‐by‐sequencing”
Genotyping-by-sequencing (GBS) dataset for genome wide associations of growth, phenology and plasticity traits in willow (Salix viminalis (L.))
<p>These vcf-files constitute underlying raw data material for the manuscript "Genome wide associations of growth, phenology and plasticity traits in willow (Salix viminalis (L.))". For more detailed information please consult the README file in the repository.</p>
SNP and indel discovery and genotyping in next-generation sequencing data
<p>Code, logs and data for discovery and genotyping of SNPs and indels, in the the D.melanogaster genome, using GATK HaplotypeCaller. Code is in the zipped folder named code.zip. Run logs for this code as in the zipped folder named logs.zip. The unfiltered vcf genotypes file is named lhm_rg_HC_2015-09-15.vcf.gz. The filtered vcf genotypes file is named f1.lhm_rg_HC_raw.vcf.gz. The vcf submitted to NCBI dbSNP (filtered, and with indels >50bp and variants with null alternate alleles both removed) is named dbSNP.lhm_rg_HC_raw.vcf.gz. The folder local_reference.zip contains the reference assembly files against which genotypes were called against, and includes the code used to format the data prior to use. Also included is genotypes data from the two in-house reference line samples sequenced (BDGP6+ISO1 mito/dm6, Bloomington <em>Drosophila</em> Stock Center no. 2057)</p> <p>Samples are 220 Sussex-LH<sub>M</sub> hemiclones, and 2 RG. The first run did not include chromosome 4 and the mitochondrial genome, so these were genotyped separately, and then added to the rest of the results.</p> <p>The link for the NCBI dbSNP record is currently https://www.ncbi.nlm.nih.gov/projects/SNP/snp_viewBatch.cgi?sbid=1062461and the submitter handle is MORROW_EBE_SUSSEX.</p> <p>At the time of writting, the NCBI D.melanogaster build is still being updated, and therefore ss identifiers, but not rs identifers are available.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
Structural variant discovery and genotyping in next-generation sequencing data
<p>Code, logs, data, and summaries for detection and genotyping of genomic structural variants in the D.melanogaster Sussex LHM hemiclones (and one in-house reference line individual), using Genomestrip/2.0</p> <p>The unfiltered CNV pipleline results are lhm_gs.cnvs.raw.vcf.gz</p> <p>Filtered CNV results (including removal of bad samples) are filtered.goodS.lhm_gs.cnvs.raw.vcf.gz</p> <p>The file uploaded to NCBI dbVAR (which comprises of the filtered CNVs and indels >50bp from the HaplotypeCaller method) is lhm_sx16.dbVAR.vcf.gz</p> <p>The NCBI dbVAR accession number is nstd134. Code, logs and summary data are in the zipped archives, named accordingly. The archive reference_data.zip contains additional input files required for Genomestrip, including a shell script for making some of them. The file gstrip_lhm_RG_bams.list is also an input for Genomestrip, indicating bam file names and paths.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p> <p> </p>
DATASET: Genotyping by sequencing of the common bean Spanish Diversity Panel
<p>Genotyping by sequencing of 308 common bean lines included in the Spanish Diversity Panel. The ApeKI restriction enzyme was used. The sequencing reads were aligned using the reference genome V2.1 (https://phytozome.jgi.doe.gov/pz/portal.html#!info?alias=Org_Pvulgaris). A total of 11,763 SNP markers are included in this dataset after filtering for missing values (< 10%) and minor allele frequency (MAF> 0.05). </p>
Genotyping of European Toxoplasma gondii strains by a new high-resolution next-generation sequencing-based method
<p>The data set comprises 164 FASTQ files generated with an Ion AmpliSeq-based genotyping method for <em>Toxoplasma gondii </em>and<em> </em>a BED file used for the design of the Ion AmpliSeq primer panel. The FASTA file named as "AmpliSeq-ME49-Reference" was used as a reference for mapping and data analysis of the FASTQ files. The GZ file named as "Tgondii_IonAmpliSeq_Results_SNPs_VCF" is a VCF file, which contains all SNPs identified within the 164 FASTQ files relative to the AmpliSeq-ME49-Reference. The VCF file was converted into a FASTA file named as "Tgondii_IonAmpliSeq_Results_SNPs", which also contains the SNPs identified within the 164 FASTQ files relative to the AmpliSeq-ME49-Reference.</p> <p>The work is published in the European Journal of Clinical Microbiology & Infectious Diseases with the title "Genotyping of European <em>Toxoplasma gondii</em> strains by a new high‑resolution next‑generation sequencing‑based method"; https://doi.org/10.1007/s10096-023-04721-7</p>
Graphing and tabulating next-generation sequencing and genotyping data
<p>Making figures and tables for publication. Each zip archive contains input data, shell script to initiate and log R script, one R script for generating several graphs and tables, and the output graphs and tables themselves.</p> <p>Data was generated by whole-genome resequencing of 22 individual D.melanogaster from Sussex-LHM population and 2 from the Sussex RG line, followed by read-mapping, then genotyping with Haplotype Caller and Genomestrip.</p> <p>Locations for raw data, code, logs, extended QC data:</p> <p>Sequence reads NCBI SRA268956</p> <p>NCBI dbSNP https://www.ncbi.nlm.nih.gov/projects/SNP/snp_viewBatch.cgi?sbid=1062461</p> <p>NCBI dbVar accession number pre-release nstd134</p> <p> </p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p> <p> </p>
Genotype reproducibility testing in next-generation sequencing data
<p>Code, log and results summary for testing the reproducibility of genotypes with three pairs of hemiclones in the Sussex LH<sub>M </sub><em>D.melanogaster </em>population sample. Discovery and genotyping of genomic sequence variants was done using GATK HaplotypeCaller, and Genomestrip. Numerical comparison of genotype calls within each pairs of hemiclone individuals was performed using GATK GenotypeConcordance.</p> <p> </p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p>
CONGA: Copy number variation genotyping in ancient genomes and low-coverage sequencing data
<p>To date, ancient genome analyses have been largely confined to the study of single nucleotide polymorphisms (SNPs). Copy number variants (CNVs) are a major contributor of disease and of evolutionary adaptation, but identifying CNVs in ancient shotgun-sequenced genomes is hampered by (i) most published genomes being <1x coverage, (ii) ancient DNA fragments being typically <80 bps. These characteristics preclude state-of-the-art CNV detection software to be effectively applied to ancient genomes. Here we present CONGA, an algorithm tailored for genotyping deletion and duplication events in genomes with low depths of coverage. Simulations and down-sampling experiments show that CONGA can genotype deletions >1 kbps with F-scores >0.75 at >=1x, and distinguish between heterozygous and homozygous states. Using CONGA, we analyse deletion events at 10,018 loci in 56 ancient human genomes spanning the last 50,000 years, with coverages 0.4x-26x. We show that inter-individual genetic diversity measured using deletions and SNPs are highly correlated, as in modern-day genomes, confirming that deletion frequencies broadly reflect demographic history. We also identify signatures of strong purifying selection on deletions in ancient-genomes, such as an excess of singletons compared to those in SNPs. CONGA paves the way for systematic studies of drift, mutation load, and adaptation in ancient and modern-day gene pools through the lens of CNVs.</p>
Data for: Range and niche expansion through multiple interspecific hybridization - a genotyping by sequencing analysis of Cherleria (Caryophyllaceae)
<p><b>Background:</b> <i>Cherleria</i> (Caryophyllaceae) is a circumboreal genus that also occurs in the high mountains of the northern hemisphere. In this study, we focus on a clade that diversified in the European High Mountains, which was identified using nuclear ribosomal (nrDNA) sequence data in a previous study. With the nrDNA data, all but one species was monophyletic, with little sequence variation within most species. Here, we use genotyping by sequencing (GBS) data to determine whether the nrDNA data showed the full picture of the evolution in the genomes of these species.</p> <p><b>Results:</b> The overall relationships found with the GBS data were congruent with those from the nrDNA study. Most of the species were still monophyletic and many of the same subclades were recovered, including a clade of three narrow endemic species from Greece and a clade of largely calcifuge species. The GBS data provided additional resolution within the two species with the best sampling, <i>C. langii</i> and <i>C. laricifolia</i>, with structure that was congruent with geography. In addition, the GBS data showed significant hybridization between several species, including species whose ranges did not currently overlap.</p> <p><b>Conclusions:</b> The hybridization led us to hypothesize that lineages came in contact on the Balkan Peninsula after they diverged, even when those lineages are no longer present on the Balkan Peninsula. Hybridization may also have helped lineages expand their niches to colonize new substrates and different areas. Not only do genome-wide data provide increased phylogenetic resolution of difficult nodes, they also give evidence for a more complex evolutionary history than what can be depicted by a simple, branching phylogeny.</p>
Structural modelling results to accompany the paper "Uncommon mutational profiles of metastatic colorectal cancer detected during routine genotyping using next generation sequencing: an update"
<p>This repository contains the results of modelling missense mutants in KRAS, NRAS and BRAF observed in our study in the corresponding protein structures. Modelling was performed using FoldX.</p>
Simultaneous genotyping of snails and infecting trematode parasites using high-throughput amplicon sequencing.
<p>Several methodological issues currently hamper the study of entire trematode communities within populations of their intermediate snail hosts. Here we develop a new workflow using high-throughput amplicon sequencing to simultaneously genotype snail hosts and their infecting trematode parasites. We designed primers to amplify 4 snail and 5 trematode markers in a single multiplex PCR. While also applicable to other genera, we focused on medically and economically important snail genera within the Superorder Hygrophila and targeted a broad taxonomic range of parasites within the Class Trematoda. We tested the workflow using 417 <i>Biomphalaria glabrata </i>specimens experimentally infected with <i>Schistosoma rodhaini</i>, two strains of<i> Schistosoma mansoni</i>,<i> </i>and combinations thereof. We evaluated the reliability of infection diagnostics, the robustness of the workflow, its specificity related to host and parasite identification, and the sensitivity to detect co-infections, immature infections, and changes of parasite biomass during the infection process. Finally, we investigated its applicability in wild-caught snails of other genera naturally infected with diverse trematode assemblages. After stringent quality control the workflow allows the identification of snails to species level, and of trematodes to taxonomic levels ranging from family to strain. It is sensitive to detect immature infections and changes in parasite biomass described in previous experimental studies. Co-infections were successfully identified, opening the possibility to examine parasite-parasite interactions such as interspecific competition. Altogether, these results demonstrate that our workflow provides a powerful tool to analyze the processes shaping trematode communities within natural snail populations.</p>
Genome-wide SNP discovery in native American and Hungarian Robinia pseudoacacia genotypes using next-generation double-digest restriction-site-associated DNA sequencing (ddRAD-Seq)
<p>Initial filtered ddRADseq dataset with highly variable SNP markers from native American and Hungarian <em>Robinia pseudoacacia</em> L. individuals</p>
Linkage maps and genotype data of strawberry produced with skim-sequencing data
<p>The following set of files contain the results and scripts to produce those results, described in Chapter 5 of the PhD thesis of Alejandro Thérèse Navarro, entitled "How to map a million markers: linkage mapping of skim-sequencing data in strawberry". In this study, a large dataset of markers produced by whole genome resequecning of a strawberry (<em>Fragaria </em>x <em>ananassa</em>) biparental population are used to generate linkage maps. To that end the software <a href="https://github.com/Alethere/SmoothDescent">Smooth Descent</a> is used, since it is oriented to obtaining linkage maps in usin low quality (error-prone) genotype data. With this methodology we were able to produce a linkage map of 27 out of 28 chromosomes of strawberry which containing 1.85M markers in ~2400 unique genetic mpositions. We also compare this map with a linkage map produced using SNP array data and with the genome sequence assembly "Camarosa".</p>
Genomic characterization and gene bank curation of Aegilops using genotyping-by-sequencing
<p>In this study, genotyping-by-sequencing (GBS) was performed on 1041 <em>Aegilops</em> accessions, representing 23 different species. These accessions have been maintained by the Wheat Genetics and Resource Center (WGRC) at Kansas State University. The GBS FASTQ files have been uploaded to the NCBI SRA public repository under the BioProject accession number # PRJNA985892. We have provided other files related to data analysis, such as the barcode key file, SNP matrices, and taxonomic information of the accessions in this Dryad repository, which can be accessed through the provided link. The aim of the study was to explore the genetic and genomic characteristics of wild wheat relatives, <em>Aegilops,</em> using a larger number of SNP markers. Here, we also curated the WGRC gene bank <em>Aegilops</em> collection via the identification of misclassified accessions and genetically identical redundant accessions. Further, we explored the genomic relationship between wheat and the different <em>Aegilops</em> species. </p>
Data for: Range and niche expansion through multiple interspecific hybridization - a genotyping by sequencing analysis of Cherleria (Caryophyllaceae)
Open the record for dataset details and reuse information.
Data from: Multiple genotypes of Phelipanche ramosa indicate repeated introductions to the Americas: Sequence alignments and phylogenetic trees
Open the record for dataset details and reuse information.
Genotyping by sequencing for estimating relative abundances of diatom taxa in mock communities
Open the record for dataset details and reuse information.
Genomic characterization and gene bank curation of Aegilops using genotyping-by-sequencing
Open the record for dataset details and reuse information.
Simultaneous genotyping of snails and infecting trematode parasites using high-throughput amplicon sequencing.
Open the record for dataset details and reuse information.
Data from: RapidRat: development, validation and application of a genotyping-by-sequencing panel for rapid biosecurity and invasive species management
<p>Invasive alien species (IAS) are among the main causes of global biodiversity loss. Invasive brown (Rattus norvegicus) and black (R. rattus) rats, in particular, are leading drivers of extinction on islands, especially in the case of seabirds where >50% of all extinctions have been attributed to rat predation. Eradication is the primary form of invasive rat management, yet this strategy has resulted in a ~10-38% failure rate on islands globally. Genetic tools can help inform IAS management, but such applications to date have been largely reactive, time-consuming, and costly. Here, we developed a Genotyping-in-Thousands by sequencing (GT-seq) panel for rapid species identification and population assignment of invasive brown and black rats (RapidRat) in Haida Gwaii, an archipelago comprising ~150 islands off the central coast of British Columbia, Canada. We constructed an optimized panel of 443 single nucleotide polymorphisms (SNPs) using previously generated double-digest restriction-site associated DNA (ddRAD) genotypic data (27,686 SNPs) from brown (n=295) and black rats (n=241) sampled throughout Haida Gwaii. The informativeness of this panel for identifying individuals to species and island of origin was validated relative to the ddRAD results; in all comparisons, admixture coefficients and population assignments estimated using RapidRat were consistent. To demonstrate application, 20 individuals from novel invasions of three islands (Agglomerate, Hotspring, Ramsay) were genotyped using RapidRat, all of which were confidently assigned (>98.5% probability) to Faraday and Murchison Islands as putative source populations. These results indicated that a previous eradication on Hotspring Island was conducted at an inappropriate geographic scale; future management should expand the eradication unit to include neighboring islands to prevent re-invasion. Overall, we demonstrated that RapidRat is an effective tool for managing invasive rat populations in Haida Gwaii and provided a clear framework for GT-seq panel development for informing biodiversity conservation in other systems.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.