Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
130
datasets available to search
ShareScore release 0.7.1
Dataset results
130 results for “whole-genome sequencing”
Supplementary dataset to publication: Oxford nanopore technologies - a valuable tool to generate whole-genome sequencing data for in silico serotyping and the detection of genetic markers in Salmonella, Thomas et al 2023
<p>Bacteria of the genus <em>Salmonella</em> pose a major risk to livestock, the food economy, and public health. <em>Salmonella</em> infections are one of the leading causes of food poisoning. The identification of serovars of <em>Salmonella</em> achieved by their diverse surface antigens is essential to gain information on their epidemiological context. Traditionally, slide agglutination has been used for serotyping. In recent years, whole-genome sequencing (WGS) followed by <em>in silico</em> serotyping has been established as an alternative method for serotyping and the detection of genetic markers for <em>Salmonella</em>. Until now, WGS data generated with Illumina sequencing are used to validate <em>in silico</em> serotyping methods. Oxford Nanopore Technologies (ONT) opens the possibility to sequence ultra-long reads and has frequently been used for bacterial sequencing. In this study, ONT sequencing data of 28 <em>Salmonella</em> strains of different serovars with epidemiological relevance in humans, food, and animals were taken to investigate the performance of the <em>in silico</em> serotyping tools SISTR and SeqSero2 compared to traditional slide agglutination tests. Moreover, the detection of genetic markers for resistance against antimicrobial agents, virulence, and plasmids was studied by comparing WGS data based on ONT with WGS data based on Illumina. Based on the ONT data from flow cell version R9.4.1, <em>in silico</em> serotyping achieved an accuracy of 96.4 and 92% for the tools SISTR and SeqSero2, respectively. Highly similar sets of genetic markers comparing both sequencing technologies were identified. Taking the ongoing improvement of basecalling and flow cells into account, ONT data can be used for <em>Salmonella in silico</em> serotyping and genetic marker detection.</p>
Inferring whole-genome histories in large population datasets: inferred tree sequences for 1000 Genomes
<p>Tree sequences inferred for the 1000 Genomes phase 3 autosomes using <a href="https://tsinfer.readthedocs.io/">tsinfer</a> version 0.1.4 and compressed using <a href="https://tszip.readthedocs.io/en/stable/">tszip</a>. Tree sequences can be decompressed as follows:</p> <pre><code class="language-bash">$ tsunzip 1kg_chr1.trees.tsz</code></pre> <p>Once decompressed, trees files can be loaded and processed using <a href="https://tskit.readthedocs.io">tskit</a>. </p> <pre><code class="language-python">import tskit ts = tskit.load("1kg_chr1.trees") # ts is an instance of tskit.TreeSequence print("Chromosome 1 contains {} trees".format(ts.num_trees))</code></pre> <p>Metadata associated with individuals and populations was derived from the original <a href="http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/working/20130606_sample_info/20130606_g1k.ped">source</a> and converted to JSON form. For example, to access individual metadata we can use:</p> <pre><code class="language-python">import tskit import json ts = tskit.load("1kg_chr1.trees") ind = ts.individual(0) metadata_dict = json.loads(ind.metadata)</code></pre> <p>The metadata_dict variable will now contain all the metadata for the individual with ID 0 as a dictionary. Metadata associated with populations can be found in a similar way. Population IDs are associated with individuals via their constituent nodes. For example,</p> <pre><code class="language-python">pop_metadata = [json.loads(pop.metadata) for pop in ts.populations()] ind_node = ts.node(ind.nodes[0]) ind_pop_metadata = pop_metadata[ind_node.population]</code></pre> <p>After this, the ind_pop_metadata variable will contain the population level metadata for individual ID 0.</p> <p>The full data pipeline used to generate these tree sequences and associated metadata is available on <a href="https://github.com/mcveanlab/treeseq-inference/tree/master/human-data">GitHub</a>.</p>
Inferring whole-genome histories in large population datasets: inferred tree sequences for Simons Genome Diversity Project
<p>Tree sequences inferred for the SGDP autosomes using <a href="https://tsinfer.readthedocs.io/">tsinfer</a> version 0.1.4 and compressed using <a href="https://tszip.readthedocs.io/en/stable/">tszip</a>. Tree sequences can be decompressed as follows:</p> <pre><code class="language-bash">$ tsunzip sgdp_chr1.trees.tsz</code></pre> <p>Once decompressed, trees files can be loaded and processed using <a href="https://tskit.readthedocs.io">tskit</a>. </p> <pre><code class="language-python">import tskit ts = tskit.load("sgdp_chr1.trees") # ts is an instance of tskit.TreeSequence print("Chromosome 1 contains {} trees".format(ts.num_trees))</code></pre> <p>Metadata associated with individuals and populations was derived from the original <a href="https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/SGDP_metadata.279public.21signedLetter.samples.txt">source</a> and converted to JSON form. For example, to access individual metadata we can use:</p> <pre><code class="language-python">import tskit import json ts = tskit.load("sgdp_chr1.trees") ind = ts.individual(0) metadata_dict = json.loads(ind.metadata)</code></pre> <p>The metadata_dict variable will now contain all the metadata for the individual with ID 0 as a dictionary. Metadata associated with populations can be found in a similar way. Population IDs are associated with individuals via their constituent nodes. For example,</p> <pre><code class="language-python">pop_metadata = [json.loads(pop.metadata) for pop in ts.populations()] ind_node = ts.node(ind.nodes[0]) ind_pop_metadata = pop_metadata[ind_node.population]</code></pre> <p>After this, the ind_pop_metadata variable will contain the population level metadata for individual ID 0.</p> <p>The full data pipeline used to generate these tree sequences and associated metadata is available on <a href="https://github.com/mcveanlab/treeseq-inference/tree/master/human-data">GitHub</a>.</p>
Whole-genome capture and sequencing of Mycobacterium tuberculosis directly from clinical samples - Design of RNA oligonucleotide baits for Agilent Technologies' SureSelect target enrichment
<p>This dataset comprises the sequence of <strong>44 278 RNA oligonucleotide "baits" (120 bp each) </strong>designed to perform <strong>whole-genome capture and sequencing of <em>Mycobacterium tuberculosis</em> directly from clinical samples</strong> (DNA) using Agilent Technologies’ SureSelect target enrichment system following the Illumina paired-end multiplexed sequencing library protocol. </p> <p>RNA oligonucleotide “baits” were designed to span the ∼4.5 Mb of the <em>M. tuberculosis</em> genome. In brief, the reference genome sequence of the MTBC H37Rv strain (Genbank #AL123456) was <em>in silico</em> fragmented into 120 bp sequences twice, to ensure an overlap of 60 bp between sequences. Due to their rich GC content, which could interfere with DNA capture, all MTBC genes of the PE, PPE and PE-PGRS family were also independently fragmented into 120 bp sequences, in order to increase capture sensitivity. All resulting sequences were BLASTn searched against the Human Genomic + Transcript database to excluded homologous sequences to the human genome. Overall, a total of 42,278 RNA probes were generated and this custom bait library was then uploaded to the SureDesign software (https://earray.chem.agilent.com/suredesign) and synthesized by Agilent Technologies. During synthesis, the 2198 sequences complementary to the PE, PPE and PE-PGRS family were unbalanced 8:1 to potentiate capture.</p> <p>More details can be found in the following publication:</p> <p>- Macedo, R., Isidro, J., Ferreira, R., Pinto, M., Borges, V., Duarte, S., Vieira, L., & Gomes, J. P. (2023). Molecular Capture of <em>Mycobacterium tuberculosis</em> Genomes Directly from Clinical Samples: A Potential Backup Approach for Epidemiological and Drug Susceptibility Inferences. <em>International journal of molecular sciences</em>, <em>24</em>(3), 2912. https://doi.org/10.3390/ijms24032912</p>
Supplementary dataset to publication: "Genomic insight into Campylobacter jejuni isolated from commercial turkey flocks in Germany using whole-genome sequencing analysis"
<p><em>Campylobacter jejuni </em>is a zoonotic bacterium of public health significance. The present investigation was designed to assess the epidemiology and genetic heterogeneity of <em>Campylobacter jejuni</em> recovered from commercial turkey farms in Germany using whole-genome sequencing. The Illumina MiSeq<sup>®</sup> technology was used to sequence 66 <em>Campylobacter jejuni </em>isolates obtained between 2010 and 2011 from commercial meat turkey flocks located in ten German federal states. Phenotypic antimicrobial resistance was determined. Phylogeny, resistome, plasmidome and virulome profiles were analyzed using whole-genome sequencing data. Genetic resistancemarkers were identified with bioinformatics tools (AMRFinder, ResFinder, NCBI and ABRicate) and compared with the phenotypic antimicrobial resistance.</p>
Data For: Identifying rare variants inconsistent with identity-by-descent in population-scale whole-genome sequencing data
<p>Simulation output and Genome-wide scan for nIBD variants in UK10K data as reported in:</p> <p>Identifying rare variants inconsistent with identity-by-descent in population-scale whole-genome sequencing data</p> <p>Johnson KE, Adams CJ, Voight BF. Methods Ecol Evol 2022 Nov;13(11): 2429–2442.</p> <p>Code available at: https://github.com/kelsj/EVICORD</p>
Sei whole-genome sequence class annotations
<p>Sei sequence class whole-genome annotations are available in the following files:</p> <ul> <li> <p>sorted.hg38.tiling.bed.ipca_randomized_300.labels.merged.bed - The sorted, merged sequence class assignments from Louvain community clustering of the 30 million sequences, uniformly tiling the whole human genome. The fourth column is the sequence class number, with any sequence classes numbering 40-61 excluded from our analyses in the publication. Sequence classes 0-39 can be mapped to the following labels: <a href="https://github.com/FunctionLab/sei-framework/blob/main/model/seqclass.names">https://github.com/FunctionLab/sei-framework/blob/main/model/seqclass.names</a></p> </li> <li> <p>sorted.hg19.tiling.bed.ipca_randomized_300.labels.merged.bed - lifted over version of the hg38 BED file.</p> </li> </ul>
Data release: Whole-genome sequencing of Schistosoma mansoni reveals extensive diversity with limited selection despite mass drug administration
<p>Source data used in the publication: Berger et al. (2021) - Provisional title: 'Whole-genome sequencing of <em>Schistosoma mansoni</em> reveals extensive diversity with limited selection despite mass drug administration'. These data were used to generate all figures used in the publication and all files are organised and labelled specifically to run with the custom code that uses these data can be found at: http://doi.org/10.5281/zenodo.4975908. </p> <p><br> <strong>File descriptions:</strong></p> <p><strong>SOURCE DATA.zip - All source data for all figures. </strong></p> <p><strong>Figure 1b:</strong></p> <ul> <li>supplementary_data_9.txt - Metadata</li> </ul> <p><strong>Figure 2a&b:</strong></p> <ul> <li>207_PCA.eigenvec - PCA eigenvectors</li> <li>207_PCA.eigenval - PCA eigenvalues</li> </ul> <p><strong>Figure 2c:</strong></p> <ul> <li>autosomes.mdist - PLINK distance matrix used to build the neighbour joining phylogeny</li> </ul> <p><strong>Figure 2d:</strong></p> <ul> <li>all.pi.pixy.schools.txt - Nucleotide diversity results for each school subpopulation.</li> </ul> <p><strong>Figure 2e:</strong></p> <ul> <li>autosomes.dxy.5kb.schools.txt - Autosomal D<sub>XY</sub> results between school subpopulations. </li> <li>autosomes.fst.5kb.schools.txt - Autosomal F<sub>ST</sub> results between school subpopulations.</li> </ul> <p><strong>Figure 2f:</strong></p> <ul> <li>admixture_all.txt - ADMIXTURE results for each sample and population sizes, column 1 represents number of populations (K), columns 3-8 represent admixture values for each population. </li> </ul> <p><strong>Figure 3a, Supplementary figure 10a:</strong></p> <ul> <li>sfs.csv - Site frequency spectra (allelic proportions at each frequency bin) for each school. </li> </ul> <p><strong>Figure 3b:</strong></p> <ul> <li>TD.all.txt - Tajima's D values calculated in 5 kb windows for each school subpopulation. </li> </ul> <p><strong>Figure 4a, Supplementary figures 13-18: </strong></p> <ul> <li>ALL.MAYUGE.IHS.ihs.out.100bins.norm.txt.zip - Normalised iHS scores for the Mayuge district parasite populations (Selscan output).</li> </ul> <p><strong>Figure 4b, Supplementary figures 13-18: </strong></p> <ul> <li>ALL.TORORO.IHS.ihs.out.100bins.norm.txt.zip -<strong> - </strong>Normalised iHS scores for the Tororo district parasite populations (Selscan output).</li> </ul> <p><strong>Figure 4c, Supplementary figures 13-18: </strong></p> <ul> <li>ALL.MAYUGEvsTORORO.xpehh.xpehh.out.norm.txt.zip - - Normalised XP-EHH scores between Mayuge and Tororo parasite populations.</li> </ul> <p><strong>Figure 4d, Supplementary figures 13-18:</strong></p> <ul> <li>MAYUGE_TORORO_2000.windowed.weir.txt.zip - F<sub>ST</sub> values calculated between Mayuge and Tororo populations in 2kb windows. </li> </ul> <p><strong>Figure 4e, Supplementary figures 12a&c:</strong></p> <ul> <li>MAYUGE_PI.windowed.pi.zip - Nucleotide diversity values calculated in 2 kb windows for Mayuge populations. </li> <li>TORORO_PI.windowed.pi.zip - Nucleotide diversity values calculated in 2 kb windows for Kocoge populations (Tororo district).</li> </ul> <p><strong>Figure 5a:</strong></p> <ul> <li>all.pi.treat.fix.txt.zip - Nucleotide diversity results for each treatment subpopulation</li> </ul> <p><strong>Figure 5b</strong></p> <ul> <li>autosomes.dxy.5kb.treatment.txt - <strong> </strong>- Autosomal D<sub>XY</sub> results between clearance phenotype subpopulations. </li> <li>autosomes.fst.5kb.treatment.txt<strong> </strong>- Autosomal F<sub>ST</sub> results between clearance phenotype subpopulations. </li> </ul> <p><strong>Figure 5c:</strong></p> <ul> <li>fst.windows.2kb.treatment.txt.zip - F<sub>ST</sub> values for comparisons between different treatment groups (Pre-treatment, post-treatment (good clearers), post-treatment (poor clearers))</li> </ul> <p><strong>Figure 5d: </strong></p> <ul> <li>assoc_err_binary.txt.zip - Results of binary trait association between miracidia sampled from hosts with good clearance phenotypes (where treatment appeared to be highly effective) and miracidia isolated post-treatment from hosts with poor clearance phenotypes (where miracidia are potentially derived from parasites that survived treatment.</li> </ul> <p><strong>Figure 5e:</strong></p> <ul> <li>assoc_err_linear.txt.zip - - Results of linear regression genome-wide association study with the ERR estimates for all 198 samples, using the mean of the posterior ERR estimates from Crellen et al. (2016) as a quantitative trait.</li> </ul> <p><strong>Supplementary figure 1:</strong></p> <ul> <li>median.coverage.txt - Normalised depth of read coverage (column 4) calculated in 25 kb windows (columns 2&3) across all samples for all chromosomes (column 1).</li> </ul> <p><strong>Supplementary figure 2a-f: </strong></p> <ul> <li>cohort.genotyped.txt.zip - <strong> </strong>- Variant quality site values (used to inform variant site retention or removal). </li> </ul> <p><strong>Supplementary figure 2g:</strong></p> <ul> <li>hard_filtered.imiss.txt - Per sample variant missingness (used to inform quality control).</li> </ul> <p><strong>Supplementary figure 2h:</strong></p> <ul> <li>hard_filtered_filtindv.lmiss.txt.zip - Per site missingness (used to inform quality control).</li> </ul> <p><strong>Supplementary figure 3a, 4a, 4b:</strong></p> <ul> <li>prunedData.eigenvec - PCA eigenvectors</li> <li>prunedData.eigenval - PCA eigenvalues</li> </ul> <p><strong>Supplementary figure 3b:</strong></p> <ul> <li>pruned_data.mdist.csv - Distance matrix used as the basis for the neighbour joining phylogeny.</li> </ul> <p><strong>Supplementary figure 5:</strong></p> <ul> <li>cv_scores.txt - ADMIXTURE coefficient of variation scores (column 2) for each population size (1).</li> </ul> <p><strong>Supplementary figure 6:</strong></p> <ul> <li>*_SMC_SE.csv - SMC++ results (from 25 subsampled replicates) for each school subpopulation and outgroup samples. </li> </ul> <p><strong>Supplementary Figure 7:</strong></p> <ul> <li>smcpp.csv - SMC++ results for each school subpopulation and outgroup samples. </li> </ul> <p><strong>Supplementary Figure 8a-d</strong></p> <ul> <li>pi.per_host.txt.zip - Nucleotide diversity values for each host infrapopulation. </li> </ul> <p><strong>Supplementary Figure 9:</strong></p> <ul> <li>sexing.csv - inferred sex (based on differential read coverage over pseudoautosomal and Z-specific regions of the Z chromosome). </li> </ul> <p><strong>Supplementary Figure 10b:</strong></p> <ul> <li>sfs_res.csv - residuals for the SFS analysis in 3a/10a.</li> </ul> <p><strong>Supplementary Figure 11:</strong></p> <ul> <li>MAYUGE_TAJIMA_D.Tajima.D.2kb.txt.zip - Tajima's D values calculated for the Mayuge population in 2kb windows. </li> <li>Tororo_TAJIMA_D.Tajima.D.2kb.txt.zip - Tajima's D values calculated for the Tororo population in 2kb windows. </li> </ul> <p><strong>Supplementary Figures 13-18:</strong></p> <ul> <li>genes.bed - Coordinates of gene models (<em>S. mansoni </em>v7 annotation).</li> <li>KOCOGE_SITE_PI.sites.pi.txt.zip - Per site nucleotide diversity values</li> <li>MAYUGE_TORORO_sites.weir.fst.txt.zip - Per site F<sub>ST</sub> values between Mayuge and Tororo populations. </li> <li>coverage_5kb.windows.txt.zip - Per sample depth of read coverage in 5 kb windows. Columns 4,5,6 represent the median, mean and sstev of coverage for each 5kb window (columns 2&3) along each chromosome (column 1). </li> <li>median.sample.coverage.txt - Median chromosomal depth of read coverage for each sample. </li> </ul> <p><strong>Supplementary Figure 19:</strong></p> <ul> <li>kocoge_median.ld.txt.zip - <strong> </strong>- The decay of linkage disequilibrium with genomic distance between all sites within 50 kb for the Kocoge parasite samples. Chromosomes are shown in column 1, distance in column 2, median values in column 3. </li> <li>mayuge_median.ld.txt.zip - The decay of linkage disequilibrium with genomic distance between all sites within 50 kb for the Mayuge parasite samples. Chromosomes are shown in column 1, distance in column 2, median values in column 3. </li> </ul> <p><strong>Misc files:</strong></p> <p>schools.list - List of samples and schools where they were sampled. </p> <p> </p>
Results from the revision of MultiSuSiE improves multi-ancestry fine-mapping in All of Us whole-genome sequencing data
<p>afr47041.zip, lat36378.zip, and eur115620.zip contain All of Us Summary Statistics used in the revised version of "MultiSuSiE improves multi-ancestry fine-mapping in All of Us whole-genome sequencing data". Summary statistics for three cohorts are included: Afr47k, Lat36k, and Eur116k. These cohorts have not been downsampled to have equal levels of missingness.</p> <p>pips.tsv contains fine-mapped variants with PIP > 0.01 via MultiSuSiE from the revised version of "MultiSuSiE improves multi-ancestry fine-mapping in All of Us whole-genome sequencing data". Subcohorts with the _unmatched suffix have not been downsampled to have equal levels of phenotyped missingness across ancestries. </p> <p>MultiSuSiE-main.zip contains the MultiSuSiE software packages (corresponds to the Github repo on 10/16/2025).</p> <p>Please cite:</p> <p>Rossen, Jordan, et al. "MultiSuSiE improves multi-ancestry fine-mapping in All of Us whole-genome sequencing data." <em>medRxiv</em> (2024): 2024-05.</p> <p> </p>
Demogaphic Reconstruction of the Western Sheep Expansion from Whole-Genome sequences
<p>Sheep (<em>Ovis aries</em>) were among the earliest livestock, domesticated in the Fertile Crescent about 12000-10000 years ago with a nearly worldwide distribution today. Most of our knowledge about the timing of their expansion stems from archaeological data but it is unclear how the genetic diversity of modern sheep fits with these dates. We used whole-genome sequencing data of 63 domestic breeds and their wild relatives, the Asiatic mouflon (<em>O. gmelini</em>), to explore the demographic history of sheep. <br> On the global scale, our analysis revealed geographic structuring among breeds with unidirectional recent gene flow from domestics into Asiatic mouflons. We then selected four representative breeds from Spain, Morocco, the UK and Iran to build a comprehensive demographic model of the western sheep expansion.<br> We inferred a single domestication event around 9,000 years ago, slightly later than archaeological evidence suggests which might reflect uncertainties in the generation time used for these estimates. The westward expansion is dated to approximately 5,000 years ago, later than the original Neolithic expansion of sheep and approximately matching the Secondary Product Revolution associated with woolly sheep. We see some signals of recent gene flow from an ancestral population into southern European breeds which could reflect admixture with feral European mouflon. Furthermore, our results indicate that many breeds experienced a reduction of their effective population size during the last centuries, probably associated with the breed development.<br> Our study provides insights into the complex demographic history of western Eurasian sheep, highlighting interactions between breeds and their wild counterparts.</p>
Whole-genome sequences of Cercospora beticola isolates from Germany and Italy.
<p>Sequences used in the manuscript <strong>"Large-scale analyses reveal the contribution of adaptive evolution in pathogenic and non-pathogenic fungal species"</strong></p>
Supplemental material of "An annotated whole-genome multilocus sequence typing schema for scalable high resolution typing of Streptococcus pyogenes"
<p>This supplemental material includes the genome assemblies, associated metadata and analysis results for five datasets used to define a publicly available annotated wgMLST schema for <em>S. pyogenes</em> and to evaluate its suitability for high resolution typing. A brief description for each file in the dataset is available in the included README file. Raw sequencing data and sample metadata for the 265 isolates included in Dataset1 have been deposited in the European Nucleotide Archive (ENA) under Project <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB49967?show=reads">PRJEB49967</a>.</p> <p>The wgMLST schema was created with <a href="https://github.com/B-UMMI/chewBBACA">chewBBACA</a> and is publicly available at <a href="https://chewbbaca.online/species/1/schemas/1">chewie-NS</a>, where a more detailed description of schema creation, annotation and curation can be found.</p>
Integrative QTL mapping and selection signatures in Groningen White Headed cattle inferred from whole-genome sequences
<p>Here, we aimed to identify and characterize genomic regions that differ between Groningen White Headed (GWH) breed and other cattle, and in particular to identify candidate genes associated with coat color and/or eye-protective phenotypes. Firstly, whole genome sequences of 170 animals from eight breeds were used to evaluate the genetic structure of the GWH in relation to other cattle breeds by carrying out principal components and model-based clustering analyses. Secondly, the candidate genomic regions were identified by integrating the findings from: a) a genome-wide association study using GWH, other white headed breeds (Hereford and Simmental), and breeds with a non-white headed phenotype (Dutch Friesian, Deep Red, Meuse-Rhine-Yssel, Dutch Belted, and Holstein Friesian); b) scans for specific signatures of selection in GWH cattle by comparison with four other Dutch traditional breeds (Dutch Friesian, Deep Red, Meuse-Rhine-Yssel and Dutch Belted) and the commercial Holstein Friesian; and c) detection of candidate genes identified via these approaches. The alignment of the filtered reads to the reference genome (ARS-UCD1.2) resulted in a mean depth of coverage of 8.7X. After variant calling, the lowest number of breed-specific variants was detected in Holstein Friesian (148,213), and the largest in Deep Red (558,909). By integrating the results, we identified five genomic regions under selection on BTA4 (70.2–71.3 Mb), BTA5 (10.0–19.7 Mb), BTA20 (10.0–19.9 and 20.0–22.7 Mb), and BTA25 (0.5–9.2 Mb). These regions contain positional and functional candidate genes associated with retinal degeneration (e.g., <em>CWC27</em> and <em>CLUAP1</em>), ultraviole<em>t</em> protection (e.g., <em>ERCC8</em>), and pigmentation (e.g. <em>PDE4D</em>) which are probably associated with the GWH specific pigmentation and/or eye-protective phenotypes, e.g. Ambilateral Circumocular Pigmentation (ACOP). Our results will assist in characterizing the molecular basis of GWH phenotypes and the biological implications of its adaptation.</p>
Whole-genome sequencing reveals asymmetric introgression between two sister species of cold-resistant leaf beetles
Open the record for dataset details and reuse information.
ParaMask, a new method to identify multicopy genomic regions, corrects major biases in whole-genome sequencing data. Additional Datasets.
<p>Data supporting the main figures of the "ParaMask, a new method to identify multicopy genomic regions, corrects major biases in whole-genome sequencing data" manuscript and a copy of the ParaMask software and scripts for analysis, and SV calls from longreads. README files are included.</p>
EGP Mitochondrial Genome Analysis on Human Genome Diversity Project Whole-Genome Sequencing Data
<p><strong>Summary: </strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the HGDP. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on Simons Genome Diversity Project</strong>: Short-read WGS CRAM files were downloaded from the EMBL-EBI Public Data Globus Endpoint from the <code>/1000g/ftp/data_collections</code> directory. Post-download, the data was run through EGP version 1.3. The results are shown below:</p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>2b388c1fa446ecec70e33ea0471e06f8</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>50b80ed32b1ae542c8967cc31986dd19</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome Variant Tables</td> <td>995f30b74c4bb094a674b1a994853246</td> </tr> <tr> <td>Human Genome Diversity Project</td> <td>Mitochondrial Genome Copy Number</td> <td>e79e61efab4c491fa2825b7d1853df58</td> </tr> </tbody> </table> </div> <div>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</div>
EGP Mitochondrial Genome Analysis on Simons Genome Diversity Project Whole-Genome Sequencing Data
<p><strong>Summary: </strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the SGDP. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on Simons Genome Diversity Project</strong>: Short-read WGS CRAM files were downloaded from the EMBL-EBI Public Data Globus Endpoint from the <code>/1000g/ftp/data_collections</code> directory. Post-download, the data was run through EGP version 1.3. The results are shown below:</p> <table> <tbody> <tr> <th>Public Dataset</th> <th>EGP Result File Type</th> <th>MD5</th> </tr> </tbody> <tbody> <tr> <td>Simons Genome Diversity Project</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>86b09553f80926c1c29c57000ec1a88f</td> </tr> <tr> <td>Simons Genome Diversity Project</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>010026d77bee81e7b8daf5836bd12da3</td> </tr> <tr> <td>Simons Genome Diversity Project</td> <td>Mitochondrial Genome Variant Tables</td> <td>f1ea3edf4a82b42f2028467fb3544dc4</td> </tr> <tr> <td>Simons Genome Diversity Project</td> <td>Mitochondrial Genome Copy Number</td> <td>4872eeb792c214ad49662e98e4b14620</td> </tr> </tbody> </table> <p>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</p>
Whole-genome sequencing reveals contribution of rare and common variation to structural kidney and urinary tract malformations
<p>Supplementary tables detailing analysis of whole-genome sequencing data from 992 patients with congenital anomalies of the kidneys and urinary tract (CAKUT). </p>
EGP Mitochondrial Genome Analysis on 1000 Genomes Project 2504 Whole-Genome Sequencing Data
<p><strong>Summary: </strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the 1000 Genomes Project 2504 Dataset. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on 1000 Genomes Project 2504 Dataset</strong>: Short-read WGS CRAM files were downloaded through the paths present in this file <code>https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1000G_2504_high_coverage/1000G_2504_high_coverage.sequence.index</code>. Please note that the index files are there as well. You just have to append a <code>.crai</code>. The results are shown below:</p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>1000 Genomes Project 2504</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>dbf39d6ff0e4389b900f9d985f2e6c64</td> </tr> <tr> <td>1000 Genomes Project 2504</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>4d53ef60ec16f3e4b566c45fdf0fb977</td> </tr> <tr> <td>1000 Genomes Project 2504</td> <td>Mitochondrial Genome Variant Tables</td> <td>16925b546051d37cce27df8ec57ccc5e</td> </tr> <tr> <td>1000 Genomes Project 2504</td> <td>Mitochondrial Genome Copy Number</td> <td>365c1b360ea327795d981356064658a6</td> </tr> </tbody> </table> <p>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</p> </div>
EGP Mitochondrial Genome Analysis on 1000 Genomes Project 698 Related Whole-Genome Sequencing Data
<div> <p><strong>Summary: </strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the 1000 Genomes Project 698 Related Dataset. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on 1000 Genomes Project 698 Related Dataset</strong>: Short-read WGS CRAM files were downloaded through the paths present in this file <code>https://ftp-trace.ncbi.nlm.nih.gov/1000genomes/ftp/1000G_2504_high_coverage/additional_698_related/1000G_698_related_high_coverage.sequence.index</code>. The results are shown below:</p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>1000 Genomes Project 698 Related</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>322038d61b4da2e937b32410613c3532</td> </tr> <tr> <td>1000 Genomes Project 698 Related</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>36c782c12245100478903f7fa191a402</td> </tr> <tr> <td>1000 Genomes Project 698 Related</td> <td>Mitochondrial Genome Variant Tables</td> <td>68b2a51361ffae4e7ad9d420b8becd38</td> </tr> <tr> <td>1000 Genomes Project 698 Related</td> <td>Mitochondrial Genome Copy Number</td> <td>1e83c8ae132b0a7ef33b090757b29063</td> </tr> </tbody> </table> <p>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</p> </div> <p> </p> </div> <h2> </h2>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.