Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,549

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,549 results for “Genome sequencing”

Learn how ShareScore rates datasets ↗
zenodo52/100

Supplementary dataset to publication: Complete Genome Sequence of Ovine Mycobacterium avium subsp. paratuberculosis Strain JIII-386 (MAP-S/type III) and Its Comparison to MAP-S/type I, MAP-C, and M. avium Complex Genomes.

<p>This is the modified supplemented material to the publication &ldquo;Complete genome sequence of ovine Mycobacterium avium subsp. paratuberculosis strain JIII-386 (MAP-S/type III) and its comparison to MAP-S/type I, MAP-C, and M. avium complex genomes&rdquo;.</p> <p>The complete circular genome of Mycobacterium avium subsp. paratuberculosis (MAP) strain JIII-386 from Germany, closed by Nanopore technology in this study, was presented and compared with the draft genome of JIII-386, previously published in [doi:10.1093/gbe/ew154], the closed genome of the MAP-S/type I strain Telford, the MAP-S/type III draft genome of strain S397, twelve closed MAP-C (type II) strains and eight closed Mycobacterium avium (M. a.) strains of subsp. hominissuis (MAH) and subsp. avium (MAA). Structural comparisons clearly revealed the mosaic nature of MAP genomes, the differences between MAP subtypes I, II and III, and the higher diversity of MAP-S compared to MAP-C genomes.&nbsp;</p> <p>The material provides a wealth of detailed results from these analyses and comparisons. These include a list of identified ncRNA and Riboswitches, as well as additional genes in finished JIII-386, the gene content of identified prophage regions, copy number of identified transposable elements and a list of selected virulence-associated genes in the different MAP-type (I - III) strains. The genomic islands identified and included genes along with their predicted functions were presented for six MAP genomes (belonging to MAP-S/type I and III, and MAP-C), one MAH genome and one MAA genome. One table shows the corresponding genomic islands in the genomes of JIII-386, Telford and three MAP-C genomes. Furthermore, homologous genes of known MAP-S specific Large Sequence Polymorphisms regions (LSP<sup>S</sup> = LSP-S) were recorded in different MAP-S type strains, one MAH and one MAA strain, as well as genes of deletions #1 (LSP<sup>A</sup>-20), #2, and s-delta-1, previously described as MAP-S-specific deletions, their presence or absence in 3 MAP-S, 12 MAP-C, 4 MAH, and 4 MAA strains were listed. Different presence or absence of genes, but also identified frameshifts or disruptions of various virulence-associated genes could lead to the different MAP-type specific phenotypic characteristics. Comprehensive core and pan genome analyses (results listed in six tables) revealed unique genes and genes likely to have been acquired by horizontal gene transfer in different MAP types and subtypes, but also emphasized the highly conserved and close relationship, and the complex evolution of M. a. strains.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Public sequence accessions from INSDC, COG-UK and CNCB and EPI_SET from GISAID for SARS-CoV-2 genome sequences in 2023-08-01 UShER tree

<p>Genome sequences and metadata for the accessions in the .tsv.gz (gzip-compressed tab-separated text) files are freely available from their corresponding sources:</p><ul><li>insdc.accessionNameDate.tsv.gz: INSDC (GenBank, ENA, DDBJ) sequences and metadata may be downloaded using NCBI Datasets: https://www.ncbi.nlm.nih.gov/datasets/taxonomy/2697049/ (7,361,734 accessions used on 2023-08-01)</li><li>cog.accessionNameDate.tsv.gz: COG-UK sequences and metadata may be downloaded from https://cog-uk.s3.climb.ac.uk/phylogenetics/latest (as of publication); most COG-UK sequences have been submitted to ENA and are available from INSDC/NCBI Datasets as well. &nbsp;(724,978 accessions used on 2023-08-01)</li><li>cncb.accessionNameDate.tsv.gz: Sequences and metadata from several databases at the China National Center for Bioinformation (CNCB) may be downloaded from GenBase: https://ngdc.cncb.ac.cn/genbase/ (26,604 accessions used on 2023-08-01)</li></ul><p>GISAID data are subject to restrictions on sharing described in https://gisaid.org/terms-of-use/. &nbsp;Genome sequences and metadata are available to registered GISAID users as part of EPI_SET_231106ax at https://doi.org/10.55876/gis8.231106ax (7,718,061 accessions used on 2023-08-01).</p>

opencc-by-sa-4.0Nov 2023View details →
zenodo48/100

Timema genome sequences and annotations. Version 8.

<p>Genome sequence (fasta) files&nbsp;and annotation (gff) files for ten <em>Timema </em>species:&nbsp;<em>T. bartmani, T. cristinae, T. poppensis, T. californicum,&nbsp; T. podura, T. tahoe, T. monikensis, T. douglasi, T. shepardi, and&nbsp; T. genevievae.</em><br> <br> Species are abbreviated as follows: Tbi =&nbsp;<em>T. bartmani</em>, Tce =&nbsp;<em>T. cristinae</em>, Tps =&nbsp;<em>T. poppensis</em>, Tcm =&nbsp;<em>T. californicum</em>, Tpa =&nbsp;<em>T. podura</em>, Tte =&nbsp;<em>T. tahoe</em>, Tms =&nbsp;<em>T. monikensis</em>, Tdi =&nbsp;<em>T. douglasi</em>, Tsi =&nbsp;<em>T. shepardi</em>, and Tge =&nbsp;<em>T. genevievae</em><br> &nbsp;</p> <p>For details of assembly and annotation see:&nbsp;<br> <br> Jaron, K. S*., Parker, D. J*., Anselmetti, Y., Tran Van, P. T., Bast, J., Dumas, &nbsp;Z., Figuet, E., Fran&ccedil;ois, C. M., Hayward, K., Rossier, V., Simion, P., Robinson-Rechavi, &nbsp;M., Galtier, N., Schwander, T. 2021. Convergent consequences of parthenogenesis on stick insect genomes. bioRxiv. doi: https://doi.org/10.1101/2020.11.20.391540</p> <p>&nbsp;</p> <p><strong>File list:</strong><br> <br> Tbi_b3v08.fasta = T. bartmani genome sequence file<br> Tbi_b3v08.max_arth_b2g_droso_b2g.gff = T. bartmani genome annotation file<br> Tce_b3v08.fasta = T. cristinae genome sequence file<br> Tce_b3v08.max_arth_b2g_droso_b2g.gff = T. cristinae genome annotation file<br> Tcm_b3v08.fasta&nbsp;&nbsp; &nbsp; = T. bartmani genome sequence file<br> Tcm_b3v08.max_arth_b2g_droso_b2g.gff = T. californicum genome annotation file<br> Tdi_b3v08.fasta = T. douglasi genome sequence file<br> Tdi_b3v08.max_arth_b2g_droso_b2g.gff = T. douglasi genome annotation file<br> Tge_b3v08.fasta = T. genevievae genome sequence file<br> Tge_b3v08.max_arth_b2g_droso_b2g.gff = T. genevievae genome annotation file<br> Tms_b3v08.fasta = T. monikensis genome sequence file<br> Tms_b3v08.max_arth_b2g_droso_b2g.gff = T. monikensis genome annotation file<br> Tpa_b3v08.fasta = T. podura genome sequence file<br> Tpa_b3v08.max_arth_b2g_droso_b2g.gff = T. podura genome annotation file<br> Tps_b3v08.fasta = T. poppensis genome sequence file<br> Tps_b3v08.max_arth_b2g_droso_b2g.gff = T. poppensis genome annotation file<br> Tsi_b3v08.fasta = T. shepardi genome sequence file<br> Tsi_b3v08.max_arth_b2g_droso_b2g.gff = T. shepardi genome annotation file<br> Tte_b3v08.fasta&nbsp;&nbsp; &nbsp; = T. tahoe genome sequence file<br> Tte_b3v08.max_arth_b2g_droso_b2g.gff = T. tahoe genome annotation file</p>

opencc-by-4.0Oct 2021View details →
zenodo48/100

FASTA file containing to the MYB encoding gene Ant1 genomic sequences corresponding to wild and cultivated tomato accessions

<p>Fasta sequence correspond to the MYB encoding gene&nbsp;<em>An2-like</em>. The genomic&nbsp;sequences correspond to&nbsp;<em>Solanum&nbsp;galagpagnese</em> accession LA1141 (this study), <em>S.&nbsp;lycopersicum</em> variety OH8245 (this study), <em>S. lycopersicum</em> variety Heinz 1706 reference genome, and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014).&nbsp;Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018), were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at <a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>).</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

FASTA file containing the MYB encoding gene An2-like genomic sequences corresponding to wild and cultivated tomato accessions

<p>FASTA sequence corresponds&nbsp;to the MYB encoding gene&nbsp;<em>An2-like</em>. The genomic&nbsp;sequences correspond to&nbsp;<em>Solanum&nbsp;galagpagnese</em> accession LA1141 (this study), <em>S.&nbsp;lycopersicum</em> variety OH8245 (this study), <em>S. lycopersicum</em> variety Heinz 1706 reference genome (Hosmani et al., 2019),&nbsp;<em>S. lycopersicum </em>variety Indigo Rose (Yan et al., 2020), <em>S. lycopersicum</em> accession LA1996 [MN242011.1&nbsp;(Colanero et al., 2020)], <em>S. chilense&nbsp;</em>accession LA1930 [MN242012.1 (Colanero et al., 2020)], and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014).&nbsp;Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018), &nbsp;Indigo Rose [MN433087 (Yan et al., 2020)], <em>S. lycopersicum </em>accession LA1996 [MN242011.1, EF433417.1 (Sapir et al., 2008; Colanero et al., 2020)], <em>S. chilense</em> accession LA1930 [MN242012.1 (Colanero et al., 2020)] were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at <a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>)&nbsp;and&nbsp;the National Center for Biotechnology Information (NCBI)(available at NCBI: <a href="https://www.ncbi.nlm.nih.gov">https://www.ncbi.nlm.nih.gov</a>).</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

FASTA file containing the MYB encoding genes at the Aft locus with genomic sequences corresponding to wild and cultivated tomato accessions

<p>FASTA sequences correspond to the MYB encoding genes&nbsp;<em>An2-like </em>and <em>Ant1</em>. The genomic&nbsp;sequences were combined correspond to&nbsp;<em>Solanum&nbsp;galagpagnese</em>&nbsp;accession LA1141 (this study),&nbsp;<em>S.&nbsp;lycopersicum</em>&nbsp;variety OH8245 (this study),&nbsp;<em>S. lycopersicum</em>&nbsp;variety Heinz 1706 reference genome (Hosmani et al., 2019),&nbsp;LA1996 [MN242011.1, EF433417.1(Sapir et al., 2008; Colanero et al., 2020)],&nbsp;and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014).&nbsp;Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018),&nbsp;<em>S. lycopersicum&nbsp;</em>accession LA1996 [MN242011.1, EF433417.1 (Sapir et al., 2008; Colanero et al., 2020)],&nbsp;<em>S. chilense</em>&nbsp;accession LA1930 [MN242012.1 (Colanero et al., 2020)] were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at&nbsp;<a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>)&nbsp;and&nbsp;the National Center for Biotechnology Information (NCBI) (available at NCBI:&nbsp;<a href="https://www.ncbi.nlm.nih.gov/">https://www.ncbi.nlm.nih.gov</a>).</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

Clonal decomposition and DNA replication states defined by scaled single cell genome sequencing

<p><strong>OV2295&nbsp;Tables</strong></p> <p>ov2295_breakpoint_counts.csv.gz: Table of breakpoint counts per cell</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>cell_id: identifier for the cell</li> <li>read_count: number of reads</li> <li>library_id: identifier for the DNA library</li> <li>sample_id: identifier for the sequenced sample</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> </ul> <p>ov2295_cell_cn.csv.gz: Table of cell specific copy number</p> <ul> <li>cell_id: identifier for the cell</li> <li>sample_id: identifier for the sequenced sample</li> <li>library_id: identifier for the DNA library</li> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>reads: number of reads</li> <li>copy: raw normalized copy number</li> <li>state: copy number state</li> <li>gc: percent gc of the bin</li> <li>map: average mappability of the bin</li> </ul> <p>ov2295_cell_metrics.csv.gz: Table of cell metrics</p> <ul> <li>cell_id: identifier of the cell</li> <li>unpaired_mapped_reads: number of unpaired mapped reads</li> <li>paired_mapped_reads: number of mapped reads that were properly paired</li> <li>unpaired_duplicate_reads: number of unpaired duplicated reads</li> <li>paired_duplicate_reads: number of paired reads that were also marked as duplicate</li> <li>unmapped_reads: number of unmapped reads</li> <li>percent_duplicate_reads: percentage of duplicate reads</li> <li>estimated_library_size: scaled total number of mapped reads</li> <li>total_reads: total number of reads, regardless of mapping status</li> <li>total_mapped_reads: total number of mapped reads</li> <li>total_duplicate_reads: number of duplicate reads</li> <li>total_properly_paired: number of properly paired reads</li> <li>coverage_breadth: percentage of genome covered by some read</li> <li>coverage_depth: average reads per nucleotide position in the genome</li> <li>median_insert_size: median insert size between paired reads</li> <li>mean_insert_size: mean insert size between paired reads</li> <li>standard_deviation_insert_size: standard deviation of the insert size between paired reads</li> <li>index_sequence: index sequence of the adaptor sequence</li> <li>column: column of the cell on the nanowell chip</li> <li>img_col: column of the cell from the perspective of the microscope</li> <li>index_i5: id of the i5 index adapter sequence</li> <li>sample_type: type of the sample</li> <li>primer_i7: id of the i5 index primer sequence</li> <li>experimental_condition: experimental treatment of the cell, includes controls</li> <li>index_i7: id of the i7 index adapter sequence</li> <li>cell_call: living/dead classification of the cell based on staining usually, C1 == living, C2 == dead</li> <li>sample_id: name of the sample</li> <li>primer_i5: id of the i5 index primer sequence</li> <li>row: row of the cell on the nanowell chip</li> <li>library_id: identifier for the DNA library</li> <li>index: ignored</li> <li>multiplier: during parameter searching, the set [1..6] that was chosen</li> <li>MSRSI_non_integerness: median of segment residuals from segment integer copy number states</li> <li>MBRSI_dispersion_non_integerness: median of bin residuals from segment integer copy number states</li> <li>MBRSM_dispersion: median of bin residuals from segment median copy number values</li> <li>autocorrelation_hmmcopy: hmmcopy copy autocorrelation</li> <li>cv_hmmcopy: ignored</li> <li>empty_bins_hmmcopy: number of empty bins in hmmcopy</li> <li>mad_hmmcopy: median absolute deviation of hmmcopy copy</li> <li>mean_hmmcopy_reads_per_bin: mean reads per hmmcopy bin</li> <li>median_hmmcopy_reads_per_bin: median reads per hmmcopy bin</li> <li>std_hmmcopy_reads_per_bin: standard deviation value of reads in hmmcopy bins</li> <li>total_halfiness: summed halfiness penality score of the cell</li> <li>total_mapped_reads_hmmcopy: total mapped reads in all hmmcopy bins</li> <li>scaled_halfiness: summed scaled halfiness penalty score of the cell</li> <li>mean_state_mads: mean value for all median absolute deviation scores for each state</li> <li>mean_state_vars: variance value for all median absolute deviation scores for each state</li> <li>mad_neutral_state: median absolute deviation score of the neutral 2 copy state</li> <li>breakpoints: number of breakpoints, as indicated by state changes not at the ends of chromosomes</li> <li>mean_copy: mean hmmcopy copy value</li> <li>state_mode: the most commonly occuring state</li> <li>log_likelihood: hmmcopy log likelihood for the cell</li> <li>true_multiplier: the exact decimal value used to scale the copy number for segmentation</li> <li>order: order of the cell in the hierarchical clustering tree</li> <li>quality: random forest classifier proability score that cell is good</li> </ul> <p>ov2295_clone_alleles.csv.gz: Table of clone specific allele data</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>hap_label: haplotype block identifier</li> <li>clone_id: clone identifier</li> <li>allele_1_sum: number of reads for allele 1 of the haplotype block</li> <li>allele_2_sum: number of reads for allele 2 of the haplotype block</li> <li>total_counts_sum: total reads for the haplotype block</li> </ul> <p>ov2295_clone_breakpoints.csv.gz:&nbsp;Table of breakpoints per clone for OV2295 samples. Columns:</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> <li>clone_id: clone identifier</li> <li>read_count: number of reads</li> <li>is_present: presence=1, absent=0</li> </ul> <p>ov2295_clone_clusters.csv.gz: Table of cell clusters as putative clones</p> <ul> <li>cell_id: identifier for the cell</li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_cn.csv.gz:&nbsp;Table of allele specific copy number per clone for OV2295 samples. Columns:</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>total_cn: HMMCopy predicted total copy number&nbsp;</li> <li>minor_cn: HMM predicted minor copy number&nbsp;</li> <li>major_cn: HMM predicted major copy number&nbsp;</li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_snvs.csv.gz:&nbsp;Table of SNVs per clone for OV2295 samples.&nbsp; Columns:</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>clone_id: clone identifier</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>total_counts: total number of reads at this position</li> <li>is_present: presence=0, absent=1</li> <li>is_het:&nbsp;is heterozygous</li> <li>is_hom: is homozygous for the alternate</li> </ul> <p>ov2295_nodes.csv.gz: Table of phylogenetic information for SNV evolution</p> <ul> <li>variant_id: identifier for the SNV as chrom:coord:ref:alt</li> <li>node: node in the phylogenetic tree</li> <li>loss: probability the SNV was lost at this node</li> <li>origin: probability the SNV originated at this node</li> <li>presence: probability the SNV is present at this node</li> <li>ml_origin: binary indicator the SNV originated at this node</li> <li>ml_presence: binary indicator the SNV is present at this node</li> <li>ml_loss: binary indicator the SNV was lost at this node</li> </ul> <p>ov2295_snv_counts.csv.gz: Table of SNV counts</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>cell_id: identifier for the cell</li> <li>total_counts: total number of reads at this position</li> <li>sample_id: identifier for the sequenced sample</li> </ul> <p>ov2295_tree.pickle: Phylogenetic tree in python pickle format.&nbsp; Requires installation of the stochastic dollo code at: https://bitbucket.org/dranew/dollo, version 0.4.2.</p> <p>Note the following sample mapping: &lsquo;SA922&rsquo;: &lsquo;OV2295(R2)&rsquo;, &lsquo;SA921&rsquo;: &lsquo;TOV2295(R)&rsquo;, &lsquo;SA1090&rsquo;: &lsquo;OV2295&rsquo;,</p> <p><strong>Plots</strong></p> <p>ov_supp_clone_allele_cn.png: Clone allele ratios for each OV2295 sample.</p> <p>ov_supp_clone_total_cn.png: Clone copy number for each OV2295 sample.</p> <p>ov_supp_sample_total_cn.png: Bulk copy number for each OV2295 sample.</p> <p>ov_supp_sample_allele_cn.png: Bulk allele ratios for each OV2295 sample.</p>

opencc-by-4.0Sep 2019View details →
zenodo48/100

Genotyping-by-sequencing (GBS) dataset for genome wide associations of growth, phenology and plasticity traits in willow (Salix viminalis (L.))

<p>These vcf-files constitute underlying raw data material for the manuscript &quot;Genome wide associations of growth, phenology and plasticity traits in willow (Salix viminalis (L.))&quot;. For more detailed information please consult the README file in the repository.</p>

opencc-by-4.0Mar 2019View details →
zenodo48/100

Genomes plasmids MDR B. fragilis ONT sequence read files in fastq format

<p>Supporting data for the manuscript <em>Complete genome assembly of clinical multidrug resistant Bacteroides fragilis isolates enables comprehensive identification of antimicrobial resistance genes and plasmids.</em></p> <p>Oxford Nanopore reads demultiplexed with <a href="https://www.google.com/url?sa=t&amp;rct=j&amp;q=&amp;esrc=s&amp;source=web&amp;cd=1&amp;cad=rja&amp;uact=8&amp;ved=2ahUKEwjNjqL5tY7iAhUawMQBHZHfDasQFjAAegQIAhAB&amp;url=https%3A%2F%2Fgithub.com%2Frrwick%2FDeepbinner&amp;usg=AOvVaw0wikvIUagLuFV38CwKZtia">Deepbinner</a> v0.2.0 and base-called (with demultiplexing) using Albacore v2.3.3. Barcodes and adapters were removed with <a href="https://github.com/rrwick/Porechop">Porechop</a> v0.2.4 with the --discard_middle option.</p> <p>Data from each isolate was produced from two runs per isolate. Data for the individual runs are included here. They can easily be concatenated eg with cat. Runs are named TVS_01,. TVS_02, TVS_03 and TVS_04.</p> <p>Fast5 (only demultiplexed with deepbinner and basecalled with albacore) as well as illumina reads and genome assemblies can be found via the NCBI bioproject accessions:</p> <p>Isolates, NCBI bioproject accession no:</p> <p>CCUG4856T,&nbsp;<a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA525024">PRJNA525024</a></p> <p>BFO17,&nbsp;<a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA244943">PRJNA244943</a></p> <p>BFO18,&nbsp;<a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA244944">PRJNA244944</a></p> <p>S01,&nbsp;<a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA244942">PRJNA244942</a></p> <p>BFO42,&nbsp;<a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA253771">PRJNA253771</a></p> <p>BFO67,&nbsp;<a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA254401">PRJNA254401</a></p> <p>BFO85,&nbsp;<a href="http://www.ncbi.nlm.nih.gov/bioproject/PRJNA254455">PRJNA254455</a></p> <p>&nbsp;</p> <p><strong>md5sum&#39;s (also found in the file md5.md5):</strong></p> <p>135d0570a1e49e25c8fde59f321cca68&nbsp; BFO17_TVS_03_99377.barcode02_trimmed.fastq.gz<br> 3c6ca800a1f735937c0cffccd263fb82&nbsp; BFO18_TVS_01_97673.barcode03_trimmed.fastq.gz<br> dc37823950f529d3a859b7a7c514af8c&nbsp; BFO18_TVS_03_99377.barcode03_trimmed.fastq.gz<br> abd707404f9ebbc38652e45ed378ca21&nbsp; BFO42_TVS_02.barcode10_trimmed.fastq.gz<br> 9761e9ab082e276624c5778dcbeaffd2&nbsp; BFO42_TVS_04.barcode10_trimmed.fastq.gz<br> ffc27009c0f7fead1af84ce045dbe5f3&nbsp; BFO67_TVS_02.barcode09_trimmed.fastq.gz<br> 06c363a7feeeb88f9d195b3769b37f2b&nbsp; BFO67_TVS_04.barcode09_trimmed.fastq.gz<br> 5553c95cc98f4b9d4cfb38c4f8f8f037&nbsp; BFO85_TVS_02.barcode08_trimmed.fastq.gz<br> b1a8013cba7a079cee6d3bdd6cd97ff2&nbsp; BFO85_TVS_04.barcode08_trimmed.fastq.gz<br> 8169225219a5fb20935d5f0304aa80c5&nbsp; CCUG4856T_TVS_01_97673.barcode01_trimmed.fastq.gz<br> a509b0ae912a798a91e677795198c1c6&nbsp; CCUG5846T_TVS_03_99377.barcode01_trimmed.fastq.gz<br> 8a9d2eb8b626a6e87ed267d31aa220e3&nbsp; S01_TVS_01_97673.barcode04_trimmed.fastq.gz<br> 69e239a17becc25a4f103dfc3bf5886a&nbsp; S01_TVS_03_99377.barcode04_trimmed.fastq.gz</p>

opencc-by-4.0May 2019View details →
zenodo48/100

Supplementary dataset to publication: Oxford nanopore technologies - a valuable tool to generate whole-genome sequencing data for in silico serotyping and the detection of genetic markers in Salmonella, Thomas et al 2023

<p>Bacteria of the genus&nbsp;<em>Salmonella</em>&nbsp;pose a major risk to livestock, the food economy, and public health.&nbsp;<em>Salmonella</em>&nbsp;infections are one of the leading causes of food poisoning. The identification of serovars of&nbsp;<em>Salmonella</em>&nbsp;achieved by their diverse surface antigens is essential to gain information on their epidemiological context. Traditionally, slide agglutination has been used for serotyping. In recent years, whole-genome sequencing (WGS) followed by&nbsp;<em>in silico</em>&nbsp;serotyping has been established as an alternative method for serotyping and the detection of genetic markers for&nbsp;<em>Salmonella</em>. Until now, WGS data generated with Illumina sequencing are used to validate&nbsp;<em>in silico</em>&nbsp;serotyping methods. Oxford Nanopore Technologies (ONT) opens the possibility to sequence ultra-long reads and has frequently been used for bacterial sequencing. In this study, ONT sequencing data of 28&nbsp;<em>Salmonella</em>&nbsp;strains of different serovars with epidemiological relevance in humans, food, and animals were taken to investigate the performance of the&nbsp;<em>in silico</em>&nbsp;serotyping tools SISTR and SeqSero2 compared to traditional slide agglutination tests. Moreover, the detection of genetic markers for resistance against antimicrobial agents, virulence, and plasmids was studied by comparing WGS data based on ONT with WGS data based on Illumina. Based on the ONT data from flow cell version R9.4.1,&nbsp;<em>in silico</em>&nbsp;serotyping achieved an accuracy of 96.4 and 92% for the tools SISTR and SeqSero2, respectively. Highly similar sets of genetic markers comparing both sequencing technologies were identified. Taking the ongoing improvement of basecalling and flow cells into account, ONT data can be used for&nbsp;<em>Salmonella in silico</em> serotyping and genetic marker detection.</p>

opencc-by-4.0Feb 2023View details →
zenodo48/100

Inferring whole-genome histories in large population datasets: inferred tree sequences for 1000 Genomes

<p>Tree sequences inferred for the 1000 Genomes phase 3&nbsp;autosomes using&nbsp;<a href="https://tsinfer.readthedocs.io/">tsinfer</a>&nbsp;version 0.1.4 and compressed using&nbsp;<a href="https://tszip.readthedocs.io/en/stable/">tszip</a>. Tree sequences can&nbsp; be decompressed as follows:</p> <pre><code class="language-bash">$ tsunzip 1kg_chr1.trees.tsz</code></pre> <p>Once decompressed, trees files can be loaded and processed using&nbsp;<a href="https://tskit.readthedocs.io">tskit</a>.&nbsp;</p> <pre><code class="language-python">import tskit ts = tskit.load("1kg_chr1.trees") # ts is an instance of tskit.TreeSequence print("Chromosome 1 contains {} trees".format(ts.num_trees))</code></pre> <p>Metadata associated with individuals and populations was derived from the original&nbsp;<a href="http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/working/20130606_sample_info/20130606_g1k.ped">source</a>&nbsp;and converted to JSON form. For example, to access individual metadata we can use:</p> <pre><code class="language-python">import tskit import json ts = tskit.load("1kg_chr1.trees") ind = ts.individual(0) metadata_dict = json.loads(ind.metadata)</code></pre> <p>The metadata_dict variable will now contain&nbsp;all the metadata for the individual with ID 0 as a dictionary. Metadata associated with populations can be found in a similar way. Population IDs are associated with individuals via their constituent nodes. For example,</p> <pre><code class="language-python">pop_metadata = [json.loads(pop.metadata) for pop in ts.populations()] ind_node = ts.node(ind.nodes[0]) ind_pop_metadata = pop_metadata[ind_node.population]</code></pre> <p>After this, the&nbsp;ind_pop_metadata variable will contain the population level metadata for individual ID 0.</p> <p>The full data pipeline used to generate these tree sequences and associated metadata is available on <a href="https://github.com/mcveanlab/treeseq-inference/tree/master/human-data">GitHub</a>.</p>

opencc-by-4.0May 2019View details →
zenodo48/100

Inferring whole-genome histories in large population datasets: inferred tree sequences for Simons Genome Diversity Project

<p>Tree sequences inferred for the SGDP autosomes using&nbsp;<a href="https://tsinfer.readthedocs.io/">tsinfer</a>&nbsp;version 0.1.4 and compressed using&nbsp;<a href="https://tszip.readthedocs.io/en/stable/">tszip</a>. Tree sequences can&nbsp; be decompressed as follows:</p> <pre><code class="language-bash">$ tsunzip sgdp_chr1.trees.tsz</code></pre> <p>Once decompressed, trees files can be loaded and processed using&nbsp;<a href="https://tskit.readthedocs.io">tskit</a>.&nbsp;</p> <pre><code class="language-python">import tskit ts = tskit.load("sgdp_chr1.trees") # ts is an instance of tskit.TreeSequence print("Chromosome 1 contains {} trees".format(ts.num_trees))</code></pre> <p>Metadata associated with individuals and populations was derived from the original&nbsp;<a href="https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/SGDP_metadata.279public.21signedLetter.samples.txt">source</a>&nbsp;and converted to JSON form. For example, to access individual metadata we can use:</p> <pre><code class="language-python">import tskit import json ts = tskit.load("sgdp_chr1.trees") ind = ts.individual(0) metadata_dict = json.loads(ind.metadata)</code></pre> <p>The metadata_dict variable will now contain&nbsp;all the metadata for the individual with ID 0 as a dictionary. Metadata associated with populations can be found in a similar way. Population IDs are associated with individuals via their constituent nodes. For example,</p> <pre><code class="language-python">pop_metadata = [json.loads(pop.metadata) for pop in ts.populations()] ind_node = ts.node(ind.nodes[0]) ind_pop_metadata = pop_metadata[ind_node.population]</code></pre> <p>After this, the&nbsp;ind_pop_metadata variable will contain the population level metadata for individual ID 0.</p> <p>The full data pipeline used to generate these tree sequences and associated metadata is available on <a href="https://github.com/mcveanlab/treeseq-inference/tree/master/human-data">GitHub</a>.</p>

opencc-by-4.0May 2019View details →
zenodo44/100

Genome sequence of the banana aphid, Pentalonia nigronervosa Coquerel (Hemiptera: Aphididae) and its symbionts

<p><strong><em>Pentalonia nigronervosa</em> v1 frozen release</strong></p> <p>Genome assembly:&nbsp;Pentalonia_nigronervosa.v1.scaffolds.fa.gz</p> <p>BRAKER2 gene models:&nbsp;Pentalonia_nigronervosa.v1.scaffolds.gff</p> <p>BRAKER2 protein&nbsp;sequences:&nbsp;Pentalonia_nigronervosa.v1.scaffolds.gff.aa.fa</p> <p>BRAKER2 protein sequences (longest transcript per gene only):&nbsp;Pentalonia_nigronervosa.v1.scaffolds.gff.aa.LTPG.fa</p> <p>BRAKER2 coding&nbsp;sequences:&nbsp;Pentalonia_nigronervosa.v1.scaffolds.gff.cds.fa</p> <p>InterProScan functional annotation:&nbsp;Pentalonia_nigronervosa.v1.scaffolds.gff.aa.LTPG.interproscan.tsv</p> <p><em>Pentalonia nigronervosa</em>&nbsp;v1 mitochondrial genome:&nbsp;Pentalonia_nigronervosa.v1.mt_genome.fa</p> <p><em>Buchnera aphidicola</em> (BPn) scaffolds:&nbsp;Buchnera_aphidicola_BPn.scaffolds.fa</p> <p><em>Wolbachia</em> (WolPenNig) scaffolds:&nbsp;Wolbachia_WolPenNig.scaffolds.fa</p> <p><strong><em>Myzus cerasi </em>v1.2 frozen release</strong></p> <p>Genome assembly:&nbsp;Myzus_cerasi.v1.2.scaffolds.fa</p> <p>BRAKER2 gene models:&nbsp;Myzus_cerasi.v1.2.scaffolds.gff</p> <p>BRAKER2 protein&nbsp;sequences:&nbsp;Myzus_cerasi.v1.2.scaffolds.gff.aa.fa</p> <p>BRAKER2 protein sequences (longest transcript per gene only):&nbsp;Myzus_cerasi.v1.2.scaffolds.gff.aa.LTPG.fa</p> <p>BRAKER2 coding&nbsp;sequences:&nbsp;Myzus_cerasi.v1.2.scaffolds.gff.cds.fa</p> <p><strong>Aphid&nbsp;orthogroups and species tree</strong></p> <p>Proteomes included in the analysis:&nbsp;proteomes.tar.gz</p> <p>Orthogroups:&nbsp;Orthogroups.txt</p> <p>Gene counts per orthogroup, per species:&nbsp;Orthogroups.GeneCount.csv</p> <p>Single copy conserved orthogroups used for species tree: Orthogroups_for_concatenated_alignment.txt</p> <p>Species tree alignment:&nbsp;SpeciesTreeAlignment.fa</p> <p>Rooted species tree:&nbsp;SpeciesTree_rooted.nwk</p> <p><strong>Bash script to run k-mer based assembly deduplication pipeline</strong></p> <p>File:&nbsp;disco_filter_dups.v1.1.sh</p> <p>This script will parse a discovar de novo assembly and remove scaffolds likely to be haplotigs based on their k-mer content and a self alignment of the assembly (see manuscript for details).</p> <p>The input discovar assembly needs to have white space in scaffold IDs replaced with &quot;_&quot; before running. Illumina reads should be unzipped before running.</p> <p>Usage:</p> <pre><code class="language-bash">sh disco_filter_dups.sh &lt;./path_to_assembly&gt; &lt;./path_to_r1&gt; &lt;./path_to_r2&gt; &lt;homozyzgous_lower_cov&gt; &lt;homozyzgous_upper_cov&gt; &lt;nucmer_id_cutoff&gt; &lt;nucmer_cov_cutoff&gt; &lt;assembly_output_prefix&gt; &lt;threads&gt; &lt;./working_dir&gt;</code></pre> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2020View details →
zenodo44/100

COG-UK Viral Genome Sequences

<p>COG-UK Consortium has published dataset contains over 10K SARS-CoV-2 viral genome sequences available as open access.&nbsp;The current COVID-19 pandemic, caused by the SARS-CoV-2 virus, represents a major threat to health in the UK and globally. To fully understand the transmission and evolution of the virus requires sequencing and analysing viral genomes at scale and speed. The numbers of samples calls for a rapid increase in the UK&rsquo;s pathogen genome sequencing capacity rapidly and robustly. To provide this increased capacity to collect, sequence and analyse the whole genomes of virus samples in the UK, the COVID-19 Genomics UK (COG-UK) consortium is pooling the world-leading knowledge and expertise in genomics of the four UK Public Health Agencies, multiple regional University hubs, and large sequencing centres such as the Wellcome Sanger Institute.</p> <ul> <li>Protocols:&nbsp;https://www.cogconsortium.uk/protocols/</li> </ul>

opencc-by-4.0Apr 2020View details →
zenodo44/100

Virus+ Sequence Masked Mouse Reference Genome (GRCm38)

<p>A version of the mouse genome (<a href="https://www.ncbi.nlm.nih.gov/assembly/327618">GRCm38</a>)&nbsp;masked for all possible viral sequences.</p> <p>See&nbsp;<a href="https://zenodo.org/record/4116107#.X5B7ti9h3UI">Virus+ Masked Human Genome</a> for a masked human reference database.</p> <p>The following commands were used to generate the additional virus sequence masked reference database:</p> <p><strong>1) Download all RefSeq and Neighbor nucleotide records:</strong></p> <p><a href="https://www.ncbi.nlm.nih.gov/nuccore/?term=Viruses[Organism]%20NOT%20cellular%20organisms[ORGN]%20NOT%20wgs[PROP]%20NOT%20gbdiv%20syn[prop]%20AND%20(srcdb_refseq[PROP]%20OR%20nuccore%20genome%20samespecies[Filter])">https://www.ncbi.nlm.nih.gov/nuccore/?term=Viruses[Organism]%20NOT%20cellular%20organisms[ORGN]%20NOT%20wgs[PROP]%20NOT%20gbdiv%20syn[prop]%20AND%20(srcdb_refseq[PROP]%20OR%20nuccore%20genome%20samespecies[Filter])</a></p> <p><strong>2) Shred the downloaded viral genomes using shred.sh from the&nbsp;<a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a>&nbsp;package</strong></p> <p>shred.sh in=refseq_virus_reformated.fasta out=virus_shred.fasta.gz length=85 minlength=75 overlap=30</p> <p><strong>3) Map shredded virus sequence to the GRCm38</strong><strong> genome using bbmap.sh&nbsp;from the&nbsp;<a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a>&nbsp;package</strong></p> <p>bbmap.sh ref=GRCm38.fa.gz in=virus_shred.fasta.gz outm=map_mouse_all_viruses.sam minid=0.90</p> <p><strong>4) Mask virus sequenced mapped regions from the&nbsp;GRCm38 genome using bbmask.sh from the&nbsp;<a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a>&nbsp;package</strong></p> <p>bbmask.sh in=GRCm38.fa.gz out=GRCm38_virus_masked.fasta.gz sam=map_mouse_all_viruses.sam</p> <p><strong>5) Remove all N&#39;s to further reduce file size using&nbsp;<a href="https://bioinf.shenwei.me/seqkit/">seqkit</a></strong><br> seqkit -is replace -p &quot;n&quot; -r &quot;&quot; GRCm38_virus_masked.fasta.gz &nbsp;&gt;&nbsp;mouse_virus_masked.fasta_Ns_removed.gz</p> <p><strong>Additional References:</strong></p> <ol> <li><a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a></li> <li><a href="https://bioinf.shenwei.me/seqkit/">seqkit</a></li> <li><a href="https://www.ncbi.nlm.nih.gov/genome/viruses/">NCBI Virus Genome RefSeq</a></li> </ol>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Virus+ Sequence Masked Human Reference Genome (hg19)

<p>A version of the human genome (hg19) originally masked for ribosomal, plant, animal, fungal and&nbsp;low-entropy sequences&nbsp;by Brian Bushnell (<a href="https://zenodo.org/record/1208052#.X5BuTy9h3UI">Bushnell Masked Human Genome</a>) additionally masked for all possible viral sequences.</p> <p>The following commands were used to generate the additional virus sequence masked reference database:</p> <p><strong>1) Download all RefSeq and Neighbor nucleotide records:</strong></p> <p><a href="https://www.ncbi.nlm.nih.gov/nuccore/?term=Viruses[Organism]%20NOT%20cellular%20organisms[ORGN]%20NOT%20wgs[PROP]%20NOT%20gbdiv%20syn[prop]%20AND%20(srcdb_refseq[PROP]%20OR%20nuccore%20genome%20samespecies[Filter])">https://www.ncbi.nlm.nih.gov/nuccore/?term=Viruses[Organism]%20NOT%20cellular%20organisms[ORGN]%20NOT%20wgs[PROP]%20NOT%20gbdiv%20syn[prop]%20AND%20(srcdb_refseq[PROP]%20OR%20nuccore%20genome%20samespecies[Filter])</a></p> <p><strong>2) Shred the downloaded viral genomes using shred.sh from the <a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a> package</strong></p> <p>shred.sh in=refseq_virus_reformated.fasta out=virus_shred.fasta.gz length=85 minlength=75 overlap=30</p> <p><strong>3) Map shredded virus sequence to the hg19-masked human genome using bbmap.sh&nbsp;from the&nbsp;<a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a> package</strong></p> <p>bbmap.sh ref=hg19_main_mask_ribo_animal_allplant_allfungus.fa.gz in=virus_shred.fasta.gz outm=map_human_all_viruses.sam minid=0.90</p> <p><strong>4) Mask virus sequenced mapped regions from the hg19-masked human genome using bbmask.sh from the&nbsp;<a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a> package</strong></p> <p>bbmask.sh in=hg19_main_mask_ribo_animal_allplant_allfungus.fa.gz out=human_virus_masked.fasta.gz sam=map_human_all_viruses<br> .sam</p> <p><strong>5) Remove all N&#39;s to further reduce file size using <a href="https://bioinf.shenwei.me/seqkit/">seqkit</a></strong><br> seqkit -is replace -p &quot;n&quot; -r &quot;&quot; human_virus_masked.fasta.gz &nbsp;&gt; human_virus_masked.fasta_Ns_removed.gz</p> <p><strong>Additional References:</strong></p> <ol> <li><a href="http://seqanswers.com/forums/showthread.php?t=42552">http://seqanswers.com/forums/showthread.php?t=42552</a> for additional information on the original masking of hg19</li> <li><a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a></li> <li><a href="https://bioinf.shenwei.me/seqkit/">seqkit</a></li> <li><a href="https://www.ncbi.nlm.nih.gov/genome/viruses/">NCBI Virus Genome RefSeq</a></li> </ol>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Genomes and full-length 16S reference sequences for 27 Alpha- and Gamma-Proteobacterial isolates from Red Sea Acropora corals

<p>Coral-associated bacteria contribute to the biology of their host, but the underlying molecular interactions are largely unknown.&nbsp;To further our functional understanding, we obtained 27&nbsp;alpha- and gamma-proteobacterial&nbsp;isolates, many of which are Rhodobacteraceae,&nbsp;from three coral species of the genus&nbsp;<em>Acropora </em>and assembled/annotated their genomes as a resource for further functional studies.&nbsp;Our results reveal the immense taxonomic and genetic diversity of common&nbsp;alpha- and gamma-proteobacterial&nbsp;coral-associated bacteria. We hope these data provide&nbsp;a framework to study the function of specific bacteria in the coral holobiont. Isolates are available upon request.</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Whole genome sequencing of Turkish genomes reveals functional private alleles and impact of genetic interactions with Europe, Asia and Africa.

<p>BACKGROUND:</p> <p>Turkey is a crossroads of major population movements throughout history and has been a hotspot of cultural interactions. Several studies have investigated the complex population history of Turkey through a limited set of genetic markers. However, to date, there have been no studies to assess the genetic variation at the whole genome level using whole genome sequencing. Here, we present whole genome sequences of 16 Turkish individuals resequenced at high coverage (32&times;-48&times;).</p> <p>RESULTS:</p> <p>We show that the genetic variation of the contemporary Turkish population clusters with South European populations, as expected, but also shows signatures of relatively recent contribution from ancestral East Asian populations. In addition, we document a significant enrichment of non-synonymous private alleles, consistent with recent observations in European populations. A number of variants associated with skin color and total cholesterol levels show frequency differentiation between the Turkish populations and European populations. Furthermore, we have analyzed the 17q21.31 inversion polymorphism region (MAPT locus) and found increased allele frequency of 31.25% for H1/H2 inversion polymorphism when compared to European populations that show about 25% of allele frequency.</p> <p>CONCLUSION:</p> <p>This study provides the first map of common genetic variation from 16 western Asian individuals and thus helps fill an important geographical gap in analyzing natural human variation and human migration. Our data will help develop population-specific experimental designs for studies investigating disease associations and demographic history in Turkey.</p>

opencc-zeroOct 2015View details →
zenodo44/100

Imputation panel for low-pass whole genome sequencing (GLIMPSE2 format)

<p>This dataset includes autosomal genotypes from the 1000 Genomes +HGDP project (<a href="https://doi.org/10.1101/2023.01.23.525248" target="_blank" rel="noopener">10.1101/2023.01.23.525248 </a>)&nbsp; as well as X chromosome genotypes from the NY Genome Center (as of yet, a comparable dataset that includes HGDP is not available for the X; see 10.1016/j.cell.2022.08.004). The genotypes were down-sampled so as to be appropriate for low-pass imputation; uncertain phase calls were removed (any PP tags), and individuals deemed to be outliers or relatives (based on autosomal data, as per the first citation) were also removed. Similarly, singleton polymorphisms were also excluded. Hemizygous genotypes on the X were converted into (quasi) diploid genotypes.</p> <p>These data were then converted into a binary imputation panel format using glimpse v2 (https://odelaneau.github.io/GLIMPSE/; using the static binaries provided). The "chunk" size was doubled from the defaults (which considers a minimum number of snps, genetic length and physical length) so as to be more performant.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

GWAS Summary Statistics for Publication: Identifying novel genetic and phenotypic associations to genomic features by leveraging off-target reads in exome sequencing data

<p>This dataset contains summary statistics for genome-wide association studies (GWAS) conducted on genomic features derived from off-target reads in whole-exome sequencing (WES) data. The study utilized tools like Seeing Beyond the Target (SBT) and ImReP to construct novel phenotypic features from unmapped reads in ~50,000 participants in the UK Biobank. Features include mitochondrial DNA (mtDNA) copy number, ribosomal DNA (rDNA) copy number (5S, 18S, 28S), immune repertoire metrics (e.g., T-cell receptor alpha diversity), and microvial genome load (viral and fungal).</p> <p>Summary statistics can be used for replication studies, meta-analyses, or further exploration of these phenotypes.</p>

opencc-by-4.0Nov 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record