Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

22,157

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

22,157 results for “Genomics”

Learn how ShareScore rates datasets ↗
zenodo48/100

Introduction to Ancient Metagenomics Textbook (Edition 2025): Genome Mapping

<p>Data and conda software environment file for the chapter &#39;Genome Mapping&#39; of the SPAAM Community&#39;s textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Public sequence accessions from INSDC, COG-UK and CNCB and EPI_SET from GISAID for SARS-CoV-2 genome sequences in 2023-08-01 UShER tree

<p>Genome sequences and metadata for the accessions in the .tsv.gz (gzip-compressed tab-separated text) files are freely available from their corresponding sources:</p><ul><li>insdc.accessionNameDate.tsv.gz: INSDC (GenBank, ENA, DDBJ) sequences and metadata may be downloaded using NCBI Datasets: https://www.ncbi.nlm.nih.gov/datasets/taxonomy/2697049/ (7,361,734 accessions used on 2023-08-01)</li><li>cog.accessionNameDate.tsv.gz: COG-UK sequences and metadata may be downloaded from https://cog-uk.s3.climb.ac.uk/phylogenetics/latest (as of publication); most COG-UK sequences have been submitted to ENA and are available from INSDC/NCBI Datasets as well. &nbsp;(724,978 accessions used on 2023-08-01)</li><li>cncb.accessionNameDate.tsv.gz: Sequences and metadata from several databases at the China National Center for Bioinformation (CNCB) may be downloaded from GenBase: https://ngdc.cncb.ac.cn/genbase/ (26,604 accessions used on 2023-08-01)</li></ul><p>GISAID data are subject to restrictions on sharing described in https://gisaid.org/terms-of-use/. &nbsp;Genome sequences and metadata are available to registered GISAID users as part of EPI_SET_231106ax at https://doi.org/10.55876/gis8.231106ax (7,718,061 accessions used on 2023-08-01).</p>

opencc-by-sa-4.0Nov 2023View details →
zenodo48/100

Genomic Epidemiology Dataset for Important Nosocomial Pathogenic Bacteria Acinetobacter baumannii

<p>The<strong>&nbsp;</strong>infections caused by various bacterial pathogens both in clinical and community settings represent a significant threat to public healthcare worldwide. The growing resistance to antimicrobial drugs acquired by bacterial species causing healthcare-associated infections has already become a life-threatening danger noticed by the World Health Organization. Several groups or lineages of bacterial isolates usually called 'the clones of high risk' often drive the spread of resistance within particular species.&nbsp;</p><p>Thus, it is vitally important to reveal and track the spread of such clones and the mechanisms by which they acquire antibiotic resistance and enhance their survival skills. Currently, the analysis of whole genome sequences for bacterial isolates of interest is increasingly used for these purposes, including epidemiological surveillance and developing of spread prevention measures. However, the availability and uniformity of the data derived from the genomic sequences often represents a bottleneck for such investigations.&nbsp;</p><p>In this dataset, we present the results of a comprehensive genomic epidemiology analysis of 17,546 genomes of a dangerous bacterial pathogen <i>Acinetobacter baumannii</i>. Important typing information including multilocus sequence typing (MLST)-based sequence types (STs), intrinsic<i> blaOXA-51-like</i> gene variants, capsular (KL) and oligosaccharide (OCL) types, CRISPR-Cas systems, and cgMLST profiles are presented, as well as the assignment of particular isolates to nine known international clones of high risk. The presence of antimicrobial resistance genes within the genomes is also reported.&nbsp;</p><p>These data will be useful for researchers in the field of <i>A. baumannii</i> genomic epidemiology, resistance analysis and prevention measure development.</p>

opencc-by-sa-4.0Nov 2023View details →
zenodo48/100

Grapegenomics.com: a web portal with genomic data and analysis tools for wild and cultivated grapevines

<p><a href="https://grapegenomics.com">Grapegenomics.com</a> is a web portal that provides public access to genome references for grapevine cultivars (<em>Vitis vinifera</em> ssp. <em>vinifera</em>), wild grapevines (<em>Vitis vinifera</em> ssp. <em>sylvestris</em>), various wild grape species (<em>Vitis</em> spp. and <em>Muscadinia</em> spp.), and major fungal pathogens affecting grapes.</p> <p>All genomes are accessible through dedicated genome browsers, and published genomes are available for complete <a href="https://www.grapegenomics.com/download.php">download</a>.</p> <p>The site hosts all genomes produced by the laboratory of Dario Cant&ugrave; in the Department of Viticulture and Enology at the University of California, Davis, along with published genome references generated by others, such as PN40024 and Pinot noir ENTAV115. Instructions for genome submission are provided <a href="https://www.grapegenomics.com/submit.php">here</a>. The portal is maintained by No&eacute; Cochetel (ndcochetel[at]ucdavis.edu). In this version 2.0, all genome browsers utilize <a href="https://jbrowse.org/jb2/">jbrowse 2</a>.&nbsp;<br><br>Link to the website: <a href="https://www.grapegenomics.com">https://www.grapegenomics.com</a>&nbsp;</p>

opencc-by-4.0Jan 2024View details →
zenodo48/100

GWAS Summary Statistics from "Sex and statin-related genetic associations at the PCSK9 gene locus – results of genome-wide association meta-analysis"

<p>GWAMA summary statistics of PCSK9 levels stratified by sex and statin useage in Europeans.</p> <p>When using this data, please cite:</p> <p>Pott, J., Kheirkhah, A., Gadin, J.R.&nbsp;<em>et al.</em> Sex and statin-related genetic associations at the <em>PCSK9</em> gene locus: results of genome-wide association meta-analysis. <em>Biol Sex Differ</em> <strong>15</strong>, 26 (2024). https://doi.org/10.1186/s13293-024-00602-6</p> <p>All txt files contain the following columns:</p> <ul> <li>markername (unique SNP ID)</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>EA (effect allele)</li> <li>OA (other allele)</li> <li>EAF (effect allele frequency)</li> <li>info (minimal info score across all used studies)</li> <li>nSamples (sample size per SNP)</li> <li>nStudies (in case of double-stratified data: number of studies; in case of single-stratified data: 2, as it is a meta-analysis of the two double-stratified data sets)</li> <li>beta (effect estimate)</li> <li>SE (standard error)</li> <li>pval (p-value)</li> <li>I2 (SNP heterogeneity across studies)</li> <li>invalidAssoc (TRUE/FALSE flag if this variant was excluded in our analysis)</li> <li>reason4exclusion (reason why this SNP was excluded)</li> <li>phenotype (phenotyp setting)</li> </ul>

opencc-by-4.0Jan 2024View details →
zenodo48/100

The evolution of genomic, transcriptomic, and single-cell protein markers of metastatic upper tract urothelial carcinoma

<p>The molecular characteristics of metastatic upper tract urothelial carcinoma (UTUC) are unknown. The genomic and transcriptomic differences between primary and metastatic UTUC is not well described either. We combined whole-exome sequencing, RNA-sequencing, and Imaging Mass Cytometry<sup>TM</sup>&nbsp;(IMC<sup>TM</sup>) of 44 tumor samples from 28 patients with high-grade primary and metastatic UTUC. IMC enables spatially resolved single-cell analyses to examine the evolution of cancer cell, immune cell, and stromal cell markers using mass cytometry with lanthanide metal-conjugated antibodies. We discovered that actionable genomic alterations are frequently discordant between primary and metastatic UTUC tumors in the same patient. In contrast, molecular subtype membership and immune depletion signature were stable across primary and matched metastatic UTUC. Molecular and immune subtypes were consistent between bulk RNA-sequencing and mass cytometry of protein markers from 340,798 single-cells. Molecular subtyping at the single cell level was highly conserved between primary and metastatic UTUC tumors within the same patient.</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

Genome data from Effrenium voratum CCMP421, RCC1521, and rt-383 and their analysis

<p>This dataset represents secondary data generated from the genomic analysis of&nbsp;three isolates of <em>Effrenium voratum</em> (CCMP421, RCC1521, and rt-383), the early-diverging, free-living lineage of Symbiodiniaceae dinoflagellates. Theese data include <strong>(A)</strong> assembled genome sequences, predicted gene models and protein sequences, and&nbsp;<strong>(B)</strong> data and scripts associated with generation of graphs and figures presented in the key genome paper (Shah et al., 2024, Massive genome reduction predates the divergence of Symbiodiniaceae dinoflagellates, under review in <em>ISME Journal</em>). An earlier preprint of this manuscript is available at <em>bioRxiv</em>: <a href="https://doi.org/10.1101/2023.03.24.534093" target="_blank" rel="noopener">https://doi.org/10.1101/2023.03.24.534093</a>.</p> <p><strong>A. Genome assemblies, annotation and gene models.&nbsp;</strong>The dataset includes, for each taxon, (a) the&nbsp;<em>de novo</em> assembled genome sequences in FASTA format (<strong>*genome.fa.tgz</strong>), (b) structural annotation&nbsp;of the&nbsp;assembled genome&nbsp;in GFF3 format (<strong>*.genome.annotation.gff3.tgz</strong>), (c) the predicted protein-coding sequences of gene models in FASTA format (<strong>*genemodel.CDS.fa.tgz</strong>), (d) the predicted protein sequences of gene models in&nbsp;FASTA format (<strong>*genemodel.PROT.fa.tgz</strong>), and (c) the associated sequences and gene annotations of organellar genomic sequences (i.e. mitochondrial and plastid) (<strong>*organellar.tgz</strong>).&nbsp;Functional annotations of all gene models from the three genomes are available in the Excel spreadsheet (<strong>*GeneModels.xlsx</strong>).</p> <p><strong>B. Data and scripts associated with generation of graphs and figures in Shah et al. (2024).&nbsp;</strong>These files are organised based on key analyses specific to main figures and supplementary figures in the paper.</p> <p>See <strong>README.txt</strong> for a more-detailed description&nbsp;of the files.</p>

opencc-by-4.0Mar 2024View details →
zenodo48/100

Genome, repeat, and functional annotation associated with the naked mole-rat genome assembly, mHetGlaV3 (GCA_964261345.1)

<p>The naked mole-rat (NMR; Heterocephalus glaber) is a eusocial subterranean rodent with a highly unusual set of physiological traits, such as extreme longevity, that has attracted great interest amongst the scientific community. However, the genetic basis of most of these traits has not been elucidated. To facilitate our understanding of the molecular mechanisms underlying NMR physiology and behaviour, we generated a long-read chromosomal-level genome assembly of the NMR. This genome, mHetGlaV2, was subsequently annotated and incorporated into a &ldquo;91 eutherian mammals&rdquo; multiple whole genome alignment in Ensembl.&nbsp;</p> <p>We identified intra-chromosomal misassemblies within mHetGlaV2. We fixed these misassemblies by comparing syntenic blocks between this assembly and the Canadian Porcupine (EreDor) genome assembly (https://www.ncbi.nlm.nih.gov/datasets/genome/GCA_028451465.1/) and a FISH-Karyotype of the naked mole-rat completed by Romanenko et al., 2023 (PMID: 380307020) to address any misassemblies and place centromeres. Chromosome numbering was identified from a composite karyogram of karyotypes from over 350 cells.&nbsp;This scaffold-corrected assembly is labelled mHetGlaV3 (https://www.ebi.ac.uk/ena/browser/view/GCA_964261345.1).</p> <p>This repository stores the repeat, genome, and epigenome annotations for HetGlaV3.</p> <p>mHetGlaV3.primary.gtf.gz. Gene structures and gene symbols are transferred from ENSEMBL annotations of mHetGlaV2 using liftOff with default parameters. Additional gene symbols were identified using TOGA and manual curation.</p> <p>mHetGlaV3.primary.gtf.gz. Simple repetitive regions and transposable elements were annotated using EarlGrey (https://github.com/TobyBaril/EarlGrey) using "Rodentia" annotations for RepeatMasker.</p> <p>mHetGlaV3.primary.genesymbol_table.txt.txt.gz. A tab-delimited file where rows are gene IDs and columns are gene symbols generated with each method. "Consensus" shows the best matching gene symbol for each gene ID.</p> <p>mHetGlaV3.primary_annotated_blacklist.bed.gz. Provides an assembly "blacklist" for mHetGlaV3. This blacklist is a bed file annotating assembly breakpoints between HetGlaV2 and HetGlaV3. This blacklist contains additional columns (e.g., closest gene, overlapping TE etc.) and should therefore be filtered to the first column before being incorporated into traditional genomic pipelines.</p> <p>mHetGlaV3.primary_hypothalamus_ABC_enhancer.bedpe.gz. Activity-By-Contact enhancers (https://github.com/broadinstitute/ABC-Enhancer-Gene-Prediction) generated in the female subordinate naked mole-rat hypothalamus using Hi-C-seq, ChIP-seq of H3K27Ac data, ATAC-seq, and RNA-seq information.</p> <p>mHetGlaV3.primary_hypothalamus_chromHMM.bed.gz. Chromatin states (using Chromhmm) annotating the female subordinate naked mole-rat hypothalamus using H3K4me3 (promoter), H4K4me2 (promoter-enhancer), H3K27Ac (active enhancer), H3K36me3 (elongated), H3K27me3 (polycomb repressed), H3K9me3 (heterochromatin), and CTCF (whole brain) ChIP-seq data, as well as ATAC-seq and RNA-seq data.</p> <p>mHetGlaV3.primary.fa.gz. Genome assembly fasta file for the naked mole-rat (V3, primary assembly). This assembly matches the primary assembly stored on ENA, however the chromosome names match these files, rather than have chromosome names processed by ENA (e.g. chr 1 instead of "OZ179169.1 Heterocephalus glaber genome assembly, chromosome: 1").</p> <p>&nbsp;</p> <p>UPDATES:</p> <p>* The 1.2 update fixed unscaffolded contig names from those used in-lab to those compatible with ENA.</p> <p>* The 1.3 update added small (50~100kbp) contigs onto mHetGlaV3.primary.fa.gz that were filtered before the ENA submission.</p> <p>* The 1.4 update fixed a small chromosome naming inconsistency spotted in the 1.3 update.</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Joint Trajectory Inference for Single-cell Genomics Using Deep Learning with a Mixture Prior

<p>The datasets used in the paper "Joint Trajectory Inference for Single-cell Genomics Using Deep Learning with a Mixture Prior". A detailed description of these datasets is available at https://github.com/jaydu1/VITAE/tree/master/data.</p>

opencc-by-4.0Dec 2020View details →
zenodo48/100

Timema genome sequences and annotations. Version 8.

<p>Genome sequence (fasta) files&nbsp;and annotation (gff) files for ten <em>Timema </em>species:&nbsp;<em>T. bartmani, T. cristinae, T. poppensis, T. californicum,&nbsp; T. podura, T. tahoe, T. monikensis, T. douglasi, T. shepardi, and&nbsp; T. genevievae.</em><br> <br> Species are abbreviated as follows: Tbi =&nbsp;<em>T. bartmani</em>, Tce =&nbsp;<em>T. cristinae</em>, Tps =&nbsp;<em>T. poppensis</em>, Tcm =&nbsp;<em>T. californicum</em>, Tpa =&nbsp;<em>T. podura</em>, Tte =&nbsp;<em>T. tahoe</em>, Tms =&nbsp;<em>T. monikensis</em>, Tdi =&nbsp;<em>T. douglasi</em>, Tsi =&nbsp;<em>T. shepardi</em>, and Tge =&nbsp;<em>T. genevievae</em><br> &nbsp;</p> <p>For details of assembly and annotation see:&nbsp;<br> <br> Jaron, K. S*., Parker, D. J*., Anselmetti, Y., Tran Van, P. T., Bast, J., Dumas, &nbsp;Z., Figuet, E., Fran&ccedil;ois, C. M., Hayward, K., Rossier, V., Simion, P., Robinson-Rechavi, &nbsp;M., Galtier, N., Schwander, T. 2021. Convergent consequences of parthenogenesis on stick insect genomes. bioRxiv. doi: https://doi.org/10.1101/2020.11.20.391540</p> <p>&nbsp;</p> <p><strong>File list:</strong><br> <br> Tbi_b3v08.fasta = T. bartmani genome sequence file<br> Tbi_b3v08.max_arth_b2g_droso_b2g.gff = T. bartmani genome annotation file<br> Tce_b3v08.fasta = T. cristinae genome sequence file<br> Tce_b3v08.max_arth_b2g_droso_b2g.gff = T. cristinae genome annotation file<br> Tcm_b3v08.fasta&nbsp;&nbsp; &nbsp; = T. bartmani genome sequence file<br> Tcm_b3v08.max_arth_b2g_droso_b2g.gff = T. californicum genome annotation file<br> Tdi_b3v08.fasta = T. douglasi genome sequence file<br> Tdi_b3v08.max_arth_b2g_droso_b2g.gff = T. douglasi genome annotation file<br> Tge_b3v08.fasta = T. genevievae genome sequence file<br> Tge_b3v08.max_arth_b2g_droso_b2g.gff = T. genevievae genome annotation file<br> Tms_b3v08.fasta = T. monikensis genome sequence file<br> Tms_b3v08.max_arth_b2g_droso_b2g.gff = T. monikensis genome annotation file<br> Tpa_b3v08.fasta = T. podura genome sequence file<br> Tpa_b3v08.max_arth_b2g_droso_b2g.gff = T. podura genome annotation file<br> Tps_b3v08.fasta = T. poppensis genome sequence file<br> Tps_b3v08.max_arth_b2g_droso_b2g.gff = T. poppensis genome annotation file<br> Tsi_b3v08.fasta = T. shepardi genome sequence file<br> Tsi_b3v08.max_arth_b2g_droso_b2g.gff = T. shepardi genome annotation file<br> Tte_b3v08.fasta&nbsp;&nbsp; &nbsp; = T. tahoe genome sequence file<br> Tte_b3v08.max_arth_b2g_droso_b2g.gff = T. tahoe genome annotation file</p>

opencc-by-4.0Oct 2021View details →
zenodo48/100

FASTA file containing to the MYB encoding gene Ant1 genomic sequences corresponding to wild and cultivated tomato accessions

<p>Fasta sequence correspond to the MYB encoding gene&nbsp;<em>An2-like</em>. The genomic&nbsp;sequences correspond to&nbsp;<em>Solanum&nbsp;galagpagnese</em> accession LA1141 (this study), <em>S.&nbsp;lycopersicum</em> variety OH8245 (this study), <em>S. lycopersicum</em> variety Heinz 1706 reference genome, and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014).&nbsp;Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018), were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at <a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>).</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

FASTA file containing the MYB encoding gene An2-like genomic sequences corresponding to wild and cultivated tomato accessions

<p>FASTA sequence corresponds&nbsp;to the MYB encoding gene&nbsp;<em>An2-like</em>. The genomic&nbsp;sequences correspond to&nbsp;<em>Solanum&nbsp;galagpagnese</em> accession LA1141 (this study), <em>S.&nbsp;lycopersicum</em> variety OH8245 (this study), <em>S. lycopersicum</em> variety Heinz 1706 reference genome (Hosmani et al., 2019),&nbsp;<em>S. lycopersicum </em>variety Indigo Rose (Yan et al., 2020), <em>S. lycopersicum</em> accession LA1996 [MN242011.1&nbsp;(Colanero et al., 2020)], <em>S. chilense&nbsp;</em>accession LA1930 [MN242012.1 (Colanero et al., 2020)], and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014).&nbsp;Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018), &nbsp;Indigo Rose [MN433087 (Yan et al., 2020)], <em>S. lycopersicum </em>accession LA1996 [MN242011.1, EF433417.1 (Sapir et al., 2008; Colanero et al., 2020)], <em>S. chilense</em> accession LA1930 [MN242012.1 (Colanero et al., 2020)] were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at <a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>)&nbsp;and&nbsp;the National Center for Biotechnology Information (NCBI)(available at NCBI: <a href="https://www.ncbi.nlm.nih.gov">https://www.ncbi.nlm.nih.gov</a>).</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

FASTA file containing the MYB encoding genes at the Aft locus with genomic sequences corresponding to wild and cultivated tomato accessions

<p>FASTA sequences correspond to the MYB encoding genes&nbsp;<em>An2-like </em>and <em>Ant1</em>. The genomic&nbsp;sequences were combined correspond to&nbsp;<em>Solanum&nbsp;galagpagnese</em>&nbsp;accession LA1141 (this study),&nbsp;<em>S.&nbsp;lycopersicum</em>&nbsp;variety OH8245 (this study),&nbsp;<em>S. lycopersicum</em>&nbsp;variety Heinz 1706 reference genome (Hosmani et al., 2019),&nbsp;LA1996 [MN242011.1, EF433417.1(Sapir et al., 2008; Colanero et al., 2020)],&nbsp;and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014).&nbsp;Local sequences databases were made and retrieved using BLAST version/2018-08 for 84 accessions from The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences corresponding to Heinz 1706 (Hosmani et al., 2018),&nbsp;<em>S. lycopersicum&nbsp;</em>accession LA1996 [MN242011.1, EF433417.1 (Sapir et al., 2008; Colanero et al., 2020)],&nbsp;<em>S. chilense</em>&nbsp;accession LA1930 [MN242012.1 (Colanero et al., 2020)] were accessed using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at&nbsp;<a href="https://solgenomics.net/tools/blast/">https://solgenomics.net/tools/blast/</a>)&nbsp;and&nbsp;the National Center for Biotechnology Information (NCBI) (available at NCBI:&nbsp;<a href="https://www.ncbi.nlm.nih.gov/">https://www.ncbi.nlm.nih.gov</a>).</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

Clonal decomposition and DNA replication states defined by scaled single cell genome sequencing

<p><strong>OV2295&nbsp;Tables</strong></p> <p>ov2295_breakpoint_counts.csv.gz: Table of breakpoint counts per cell</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>cell_id: identifier for the cell</li> <li>read_count: number of reads</li> <li>library_id: identifier for the DNA library</li> <li>sample_id: identifier for the sequenced sample</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> </ul> <p>ov2295_cell_cn.csv.gz: Table of cell specific copy number</p> <ul> <li>cell_id: identifier for the cell</li> <li>sample_id: identifier for the sequenced sample</li> <li>library_id: identifier for the DNA library</li> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>reads: number of reads</li> <li>copy: raw normalized copy number</li> <li>state: copy number state</li> <li>gc: percent gc of the bin</li> <li>map: average mappability of the bin</li> </ul> <p>ov2295_cell_metrics.csv.gz: Table of cell metrics</p> <ul> <li>cell_id: identifier of the cell</li> <li>unpaired_mapped_reads: number of unpaired mapped reads</li> <li>paired_mapped_reads: number of mapped reads that were properly paired</li> <li>unpaired_duplicate_reads: number of unpaired duplicated reads</li> <li>paired_duplicate_reads: number of paired reads that were also marked as duplicate</li> <li>unmapped_reads: number of unmapped reads</li> <li>percent_duplicate_reads: percentage of duplicate reads</li> <li>estimated_library_size: scaled total number of mapped reads</li> <li>total_reads: total number of reads, regardless of mapping status</li> <li>total_mapped_reads: total number of mapped reads</li> <li>total_duplicate_reads: number of duplicate reads</li> <li>total_properly_paired: number of properly paired reads</li> <li>coverage_breadth: percentage of genome covered by some read</li> <li>coverage_depth: average reads per nucleotide position in the genome</li> <li>median_insert_size: median insert size between paired reads</li> <li>mean_insert_size: mean insert size between paired reads</li> <li>standard_deviation_insert_size: standard deviation of the insert size between paired reads</li> <li>index_sequence: index sequence of the adaptor sequence</li> <li>column: column of the cell on the nanowell chip</li> <li>img_col: column of the cell from the perspective of the microscope</li> <li>index_i5: id of the i5 index adapter sequence</li> <li>sample_type: type of the sample</li> <li>primer_i7: id of the i5 index primer sequence</li> <li>experimental_condition: experimental treatment of the cell, includes controls</li> <li>index_i7: id of the i7 index adapter sequence</li> <li>cell_call: living/dead classification of the cell based on staining usually, C1 == living, C2 == dead</li> <li>sample_id: name of the sample</li> <li>primer_i5: id of the i5 index primer sequence</li> <li>row: row of the cell on the nanowell chip</li> <li>library_id: identifier for the DNA library</li> <li>index: ignored</li> <li>multiplier: during parameter searching, the set [1..6] that was chosen</li> <li>MSRSI_non_integerness: median of segment residuals from segment integer copy number states</li> <li>MBRSI_dispersion_non_integerness: median of bin residuals from segment integer copy number states</li> <li>MBRSM_dispersion: median of bin residuals from segment median copy number values</li> <li>autocorrelation_hmmcopy: hmmcopy copy autocorrelation</li> <li>cv_hmmcopy: ignored</li> <li>empty_bins_hmmcopy: number of empty bins in hmmcopy</li> <li>mad_hmmcopy: median absolute deviation of hmmcopy copy</li> <li>mean_hmmcopy_reads_per_bin: mean reads per hmmcopy bin</li> <li>median_hmmcopy_reads_per_bin: median reads per hmmcopy bin</li> <li>std_hmmcopy_reads_per_bin: standard deviation value of reads in hmmcopy bins</li> <li>total_halfiness: summed halfiness penality score of the cell</li> <li>total_mapped_reads_hmmcopy: total mapped reads in all hmmcopy bins</li> <li>scaled_halfiness: summed scaled halfiness penalty score of the cell</li> <li>mean_state_mads: mean value for all median absolute deviation scores for each state</li> <li>mean_state_vars: variance value for all median absolute deviation scores for each state</li> <li>mad_neutral_state: median absolute deviation score of the neutral 2 copy state</li> <li>breakpoints: number of breakpoints, as indicated by state changes not at the ends of chromosomes</li> <li>mean_copy: mean hmmcopy copy value</li> <li>state_mode: the most commonly occuring state</li> <li>log_likelihood: hmmcopy log likelihood for the cell</li> <li>true_multiplier: the exact decimal value used to scale the copy number for segmentation</li> <li>order: order of the cell in the hierarchical clustering tree</li> <li>quality: random forest classifier proability score that cell is good</li> </ul> <p>ov2295_clone_alleles.csv.gz: Table of clone specific allele data</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>hap_label: haplotype block identifier</li> <li>clone_id: clone identifier</li> <li>allele_1_sum: number of reads for allele 1 of the haplotype block</li> <li>allele_2_sum: number of reads for allele 2 of the haplotype block</li> <li>total_counts_sum: total reads for the haplotype block</li> </ul> <p>ov2295_clone_breakpoints.csv.gz:&nbsp;Table of breakpoints per clone for OV2295 samples. Columns:</p> <ul> <li>prediction_id: identifier for the breakpoint</li> <li>chromosome_1: chromosome of breakend 1</li> <li>strand_1: orientation of break end 1</li> <li>position_1: position of break end 1</li> <li>chromosome_2: chromosome of breakend 2</li> <li>strand_2: orientation of break end 2</li> <li>position_2: position of break end 2</li> <li>clone_id: clone identifier</li> <li>read_count: number of reads</li> <li>is_present: presence=1, absent=0</li> </ul> <p>ov2295_clone_clusters.csv.gz: Table of cell clusters as putative clones</p> <ul> <li>cell_id: identifier for the cell</li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_cn.csv.gz:&nbsp;Table of allele specific copy number per clone for OV2295 samples. Columns:</p> <ul> <li>chr: chromosome of bin</li> <li>start: start of bin</li> <li>end: end of bin</li> <li>total_cn: HMMCopy predicted total copy number&nbsp;</li> <li>minor_cn: HMM predicted minor copy number&nbsp;</li> <li>major_cn: HMM predicted major copy number&nbsp;</li> <li>clone_id: clone identifier</li> </ul> <p>ov2295_clone_snvs.csv.gz:&nbsp;Table of SNVs per clone for OV2295 samples.&nbsp; Columns:</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>clone_id: clone identifier</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>total_counts: total number of reads at this position</li> <li>is_present: presence=0, absent=1</li> <li>is_het:&nbsp;is heterozygous</li> <li>is_hom: is homozygous for the alternate</li> </ul> <p>ov2295_nodes.csv.gz: Table of phylogenetic information for SNV evolution</p> <ul> <li>variant_id: identifier for the SNV as chrom:coord:ref:alt</li> <li>node: node in the phylogenetic tree</li> <li>loss: probability the SNV was lost at this node</li> <li>origin: probability the SNV originated at this node</li> <li>presence: probability the SNV is present at this node</li> <li>ml_origin: binary indicator the SNV originated at this node</li> <li>ml_presence: binary indicator the SNV is present at this node</li> <li>ml_loss: binary indicator the SNV was lost at this node</li> </ul> <p>ov2295_snv_counts.csv.gz: Table of SNV counts</p> <ul> <li>chrom: chromosome</li> <li>coord: genome position</li> <li>ref: reference nucleotide</li> <li>alt: alternate nucleotide</li> <li>ref_counts: number of reads at this position matching the reference nucleotide</li> <li>alt_counts: number of reads at this position matching the alternate nucleotide</li> <li>cell_id: identifier for the cell</li> <li>total_counts: total number of reads at this position</li> <li>sample_id: identifier for the sequenced sample</li> </ul> <p>ov2295_tree.pickle: Phylogenetic tree in python pickle format.&nbsp; Requires installation of the stochastic dollo code at: https://bitbucket.org/dranew/dollo, version 0.4.2.</p> <p>Note the following sample mapping: &lsquo;SA922&rsquo;: &lsquo;OV2295(R2)&rsquo;, &lsquo;SA921&rsquo;: &lsquo;TOV2295(R)&rsquo;, &lsquo;SA1090&rsquo;: &lsquo;OV2295&rsquo;,</p> <p><strong>Plots</strong></p> <p>ov_supp_clone_allele_cn.png: Clone allele ratios for each OV2295 sample.</p> <p>ov_supp_clone_total_cn.png: Clone copy number for each OV2295 sample.</p> <p>ov_supp_sample_total_cn.png: Bulk copy number for each OV2295 sample.</p> <p>ov_supp_sample_allele_cn.png: Bulk allele ratios for each OV2295 sample.</p>

opencc-by-4.0Sep 2019View details →
zenodo48/100

Genomic evidence for the parallel regression of melatonin synthesis and signaling pathways in placental mammals

<p><strong>Supplementary Material for:</strong></p> <p>Emerling C.A., Springer M.S., Gatesy J., Jones Z., Hamilton D., Xia-Zhu D., Collin M.A.,&nbsp;and Delsuc F. (2021).&nbsp;Genomic evidence for the parallel regression of melatonin synthesis and signaling pathways in placental mammals.<strong><em> Open Research Europe</em></strong> 1:75. doi:10.12688/openreseurope.13795.1.</p> <p>&nbsp;</p> <p><strong>Supplementary File Legends:</strong></p> <p><strong>- Supplementary_Figure_S1.pdf:</strong>&nbsp;<em>AANAT</em> PAML &lsquo;master model&rsquo; showing branch categories, corresponding to &ldquo;Model 1: 24 ratio&rdquo; in Supplementary Table S7.</p> <p><strong>- Supplementary_Figure_S2.pdf:</strong>&nbsp;<em>ASMT</em> PAML &lsquo;master model&rsquo; showing branch categories, corresponding to &ldquo;Model 2: 24 ratio&rdquo; in Supplementary Table S8.</p> <p><strong>- Supplementary_Figure_S3.pdf:</strong>&nbsp;<em>MTNR1A</em> PAML &lsquo;master model&rsquo; showing branch categories, corresponding to &ldquo;Model 1: 27 ratio&rdquo; in Supplementary Table S9.</p> <p><strong>- Supplementary_Figure_S4.pdf:</strong>&nbsp;<em>MTNR1B</em> PAML &lsquo;master model&rsquo; showing branch categories, corresponding to &ldquo;Model 1: 46 ratio&rdquo; in Supplementary Table S10.</p> <p><strong>- Supplementary_Figure_S5.pdf:</strong>&nbsp;RAxML <em>AANAT</em> gene tree. Numbers at nodes correspond to bootstrap support values.</p> <p><strong>- Supplementary_Figure_S6.pdf:&nbsp;</strong>RAxML <em>ASMT</em> gene tree. Numbers at nodes correspond to bootstrap support values.</p> <p><strong>- Supplementary_Figure_S7.pdf:&nbsp;</strong>RAxML <em>MTNR1A</em>+<em>MTNR1B</em>&nbsp;tree. Numbers at nodes correspond to bootstrap support values.</p> <p><strong>- Supplementary_Figure_S8.pdf:&nbsp;</strong>Supporting data showing the inactivation of <em>MTNR1A</em> exon 2 in cetaceans. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S9.pdf:&nbsp;</strong>Supporting data showing the inactivation of <em>ASMT</em> in spalacids and <em>Fukomys damarensis</em>. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S10.pdf:&nbsp;</strong>Supporting data showing the inactivation of <em>MTNR1A</em> in hyracoids and <em>Cyclopes didactylus</em>. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S11.pdf:&nbsp;</strong>Supporting data showing the inactivation of <em>MTNR1A</em> in sirenians. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S12.pdf:&nbsp;</strong>Supporting data showing the inactivation of <em>AANAT</em> in sirenians and a polymorphic premature stop codon in exon 5 of <em>ASMT</em> in <em>Trichechus manatus</em>. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S13.pdf:&nbsp;</strong>Supporting data showing the inactivation of <em>MTNR1A</em> in <em>Condylura cristata</em>. Read Supplementary Table S13 for further details.</p> <p><strong>- Supplementary_Figure_S14.pdf:&nbsp;</strong>Supporting data showing the inactivation of <em>MTNR1A</em> in <em>Phataginus tricuspis</em>. Read Supplementary Table S14 for further details.</p> <p><strong>- Supplementary_Figure_S15.pdf:&nbsp;</strong>PAML <em>AANAT</em> results, Model 1: 24 ratio (see Supplementary Table S7).</p> <p><strong>- Supplementary_Figure_S16.pdf:&nbsp;</strong>PAML <em>ASMT</em> results, Model 2: 24 ratio (see Supplementary Table S8).</p> <p><strong>- Supplementary_Figure_S17.pdf:&nbsp;</strong>PAML <em>MTNR1A</em> results, Model 1: 27 ratio (see Supplementary Table S9).</p> <p><strong>- Supplementary_Figure_S18.pdf:&nbsp;</strong>PAML <em>MTNR1B</em> results, Model 1: 46 ratio (see Supplementary Table S10).</p> <p><strong>- Supplementary_Table_S1.xlsx:&nbsp;</strong>List of species examined in this study and the sources of the genes. Source key: WGS: Sequences derived from NCBI&#39;s Whole Genome Shotgun database, with accession prefix provided; Whole Genome Sequencing of Short Reads: whole genomes were sequenced using short-read technologies. The methodologies&nbsp;varied for the species, and will be or have been published with other projects, so please contact the author(s) for information on the specific methodology and samples used (Xenarthrans, <em>Proteles cristatus</em>, <em>Otocyon megalotis</em>: Fr&eacute;d&eacute;ric Delsuc, e-mail: Frederic.Delsuc@umontpellier.fr; Crocodylians: John Gatesy, e-mail: jgatesy@amnh.org; <em>Dugong dugon</em>: Mark Springer, e-mail: mark.springer@ucr.edu; SRA: sequences derived from NCBI&#39;s Sequence Read Archive; GenBank: sequences derived from NCBI&#39;s nucleotide collection; Bowhead Whale Genome Resource: sequences derived from http://www.bowhead-whale.org; Ensembl: sequences derived from Ensembl genome browser (www.ensembl.org)l; Discovar de novo: sequences derived genomes assembled via Discovar de novo&nbsp; (<a href="https://software.broadinstitute.org/software/discovar/blog/">https://software.broadinstitute.org/software/discovar/blog/</a>). Coverage: indicates coverage of the whole genome (reported in NCBI or other source) or individual genes (derived from short read mapping). Scaffold and contig N50: reported in NCBI or other source.</p> <p><strong>- Supplementary_Table_S2.xlsx:&nbsp;</strong>Accession numbers and functionality of <em>AANAT</em> in species examined. If Accession # indicated as &ldquo;New&rdquo;, sequence generated for this study and can be found in Supplementary Dataset S1. Parentheses after accession number indicates coordinates for sequence on the contig / scaffold. Exon colors code for the following: green = putatively functional; yellow = missing (e.g., negative BLAST results, negative mapping results); pink = one or more inactivating mutations found. Abbreviations for mutations are as follows: del = deletion; ins = insertion; start = start codon mutation; stop = premature stop codon; ? = ambiguity whether the mutation is shared among all members of the clade. Abbreviations in brackets following an inactivating mutation indicate shared inactivating mutation. Key for each abbreviation follows: Bacu =&nbsp;<em>Balaenoptera acutorostrata</em>; BALA = Balaenidae; BALAEN = Balaenopteridae; Bbon =&nbsp;<em>Balaenoptera bonaerensis</em>; CAB =&nbsp;<em>Cabassous</em>; Ccap =&nbsp;<em>Cebus capucinus</em>; CETA = Cetacea; CHLAM = Chlamyphoridae; CHOL =&nbsp;<em>Choloepus</em>; Cjac =&nbsp;<em>Callithrix jacchus</em>; CING = Cingulata; DASY = Dasypodidae; DELP = Delphinidae; DERM = Dermoptera; Erob =&nbsp;<em>Eschrichtius robustus</em>; INIA =&nbsp;<em>Inia</em>; FOLI = Folivora; GALE =&nbsp;<em>Galeopterus</em>; LIPO =&nbsp;<em>Lipotes</em>; Lobl =&nbsp;<em>Lagenorhynchus obliquidens</em>; MANI = Manidae; MONO = Monodontidae; MYRM = Myrmecophagidae; MYST = Mysticeti; NPP = Not present in&nbsp;<em>Platanista</em>&nbsp;or Physeteroidea, but present in other Odontocetes; NPZ = Not present in Ziphiidae, but present in other Odontocetes; Oorc =&nbsp;<em>Orcinus orca</em>; PEUT = Tolypeutinae; PHOC = Phocoenidae; PHOL = Pholidota; PHOR = Chlamyphorinae; PILO = Pilosa; PHYS = Physeteroidea; PONT =&nbsp;<em>Pontoporia</em>; Schi =&nbsp;<em>Sousa chinensis</em>; SIRE = Sirenia; Tadu =&nbsp;<em>Tursiops aduncus</em>; TOLY =&nbsp;<em>Tolypeutes</em>; VERM = Vermilingua; XEN = Xenarthra.</p> <p><br> <strong>- Supplementary_Table_S3.xlsx:&nbsp;</strong>Accession numbers and functionality of <em>ASMT</em> in species examined. See Table S2 caption for details.</p> <p><strong>- Supplementary_Table_S4.xlsx:&nbsp;</strong>Accession numbers and functionality of <em>MTNR1A</em> in species examined. See Table S2 caption for details.</p> <p><strong>- Supplementary_Table_S5.xlsx:&nbsp;</strong>Accession numbers and functionality of <em>MTNR1B</em> in species examined. See Table S2 caption for details.</p> <p><strong>- Supplementary_Table_S6.xlsx:&nbsp;</strong>Codon frequency model selection. These are the results from one ratio dN/dS analyses using different codon frequency models.&nbsp;AIC = Akaike Information Criterion.</p> <p><strong>- Supplementary_Table_S7.xlsx:&nbsp;</strong>Results of <em>AANAT</em> PAML dN/dS analyses for mammals. Model: BG = branch(es) grouped with background; fixed 1 = branch(es) fixed at 1. p&rsquo;-value: p-value after Holm-Bonferroni correction for multiple testing. Model Comparison: if model comparison yields statistically significant differences (p &lt; 0.05), model comparison bolded and given green background; if model comparison is still significant after Holm-Bonferroni correction, asterisk (*) added. For most models, w only shown for branch(es) of interest. Numbers in front of taxonomic names in first row correspond to numbers in the master model shown in Supplementary Figure S1.</p> <p><strong>- Supplementary_Table_S8.xlsx:&nbsp;</strong>Results of <em>ASMT</em> PAML dN/dS analyses for mammals. Refer to Table S7 caption for additional details. Numbers in front of taxonomic names in first row correspond to numbers in the master model shown in Supplementary Figure S2.</p> <p><strong>- Supplementary_Table_S9.xlsx:&nbsp;</strong>Results of <em>MTNR1A</em> PAML dN/dS analyses for mammals. Refer to Table S7 caption for additional details. Numbers in front of taxonomic names in first row correspond to numbers in the master model shown in Supplementary Figure S3.</p> <p><strong>- Supplementary_Table_S10.xlsx:&nbsp;</strong>Results of <em>MTNR1B</em> PAML dN/dS analyses for mammals. Refer to Table S7 caption for additional details. Numbers in front of taxonomic names in first row correspond to numbers in the master model shown in Supplementary Figure S4.</p> <p><strong>- Supplementary_Table_S11.xlsx:&nbsp;</strong>Results of PAML analyses for sauropsids.</p> <p><strong>- Supplementary_Table_S12.xlsx:&nbsp;</strong>Results of BLASTing and mapping short reads from&nbsp;<em>Alligator mississippiensis</em>&nbsp;RNA sequencing experiments.</p> <p><strong>- Supplementary_Table_S13.xlsx:&nbsp;</strong>Supporting data for validating putative inactivating mutations. Validating data came from four general sources of information: mutations shared by more than one species within a clade, mutations shared by two sources of sequencing data for the same species, mutations validated by coverage of mapped short reads and statistically elevated dN/dS ratio estimates. For additional details, see Supplementary Tables S2&ndash;S5 and S7&ndash;S10, as well as Figure 2 and Supplementary Figures S8&ndash;S18.</p> <p><strong>- Supplementary_Dataset_S1.txt:</strong><strong>&nbsp;</strong>Genomic alignments in fasta format used to determine the pseudogene/functional&nbsp;status of all four melatonin genes in different taxonomic groups.</p> <p><strong>- Supplementary_Dataset_S2.txt:</strong><strong>&nbsp;</strong>Alignment of <em>AANAT</em>&nbsp;in phylip format used in maximum likelihood phylogenetic reconstruction with RAxML.&nbsp;</p> <p><strong>- Supplementary_Dataset_S3.txt:&nbsp;</strong>Alignment of <em>ASMT</em> in phylip format used in maximum likelihood phylogenetic reconstruction with RAxML.&nbsp;</p> <p><strong>- Supplementary_Dataset_S4.txt:&nbsp;</strong>Alignment of <em>MTNR1A</em> and <em>MTNR1B</em> in phylip format used in maximum likelihood phylogenetic reconstruction with RAxML.&nbsp;</p> <p><strong>- Supplementary_Dataset_S5.txt:</strong><strong>&nbsp;</strong>Codon&nbsp;alignments of <em>AANAT</em> used in selection pressure analyses&nbsp;with PAML.&nbsp;</p> <p><strong>- Supplementary_Dataset_S6.txt:&nbsp;</strong>Codon&nbsp;alignments of <em>ASMT</em> used in selection pressure analyses&nbsp;with PAML.</p> <p><strong>- Supplementary_Dataset_S7.txt:</strong><strong>&nbsp;</strong>Codon&nbsp;alignments of <em>MTNR1A</em> used in selection pressure analyses&nbsp;with PAML.</p> <p><strong>- Supplementary_Dataset_S8.txt: </strong>Codon&nbsp;alignments of <em>MTNR1B</em> used in selection pressure analyses&nbsp;with PAML.</p> <p><strong>- Supplementary_Dataset_S9.txt:&nbsp;</strong>Tree topologies in newick format used in selection pressure analyses&nbsp;with PAML.</p>

opencc-by-4.0Jun 2021View details →
zenodo48/100

Genome- and transcriptome-wide association summary statistics for outcome from traumatic brain injury

<p>The dataset contains summary statistics for the genome- and transcriptome-wide association studies (GWAS, TWAS) of genetic effects on outcome in traumatic brain injury (TBI). The study participants attended hospital within 24 hours of TBI, and underwent head computed tomography imaging.</p> <p><strong>Study participants</strong></p> <p>European ancestry data set contains 4710 individuals; multi-ethnic cohort 5268 individuals, including Europeans (n = 4710), Africans (n = 245) and Admixed Americans (n = 313).</p> <p>The largest European population contribution was from CENTER-TBI (Collaborative European NeuroTrauma Effectiveness Research, https://www.center-tbi.eu), where each participating center (60 centers from 20 countries in Europe) recruited patients between December 2013 and December 2017. The patients recruited in CENTER-TBI were supplemented by subjects from cohorts recruited at two European centres (Cambridge, UK, and Turku, Finland).</p> <p>The majority of patients in the US cohort were recruited between 2014 and 2018 to TRACK-TBI (Transforming Research and Clinical Knowledge in TBI, https://tracktbi.ucsf.edu) by the 18 US participant sites. The subjects recruited to the US cohort from TRACK-TBI were supplemented by patients recruited to an institutional research initiative at Mass General Brigham (MGB).</p> <p><strong>Outcome definition</strong></p> <p>Outcomes were measured using the extended Glasgow Outcome Scale (GOSE), ranging from 1 (dead) to 8 (upper good recovery), measured 6 months post-TBI. TBI severity was specified using the Glasgow Coma Score (GCS), with TBI classified as mild (GCS 13-15), moderate (GCS 9-12), or severe (GCS 3-8).</p> <p>To account for the effect of injury severity on outcome, sliding dichotomization was used to categorize outcome as favourable or unfavourable. A GOSE &le; 4 was used to define an unfavourable outcome for patients with either moderate (GCS 9-12) or severe (GCS 3-8) TBI, while the unfavourable group was extended to patients with GOSE &le; 7 if they had mild (GCS 13-15) TBI.</p> <p><strong>Genotype data and imputation</strong></p> <p>Genotyping was completed at FIMM Technology Center for CENTER-TBI, Cambridge, Turku patients and the Broad Institute for TRACK-TBI, using the Illumina Global Screening Array (GSA-24v2-0 + Multi-Disease). The MGB cohort were genotyped using Illumina&rsquo;s Multi-Ethnic Global array (MEGA) and the pre-releases forms, including MEGA and MEGA-Ex arrays at Illumina at the MGB Translational Genomics Core.</p> <p>A unified quality control procedure was applied for each study cohort and the array-based genotypes were imputed using the Haplotype Reference Consortium panel. Autosomal chromosomes were considered, post-imputation data was filtered by imputation quality (INFO &gt; 0.4 for CENTER-TBI, Cambridge and Turku;&nbsp;R2 &gt; 0.4 for TRACK-TBI and MGB) and MAF &gt; 1%.</p> <p><strong>Genome-wide association analysis and meta-analysis</strong></p> <p>Genome-wide single-marker scans were performed using a penalized likelihood-based Firth logistic regression, and implemented in PLINK v2.0. Using favourable outcome as reference, models were fitted on the basis of imputed allelic dosages. Age, sex, major extracranial injury, pupillary reactivity, and the first 10 principal components were included as covariates. Study cohort (CENTER-TBI, Cambridge, Turku) was an additional covariate in the CENTER-TBI GWAS.</p> <p>Fixed-effects meta-analysis of the three European ancestry GWAS was performed using METAL. For trans-ethnic meta-analysis, summary statistics of five GWASs in patients of European, African and Admixed Americans were aggregated via MR-MEGA.</p> <p><strong>Transcriptome-wide association study</strong></p> <p>Genetically regulated gene expression (GREx) was imputed using a regression model fitted on a separate gene expression database. Elastic net models provided by PrediXcan for all available GTEx brain tissues and whole blood were used. For TWAS, the same sliding dichotomy model for outcome with the same set of covariates as in the GWAS, but PCA components were replaced with the top five principal components of the respective gene expression data.&nbsp;</p> <p><strong>Column headers - GWAS</strong></p> <p>rsID: variant rsID<br> Chrom: chromosome<br> Pos: position (build GRCh38)<br> A1: effect allele<br> A2: reference allele<br> EAF: allele frequency of effect allele<br> Effect: effect size of effect allele<br> StdErr: standard error of effect size<br> P: p value of association (with genomic correction)<br> N: sample size</p> <p>Note. &#39;Effect&#39; and &#39;StdErr&#39; are only available for the European ancestry meta-analysis.</p> <p><br> <strong>Column headers - TWAS</strong></p> <p>tissue: GTEx tissue type<br> id: ensembl gene id<br> coef: model coefficient<br> se: model standard error for coefficient<br> p: model-based p value<br> symbol: gene symbol<br> name: gene name written out<br> chr: chromosome<br> start: gene start position (build GRCh38)</p>

opencc-by-4.0Jan 2022View details →
zenodo48/100

A cross-disorder dosage sensitivity map of the human genome

<p>This repository contains data from Collins et al., <em>A cross-disorder dosage sensitivity map of the human&nbsp;genome</em> (2022), including:</p> <p>1. <strong>Collins_rCNV_2022.dosage_sensitivity_scores.tsv.gz</strong>: This file contains predicted probabilities of haploinsufficiency (pHaplo) and triplosensitivity (pTriplo) for 18,641 autosomal protein-coding genes as defined in Gencode v19.</p> <p>2. <strong>Collins_rCNV_2022.sliding_window_sumstats.tar.gz</strong>: This compressed directory contains rCNV association summary statistics for 54 phenotypes from genome-wide sliding window meta-analyses. Please refer to the README file included in this compressed directory for more details.</p> <p>3. <strong>Collins_rCNV_2022.gene_association_sumstats.tar.gz</strong>: This compressed directory contains rCNV association summary statistics for 54 phenotypes from exome-wide gene-based meta-analyses. Please refer to the README file included in this compressed directory for more details.&nbsp;</p> <p>4. <strong>Collins_rCNV_2022.gene_features_matrix.tar.gz</strong>: This compressed directory contains gene-level feature annotations for 145 features and 18,641 autosomal protein-coding genes. Please refer to the README file included in this compressed directory for more details.&nbsp;</p> <p>Smaller data files have been provided as supplemental tables alongside the publication online.</p> <p>Please also refer to the original publication for details on data sources, study design, methods, and other analyses.</p>

opencc-by-4.0Mar 2022View details →
zenodo48/100

Genome-wide association study suggests that variation at the RCOR1 locus is associated with tinnitus in UK Biobank

<p>The dataset contains results of a genome-wide association studies for age-related hearing impairment (ARHI)-related traits as described in the following publication:<br> Wells, H.R.R., Abidin, F.N.Z., Freidin, M.B. et al. Genome-wide association study suggests that variation at the RCOR1 locus is associated with tinnitus in UK Biobank. Sci Rep 11, 6470 (2021). https://doi.org/10.1038/s41598-021-85871-6</p>

opencc-by-4.0Jun 2022View details →
zenodo48/100

Summary statistics accompanying the article "Genome-wide association study of the human brain functional connectome reveals strong vascular component underlying global network efficiency" in Scientific Reports (2022)

<p>Summary statistics for genome-wide association studies reported in:</p> <p>Bell, S., Tozer, D.J., &amp; Markus H.S. (2022). Genome-wide association study of the human brain functional connectome reveals strong vascular component underlying global network efficiency. <em>Scientific Reports</em>, DOI: <a href="https://dx.doi.org/10.1038/s41598-022-19106-7">10.1038/s41598-022-19106-7</a>.&nbsp;</p> <p><strong>Abstract</strong></p> <p>Complex brain networks play a central role in integrating activity across the human brain, and such networks can be identified in the absence of any external stimulus. We performed 10 genome-wide association studies of resting state network measures of intrinsic brain activity in up to 36,150 participants of European ancestry in the UK Biobank. We found that the heritability of global network efficiency was largely explained by blood oxygen level-dependent (BOLD) resting state fluctuation amplitudes (RSFA), which are thought to reflect the vascular component of the BOLD signal. RSFA itself had a significant genetic component and we identified 24 genomic loci associated with RSFA, 157 genes whose predicted expression correlated with it, and 3 proteins in the dorsolateral prefrontal cortex and 4 in plasma. We observed correlations with cardiovascular traits, and single-cell RNA specificity analyses revealed enrichment of vascular related cells. Our analyses also revealed a potential role of lipid transport, store-operated calcium channel activity, and inositol 1,4,5-trisphosphate binding in resting-state BOLD fluctuations. We conclude that that the heritability of global network efficiency is largely explained by the vascular component of the BOLD response as ascertained by RSFA, which itself has a significant genetic component.</p> <p>&nbsp;</p> <p>Further information on the files uploaded here can be found in the README. Users interested in bulk downloading these summary statistics may find <a href="https://github.com/dvolgyes/zenodo_get">zenodo_get</a> helpful.</p>

opencc-by-4.0Aug 2022View details →
zenodo48/100

deGeco genomic compartments model fit results on whole genome at resolution of 50kb

<p>These files contain the fitted parameters for the <a href="https://www.biorxiv.org/content/10.1101/2022.10.01.510432v1.article-info">deGeco</a> model for genomic compartments. Fits were done at 50kb on four cell lines: GM12878 (from Rao, et al., 2014),&nbsp; H1, HFF (both from&nbsp;Krietenstein, et al., 2020) and mESC (from Bonev, et al., 2017). Each Hi-C file was zoomified again using cooler, to prevent duplicate entries in the pixel table.</p> <p>The file format is NumPy&#39;s npz object that has two main keys:</p> <ol> <li>Metadata - an object containing various information on the run: command line parameters, duration of run, etc</li> <li>Parameters - an object containing the actual fitted parameters: <ol> <li>state_probabilities - an NxS matrix of state probabilities, where N is the number of bins and S the number of states the model was run with</li> <li>cis_weights - an SxS matrix of cis state affinities</li> <li>trans_weights - an SxS matrix of trans state affinities</li> <li>cis_dd_power - the exponent of the power law decay of interaction intensity in cis (also denoted as alpha)</li> <li>trans_dd - the constant background level of trans interaction (also denoted as beta)</li> <li>cis_lengths - Number of bins for each chromosome. Sum of cis_lengths is N, the total number of bins.</li> </ol> </li> </ol> <p>To read using numpy:</p> <pre><code class="language-python">import numpy as np fit = np.load(filename, allow_pickle=True) metadata = fit['metadata'][()] parameters = fit['parameters'][()]</code></pre> <p>or use the gc_datafile module from the deGeco <a href="https://github.com/KaplanLab/deGeco">repository</a>:</p> <pre><code class="language-python">import gc_datafile parameters = gc_datafile.load_params(filename)</code></pre> <p>&nbsp;</p>

opencc-by-4.0Oct 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record