Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
29,880
datasets available to search
ShareScore release 0.7.1
Dataset results
29,880 results for “gene expression”
MAGE: Multi-ancestry Analysis of Gene Expression
<p>MAGE comprises RNA-seq data from lymphoblastoid cell lines derived from 731 individuals from the <a href="https://doi.org/10.1038/nature15393" rel="nofollow">1000 Genomes Project (1KGP)</a>, representing 26 globally-distributed populations across five continental groups. These data offer a large, geographically diverse, open access resource to facilitate studies of the distribution, genetic underpinnings, and evolution of variation in human transcriptomes and include data from several ancestry groups that were poorly represented in previous studies.</p> <p>Briefly, this repo contains the following data:</p> <ol> <li>Sample metadata and sequencing metrics</li> <li>Gene expression and splicing matrices used for e/sQTL mapping and analyses of global trends of expression/splicing diversity</li> <li>cis-e/sQTL mapping results, including aFC estimates for cis-eQTLs</li> <li>Functional annotations of cis-e/sQTLs</li> <li>Results of colocalization analysis between MAGE e/sQTLs and complex trait GWAS from the <a href="https://doi.org/10.1038/s41586-019-1310-4" rel="nofollow">PAGE</a> study</li> <li>Results of analyses of global trends of expression/splicing diversity</li> <li>Jointly-generated top genotype PCs for samples in MAGE and other resources with paired WGS/RNA-seq data (Geuvadis, GTEx, AFGR)</li> </ol> <p>READMEs are provided for all data in the repo.</p>
COLONOMICS - predictive models for normal colon gene expression and DNA methylation for TWAS and MWAS
<p>We provide significant SNP prediction models derived from the COLONOMICS data (<a href="https://www.colonomics.org">https://www.colonomics.org</a>). Genotypes were obtained by Affymetrix 6.0 array, imputed to TopMed panel. Gene expression was obtained from Affymetrix U219 array, DNA methylation was obtained with Illuminan 450K array and miRNA expression was obtained by NGS. We provide SNP prediction models for 1,758 genes, 30,530 CpG probes and 38 miRNAs obtained from colon normal biopsy samples. These features can be predicted from SNPs located within ±1Mb, which we assumed they act through cis mechanisms. We include the model’s summary statistics and corresponding SNP weights in SQLite objects. Models were trained using the elastic net procedure employed in the PredictDB pipeline (<a href="https://predictdb.org/">https://predictdb.org</a>), according to which only models with a predictive performance p-value < 0.05 and R<sup>2</sup> > 0.1 are considered significant. We adjusted the models by basic covariates, i.e., sex, age, tissue type and colon anatomic location where biopsies were collected (left and right colon). Genome coordinates refer to GRCh37/hg19.</p>
Gene expression count data from human post-mortem spinal cord
<p>Gene expression data from human post-mortem tissue for three spinal cord sections (cervical, thoracic and lumbar) from amyotrophic lateral sclerosis (ALS) patients and non-neurological disease controls. RNA sequencing performed as part of the New York Genome Center ALS Consortium.</p> <p>Analysis workbooks: <a href="https://jackhump.github.io/ALS_SpinalCord_QTLs/">https://jackhump.github.io/ALS_SpinalCord_QTLs/</a> </p> <p>Preprint describing results: <a href="https://www.medrxiv.org/content/10.1101/2021.08.31.21262682v1">https://www.medrxiv.org/content/10.1101/2021.08.31.21262682v1</a> </p> <p><strong>Sample sizes:</strong></p> <table align="left"> <tbody> <tr> <td> <p><strong>Region</strong></p> </td> <td> <p><strong>Control</strong></p> </td> <td> <p><strong>ALS</strong></p> </td> </tr> <tr> <td> <p>Cervical</p> </td> <td> <p>35</p> </td> <td> <p>139</p> </td> </tr> <tr> <td> <p>Thoracic </p> </td> <td> <p>10</p> </td> <td> <p>42</p> </td> </tr> <tr> <td> <p>Lumbar</p> </td> <td> <p>32</p> </td> <td> <p>122</p> </td> </tr> </tbody> </table> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p><strong>Library preparation</strong></p> <p>RNA was extracted from flash-frozen postmortem tissue using TRIzol (Thermo Fisher Scientific) chloroform, followed by column purification (RNeasy Minikit, QIAGEN). RNA integrity number (RIN) was assessed on a Bioanalyzer (Agilent Technologies). RNA-Seq libraries were prepared from 500ng total RNA using the KAPA Stranded RNA-Seq Kit with RiboErase (KAPA Biosystems) for rRNA depletion and Illumina-compatible indexes (NEXTflex RNA-Seq Barcodes, NOVA-512915, PerkinElmer, and IDT for Illumina TruSeq UD Indexes, 20022370). Pooled libraries (average insert size: 375 bp) passing the quality criteria were sequenced either on an Illumina HiSeq 2500 (125 bp paired end) or an Illumina NovaSeq (100 bp paired-end). The samples had a median sequencing depth of 42 million read pairs, with a range between 16 and 167 million read pairs.</p> <p><strong>Data processing</strong></p> <p>Samples were uniformly processed using RAPiD-nf, an efficient RNA-Seq processing pipeline implemented in the NextFlow framework. Following adapter trimming with Trimmomatic (version 0.36), all samples were aligned to the hg38 build (GRCh38.primary_assembly) of the human reference genome using STAR (2.7.2a), with indexes created from GENCODE, version 30. Gene expression was quantified using RSEM (1.3.1) using GENCODE v30. Quality control was performed using SAMtools and Picard, and the results were collated using MultiQC. Various technical metrics for sequencing quality control are provided in the metadata. Estimated read counts and normalised transcripts per million (TPM) matrices provided for each tissue.</p> <p><strong>Provided data:</strong></p> <p><em>gencode.v30.gene_meta.tsv.gz</em> - tab separated table with columns "genename", the HGNC gene symbol, and "geneid" the Ensembl ID, as set in the GENCODE v30 comprehensive annotation.</p> <p>For {tissue} in Cervical_Spinal_Cord, Thoracic_Spinal_Cord, Lumbar_Spinal_Cord:</p> <p><em>{tissue}_metadata.tsv.gz </em>- metadata describing each sample. Each row describes a sample. Descriptions of each column below.</p> <p><em>{tissue}_gene_tpm.tsv.gz</em> - the normalised TPM values from RSEM for all 58,884 genes in GENCODE v30. Each row describes a gene and each column describes a sample.</p> <p><em>{tissue}_gene_counts.tsv.gz</em> - the estimated read counts from RSEM for all 58,884 genes in GENCODE v30. Each row describes a gene and each column describes a sample.</p> <p><strong>Metadata Column Description</strong></p> <p><em>rna_id</em> - de-identified sample ID for each unique RNA-seq sample</p> <p><em>dna_id</em> - de-identified donor ID for each patient enrolled in the study</p> <p><em>site_id</em> - de-identified site name for each contributing site</p> <p><em>tissue</em> - name of tissue/region</p> <p><em>age_rounded</em> - age at death, rounded to nearest decade</p> <p><em>sex</em> - biological sex of donor</p> <p><em>subject_group</em> - long form disease group</p> <p><em>disease</em> - short form disease group</p> <p><em>site_of_motor_onset </em>- for ALS donors, where did symptoms start?</p> <p><em>disease_duration</em> - for ALS donors, how long did donor live with disease? </p> <p><em>mutations </em>- any known ALS gene mutations</p> <p><em>library_prep</em> - type of library preparation method used</p> <p><em>seq_platform </em>- sequencing platform used for sequencing</p> <p><em>rin</em> - RNA integrity number, 0-10</p> <p><em>c9orf72_repeat_size</em> - estimated C9orf72 repeat expansion size</p> <p><em>gPC1 - gPC5 </em>- principal component of genetic ancestry from whole genome sequencing</p> <p>Remaining metadata columns are from Picard - see here: <a href="http://broadinstitute.github.io/picard/picard-metric-definitions.html#RnaSeqMetrics">http://broadinstitute.github.io/picard/picard-metric-definitions.html#RnaSeqMetrics</a> </p> <p> </p>
SEESAW quantification data for temporal gene expression across osteoblastogenesis (B6xCAST), n=9
<p>Osteoblast cells mature from a mesenchymal stem cell pool to become cells capable of forming bone matrix and mineralizing this matrix. The goal of this study was to characterize temporal changes in the transcriptome across osteoblast maturation, starting with committed mesenchymal stem cell/ early pre-osteoblast stage through to mature osteoblasts capable of matrix mineralization. Methods: Enriched populations of pre-osteoblast-like cells were obtained from neonatal calvaria from B6xCAST mice expressing CFP under the control of the Col3.6 promoter. These cells were placed into culture for 4 days, removed from culture and subjected to FACS sorting based on the presence/absence of CFP expression. Cells expressing CFP were returned to culture, subjected to an osteoblast differentiation cocktail and RNA was collected at 2, 4, 6, 8, 10, 12, 14, 16 and 18 days post differentiation. Methods II: mRNA profiles for each time point were generated by next generation RNA sequencing, using an Illumina HiSeq 2000. Three technical replicates per sample were sequenced. Overall design: Gene expression in calvarial osteoblasts from neonatal B6xCAST-Col3.6 CFP mice at 9 time points post differentiation.</p> <p>File description: the R data files (.rda) provide outputs of the scripts in the mikelove/osteoblast-quant GitHub repo (July 2022, commit 01d96490), having run the fishpond package function importAllelicCounts() followed by minimal filtering. The `_counts.rda` files contain SummarizedExperiment objects with estimated count, TPM abundance, and effective length, but do not contain inferential replicates (bootstrap counts), although the transcript-level allelic counts object contains bootstrap mean and variances for every isoform, sample, and allele. The other two `.rda` files are summarized to gene level.</p> <p>The `_quant_dirs.tgz` files contain all the Salmon quantification data including bootstraps for the 9 time points. They are grouped into sets of three for convenience. The `CAST_EiJ.diploid.fa.gz` file provides the transcript sequences that were used for Salmon quantification.</p> <p>The `B6xCAST_discordant_global_AI.csv` file contains the same information as presented in Table S1 of Wu et al (2022). These are TSS-level results for 134 genes showing significant and discordant patterns within gene.</p> <p>The other 6 CSV files provide global and dynamic AI testing results at three levels of resolution: gene level, isoform level (txp), and TSS level where TSS within 50bp are combined into a single TSS-group. The significance cutoff is a q-value of 0.05. The code used for generating these results is provided in the GitHub repo: FennecFish/osteoblast-test.</p>
Cytochrome P450 Genes Expressed in Phasmatodea Midguts
<p>In June 2021, representative sequences for insect cytochrome P450s were downloaded from NCBI, limiting the search to those in the UniProtKB database. The resulting 111 sequences were used as a query to mine the above transcriptomes using tblastn with an expect value threshold of e-10. These were manually annotated by removing truncated sequences, using the ExPASy online translation tool to obtain the complete amino acid sequences, removing duplicates using the sRNAtoolbox webserver, and confirming that the sequences were cytochrome P450s by identifying them using blastp against the NCBI database. The resulting sequences were combined with the representatives from NCBI, aligned using the Clustal W program built into the software MEGA version X. Any sequences missing the heme-binding domain FXXGXXXCXG/A, which is a signature motif for CYPs, were deleted. The sequences were also checked for the presence and absence of other four signature motifs from insect CYPs: helix C (WxxxR), helix I (GxE/DTT/S), helix K (ExLR), and PERF (PxxFxPE/DRE/F).</p>
Changes in gene expression during germination reveal pea genotypes with either 'quiescence' or 'escape' mechanisms of waterlogging tolerance
<p>Waterlogging causes germination failure in pea (<em>Pisum sativum</em> L.). Three genotypes (BM-3, NL-2 and Kaspa) contrasting in ability to germinate in waterlogged soil were exposed to different durations of waterlogging. Whole genome RNAseq was employed to capture differentially expressing genes. The ability to germinate in waterlogged soil was associated with testa colour and testa membrane integrity as confirmed by electrical conductivity measurements. Among the most differentially regulated genes, upregulated gene tyrosine protein kinase responsible for metabolic regulation and downregulated LOX5 involved in fat metabolism indicated energy preservation in tolerant Kaspa, while in the other tolerant NL-2 subtilase family protein and PNC2 involved in protein and fat metabolism respectively showed upregulated expression suggesting energy utilization during waterlogging. By contrast, in sensitive genotype BM-3 high upregulation was recorded for the kunitz-type trypsin/protease inhibitor whose role is blocking the activity of protein metabolism leading to excessive lipid metabolism causing membrane leakage and subsequent seed damage. Pathway analyses based on gene ontologies showed seed storage protein metabolism as upregulated in tolerant genotypes and downregulated in the sensitive genotype. Understanding the tolerance mechanism provides a platform to breed for adaptation to waterlogging stress at germination in pea. </p>
Changes in gene expression during germination reveal pea genotypes with either 'quiescence' or 'escape' mechanisms of waterlogging tolerance
<p>Waterlogging causes germination failure in pea (<em>Pisum sativum</em> L.). Three genotypes (BM-3, NL-2 and Kaspa) contrasting in ability to germinate in waterlogged soil were exposed to different durations of waterlogging. Whole genome RNAseq was employed to capture differentially expressing genes. The ability to germinate in waterlogged soil was associated with testa colour and testa membrane integrity as confirmed by electrical conductivity measurements. Among the most differentially regulated genes, upregulated gene tyrosine protein kinase responsible for metabolic regulation and downregulated LOX5 involved in fat metabolism indicated energy preservation in tolerant Kaspa, while in the other tolerant NL-2 subtilase family protein and PNC2 involved in protein and fat metabolism respectively showed upregulated expression suggesting energy utilization during waterlogging. By contrast, in sensitive genotype BM-3 high upregulation was recorded for the kunitz-type trypsin/protease inhibitor whose role is blocking the activity of protein metabolism leading to excessive lipid metabolism causing membrane leakage and subsequent seed damage. Pathway analyses based on gene ontologies showed seed storage protein metabolism as upregulated in tolerant genotypes and downregulated in the sensitive genotype. Understanding the tolerance mechanism provides a platform to breed for adaptation to waterlogging stress at germination in pea. </p>
Phaeocystis globosa colonial gene expression
<p>Data and analysis for the paper: </p> <p><strong>Differential gene expression supports a resource-intensive, defensive role for colony production in the bloom-forming haptophyte, <em>Phaeocystis globosa</em></strong></p> <p>by: Margaret Mars Brisbin and Satoshi Mitarai</p> <p>The <em>Phaeocystis globosa</em> CCMP1528 transcriptome used in the study (phaeocystisglobosa_euk_seqs.fasta or pg_euk_seqs_altnames.fasta) was assembled with trimmed sequencing reads from 8 biological replicates (4 colonial replicates and 4 solitary replicates) with the Trinity software (v2.3.2).</p> <p>Raw sequencing reads are available from the NCBI SRA with accession numbers: SRR7811979–SRR7811986.</p> <p>Before assembling the transcriptome, reads were quality filtered and trimmed with the Trimmomatic software (v3.36) using the command:</p> <pre><code>java -jar $TRIM/trimmomatic-0.36.jar PE -phred33 $DATA2/S${SLURM_ARRAY_TASK_ID}_S*_R1_001.fastq.gz \ $DATA2/S${SLURM_ARRAY_TASK_ID}_S*_R2_001.fastq.gz \ $OUT/S${SLURM_ARRAY_TASK_ID}_1_paired.fq $OUT/S${SLURM_ARRAY_TASK_ID}_1_unpaired.fq \ $OUT/S${SLURM_ARRAY_TASK_ID}_2_paired.fq $OUT/S${SLURM_ARRAY_TASK_ID}_2_unpaired.fq \ ILLUMINACLIP:$TRIM/adapters/NexteraPE-PE.fa:2:30:10 LEADING:3 TRAILING:3 SLIDINGWINDOW:4:15 MINLEN:36</code></pre> <p>Trimmed reads were mapped to the ERCC reference sequences for Mix1 and mapped reads were filtered using the following commands from bowtie2 (v2.2.6), samtools, and bedtools: </p> <pre><code>bowtie2 -t -x $REF \ -1 $DATA/S${SLURM_ARRAY_TASK_ID}_1_paired.fq \ -2 $DATA/S${SLURM_ARRAY_TASK_ID}_2_paired.fq \ -S $OUT/S${SLURM_ARRAY_TASK_ID}_ercc.sam samtools view -bS $DATA/S${SLURM_ARRAY_TASK_ID}_ercc.sam >$DATA/S${SLURM_ARRAY_TASK_ID}.bam samtools sort $DATA/S${SLURM_ARRAY_TASK_ID}.bam $DATA/S${SLURM_ARRAY_TASK_ID}_sorted samtools view -b -f 13 S${SLURM_ARRAY_TASK_ID}_sorted.bam > S${SLURM_ARRAY_TASK_ID}_unmapped.bam samtools sort -n $DATA/S${SLURM_ARRAY_TASK_ID}_unmapped.bam $DATA/S${SLURM_ARRAY_TASK_ID}.qsort bedtools bamtofastq -i $DATA/S${SLURM_ARRAY_TASK_ID}.qsort.bam -fq $DATA/S${SLURM_ARRAY_TASK_ID}_1_paired.fq -fq2 $DATA/S${SLURM_ARRAY_TASK_ID}_2_paired.fq</code></pre> <p>The resulting Trimmed reads without ERCC sequences were used to make the transcriptome assembly: </p> <pre><code>Trinity --seqType fq --max_memory 475G \ --left $DATA2/C1_1_paired.fq,$DATA2/C2_1_paired.fq,$DATA2/C3_1_paired.fq,$DATA2/C4_1_paired.fq,$DATA2/S1_1_paired.fq,$DATA2/S2_1_paired.fq,$DATA2/S3_1_paired.fq,$DATA2/S4_1_paired.fq \ --right $DATA2/C1_2_paired.fq,$DATA2/C2_1_paired.fq,$DATA2/C3_2_paired.fq,$DATA2/C4_2_paired.fq,$DATA2/S1_2_paired.fq,$DATA2/S2_2_paired.fq,$DATA2/S3_2_paired.fq,$DATA2/S4_2_paired.fq \ --CPU 12</code></pre> <p>The Trinity assembly was dereplicated with CD-HIT-EST (v2016-0304) at 95% : </p> <pre><code>cd-hit-est -i $DATA/Trinity.fasta -o Trinity_Pg_clustered_95 -c 0.95 -n 8 -p 1 -g 1 -M 200000 -T 8 -d 40</code></pre> <p>The Trinity assembly was filtered to remove bacterial contamination by first running a blastn(v2.6.0+) against the nr/nt NCBI database:</p> <pre><code>blastn -query $DATA/Trinity_Pg_clustered_95.fasta -task blastn -db $REF -num_threads 12 -max_target_seqs 1 -outfmt 5 > TrinityBlast.xml</code></pre> <p>and then removing bacterial reads with custom python scripts included here: TrinityBlastXML.ipynb and FIlterTrinityEukNotEuk.ipynb </p> <p>RSEM (v1.2.22) was run with the final transcriptome assembly (phaeocystisglobosa_euk_seqs.fasta or pg_euk_seqs_altnames.fasta): </p> <pre><code>rsem-calculate-expression --bowtie2 --paired-end \ $DATA/C${SLURM_ARRAY_TASK_ID}_1_paired.fq \ $DATA/C${SLURM_ARRAY_TASK_ID}_2_paired.fq \ $REF/rsemref_longISO/pg_euks_RSEMref \ $REF/rsemout_longISO/C${SLURM_ARRAY_TASK_ID} rsem-calculate-expression --bowtie2 --paired-end \ $DATA/S${SLURM_ARRAY_TASK_ID}_1_paired.fq \ $DATA/S${SLURM_ARRAY_TASK_ID}_2_paired.fq \ $REF/rsemref_longISO/pg_euks_RSEMref \ $REF/rsemout_longISO/S${SLURM_ARRAY_TASK_ID} </code></pre> <p>The resulting data files are: C*.genes.results and S*.genes.results which were used with DESeq2 in the R environment to analyze different gene expression. The code for these analyses is available in html and R markdown (PhaeoColSol_DE.html, PhaeoColSol_DE.Rmd). </p> <p>The transcriptome assembly was annotated with the Dammit software (v1.0rc2), which wraps Transdecoder, HMMER, and BUSCO, and by submitting the translated amino acid sequences to GhostKOALA. </p> <p>The raw pfam Dammit annotation results are included: pg_euk_seqs.fasta.x.pfam.gff3. These results were parsed with the script: Pfam_gffParsing.ipynb. The resulting file, pfam_parsed_annotation.csv, is used in the script PhaeoColSol_DE.Rmd with pfam2go4R.txt for GO enrichment analysis. The script shinycolsol.Rmd creates an interactive plot of GO enrichment results. </p> <p>The GhostKOALA results are user_ko.csv, and are used in the script PhaeoColSol_DE.Rmd for KEGG pathway enrichment analysis. </p>
Processed data for the study on "Chromatin 3D interactions mediate genetic effects on gene expression"
<p>This repository contains the processed data that was generated as part of the following study:</p> <p>Delaneau et al. (2019) <strong>Chromatin 3D interactions mediate genetic effects on gene expression.</strong></p> <p><em>Abstract:</em> Studying the genetic basis of gene expression and chromatin organization is key to characterize the effect of genetic variability on the function and structure of the human genome. Here, we unravel how genetic variation perturbs gene regulation using a dataset combining activity of regulatory elements, gene expression and genetic variants across 317 individuals and two cell types. We show that variability in regulatory activity is structured at the intra- and inter-chromosomal levels within 12,583 Cis Regulatory Domains and 30 Trans Regulatory Hubs that highly reflect the local (i.e. Topologically Associating Domains) and global (i.e. open/close chromatin compartments) nuclear chromatin organization. These structures delimit cell type specific regulatory networks that control gene expression/co-expression and mediate the genetic effects of <em>cis</em>- and <em>trans</em>-acting regulatory variants on genes.</p> <p> </p> <p>This repository contains:</p> <ol> <li>Chromatin QTLs for H3K27ac, H3K4me1 and H3K4me3 discovered in 317 Lymphoblastoids Cell Lines (LCLs) and 78 Fibroblasts.</li> <li>Molecular QTLs affecting the activity and structure of Cis Regulatory Domains (CRDs) in LCLs.</li> <li>Basic information about the full set of genetic variants being analyzed in the study.</li> <li>The peak coordinates, their hierarchy based on inter-individual correlation and the CRD calls for both LCLs and Fibroblasts.</li> <li>The functional links discovered in LCLs between CRDs and genes.</li> <li>eQTLs for LCLs.</li> <li>A README file containing the description of the file format for each file.</li> </ol>
GTEx gene-level expression summary for Snaptron
<p>GTEx gene-level expression summary from the Snaptron collection. Format is a tab-separated text file compressed and indexed using BGZip, along with supplementary files containing a Tabix index for the data (ending in tbi) as well as two files describing the samples represented in the columns of the data file (ending in tsv). Uses GENCODE v25 annotation for quantification. Source data for the quantification are the bigWig files produced as part of recount2. More information at http://snaptron.cs.jhu.edu.</p>
Defective HNF4alpha-dependent gene expression as a driver of hepatocellular failure in alcoholic hepatitis [Suppl Data]
<p>Alcoholic hepatitis (AH) is a life-threatening condition characterized by profound hepatocellular dysfunction for which targeted treatments are urgently needed. Identification of molecular drivers is hampered by the lack of suitable animal models. By performing RNA sequencing in livers from patients with different phenotypes of alcohol-related liver disease (ALD), we show that the development of AH is characterized by the defective activity of liver-enriched transcription factors (LETFs). TGFb1is a key upstream transcriptome regulator in AH and induces the use of HNF4aP2 promoter in hepatocytes, which results in defective metabolic and synthetic functions. Gene polymorphisms in LETFs including HNF4aare not associated with the development of AH. In contrast, epigenetic studies show that AH livers have profound changes in DNA methylation state and chromatin remodeling, affecting HNF4a-dependent gene expression. We conclude that targeting TGFb1and epigenetic drivers that modulate HNF4a-dependent gene expression could be beneficial to improve hepatocellular function in patients with AH.</p>
Multi-study reanalysis of 2213 acute myeloid leukemia patients reveals age- and sex-dependent gene expression signatures
<p>Supplemental Information, Supplementary Tables, and Supplementary Files, as well as accompanying data for the manuscript "Stratified computational meta-analysis of 2213 acute myeloid leukemia patients reveals age- and sex-dependent gene expression signatures" by Raeuf Roushangar and George I. Mias. </p>
Arena assay and gene expression of stingless bees
<p>The dataset includes behavioral data collected from arena experiments with workers of stingless bee species <em>Scaptotrigona postica</em>, <em>Frieseomelitta varia</em>, and <em>Melipona quadrifasciata</em>. Additionally, it contains gene expression data (CT values) for these species across different life stages and tissues.</p>
Data from: Polymorphic tandem repeats shape single-cell gene expression across the immune landscape
<p>This dataset contains the association summary statistics (v0.1) for genome-wide tandem repeat (TR) expression quantitative trait (eQTL) analysis of TenK10K Phase 1 (https://doi.org/10.1101/2024.11.02.621562). </p> <p>Please access the README for a detailed description of file contents. </p> <p> </p>
Tissue heterogeneity is prevalent in gene expression studies
<p>This archive contains results associated with the publication</p> <p><em>Tissue heterogeneity is prevalent in gene expression studies. Gregor Sturm, Markus List and Jitao David Zhang.</em></p> <p> </p> <ul> <li>expr.tissuemark.affy.roche.symbols.gmt: The tissue signatures from the BioQC publication used in this study</li> <li>gtex_v6_gini_solid.gmt: The cross-platform cross-species validated tissue signatures produced in this study</li> <li>heterogeneity_results.tsv.gz: Signature scores and heterogeneity calls for each tested signature</li> <li>heterogeneity_fractions.tsv: Fraction of heterogeneous and severely heterogeneous samples per tissue</li> </ul>
The spatial landscape of gene expression isoforms in tissue sections
<p><strong>This upload provides raw in situ sequencing (ISS) data used to validate Spatial Isoform Transcriptomics (SiT), as well as R scripts required for SiT analysis.</strong></p> <p><strong>GenePlots.zip and Reads.zip are ISS data </strong><strong>generated and collected by the CARTANA ISS service</strong>. <strong>The following data description is cited from the report provided by CARTANA ISS service:</strong></p> <p><em>"Folder "Reads" contains coordinates and gene information of segmented spots.<br> The coordinates are in pixel unit. Scaling factor is 0.32 um/pixel. (0,0) is at northwest (top-left corner).<br> With Low/High Threshold, we refer to the quality thresholding. Our technology is fluorescence based, i.e. with the thresholding one can balance how certain the signals are.</em></p> <p><em>Files ending with _LowThreshold: reads not matching with any known barcode were already discarded.</em></p> <p><em>Files ending with _HighThreshold: has information only about spots that passed additional quality check.</em></p> <p><em>Folder "GenePlots" has plotted images in static .png format, fully zoomed out. LowThreshold and HighThreshold follow the same thresholding strategy as in reads files."</em></p> <p> </p> <p><strong>SiT-master.zip is a download of the GitHub repository </strong><a href="https://github.com/ucagenomix/SiT">https://github.com/ucagenomix/SiT</a>, <strong>providing figures and analysis scripts for SiT.</strong></p> <p> </p> <p><strong>Related SiT data are deposited through GEO, accession number <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE153859">GSE153859</a></strong></p>
KG for heart failure gene expression data
<p>Pre processed gene expression data for different heart failure. Includes count table, gene patiens metadata, gene lenght</p>
Master Coral database used in USVI SCTLD Transmission Experiment Gene Expression Analysis
<p>The Master Coral Database fasta file is comprised of previously published genome-derived predicted gene models and transcriptomes spanning a wide diversity of coral families. Transcriptomes are from Davies et al., 2016 (doi: 10.3389/fmars.2016.00112), Kirk et al., 2018 (DOI: 10.1111/mec.14934); Moya et al., 2012 (doi: 10.1111/j.1365-294X.2012.05554.x); van de Water et al., 2018 (DOI: 10.1111/mec.14489).</p>
Genome-wide characterization of human minisatellite VNTRs: population-specific alleles and gene expression differences
<p>This repository consists of minisatellite VNTR genotypes for 2,800 samples (2,770 individuals). The raw VCF files were produced using <a href="https://github.com/yzhernand/VNTRseek">VNTRseek</a> on xxx data sources: <a href="http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1000_genomes_project/">30 high coverage WGS datasets</a> from the 1000 Genomes Project phase 3, <a href="https://www.internationalgenome.org/data-portal/data-collection/30x-grch38">2,504 unrelated genomes</a> from New York Genome Center (NYGC), <a href="https://www.internationalgenome.org/data-portal/data-collection/sgdp">253 genomes from Simons Diversity Genome Project</a> (SGDP), <a href="https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/tumor-normal.html">two tumor-normal breast cancer samples</a> from Illumina Basespace, haploid genomes <a href="https://www.ncbi.nlm.nih.gov/sra/SRX652547">CHM1 </a>and <a href="https://www.ncbi.nlm.nih.gov/sra/SRX1009644">CHM13</a>, and seven genomes from the Personal Genome Project from the Genome In A Bottle Consortium (GIAB). Raw VCF files are provided for each data source separately.</p> <p>The raw VCF files were preprocessed (preprocess.sh) to extract genotypes and provided in VNTRseek_preprocessed_data.tar.gz (uncompressed size 10G). The R Markdown code to analyze the preprocessed data and produce figures and tables is also provided (tables_and_figures.Rmd). For more information see the ReadMe file.</p> <p>This work was supported in part by NSF grants IIS-1423022 and DBI-1559829.</p>
Supplementary Tables for Expression of cell-wall related genes is highly variable and correlates with sepal morphology
<p>Supplementary Information and script for "Expression of cell-wall related genes is highly variable and correlates with sepal morphology"</p> <p>R scripts for analysis</p> <p>Data necessary to run analyses</p> <p>Generated data</p> <p>Supplementary Tables</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.