Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,801
datasets available to search
ShareScore release 0.9.0
Dataset results
3,801 results for “ATAC-seq”
Test data for running snakePipes : ATAC-seq workflow
<p><strong>Test files for running snakePipes workflows</strong></p> <p><strong>snakePipes</strong> are pipelines built using snakemake and python for the analysis of epigenomic datasets. Please refer to <a href="https://snakepipes.readthedocs.io/en/latest/">this link</a> for further information on snakePipes.</p> <p>This folder contains test files that can be used to run the ATAC-seq workflow under snakePipes. To test the workflow, follow the following steps : </p> <ul> <li>Download or prepare genome fasta, indices and annotations for fruit fly (<strong>dm6</strong>) genome.</li> <li>Download and install snakePipes via `conda create -n snakePipes -c mpi-ie -c bioconda -c conda-forge snakePipes`</li> <li>Update <a href="https://snakepipes.readthedocs.io/en/latest/content/running_snakePipes.html#genome-configuration-file">Genome configuration file</a> with path to indices and annotations.</li> <li>Move to this repository and run the example <strong>command.sh</strong></li> </ul>
Supporting data for "A simple ATAC-seq protocol for population epigenetics"
<p>This is supporting data for an article in which we describe a protocol for the generation of sequence-ready libraries for population epigenomics studies. The protocol is a streamlined version of the Assay for transposase accessible chromatin with high-throughput sequencing (ATAC-seq) that provides a positive display of accessible, presumably euchromatic regions. The protocol is straightforward and can be used with small individuals such as daphnia and schistosome worms, and probably many other biological samples of comparable size, and it requires little molecular biology handling expertise.</p> <p>In "Agarose picture.Tif" the left lane shows the 100 bp size marker, first 10 bands from down to top: 100bp, 200bp, 300bp, 400bp, 500bp, 600bp, 700bp, 800bp, 900bp and 1kbp.</p> <p>Produced at IHPE (http://ihpe.univ-perp.fr/)</p>
Schistosoma mansoni ATAC-seq results for IGV (female and male worms with and without LSD1 inhibitor)
<p>In this study, the anti-schistosomal activity of 39 <em>Homo sapiens</em> Lysine Specific Demethylase 1 (HsLSD1) inhibitors was investigated on parasitic life cycle stages associated with both definitive and intermediate host infection. Amongst this collection of small molecules, compound <strong>33</strong> was the most potent and reduced <em>ex vivo</em> viabilities of schistosomula, juveniles, miracidia and adults. At its sub-lethal concentration to adults (3.13 µM), compound <strong>33 </strong>also significantly impacted oviposition, ovarian as well as vitellarian architecture and gonadal/neoblast stem cell proliferation. ATAC-seq analysis of adults demonstrated that compound <strong>33</strong> significantly affected chromatin structure (intragenic regions > intergenic regions), especially in genes differentially expressed in cell populations (e.g., germinal stem cells, hes2<em><sup>+</sup></em>stem cell progeny, S1 cells and late female germinal cells) linked to these <em>ex vivo</em> phenotypes.</p> <p>The data presented here allow for visualisation in IGV https://igv.org/app/</p> <p>Produced in collaboration with IHPE. </p>
Robust estimation of cancer and immune cell-type proportions from bulk tumor ATAC-Seq data.
<p>Bulk ATAC-seq data of tumour samples result in an averaged signal across different cell-types (cancer, stromal, vascular and immune cells). We propose a deconvolution framework called EPIC-ATAC (<a href="https://doi.org/10.7554/eLife.94833.1">https://doi.org/10.7554/eLife.94833.1</a>), which relies on newly identified cell-type specific ATAC-Seq marker peaks and reference profiles for all major cancer-relevant cell-types to predict the proportions of each cell-type.</p> <p>To evaluate EPIC-ATAC, we generated a bulk ATAC-Seq dataset from peripheral blood mononuclear cells (PBMCs) samples, from which the number of cells in each cell-type has been estimated using flow cytometry, as ground truth for cell proportions. The data provided in this Zenodo deposit correspond to:</p> <p>- The raw counts matrix for each peak called in this ATAC-Seq dataset: PBMC_counts.txt</p> <p>- The normalized (TPM-like) counts matrix for each peak called in this ATAC-Seq dataset: PBMC_counts_norm.txt</p> <p>- The cell fractions of each cell type in each sample: PBMC_cell_fractions.txt</p> <p>- The peaks called in each sample using MACS2 (*narrow.peaks): *_normalized.narrowPeak</p> <p>- Bed files listing ATAC-Seq fragments for each sample: *.bed</p> <p>We also evaluated EPIC-ATAC on multiple pseudobulks generated from single-cell ATAC-Seq data. We provide rds files containing the pseudobulks data used in our work for the evaluation of EPIC-ATAC. The rds files are located in the zip file "pseudobulks.zip".</p> <p>The file "additional_data.zip" contains additional files used to generate the reference profiles in EPIC-ATAC and to reproduce the main analyses performed in the manuscript: <a href="https://doi.org/10.7554/eLife.94833.1">https://doi.org/10.7554/eLife.94833.1</a>. These files are required to run the code available on the following GitHub repository: GfellerLab/EPIC-ATAC_manuscript. </p>
ATAC-seq dataset: Chromatin accessibility landscapes activated by cell-surface and intracellular immune receptors
<p>The dataset encompasses raw sequencing reads, identified peaks, and regions of differential accessibility derived from ATAC-seq experiments conducted under various immune activation conditions. For additional technical details regarding data collection, please refer to the published source at https://doi.org/10.1093/jxb/erab373.</p>
ATAC-seq processing resources for the GRCh38 (hg38) assembly of the human genome
<p>A collection of publicly available, but preprocessed, reference data for the analysis of ATAC-seq samples using the GRCh38 (hg38) assembly of the human genome using the <a href="https://doi.org/10.5281/zenodo.6323634">Ultimate ATAC-seq Data Processing & Analysis Pipeline</a> (details in the documentation on GitHub).</p>
Data for bulk ATAC-seq normalization paper
<p>This zipped file contains all public datasets used in our benchmark of bulk ATAC-seq normalization methods.</p>
Genomica ed Epigenomica Biocomputazionale: ATAC-seq Mesoderm differentiation data1
<p>Data collection of FastQ files used for teaching purposes.</p> <p>This subset was obtained from Koh, P., Sinha, R., Barkal, A. <em>et al.</em> An atlas of transcriptional, chromatin accessibility, and surface marker changes in human mesoderm development. <em>Sci Data</em> <strong>3</strong>, 160109 (2016). https://doi.org/10.1038/sdata.2016.109</p>
ATAC-seq analysis for intravesical BCG in bladder cancer
<p>BCG vaccination can boost innate immune responses via trained immunity (TI), resulting in an increased resistance to respiratory viral infections. Assay for transposase accessible chromatin (ATAC), including tagmentation, library preparation and sequencing were performed by Genewiz (Azenta Life Sciences, MA, USA) on PBMCs from two BCG-treated NMIBC patients at baseline and during BCG (mid) as well as from three healthy donors. This dataset include the pipeline for preprocessing of raw data including mapping to the hg38 reference genome using bowtie2 and peak calling by MACS2. Differentially accessible regions in proximity to annotated genes between the two time points (during BCG versus baseline) were identified using the R packages csaw and edgeR and resulting data files are provided.</p> <p> </p> <p> </p> <p> </p> <p> </p>
Training material for the mapping and quantification of single-cell ATAC-seq 10X Datasets
<p>The data provided here is part of the Galaxy Training Network tutorial that analyses 10x genomics single-cell ATAC-seq data from the 10x platform. The original data is from 1k Peripheral Blood Mononuclear Cells (PBMCs) from a Healthy Donor.</p> <p>Due to time constraints during training, the datasets were subsampled to reads that map to chromosome 21 only.</p> <p>The 10x Genomics Datasets follow the <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution</a> license.</p> <p>There is an additional count matrix in Anndata format created from full datasets.</p>
Data from: Single-nucleus RNA-seq and ATAC-seq in outbred rats with divergent cocaine addiction behaviors reveal long-term changes in gene regulation and GABAergic inhibition in the amygdala
<p>This dataset accompanies our publication titled: "Single-nucleus RNA-seq and ATAC-seq in outbred rats with divergent cocaine addiction behaviors reveal long-term changes in gene regulation and GABAergic inhibition in the amygdala."</p> <p><strong>Files Included:</strong></p> <p><strong>1. geno.N26.vcf.gz</strong><br> - Description: Contains genotypes for 26 Heterogeneous Stock rats whose gene expression was predicted.</p> <p><strong>2. pred_expr.Brain.N26.tsv</strong><br> - Description: This tab-delimited table contains predicted relative gene expression in the brain for 26 Heterogeneous Stock rats. <br> - Details: Predictions were made for 8,997 genes from linear models based on cis-eQTLs from whole brain hemisphere tissue downloaded from the <a href="https://ratgtex.org/download/">RatGTEx Portal</a>. A gene is included in the table if it had at least one significant cis-eQTL, and if its predicted expression in these 26 animals had nonzero variance. The values in the table give the predicted log2(relative expression), where log2(2) = 1 is the baseline expression from the two haplotypes of the gene if it had only reference alleles at all its regulatory loci.<br> - Additional Info: Predictions were generated using <strong>gene_expr_pred.py</strong> available at https://github.com/PejLab/gene_expr_pred<br> An explanation of the prediction model is given in https://doi.org/10.1101/2022.01.28.478116</p> <p><strong>3. Behavioral data.xlsx</strong><br> - Description: Contains behavioral data for the Heterogeneous Stock (HS) rats.<br> - Organization: Each sheet in the file corresponds to data for a specific figure.</p> <p><strong>Additional Dataset Locations:</strong></p> <p>The primary datasets generated during this study can be found on the Gene Expression Omnibus under accession number <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE212417">GSE212417</a></p> <p><strong>Publicly Available Datasets Utilized:</strong></p> <p>- Rattus norvegicus Ensembl v98 reference genome and genome assembly: <a href="http://useast.ensembl.org/Rattus_norvegicus/Info/Index">Rnor_6.0 </a><br> - JASPAR2022 transcription factor binding profiles for vertebrates: <a href="https://jaspar.genereg.net/">JASPAR</a><br> - ENCODE Honeybadger 2 ChIP-seq: <a href="https://personal.broadinstitute.org/meuleman/reg2map/">Broad Institute</a><br> - Liu et al. 2019106 GWAS for tobacco and nicotine addiction summary statistics: <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6358542/">PubMed</a><br> - RatGTEx Portal tissue-specific cis-eQTLs: <a href="https://ratgtex.org/download/">RatGTEx Portal </a><br> - 1000 Genomes European reference panel: <a href="https://alkesgroup.broadinstitute.org/LDSCORE/">Alkes Group</a><br> - KEGG pathways: <a href="https://www.kegg.jp/kegg/rest/keggapi.html">KEGG API</a></p>
KLF2 maintains lineage fidelity and suppresses CD8 T cell exhaustion during acute LCMV infection (LCMV DSM scRNA data and ATAC-seq)
Open the record for dataset details and reuse information.
Inputs for Galaxy Trainig ATAC-seq
<p>The fastq.gz are a subset of SRR891268 but enriched into pairs which map to chr22.</p> <p>The bed is from ENCODE.</p>
Single-cell ATAC-seq control of cross-contaminations (experiment 2)
<p>On the Fluidigm C1 platform for single-cell analysis, the cells are captured in 96 chambers arranged serially, and then washed before further processing. Thus, debris present from the loading medium or released by captured cells upstream are a possible source of contamination. We generated a control datasets using the single-cell ATAC-seq protocol available from Fluidigm's ScriptHub. We cultivated human Hep G2 and mouse Hepa 1-6 (both are liver cancer cell lines), stained them with green and red calceins (respectively), and loaded them at equal concentration in a Fluidigm medium flow cell (old design), before running the C1 single-cell ATAC-seq program. To evaluate damage and carry-over of debris from FACS-sorting, two IFCs were run in two C1 machines in parallel. In the first (flowcell ID 1772-123-155), the cells not washed and in the second, they were washed (ID 1772-123-158).</p> <p>The data deposited here is a sequencing run (Illumina MiSeq) of these ATAC-seq libraries. The metadata indicating the contents of each well is being uploaded separately and this record will be updated once the DOIs are available.</p>
Single-cell ATAC-seq control of cross-contaminations (experiment 1)
<p>On the Fluidigm C1 platform for single-cell analysis, the cells are captured in 96 chambers arranged serially, and then washed before further processing. Thus, debris present from the loading medium or released by captured cells upstream are a possible source of contamination. We generated a control datasets using the single-cell ATAC-seq protocol available from Fluidigm's ScriptHub. We cultivated human Hep G2 and mouse Hepa 1-6 (both are liver cancer cell lines), stained them with green and red calceins (respectively), and loaded them at equal concentration in a Fluidigm medium flow cell (old design), before running the C1 single-cell ATAC-seq program.</p> <p>The data deposited here is a sequencing run (Illumina MiSeq) of these ATAC-seq libraries. The metadata indicating the contents of each well is being uploaded separately and this record will be updated once the DOIs are available.</p>
ATAC-seq raw reads, MACS3 outputs
<p>The dataset contains raw reads that were mapped to the nuclear genome and all outputs of MACS3 in <a href="https://doi.org/10.5281/zenodo.10972575">10.5281/zenodo.10972575</a></p>
ATAC-seq processing resources for the GRCm38 (mm10) assembly of the mouse genome
<p>A collection of publicly available, but preprocessed, reference data for the analysis of ATAC-seq samples using the GRCm38 (mm10) assembly of the mouse genome using the <a href="https://doi.org/10.5281/zenodo.6323634">Ultimate ATAC-seq Data Processing & Analysis Pipeline</a> (details in the documentation on GitHub).</p>
Processing of 10X 500 PBMC single cell ATAC-seq data with SnapATAC2
<p>Input files for SnapATAC2 tutorial on Galaxy</p>
Processed read counts from macrophage RNA-seq and ATAC-seq experiments
<p><strong>RNA-seq files:</strong></p> <ol> <li>RNA_count_matrix.txt.gz - raw read counts</li> <li>RNA_cqn_matrix.txt.gz - read counts quantile normalised with the cqn R package</li> <li>RNA_gene_metadata.txt.gz - information about the genes</li> <li>RNA_sample_metadata.txt.gz - information about the samples</li> </ol> <p><strong>ATAC-seq files:</strong></p> <ol> <li>ATAC_count_matrix.txt.gz - raw read counts</li> <li>ATAC_cqn_matrix.txt.gz - read counts quantile normalised with the cqn R package</li> <li>ATAC_peak_metadata.txt.gz - peak coordinates and other metadata</li> <li>ATAC_sample_metadata.txt.gz - sample metadata</li> <li>ATAC_consensus_peaks.gff3.gz - GFF3 file containing the GRCh38 coordinates of the peaks</li> </ol>
quaqc: Efficient and quick ATAC-seq quality control and filtering at any scale
<p>This repository contains data and code necessary to replicate the analyses for the associated manuscript. Copies of the quaqc and quaqcr programs are also included.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.