Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

137

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

137 results for “transcription start site”

Learn how ShareScore rates datasets ↗
dryad40/100

Transcription start site analysis for heterogenous CD4+ T cells using 5′ scRNA-seq

<p>These datasets are generated by ReapTEC (read-level pre-filtering and transcribed enhancer call) using 5' single-cell RNA-seq data on human heterogenous CD4+ T cells. By taking advantage of a unique "cap signature" derived from the 5′-end of a transcript, ReapTEC simultaneously profiles gene expression and enhancer activity at nucleotide resolution using 5′-end single-cell RNA-sequencing (5′ scRNA-seq). The detail of ReapTEC pipeline is described in https://github.com/MurakawaLab/ReapTEC.</p>

opencc-zeroApr 2024View details →
dryad40/100

Transcription start site analysis for heterogenous CD4+ T cells using 5′ scRNA-seq

Open the record for dataset details and reuse information.

publicApr 2024View details →
zenodo36/100

Transcription start sites from capped small RNA-seq of rat nucleus accubmens and prefrontal cortex

<p>Small RNAs of &sim;15&ndash;60 nt were size selected by denaturing gel electrophoresis starting from total RNA extracted from 14 rat brain tissue dissections. For csRNA libraries, cap selection was followed by decapping, adapter ligation, and sequencing. For input libraries, 10% of small RNA input was used for decapping, adapter ligation, and sequencing. After library quality check by gel electrophoresis, the samples were sequenced using the Illumina NextSeq 500 platform using 75 cycles single end. Sequencing reads were aligned to the rat mRatBN7.2 genome assembly using STAR v2.5.3a aligner with default parameters. Transcriptional start regions were defined using HOMER&rsquo;s findPeaks tool.&nbsp;</p> <p>Duttke, S.H., Montilla-Perez, P., Chang, M.W., Li, H., Chen, H., Carrette, L.L.G., de Guglielmo, G., George, O., Palmer, A.A., Benner, C., et al. (2022). Glucocorticoid Receptor-Regulated Enhancers Play a Central Role in the Gene Regulatory Networks Underlying Drug Addiction. Front. Neurosci. 16, 858427.</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

Datasets for : High-resolution detection and differential expression analysis of transcription start sites using MAPCap

<p>This dataset corresponds to the study:&nbsp;High-resolution detection and differential expression analysis of transcription start sites using MAPCap (Bhardwaj&nbsp;et. al. 2018)</p> <p>It includes:</p> <p>&nbsp;- TSS identified using MAPCap in stage 15 embryos and larvae.</p> <p>&nbsp;- Differentially expressed TSS using MAPCap (FDR &lt; 0.05) in larvae.</p> <p>&nbsp;- Common and stage-specific enhancer TSS identified in this study</p>

opencc-by-4.0Apr 2019View details →
dryad28/100

Data from: Paired-end analysis of transcription start sites in Arabidopsis reveals plant-specific promoter signatures

Understanding plant gene promoter architecture has long been a challenge due to the lack of relevant large-scale data sets and analysis methods. Here, we present a publicly available, large-scale transcription start site (TSS) data set in plants using a high-resolution method for analysis of 5′ ends of mRNA transcripts. Our data set is produced using the paired-end analysis of transcription start sites (PEAT) protocol, providing millions of TSS locations from wild-type Columbia-0 Arabidopsis thaliana whole root samples. Using this data set, we grouped TSS reads into "TSS tag clusters" and categorized clusters into three spatial initiation patterns: narrow peak, broad with peak, and weak peak. We then designed a machine learning model that predicts the presence of TSS tag clusters with outstanding sensitivity and specificity for all three initiation patterns. We used this model to analyze the transcription factor binding site content of promoters exhibiting these initiation patterns. In contrast to the canonical notions of TATA-containing and more broad "TATA-less" promoters, the model shows that, in plants, the vast majority of transcription start sites are TATA free and are defined by a large compendium of known DNA sequence binding elements. We present results on the usage of these elements and provide our Plant PEAT Peaks (3PEAT) model that predicts the presence of TSSs directly from sequence.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Paired-end analysis of transcription start sites in Arabidopsis reveals plant-specific promoter signatures

Open the record for dataset details and reuse information.

publicJun 2015View details →
geo24/100

DNA methylation controls unmethylated transcription start sites in the genome in trans

GEO Series GSE85313. Homo sapiens. 9 samples. Type: Expression profiling by array.

openGEO-OpenJul 2025View details →
geo24/100

The transcription start site and enhancer landscape of the descending colon in inflammatory bowel disease

GEO Series GSE95437. Homo sapiens. 94 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMar 2018View details →
geo24/100

Genome-wide mapping of transcription start sites in a ∆set2 strain

GEO Series GSE62735. Saccharomyces cerevisiae. 4 samples. Type: Other.

openGEO-OpenJan 2015View details →
geo24/100

The TRIPLE PHD FINGERS proteins are required for SWI/SNF complex-mediated +1 nucleosome positioning and transcription start site selection in Arabidopsis [RNA-seq]

GEO Series GSE205111. Arabidopsis thaliana. 9 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenSep 2022View details →
geo24/100

Transcript 5'-end mapping was used to identify transcriptional start sites and RNA processing sites genome-wide in Mycobacterium tuberculosis

GEO Series GSE62152. Mycobacterium tuberculosis. 9 samples. Type: Expression profiling by high throughput sequencing; Other.

openGEO-OpenMay 2015View details →
geo24/100

G2/M DNA synthesis at transcription start sites because of RNA transcription persistence during DNA replication [Chr-RNA-seq]

GEO Series GSE136292. Homo sapiens. 12 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenFeb 2021View details →
geo24/100

Alternative transcription start sites contribute to acute-stress–induced transcriptome response in human skeletal muscle

GEO Series GSE164081. Homo sapiens. 50 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenDec 2020View details →
geo24/100

Analysis of Drosophila melanogaster nascent transcription start sites

GEO Series GSE203135. Drosophila melanogaster. 4 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenOct 2022View details →
geo24/100

Transcription start site profiling of HSV-1 infected cells using cRNA-seq and dRNA-seq

GEO Series GSE128323. Homo sapiens; Human alphaherpesvirus 1 strain 17. 18 samples. Type: Other.

openGEO-OpenApr 2020View details →
geo24/100

G2/M DNA synthesis at transcription start sites because of RNA transcription persistence during DNA replication [ChIP-seq]

GEO Series GSE136291. Homo sapiens. 12 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenFeb 2021View details →
geo24/100

G2/M DNA synthesis at transcription start sites because of RNA transcription persistence during DNA replication [Chip-Seq H2AX]

GEO Series GSE160833. Homo sapiens. 3 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenFeb 2021View details →
geo24/100

Quantitative analysis of transcription start site selection in Saccharomyces cerevisiae reveals control by DNA sequence, RNA Polymerase II activity, and NTP levels

GEO Series GSE185290. Escherichia coli; Saccharomyces cerevisiae. 90 samples. Type: Other.

openGEO-OpenNov 2023View details →
geo24/100

Genome-wide Identification of Transcription Start Sites in Two Alphaproteobacteria

GEO Series GSE150944. Novosphingobium aromaticivorans; Cereibacter sphaeroides 2.4.1. 24 samples. Type: Other.

openGEO-OpenJul 2020View details →
geo24/100

AMPKα2 knockout mice and usage of alternative transcription start sites

GEO Series GSE112325. Mus musculus. 4 samples. Type: Expression profiling by high throughput sequencing; Other.

openGEO-OpenMay 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record