Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
21,320
datasets available to search
ShareScore release 0.9.0
Dataset results
21,320 results for “Transcription”
Data from: Transcriptional profiling of lung macrophages following ozone exposure in mice identifies signaling pathways regulating immunometabolic activation
<p>Macrophages play a key role in ozone-induced lung injury by regulating both the initiation and resolution of inflammation. These distinct activities are mediated by pro-inflammatory and anti-inflammatory/pro-resolution macrophages which sequentially accumulate in injured tissues. Macrophage activation is dependent, in part, on intracellular metabolism. Herein, we used RNA-sequencing (seq) to identify signaling pathways regulating macrophage immunometabolic activity following exposure of mice to ozone (0.8 ppm, 3 hr) or air control. Analysis of lung macrophages using an Agilent Seahorse showed that inhalation of ozone increased macrophage glycolytic activity and oxidative phosphorylation at 24 and 72 hr post exposure. An increase in the percentage of macrophages in the S phase of the cell cycle was observed 24 hr post ozone. RNA-seq revealed significant enrichment of pathways involved in innate immune signaling and cytokine production among differentially expressed genes at both 24 and 72 hr after ozone, while pathways involved in cell cycle regulation were upregulated at 24 hr and intracellular metabolism at 72 hr. An interaction network analysis identified tumor suppressor 53 (TP53), E2F family of transcription factors (E2Fs), Cyclin Dependent Kinase Inhibitor 1A (CDKN1a/p21), and Cyclin D1 (CCND1) as upstream regulators of cell cycle pathways at 24 hr and TP53, nuclear receptor subfamily 4 group a member 1 (NR4A1/Nur77), and estrogen receptor alpha (ESR1/ERα) as central upstream regulators of mitochondrial respiration pathways at 72 hr. These results highlight the complex interaction between cell cycle, intracellular metabolism, and macrophage activation which may be important in the initiation and resolution of inflammation following ozone exposure.</p>
Corallium rubrum genome and transcripts alignment
<p><em>Corallium rubrum</em>, the precious red coral, is an octocoral endemic to the western Mediterranean Sea. It plays a key role in the well-known Mediterranean coralligenous ecosystem, a biodiversity hotspot. We sequenced its genome for further investigations into its biology (skeleton formation and color, genetic diversity), ecology, and evolutionary history.</p>
Taxonomic and functional annotations of transcripts and proteins derived from a Metatranscriptomic study of microbial eukaryotes from Lake Pavin
<p>These data were obtain as part of a metatranscriptomic study (Monjot <em>et al.,</em> 2023, 2024). All scripts to obtain these annotations are available at https://github.com/amonjot/SSN_Monjot_2024. The sequencing data (i.e. metatranscriptomic) used to obtain this taxonomic and functional information are archived at ENA under accession number PRJEB61515.</p> <p>This repository also contains various protein sequence similarity networks (Lagoon_output.zip). As these files are very time-consuming to produce, we have provided them to complete all the steps in Monjot <em>et al,</em> 2024. All procedures to produce them are present on the following github repository : https://github.com/amonjot/SSN_Monjot_2024.</p> <p> </p>
Post-transcriptional regulation supports the homeostatic expression of mature RNA
<h3>Overview</h3> <p>This dataset consists of RNA sequencing data stored in HDF5 files. The data has been processed to quantify gene expression changes at the precursor RNA (preRNA) and mature RNA (matureRNA) levels across various biological conditions, including normal tissues and diseases. The dataset supports the study "Post-transcriptional regulation supports the homeostatic expression of mature RNA," providing insight into the general influence of post-transcriptional regulation (PTR) on gene expression homeostasis.</p> <h3>File Contents</h3> <p>Each HDF5 file contains the following datasets:</p> <ul> <li><strong>logFC_MatureRNA</strong>: Represents the log fold change of mature RNA expression levels between different conditions or tissues.</li> <li><strong>logCPM_MatureRNA</strong>: Represents the log counts per million for mature RNA.</li> <li><strong>PValue_MatureRNA</strong>: Represents the p-value for the statistical test performed on mature RNA expression levels.</li> <li><strong>FDR_MatureRNA</strong>: Represents the false discovery rate for the mature RNA expression levels.</li> <li><strong>logFC_preRNA</strong>: Represents the log fold change of precursor RNA expression levels between different conditions or tissues.</li> <li><strong>logCPM_preRNA</strong>: Represents the log counts per million for precursor RNA.</li> <li><strong>PValue_preRNA</strong>: Represents the p-value for the statistical test performed on precursor RNA expression levels.</li> <li><strong>FDR_preRNA</strong>: Represents the false discovery rate for the precursor RNA expression levels.</li> <li><strong>genes</strong>: A list of gene identifiers used in the study.</li> <li><strong>studies</strong>: A list of study or condition names included in the dataset.</li> </ul> <h3>File Naming Convention</h3> <p>The HDF5 files are named to reflect the filtering applied to exclude lowly expressed genes (logCPM > 5), and DESeq2 and edgeR refer to the tools used for differential expression analysis:</p> <ul> <li><code>DESeq2_filtered_exon_intron_abnormal_conditions_all_reads.h5</code></li> <li><code>edgeR_filtered_exon_intron_abnormal_conditions_all_reads.h5</code></li> <li><code>DESeq2_filtered_exon_intron_gtex_tissue_all_reads.h5</code></li> <li><code>edgeR_filtered_exon_intron_gtex_tissue_all_reads.h5</code></li> </ul> <h3>Example Usage</h3> <p>To query the HDF5 files, you can use the provided `<a href="https://github.com/suzheng/PTR_RNA_seq/blob/main/notebooks/parse_h5_file.ipynb">parse_h5_file.ipynb</a>` example notebook to parse the HDF5 files.</p>
Single-molecule dynamics and genome-wide transcriptomics reveal that NF-kB (p65)-DNA binding times can be decoupled from transcriptional activation
<p>Data and analysis code repository for: </p> <p>Single-molecule dynamics and genome-wide transcriptomics reveal that NF-kB (p65)-DNA binding times can be decoupled from transcriptional activation</p> <p>https://doi.org/10.1101/255380</p>
Datasets and Scripts: Full-length mRNA sequencing uncovers a widespread coupling between transcription and mRNA processing
<p>The multilayered control of gene expression requires tight coordination of regulatory mechanisms at the transcriptional and post-transcriptional level. In this study, we studied the interdependence of transcription, splicing and polyadenylation events on single mRNA molecules by full-length mRNA sequencing. In MCF-7 breast cancer cells and three human tissues, we found an unforeseen number of genes that demonstrate mutually inclusive or exclusive alternative transcription and mRNA processing events, which can span the entire length of mRNA molecules. Furthermore, alternative poly(A) sites that are coupled with alternative splicing events are depleted for known poly(A) signals and enriched for MBNL binding motifs, supporting a dual role of MBNL proteins in regulating splicing and polyadenylation. We predict thousands of open-reading frames from the sequence of full-length mRNAs, allowing for a more sensitive proteogenomics analysis of MCF-7 mass-spectrometry data. Our findings demonstrate that our understanding of transcriptome complexity is far from complete and provides a framework to reveal largely unresolved mechanisms that coordinate transcription and mRNA processing.</p>
Annexe 22 - Texte de la transcription d'entrevue semi-dirigée réalisée auprès des usagers réels ou potentiels de la traduction ou de l'interprétation
<p><strong>Annexe 22- Texte de la transcription d’entrevue semi-dirigée réalisée auprès des usagers réels ou potentiels de la traduction ou de l’interprétation</strong></p>
Summary statistics of transcript usage QTLs in naive and stimulated macrophages (part 2)
<p>Column names:</p> <ol> <li>phenotype_id</li> <li>pheno_chr</li> <li>pheno_start</li> <li>pheno_end</li> <li>strand - strand of the phenotype</li> <li>n_snps - number of SNPs tested per phenotype</li> <li>distance - distance from the variant to the phenotype</li> <li>snp_id</li> <li>snp_chr</li> <li>snp_start - position of the SNP</li> <li>snp_end - same as snp_start</li> <li>p_nominal - nominal p-value from QTLTools</li> <li>beta - effect size from QTLTools</li> <li>is_lead - is the variant the lead QTL for the phenotype?</li> </ol>
Summary statistics of transcript usage QTLs in naive and stimulated macrophages (part 1)
<p>Column names:</p> <ol> <li>phenotype_id</li> <li>pheno_chr</li> <li>pheno_start</li> <li>pheno_end</li> <li>strand - strand of the phenotype</li> <li>n_snps - number of SNPs tested per phenotype</li> <li>distance - distance from the variant to the phenotype</li> <li>snp_id</li> <li>snp_chr</li> <li>snp_start - position of the SNP</li> <li>snp_end - same as snp_start</li> <li>p_nominal - nominal p-value from QTLTools</li> <li>beta - effect size from QTLTools</li> <li>is_lead - is the variant the lead QTL for the phenotype?</li> </ol>
Transcript annotations for the spiny mouse (ref https://doi.org/10.5281/zenodo.1188364)
<p>Annotations for spiny mouse transcripts ref project https://doi.org/10.1101/280412</p> <p>Transcript ID, Corset cluster ID, BLASTx results, UniProtKB accession ID, UniProt ID, Gene name and Protein name included.</p>
GENCODE mouse & human transcript to gene files for txImport
<p>GENCODE data downloaded for human (GRCh38p12 v28) and mouse (GRCm38p6 vM18). For each organism (human=v28, mouse=vM18), there is an R script with the commands necessary to generate the tx2gene CSV that can be read in and used with the tximport Bioconductor package for summarizing transcript abundances to the gene level.</p> <p><strong>gencode.v28 = human</strong></p> <p><strong>gencode.vM18 = mouse</strong></p> <p> </p>
Flute audio labelled database for Automatic Music Transcription
<p>Automatic Music Transcription (ATM) is a well-known task in the Music Information Retrieval (MIR) domain and consists on the computation of a symbolic music representation from an audio recording. In this work, our focus is to adapt algorithms that extract musical information from an audio file for a particular instrument. The main objective is to study the automatic transcription of digitized music support systems. Currently, these techniques are applied to a generic sound timbre, to sounds to any instrument for further analysis and conversion to a digital music encoding and final score format. The results of this project add new knowledge in this automatic transcription field, since traverse flute has been selected as the instrument on which to focus all the process and, until now, there is no database of flute sounds for this purpose.</p> <p>For so, we have recorded some sounds, both monophonic and polyphonic music. These audio files have been processed by the chosen transcription algorithm and converted to a digital music encoding format for its posterior alignment with the original recordings. Once all these data have been converted to text, the resulting labeled database its constituted by the initial audios and final aligned files.</p> <p>Furthermore, after this process and from the obtained data, an evaluation of the transcriptor behavior has been made based on two main techniques: note and frame level.</p> <p>This database includes the original audio files (.wav), transcribed MIDI files (.mid), aligned MIDI files (.mid), aligned text files (.txt) and evaluation files (.csv).</p>
Strong gene activation in plants with genome-wide specificity using a new orthogonal CRISPR/Cas9-based Programmable Transcriptional Activator.
<p>This data set correspond to the supporting data generated in the manuscript: Strong gene activation in plants with genome-wide specificity using a new orthogonal CRISPR/Cas9-based Programmable Transcriptional Activator. </p>
Interaction of N-3-oxododecanoyl homoserine lactone with transcriptional regulator LasR of Pseudomonas aeruginosa: Insights from molecular docking and dynamics simulations
<p>Dataset and supplementary files of the research: Interaction of N-3-oxododecanoyl homoserine lactone with transcriptional regulator LasR of Pseudomonas aeruginosa: Insights from molecular docking and dynamics simulations (https://doi.org/10.1101/121681)</p> <p>- Supporting Information</p> <p>- Input: Parameters and initial structures</p> <p>- Output: Trajectories, Docking poses</p> <p>Gromacs (multi-core with CUDA) was used for the simulations.</p> <p>Autodock Vina, FlexAid and rDock were used for molecular docking.</p>
Pilot 3 interview transcripts
<p>This dataset provides the user responses to the interviews that will be performed during the evaluation phase of the project and will be designed around the unified theory of user’s acceptance of technology.</p> <p>A simple Dublin core set of metadata will be used to describe the dataset.</p> <p>It will be created during the pilot 3 tests in order to understand how users accept and perceive the concept. No reuse of additional information is foreseen, as this information cannot be enhanced by any external contribution.</p> <p>The information volume that will be stored is quite small, several Kbytes for each participant at the most, yielding up to a few Mbytes per experimentation session.</p> <p>The dataset could be used by any interested researchers. It is expected that the dataset can contribute to an understanding of the users’ acceptance of the CROSSCULT technology and to an article based on this study. It will also inform the evaluation framework.</p>
Pilot 2 interview transcripts
<p>This dataset provides the user responses to the interviews that will be performed during the evaluation phase of the project and will be designed around the unified theory of user’s acceptance of technology.</p> <p>A simple Dublin core set of metadata will be used to describe the dataset.</p> <p>It will be created during the pilot 2 tests in order to understand how users accept and perceive the concept. No reuse of additional information is foreseen, as this information cannot be enhanced by any external contribution.</p> <p>The information volume that will be stored is quite small, several Kbytes for each participant at most, yielding up to a few Mbytes per experimentation session.</p> <p>The dataset could be used by any interested researchers. It is expected that the dataset can contribute to an understanding of the users’ acceptance of the CROSSCULT technology and to an article based on this study. It will also inform the evaluation framework.</p>
cldf/clts: Cross-Linguistic Transcription Systems
<p>Cross-Linguistic Transcription Systems</p>
Research data supporting "Rolling Circle Transcription-Amplified Hierarchically Structured Organic-Inorganic Hybrid RNA Flowers for Enzyme Immobilization""
<p>Raw research data supporting the publication:</p> <p>Wang Y. et al., 2019, ACS Applied Materials and Interfaces, DOI: 10.1021/acsami.9b04663</p>
རྦ་བཞེད་ (Ziling 2011) : Annotated Transcription
<p>རྦ་བཞེད་ (Ziling 2011) : Annotated Transcription.</p>
Chromatin fiber invasion and nucleosome displacement by the Rap1 transcription factor_Figure5e_2
<p>Raw microscopy movies for smFRET experiments with various chromatin templates for Mivelaz M., et al, 2019 <a href="https://doi.org/10.1016/j.molcel.2019.10.025">https://doi.org/10.1016/j.molcel.2019.10.025</a></p> <p>for Figure 5e (second part)</p> <p>see attached documentation for more details</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.