Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
21,320
datasets available to search
ShareScore release 0.9.0
Dataset results
21,320 results for “Transcription”
Chromatin fiber invasion and nucleosome displacement by the Rap1 transcription factor_Figure5f_1
<p>Raw microscopy movies for smFRET experiments with various chromatin templates for Mivelaz M., et al, 2019 <a href="https://doi.org/10.1016/j.molcel.2019.10.025">https://doi.org/10.1016/j.molcel.2019.10.025</a></p> <p>for Figure 5f (first part)</p> <p>see attached documentation for more details</p>
Chromatin fiber invasion and nucleosome displacement by the Rap1 transcription factor_Figure5f_2
<p>Raw microscopy movies for smFRET experiments with various chromatin templates for Mivelaz M., et al, 2019 <a href="https://doi.org/10.1016/j.molcel.2019.10.025">https://doi.org/10.1016/j.molcel.2019.10.025</a></p> <p>for Figure 5f (second part)</p> <p>see attached documentation for more details</p>
Chromatin fiber invasion and nucleosome displacement by the Rap1 transcription factor_FigureS6
<p>Raw microscopy movies for smFRET experiments with various chromatin templates for Mivelaz M., et al, 2019 <a href="https://doi.org/10.1016/j.molcel.2019.10.025">https://doi.org/10.1016/j.molcel.2019.10.025</a></p> <p>for FigureS6</p> <p>see attached documentation for more details</p>
Chromatin fiber invasion and nucleosome displacement by the Rap1 transcription factor_Figure5e_1
<p>Raw microscopy movies for smFRET experiments with various chromatin templates for Mivelaz M., et al, 2019</p> <p><a href="https://doi.org/10.1016/j.molcel.2019.10.025">https://doi.org/10.1016/j.molcel.2019.10.025</a></p> <p> </p> <p>for Figure 5e (first part)</p> <p>see attached documentation for more details</p>
Chromatin fiber invasion and nucleosome displacement by the Rap1 transcription factor_Fig.2,3,S3,S4,S7
<p>Raw microscopy movies for colocalization TIRF experiments (Rap1 binding) with various chromatin templates for Mivelaz M., et al, 2019 (<a href="https://doi.org/10.1016/j.molcel.2019.10.025">https://doi.org/10.1016/j.molcel.2019.10.025</a>)</p> <p>for Figures Fig.2,3,S3,S4,S7</p> <p>see attached documentation for more details</p>
Bigwig files for paper "STK19 is a transcription-coupled repair factor that participates in UVSSA ubiquitination and TFIIH loading"
<p>Bigwig files for paper "STK19 is a transcription-coupled repair factor that participates in UVSSA ubiquitination and TFIIH loading". </p>
Assessment of changes in circRNA expression based on transcripts of genes encoding ADAMTS proteins in patients with non-small cell lung carcinoma compared to normal tissue
Open the record for dataset details and reuse information.
Benchmarking Illumina RNA-seq fusion transcript detection methods - cancer cell lines RNA-seq
<p>Cancer cell line RNA-seq data (reads or names of reads from CCLE data) used for benchmarking Illumina-based fusion detection methods as used in:</p> <p>Haas, B.J., Dobin, A., Li, B. <em>et al.</em> Accuracy assessment of fusion transcript detection via read-mapping and de novo fusion transcript assembly-based methods. <em>Genome Biol</em> <strong>20</strong>, 213 (2019). https://doi.org/10.1186/s13059-019-1842-9</p> <p> </p> <p>For CCLE data, direct sharing of fastq files was not possible. CCLE data must be obtained from:</p> <p> https://portals.broadinstitute.org/ccle/home</p> <p>Instead, the identifiers for the reads leveraged as part of our study are made available, and these reads can be extracted from the CCLE fastq files directly once obtained from the primary source.</p> <p><br>For the non-CCLE data, the exact reads leveraged by our study are made directly available here in fastq format.</p>
Transcriptional stochasticity reveals multiple mechanisms of long non-coding RNA regulation at the Xist – Tsix locus
<p>Data and codes for Figure reprodicbility, image processing and demo.</p>
Decoding the Epigenetics and Chromatin Loop Dynamics of Androgen Receptor-Mediated Transcription
<p>This repository stores the datasets for the "Decoding the Dynamic Regulation and Chromatin Architecture of Androgen Receptor-mediated Gene Expression" paper.</p>
GRAS Family Transcription Factor Binding Behaviors in Sorghum bicolor, Oryza, and Maize
<p>Supplemental Data Files for the manuscript entitled "GRAS Family Transcription Factor Binding Behaviors in Sorghum bicolor, Oyrza, and Maize".</p>
Specifying cellular context of transcription factor regulons for exploring context-specific gene regulation programs
<p>This repository contains the raw and processed files used in Minaeva et al. 2024.</p> <p>In this version, we have revised the regulon construction pipeline and expanded the dataset to cover 40 common cell lines.</p> <p>The code used to generate these files is available at <a href="https://github.com/LappalainenLab/chip_seq_regulons" target="_new" rel="noreferrer">GitHub - LappalainenLab/chip_seq_regulons</a>.</p> <p>The descriptions of the files contained within each subdirectory are as follows:</p> <h3>1-dataset_stats</h3> <ul> <li><code>per_gene_stats_{approach}_{cell_line}.tsv</code>: Number of TFs regulating a gene according to the respective approach (S2Mb, M2Kb, or S2Kb) in a given cell line.</li> <li><code>per_tf_stats_{approach}_{cell_line}.tsv</code>: Number of target genes regulated by a TF according to the respective approach (S2Mb, M2Kb, or S2Kb) in a given cell line.</li> </ul> <h3>1-network_enrichment</h3> <ul> <li><code>enrich_scores_remap_all_tfs_K562.tsv</code>: Results of fitting logistic regression for testing the enrichment of the K562 regulon in other biological networks (PPI, coexpression, experimental trans-networks).</li> </ul> <h3>2-plot_decoupler_comparison_benchmark_across_cells</h3> <ul> <li><code>{cell_line}_comparison_benchmark.tsv</code>: Results of benchmarking S2Mb, M2Kb, CollecTri, Dorothea, ChIP-Atlas, RegNet, and TRRUST regulons using the decoupler package and the KnockTF database. Cell lines considered are K562, HepG2, and MCF7 (see Methods for benchmarking pipeline details).</li> </ul> <h3>2-plot_decoupler_filter_benchmark_across_methods</h3> <ul> <li><code>{cell_line}_filtering_benchmark.tsv</code>: Results of benchmarking S2Mb, M2Kb, and S2Kb regulons with different filters applied using the decoupler package and the KnockTF database. Cell lines considered are K562, HepG2, and MCF7 (see Methods for benchmarking pipeline details).</li> </ul> <h3>3-tf_activity</h3> <ul> <li><code>aml_k562_activity_{regulon}_sc.tsv</code>: Results of TF activity analysis based on a respective regulon between healthy hematopoietic stem cells (HSCs) and abnormal AML progenitor cells following the decoupler pipeline. Regulons considered are K562-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>aml_activity_estimates_hsc_sc.tsv</code>: Summary of the TF activity analysis for statistically significantly dysregulated TFs between healthy HSCs and abnormal AML progenitor cells across regulons.</li> <li><code>aml_dhsc_ahsc_activity_{regulon}_sc.tsv</code>: Results of TF activity analysis based on a respective regulon between leukemic activated and dormant HSCs following the decoupler pipeline. Regulons considered are K562-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>aml_activity_estimates_dhsc_ahsc_sc.tsv</code>: Summary of the TF activity analysis for statistically significantly dysregulated TFs between leukemic activated and dormant HSCs across regulons.</li> <li><code>bc_bas_activity_{regulon}.tsv</code>: Results of TF activity analysis based on a respective regulon between healthy epithelial breast cells and malignant epithelial cells from basal breast cancer following the decoupler pipeline. Regulons considered are MCF7-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>bc_activity_estimates_bas.tsv</code>: Summary of the TF activity analysis for statistically significantly dysregulated TFs between healthy epithelial breast cells and malignant epithelial cells from basal breast cancer across regulons.</li> <li><code>bc_lum_activity_{regulon}.tsv</code>: Results of TF activity analysis based on a respective regulon between healthy epithelial breast cells and malignant epithelial cells from luminal type A breast cancer following the decoupler pipeline. Regulons considered are MCF7-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>bc_activity_estimates_lum.tsv</code>: Summary of the TF activity analysis for statistically significantly dysregulated TFs between healthy epithelial breast cells and malignant epithelial cells from luminal type A breast cancer across regulons.</li> <li><code>hep_activity_{regulon}.tsv</code>: Results of TF activity analysis based on a respective regulon between neoplastic and healthy liver cells following the decoupler pipeline. Regulons considered are HepG2-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>hep_activity_estimates.tsv</code>: Summary of the TF activity analysis for statistically significantly dysregulated TFs between neoplastic and healthy liver cells across regulons.</li> </ul> <h3>3-tf_disease_enrichment</h3> <ul> <li><code>aml_{database}_enrich_{regulon}_hsc_sc.tsv</code>: Results of enrichment analysis of dysregulated TFs identified based on a respective regulon between healthy HSCs and abnormal AML progenitor cells following the decoupler pipeline. Databases considered are COSMIC, DisGeNet, OMIM, and KEGG. Regulons considered are K562-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>aml_{database}_enrich_{regulon}_dhsc_ahsc_sc.tsv</code>: Results of enrichment analysis of dysregulated TFs identified based on a respective regulon between leukemic activated and dormant HSCs following the decoupler pipeline. Databases considered are COSMIC, DisGeNet, OMIM, and KEGG. Regulons considered are K562-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>bc_{database}_enrich_{regulon}_bas.tsv</code>: Results of enrichment analysis of dysregulated TFs identified based on a respective regulon between healthy epithelial breast cells and malignant epithelial cells from basal breast cancer following the decoupler pipeline. Databases considered are COSMIC, DisGeNet, and OMIM. Regulons considered are MCF7-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>bc_{database}_enrich_{regulon}_lum.tsv</code>: Results of enrichment analysis of dysregulated TFs identified based on a respective regulon between healthy epithelial breast cells and malignant epithelial cells from luminal type A breast cancer following the decoupler pipeline. Databases considered are COSMIC, DisGeNet, and OMIM. Regulons considered are MCF7-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>hep_{database}_enrich_{regulon}.tsv</code>: Results of enrichment analysis of dysregulated TFs identified based on a respective regulon between neoplastic and healthy liver cells following the decoupler pipeline. Databases considered are COSMIC, DisGeNet, OMIM, and KEGG. Regulons considered are HepG2-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> </ul> <h3>regulons</h3> <ul> <li><code>{cell_line}_regulon.tsv</code>: S2Mb, M2Kb, and S2Kb regulons generated in this study with all acquired annotations (see Methods for details).</li> </ul> <p>External regulons used for comparison. Cell lines considered are K562, HepG2, MCF7, and GM12878:</p> <ul> <li><code>ChIP-Atlas_target_genes_{cell_line}.tsv</code>: Customized ChIP-Atlas regulons (see Methods for details).</li> <li><code>Revised_Supplemental_Table_S3_Normal.csv</code>: Dorothea regulon collected from supplementary materials of Garcia-Alonso et al. (2019).</li> </ul> <h3>s3-network_enrichment</h3> <ul> <li><code>enrich_scores_remap_all_tfs_{cell_line}.tsv</code>: Results of fitting logistic regression for testing the enrichment of cell-line-specific regulons in PPI networks (see Methods and corresponding GitHub repository for details). Cell lines considered are K562, HepG2, MCF7, and GM12878.</li> </ul> <p> </p>
Towards Musically Informed Evaluation of Piano Transcription Models
<p>We provide here the evaluation set employed in our experiments described in "Towards Musically Informed Evaluation of Piano Transcription Models", published in the Proceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR), San Francisco, United States, 2024.</p> <p>In this work, we demonstrate musically informed piano transcription metrics using transcriptions derived from three state-of-the-art transcriptions ([1], [2], [3]). To this end, we create an evaluation set that includes (1) a subset of the original audio recordings from the MAESTRO dataset [1], (2) a re-recorded version that subset, and (3) a perturbed version of recordings from both (1) and (2). In this data repository, we provide components (2) and (3).</p> <p>[1] Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse Engel, and Douglas Eck, “Enabling factorized piano music modeling and generation with the MAESTRO dataset,” in International Conference on Learning Representations, 2019. </p> <p>[2] Qiuqiang Kong, Bochen Li, Xuchen Song, Yuan Wan, and Yuxan Wang, “High-resolution piano transcription with pedals by regressing onset and offset times,” IEEE/ACM Transactions on Audio, Speech and Language Processing, vol. 29, pp. 3707–3717, 2021. </p> <p>[3] Curtis Hawthorne, Ian Simon, Rigel Swavely, Ethan Manilow, and Jesse Engel. “Sequence-to-sequence piano transcription with transformers,” in Proceedings of the 22nd International Society for Music Information Retrieval Conference, ISMIR 2021.</p>
Gene, Splice and Transcript based QTL association in the human sACC and Amygdala
Open the record for dataset details and reuse information.
Transcripts (+attributions) from a mixed-presence user study with two wall-sized displays
<p>Transcripts from a mixed-presence experiment with two wall-sized displays.<br>Automatic transcription (and translation when necessary) using Whisper large v3. Resulting sentences were then attributed to individuals.<br>Comes from a study ran in Q4 2023. Accompanies a paper.</p> <p>As for the abbreviations/acronyms used in the file:</p> <ul> <li>Conditions C0 and C1 correspond respectively to "no cues" and "cues enabled" (see paper)</li> <li>Sides A and V respectively correspond to Arena (=circular display) and Viswall (flat display)</li> <li>Speakers CEO, FIN, ICU and LOG respectively correspond to the roles given to these speakers (i.e. CEO, Head of Finance, Head of Intensive Care Unit, and Head of Logistics)</li> <li>Speakers TECH_VIZWALL and FACILITATOR_VIZWALL correspond to members of the research team (see protocol)</li> </ul>
A machine-readable provisional transcription of the Piers Plowman text of Takamiya MS 23
<p>A machine-readable provisional transcription of the Middle English poem <em>Piers Plowman</em> as transmitted in New Haven, Beinecke Library, MS Takamiya 23. See the <a title="Technical Introduction" href="https://github.com/icornelius/s-takamiya-23/releases/latest/download/documentation.pdf">documentation</a>.</p>
Forsthofel-Lab/Identification of EV-responsive transcripts in the planarian Schmidtea mediterranea
Open the record for dataset details and reuse information.
TABLE 1 in High-quality herbarium-label transcription by citizen scientists improves taxonomic and spatial representation of the tropical plant family Annonaceae
<p>TABLE 1. — Species per dataset and continent.</p><table><tbody><tr><th><b>Region</b></th><th><b>GBIF</b></th><th><b>Herbonautes</b></th><th><b>Herbonautes and not</b> <b>found in GBIF</b></th></tr></tbody><tbody><tr><th>Americas</th><td>700</td><td>248</td><td>7</td></tr><tr><th>Africa</th><td>268</td><td>189</td><td>17</td></tr><tr><th>Madagascar</th><td>79</td><td>79</td><td>10</td></tr><tr><th>Asia and Oceania</th><td>621</td><td>492</td><td>125</td></tr></tbody></table>
Enhancer-promoter interactions and transcription are maintained upon acute loss of CTCF, cohesin, WAPL, and YY1
<p>1. spaSPT data for endogenously tagging HaloTag-YY1 in mESCs (clones YN11 and YN31).</p> <p>2. spaSPT data for HaloTag-YY1 in CTCF- (clone CD1) or RAD21- (clone RD35) depleted mESCs.</p> <p>3. sptSPT data for stably expressing HaloTag-YY1 in clone PBYN2 and YY1-HaloTag in clone PBYC3.</p> <p> </p> <div> </div>
Automatic translation and multilingual cultural heritage retrieval: a case study with transcriptions in Europeana (dataset)
<p>The dataset contains all the data required to reproduce the experiments done in the paper "Automatic translation and multilingual cultural heritage retrieval: a case study with transcriptions in Europeana", published in the 25th International Conference on Theory and Practice of Digital Libraries (<a href="http://www.tpdl.eu/tpdl2021/">TPDL'21</a>). In that work we run an experiment using the Europeana CH digital library as a use case, and we evaluated the effectiveness of a multilingual information retrieval strategy using machine translations to English as pivot language. We used the CEF translation service (eTranslation) for the translation of queries and content to English (<a href="https://ec.europa.eu/cefdigital/wiki/display/CEFDIGITAL/eTranslation">https://ec.europa.eu/cefdigital/wiki/display/CEFDIGITAL/eTranslation</a>).</p> <p>The dataset is also available at <a href="https://rnd-2.eanadev.org/share/crosslingual-search/">https://rnd-2.eanadev.org/share/crosslingual-search/</a>, and it is organized in four main folders:</p> <ul> <li><strong>queries</strong>: sample of 68 queries and their translations to English. The queries were issued in languages other than English from the Europeana Portal, using the Europeana’s 1914-1918 thematic collection, between January and August 2019.</li> <li><strong>transcriptions</strong>: sample of 18,257 handwriting transcriptions and its translations to English. The transcriptions are taken from the Europeana 1914-1918 thematic collection, and obtained from the Transcribathon crowdsourcing platform (https://europeana.transcribathon.eu/).</li> <li><strong>solr_configuration</strong>: Apache Solr search engine configuration used in the experiments (which replicates the one used in Europeana).</li> <li><strong>results</strong>: manual evaluation of the query translations, and automatic evaluation of the multilingual retrieval.</li> </ul> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.