Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

21,320

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

21,320 results for “Transcript”

Learn how ShareScore rates datasets ↗
zenodo44/100

An Empirical Characterization of Event Sourced Systems and Their Schema Evolution - Lessons from Industry - Accompanying Anonymized Transcripts

<p>Anonymized interviews with 25 engineers on their experience applying Event Sourcing, with accompanying classifications.&nbsp;These transcripts are used in our publication &quot;An Empirical Characterization of Event Sourced Systems and Their Schema Evolution - Lessons from Industry&quot;.</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

Widespread Polycistronic Transcripts in Fungi Revealed by Single-Molecule mRNA Sequencing

<p>Genes in prokaryotic genomes are often arranged into clusters and co-transcribed into poly- cistronic RNAs. Isolated examples of polycistronic RNAs were also reported in some higher eukaryotes but their presence was generally considered rare. Here we developed a long- read sequencing strategy to identify polycistronic transcripts in several mushroom forming fungal species including Plicaturopsis crispa, Phanerochaete chrysosporium, Trametes ver- sicolor, and Gloeophyllum trabeum. We found genome-wide prevalence of polycistronic transcription in these Agaricomycetes, involving up to 8% of the transcribed genes. Unlike polycistronic mRNAs in prokaryotes, these co-transcribed genes are also independently transcribed. We show that polycistronic transcription may interfere with expression of the downstream tandem gene. Further comparative genomic analysis indicates that polycis- tronic transcription is conserved among a wide range of mushroom forming fungi. In sum- mary, our study revealed, for the first time, the genome prevalence of polycistronic transcription in a phylogenetic range of higher fungi. Furthermore, we systematically show that our long-read sequencing approach and combined bioinformatics pipeline is a generic powerful tool for precise characterization of complex transcriptomes that enables identifica- tion of mRNA isoforms not recovered via short-read assembly.</p>

opencc-by-4.0Nov 2016View details →
zenodo44/100

Transcription initiation peaks based on FANTOM5 CAGE data on hg38 and mm10

<p><strong>Overview</strong></p> <p>Decomposition-based peak identification (DPI, https://github.com/hkawaji/dpi1) is applied to the re-processed (re-aligned) FANTOM5 data, upon hg38 (GRCh38) and mm10 (GRCm38), obtained from below:</p> <ul> <li>http://fantom.gsc.riken.jp/5/datafiles/reprocessed/hg38_v1/basic/</li> <li>http://fantom.gsc.riken.jp/5/datafiles/reprocessed/mm10_v1/basic/</li> </ul> <p>The same parameters to the ones used in the previous paper (Forrest ARR, Kawaji H, Rehli M, et al. Nature 507: 462–470, 2014) was used.</p> <p> </p> <p><strong>Data files</strong></p> <p>Four data files per assembly are prepared as below.</p> <ol> <li>tag cluster in the original definition (*.tc.bed.gz)</li> <li>full set of DPI peaks (*.tc.decompose_smoothing_merged.bed.gz)</li> <li>permissive set of DPI peaks (*.tc.decompose_smoothing_merged.ctssMaxCounts3.bed.gz)</li> <li>robust set of DPI peaks (*.tc.decompose_smoothing_merged.ctssMaxCounts11_ctssMaxTpm1.bed.gz)</li> </ol> <p> </p> <p><strong>Acknowledgement</strong></p> <p>This data set is supported by Research Grant from MEXT to RIKEN Preventive Medicine and Diagnosis Innovation Program, RIKEN Center for Life Science Technologies, and JSPS KAKENHI Grant-in-Aid for Scientific Research No. 16H02902.</p>

opencc-by-4.0Apr 2017View details →
zenodo44/100

S1Data: ChIP-seq Data from Ferrie et. al. "p300 Is an Obligate Integrator of Combinatorial Transcription Factors Inputs"

<p>ChIP data from Ferrie et. al. "p300 Is an Obligate Integrator of Combinatorial Transcription Factors Inputs"</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

N20EMv2 dataset for automatic music transcription from multimodal singing

<p>N20EMv2 dataset for multimodal automatic music transcription from multimodal singing, presented in our TOMM 2024 paper, Automatic Lyric Transcription and Automatic Music Transcription from Multimodal Singing. This dataset contains recordings of two modalities: audio and video.&nbsp;</p> <p>Our paper is available at: https://dl.acm.org/doi/10.1145/3651310.</p> <p>Code is available at: https://github.com/guxm2021/SVT_SpeechBrain</p> <p>Please cite our work as:</p> <pre>@article{gu2024automatic, title={Automatic Lyric Transcription and Automatic Music Transcription from Multimodal Singing}, author={Gu, Xiangming and Ou, Longshen and Zeng, Wei and Zhang, Jianan and Wong, Nicholas and Wang, Ye}, journal={ACM Transactions on Multimedia Computing, Communications and Applications}, publisher={ACM New York, NY}, year={2024} }</pre>

opencc-by-sa-4.0Mar 2024View details →
zenodo44/100

Supplementary datasets for manuscript titled: Seasonal tissue-specific gene expression in wild crown-of-thorns starfish reveals reproductive and stress-related transcriptional systems

<p>Supplementary datasets for manuscript titled: Seasonal tissue-specific gene expression in wild crown-of-thorns starfish reveals reproductive and stress-related transcriptional systems</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Chinese Transcription of Buddhist Terms in the Late Hàn Dynasty - Dataset

<p>This dataset is a compilation of Chinese transcriptions of Buddhist terms produced by translators<br> from the late H&agrave;n period. It is a compilation of the previous works of (Coblin, 1983), (Karashima,<br> 2010), (Vetter, 2012), (Hill, Nattier, Granger, &amp; Kollmeier, 2020) for the Chinese transcriptions. To<br> these were added phonological reconstructions of the Chinese terms for late H&agrave;n from (Schuessler,<br> 2009) and Middle Chinese from (Baxter &amp; Sagart, 2014a), as well as the Gandhari equivalents of<br> Sanskrit and Pāli terms from (Baums &amp; Glass, 2002). This dataset, shared on Zenodo, aims at<br> being the new state-of-the-art dataset on Buddhist transcription material and can be used by anyone<br> working on H&agrave;n Chinese phonology and will help better understanding the possible language sources<br> of the Chinese transcriptions, as well as the phonology of the target Chinese dialects.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

The genomic and transcriptional landscape of primary central nervous system lymphoma

<p>Primary lymphomas of the central nervous system (PCNSL) are mainly diffuse large B-cell lymphomas (DLBCLs) confined to the central nervous system (CNS). Despite extensive research, the molecular alterations leading to PCNSL have not been fully elucidated. In order to provide a comprehensive description of the genomic and transcriptional landscape of PCNSL, we here performed whole-genome and transcriptome sequencing and integrative analysis of 51 lymphomas presenting in the CNS, including 42 EBV-negative PCNSL, 6 secondary CNS lymphomas (SCNSL) and 3 EBV+ CNSL and matched controls. The results were compared to an independent validation cohort of 31 FFPE CNSL specimens (PCNSL, n = 19; SCNSL, n = 9; EBV+ CNSL, n = 3) and 36 systemic DLBCL cases outside the CNS.</p> <p>This repository contains tab&nbsp;separated value text files:</p> <p>-&nbsp;Radke_et_al_supplementary_somatic_CNVs.tsv (somatic copy number variations predicted by ACEseq)<br> - Radke_et_al_supplementary_somatic_indels.tsv (somatic indels predicted by the DKFZ platypus workflow)<br> - Radke_et_al_supplementary_somatic_indels_exonic.tsv&nbsp;(somatic exonic indels predicted by the DKFZ platypus workflow)<br> - Radke_et_al_supplementary_somatic_mutations_integrated.tsv (table of gene by patients, stating which mutations were observed)&nbsp; &nbsp;&nbsp;<br> - Radke_et_al_supplementary_somatic_mutations_integr_integrated_including_kataegis_counts.tsv&nbsp;&nbsp; (table of gene by patients, stating which mutations were observed, including the count of mutations falling into kataegis hotspots)&nbsp; &nbsp;&nbsp;<br> - Radke_et_al_supplementary_somatic_SNVs.tsv&nbsp;(somatic SNVs predicted by the DKFZ mpileup workflow)<br> - Radke_et_al_supplementary_somatic_SNVs_exonic_functional.tsv&nbsp;(somatic exonic SNVs predicted by the DKFZ mpileup workflow)<br> - Radke_et_al_supplementary_somatic_SNVs_rescued_by_TiNDA&nbsp;(mutations initially classified as germline, but likely tumor mutations based on VAF modelling by TiNDA)<br> - Radke_et_al_supplementary_somatic_SVs.tsv (somatic structural variations predicted by the DKFZ Sophia workflow)<br> - Radke_et_al_supplementary_RNAseq_numReads_CNSLs.tsv (RNAseq read counts calculated by the DKFZ RNAseq workflow)</p> <p>This repository contains the raw unedited images from the manuscript:</p> <p>- Radke_et_al_Main_Figure_1c_BCL6.tif &nbsp;(raw unedited image for Main Figure 1c - BCL6)<br> - Radke_et_al_Main_Figure_1c_CD10.tif &nbsp;(raw unedited image for Main Figure 1c - CD10)<br> - Radke_et_al_Main_Figure_1c_MUM1.tif &nbsp;(raw unedited image for Main Figure 1c - MUM1)<br> - Radke_et_al_Supplementary_Figure_1a_CD20.tif &nbsp;(raw unedited image for Supplementary Figure 1a - CD20)<br> - Radke_et_al_Supplementary_Figure_1a_EBV.tif &nbsp;(raw unedited image for Supplementary Figure 1a - EBV)<br> - Radke_et_al_Supplementary_Figure_1a_HE_(FFPE).tif &nbsp;(raw unedited image for Supplementary Figure 1a - HE (FFPE))<br> - Radke_et_al_Supplementary_Figure_1a_HE_(frozen)_1.tif &nbsp;(raw unedited image for Supplementary Figure 1a - HE (frozen) 1)<br> - Radke_et_al_Supplementary_Figure_1a_HE_(frozen)_2.tif &nbsp;(raw unedited image for Supplementary Figure 1a - HE (frozen) 2)<br> - Radke_et_al_Supplementary_Figure_1a_HE_(frozen)_3.tif &nbsp;(raw unedited image for Supplementary Figure 1a - HE (frozen) 3)<br> - Radke_et_al_Supplementary_Figure_1a_Ki67.tif &nbsp;(raw unedited image for Supplementary Figure 1a - Ki67)<br> - Radke_et_al_Supplementary_Figure_1b_EBV_PCR.pptx &nbsp;(raw unedited image for Supplementary Figure 1b - EBV PCR&nbsp;)<br> - Radke_et_al_Supplementary_Figure_1c_CDKN2A_FISH_1.jpg &nbsp;(raw unedited image for Supplementary Figure 1c - CDKN2A FISH 1)<br> - Radke_et_al_Supplementary_Figure_1c_CDKN2A_FISH_2.jpg &nbsp;(raw unedited image for Supplementary Figure 1c - CDKN2A FISH 2)<br> - Radke_et_al_Supplementary_Figure_7h_PD-L1_LS-033.tif &nbsp;(raw unedited image for Supplementary Figure 7h - PD-L1 LS-033)<br> - Radke_et_al_Supplementary_Figure_7h_PD-L1_LS-031.tif &nbsp;(raw unedited image for Supplemnetary Figure 7h - PD-L1 LS-031)</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

FEDORA. Excerpts from essays, transcript of interviews and group discussions on students' future perception. Part 1: Essays, Finland.

<p><strong>Version 1.1.</strong></p> <p><strong>Updated from&nbsp;</strong>https://zenodo.org/record/5517595</p> <p><strong>Changes:&nbsp;</strong>added .csv copy of the dataset. Clarified the README below, and added name of publishing journal.&nbsp;No other changes.</p> <p>Added a FEDORA project README below.</p> <p>&nbsp;</p> <p><strong>Description of dataset:</strong></p> <p>This&nbsp;matrix, presented in two formats (.xlsx and .csv), contains an&nbsp;English-language dataset (translated from original&nbsp;Finnish). The data relate&nbsp;to a research article&nbsp;<em>Students&rsquo; technological images of the future: implications for science and technology education, </em>accepted to be published in European Journal of Futures Research.</p> <p>As per ethical concerns and participants&#39; consent, the dataset is given in a fully anonymised form. Here, excerpts from&nbsp;students&#39; essays&nbsp;(the context of which is given in the article) are given. The excerpts are the ones&nbsp;that have been used in analysis for the article identified above. Further details will be available in the published article.</p> <p>385 such excerpts are given, originating in&nbsp;57 essays in which upper-secondary&nbsp;students imagine the year 2035 or 2040 and the technological environment in which they would like to live at that time. The numbering was used to group codes for the analysis: type of technology (1), effect of technology (1E), and positive/negative framing (2A-C).</p> <p>The dataset is intended for providing transparency, but it may also be used for further research. Assistance may be available from the authors at reasonable request. Please note that the dataset presented here contains redundancies and a few additional codes that were not used in the analysis. The redundant quotations from the essays were not duplicated in the analysis, but were not removed from this spreadsheet export. Apologies for any inconvenience.</p> <p>To preserve full anonymity, students are not identified by any marker or pseudonym here; rather, the quotations are given alphabetically. The start and end of passages has not been checked for additional or missing first and last characters, as these can easily be inferred.</p> <p>The related research article gives a fuller description of the dataset and analysis.</p> <p>Please contact the corresponding author for more information.</p> <p>&nbsp;</p> <p>--</p> <p>&nbsp;</p> <p><a href="https://zenodo.org/communities/futuresthinking?page=1&amp;size=20">FEDORA Project</a>&nbsp;README:</p> <p>&nbsp;</p> <p><strong>README</strong></p> <p><strong>Data Set Title:</strong>&nbsp;&ldquo;FEDORA. Excerpts from essays, transcript of interviews and group discussions on students&rsquo; future perception. Finland&quot;</p> <p><strong>Data Set Author/s:</strong>&nbsp;Antti Laherto, Tapio Rasa,&nbsp;(University of Helsinki)</p> <p><strong>Data Set Contact Person/s</strong>: Tapio Rasa<strong>&nbsp;</strong>(University of Helsinki), ORCID 0000-0003-1315-5207, tapio.rasa@helsinki.fi;</p> <p><strong>Data Set License</strong>: this data set is distributed under the Creative Commons Attribution&nbsp;4.0 International (CC BY 4.0) license.</p> <p><strong>Publication Year</strong>: 2021</p> <p><strong>Project Info</strong>: FEDORA<strong>&nbsp;</strong>(Future-oriented Science EDucation to enhance Responsibility and engagement in the society of Acceleration and uncertainty<strong>&nbsp;,&nbsp;</strong>funded by European Union, Horizon 2020 Programme. Grant Agreement num.<strong>&nbsp;</strong>872841,<br> www.fedora-project.eu)</p> <p>&nbsp;</p> <p><strong>Data set Contents</strong></p> <p>The data set consists of:</p> <p>One spreadsheet file, provided in two alternative formats (CSV and XLSX).</p> <p>Students_images_of_technological_futures_DATA_Zenodo_csv.csv</p> <p>Students_images_of_technological_futures_DATA_Zenodo_xlsx.xlsx</p> <p>&nbsp;</p> <p><strong>Data set Documentation</strong></p> <p><em>Given above this README, on the ZENODO repository.&nbsp;https://zenodo.org/record/6397196</em></p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Data and code for the publication "DNA methylation underpins the epigenomic landscape regulating genome transcription in Arabidopsis"

<p>The zipped file of this repository contains code and data to reproduce the results of the publication:</p> <p>Zhao et al, DNA methylation underpins the epigenomic landscape regulating genome transcription in Arabidopsis. Genome Biology (2022).&nbsp;</p> <p>All sequence data have been deposited in NCBI GEO accession codes GSE183987 and&nbsp;GSE169497.</p> <p>&nbsp;</p> <p>Please see the README document for detailed:</p> <p>- Descriptions of the code and data provided</p> <p>- Lists of the required dependencies</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Toward a base-resolution panorama of the in vivo impact of cytosine methylation on transcription factor binding

<p>TF binding models built by JAMS (https://github.com/csglab/JAMS), ChIP-seq peak files (from ENCODE, Najafabadi et al. 2015, Schmitges et al. 2016, and Imbeault et al. 2017; called by MACS 1.4v),&nbsp;ChIP-seq pulldown and control tags from said peaks, input data for JAMS, and RCADE2 motifs for C2H2 zinc finger proteins.&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Dataset for the article "Spatially coherent diffusion of human RNA Pol II depends on transcriptional state rather than chromatin motion" by Roman Barth and Haitham Shaban

<p>The data set comprises all raw microscopy images and DFCC analyses as presented in&nbsp;</p> <p><strong>Spatially coherent diffusion of human RNA Pol II depends on transcriptional state rather than chromatin motion</strong></p> <p>by Roman Barth and Haitham Shaban, published in Nucleus (https://doi.org/10.1080/19491034.2022.2088988)</p> <p>There are two folders for RNAPII and DNA each, one for the raw images and one for the processed DFCC data, supplied as .mat files.</p> <p>Every folder contains three sub-folders containing the data for the conditions: +Serum, -Serum, and +DRB.</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Benchmarking tools for transcription factor prioritization

<p><strong>Abstract:</strong></p> <p>Spatiotemporal regulation of gene expression is controlled by transcription factor (TF) binding to regulatory elements, resulting in a plethora of cell types and cell states from the same genetic information.&nbsp; Due to the importance of regulatory elements, various sequencing methods have been developed to localise them in genomes, for example using ChIP-seq profiling of the histone mark H3K27ac that marks active regulatory regions. Moreover, multiple tools have been developed to predict TF binding to these regulatory elements based on DNA sequence. As altered gene expression is a hallmark of disease phenotypes, identifying TFs driving such gene expression programs is critical for the identification of novel drug targets.In this study, we curated 84 chromatin profiling experiments (H3K27ac ChIP-seq) where TFs were perturbed through e.g., genetic knockout or overexpression. We ran nine published tools to prioritize TFs using these real-world data sets and evaluated the performance of the methods in identifying the perturbed TFs. This allowed the nomination of three frontrunner tools, namely RcisTarget, MEIRLOP and monaLisa. Our analyses revealed opportunities and commonalities of tools that will help to guide further improvements and developments in the field.</p> <p><strong>Dataset description:</strong></p> <ul> <li>tf_tool_benchmark_atacseq_diffPeaks.tar.gz -Archive containing differential peak statistics, tool diff peak input files (fore- and background) for all currated ATAC-seq datasets.&nbsp;</li> <li>tf_tool_benchmark_h3K27ac_chipseq_diffPeaks.tar.gz - Archive containing differential peak statistics, tool diff peak input files (fore- and background) for all currated H3K27ac ChIP-seq datasets.&nbsp;</li> <li>tf_tool_benchmark_atacseq_results.tar.gz - Archive containing the raw tool results for each ATAC-seq dataset.</li> <li>tf_tool_benchmark_chipseq_results.tar.gz - Archive containing the raw tool results for each H3K27ac ChIP-seq dataset.</li> <li>tf_tool_benchmark_results.tar.gz - Archive containing tool results summary for plotting (rds files).</li> </ul> <p><strong>Contact:&nbsp; </strong>Sebastian Steinhauser - sebastian.steinhauser@novartis.com</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Frauen* im Fokus. Transcriptions and full texts of letters and works of women's rights activists

<p>In winter 2023/24 the Berlin State Library and Potsdam University (chair for Comparative Literature) organized the citizen science workshop&nbsp;<a href="https://lab.sbb.berlin/events/frauen-im-fokus/">"Frauen* im Fokus"</a> (women* in focus).</p> <p>Within the project 48 participants transcribed 85 letters and documents by 19th- and early 20th-century women's rights activists held in the collections of the State Library. The transcriptions provided by the participants were aligned with the digital images of the items using the <a href="https://ocr-bw.bib.uni-mannheim.de/escriptorium/" target="_blank" rel="noopener">eScriptorium</a> platform and software and checked for potential errors. From eScriptorium, the transcriptions were exported as ALTO and PAGE files. From the PAGE files, TEI/XML files were created with added metadata about the correspondence (where applicable).</p> <p>In addition to these MS sources, 55 printed works of the same activists were OCRed and are included in the dataset as PAGE and ALTO-files; these files were not manually checked for quality.</p> <p>We would like to thank our trainees Lilly Bucksteeg and Lilly Welz for their valuable contribution to the creation of the data set.</p> <p>The data set consists of five zip-files containing:</p> <ul> <li>the print sources in PAGE format</li> <li>the print sources in ALTO format</li> <li>the manuscript sources in PAGE format</li> <li>the manuscript sources in ALTO format</li> <li>the manuscript sources in TEI format</li> </ul> <p>------------------</p> <p>Im Wintersemester 2023/24 f&uuml;hrten die Staatsbibliothek zu Berlin und die Universit&auml;t Potsdam (Professur f&uuml;r Allgemeine und Vergleichende Literaturwissenschaft) das&nbsp;<a href="https://lab.sbb.berlin/events/frauen-im-fokus/">Projekt "Frauen* im Fokus"</a> durch.</p> <p>Im Rahmen des Projekts transkribierten 48 Personen 85 Briefe und andere Nachlassdokumente von Frauenrechtlerinnen des 19. und fr&uuml;hen 20. Jahrhunderts. Das Organisationsteam nahm auf der Plattform <a href="https://ocr-bw.bib.uni-mannheim.de/escriptorium/" target="_blank" rel="noopener">eScriptorium</a> eine teilweise automatisierte Layoutanalyse und Zeilensegmentierung der Digitalisate vor und f&uuml;gte nach erfolgter inhaltlicher Qualit&auml;tskontrolle die von den Teilnehmenden erstellten Transkriptionen dort ein, um sie dann als PAGE- und ALTO-Dateien zu exportieren; zus&auml;tzlich wurden aus den PAGE-Dateien TEI-Dateien der einzelnen Dokumente generiert, die (wo passend) mit Brief-Metadaten zu Absendern, Empf&auml;ngern und Orten angereichert wurden.</p> <p>Im Rahmen des Projekts wurden zudem 55 Druckwerke der Frauenrechtlerinnen aus dem Bestand der Staatsbibliothek als Volltexte erschlossen. Die Segmentierung und Volltexterkennung der Druckwerke erfolgte automatisch in eScriptorium ohne zus&auml;tzliche manuelle Qualit&auml;tskontrolle. Die Daten liegen exportiert in den Formaten PAGE und ALTO vor.</p> <p>Besonderer Dank geb&uuml;hrt Lilly Bucksteeg und Lilly Welz, die als Praktiantinnen im Projekt ma&szlig;geblich zur Erstellung des Datensets beigetragen haben.</p> <p>Das Datenset besteht aus f&uuml;nf zip-Dateien die folgende Dateien enthalten:</p> <ul> <li>die gedruckten Werke im PAGE-Format</li> <li>die gedruckten Werke im ALTO-Format</li> <li>die Manuskript-Transkriptionen im PAGE-Format</li> <li>die Manuskript-Transkriptionen im ALTO-Format</li> <li>die Manuskript-Transkriptionen im TEI-Format</li> </ul>

opencc-zeroApr 2024View details →
zenodo44/100

Anonymised transcriptions (local and translated versions) of 18 Focus Groups with RWPP voters in Spain, UK, Denmark, Germany, Hungary, Switzerland

<p><strong>Anonymised transcriptions (local and translated versions) of 18 Focus Groups with RWPP voters in Spain, UK, Denmark, Germany, Hungay, Switzerland</strong></p> <p>In the UNTWIST project, we have carried out a total of 18 focus groups in Denmark, Germany, Hungary, Spain, Switzerland and the United Kingdom. They explore RWPP voters&rsquo; subjective perceptions of their needs and demands, their horizon of expectations, and their level of &lsquo;gender fatigue&rsquo;. Groups&rsquo; design followed two minimum criteria: same-sex composition (with a minimum of two same sex -male and female- groups per country) and voting behaviour (current voters of RWPP who have previously voted for mainstream parties or abstained or have doubts about RWPP and mainstream or abstain in case of voting for the first time).</p> <p>&nbsp;The composition of the groups varied between 6 and 10 participants per group in all but one partner&rsquo;s country. In Denmark, all focus groups experienced dropouts. These unforeseen issues led to conducting the focus groups with fewer participants than was initially designed.</p> <p>&nbsp;In 83% of countries, the empirical composition of focus groups was considered and controlled for participants&rsquo; age, social class position and level of education.</p> <p>Finally, groups were same-sex moderated.</p> <p>Two comprised folders are provided. One contains the anonymised transcriptions of 18 Focus Groups carried out for WP2 of the UNTWIST project in their local languages. The other contains the IA-translated (Deepl) version of the same focus groups. Please note that the translations have not been human-supervised.&nbsp;</p> <p>FG_CHE_1 Female &nbsp; &nbsp;<br>Female Group, Switzerland</p> <p>FG_CHE_2 Male<br>Male Group, Switzerland &nbsp; &nbsp;</p> <p>FG_DEN_1 Female &nbsp; &nbsp;<br>Female Groups, Denmakr</p> <p>FG_DEN_2 Male<br>Male Group, Denmark &nbsp; &nbsp;</p> <p>FG_DEN_3 Male &nbsp; &nbsp;<br>Male Group, Denmark</p> <p>FG_DEN_4 Mixed &nbsp; &nbsp;<br>Mix Male and Female Group, Switzerland</p> <p>FG_ESP_1 Male<br>Male Group, Spain</p> <p>FG_ESP_2 Male<br>Male Group, Spain &nbsp; &nbsp;</p> <p>FG_ESP_3 Female &nbsp; &nbsp;<br>Female Group, Spain</p> <p>FG_ESP_4 Female<br>Female Group, Spain</p> <p>FG_GBR_1 Female &nbsp; &nbsp;<br>Female Group, UK</p> <p>FG_GBR_2 Male &nbsp; &nbsp;<br>Male Group, UK</p> <p>FG_GER_1 Female &nbsp; &nbsp;<br>Female Group, Germany</p> <p>FG_GER_2 Male<br>Male Group, Germany</p> <p>FG_HUN_1 Female &nbsp; &nbsp;<br>Female Group, Hungary</p> <p>FG_HUN_2 Female &nbsp; &nbsp;<br>Female Group, Hungary</p> <p>FG_HUN_3 Male<br>Male Group, Hungary &nbsp; &nbsp;</p> <p>FG_HUN_4 Male<br>Male Group, Hungary &nbsp; &nbsp;</p>

opencc-by-sa-4.0May 2024View details →
zenodo44/100

Evaluation of transcription factor knockout impact on paclitaxel response for Triple Negative Breast Cancer

<div>Data and code related to Zenodo repository: 10.5281/zenodo.11238552</div> <div>&nbsp;</div> <div>Two experimental formats included:</div> <div>'fixed' prefix: data from terminal time point of siRNA screen applied to HCC1143, HCC1806, and MDA-MB-468 Triple Negative Breast Cancer cell lines.</div> <div>'live' prefix: data from live-cell imaging of cell cycle reporter (HDHB-mClover/NLS-mCherry) HCC1143 Triple Negative Breast Cancer cell line.</div> <div>Note: 'live' level 1 data is available upon request (heiserl@ohsu.edu, calistri@ohsu.edu).</div> <div>&nbsp;</div> <div>Experimental goal:</div> <div>Evaluate whether siRNA knockdown of transcription factors elevated during paclitaxel response impact cell count, cell morphology or cycling dynamics.</div> <div>&nbsp;</div> <div>Methods:</div> <div>siRNA Knockdown: Cells were plated in 90ul of serum free media per well of a 96 well plate. 24 hours later, siRNA knockdown mixture was prepared using a cell-line optimized concentration of Lipofectamine RNAiMAX (cat 13778075-075, Invitrogen) and siRNA (Horizon Discovery ON-TARGETplus) following RNAiMAX recommended protocol. The final concentration of siRNA per well was 1pmol and the final volume of RNAiMAX per well was 75nL for HCC1143, and 37.5nL for HCC1806 or MDA-MB-468 in 100uL of cell containing volume. 24 hours after siRNA transfection cells were treated with an addition of 100uL complete media containing either DMSO vehicle control or paclitaxel.&nbsp;</div> <div>&nbsp;</div> <div>Fixed-cell assays: Cells were plated at 3000 cells in 100ul of complete media per well in a 96 well plate (#08-772-225, FisherScientific). After 24 hours, an additional 100ul of either vehicle (0.1% DMSO) or paclitaxel containing complete media was added. After 72 hours cells were fixed with 4% Formaldehyde (#28908, ThermoFisher Scientific) for 15 minutes at room temperature, then permeabilized with 0.3% Triton X-100 (#X100-100ML, Sigma Aldrich) for 10 minutes at room temperature, then washed twice with PBS. Fixed cells were then stained with 0.5ug/mL DAPI (4083S, Cell Signaling Technology) in PBS for 15 minutes at room temperature. Following DAPI staining, wells were washed once with PBS, then stained with 1:20,000 HCS CellMask Green in PBS (H32714, Invitrogen) for 15 minutes at room temperature. Wells were washed twice with room temperature PBS and then 4 fields of view per well imaged on an InCell 6000 (GE Healthcare). Images were segmented with two custom Cellpose models to segment the nucleus (using parameters: diameter = 50, chan = DAPI, chan2 = Cellmask Orange) and cytoplasm (using parameters: diameter = 90, chan = Cellmask Orange, chan2 = DAPI). Image quantification was performed in R (v4.3.1) using EBImage (v4.42.0), and cells were annotated based on the number of distinct nuclei segmented within each cytoplasmic mask.&nbsp;</div> <div>&nbsp;</div> <div>HDHB reporter live-cell assays: siRNA knockdown and drug treatment was performed as described above, and then the plate was loaded on an Incucyte S3 (Sartorious) and cells imaged every 15 minutes for 72 hours post drug treatment. At each timepoint 4 fields of view were captured at 20x magnificantion in each well using the phase, red and green channels. A cytoplasmic mask was computed from the mean of normalized red/green channel (cellpose parameters: diameter = 57, chan = mean(normalized(red), normalized(green)), and a nuclear mask was computed from the red channel (cellpose parameters: diameter = 30, chan = DAPI) using custom trained Cellpose models. Image quantification was performed in R (v4.3.1) using EBImage (v4.42.0). An additional perinuclear ring mask was computed as the 11 pixel dilation from the nuclear mask, but still bound by the cytoplasmic mask. To determine mClover localization thresholds for cell cycle assignment, 250 cell images were randomly selected and manually assigned to the G1, S/G2 or M cell cycle state based on mClover localization. The mClover intensity ratios were then used to determine thresholds for automated cell cycle phase calling which was applied to the rest of the data set (Supplemental Figure 5A). Mononuclear cells with a Perinuclear:Nuclear mean intensity ratio greater than 0.8 and Nuclear:Cytoplasmic total intensity less than 0.5 were assigned to the S/G2 phase. Mononuclear and Multinuclear cells with a Nuclear:Cytoplasmic total intensity ratio greater than 0.8 and Perinuclear:Nuclear mean intensity ratio less than 0.8 were assigned to the &lsquo;M&rsquo; phase. The remainder of mononuclear cells were assigned &lsquo;G1&rsquo;, and the remainder of multinucleated cells were assigned &lsquo;Multinucleated&rsquo;.&nbsp;</div> <div>&nbsp;</div> <div>Included files:</div> <div>fixed_level_1-plate_#.zip : Six .zip archives containing the raw images (DAPI/CellMask/Brightfield) from fixed-cell experiments.</div> <div>plate 1: HCC1143 cells treated with plate A schema</div> <div>plate 2: HCC1143 cells treated with plate B schema</div> <div>plate 3: HCC1806 cells treated with plate A schema</div> <div>plate 4: HCC1806 cells treated with plate B schema</div> <div>plate 5: MDA-MB-468 cells treated with plate A schema</div> <div>plate 6: MDA-MB-468 cells treated with plate B schema</div> <div>fixed_level_2: Data quantified from cellpose masks at the single-nuclei level (redundant cytoplasm information)</div> <div>fixed_level_3: Data from 'fixed_level_2.csv' collapsed to the single cell level, including staining intensity and aggregate nuclear information</div> <div>fixed_incell_to_cellpose.rmd: R markdown code for converting original incell files (fixed_level_1) to RGB images for cellpose segmentation</div> <div>fixed_image_quantification.rmd: R markdown code for quantifying images using cellpose segmentation masks and original images (fixed_level_1)</div> <div>fixed_cellpose_models.zip: Archive including cellpose models used for fixed experiment</div> <div>live_level_2: Data quantified from cellpose masks at the single-nuclei level (redundant cytoplasm information)</div> <div>live_level_3: Data from 'live_level_2.csv' collapsed to the single cell level, including staining intensity and aggregate nuclear information</div> <div>live_level_4: Data from 'live_level_3.csv' collapsed to the single condition level summarizing the number, multinucleation status and phase of cells at each time point.</div> <div>live_image_quantification.rmd: R markdown code for quantifying images using cellpose segmentation masks and original images (live_level_1).</div> <div>l ive_incu_archive2rgb.rmd: R markdown code for converting incucyte archive formatted data into RGB images, where the blue channel is the arithmetic mean of the min-max (0-1) normalized red and green channels.</div> <div>live_cellpose_models.zip: Archive including cellpose models used for live experiment.</div> <div>&nbsp;</div> <div>&nbsp;</div>

opencc-by-4.0May 2024View details →
zenodo44/100

FONDUE-FR-PRINT-17 - Transcriptions of French 17th c. prints

<p>HTR Groundtruth for French 17th c. prints, produced with&nbsp;<a href="https://github.com/mittagessen/kraken">Kraken</a> and <a href="https://gitlab.com/scripta/escriptorium">eScriptorium</a>.</p> <p>Original data is available on <a href="https://github.com/FoNDUE-HTR/FONDUE-FR-PRINT-17">GitHub</a>.</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Dataset of "Single-Cell RNA-Seq Reveals Transcriptional Heterogeneity in Latent and Reactivated HIV-infected Cells"

<p><strong>Detailed quantitative analysis of GFP expression in SAHA and TCR-treated cells &amp; Computational analysis of&nbsp;bulk and single-cell RNA-Seq data.</strong></p> <p>&nbsp;</p> <p><em><strong>Detailed quantitative analysis of GFP expression in SAHA and TCR-treated cells.</strong></em></p> <p>Cells were prepared for single cell analysis at the Genome Technology Facility (GTF) of the University of Lausanne. Cells were loaded on Fluidigm C1 IFC plates (5-10 &mu;m), with run ID smart33, smart34 and smart35, corresponding to untreated, SAHA- and TCR-treated conditions respectively. After single cell capture on the Fluidigm C1 IFC plate, each chamber was inspected visually by microscopy and pictures were captured with a Zeiss Axiovert 200 M fluorescence microscope equipped with a Roper Scientific CoolSnap HQ camera using a Plan-Neofluar 10X lens (smart34 run) or 20X lens (for smart35 run). For each capture chamber, pictures in bright field and FITC channel were taken with the MetaMorph 6.3 software. Picture analysis was then performed using ImageJ 1.50b software (open access software: website). Brightness and contrast were adjusted for qualitative assessment of the pictures.</p> <p><em><strong>Computational analysis of&nbsp;bulk and single-cell RNA-Seq data.</strong></em></p> <p>Upon bulk or single cell isolation, RNA extraction and library preparation was performed according to Illumina protocols. Bulk and single-cell RNA-Seq data analysis are detailed here.</p> <p>&nbsp;</p> <p>Linked to the paper published in Cell Reports (doi:10.1016/j.celrep.2018.03.102):&nbsp;</p> <p><strong>Single-Cell RNA-Seq Reveals Transcriptional Heterogeneity&nbsp;in Latent and Reactivated HIV-infected Cells</strong></p> <p>Despite effective treatment, HIV can persist in latent reservoirs, which represent a major obstacle towards HIV eradication. Targeting and reactivating latent cells is challenging due to the heterogeneous nature of HIV infected cells. Here, we used a primary model of HIV latency and single-cell RNA sequencing to characterize transcriptional heterogeneity during HIV latency and reactivation. Our analysis identified transcriptional programs leading to successful reactivation of HIV expression.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2018View details →
zenodo44/100

Virtual ChIP-seq predictions of binding of 36 transcription factor in Roadmap Epigenomics Project tissues

<p>This dataset contains predictions of Virtual ChIP-seq for binding of 36&nbsp;transcription factors in Roadmap Epigenomics dataset tissues with matched DNase-seq and RNA-seq data.</p> <p>Tarball contains subfolders for each of the 36&nbsp;TFs where Virtual ChIP-seq median MCC&nbsp;in validation cell types was &gt; 0.3.</p> <p>Each subfolder contains gzipped BED files. Each file is named as &lt;Tissue&gt;_&lt;Age&gt;_&lt;TF&gt;_&lt;Accession&gt;_Predictions.bed.gz. Columns correspond to Chromosome, Start, End,&nbsp;&lt;Tissue&gt;_&lt;Age&gt;_&lt;TF&gt;_&lt;Accession&gt;, Posterior probability</p> <p>You can use the posterior probabilities provided in Virchip_PosteriorCutoffs_V3.0.0.tsv. These are posterior probability cutoffs which maximized MCC in H1-hESC cell type, or are set to 0.4 if there was no ChIP-seq data of that TF in H1-hESC (0.4 is the mode of all optimal posterior probability cutoffs in H1-hESC).</p>

opencc-zeroOct 2018View details →
zenodo44/100

Counting Words That Count: NLP for exploring Romanian Parliament Transcripts

<p>The data is obtained by scraping the cdep.ro website and contains 500k+ instances of speech from the parliament podium from 1996 to 2019. (Up to 2001 only the Chamber of Deputies published transcripts, after jan. 2001&nbsp;Senate data is also included.)&nbsp;<br> <br> Columns:&nbsp;</p> <p>&#39;index&#39; - incremented integer as row number in order of scraping</p> <p>&#39;title&#39;, - title of the scraped page, usually contains the name of the chamber and the exact data</p> <p>&#39;name&#39;, - the name of the speaker, preappended with Mr. or Mrs.&nbsp;</p> <p>&#39;speech&#39;, - the content of the speech,&nbsp;&nbsp;</p> <p>&#39;gender&#39;, - the gender of the speaker</p> <p>&#39;url&#39; - the url to the profile of the speaker (useful for extending the data)</p> <p>&nbsp;</p> <p>CDEPs2.csv - Contains all transcripts, prone to parsing errors. 100% of data.</p> <p>validated-1.csv - Consists of 99% of original data. Less than 1% dropped for convenience. Ready to use.</p>

opencc-by-4.0Jul 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record