Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

21,320

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

21,320 results for “Transcript”

Learn how ShareScore rates datasets ↗
zenodo44/100

BaRTv1.0: an improved barley reference transcript dataset to determine accurate changes in the barley transcriptome using RNA-seq

<p>Background<br> Time consuming computational assembly and quantification of gene expression and splicing analysis from RNA-seq data vary considerably. Recent fast non-alignment tools such as Kallisto and Salmon overcome these problems, but these tools require a high quality, comprehensive reference transcripts dataset (RTD), which are rarely available in plants.</p> <p>Results<br> A high-quality, non-redundant barley gene RTD and database (Barley Reference Transcripts &ndash; BaRTv1.0) has been generated. BaRTv1.0, was constructed from a range of tissues, cultivars and abiotic treatments and transcripts assembled and aligned to the barley cv. Morex reference genome (Mascher et al., 2017). Full-length cDNAs from the barley variety Haruna nijo (Matsumoto et al., 2011) determined transcript coverage, and high-resolution RT-PCR validated alternatively spliced (AS) transcripts of 86 genes in five different organs and tissue. These methods were used as benchmarks to select an optimal barley RTD. BaRTv1.0-Quantification of Alternatively Spliced Isoforms (QUASI) was also made to overcome inaccurate quantification due to variation in 5&rsquo; and 3&rsquo; UTR ends of transcripts. BaRTv1.0-QUASI was used for accurate transcript quantification of RNA-seq data of five barley organs/tissues. This analysis identified 20,972 significant differentially expressed genes, 2,791 differentially alternatively spliced genes and 2,768 transcripts with differential transcript usage.</p> <p>Conclusion<br> A high confidence barley reference transcript dataset consisting of 60,444 genes with 177,240 transcripts has been generated. Compared to current barley transcripts, BaRTv1.0 transcripts are generally longer, have less fragmentation and improved gene models that are well supported by splice junction reads. Precise transcript quantification using BaRTv1.0 allows routine analysis of gene expression and AS.</p>

opencc-by-4.0Dec 2018View details →
zenodo44/100

CLDF dataset on Panoan Languages in Standardized Transcription derived from Key and Comrie's "Intercontinental Dictionary Series" from 2023

<p>Cite the source of the dataset as:</p> <blockquote> <p>Miller, J. and List, J.-M. (2024): Providing standardized phonetic transcriptions for the Panoan languages in the Intercontinental Dictionary Series. Computer-Assisted Language Comparison in Practice 7.2. URL: https://calc.hypotheses.org/7503.</p> </blockquote>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Nuremberg Letterbooks: A Multi-Transcriptional Dataset of Early 15th Century Manuscripts for Document Analysis

<p>This dataset contains the images and labels of the Nuremberg Letterbooks dataset.</p> <p>It consists of four books (books 2 - 5) with line-wise transcriptions. Three kinds of transcriptions are reported: basic, regularized, and diplomatic, with additional expanded abbreviations.&nbsp;</p> <p>Code templates for text verification and writer verification are available at:</p> <ul> <li><a href="https://github.com/M4rt1nM4yr/letterbooks_text_verification">https://github.com/M4rt1nM4yr/letterbooks_text_verification</a></li> <li><a href="https://github.com/M4rt1nM4yr/letterbooks_writer_verification">https://github.com/M4rt1nM4yr/letterbooks_writer_verification</a></li> </ul> <p>When using this dataset, please cite:&nbsp;<br>M. Mayr, J. Krenz, K. Neumeier, A. Bub, S. B&uuml;rcky, N. Brolich, K. Herbers, M. Habermann, P. Fleischmann, A. Maier, and V. Christlein<em>.</em> <br>Nuremberg Letterbooks: A Multi-Transcriptional Dataset of Early 15th Century Manuscripts for Document Analysis. <em>Sci Data</em> <strong>12</strong>, 811 (2025).<br><a href="https://doi.org/10.1038/s41597-025-05144-z">https://doi.org/10.1038/s41597-025-05144-z</a></p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Direct molecular evidence for an ancient, conserved developmental toolkit controlling post-transcriptional gene regulation in land plants

<p>In plants, miRNA production is orchestrated by a suite of proteins that control transcription of the pri-miRNA gene, post-transcriptional processing and nuclear export of the mature miRNA. Post-transcriptional processing of miRNAs is controlled by a pair of physically-interacting proteins, HYL1 and DCL1. However, the evolutionary history and structural basis of the HYL1-DCL1 interaction is unknown. Here we use ancestral sequence reconstruction and functional characterization of ancestral HYL1 <em>in vitro</em> and in <em>Arabidopsis thaliana </em>to better understand the origin and evolution of the HYL1-DCL1 interaction and its impact on miRNA production and plant development. We found the ancestral plant HYL1 evolved high affinity for both double-stranded RNA (dsRNA) and its DCL1 partner before the divergence of mosses from seed plants (~500 Ma), and these high-affinity interactions remained largely conserved throughout plant evolutionary history. Structural modeling and molecular binding experiments suggest that the second of two double-stranded RNA-binding motifs (DSRMs) in HYL1 may interact tightly with the first of two C-terminal DCL1 DSRMs to mediate the HYL1-DCL1 physical interaction necessary for efficient miRNA production. Transgenic expression of the nearly 200 Ma-old ancestral flowering-plant HYL1 in <em>A. thaliana</em> was sufficient to rescue many key aspects of plant development disrupted by HYL1<sup>-</sup> knockout and restored near-native miRNA production, suggesting that the functional partnership of HYL1-DCL1 originated very early in and was strongly conserved throughout the evolutionary history of terrestrial plants. Overall, our results are consistent with a model in which miRNA-based gene regulation evolved as part of a conserved plant &lsquo;developmental toolkit&rsquo;.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Generation of transcriptional novelty by transposable element insertions in Arabidopsis, Genome Sequencing and eccDNA Data

<p><strong>Raw Illumina sequencing data from the Manuscript entitled &quot;Generation of transcriptional novelty by transposable element insertions in Arabidopsis&quot;</strong></p> <p><strong>A. Illumina genome sequencing reads of Arabidopsis control and hcLines that contain novel transposable element insertions.</strong></p> <p>To identify the genomic position of the new <em>ONSEN</em> insertions, the extracted DNA of the 11 selected lines (nine lines with new insertions and two control lines) was sent to BGI, Hong-Kong for Illumina paired-end 150 bp sequencing, aiming for a minimum of 20X sequencing coverage. Quality control of the raw reads was done using FastQC (Andrews S. (2010). FastQC: a quality control tool for high throughput sequence data. Available online at: <a href="http://www.bioinformatics.babraham.ac.uk/projects/fastqc">http://www.bioinformatics.babraham.ac.uk/projects/fastqc</a>) and trimming/clipping was done using Trimmomatic with parameters ILLUMINACLIP: TruSeq3:2:30:10 LEADING:20 TRAILING:20 SLIDINGWINDOW:4:20 and MINLEN:36. Quality of the reads was deemed excellent and no further actions were taken.</p> <p>Samples identifications: genome_hcLineX with &quot;_1&quot; indicating the forward and &quot;_2&quot; the reverse reads.</p> <p><strong>B. Illumina eccDNA sequencing&nbsp;of Arabidopsis control and hcLines following stress treatments</strong></p> <p>Extrachromosomal circular DNA was prepared and sequenced as follows:&nbsp;twenty plants from each petri dish were pooled separately and DNA was extracted using the CTAB method (<a href="https://dx.doi.org/10.17504/protocols.io.quidwue">dx.doi.org/10.17504/protocols.io.quidwue</a>). Following the mobilome-seq method described in (Lanciano et al., 2017), for all samples, we digested linear DNA from 2 &micro;g of total DNA for 17 hours at 37<sup>o</sup>C using 10 U of PlasmidSafe (<em>LubioScience cat# E3101K</em>), followed by enzyme denaturation (30 mins at 70<sup>o</sup>C). Digested DNA was precipitated with isopropanol supplemented with 1 &micro;g of GlycoBlue coprecipitant (<em>Fisher Scientific cat# 10391565</em>). Circular DNA was then amplified through rolling circle amplification (RCA) with the Illustra TempliPhi kit (<em>GE Healthcare cat# 25-6400-10</em>), following the manufacturer recommendation and leaving the reaction for 16h at 30<sup>o</sup>C. DNA was once again precipitated with isopropanol and sent for Illumina paired end 150 bp sequencing at BGI, Hong Kong.&nbsp;</p> <p>Samples identification:&nbsp;</p> <p>eccDNA_A.thaliana_ctrl:&nbsp;control reads</p> <p>eccDNA_A.thaliana_HS: heat stressed plants reads</p> <p>eccDNA_A.thaliana_AZ_HS: reads of&nbsp;alpha-amanitin, zebularine and heat-stressed plants</p> <p>&quot;R1&quot; indicates forward and &quot;R2&quot; reverse reads.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Diel-regulated transcriptional cascades of microbial eukaryotes in the North Pacific Subtropical Gyre

<p>Trinity <em>de novo </em>assemblies of 24 poly-A+ selected, combined-replicate metatranscriptomes from&nbsp;HOE-Legacy 2 cruise KM1513&nbsp;(Jul 24 - Aug 6, 2015).&nbsp;KM1513 cruise information, plots, and associated environmental data for the HOE Legacy II&nbsp;cruise can be found online at <a href="http://hahana.soest.hawaii.edu/hoelegacy/hoelegacy.html">http://hahana.soest.hawaii.edu/hoelegacy/hoelegacy.html</a>. Raw&nbsp;metatranscriptome short-read sequence data is available in the NCBI Sequence Read Archive&nbsp;under BioProject ID PRJNA492142. Code associated with this project is available on&nbsp;Github (<a href="https://github.com/armbrustlab/diel_eukaryotes">https://github.com/armbrustlab/diel_eukaryotes</a>).</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

TF-Marker: A comprehensive manually curated database for transcription factors and related markers in specific cell and tissue types in human.

<p>Here, we developed the TF-Marker database (TF-Marker, http://bio.liclab.net/TF-Marker/) which is committed to a comprehensive manual curation of TFs and related markers with experimental evidence in specific cell and tissue types in human. Currently, through reviewing <strong>2,091</strong> published literature, we have manually classified TFs and related markers into five types according to their functions: 1) <strong>TF</strong>: TFs, which regulate the expression of markers; 2) <strong>T Marker</strong>: markers, which are regulated by TFs (TF and T Marker pairs can identify cell types more specifically); 3) <strong>I Marker</strong>: markers, which influence the activity of TFs (I Markers can also influence the development of specific cells and tissues); 4) <strong>TFMarker</strong>: TFs, which play roles as markers (TFMarkers are cell/tissue-specific TFs used as cell markers in biology experiments); and 5) <strong>TF Pmarker</strong>: TFs, which play roles as potential markers. By curating thousands of published literature, <strong>5,905</strong> entries including <strong>1,316</strong> TFs, <strong>1,092</strong> T Markers, <strong>473</strong> I Markers, <strong>1,600</strong> TFMarkers and <strong>1,424</strong> TF Pmarkers, were annotated in <strong>383</strong> cell types and <strong>95</strong> tissue types in human. Moreover, TF-Marker divided markers into disease markers and tissue/cell-specific markers. TF-Marker is an elaborate database, which provides TFs and related markers supported by experimental evidence. We believe TF-Marker will provide strong support for research into cell/tissue-specific TFs and related markers.</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

CrowdSpeech and Vox DIY: Benchmark Dataset for Crowdsourced Audio Transcription

<p>We collect and release CrowdSpeech &mdash;&nbsp;the first publicly available large-scale dataset of crowdsourced audio transcriptions.&nbsp;e show its applicability on an under-resourced language by constructing VoxDIY &mdash;&nbsp;a counterpart of CrowdSpeech for the Russian language.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Kremlin.ru transcripts 1999–2019, Russian

<p>Document collection scraped from the Russian governmental website kremlin.ru. Includes all items listed on&nbsp;http://kremlin.ru/events/president/transcripts and following pages (e.g. http://kremlin.ru/events/president/transcripts/2) from 31 December 1999 until the end of 2019.</p> <p>10,723 documents. One document in each row.</p> <p>Columns:</p> <p>- Id: format Kremlin-1&nbsp;&nbsp;<br> - Id_no: format 1&nbsp;&nbsp;<br> - Date: Document date, format 1999-12-31&nbsp;&nbsp;<br> - Title: Document title<br> - Text: Document text including title&nbsp;&nbsp;<br> - URL: URL from which the document is downloaded&nbsp;&nbsp;<br> - Downloaded: Date of download, format 2019-12-31</p> <p>Formats: rds and json.</p> <p>Version 1.1: edited column names.</p> <p>Kremlin.ru content is licensed&nbsp;under Creative Commons Attribution 4.0 International.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Integrative in situ mapping of single-cell transcriptional states and tissue histopathology in an Alzheimer disease model

<p>Amyloid-&beta; plaques and neurofibrillary tau tangles are the neuropathologic hallmarks of Alzheimer&rsquo;s disease (AD), but the spatiotemporal cellular responses and molecular mechanisms underlying AD pathophysiology remain poorly understood. Here we introduce STARmap PLUS to simultaneously map single-cell transcriptional states and disease marker proteins in brain tissues of AD mouse models at a voxel size of 95  95  350 nm. This high-resolution spatial transcriptomics map revealed a core-shell structure where disease-associated microglia (DAM) closely contact amyloid-&beta; plaques, whereas disease-associated astrocyte-like cells (DAA-like) and oligodendrocyte precursor cells (OPC) are enriched in the outer shells surrounding the plaque-DAM complex. Hyperphosphorylated tau emerged mainly in excitatory neurons in the CA1 region accompanied by infiltration of oligodendrocyte subtypes into the axon bundles of hippocampal alveus. The integrative STARmap PLUS method bridges single-cell gene expression profiles with tissue histopathology at subcellular resolution, providing an unprecedented roadmap to pinpoint the molecular and cellular mechanisms of AD pathology and neurodegeneration.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Increased egg shell temperature during incubation leads to changes in transcriptional and epigenetic profiles in chicken lungs

<p>These RDS files contain <strong>DESeqDataSet </strong>objects subsets per broiler age and treatment. These objects are the result of DESeq2::DESeq( &hellip; ,betaPrior=FALSE).The .txt-objects contain the normalized sequencing counts per broiler age and treatment group. These objects are the result of DESeq2::counts( &hellip; , normalized=TRUE). Data was generated using STAR v2.7.10a and DESeq2 v1.36. Metadata is included as Excel file.</p> <p>Sequencing data is deposited at NCBI-SRA under BioProject: PRJNA949139.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</p> <p>&nbsp;</p> <p><strong>Study abstract</strong></p> <p>D. Schokker, J. de Vos, P.B. Stege, O. Madsen, H.J. Wijnen, S.K. Kar, and J.M.J. Rebel</p> <p>Health and resilience against respiratory diseases are important features for broiler chicken. In this study, epigenetic and transcriptomic changes in the lungs of broiler chickens of different ages during rearing that were either exposed to elevated egg shell temperature (HIGH) of 38.9&deg;C during mid-incubation or normal egg shell temperature (control; CON). The objective was to better understand how environmental challenges, such as heat stress during egg incubation, affect the development of the immune system and health of broiler chicken at later age. To this end we generated both epigenetic and transcriptomic data of lung tissue of elevated HIGH and CON chicken, furthermore these chicken were challenged by introducing either an infectious E. coli or an IBV vaccination to monitor the respiratory response. Thousands of differential methylated sites were observed at days 15 and 33, when comparing HIGH vs. CON. Pathway enrichment analysis of HIGH vs. CON showed that differentially expressed genes were mainly involved in cilium, cytoskeleton, and immune processes. These findings provide insight into the underlying biological mechanisms of early life conditions, like elevated EST, and their potential role in health of broilers.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Analysis accompanying "Dynamically regulated transcription factors are encoded by highly unstable mRNAs in the Drosophila larval brain"

<p>This repository documents the raw data processing and figure generation for the article &ldquo;Dynamically regulated transcription factors are encoded by highly unstable mRNAs in the <em>Drosophila </em>larval brain&rdquo;, doi: 10.1261/rna.079552.122.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Epigenetic and transcriptional landscape of stress memory in woodland strawberry

<p>Bedfiles of differentially methylated regions (DMRs) that were detected in stressed mother plants (M) and their (themselves unstressed) clonal daughter plants that were formed <em>via</em> stolon formation (St1, St2, St3). The M plants were grown and sampled <em>in vitro</em> and the St1, St2, St3 plants in the green house. The DMRs were called using the the EpiDiverse/dmr bioinformatic analysis pipeline (Nunn et al., 2021).</p> <p><strong>Stress assays <em>in vitro</em></strong></p> <p>One-month-old seedlings were transferred to a fresh MS media and growth chambers at 24<sup>o</sup>C/21<sup>o</sup>C (day/night),16 h light/8 h dark, as control conditions<em>. </em>For heat-stress, plants were exposed to 30<sup>o</sup>C (day/night) for one week followed by 2 days of recovery (24<sup>o</sup>C/21<sup>o</sup>C) on fresh medium as well as the control plants. Then, the plates were transferred to 37<sup>o</sup>C (day/night) for 1 week with 2 recovery days (Figure 1A). We sampled aerial parts of plants for the molecular analyses. To reduce variability resulting from individual plants, three biological replicates of 5 pooled plants were collected per condition. Samples were harvested in 1.5 mL tubes between 9:00-11:00 a.m. and immediately frozen in liquid nitrogen and stored at -80<sup>o</sup>C until required.</p> <p><strong>Greenhouse propagation assays</strong></p> <p><em>In vitro</em> plants after heat and control treatment were transferred to soil (one plant per pot) in square plastic pots (size: 12x12x10 cm) and to a greenhouse with long day conditions (24<sup>o</sup>C/21<sup>o</sup>C day/night and 60%-70% humidity).</p> <p>Twelve mother plants (M) from control (CM; n=12) and heat-stress (HM; n=12) conditions were used for asexual propagation. From each mother plant, the two first stolons (St) were kept for producing the daughter plants of the first asexual propagation (St1) in individual pots. After two weeks, following root formation, the stolons were cut to get independent daughter plants from their mother plant (M). This process was continued until St3.</p> <p>The DMRs can also be visualized here: <a href="https://jbrowse.agroscope.info/jbrowse/?data=fragaria_sub">https://jbrowse.agroscope.info/jbrowse/?data=fragaria_sub</a></p> <p>And the raw bisulfite sequencing data can be found here: <a href="https://www.ebi.ac.uk/ena/browser/text-search?query=ERP135585">https://www.ebi.ac.uk/ena/browser/text-search?query=ERP135585</a></p> <p>&nbsp;</p> <p>&nbsp;How many daughter plants per stolon? One stolon produce a chain of daughter plants.</p>

opencc-by-4.0May 2024View details →
zenodo44/100

TRANSCRIPT drug repurposing dataset

<p>Version 2.0.0 (05/29/2023)</p> <p>This is a drug repurposing dataset under MIT licence, compiled by Dr. Cl&eacute;mence R&eacute;da &lt;clemence.reda@uni-rostock.de&gt; at Universit&auml;t Rostock, comprising a drug-disease association matrix, and several drug-drug and disease-disease similarity matrices. It only uses transcriptomic data (i.e., gene activity/expression). The sparsity number is the percentage of nonzero values in the association matrix.</p> <p># drugs | # diseases | Sparsity number | # positive associations | # negative associations | # genes<br> ------- | ---------- | --------------- | ----------------------- | ----------------------- | -------<br> 204&nbsp;&nbsp;&nbsp;&nbsp; | 116&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; | 0.44%&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; | 401&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; | 11&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; | 12,096</p> <p>All drugs (resp., diseases) are associated with a gene expression feature vector of length 12,096 (that is, all drugs and diseases in the feature matrices appear in the association matrix, and vice versa).</p> <p>----------</p> <p>This dataset consists of three .CSV files:</p> <p>* Drug-Disease Association Matrix</p> <p>1. &quot;ratings_mat.csv&quot;</p> <p>This matrix contains values in {-1,0,1} where -1 stands for a negative association (i.e., the drug failed for some reason to treat the considered disease: e.g., lack of accrual in the associated clinical trial, or proven toxicity), 1 for a positive association (i.e., the drug was shown to treat the disease), and 0 for unknown associated status. The columns are diseases, identified by their MedGen Concept ID, whereas rows are drugs, identified by their DrugBank IDs or PubChem CIDs.</p> <p>* Drug Feature Matrix</p> <p>1. &quot;items.csv&quot;</p> <p>This matrix has drugs in its columns, identified by their DrugBank IDs or PubChem CIDs, and genes in its rows, identified by their HUGO Gene Symbol. Genewise transcriptomic variation induced by drug treatment, from the CREEDS or the LINCS L1000 databases.</p> <p>* Disease Feature Matrix</p> <p>1. &quot;users.csv&quot;</p> <p>This matrix has diseases in its columns, identified by their MedGen Concept IDs, and genes in its rows, identified by their HUGO Gene Symbol. Genewise transcriptomic variation induced by the disease, from the CREEDS database.</p> <p>----------</p> <p>Further information about the generation of those matrices is available by running the Jupyter notebook TRANSCRIPT_dataset.ipynb on the following GitHub repository: https://github.com/RECeSS-EU-Project/drug-repurposing-datasets. For any questions, please contact the author at &lt;clemence.reda@uni-rostock.de&gt; or the RECeSS project contributors at &lt;recess-project@proton.me&gt;.</p>

openmit-licenseMay 2023View details →
zenodo44/100

FEDORA. Excerpts from essays, transcript of interviews and group discussions on students' future perception. Part 3: Interviews, Finland.

<p>Finnish-language dataset. Related to a research article that is awaiting acceptance for publication: <em>Future, technology and agency: Students&rsquo; experiences from a course on futures thinking and quantum computing</em>.</p> <p>As per ethical concerns and participants&#39; consent, the dataset is given in a fully anonymised form. Instead of students&#39; interviews (the context of which is given in the article). In a nutshell, 21 upper-secondary school students were interviewed in 2018 regarding their experiences on taking an experimental science course that combined ideas from futures thinking and quantum computing. The present dataset contains all 245 transcribed passages from 21 student interviews that were initially marked as relevant to the research goals (i.e. how students saw their conceptions change over the course). Additionally, for each passage the final coding that was used in the analysis for the research paper is shown. The &quot;number-letter codes&quot; were used as shorthands; the full names of the codes correspond closely with the final, English-language codes in the paper.</p> <p>To preserve full anonymity, students are not identified by any marker or pseudonym here; rather, the passages are given alphabetically. The start and end of passages has not been checked for additional or missing first and last characters. Please also note that the character &gt; marks change of speaker. Identifying the interviewer and interviewee should be straighforward based on the context.</p> <p>Please contact the corresponding author for more information.</p> <p>&nbsp;</p> <p>--</p> <p>&nbsp;</p> <p><a href="https://zenodo.org/communities/futuresthinking?page=1&amp;size=20">FEDORA Project</a>&nbsp;README:</p> <p>&nbsp;</p> <p><strong>README</strong></p> <p><strong>Data Set Title:</strong>&nbsp;&ldquo;FEDORA. Excerpts from essays, transcript of interviews and group discussions on students&rsquo; future perception. Finland&quot;</p> <p><strong>Data Set Author/s:</strong>&nbsp;Antti Laherto, Tapio Rasa, Elina Palmgren&nbsp;(University of Helsinki)</p> <p><strong>Data Set Contact Person/s</strong>: Tapio Rasa<strong>&nbsp;</strong>(University of Helsinki), ORCID 0000-0003-1315-5207, tapio.rasa@helsinki.fi;</p> <p><strong>Data Set License</strong>: this data set is distributed under the Creative Commons Attribution 4.0 International (CC BY&nbsp;4.0) license.</p> <p><strong>Publication Year</strong>: 2021</p> <p><strong>Project Info</strong>: FEDORA<strong>&nbsp;</strong>(Future-oriented Science EDucation to enhance Responsibility and engagement in the society of Acceleration and uncertainty<strong>&nbsp;,&nbsp;</strong>funded by European Union, Horizon 2020 Programme. Grant Agreement num.<strong>&nbsp;</strong>872841,<br> www.fedora-project.eu)</p> <p>&nbsp;</p> <p><strong>Data set Contents</strong></p> <p>The data set consists of:</p> <p>One spreadsheet file, provided in two alternative formats (CSV and XLSX).</p> <p>Future_technology_agency_DATA_CSV.csv</p> <p>Future_technology_agency_DATA_XLSX.xlsx</p> <p>&nbsp;</p> <p><strong>Data set Documentation</strong></p> <p><em>Given above this README, on the ZENODO repository.&nbsp;https://zenodo.org/record/4734161</em></p>

opencc-by-4.0May 2021View details →
zenodo44/100

CollecTRI Data for Investigation of SETBP1 gene expression and transcription factor activity across human tissues

<p>Here we provide the human CollecTRI prior (accessed May 2023) for inference of TF activity across 31 GTEx tissues using multivariate linear modeling method decoupleR.<br> <br> The `human_prior_tri.csv` includes 1,178 unique TFs (referred to as the source) that target 6,627 unique genes (referred to as targets) to give us 42,595 interactions in the CollecTRI prior input. Interactions are represented as a + or - 1 (mor).</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

H4K20me3 is important for Ash1-mediated H3K36me3 and transcriptional silencing in facultative heterochromatin in a fungal pathogen

<p>Normalized ChIP-seq datasets&nbsp;for visualization in IGV. The tracks contain means of pooled replicate datasets.</p> <p>ChIP-seq data were quality-filtered and adapters removed with trimmomatic v.0.39&nbsp;(Bolger et al., 2014). Mapping was performed with bowtie2 v.2.4.4&nbsp;(Langmead and Salzberg, 2012), and sorting and indexing with samtools v.1.9&nbsp;(Li, 2011). Normalized coverage bigwig files and heatmaps were created with deeptools v.3.5.1&nbsp;(Ram&iacute;rez et al., 2016). Wiggletools v.1.2 and the UCSC Genome Browser tools were used to calculate means for replicates and converting wig to bigwig files.</p> <p>Reference genome file is modified from Goodwin et al., 2011. Chromosome 18 was removed from the genome as our reference isolate&nbsp;is missing chromosome 18.&nbsp;</p> <p>Gene annotation file was obtained from FungiDB (release&nbsp;53) and is based on the annotation published by Grandaubert et al., 2015.</p> <p>In this version, we have added new ChIP-seq bw tracks for ∆ash1::ash1-gfp-V5&nbsp;and ∆kmt5::kmt5 complementation experiments. All tracks coming from this experiment are labeled *_compl_exp_mean.bw.</p> <p>We also added ChIP peak files (peaks called with HOMER: Heinz et&nbsp;al., 2010) for H4K20me3, H3K36me3 and H3K27me3 in WT, ∆kmt5 and ∆ash1, as well as H3K36me3 peak files for Set2- and Ash1-mediated H3K36me3.</p> <p>We have also added bed files (500 bp windows)&nbsp;containing facultative heterochromatin clusters 1 (Zt09_500bp_K27filtered_K36_K20_cluster1.bed) and 2 (Zt09_500bp_K27filtered_K36_K20_cluster2.bed).&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

COALA voice data and transcripts Dutch

<p>This dataset contains audio files and transcripts in Dutch and related to manufacturing. We collected the scripts during the Horizon Europe RIA COALA (GA 957296, <a href="https://cordis.europa.eu/project/id/957296">project reference website</a>) from industrial use cases and hired a service provider to generate the related audio files (BIBA - Bremer Institut f&uuml;r Produktion und Logistik GmbH ordered the service). The service provider checked the audio files for quality.</p> <p>The service provider recruited crowd workers, and gathered their audio records, informed consent (privacy) and agreement that their records become public domain (Creative Commons 0; https://creativecommons.org/share-your-work/public-domain/cc0/). The service provider declared to follow a Crowd Code of Ethics and a Fair Pay policy.</p> <p>The metadata file contains the following information:</p> <ul> <li><strong>file_name</strong>: name of the audio file</li> <li><strong>script</strong>: script the speaker had to speak</li> <li><strong>scriptId</strong>: the numeric identifier of the script</li> <li><strong>participantId</strong>: the numeric identifier of the participant (speaker)</li> <li><strong>gender</strong>: the gender as indicated by the participant (MALE or FEMALE)</li> <li><strong>age</strong>: the age in years as indicated by the participant</li> <li><strong>age_range</strong>: the age range in years&nbsp; (18-30, 31-45, 46+)</li> <li><strong>country</strong>: the birth country indicated by the participant</li> <li><strong>current_country</strong>: the country of residence indicated by the participant</li> <li><strong>primary_language</strong>: the language indicated as primary by the participant</li> <li><strong>ever_worked_factory</strong>: answer to the question: &quot;Have you ever worked in a factory, manufacturing setting?&quot; (Yes/No)</li> <li><strong>years_worked_factory</strong>: answer to the question: &quot;If yes, for how many years?&quot; (1-10, 10+)</li> <li><strong>background_noise_type</strong>: background noise in the audio as indicated by the participant (mild, humming/technical, no noise)</li> <li><strong>gdpr_and_ipr_consent</strong>: answer to the privacy notice and the ipr transfer to CC-0 (Yes)</li> <li><strong>date_signed</strong>: date when the participant signed the consent form (US format, MM.DD.YYYY)</li> </ul>

opencc-zeroOct 2023View details →
zenodo44/100

COALA voice data and transcripts Italian

<p>This dataset contains audio files and transcripts in Italian and related to manufacturing. We collected the scripts during the Horizon Europe RIA COALA (GA 957296, <a href="https://cordis.europa.eu/project/id/957296">project reference website</a>) from industrial use cases and hired a service provider to generate the related audio files (BIBA - Bremer Institut f&uuml;r Produktion und Logistik GmbH ordered the service). The service provider checked the audio files for quality.</p> <p>The service provider recruited crowd workers, and gathered their audio records, informed consent (privacy) and agreement that their records become public domain (Creative Commons 0; https://creativecommons.org/share-your-work/public-domain/cc0/). The service provider declared to follow a Crowd Code of Ethics and a Fair Pay policy.</p> <p>The metadata file contains the following information:</p> <ul> <li><strong>file_name</strong>: name of the audio file</li> <li><strong>script</strong>: script the speaker had to speak</li> <li><strong>scriptId</strong>: the numeric identifier of the script</li> <li><strong>participantId</strong>: the numeric identifier of the participant (speaker)</li> <li><strong>gender</strong>: the gender as indicated by the participant (MALE or FEMALE)</li> <li><strong>age</strong>: the age in years as indicated by the participant</li> <li><strong>age_range</strong>: the age range in years&nbsp; (18-30, 31-45, 46+)</li> <li><strong>country</strong>: the birth country indicated by the participant</li> <li><strong>current_country</strong>: the country of residence indicated by the participant</li> <li><strong>primary_language</strong>: the language indicated as primary by the participant</li> <li><strong>ever_worked_factory</strong>: answer to the question: &quot;Have you ever worked in a factory, manufacturing setting?&quot; (Yes/No)</li> <li><strong>years_worked_factory</strong>: answer to the question: &quot;If yes, for how many years?&quot; (1-10, 10+)</li> <li><strong>background_noise_type</strong>: background noise in the audio as indicated by the participant (mild, humming/technical, no noise)</li> <li><strong>gdpr_and_ipr_consent</strong>: answer to the privacy notice and the ipr transfer to CC-0 (Yes)</li> <li><strong>date_signed</strong>: date when the participant signed the consent form (US format, MM.DD.YYYY)</li> </ul>

opencc-zeroOct 2023View details →
zenodo40/100

Robustness and applicability of transcription factor and pathway analysis tools on single-cell RNA-seq data

<p>Data used to test the robustness and applicability of transcription factor and pathway analysis tools on single-cell RNA-seq data, described in <a href="https://doi.org/10.1186/s13059-020-1949-z">Holland et al. 2020</a>.</p> <p>The folder&nbsp;<em>data </em>contains<em>&nbsp;</em>raw data and the folder <em>output</em> contains intermediate and final results of all analyses.&nbsp;</p> <p>The associated analyses code and more information are available on&nbsp;<a href="https://github.com/saezlab/FootprintMethods_on_scRNAseq">GitHub</a>.</p> <p>&nbsp;</p> <p><strong>Abstract</strong></p> <p><strong>Background</strong></p> <p>Many functional analysis tools have been developed to extract functional and mechanistic insight from bulk transcriptome data. With the advent of single-cell RNA sequencing (scRNA-seq), it is in principle possible to do such an analysis for single cells. However, scRNA-seq data has characteristics such as drop-out events and low library sizes. It is thus not clear if functional TF and pathway analysis tools established for bulk sequencing can be applied to scRNA-seq in a meaningful way.</p> <p><strong>Results</strong></p> <p>To address this question, we perform benchmark studies on simulated and real scRNA-seq data. We include the bulk-RNA tools PROGENy, GO enrichment, and DoRothEA that estimate pathway and transcription factor (TF) activities, respectively, and compare them against the tools SCENIC/AUCell and metaVIPER, designed for scRNA-seq. For the in silico study, we simulate single cells from TF/pathway perturbation bulk RNA-seq experiments. We complement the simulated data with real scRNA-seq data upon CRISPR-mediated knock-out. Our benchmarks on simulated and real data reveal comparable performance to the original bulk data. Additionally, we show that the TF and pathway activities preserve cell type-specific variability by analyzing a mixture sample sequenced with 13 scRNA-seq protocols. We also provide the benchmark data for further use by the community.</p> <p><strong>Conclusions</strong></p> <p>Our analyses suggest that bulk-based functional analysis tools that use manually curated footprint gene sets can be applied to scRNA-seq data, partially outperforming dedicated single-cell tools. Furthermore, we find that the performance of functional analysis tools is more sensitive to the gene sets than to the statistic used.</p> <p>&nbsp;</p> <p>For questions related to the data please write an email to christian.holland@bioquant.uni-heidelberg.de or use the <a href="https://github.com/saezlab/FootprintMethods_on_scRNAseq/issues">GitHub issue system</a>.</p>

opencc-by-4.0Dec 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record