Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
241
datasets available to search
ShareScore release 0.9.0
Dataset results
241 results for “Transcriber”
FANTOM5 transcribed enhancers in hg38
<p><strong>Overview</strong></p> <p>Transcribed enhancers were identified and their expression was quantified across all human FANTOM5 libraries, following the re-aligned FANTOM5 CAGE data upon hg38 (GRCh38) (obtained from http://fantom.gsc.riken.jp/5/datafiles/reprocessed/hg38_v1/basic/), and decomposition-based peak identification (obtained from https://zenodo.org/record/545682#.WPuNy1Pyv2Q) by Kawaji, Hideya.</p> <p> </p> <p><strong>Description</strong></p> <p>Transcribed enhancers were called based on bidirectional balanced RNA signatures as per Andersson et al (2014). Enhancers were only identified distal to known exons (+/-100bp region from boundaries) and transcription start sites (+/-300bp), defined by GENCODE v24 annotation. In total, 63,285 enhancers were identified across 1,829 libraries. The expression was quantified and TPM (tags per million) normalised according to the total number of mapped reads within the full set of TCs. For details regarding the identification of transcribed enhancers from CAGE data, please see Andersson et al (2014) and blog post.</p> <p>Due to varying noise levels across FANTOM5 libraries and the intrinsic low expression levels of transcribed enhancers, library-specific noise levels were estimated to define of robust set of enhancers in each sample. In summary, for each library, expression was quantified in randomly sampled genomic regions distal to assembly gaps, DNase hypersensitive sites (ENCODE), known exons and gene TSSs (GENCODE) to create a genomic background expression distribution. For each library, we called an enhancer active (used) if its expression was above the 99.9th quantile of the library’s genomic background expression distribution. The robust set of enhancers consist of 60,215 over 1,829 libraries, being significantly expressed in at least one library.</p> <p>While this approach ensures less permissive enhancer calling in noisy libraries, for some libraries the noise threshold is zero meaning that a single CAGE tag is sufficient for calling an enhancer active. Furthermore, the possibility of detecting enhancer transcription is affected by sequencing depth, so the number of active enhancers per library might not be biologically meaningful to compare when sequencing depths differ.</p> <p> </p> <p><strong>Data files</strong><br> Each predicted enhancer is described in BED12 format with two blocks denoting the merged regions of transcription initiation on the minus and plus strands. The thickStart and thickEnd columns denote the inferred mid position between blocks of transcription initiation events. Expression and usage matrices are tab delimited and the first row gives the FANTOM5 CNhs IDs and the first column the enhancer ID (same as column 4 in BED file). Usage matrices contain zeroes and ones (0:not used, 1:used).</p> <ul> <li>enhancers (BED12 format)</li> <li>enhancer expression matrix (tab delimited, first row: CNhs IDs, first column: enhancer ID (coordinate))</li> <li>enhancer expression matrix TPM normalized (tab delimited, first row: CNhs IDs, first column: enhancer ID (coordinate))</li> <li>binary enhancer usage matrix (0:not used, 1:used, tab delimited, first row: CNhs IDs, first column: enhancer ID (coordinate))</li> </ul>
Data for: Raw count data, transcribed variant count data, and reference genomic annotation files for Boocock et al. 2024
<p>Expression quantitative trait loci (eQTLs) provide a key bridge between noncoding DNA sequence variants and organismal traits. The effects of eQTLs can differ among tissues, cell types, and cellular states, but these differences are obscured by gene expression measurements in bulk populations. We developed a one-pot approach to map eQTLs in <em>Saccharomyces cerevisiae</em> by single-cell RNA sequencing (scRNA-seq) and applied it to over 100,000 single cells from three crosses. We used scRNA-seq data to genotype each cell, measure gene expression, and classify the cells by cell-cycle stage. We mapped thousands of local and distant eQTLs and identified interactions between eQTL effects and cell-cycle stages. We took advantage of single-cell expression information to identify hundreds of genes with allele-specific effects on expression noise. We used cell-cycle stage classification to map 20 loci that influence cell-cycle progression. One of these loci influenced the expression of genes involved in the mating response. We showed that the effects of this locus arise from a common variant (W82R) in the gene <em>GPA1</em>, which encodes a signaling protein that negatively regulates the mating pathway. The 82R allele increases mating efficiency at the cost of slower cell-cycle progression and is associated with a higher rate of outcrossing in nature. Our results provide a more granular picture of the effects of genetic variants on gene expression and downstream traits.</p>
FANTOM5 transcribed enhancers in mm10
<p><strong>Overview</strong></p> <p>Transcribed enhancers were identified and their expression was quantified across all human FANTOM5 libraries, following the re-aligned FANTOM5 CAGE data upon mm10 (GRCm38) (obtained from http://fantom.gsc.riken.jp/5/datafiles/reprocessed/mm10_v1/basic/), and decomposition-based peak identification (obtained from https://zenodo.org/record/545682#.WPuNy1Pyv2Q) by Kawaji, Hideya.</p> <p><strong>Description</strong></p> <p>Transcribed enhancers were called based on bidirectional balanced RNA signatures as per Andersson et al (2014). Enhancers were only identified distal to known exons (+/-100bp region from boundaries) and transcription start sites (+/-300bp), defined by GENCODE vM7 annotation. In total, 44,138 enhancers were identified across 1,068 libraries and the expression was quantified. For details regarding the identification of transcribed enhancers from CAGE data, please see Andersson et al (2014) and blog post.</p> <p>Due to varying noise levels across FANTOM5 libraries and the intrinsic low expression levels of transcribed enhancers, library-specific noise levels were estimated to define of robust set of enhancers in each sample. In summary, for each library, expression was quantified in randomly sampled genomic regions distal to assembly gaps, DNase hypersensitive sites (ENCODE), known exons and gene TSSs (GENCODE vM7) to create a genomic background expression distribution. For each library, we called an enhancer active (used) if its expression was above the 99.9th quantile of the library’s genomic background expression distribution. The robust set of enhancers consist of those significantly expressed in at least one library.</p> <p>While this approach ensures less permissive enhancer calling in noisy libraries, for some libraries the noise threshold is zero meaning that a single CAGE tag is sufficient for calling an enhancer active. Furthermore, the possibility of detecting enhancer transcription is affected by sequencing depth, so the number of active enhancers per library might not be biologically meaningful to compare when sequencing depths differ.</p> <p><strong>Data files</strong><br> Each predicted enhancer is described in BED12 format with two blocks denoting the merged regions of transcription initiation on the minus and plus strands. The thickStart and thickEnd columns denote the inferred mid position between blocks of transcription initiation events. Expression and usage matrices are tab delimited and the first row gives the FANTOM5 CNhs IDs and the first column the enhancer ID (same as column 4 in BED file). Usage matrices contain zeroes and ones (0:not used, 1:used).</p> <ul> <li>enhancers (BED12 format)</li> <li>enhancer expression matrix (tab delimited, first row: CNhs IDs, first column: enhancer ID (coordinate))</li> <li>binary enhancer usage matrix (0:not used, 1:used, tab delimited, first row: CNhs IDs, first column: enhancer ID (coordinate))</li> </ul>
The Smc5/6 complex counteracts R-loop formation at highly transcribed genes in cooperation with RNase H2
Open the record for dataset details and reuse information.
An atlas of transcribed enhancers across helper T cell diversity for decoding human diseases
Open the record for dataset details and reuse information.
Data for: Raw count data, transcribed variant count data, and reference genomic annotation files for Boocock et al. 2024
Open the record for dataset details and reuse information.
At-home Jomo Propitiation in Chug valley: Toolbox, transcriber and audio file
<p>This dataset contains the .wav sound files, .trs Transcriber files, .txt Toolbox-compatible Notepad files and .pdf files with the completely transcribed, glossed, parsed and translated examples of the following recordings that belong to the following publication:</p> <p>Bodt, Timotheus Adrianus. 2020. Grammar of Duhumbi. Leiden: Brill. ISBN 978-90-04-40947-7. <a href="https://brill.com/view/title/55767">https://brill.com/view/title/55767</a></p> <p>CHUK221212D2A Bonpo prediction text.</p> <p>The explanation of all the grammatical features that occur in these sound files can be found in the Grammar of Duhumbi.</p> <p>The main Toolbox files can be found in the zip file “Settings”, this includes the IPA keys for Duhumbi, the entire setup of the Toolbox database, and the Duhumbi dictionary and Parsing dictionary.</p> <p>The .wav, .txt and .trs files combined in the same folder will enable to open Toolbox and work with the recordings, e.g. play them sentence for sentence and see the transcriptions and translations.</p> <p>Transcriber version 1.5.1: <a href="http://trans.sourceforge.net/en/presentation.php">http://trans.sourceforge.net/en/presentation.php</a> or <a href="https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/">https://osdn.net/projects/sfnet_trans/downloads/transcriber/1.5.1/Transcriber-1.5.1-Windows.exe/</a></p> <p>Toolbox version 1.6.1: <a href="https://software.sil.org/toolbox/download/">https://software.sil.org/toolbox/download/</a></p> <p>For the metadata of the sound files in this data set, I refer to Chapter 13 Texts in the Grammar of Duhumbi. This Chapter has a complete listing of the texts, their topics, the speakers and their background etc.</p> <p>Kesang was one of the last <em>bon-po</em>, practitioners of the traditional religious system in the Chug valley. This religious system has no ‘name’, but comes under what is often called ‘Bon’. The locus of propitiation is on the opposition between the high, in winter snow-clad mountain peaks <em>phu</em> (Tib. phu) that represent purity, cleanliness, goodness and beneficial powers versus the low-lying marshy, swampy areas <em>da</em> (tib. mdaḥ) that represent pollution, disease, evil and malevolent forces. In between these two there is a plethora of other local deities, many of which are local representations of the <em>lha-srin bde-brgyad</em> ‘eight classes of deities and demons’ also found in Tibetan Buddhism, but many of whom also represent deified human beings who have taken on some negative or positive force. It is the role of the <em>bon-po</em> to maintain the balance between the <em>phu</em>, the ‘good’ and the <em>da</em>, the ‘evil’ and hence prevent damage to humans and their livelihoods in the form of diseases, natural disasters, death etc. The <em>phu-da</em> religious system is not limited to the Chug valley, but, in various forms, can also be found among the related Khispi, Sartang and Sherdukpen people, as well as among the Tshangla speakers of West Kameng and eastern Bhutan.</p> <p>The main ritual conducted by the bon-po is called <em>zhiwa</em> (Tib. źi-ba ‘peace’) or <em>jomo soykha</em> (Tib. jo-mo gsol-kha ‘propitiation of the Jomo). It is conducted once before the 20th day of every Tibetan month. Jomo is the main female deity in the area (see also the files concerning the on-site Jomo propitiation). During this ritual, the <em>bon-po</em> first invites the Jomo and all other deities and spirits to attend the offering. He then offers <em>nyingba</em> (Tib. sñiṅ-ba ‘old’), also called <em>lemchang</em>, a mixture of rice, maize, finger millet (traditionally also wheat, barley, broomcorn millet, foxtail millet, buckwheat and amaranth) that has been kept fermenting for a long time. After that, he offers <em>tochang</em>, freshly cooked rice (Tib. lto-chaṅ ‘food-liquor’), and after that <em>darcok</em> (Tib. dar-lcog ‘prayer flags’), small twigs with triangular-shaped flags made of traditional paper. He then conducts a prediction for the coming month, by making three heaps of a mixture of grains, and interprets the way in which these grains pattern. He then sends off the assembled deities.</p> <p><em>Bon-po</em> Kesang died in early 2016. His son Tow Tsering has taken over his role, but does not seem to know the ritual as well as his father. He may well be the last of the <em>bon-po</em> in Chug valley, given that no one has come forward to learn from him.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <strong><em>of any kind</em></strong><em>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p>
Data from: Nuclear internal transcribed spacer-1 as a sensitive genetic marker for environmental DNA studies in common carp Cyprinus carpio
The recently developed environmental DNA (eDNA) analysis has been used to estimate the distribution of aquatic vertebrates by using mitochondrial DNA (mtDNA) as a genetic marker. However, mtDNA markers have certain drawbacks such as variable copy number and maternal inheritance. In this study, we investigated the potential of using nuclear DNA (ncDNA) as a more reliable genetic marker for eDNA analysis by using common carp (Cyprinus carpio). We measured the copy numbers of cytochrome b (CytB) gene region of mtDNA and internal transcribed spacer 1 (ITS1) region of ribosomal DNA of ncDNA in various carp tissues and then compared the detectability of these markers in eDNA samples. In the DNA extracted from the brain and gill tissues and intestinal contents, CytB was detected at 95.1 ± 10.7 (mean ± 1 standard error), 29.7 ± 1.59 and 24.0 ± 4.33 copies per cell, respectively, and ITS1 was detected at 1760 ± 343, 2880 ± 503 and 1910 ± 352 copies per cell, respectively. In the eDNA samples from mesocosm, pond and lake water, the copy numbers of ITS1 were about 160, 300 and 150 times higher than those of CytB, respectively. The minimum volume of pond water required for quantification was 33 and 100 mL for ITS1 and CytB, respectively. These results suggested that ITS1 is a more sensitive genetic marker for eDNA studies of C. carpio.
FIGURE. Comparison of the sequences of the internal transcribed spacer (ITS) region of Hedysarum sunhangii and Hedysarum nuratense. The variable region is shown in a red frame. in Hedysarum sunhangii (Fabaceae, Hedysareae), a new species from Pamir-Alay (Babatag Ridge - Uzbekistan)
FIGURE. Comparison of the sequences of the internal transcribed spacer (ITS) region of Hedysarum sunhangii and Hedysarum nuratense. The variable region is shown in a red frame.
Supplementary material 6 from: Ceballos-Escalera A, Richards J, Arias MB, Inward DJG, Vogler AP (2022) Metabarcoding of insect-associated fungal communities: a comparison of internal transcribed spacer (ITS) and large-subunit (LSU) rRNA markers. MycoKeys 88: 1-33. https://doi.org/10.3897/mycokeys.88.77106
Table S2. Class level identification of OTUs showing the number of OTUs produced with ITS2 and LSU and the proportion of the total OTU set on the rarefied data
Supplementary material 5 from: Ceballos-Escalera A, Richards J, Arias MB, Inward DJG, Vogler AP (2022) Metabarcoding of insect-associated fungal communities: a comparison of internal transcribed spacer (ITS) and large-subunit (LSU) rRNA markers. MycoKeys 88: 1-33. https://doi.org/10.3897/mycokeys.88.77106
Table S1. Accession numbers corresponding with the reference sequences used to build the phylogenetic trees
Supplementary material 3 from: Ceballos-Escalera A, Richards J, Arias MB, Inward DJG, Vogler AP (2022) Metabarcoding of insect-associated fungal communities: a comparison of internal transcribed spacer (ITS) and large-subunit (LSU) rRNA markers. MycoKeys 88: 1-33. https://doi.org/10.3897/mycokeys.88.77106
Figure S3. Maximum-likelihood tree constructed in IQ-Tree2 based on three-gene (LSU D1-D2, SSU, ITS2) reference sequence alignments and OTUs for both markers (clustering thresholds: 99% LSU D1-D2 and 98% ITS2)
Supplementary material 2 from: Ceballos-Escalera A, Richards J, Arias MB, Inward DJG, Vogler AP (2022) Metabarcoding of insect-associated fungal communities: a comparison of internal transcribed spacer (ITS) and large-subunit (LSU) rRNA markers. MycoKeys 88: 1-33. https://doi.org/10.3897/mycokeys.88.77106
Figure S2. Species accumulation curves of the OTUs generated from the ITS (panel right) and LSU (panel left) metabarcodes
FIGURE 2 in Molecular Phylogeny of Ethiopian Artemisia (Asteraceae) Species Based on Nuclear External Transcribed Spacer (ETS) and Internal Transcribed Spacer (ITS)
FIGURE 2. The maximum likelihood (ML) tree inferred from 1000 replicates is taken to represent the evolutionary history of the combined nuclear datasets (ITS and ETS). Contrary to this, the branches corresponding to partitions reproduced in less than 50% bootstrap replicates were collapsed. The values indicated above and below branches are the Bootstrap values (> 50%) obtained from ML and MP analysis respectively with 1000 replicates. The species names are colored according to their subgeneric affiliation.
FIGURE 1 in Molecular Phylogeny of Ethiopian Artemisia (Asteraceae) Species Based on Nuclear External Transcribed Spacer (ETS) and Internal Transcribed Spacer (ITS)
FIGURE 1. Map of Ethiopia indicating the geographic distribution of Artemisia samples included in this study.
Supplementary material 4 from: Rosenblad MA, Martín MP, Tedersoo L, Ryberg M, Larsson E, Wurzbacher C, Abarenkov K, Nilsson RH (2016) Detection of signal recognition particle (SRP) RNAs in the nuclear ribosomal internal transcribed spacer 1 (ITS1) of three lineages of ectomycorrhizal fungi (Agaricomycetes, Basidiomycota). MycoKeys 13: 21-33. https://doi.org/10.3897/mycokeys.13.8579
SRP RNA multiple sequence alignment : Explanation note: Multiple sequence alignment with the SRP RNA sequences of Dumesic et al. (2015; Stereum hirsutum, Heterobasidion irregulare, and Heterobasidion annosum) aligned to our newly generated ITS sequences of Russula and Lactarius.
Supplementary material 3 from: Rosenblad MA, Martín MP, Tedersoo L, Ryberg M, Larsson E, Wurzbacher C, Abarenkov K, Nilsson RH (2016) Detection of signal recognition particle (SRP) RNAs in the nuclear ribosomal internal transcribed spacer 1 (ITS1) of three lineages of ectomycorrhizal fungi (Agaricomycetes, Basidiomycota). MycoKeys 13: 21-33. https://doi.org/10.3897/mycokeys.13.8579
ITS/SRP RNA multiple sequence alignment : Explanation note: Multiple sequence alignment comprising the 63 public ITS1 sequences with SRP RNA found in them, the three newly generated sequences, and the SRP RNA sequences from Dumesic et al. (2015) (Stereum hirsutum, Heterobasidion irregulare, and Heterobasidion annosum).
Supplementary material 2 from: Rosenblad MA, Martín MP, Tedersoo L, Ryberg M, Larsson E, Wurzbacher C, Abarenkov K, Nilsson RH (2016) Detection of signal recognition particle (SRP) RNAs in the nuclear ribosomal internal transcribed spacer 1 (ITS1) of three lineages of ectomycorrhizal fungi (Agaricomycetes, Basidiomycota). MycoKeys 13: 21-33. https://doi.org/10.3897/mycokeys.13.8579
ITS multiple sequence alignment : Explanation note: A multiple sequence alignment in the NEXUS format (Maddison et al. 1997) comprising all 63 matching ITS sequences, plus the three newly generated ones (KU356730, KU356731, and KU356732). The alignment was produced in MAFFT without manual adjustment (Katoh and Standley 2013). The alignment is composed of partial nSSU (bases 1-34 in the alignment), the full ITS1 (bases 35-678), the full 5.8S (bases 679-838), the full ITS2 (bases 839-1395), and partial nLSU (bases 1396-end). The SRP RNA occupies position 203-474 in the alignment. The alignment is provided for overview purposes only; the two-order nature of the taxa (Boletales and Russulales) coupled with the high variability of the ITS region jointly mean that the alignment will not be suited for phylogenetic inference.
Supplementary material 1 from: Rosenblad MA, Martín MP, Tedersoo L, Ryberg M, Larsson E, Wurzbacher C, Abarenkov K, Nilsson RH (2016) Detection of signal recognition particle (SRP) RNAs in the nuclear ribosomal internal transcribed spacer 1 (ITS1) of three lineages of ectomycorrhizal fungi (Agaricomycetes, Basidiomycota). MycoKeys 13: 21-33. https://doi.org/10.3897/mycokeys.13.8579
Output from cmsearch and primers used : Explanation note: A) The output from cmsearch showing all 63 relevant matches to the three ectomycorrhizal lineages. B) Detail of the primers used to re-amplify the specimens.
Leafcutter ants of the genus Atta in the Insects Collection at the Field Museum of Natural History. The field data on the attached tags are transcribed for entry into databases such as AntWeb and the Global Biodiversity Information Facility. Photograph: Matthew Nelsen. in The Evolution of Natural History Collections
Leafcutter ants of the genus Atta in the Insects Collection at the Field Museum of Natural History. The field data on the attached tags are transcribed for entry into databases such as AntWeb and the Global Biodiversity Information Facility. Photograph: Matthew Nelsen.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.