Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
21,320
datasets available to search
ShareScore release 0.9.0
Dataset results
21,320 results for “Transcription”
Transcription events constructed by txrevise
<p>List of pre-computed transcription events constructed by txrevise. See the txrevise home page for more details: https://github.com/kauralasoo/txrevise.</p> <p>If you use txrevise in your research, please cite the following paper: <a href="https://doi.org/10.7554/eLife.41673">Alasoo, Kaur, et al. "Genetic effects on promoter usage are highly context-specific and contribute to complex traits." Elife 8 (2019): e41673</a></p> <p>We have constructed two types of events. In the main event files we have masked the alternative internal exons in the promoter and 3' end events (labeled upstream and downstream). This ensures that we can reliably distinguish promoter usage and 3' end usage from alternative splicing, which are likely to be driven by distinct molecular mechanisms. The main annotation files are:</p> <ul> <li>Homo_sapiens.GRCh37.87.version_1.tar.gz </li> <li>Homo_sapiens.GRCh38.92.version_1.tar.gz </li> <li>Homo_sapiens.GRCh38.96.version_1.tar.gz</li> </ul> <p>We have also constructed an alternative set of "raw" annotation files where the alternative internal exons in promoter and 3' end events have not masked. This maximises QTL discovery, because we can now also detect splicing event near promoters and 3 ends, but it comes at the expense that we are no longer able to distinguish between different molecular mechanisms. The raw annotation files are:</p> <ul> <li>Homo_sapiens.GRCh37.87.raw_events.version_1.tar.gz </li> <li>Homo_sapiens.GRCh38.92.raw_events.version_1.tar.gz </li> </ul>
Transcriptional networks underlying a primary ovarian insufficiency disorder in alligators naturally exposed to EDCs: Transformed read counts and supplementary materials
<p>Interactions between the endocrine system and environmental contaminants are responsible for impairing reproductive development and function. Despite the taxonomic diversity of affected species and attendant complexity inherent to natural systems, the underlying signaling pathways and cellular consequences are mostly studied in lab models. To resolve the genetic and endocrine pathways that mediate affected ovarian function in organisms exposed to endocrine disrupting contaminants in their natural environments, we assessed broad-scale transcriptional and steroidogenic responses to exogenous gonadotropin stimulation in juvenile alligators (<em>Alligator missippiensis</em>) originating from a lake with well-documented pollution (Lake Apopka, FL) and a nearby reference site (Lake Woodruff, FL). We found that individuals from Lake Apopka are charachterized by hyperandrogenism and display hyper-sensitive transcriptional responses to gonadotropin stimulation when compared to individuals from Lake Woodruff. Site-specific transcriptomic divergence appears to be driven by wholly distinct subsets of transcriptional regulators, indicating alterations to fundamental genetic pathways governing ovarian function. Consistent with broad-scale transcriptional differences, ovaries of Lake Apopka alligators displayed impediments to folliculogenesis, with larger germinal beds and decreased numbers of late-stage follicles. After resolving the ovarian transcriptome into clusters of co-expressed genes, most site-associated modules were correlated to ovarian follicule phenotypes across individuals. However, expression of two site-specific clusters were independent of ovarian cellular architecture and are hypothesized to represent alterations to cell-autonomous transcriptional programs. Collectively, our findings provide high resolution mapping of transcriptional patterns to specific reproductive function and advance our mechanistic understanding regarding impaired reproductive health in an established model of environmental endocrine disruption.</p>
Quantification of transgene expression in GSH AAVS1 with a novel CRISPR/Cas9 based approach reveals high transcriptional variation
<p>Inderbitzin, Loosli and colleagues employ a novel method to construct and extensively characterize DNA barcode libraries and apply CRISPR/Cas9 technology for targeted insertion into the safe harbor gene AAVS1 in Jurkat cells. This technique revealed high fluctuations in gene expression in AAVS1, spanning over two logs.</p>
Transcription feedback dynamics in the wake of cytoplasmic mRNA degradation shutdown
<p>In the last decade, multiple studies demonstrated that cells maintain a balance of mRNA production and degradation, but the mechanisms by which cells implement this balance remain unknown. Here, we monitored cells’ total and recently-transcribed mRNA profiles immediately following an acute depletion of Xrn1—the main 5′-3′ mRNA exonuclease—which was previously implicated in balancing mRNA levels. We captured the detailed dynamics of the adaptation to rapid degradation of Xrn1 and observed a significant accumulation of mRNA, followed by a delayed global reduction in transcription and a gradual return to baseline mRNA levels. We found that this transcriptional response is not unique to Xrn1 depletion; rather, it is induced earlier when upstream factors in the 5′-3′ degradation pathway are perturbed. Our data suggest that the mRNA feedback mechanism monitors the accumulation of inputs to the 5′-3′ exonucleolytic pathway rather than its outputs.</p>
Modeling Word Importance in Conversational Transcripts from the Perspective of Deaf and Hard of Hearing Viewers
<ol> <li> <p><strong>MaskedPOSAugmentedData.csv</strong></p> </li> </ol> <p><strong>This file contains word embeddings of 6659 tokens augmented with POS tagging and word importance score. The embedding size for each token is a one dimensional vector of length 768 by 1. Due to augmenting POS tagging, the new feature vector becomes a size of 769 by 1. So now the total feature matrix size is 6659 by 769. In the dataset the final column represents the word importance score. </strong></p> <p><br> </p> <ol> <li> <p><strong>MaskedSentenceWithImportanceTag.csv</strong></p> </li> </ol> <p><strong>This dataset contains the newly generated sentence using standard Masking technique and their corresponding token-wise importance score.</strong><br> </p> <ol> <li> <p><strong>Data cleaning procedure</strong></p> </li> </ol> <p><strong>Using several steps, we have cleaned the “switchboard corpus” that we are using for this particular study. Here are the steps we followed:</strong></p> <p><strong>– Convert all the letter into lower case</strong></p> <p><strong>– Punctuation has be removed</strong></p> <p><strong>– Numbers or Cardinal values have been converted to text representation </strong></p> <p><strong>– We use lemmatization techniques to eradicate the possibility of multiple versions of the same word token.</strong><br> <br> </p> <p> </p> <ol> <li> <p><strong>Annotation Instruction</strong></p> </li> </ol> <p><strong>If researchers intend to produce additional masked text for further the size of the dataset, we recommend using the method we described in the paper.</strong></p> <p><strong>However, if someone wants to manually annotate the words or tokens in a dataset, it is important to remember that annotators need to put a score on each word based on its relative importance within a sentence or the information available around that text. It may not be appropriate to allow annotators to read the whole document first and then conduct annotation. Also using multiple annotators is recommended otherwise interrater agreement may not work well.</strong></p> <p> </p> <ol> <li> <p><strong>Test and training data splitting</strong></p> </li> </ol> <p><strong>As described in the paper, during our experiment, we have retained 10% of data as test data and use 90% data to train the models. For proper replication, we recommend reading our paper thoroughly. </strong></p> <p> </p> <ol> <li> <p><strong>We understand that the dataset size is relatively small for training and testing a model that might be reliable. It is important to remember that data annotation with this particular user group might be challenging. </strong></p> </li> </ol> <p> </p> <p><strong>N.B: While augmenting the new feature within the dataset, a portion of data has been excluded during the data curation phase. For validation, we have replicated the previous models so that we can measure how the dataset can perform with this newly formed dataset. </strong></p>
The transcription factor network of E. coli steers global responses to shifts in RNAP concentration
<p><span>The robustness and sensitivity of gene networks to environmental changes</span><span> is critical for cell survival. How gene networks produce specific, chronologically ordered responses to genome-wide perturbations, while robustly maintaining homeostasis, remains an open question. We analysed if short- and mid-term genome-wide responses to shifts in RNA polymerase (RNAP) concentration are influenced by the <em>known</em> topology and logic of the transcription factor network (TFN) of <em>Escherichia coli</em>. We found that, at the gene cohort level, the magnitude of the single-gene, mid-term transcriptional responses to changes in RNAP concentration can be explained by the absolute difference between the gene's numbers of activating and repressing input transcription factors (TFs)</span><span>. Interestingly, this difference is strongly positively correlated with the number of input TFs of the gene. Meanwhile, short-term responses showed only weak influence from the TFN. </span><span>Our results suggest that the global topological traits of the TFN of <em>E. coli</em> shape which gene cohorts respond to genome-wide stresses.</span></p>
Transcriptional response of mushrooms to artificial sun exposure
<p>Climate change causes increased tree mortality leading to canopy loss and thus sun-exposed forest floors. Sun exposure creates extreme temperatures and radiation, with potentially more drastic effects on forest organisms than the current increase in mean temperature. Such conditions might potentially negatively affect the maturation of mushrooms of forest fungi. A failure of reaching maturation would mean no sexual spore release and, thus, entail a loss of genetic diversity. However, we currently have a limited understanding of the quality and quantity of mushroom-specific molecular responses caused by sun exposure. Thus, to understand the short-term responses towards enhanced sun exposure, we exposed mushrooms of the wood-inhabiting forest species <i>Lentinula edodes, </i>while still attached to their mycelium and substrate, to artificial solar light (ca. 30 °C and 100.000 lux) for 5, 30, and 60 minutes. We found significant differentially expressed genes at 30 and 60 minutes. Eukaryotic Orthologous Groups (KOG) class enrichment pointed to defense mechanisms. The 20 most significant differentially expressed genes showed the expression of heat-shock proteins, an important family of proteins under heat stress. Although preliminary, our results suggest mushroom-specific molecular responses to tolerate enhanced sun exposure as expected under climate change. Whether mushroom-specific molecular responses are able to maintain fungal fitness under opening forest canopies remains to be tested.</p>
Data for The genetics monopolistic industry, as expected, is missing information two inches beyond their nose: in this case, transcripts
<p><strong>De novo transcriptome assembly is one of the many fundamental pieces of new research in genomics. It is, for example, the preferred method for studying non-model organisms, since it is easier and cheaper than building a genome, and reference methods are not possible without an existing genome. The transcriptomes of these organisms can thus reveal novel proteins and their isoforms that are implicated in such unique biological phenomena. This technique is also useful in cancer research as it makes possible to detect potentially significant chimeric transcripts in cancer and normal somatic tissues.</strong></p> <p><strong>Given that the genetics industry is organized in the form of a monopoly controlled by hidden lobbies who also control Academia, all the available software for de novo transcriptome assembly is being developed by academic researchers under the open-source paradigm. We report here, that as anyone could have very easily deduced from past experiences in other industries, this unethical form of organization in the industry is resulting in incompetence whose effects include missing a significant portion of the available information that could be obtained from some genomics studies. In this case, missing transcripts in transcriptome studies. We won't deep in on the consequences, but these could include overpricing, over costs, and failing to achieve the goals of some studies.</strong></p>
Transcriptional and metabolic changes in Candida albicans during yeast-to-hypha transition
<p>In this study hypha-associated transcriptional and metabolic changes were investigated in the opportunistich human pathogen <em>Candida albicans</em>. Three different strains were investigated the SC5314 wildtype and two filamentation-affected mutant strains <em>cph1</em>Δ<em>efg1</em>Δ and <em>hgc1</em>Δ. Three different hypha-inducing conditions were compared with two respective yeast conditions: SD minimal medium with 10% human serum or N-acetylglucosamin vs. SD minimal medium, and M199 pH7.4 vs. M199 pH4. Sampling time points were 0, 90 and 240 min. ON cultures were prepared in SD and re-inoculated in SD and grown to early log-phase prior to the experiment. Cells were collected, washed and used to inoculate the test medium at OD0.3. All cultures were generally maintained at 37°C. Obtained samples were submitted for transcriptional profiling (RNA-Seq with 10 million reads per sample and 150 bp paired-ends) and metabolomics.</p> <p>Media list:</p> <p>SD: 2% glucose, 0.17% YNB without amino acids, 0.5% ammonium sulfate, pH 4.5</p> <p>HS - SD with 10 vol% human serum (Male, AB negative)</p> <p>GlcNAc - 2% <em>N</em>-acetylglucosamine, 0.17% YNB without amino acids, 0.5% ammonium sulfate, pH 4.5</p> <p>Medium 199 with Earle's salts and L-Glu buffered to either pH 4 or 7.4 by addition of 1/5 volume of 0.1 M citric acid and 0.2 M Na<sub>2</sub>HPO<sub>4</sub> buffer</p>
Data and scripts for the manuscript of svaRetro and svaNUMT: modular packages for annotating retrotransposed transcripts and nuclear integration of mitochondrial DNA in genome sequencing data
<p>This upload include data and scripts supporting the results described in the manuscript of <em>svaRetro and svaNUMT: modular packages for annotating retrotransposed transcripts and nuclear integration of mitochondrial DNA in genome sequencing data</em><em>. </em>Detailed description of the contents can be found in README.txt.</p>
Mitigating transcription-replication conflicts: In early Drosophila embryos, rapid onset of transcription after mitosis depends on DNA replication
<p>This dataset includes the raw imaging data (exported from Volocity as TIF in Z-stacks or maximal projections) related to the publication: Mitigating transcription-replication conflicts: In early Drosophila embryos, rapid onset of transcription after mitosis depends on DNA replication, Cell Reports 2022.</p>
tidyqpcr User Interview Transcripts
<p>Transcripts of a series of tidyqpcr user interviews.</p> <p>tidyqpcr is an R package for the analysis and design of qPCR assays. https://github.com/ropensci/tidyqpcr</p> <p>The interviews were conducted over zoom and transcribed by otter.io</p>
Evolved transcriptional responses and their trade-offs after long-term adaptation of Bemisia tabaci to a marginally-suitable host
<p>Scripts and data used in the work "Evolved transcriptional responses and their trade-offs after long-term adaptation of Bemisia tabaci to a marginally-suitable host".</p> <p><strong>Abstract: </strong>Although generalist insect herbivores can migrate and rapidly adapt to a broad range of host plants, they can face significant difficulties when accidentally migrating to novel and marginally-suitable hosts. What happens, at both the performance and transcriptional levels, if these marginally-suitable hosts must be used for multiple generations before migration to a suitable host can take place, largely remains unknown. In this study, we established multigenerational colonies of the whitefly <em>Bemisia tabaci</em>, a generalist phloem-feeding species, adapted to a marginally-suitable host (habanero pepper) or an optimal host (cotton). We used reciprocal host tests to estimate the differences in performance of the populations on both hosts under optimal (30 <sup>o</sup>C) and mild-stressful (24 <sup>o</sup>C) temperature conditions, and documented the associated transcriptomic changes. The habanero pepper-adapted population greatly improved its performance on habanero pepper but did not reach its performance level on cotton, the original host. It also showed reduced performance on cotton, relative to the non-adapted population, and an antagonistic effect of the lower-temperature stressor. The transcriptomic data revealed that most of the expression changes, associated with long-term adaptation to habanero pepper, can be categorized as “evolved” with no initial plastic response. Three molecular functions dominated: enhanced formation of cuticle structural constituents, enhanced activity of oxidation-reduction processes involved in neutralization of phytotoxins and reduced production of proteins from the cathepsin B family. Taken together, these findings indicate that generalist insects can adapt to novel host plants by modifying the expression of a relatively small set of specific molecular functions.</p>
A dual selection system for directed evolution to identify allosteric transcription factor PobR variants responsive to different aromatic compounds
<p>This dataset includes all the raw data of our characterization experiments during the work titled “A dual selection system for directed evolution to identify allosteric transcription factor PobR variants responsive to different aromatic compounds”.</p>
Lietuvos Respublikos Seimo posėdžių debatų stenogramų tekstynas nuo 1990 m. kovo mėn. 10 d. = Corpus of the Transcripts of the Plenary Debates of the Seimas of the Republic of Lithuania starting from March 10th, 1990
<p>Šiame duomenų rinkinyje kaupiamos Lietuvos Respublikos <strong>Seimo posėdžių debatų stenogramos</strong>. Stenogramos parsiunčiamos automatizuotu būdu iš LR Seimo <a href="https://www.lrs.lt/sip/portal.show?p_r=35727&p_k=1&p_a=sale_ses_pos&p_kade_id=1&p_ses_id=1">portalo</a> ir/arba <a href="https://e-seimas.lrs.lt/portal/documentSearch/lt">paieškos įrankių</a> (abiejų sąrašų įrašai sutikrinami ir sudaromas bendras stenogramų sąrašas (su nuorodomis į šaltinius), kuris pridėtas prie šio duomenų rinkinio). Duomenų rinkinys apima stenogramas <strong>nuo 1990 m. kovo mėn. 10 d.</strong> iki paskutinės pilnos eilinės LR Seimo sesijos. Duomenų rinkinio atnaujinimas vykdomas pasibaigus paskutinei eilinei LR Seimo sesijai.</p> <p>Stenogramos parsiunčiamos DOC/DOCX formatais ir <strong>transformuojamos į TXT bei CSV ir XLSX formatus</strong>:</p> <p>1. Konvertavimas į TXT formatą vykdomas naudojant du įrankius: MultiDoc Converter (www.multidoc-converter.com/en/index.html) ir EmEditor (www.emeditor.com).</p> <p>2. TXT formato stenogramos konvertuojamos į struktūruotus CSV ir XLSX failus naudojant R skriptus, kurie pridedami prie šio duomenų rinkinio.</p> <p>3. Prie duomenų rinkinio pridėtas dokumentas, kuri aprašo CSV ir XLSX failų struktūrą.</p>
Revised transcript annotations for GRCh38 reference genome and Ensembl v87.
<p>Custom transcript annotations generated using the reviseAnnotations package. </p> <p>Reference genome: GRCh38<br> Ensembl version: 87</p> <p>See the GitHub page of reviseAnnotations for more details:<br> https://github.com/kauralasoo/reviseAnnotations</p>
Revised transcript annotations for GRCh37 (hg19) reference genome and Ensembl v90.
<p>Custom transcript annotations generated using the reviseAnnotations package. </p> <p>Reference genome: GRCh37<br> Ensembl version: 90</p> <p>See the GitHub page of reviseAnnotations for more details:<br> https://github.com/kauralasoo/reviseAnnotations</p>
Dataset of Brazilian Federal Senate Session Transcriptions From 2023 with Relevant Topics and Stance Detection Annotations.
<p><strong>[pt-BR] Conjunto de dados de transcrições de sessões do Senado Federal brasileiro de 2023 com anotações de tópicos relevantes e de deteção de posicionamento.</strong></p> <p><strong>Dataset description</strong></p> <p>This set contains transcript data from 203 Federal Senate sessions from the year 2023, with annotations of relevant topics and positioning detection.</p> <p>The file <strong><a href="../api/records/11106904/draft/files/meetings.csv/content" target="_blank" rel="noopener">meetings.csv</a></strong> has the following structure:</p> <p>session_id: Unique event identifier<br>speaker_name: Name of the person who gave the speech<br>party: Political party<br>speech: Speech given</p> <p>The folder <a href="../api/records/11106904/draft/files/ground_truth.zip/content" target="_blank" rel="noopener noreferrer"><strong>ground_truth.zip</strong></a> contains 6 JSON annotation files related to the detection of relevant topics and positions, and its main keys are the following:</p> <p>id_session: Unique event identifier</p> <p>response: Model response<br> list_latent_topics: List of latent topics considered by the model<br> stances: The stances of each person, according to the model, on a given topic<br>model_response_evaluation: Evaluation of the model's response according to a human annotation<br> mapping: Mapping between the topic named by the model and the topic named in the annotation<br> list_latent_topics: It contains four keys which are the lists of topics considered true positives, false positives, true negatives and false negatives.</p> <p><strong>Code used</strong></p> <p>All the code used to process the data can be found at:</p> <p><a href="https://github.com/helenbc/tcc-notas-taquigraficas">https://github.com/helenbc/tcc-notas-taquigraficas</a></p> <p><strong>[pt-BR] Descrição do conjunto de dados</strong></p> <p>Este conjunto contém dados de transcrição de 203 sessões do Senado Federal do ano de 2023, com anotações de tópicos relevantes e detecção de posicionamento.</p> <p>O arquivo meetings.csv possui a seguinte estrutura:</p> <p>session_id: Identificador único do evento </p> <p>speaker_name: Nome da pessoa que fez o discurso </p> <p>party: Partido político</p> <p>speech: Discurso proferido</p> <p>A pasta ground_truth.zip contém 6 arquivos de anotação JSON relacionados à detecção de tópicos relevantes e posições, e suas principais chaves são as seguintes:</p> <p>id_session: Identificador único do evento</p> <p>response: Resposta do modelo</p> <p> list_latent_topics: Lista de tópicos latentes considerados pelo modelo</p> <p> stances: As posições de cada pessoa, de acordo com o modelo, sobre um determinado tópico</p> <p>model_response_evaluation: Avaliação da resposta do modelo de acordo com uma anotação humana</p> <p> mapping: Mapeamento entre o tópico nomeado pelo modelo e o tópico nomeado na anotação</p> <p> list_latent_topics: Contém quatro chaves que são as listas de tópicos considerados verdadeiros positivos, falsos positivos, verdadeiros negativos e falsos negativos.</p> <p><strong>Código utilizado</strong></p> <p>Todo o código utilizado para processar os dados pode ser encontrado em:</p> <p><a href="https://github.com/helenbc/tcc-notas-taquigraficas" target="_new" rel="noreferrer">https://github.com/helenbc/tcc-notas-taquigraficas</a></p> <p> </p> <p> </p>
Fertility decline in Aedes aegypti mosquitoes is associated with reduced maternal transcript deposition and does not depend on female age
<p>Female mosquitoes undergo multiple rounds of reproduction known as gonotrophic cycles. A gonotrophic cycle spans the period from blood meal intake to egg laying. Nutrients from vertebrate host blood are necessary for completing egg development. During oogenesis, a female pre-packages mRNA into her oocytes, and these maternal transcripts drive the first two hours of embryonic development before zygotic genome activation. In this study, we profiled transcriptional changes in 1-2 hour-old <em>Aedes aegypti</em> embryos across two gonotrophic cycles. We found that homeotic genes which are regulators of embryogenesis are downregulated in embryos from the second gonotrophic cycle. Interestingly, embryos produced by <em>Ae. aegypti</em> females progressively reduced their ability to hatch as the number of gonotrophic cycles increased. We show that this fertility decline is due to increased reproductive output and not the mosquitoes' age. Moreover, we found a similar decline in fertility and fecundity across three gonotrophic cycles in <em>Ae. albopictus</em>. Our results are useful for predicting mosquito population dynamics to inform vector control efforts.</p>
"Systematic Identification of Post-Transcriptional Regulatory Modules": preprocessed data
<p>Processed functional genomic data supplementing the paper "<strong>Systematic Identification of Post-Transcriptional Regulatory Modules"</strong></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.