Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
21,320
datasets available to search
ShareScore release 0.9.0
Dataset results
21,320 results for “transcription”
Antigen-specific CD4+ T cells exhibit distinct transcriptional phenotypes in the lymph node and blood following vaccination in humans
<p><strong>Abstract: </strong><br>SARS-CoV-2 infection and mRNA vaccination induce robust CD4+ T cell responses that are critical for the development of protective immunity. Here, we evaluated spike-specific CD4+ T cells in the blood and draining lymph node (dLN) of human subjects following BNT162b2 mRNA vaccination using single-cell transcriptomics. We analyze multiple spike-specific CD4+ T cell clonotypes, including novel clonotypes we define here using Trex, a new deep learning-based reverse epitope mapping method integrating single-cell T cell receptor (TCR) sequencing and transcriptomics to predict antigen-specificity. Human dLN spike-specific T follicular helper cells (TFH) exhibited distinct phenotypes, including germinal center (GC)-TFH and IL-10+ TFH, that varied over time during the GC response. Paired TCR clonotype analysis revealed tissue-specific segregation of circulating and dLN clonotypes, despite numerous spike-specific clonotypes in each compartment. Analysis of a separate SARS-CoV-2 infection cohort revealed circulating spike-specific CD4+ T cell profiles distinct from those found following BNT162b2 vaccination. Our findings provide an atlas of human antigen-specific CD4+ T cell transcriptional phenotypes in the dLN and blood following vaccination or infection.</p> <p><strong>More Information:</strong></p> <ul> <li><strong>Preprint:</strong> <a href="https://www.researchsquare.com/article/rs-3304466/v1">Research Square.</a></li> <li><strong>Sample information</strong>: data_inventory.csv file.</li> <li><strong>Code</strong> code_github_repo.zip or at the <a href="https://github.com/ncborcherding/COVID_TCR">original github repo</a></li> <li><strong>Interactive Portal</strong>: <a href="https://cellpilot.emed.wustl.edu/">CellPilot</a></li> </ul>
Chemical structures, Cell Painting and transcriptional profiles for compound bioactivity prediction.
<p>This is the related data, both input and produced for the paper <a href="https://doi.org/10.1101/2020.12.15.422887">"Predicting compound activity from phenotypic profiles and chemical structures"</a>.</p> <p>This data can be merged with <a href="https://github.com/CaicedoLab/2023_Moshkov_NatComm">paper's GitHub repository</a> for reproduction.</p> <p>Folders and files and are described below:</p> <pre><code>├── assay_data ├── assay_matrix_discrete_270_assays.csv Assay matrix with hits for assays (270) and compounds (16170). Note that this is the final file that we used to produce splits. ├── assay_metadata.csv Assay metadata ├── broad_ids.txt List of broad ids used in this study. That is an unfiltered list of compounds required by some analysis scripts. ├── smiles.txt Same as broad_ids.txt, but SMILES strings. ├── feature_data (for 16978 compounds, can be masked with ./misc/compounds16978to16170.npy) ├── cp.npz Classical chemical features ├── ge.npz Gene expression features ├── ge_scale.npz Gene expression scaled features ├── mo.npz Morphology features (not batch corrected) ├── mobc.npz Morphology features (batch corrected) ├── misc ├── compound_analysis.npz Compounds in the dataset identified as PAINS ├── compounds16978to16170.npy Used to filter features from the bigger set of compounds to the final one ├── fingerprints.npz Calculated fingerprints of compounds, those were then used to calculate similarity ├── similarity_fingerprints.npz Similarity matrix for compounds (16978) ├── population_normalized.csv.gz Well-level morphological profiles that were used for batch-correction ├── Table for PUMA Excel file with additional data and plots ├── predictions ├── scaffold_median(mean)_AUC.csv Aggregated median(mean) AUC scores over scaffold-based cross-validation splits. In the paper, median results were reported. ├── scaffold_median(mean)_EF.csv Aggregated median(mean) enrichment factor (EF) over scaffold-based cross-validation splits. In the paper, median results were reported. ├── toprank_chemical_cv{}_hitsnorm.csv Those files are needed to create enrichment plots and contain hit rate and top rank hit rate. ├── Each folder here stands for an experiment type, the number in the folder name is a number of the split. Inside each folder there are the following elements: ├── predictions Folder with predictions for each assay-compound pair for each modality ├── 2022_01_evaluation_all_data.csv File with AUC scores for each assay for the test set in the split ├── 2022_01_evaluation_all_data_EF.csv File with enrichment factor (EF) values for each assay for the test set in the split. Those files exist only for *chemical* folders. ├── assay_matrix_discrete_train(test)_old_scaff.csv Training and test subsets of data for the split. The first column contains broad_id. ├── assay_matrix_discrete_train(test)_old_scaff.csv Same, but SMILES strings in the first column. Those files are used as input to ChemProp! Experiments in this folder are the following: - chemical Scaffold-based 5-fold cross-validation splits, the main results in the paper are reported with this series of experiments. - chemical_bal Same splits as in chemical, but training were run with ChemProp built-in data balancing. - chemical_st Same splits as in chemical, but separate models were trained for each assay. - CV Random 5-fold cross-validation splits. - GE 5-fold cross-validation splits based on same-size clustering of gene expression features. - MOBC 5-fold cross-validation splits based on same-size clustering of batch-corrected morphology features. - random 10 random splits, ~80% of compounds in the training set and the rest in the test set. ├── splitting This folder contains numpy files which help to match compounds and features to create training and test sets for a split, which can be reused in the analysis notebook for data preparation. ├── scaffold_based_split.npz Splitting for scaffold-based splits. ├── random_split_{}.npz Random split indices of test set compounds (10 files). ├── cross_validation_indicies.npz Indices for random cross-validation splits ├── GE_clusters_size_constrained.npz Indicies of clusters of same-size clustering for gene-expression features. ├── MOBC_clusters_size_constrained.npz Indices of clusters of same-size clustering for batch-corrected morphology features.</code></pre> <p> </p>
Global consensus map of human transcription factor footprints
<p>Vierstra, J. <em>et al.</em> <strong>Global reference mapping of human transcription factor footprints.</strong> <em>Nature</em><strong> </strong>583, 729–736 (2020). <a href="https://doi.org/10.1038/s41586-020-2528-x">https://doi.org/10.1038/s41586-020-2528-x</a></p> <p>Preprint @ bioRxiv: <a href="https://doi.org/10.1101/2020.01.31.927798">https://doi.org/10.1101/2020.01.31.927798</a></p> <p><strong>Contact:</strong> Jeff Vierstra (<a href="mailto:jvierstra@altius.org?subject=Consensus%20DNase%20I%20footprints">jvierstra@altius.org</a>)</p> <p>Genomic DNase I footprinting enables quantitative, nucleotide-resolution delineation of sites of transcription factor occupancy within native chromatin. We combined sampling of >67 billion uniquely mapping DNase I cleavages from >240 human cell types and states to index, with unprecedented accuracy and resolution, human genomic footprints and thereby the sequence elements that encode transcription factor recognition sites.</p> <p>Please see <a href="http://vierstra.org/resources/dgf">http://vierstra.org/resources/dgf </a>for additional information and a complete set of raw DNase I data for individual datasets. Additionally, raw data can also be accessed via the ENCODE data portal (<a href="http://encodeproject.org">http://encodeproject.org</a>) using the dataset accessions found in Supplementary Table 1.</p> <p>Code for footprint analysis and tutorials on how to access and manipulate digital genomic footprint data can be found at <a href="https://footprint-tools.readthedocs.io/en/latest/">https://footprint-tools.readthedocs.io/en/latest/</a>.</p> <p>All files herein correspond to human genome build version GRCh38 (UCSC hg38).</p> <p><strong>Dataset contents:</strong></p> <ul> <li><strong>Biosample metadata</strong> – Supplementary_Table_1.xlsx</li> <li><strong>Motif clustering metadata </strong>– Supplementary_Table_2.xlsx</li> <li><strong>ChIP-seq validation metadata </strong>–<strong> </strong>Supplementary_Table_3.xlsx</li> <li><strong>Consensus footprint coordinates and assigned motif archetypes</strong><br> TSV file (BED-format) with consensus footprint (posterior probability>0.99) coordinates and overlaps with matches to motif model clusters. The legend file contains column definitions in detail. <ul> <li>consensus_footprints_and_motifs_hg38.bed.gz</li> <li>consensus_footprints_and_motifs_legend.txt</li> </ul> </li> <li><strong>Motif archetype matches overlapping consensus footprints</strong><br> TSV file (BED-format) containing the coordinates for clustered motif model matches that overlap consensus footprints <ul> <li>collapsed_motifs_overlaping_consensus_footprints.bed.gz</li> <li>collapsed_motifs_overlaping_consensus_footprints_legend.txt</li> </ul> </li> <li><strong>Footprint occupancy matrix of consensus footprints</strong><br> Rows are same order as the consensus footprint file and columns are same order as in the metadata files. <ul> <li>consensus_index_matrix_full_hg38.txt.gz (Values are –log(1-posterior))</li> <li>consensus_index_matrix_binary_hg38.txt.gz (binary occupancy matrix, where footprints with posterior footprint probability >0.99 are considered occupied)</li> </ul> </li> <li><strong>Single nucleotide variants tested for allelic imbalance </strong><br> The legend file contains column definitions in detail. <ul> <li>genotypes.vcf.gz - Genotyping and allelic read depth for each biosample (see header for more information)</li> <li>tested_snvs_padj.bed.gz - SNVs tested for imbalance (TSV, BED-format)</li> <li>tested_snvs_padj_legend.txt</li> </ul> </li> </ul>
Pan-cancer analysis of mRNA stability for decoding tumour post-transcriptional programs
<p>Supplemental data and analysis files for Perron et al.: "Pan-cancer analysis of mRNA stability for decoding tumour post-transcriptional programs" (<a href="https://www.nature.com/articles/s42003-022-03796-w">https://www.nature.com/articles/s42003-022-03796-w</a>). The .tar.gz files contain read counts associated with various RNA-seq analyses. The .rds files are single R object files that contain various analysis results tables. The .csv files also contain analysis results or sample metadata tables. See <a href="http://csg.lab.mcgill.ca/sup/pancancer_stability/">http://csg.lab.mcgill.ca/sup/pancancer_stability/</a> for a full description of the files.</p>
Sex affects transcriptional associations with schizophrenia across the dorsolateral prefrontal cortex, hippocampus, and caudate nucleus
<p>This is supplementary data and source data for the manuscript, <em>"Sex affects transcriptional associations with schizophrenia across the dorsolateral prefrontal cortex, hippocampus, and caudate nucleus"</em>.</p> <p><strong>Abstract</strong>: Schizophrenia is a complex neuropsychiatric disorder with sexually dimorphic features, including differential symptomatology, drug responsiveness, and male incidence rate. Prior large-scale transcriptome analyses for sex differences in schizophrenia have focused on the prefrontal cortex. Analyzing BrainSeq Consortium data (caudate nucleus: n=399, dorsolateral prefrontal cortex: n=377, and hippocampus: n=394), we identified 831 unique genes that exhibit sex differences across brain regions, enriched for immune-related pathways. We observed X-chromosome dosage reduction in the hippocampus of male individuals with schizophrenia. Our sex interaction model revealed 148 junctions dysregulated in a sex-specific manner in schizophrenia. Sex-specific schizophrenia analysis identified dozens of differentially expressed genes, notably enriched in immune-related pathways. Finally, our sex-interacting expression quantitative trait loci analysis revealed 704 unique genes, nine associated with schizophrenia risk. These findings emphasize the importance of sex-informed analysis of sexually dimorphic traits, inform personalized therapeutic strategies in schizophrenia, and highlight the need for increased female samples for schizophrenia analyses.</p>
Transcribing audio data: overview and transcripts of several automatic transcription tools
<p>Throughout institutions, audio recordings are being made regularly. To be able to further process these recordings, the audio often needs to be transcribed. In order to avoid having to transcribe the audio manually, there is a wealth of tools available for doing so automatically. In this record, we present an overview of several often-used tools to automatically transcribe pre-recorded audio data, including their features, costs, and security.</p> <p>To check the quality of the tool, we also recorded an audio fragment in Dutch that we ran through all tools in this overview in March of 2022. This original audio fragment (Test_interview_20220203.mp3), the cleaned-up transcription (Test_interview_cleaned_transcript.odt) and each tool’s raw transcript of the audio fragment (Test_interview_[name-tool]_raw_[date-run]) are included in this record as well. The raw transcripts were downloaded as .docx or .txt files and the .docx files saved as .odt. No edits to the transcripts were made before saving them, except an incidental removal of a personal email address or hyperlink.</p> <p>The overview contains information and transcripts of following transcription tools:</p> <ul> <li>Amberscript</li> <li>HappyScribe</li> <li>Kaldi</li> <li>NVIVO transcription</li> <li>Sonix</li> <li>SpokenOnline</li> <li>Transcribe</li> <li>Trint</li> <li>Microsoft Word 365 Online</li> </ul> <p><strong>About</strong></p> <p>This overview was created through a collaboration between Utrecht University’s Research Data Management (RDM) Support and the <a href="https://datahub.sites.uu.nl/">DataHub SSH</a> programme situated at the faculty of Humanities.</p> <p>The details in the overview have last been updated April 19, 2022. Please note that at the time you are downloading these files, the quality of the (Dutch) speech-to-text conversion may have been improved by the respective supplier.</p>
Videos, audio transcriptions and closed captions for the workshop: Introduction to Wikidata for Maastricht University, Theory and Pratice - 15 October 2024
<h1><strong>Videos, audio transcriptions and closed captions for the workshop: Introduction to Wikidata for Maastricht University, Theory and Pratice - 15 October 2024</strong></h1> <h3><a href="https://theplant.maastrichtuniversity.nl/event/navigating-the-world-of-wikidata-for-research-science-and-cultural-heritage-2"><em>Navigating the World of Wikidata for Research, Science and Cultural Heritage</em></a></h3> <h3><em><a href="https://www.wikidata.org/wiki/Wikidata:Twelfth_Birthday/Workshop_in_Maastricht" target="_blank" rel="noopener">Wikidata's Twelfth Birthday: Workshop in Maastricht </a></em></h3> <p>Wikidata is a free, collaborative, multilingual database, collecting structured open data for anyone in the world to use. It also plays a crucial role in supporting Wikimedia projects, such as Wikipedia and Wikimedia Commons. Over the last 12 years it has strongly increased in popularity among the scientific and cultural heritage communities.</p> <p>In this 2,5 hours workshop you will learn the basics of working with Wikidata, both in theory and practice. You will learn</p> <ol> <li>The basics of Wikidata: A first look at what Wikidata is and how it works, both technically and socially (Wikidata community)</li> <li>How Wikidata can be relevant for research, science and cultural heritage (GLAM), and</li> <li>First steps in contributing to Wikidata yourself, with a focus on the topic of UM professors from past and present.</li> </ol> <p>As part of the <a title="Wikidata:Twelfth Birthday" href="https://www.wikidata.org/wiki/Wikidata:Twelfth_Birthday">Wikidata 12th Birthday celebrations</a> this workshop is open to academics, researchers, students, and professionals interested in working with Wikidata in the intersection of open data, research, and science. Whether you are new to Wikidata or looking to deepen your understanding, this session will provide valuable insights for improving your work.</p> <h2><strong>Workshop outline</strong></h2> <h3><strong>Part 1: Theory, Wikidata basics (45-60 minutes) </strong></h3> <ul> <li><strong><a href="https://zenodo.org/records/13984149/files/Wikidata%20Workshop%20-%20Theoretical%20part%20-%20Maastricht%20University%20-%2015%20October%202024.webm" target="_blank" rel="noopener">Video</a> (.webm) including <a href="https://zenodo.org/records/13984149/files/WikidataWorkshop_MaastrichtUniversity_15October2024_TheoreticalPart.txt?download=1" target="_blank" rel="noopener">audio transcription </a>(.txt) and <a href="https://zenodo.org/records/13984149/files/Wikidata%20Workshop%20-%20Theoretical%20part%20-%20Maastricht%20University%20-%2015%20October%202024.webm.en.srt?download=1" target="_blank" rel="noopener">closed captions</a> (.srt) are available below</strong></li> </ul> <p><strong>Additional materials</strong></p> <ul> <li><strong>Slides in <a href="https://zenodo.org/records/13837957/files/WikidataWorkshop_MaastrichtUniversity_15October2024_TheoreticalPart.pptx?download=1" rel="nofollow">PowerPoint</a> or <a href="https://zenodo.org/records/13837957/files/Wikidata%20Workshop%20-%20Theoretical%20part%20-%20Maastricht%20University%20-%2015%20October%202024.pdf?download=1" rel="nofollow">PDF</a> are available from <a href="https://zenodo.org/records/13837957" target="_blank" rel="noopener">https://zenodo.org/records/13837957</a></strong></li> </ul> <p><em>1) Wikidata basics</em></p> <ul> <li>What is Wikidata?</li> <li>What are the principles of Wikidata?</li> <li>How are things described in Wikidata?</li> <li>Who builds Wikidata? - The Wikidata community</li> </ul> <p><em>2) Wikidata for research, science and cultural heritage</em></p> <ul> <li>To what extent is Wikidata used throughout science, research and GLAM?</li> <li>Six anecd<em>a</em>tic cases <ol> <li>Scientometrics - Scholia</li> <li>Life and biomedical sciences</li> <li>Astronomy</li> <li>Language technology / AI / LLMs</li> <li>GLAM – KB collection highlights</li> <li>Representation of (female) scientists</li> </ol> </li> </ul> <h3><strong>Break (15 minutes)</strong></h3> <h3><strong>Part 2: Practice, contributing to Wikidata (75-90 minutes) </strong></h3> <ul> <li><strong><a href="https://zenodo.org/records/13984149/files/Wikidata%20Workshop%20-%20Practical%20part,%20UM%20professors%20-%20Maastricht%20University%20-%2015%20October%202024.webm?download=1" target="_blank" rel="noopener">Video</a> (.webm) including <a href="https://zenodo.org/records/13984149/files/WikidataWorkshop_MaastrichtUniversity_15October2024_PracticalPart_UMprofessors.txt?download=1" target="_blank" rel="noopener">audio transcription </a>(.txt) and <a href="https://zenodo.org/records/13984149/files/Wikidata%20Workshop%20-%20Practical%20part,%20UM%20professors%20-%20Maastricht%20University%20-%2015%20October%202024.webm.en.srt?download=1" target="_blank" rel="noopener">closed captions</a> (.srt) are available below</strong></li> </ul> <p><strong>Additional materials</strong></p> <ul> <li><strong>Slides in <a href="https://zenodo.org/records/13837957/files/WikidataWorkshop_MaastrichtUniversity_15October2024_PracticalPart_UMprofessors.pptx?download=1" rel="nofollow">PowerPoint</a> or <a href="https://zenodo.org/records/13837957/files/Wikidata%20Workshop%20-%20Practical%20part,%20UM%20professors%20-%20Maastricht%20University%20-%2015%20October%202024.pdf?download=1" rel="nofollow">PDF</a> are available from <a href="https://zenodo.org/records/13837957" target="_blank" rel="noopener">https://zenodo.org/records/13837957</a></strong></li> <li><strong>Handout for participants in <a href="https://zenodo.org/records/13837957/files/WikidataWorkshop_MaastrichtUniversity_15October2024_PracticalPart_HandoutForParticipants.docx?download=1" rel="nofollow">Word</a> or <a href="https://zenodo.org/records/13837957/files/WikidataWorkshop_MaastrichtUniversity_15October2024_PracticalPart_HandoutForParticipants.pdf?download=1" rel="nofollow">PDF</a></strong> <strong>are available from <a href="https://zenodo.org/records/13837957" target="_blank" rel="noopener">https://zenodo.org/records/13837957</a></strong></li> </ul> <p><strong> </strong>The goals of this hands-on part are:</p> <ul> <li>Get familiar with basic data editing via the Wikidata interface</li> <li>Understand WD data models and structures related to professors (of Maastricht University)</li> <li>Extend existing <a title="Wikidata:Wiki-wetenschappers/Universiteit Maastricht/hoogleraren" href="https://www.wikidata.org/wiki/Wikidata:Wiki-wetenschappers/Universiteit_Maastricht/hoogleraren">Wikidata items about UM professors</a>, based on information in public sources.</li> <li>If time allows: Create new Wikidata items about UM professors</li> </ul> <p>The visual slides and the textual handout explain the same content, blocks and exercises, albeit in a slightly different order.<strong> </strong></p> <h2><strong>Required preparation</strong></h2> <p>To make optimal use of our time, participants must create a Wikidata account in the weeks before the workshop. See <a href="https://www.wikidata.org/w/index.php?title=Special:CreateAccount" target="_blank" rel="noopener">https://www.wikidata.org/w/index.php?title=Special:CreateAccount</a>.</p> <p>This is important because very fresh accounts may have limited editing rights. Furthermore only 6 Wikidata accounts can be created per day from UM IP addresses, so creating a lot of new accounts during the workshop might overstretch this limit.</p> <h2><strong>Workshop leader</strong></h2> <p>This workshop was given by <a href="https://www.kb.nl/over-ons/experts/olaf-janssen">Olaf Janssen</a>, the Wikimedia coordinator of the <a href="https://www.kb.nl/over-ons/experts/olaf-janssen">Koninklijke Bibliotheek</a>, the national library of the Netherlands.</p> <p>In this role he stimulates and facilitates collaboration between the collections, knowledge, open data and staff of the KB on the one hand, and the projects of the Wikimedia movement, such as Wikipedia, Wikimedia Commons and Wikidata on the other. He is also active as a volunteer within the community. Feel free to contact Olaf via olaf.janssen(at)<a href="http://kb.nl">kb.nl</a></p> <h2><strong>Materials on Wikimedia Commons</strong></h2> <p>Photos , videos and presentations related to this event can be found on Wikimedia Commons: <a title="c:Category:Wikidata Workshop at Maastricht University, 15 October 2024" href="https://commons.wikimedia.org/wiki/Category:Wikidata_Workshop_at_Maastricht_University,_15_October_2024">Category:Wikidata Workshop at Maastricht University, 15 October 2024</a></p> <h2>Relevant URLs </h2> <ul> <li><a href="https://www.wikidata.org/wiki/Wikidata:Twelfth_Birthday/Workshop_in_Maastricht">https://www.wikidata.org/wiki/Wikidata:Twelfth_Birthday/Workshop_in_Maastricht </a></li> <li><a href="https://theplant.maastrichtuniversity.nl/event/navigating-the-world-of-wikidata-for-research-science-and-cultural-heritage-2" target="_blank" rel="noopener">https://theplant.maastrichtuniversity.nl/event/navigating-the-world-of-wikidata-for-research-science-and-cultural-heritage-2</a> + <a href="https://web.archive.org/web/20240926154021/https://theplant.maastrichtuniversity.nl/event/navigating-the-world-of-wikidata-for-research-science-and-cultural-heritage-2/">archived version</a></li> <li><a href="https://www.linkedin.com/feed/update/urn:li:activity:7244614063102529536/">https://www.linkedin.com/feed/update/urn:li:activity:7244614063102529536/</a></li> <li><a href="https://www.linkedin.com/feed/update/urn:li:activity:7245329234385059843/" target="_blank" rel="noopener">https://www.linkedin.com/feed/update/urn:li:activity:7245329234385059843/</a></li> </ul> <p>Earlier LinkedIn posts (April-May 2024, before rescheduling the worlshop to October)</p> <ul> <li><a href="https://www.linkedin.com/feed/update/urn:li:activity:7188470390858428416/" target="_blank" rel="noopener">https://www.linkedin.com/feed/update/urn:li:activity:7188470390858428416/</a></li> <li><a href="https://www.linkedin.com/feed/update/urn:li:activity:7188180802021543936/" target="_blank" rel="noopener">https://www.linkedin.com/feed/update/urn:li:activity:7188180802021543936/</a></li> <li><a href="https://www.linkedin.com/feed/update/urn:li:activity:7189276985267826691/" target="_blank" rel="noopener">https://www.linkedin.com/feed/update/urn:li:activity:7189276985267826691/</a></li> </ul> <h3> </h3>
Correction by focus: Sound files and transcriptions
<p>This data was collected to examine the prosodic reflexes of corrective focus in canonical and cleft clauses of Chinese, English, French, German. Experimental factors: FOCUS: subject|object, CONSTRUCTION: canonical|cleft; ITEMS: 4, SPEAKERS: 16 (per language). See text file ALI.txt for further details.</p> <p> </p>
Censored Books during the Portuguese Estado Novo: Transcription Dataset of the Censorship Commission's Card Files (1934-74)
<p>This spreadsheet contributes to a new bibliography of censored books under the Portuguese Estado Novo dictatorial regime.</p> <p>It contains the transcription of the data fields of 1,015 card files of censored books, which are indexed by author surname in letters A and B. These files are available at the Arquivo Nacional da Torre do Tombo, in Lisbon, Portugal (PT/TT/SNI-DSC/7, "Fichas de Autores de Obras Proibidas e Autorizadas", <a href="https://digitarq.arquivos.pt/details?id=4326912">https://digitarq.arquivos.pt/details?id=4326912</a>).</p> <p>The card files document data about the books censored by the Estado Novo Censorship Commission (1934-74). Data fields include file number, book report number, decision, date, author, title, origin, destination, observations, notices, and author or book process number.</p> <p>All card files have been photographed from very poor-quality photocopies and manually transcribed by Álvaro Seiça during 2020/21. Letters C-Z are ongoing work and will be added to this dataset.</p> <p>This project received funding from the European Union’s Horizon 2020 research and innovation programme under the Marie Skłodowska-Curie grant agreement no. 793147, ARTDEL.</p> <p>More info at https://artdel.net</p>
THCHS-30 - Aligned IPA transcriptions
<p>This upload contains aligned IPA transcriptions for the <a href="https://www.openslr.org/18/">THCHS-30 dataset from OpenSLR</a>. Thereby, punctuation is added, silence marked and duration markers for each phoneme are assigned. Furthermore, the silence on the beginning and ending of each file is marked.</p> <p>The words were transcribed using <a href="https://pypi.org/project/pypinyin/">pypinyin</a> (v0.47.1) via <a href="https://pypi.org/project/dict-from-pypinyin/">dict-from-pypinyin</a> (v0.0.1) and mapped to IPA using the <code>pinyin-ipa-map-TONE3-all.json</code> mapping <a href="https://zenodo.org/record/7525638">from here</a>. The alignment was done using <a href="https://zenodo.org/record/6796264">Montreal Forced Aligner</a> (v2.0.5) and the acoustic model <a href="https://mfa-models.readthedocs.io/en/latest/acoustic/Mandarin/Mandarin%20MFA%20acoustic%20model%20v2_0_0a.html">Mandarin MFA</a> (v2.0.0a).</p> <p>Phoneme duration markers:</p> <ul> <li><code>˘</code> -> [0, 20) percentile (speaker-wise), e.g., <code>a˥˩˘</code></li> <li>(none) -> [20, 80) percentile (speaker-wise), e.g., <code>a˥˩</code></li> <li><code>ˑ</code> -> [80, 90) percentile (speaker-wise), e.g., <code>a˥˩ˑ</code></li> <li><code>ː</code> -> [90, inf) percentile (speaker-wise), e.g., <code>a˥˩ː</code></li> </ul> <p>Thereby each phoneme (including tones) was considered on its own, i.e., phonemes with different tones were not considered together for the percentile calculation.</p> <p>Silence markers:</p> <ul> <li><code>SILX</code> -> silence at start/end of a recording (aligned on all tiers)</li> <li><code>SIL0</code> -> no silence</li> <li><code>SIL1</code> -> [0, 33.33333333) percentile of all silences (speaker-wise)</li> <li><code>SIL2</code> -> [33.33333333, 66.66666666) percentile of all silences (speaker-wise)</li> <li><code>SIL3</code> -> [66.66666666, inf) percentile of all silences (speaker-wise)</li> </ul> <p>Files:</p> <ul> <li><code>grids.zip</code> <ul> <li>contains TextGrids for all audio files containing three tiers <code>words</code>, <code>phonemes</code> and <code>transcription</code> <ul> <li><code>words</code> contains the aligned Chinese words</li> <li><code>phonemes</code> contains the IPA pronunciations including silence markers at start and end (<code>SILX</code>)</li> <li><code>transcription</code> contains unaligned phonemes including punctuation and word boundary labels (<code>SIL0</code>)</li> </ul> </li> <li>the folder structure is equal to one from the dataset</li> </ul> </li> <li><code>grids-sdp.zip</code> <ul> <li>same as <code>grids.zip</code> except the folder structure is the one from <a href="https://pypi.org/project/speech-dataset-parser/">speech-dataset-parser</a> (v0.0.4)</li> </ul> </li> <li><code>preview-start/middle/end.png</code> <ul> <li>preview of the first TextGrid from speaker <code>A2</code> opened in Praat in different positions</li> </ul> </li> <li><code>words-vocabulary.txt</code> <ul> <li>contains all Chinese words from tier <code>words</code></li> </ul> </li> <li><code>phonemes-vocabulary.txt</code> <ul> <li>contains all phonemes from tier <code>phonemes</code></li> </ul> </li> <li><code>transcription-vocabulary.txt</code> <ul> <li>contains all phonemes/punctuation from tier <code>transcription</code></li> </ul> </li> <li><code>phonemes-durations.pdf</code> <ul> <li>contains the plotted phoneme duration distribution of tier <code>phonemes</code></li> </ul> </li> <li><code>phonemes-durations-simple.pdf</code> <ul> <li>contains the plotted phoneme duration distribution of tier <code>phonemes</code> if all duration markers are ignored</li> </ul> </li> <li><code>phonemes-durations-simple-toneless.pdf</code> <ul> <li>contains the plotted phoneme duration distribution of tier <code>phonemes</code> if all duration markers and tones are ignored</li> </ul> </li> <li><code>pronunciations-broad.dict</code> <ul> <li>contains the broad pronunciations for each word including punctuation but not duration markers</li> <li>e.g., <code>一下。 i˥ ɕ j a˥˩ 。</code></li> </ul> </li> <li><code>pronunciations-narrow.dict</code> <ul> <li>contains the narrow pronunciations for each word including punctuation, duration markers and weights (= occurrence) over all speakers</li> <li>e.g., <code>一下。 3 i˥ ɕ j a˥˩ː 。</code></li> </ul> </li> <li><code>pronunciations-narrow-speakers.zip</code> <ul> <li>contains the narrow pronunciations separated for each speaker</li> </ul> </li> <li><code>script.sh</code> <ul> <li>contains the script to reproduce all results</li> <li>error in line 32: replace with <code>speech-dataset-parser==0.0.4</code></li> </ul> </li> </ul>
Supplemental data files: Beyond the reference: gene expression variation and transcriptional response to RNAi in C. elegans
<p>This dataset holds all non-GEO-hosted supplemental data files for manuscript "Beyond the reference: gene expression variation and transcriptional response to RNAi in <em>C. elegans</em>". Please see the linked preprint/publication for full details.</p> <p>The PDF _guide_to_datafiles.pdf gives details on the format and content of each of the included files.</p>
Data to reproduce analysis in "Systematic analysis of transcriptional and epigenetic effects of genetic variation in Kupffer cells enables discrimination of cell intrinsic and environment-dependent mechanisms"
<p>Here you can find the datasets necessary to reproduce all analyses described in the Glass lab paper by <a href="https://www.biorxiv.org/content/10.1101/2022.09.22.509046v1">Bennett et al</a>. The python and R code for reproducing analysis and figures can be found on our linked <a href="https://github.com/HunterBennett/KupfferCell_NaturalGeneticVariation">github repository.</a></p> <p>Briefly, this paper explores the effect of natural genetic variation <em>in vivo</em>, using Kupffer cells as a model cell type. We collect and analyze transcriptional and epigenetic data (ATAC-seq, H3K27Ac ChIP-seq) to identify putative <em>trans</em> regulators driving differential gene expression across inbred strains of mice. Additionally, we provide evidence that <em>trans</em> effects control a majority of strain differential genes at homeostasis while <em>cis</em> effects dominate the transcriptional response to an external signal (lipopolysaccharide).</p> <p>References:</p> <p>Hunter Bennett, Ty D. Troutman, Enchen Zhou, Nathanael J. Spann, Verena M. Link, Jason S. Seidman, Christian K. Nickl, Yohei Abe, Mashito Sakai, Martina P. Pasillas, Justin M. Marlman, Carlos Guzman, Mojgan Hosseini, Bernd Schnabl, Christopher K. Glass bioRxiv 2022.09.22.509046; doi: <a href="https://doi.org/10.1101/2022.09.22.509046">https://doi.org/10.1101/2022.09.22.509046</a></p> <p> </p>
Allele-specific quantitation of ATXN3 and HTT transcripts in polyQ disease models.
<p>Precise values obtained during the research that led to the publishing of scientific paper entitled 'Allele-specific quantitation of ATXN3 and HTT transcripts in polyQ disease models'.</p>
Oral history transcripts from the H.J. Andrews Experimental Forest Program, 1996 to 2018
Oral history interviews have been conducted over the last two decades with members of the Andrews Forest community who provided valuable historical information about the program and related issues. On the occasion of the 50th anniversary of the experimental forest (1998), history professor Max Geier (Western Oregon University) conducted 33 interviews with individuals (or pairs of people) and five research groups from 1996-1998. About 20 years later (2013-2018) historian Sam Schmieding (Oregon State University) conducted an additional 10 oral histories, including some with people who had been interviewed by Geier 20 years earlier. Several additional relevant oral histories with people who have been important in the history of the Andrews Forest have been conducted and are also included in this collection. This data package includes an inventory and transcripts of these oral history interviews including brief biosketches of interviewees.
Supplementary Data to "Disparate regulation of Smad3 phosphorylation and collagen gene transcription by full-length IL-33"
<p>These are Supplementary Figures for the article "Disparate regulation of Smad3 phosphorylation and collagen transcription by full-length IL-33"</p>
Human T-box transcription factor T (Brachyury); A Target Enabling Package
<p>Chordoma is a rare cancer occurring along the spinal cord (OMIM: <a href="https://www.omim.org/entry/215400">215400</a>). Chordoma is derived from an embryonic tissue, the notochord, and over-expresses the embryonic transcription factor T-box transcription factor T, the homologue of mouse Brachyury. Chordomas are “genomicaly silent” cancers that do not carry an extensive mutation load. Recent studies indicate that expression of TBXT is essential for persistence and growth of chordoma cells. As TBXT is not expressed in any post-embryonic tissues, it could be an excellent target for treatment of chordoma. The long-term aim of this project is to test whether TBXT can be targeted with small molecules with sufficient affinity and specificity to be therapeutically useful. In this TEP we have determined crystal structures of the DNA-binding domain (DBD) of TBXT with and without cognate DNA oligonucleotides. The DNA-free protein crystals were used in a high-throughput fragment screen to identify 29 fragments bound in 6 clusters. The crystal structures of the bound fragments provide starting points for development of stronger binders which could be used to disrupt TBXT activity or to induce the degradation of the protein through a Proteolysis-targeting chimeric molecule (PROTAC) approach.</p>
Dataset for Automated Medical Transcription
<p>We generated this dataset to train a machine learning model for automatically generating psychiatric case notes from doctor-patient conversations. Since, we didn't have access to real doctor-patient conversations, we used transcripts from two different sources to generate audio recordings of enacted conversations between a doctor and a patient. We employed eight students who worked in pairs to generate these recordings. Six of the transcripts that we used to produce this recordings were hand-written by Cheryl Bristow and rest of the transcripts were adapted from Alexander Street which were generated from real doctor-patient conversations. Our study requires recording the doctor and the patient(s) in seperate channels which is the primary reason behind generating our own audio recordings of the conversations. </p> <p>We used Google Cloud Speech-To-Text API to transcribe the enacted recordings. These newly generated transcripts are auto-generated entirely using AI powered automatic speech recognition whereas the source transcripts are either hand-written or fine-tuned by human transcribers (transcripts from Alexander Street). </p> <p>We provided the generated transcripts back to the students and asked them to write case notes. The students worked independently using a software that we developed earlier for this purpose. The students had past experience of writing case notes and we let the students write case notes as they practiced without any training or instructions from us.</p> <p><strong>NOTE:</strong> Audio recordings are not included in Zenodo due to large file size but they are available in the <a href="https://github.com/nazmulkazi/dataset_automated_medical_transcription">GitHub</a> repository.</p>
EoRNA, a barley gene and transcript abundance database
<p>A high-quality, barley gene reference transcript dataset (BaRTv1.0, Rapazote-Flores et al. 2019), was used to quantify gene and transcript abundances from 22 RNA-seq experiments, covering 843 separate samples. Using the abundance data we developed a Barley Expression Database (EoRNA* – Expression of RNA) to underpin a visualisation tool that displays comparative gene and transcript abundance data on demand as transcripts per million (TPM) across all samples and all the genes. EoRNA provides gene and transcript models for all of the transcripts contained in BaRTV1.0, and these can be conveniently identified through either BaRT or HORVU gene names, or by direct BLAST of query sequences. Browsing the quantification data reveals cultivar, tissue and condition specific gene expression and shows changes in the proportions of individual transcripts that have arisen via alternative splicing. TPM values can be easily extracted to allow users to determine the statistical significance of observed transcript abundance variation among samples or perform meta analyses on multiple RNA-seq experiments. * Eòrna is the Scottish Gaelic word for Barley</p>
Supporting data to M. Cavallaro, et al., 3'-5' crosstalk contributes to transcriptional bursting, 2019
<p>This repository contains supporting data to reference [1]. Please cite [1] if you find this repository useful. The data include:</p> <ul> <li>Flow cytometry data of HBB and HIV transgenes' expression in `.fcs` format.</li> <li>NanoString data for HIV expression.</li> <li>smFISH data for HBB and Akt1 gene expression.</li> </ul> <p>[1] M. Cavallaro, <em>et al.</em>, 3'-5' interactions contribute to transcriptional bursting, bioR$\chi$iv 514174. <a href="https://doi.org/10.1101/514174">https://doi.org/10.1101/514174</a></p>
Transcript quantification data from Varabyou et al. 2020 simulations
<p>Transcript quantification data from different methods and configurations and simulated counts for simulated samples. The quantification results are in quants.tar.gz, and the true transcript fragment counts are in true_counts.tar.gz.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.