Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
21,320
datasets available to search
ShareScore release 0.9.0
Dataset results
21,320 results for “Transcription”
Temporal iCLIP captures co-transcriptional RNA-protein interactions
<p>Minimum datasets required to perform analysis using code from github. </p>
Stochastic motion and transcriptional dynamics of pairs of distal DNA loci on a compacted chromosome
<p>This repository contains the trajectory data from "Stochastic motion and transcriptional dynamics of pairs of distal DNA loci on a compacted chromosome".</p> <p><strong>Data sets</strong></p> <p>We provide data sets for the following constructs and imaging conditions, where we give the name of the construct and the MS2-parS genomic separation in kb:</p> <ul> <li>data_line0.csv: parS-homie-evePr-PP7, 58 kb</li> <li>data_line1.csv: parS-homie-evePr-PP7, 82 kb</li> <li>data_line2.csv: parS-homie-evePr-PP7, 88 kb</li> <li>data_line3.csv: parS-homie-evePr-PP7, 149 kb</li> <li>data_line4.csv: parS-homie-evePr-PP7, 190 kb</li> <li>data_line5.csv: parS-homie-evePr-PP7, 595 kb</li> <li>data_line6.csv: parS-homie-evePr-PP7, 3.3 Mb</li> <li>data_line3_5s.csv: parS-homie-evePr-PP7, 149 kb, 5 second time interval</li> <li>data_line0_nohomie.csv: parS-lambda-evePr-PP7, 58 kb</li> <li>data_line3_nohomie.csv: parS-lambda-evePr-PP7, 149 kb</li> </ul> <p><strong>Structure of Data</strong></p> <p>The trajectory data are provided as csv files consisting of 13 columns. The column headers are:</p> <ul> <li>cell_id: a unique cell index</li> <li>time_point: the time frame</li> <li>x_blue: x-coordinate of the locus in the blue channel (units in nm)</li> <li>y_blue: y-coordinate of the locus in the blue channel (units in nm)</li> <li>z_blue: z-coordinate of the locus in the blue channel (units in nm)</li> <li>x_green: x-coordinate of the locus in the green channel (units in nm) </li> <li>y_green: y-coordinate of the locus in the green channel (units in nm)</li> <li>z_green: z-coordinate of the locus in the green channel (units in nm) </li> <li>x_Rij: x-component of the aberration corrected 3D distance vector connecting the blue and green loci (units in nm)</li> <li>y_Rij: y-component of the aberration corrected 3D distance vector connecting the blue and green loci (units in nm)</li> <li>z_Rij: z-component of the aberration corrected 3D distance vector connecting the blue and green loci (units in nm)</li> <li>red: intensity in the red channel (a.u.)</li> <li>state: inferred state using a 3-state HMM, with entries 0 (O_off), 1 (P_off), 2 (P_on)</li> </ul>
Interview Transcriptions related to PhD Thesis "Degrowth at a Global Scale? Geographies of Chile's Fruit Production and Export between Extractivism and Socio-Ecological Transformation"
<p>The material is composed of transcriptions of interviews conducted for the empirical work of this PhD Thesis.</p> <p>Not all conducted interviews are included (which can be seen from the accompanying table); those not included are not available mostly due to lack of consent of interviewees or because certain interviews were not registered and only hand-written notes were taken.</p>
Precise modulation of transcription factor levels identifies features underlying dosage sensitivity
<p><strong>Processed data and code for "Precise modulation of transcription factor levels reveals drivers of dosage sensitivity," Naqvi et al 2022.</strong></p> <p><strong>Count/expression data</strong></p> <ul> <li>all.sub.150bpclust.greater2.500bp.merge.ATAC.counts.fulldep.3h.24h.txt.gz - ATAC-seq counts from all samples (SOX9 titration and depletion) over all reproducible ATAC-seq peak regions</li> <li>all.sub.150bpclust.greater2.500bp.merge.k27ac.txt.gz - H3K27ac ChIP-seq counts from SOX9 depletion samples over all reproducible peak regions</li> <li>all.sub.150bpclust.greater2.500bp.merge.SOX9titr.V5.counts.txt.gz - V5 (SOX9) ChIP-seq counts from partial SOX9 titration (100%, 60%, 30%, 0%) over all reproducible peak regions</li> <li>all.sub.150bpclust.greater2.500bp.merge.SOX9titr.TWIST1.in.counts.tab.txt.gz - TWIST1 and input ChIP-seq counts from partial SOX9 titration (100%, 60%, 30%, 0%) over all reproducible peak regions</li> <li>rna.salmon.7rep.txi.counts.txt.gz - RNA-seq counts from SOX9 titration samples </li> <li>rna.salmon.7rep.txi.abundance.txt.gz - RNA-seq TPM values from SOX9 titration samples </li> <li>slam.tcreadcount.txt.gz - SLAM-seq T-C conversion-containing read counts (representing newly transcribed mRNAs) from SOX9 depletion samples</li> <li>slam.readcount.txt.gz - SLAM-seq read counts (representing all mRNAs) from SOX9 depletion samples</li> </ul> <p><strong>Metadata</strong></p> <ul> <li>all.protcod.gene.features.txt.gz - Features of interest for all protein-coding genes</li> <li>all.sub.150bpclust.greater2.500bp.merge.features.txt.gz - Features of interest for all reproducible peak regions</li> <li>atac_depletion_3h_24h_design.txt - design matrix for ATAC-seq SOX9 depletion samples</li> <li>atac_titration_48h_design.txt - design matrix for ATAC-seq SOX9 titration samples</li> <li>Homo_sapiens.GRCh38.cdna.all.txid2gene.id.symbol.type.txt.gz - Ensembl transcript types (for filtering to protein-coding genes in various analyses)</li> <li>k27_depletion_3h_24h_design.txt - design matrix for H3K27ac ChIP-seq SOX9 depletion samples</li> <li>v5_sox9titration_design.txt - design matrix for V5 (SOX9) ChIP-seq SOX9 partial titration samples</li> <li>twist1_sox9titration_design.txt - design matrix for TWIST1 ChIP-seq SOX9 partial titration samples</li> <li>rna_titration_48h_design.txt - design matrix for RNA-seq SOX9 titration samples</li> <li>slam_depletion_3h_24h_design.txt - design matrix for SLAM-seq SOX9 depletion samples</li> <li>facialgwas_snpia_ld0.5.hg38.bed - SNPs in LD (r2 > 0.5) with any of the facial GWAS lead SNPs in Supplementary Table 2 of Naqvi, Hoskens, et al, Annu Rev. Hum Genet. Genom. 2022. </li> <li>facialgwas_prsendo_7e5_either_snpia_ld0.5.hg38.bed - SNPs in LD (r2 > 0.5) with the subset of the same facial GWAS SNPs that show significant (p-value < 7e-05, ~corresponding to Bonferonni-corrected p-value of 0.01) association with the PRS endophenotype GWAS in either US or UK cohort. </li> </ul> <p><strong>Scripts</strong></p> <ul> <li>atac_deseq_fitmodels_bs_parallel.R - R code for fitting bootstrapped Hill equations to all SOX9-dependent REs (computationally intensive, so has been coded for parallelization over multiple cores) <ul> <li>Input: all.sub.150bpclust.greater2.500bp.merge.ATAC.counts.fulldep.3h.24h.txt.gz, atac_titration_48h_design.txt</li> <li>Output: enh_linear_sig_aic_bsmat.txt, enh_linear_sig_aic_bsmat_enhind.txt</li> </ul> </li> <li>atac_deseq_fitmodels.R - R code for fitting Hill equations (no bootstrap) to all SOX9-dependent REs <ul> <li>Input: all.sub.150bpclust.greater2.500bp.merge.ATAC.counts.fulldep.3h.24h.txt.gz, atac_titration_48h_design.txt</li> <li>Output: enh_linear_sig_aic.rds</li> </ul> </li> <li>atac_k27_depletion_deseq.R - R code for DESeq2 analysis of ATAC and H3K27ac ChIP SOX9 depletion (3h and 24h) <ul> <li>Input: all.sub.150bpclust.greater2.500bp.merge.ATAC.counts.fulldep.3h.24h.txt.gz, all.sub.150bpclust.greater2.500bp.merge.k27ac.txt.gz, atac_depletion_3h_24h_design.txt</li> <li>Output: atac_depletion_3h_24h_deseq.txt, k27_depletion_3h_24h_deseq.txt </li> </ul> </li> <li>v5_twist1_sox9titration_deseq.R - R code for DESeq2 analysis of V5 (SOX9) and TWIST1 ChIP in partial SOX9 titration (100%, 60%, 30%, 0%) <ul> <li>Input: all.sub.150bpclust.greater2.500bp.merge.SOX9titr.V5.counts.txt.gz, all.sub.150bpclust.greater2.500bp.merge.SOX9titr.TWIST1.in.counts.tab.txt.gz, v5_sox9titration_design.txt, twist1_sox9titration_design.txt</li> <li>Output: v5_sox9titration_deseq.txt, twist1_sox9titration_deseq.txt </li> </ul> </li> <li>drm.R - Modified version of code from drc() package to prevent errors, install drc() with this version to avoid errors</li> <li>group_comparisons.Rmd - R code to compare computed parameters (i.e. ED50, Hill) between sets of REs/genes utilizing bootstrap information <ul> <li>Input: enh_linear_sig_aic_bsmat_enhind.txt, enh_linear_sig_aic_bsmat.txt.gz, enh_linear_sig_aic.rds, gene_linear_sig_aic_bsmat_enhind.txt, gene_linear_sig_aic_bsmat.txt.gz, gene_linear_sig_aic.rds, all.sub.150bpclust.greater2.500bp.merge.features.txt.gz, all.protcod.gene.features.txt.gz</li> <li>Uses: summarize_bs_helper.R</li> </ul> </li> <li>plot_re_gene_fits.Rmd - R code for plotting individual RE/gene counts and Hill/linear fits <ul> <li>Input: all.sub.150bpclust.greater2.500bp.merge.ATAC.counts.fulldep.3h.24h.txt.gz, atac_titration_48h_design.txt, rna.salmon.7rep.txi.counts.txt.gz, rna.salmon.7rep.txi.abundance.txt.gz, rna_titration_48h_design.txt</li> </ul> </li> <li>rna_deseq_fitmodels_bs_parallel.R - R code for fitting bootstrapped Hill equations to all SOX9-dependent genes (computationally intensive, so has been coded for parallelization over multiple cores) <ul> <li>Input: rna.salmon.7rep.txi.counts.txt.gz, rna.salmon.7rep.txi.abundance.txt.gz, rna_titration_48h_design.txt</li> <li>Output: gene_linear_sig_aic_bsmat.txt, gene_linear_sig_aic_bsmat_enhind.txt</li> </ul> </li> <li>rna_deseq_fitmodels.R - R code for fitting Hill equations (no bootstrap) to all SOX9-dependent genes <ul> <li>Input: rna.salmon.7rep.txi.counts.txt.gz, rna.salmon.7rep.txi.abundance.txt.gz, rna_titration_48h_design.txt</li> <li>Output: gene_linear_sig_aic.rds</li> </ul> </li> <li>slam_depletion_deseq.R - R code for DESeq2/sva analysis of SLAM-seq SOX9 depletion (3h and 24h) <ul> <li>Input: slam.tcreadcount.txt.gz, slam_depletion_3h_24h_design.txt</li> <li>Output: slam_depletion_3h_24h_deseq.txt</li> </ul> </li> <li>summarize_bs_helper.R - Helper functions for group_comparisons.Rmd</li> </ul> <p><strong>Intermediate/output files (some files are gzipped to save space, the Rscripts that output them won't gzip but they expect gzipped input when indicated)</strong></p> <ul> <li>atac_depletion_3h_24h_deseq.txt.gz - DESeq2 output of ATAC SOX9 depletion (3h and 24h)</li> <li>enh_linear_sig_aic_bsmat_enhind.txt - index to name file for SOX9-dependent RE bootstrap output</li> <li>enh_linear_sig_aic_bsmat.txt.gz - SOX9-dependent RE bootstrap output</li> <li>enh_linear_sig_aic.rds - Parameters from Hill equation fit on all SOX9-dependent REs (no bootstrap) (RDS file)</li> <li>gene_linear_sig_aic_bsmat_enhind.txt - index to name file for SOX9-dependent gene bootstrap output</li> <li>gene_linear_sig_aic_bsmat.txt.gz - SOX9-dependent gene bootstrap output</li> <li>gene_linear_sig_aic.rds - Parameters from Hill equation fit on all SOX9-dependent gene (no bootstrap) (RDS file)</li> <li>k27_depletion_3h_24h_deseq.txt.gz - DESeq2 output of H3K27ac ChIP-seq SOX9 depletion (3h and 24h)</li> <li>slam_depletion_3h_24h_deseq.txt.gz - DESeq2 output of SLAM-seq SOX9 depletion (3h and 24h)</li> <li>v5_sox9titration_deseq.txt.gz - DESeq2 output of V5 (SOX9) ChIP-seq in partial SOX9 titration (100%, 60%, 30%, 0%)</li> <li>twist1_sox9titration_deseq.txt.gz - DESeq2 output of TWIS1 ChIP-seq in partial SOX9 titration (100%, 60%, 30%, 0%)</li> </ul> <p><strong>chromatin_predictions.tar.gz (self-contained folder for chromatin-based predictions of gene expression change) contains:</strong></p> <ul> <li>ABC_6conc.sh - Bash script to calculate predicted gene expression change based on ATAC-seq fold-change at each of five SOX9 concentrations (warning: creates a number of very large additional intermediate output files). Requires as input all files in this folder except for all.sub.150bpclust.greater2.500bp.merge.ABC.5Mb.power-0.7.norm.6conc.all.total.txt</li> <li><strong> </strong>all.sub.150bpclust.greater2.500bp.merge.ABC.5Mb.power-0.7.norm.6conc.all.total.txt - ATAC-based predicted fold-change of all genes each of five SOX9 concentrations (78, 52, 25, 8, 0, in that order), relative to 100% SOX9</li> <li>all.sub.150bpclust.greater2.500bp.merge.ATAC.DMSO.counts.txt - ATAC-seq counts over all reproducible peak regions in updepleted samples</li> <li>all.sub.150bpclust.greater2.500bp.merge.bed - bed file of all reproducible peak regions</li> <li>all.sub.150bpclust.greater2.500bp.merge.deseq.allconc.lfc.txt - DESeq2 output from ATAC SOX9 titration, comparing each lowered SOX9 concentration to 100% SOX9</li> <li>all.sub.150bpclust.greater2.500bp.merge.k27ac.dmso.counts.txt - H3K27ac ChIP-seq counts over all reproducible peak regions in updepleted samples</li> <li>hg38_refGene_TSS_collapsed.bed - collapsed TSSs for all genes</li> <li>hg38.genome - genome file</li> </ul> <p> </p>
Extended Data for Single-cell Transcriptional Uncertainty Landscape of Cell Differentiation
<p>The dataset includes supporting information for the study titled "Single-cell Transcriptional Uncertainty Landscape of Cell Differentiation".</p>
Osteosarcoma-enriched transcripts paradoxically generate osteosarcoma-suppressing extracellular proteins
<p>Osteosarcoma (OS) is the common primary bone cancer that affects mostly children and young adults. To augment the standard-of-care chemotherapy, we examined the possibility of protein-based therapy using mesenchymal stem cells (MSCs)-derived proteomes and osteosarcoma-elevated proteins. While a conditioned medium (CM), collected from MSCs, did not present tumor-suppressing ability, the activation of PKA converted MSCs into induced tumor-suppressing cells (iTSCs). In a mouse model, the direct and hydrogel-assisted administration of CM inhibited tumor-induced bone destruction, and its effect was additive with Cisplatin. CM was enriched with proteins such as Calreticulin, which acted as an extracellular tumor suppressor by interacting with CD47. Notably, the level of Calr transcripts was elevated in OS tissues, together with other tumor-suppressing proteins, including histone H4, and PCOLCE. PCOLCE acted as an extracellular tumor-suppressing protein by interacting with amyloid precursor protein (APP), a prognostic OS marker with poor survival. The results supported the possibility of employing a paradoxical strategy of utilizing OS transcriptomes for the treatment of OS.</p>
Genetic control of the dynamic transcriptional response to immune stimuli and glucocorticoids at single cell resolution
<p>Supplementary Tables for article "Genetic control of the dynamic transcriptional response to immune stimuli and glucocorticoids at single cell resolution"<br> </p>
D2.1 - Demonstrator baseline and market characteristics report - Transcripts of interviews with a French company
<p>This is the transcript of the interviews with a French company for D2.1 "Demonstrator baseline and market characteristics report" defining the current baseline and the target/improved circular business models for two demonstrators and analyzing both demonstrators´ market characteristics and their impact on the target circular business models.</p>
Desmodemsus cf. armatus assembled transcripts
<p>Transcripts from <em>Desmodesmus </em>cf. <em>armatus</em> (isolated from a waste stabilisation pond and initially identified as <em>Scenedesmus</em> sp.) assembled from Illumina RNA-Seq (150 bp paired-end reads) using Trinity.</p>
Chlorella vulgaris (UTEX 259) assembled transcripts
<p>Transcripts from <em>Chlorella vulgaris</em> (UTEX 259) assembled from Illumina RNA-Seq (150 bp paired-end reads) using Trinity.</p>
Transcription factor dataset in the pea genome
<p><strong>Table 1:</strong> TF dataset identified in the pea genome. The 1807 TF dataset is shown with TF family prediction (from PlantTFDB) and the gene description available from the <em>Pisum sativum</em> genome (v1a). The best BlastP hit against <em>M. truncatula </em>(v4) and <em>A. thaliana</em> (TAIR10) proteomes is presented with both coverage (%) and e-value.</p> <p><strong>Table 2:</strong> Pea, Medicago and Arabidopsis TF datasets with unique gene identifier, predicted TF family (from PlantTFDB) and gene description (retrieved from respective genome annotation) is displayed on a second sheet.</p> <p> </p>
Single-cell transcriptional atlas of the developing Drosophila visual system (V1.1)
<p>An updated version of a single-cell transcriptional atlas of the developing Drosophila visual system (V1.1): https://doi.org/10.5281/zenodo.8097374</p> <p>Changes are described in Yoo et al. 2023 (https://doi.org/10.1101/2023.04.03.534791)</p> <pre> </pre>
Code and data for "Bordetella adenylate cyclase toxin elicits chromatin remodeling and transcriptional reprogramming that blocks differentiation of monocytes into macrophages"
<p># Code and data for "Bordetella adenylate cyclase toxin elicits chromatin remodeling and transcriptional reprogramming that blocks differentiation of monocytes into macrophages"</p> <p><br> - `pipeline` contains scripts for processing raw reads with `nf-core/rnaseq` Nextflow pipeline<br> - `setup_conda.sh` sets up a conda environment for Nextflow<br> - `samplesheet.csv` describes the samples<br> - `rna-seq-nextflow.sh` runs the analysis<br> - `data` contains the quantifications by Salmon from the pipeline and data for the CRE motif (see `analysis.Rmd` for more details)<br> - `analysis.Rmd` an R notebook containing a code to reproduce conclusions and figures in the paper and additional comments.</p> <p>The R environment (all packages) is reproducibly reconstructed via `renv` - run `renv::restore()` after loading the project.</p> <p>The generated figures and tables are stored in the `outputs` directory.</p> <p> </p>
Appendix: Ferdinando Calori Cesis' transcription errors in the Modena inventory of Pico della Mirandola's library
<p>This Appendix presents the complete list of Ferdinando Calori Cesis’ transcription errors in the Modena inventory of Pico della Mirandola’s library. The Appendix accompanies the article titled "First-language transfer in the copying of Latin manuscripts: the case of Ferdinando Calori Cesis’ transcription of the Modena inventory of Pico della Mirandola’s library", to be published in the Nordic Journal of Renaissance Studies (2024). The columns contain the following information: 1) identifier number, 2) the reading of the inventory followed by Calori Cesis' transcription, separated by a square bracket, 3) whether the error is related to a(n Arabic) number or whether it is an omission, 4) the linguistic classification of the error, 5) whether the error is due to a hypercorrection, 6) whether it involves the transcription of j as i, 7) whether it involves the transcription of u as v, 8) whether it involves the transcription of y as i, and 9) whether the word involves h.</p>
Transcription Factor Binding Regulates Chromatin Architecture
<p>Matlab scripts to calculate fiber packing ratio, sedimentation coefficient, volume, and radius of gyration.</p> <p>Representative structures for chromatin fibers in pdb format. </p>
Development of a pipeline for analyzing the gene expression profile of transcripts
<pre>Data corresponding to the scientific initiation project in bioinformatics. </pre> <pre>The data in the <em>parcial_report_input</em> folder are equivalent to raw counts, metadata and data of FPKM values of genes from the analysis performed during the first six months of the project. They correspond the counts and metadata from a previous study from renal cell carcinoma. As for the <em>F</em><em>inal_report_input</em>, it also contains counts and metadata, but from a previous study of ostesarcoma. The metadata and raw data files can be found under the accesion number hs000699.v1.p1 in dbGAP. The scripts wrote to perform pre-processing of samples, differential expression analysis, network analysis and functional annotation can be found in <a href="https://github.com/meidanis-lab/trans-pipeline">GitHub repository</a>. </pre>
Mapped data: Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency
<p>This record contains mapped sequencing data for the paper "Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency" by Nair, Ameen <em>et al</em>. It contains single-cell RNA-seq (scRNA) and single-cell ATAC-seq (scATAC) data from a time course of human dermal fibroblasts induced with Yamanaka factors OSKM using a Sendai virus based delivery system. The scRNA and scATAC data is performed at days 0, 2, 4, 6, 8, 10, 12, 14 and the final iPSCs. The experiment was re-performed and single-nucleus multiome (ATAC+RNA) was collected on days 1 and 2. </p> <p>The data is as follows:</p> <p><strong>scATAC</strong>: We used Chromap (commit <a href="https://github.com/haowenz/chromap/tree/6e97125b9">https://github.com/haowenz/chromap/tree/6e97125b9</a>, <a href="https://doi.org/10.1038/s41467-021-26865-w">https://doi.org/10.1038/s41467-021-26865-w</a>) to perform barcode correction, alignment and filtering for each of our samples. The corresponding fragment files (tab separated file containing mapped fragments with columns: chr, start, end, barcode, number of reads) and their tabix indices are available for each sample.</p> <p><strong>scRNA</strong>: We used cellranger v6.0.2 for read mapping and quantification to obtain the counts matrix. We used the GRCh38 2020-A reference. For each sample, the raw and filtered counts matrices are provided. E.g. `D0/raw_feature_bc_matrix.h5` contains an HDF5 object containing gene counts for each barcode and associated metadata for the Day 0 sample. Similarly, the files in `D0/raw_feature_bc_matrix/` contain the same gene x barcode matrix, with the counts matrix in Matrix Market format (`matrix.mtx.gz`), and gene (`features.tsv.gz`) and barcode names (`barcodes.tsv.gz`). </p> <p><strong>multiome</strong>: The ATAC and RNA components are separately processed using the same tools as mentioned above for scATAC and scRNA. Outputs are in the `snATAC` and `snRNA` subdirectories respectively. In addition, the `ATAC.RNA.bc.map.tsv` file contains a map to link snATAC barcodes to snRNA barcodes. </p>
ChromBPNet models and data: Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency
<p>This record contains ChromBPNet models and data used to train the models for the paper "Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency" by Nair, Ameen <em>et al</em>.</p> <p>`data` contains bigwigs and regions (peaks + non-peaks) used for training each of the models. See `data/README.txt` for more details.</p> <p><strong>Models:</strong></p> <p><em>Loading the model:</em></p> <p>The models were trained using tf1.14. The models are provided in h5 format for tf1.14 (py3.7) and SavedModel format for tf2.X. tf2.X tested only for py3.8-11, tf2.8-13.</p> <p>To load the models in tf1.14:</p> <pre><code class="language-python">model = tf.keras.models.load_model("path/to/model.h5")</code></pre> <p>In tf2:</p> <pre><code class="language-python">model = tf.keras.models.load_model("path/to/model_dir")</code></pre> <p>If all fails, you can load the architecture as provided in `model_arch.py` with default parameters (`bpnet_seq` for bias model and `chrombpnet` for chrombpnet model), and then load the weights using `model.load_weights` from the weights provided in the `weights` directory.</p> <p> </p> <p><em>Usage:</em></p> <p>The bias models take as input one-hot sequence of length 2000. It has 2 outputs, a vector of logits of length 2000, and 1 logcounts scalar:</p> <pre><code class="language-python"># seq_one_hot of length B x 2000 x 4 out_bias_logits, out_bias_logcounts = bias_model.predict(seq_one_hot) # out_bias_logits: B x 2000 # out_bias_logcounts: B x 1</code></pre> <p>The ChromBPNet model takes as input a one-hot sequence of length 2000, bias logits of length 2000 and bias log-counts scalar. It has the same output types as the bias model. To run the chrombpnet model to obtain predictions:</p> <pre><code class="language-python">pred_profile, pred_logcounts = chrombpnet_model.predict([seq_one_hot, out_bias_logits, out_bias_logcounts]) # pred_profile: B x 2000 # pred_logcounts: B x 1 </code></pre> <p>If you wish to obtain the "de-biased" predictions (see Methods), simply pass in zeros instead of the bias model predictions as:</p> <pre><code class="language-python">pred_profile_debiased, pred_logcounts_debiased = chrombpnet_model.predict([seq_one_hot, np.zeros((seq_one_hot.shape[0], 2000)), np.zeros((seq_one_hot.shape[0], 1))])</code></pre> <p>To obtain predicted per-base predicted counts (with or without bias):</p> <pre><code class="language-python">pred_per_base_counts = scipy.special.softmax(pred_profile, axis=-1) * (np.exp(pred_logcounts)-1) # pred_per_base_counts: B x 2000 </code></pre> <p>Note that in general predicted counts can't be compared across models as they are not corrected for sequencing depth.</p> <p> </p> <p><em>Note:</em></p> <p>All bias models used across folds are identical, except for the final intercept term in the counts output (see Methods), that is specific to each cell state, fold combination.</p> <p> </p> <p><em>Folds:</em></p> <p>The splits used for training the different folds are as below:</p> Fold Test Chromosomes Validation Chromosomes 0 chr1 chr8, chr10 1 chr2, chr19 chr1 2 chr3, chr20 chr2, chr19 3 chr6, chr13, chr22 chr3, chr20 4 chr5, chr16, chrY chr6, chr13, chr22 5 chr4, chr15, chr21 chr5, chr16, chrY 6 chr7, chr18, chr14 chr4, chr15, chr21 7 chr11, chr17, chrX chr7, chr18, chr14 8 chr9, chr12 chr11, chr17, chrX 9 chr8, chr10 chr9, chr12 <p>Remaining chromosomes were used as the training chromosome for each fold.</p>
FEDORA. Transcription and photos from the Intensive Creative Workshop II on Futurisation
<table> <tbody> <tr> <td> <p>The dataset contains transcription and photos from the Intensive Creative Workshop II on Futurisation organised within the FEDORA project. The workshop was held in Bologna on the 21st and 22nd of April 2022 and involved 9 participants from 4 European countries with diverse backgrounds: science communicators, artists, graphic designers, musicians, poets, and researchers. Participants worked in three groups, and each of them received a problem to be solved based on the current situations that young students face. Each group was also given two sheets with a summary of “The Future thinking skills” and “Stimulating engagement and critical thinking around the future and the role of science in it: the characteristics of a good activity or prototype”.</p> <p>The results are detailed in the deliverable D2.4 Multimedia report - Intensive Creative<br> workshop II on Futurisation (Confidential Deliverable), and they are part of the framework developed for deliverable D2.5: Framework for aligning science education with society: the search for new languages and narratives to enhance imagination and the capacity to talk about contemporary challenges:<a href="https://doi.org/10.5281/zenodo.7519100"> https://doi.org/10.5281/zenodo.7519100</a></p> </td> </tr> </tbody> </table>
Transcriptional patterns of sexual dimorphism and in host developmental programs in the model parasitic nematode Heligmosomoides bakeri
<p><strong>Background</strong></p> <p><em>Heligmosomoides bakeri </em>(often mistaken for <em>Heligmosomoides</em> <em>polygyrus</em>) is a promising model for parasitic nematodes with the key advantage of being amenable to study and manipulation within a controlled laboratory environment. While draft genome sequences are available for this worm, which allow for comparative genomic analyses between nematodes, there is a notable lack of information on its gene expression.</p> <p><strong>Methods </strong></p> <p>We generated biologically replicated RNA-seq datasets from samples taken throughout the parasitic life of <em>H. bakeri</em>. RNA from tissue-dwelling and lumen-dwelling worms, collected under a dissection microscope, was sequenced on an Illumina platform. <strong> </strong></p> <p><strong>Results</strong></p> <p>We find extensive transcriptional sexual dimorphism throughout the fourth larval and adult stages of this parasite and identify alternative splicing, glycosylation, and ubiquitination as particularly important processes for establishing and/or maintaining sex-specific gene expression in this species. We find sex-linked differences in transcription related to aging and oxidative and osmotic stress responses. We observe a starvation-like signature among transcripts whose expression is consistently upregulated in males, which may reflect a higher energy expenditure by male worms. We detect evidence of increased importance for anaerobic respiration among the adult worms, which coincides with the parasite's migration into the physiologically hypoxic environment of the intestinal lumen. Furthermore, we hypothesize that oxygen concentration may be an important driver of the worms encysting in the intestinal mucosa as larvae, which not only fully exposes the worms to their host's immune system but also shapes many of the interactions between the host and parasite. We find stage- and sex-specific variation in the expression of immunomodulatory genes and in anthelmintic targets. <strong> </strong></p> <p><strong>Conclusions</strong></p> <p>We examine how different the male and female worms are at the molecular level and describe major developmental events that occur in the worm, which extend our understanding of the interactions between this parasite and its host. In addition to generating new hypotheses for follow-up experiments into the worm's behavior, physiology, and metabolism, our datasets enable future more in-depth comparisons between nematodes to better define the utility of <em>H. bakeri</em> as a model for parasitic nematodes in general. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.