Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
48,977
datasets available to search
ShareScore release 0.7.1
Dataset results
48,977 results for “Genes”
Selection against admixture and gene regulatory divergence in a long-term primate field study
<p><strong>Selection against admixture and gene regulatory divergence in a long-term primate field study</strong><br> <em>Vilgalys & Fogel et al. (bioRxiv)</em></p> <ul> <li><a href="https://zenodo.org/api/files/7cb721ac-b8e0-4dc0-a91b-fd54cc70f8d7/Panubis1.0_to_hg38.chain.gz">Panubis1.0_to_hg38.chain.gz</a>; <a href="https://zenodo.org/api/files/7cb721ac-b8e0-4dc0-a91b-fd54cc70f8d7/hg38_to_Panubis1.0.chain.gz">hg38_to_Panubis1.0.chain.gz</a>: Liftover chain files between Panubis1.0 and hg38. </li> <li><a href="https://zenodo.org/api/files/7cb721ac-b8e0-4dc0-a91b-fd54cc70f8d7/amboseli_LCLAE_tracts.txt.gz">amboseli_LCLAE_tracts.txt.gz</a>: Local ancestry calls for 442 wild, hybrid baboons studied as part of the Amboseli Baboon Research Project. Local ancestry was called using LCLAE and is represented by a 0 for homozygous yellow ancestry, 2 for homozygous anubis ancestry, and 1 for heterozygous ancestry. Each row in the file has a genomic position (chromosome, start, and end), local ancestry call, and the individual for whom the call was made. </li> <li><a href="https://zenodo.org/api/files/7cb721ac-b8e0-4dc0-a91b-fd54cc70f8d7/masked_yellow_and_anubis.vcf.gz?versionId=25e26878-bca3-4667-9668-9e19424bc23e">masked_yellow_and_anubis.vcf.gz</a>: Genotype calls for non-Amboseli yellow and anubis baboons, after masking to remove putative introgressed ancestry. </li> <li>A time-stamped version of the code is included here, and also available on GitHub at <a href="http://github.com/TaurVil/VilgalysFogel_Amboseli_admixture">github.com/TaurVil/VilgalysFogel_Amboseli_admixture</a>. </li> </ul>
Supplementary data for publication Global distribution of mcr gene variants in 214K metagenomic samples
<p># Supplementary data for the manuscript "Global distribution of mcr gene variants in 214,095 metagenomic samples"</p> <p>SD1_mapped_runids.csv : tab-separated file with columns of run_accessions downloaded from ENA and whether the metagenome were positive for at least one of the mcr genes.</p> <p>SD2_mcr_df.csv : compositional table of mcr-positive metagenomes with associated metadata (collection_year, country, and host) for each run_accession, as well as mapping results.</p> <p>SD3_mcr_contigs.fa : FASTA file with contigs carrying mcr genes. The header contains the run_accession ID.</p> <p>SD4_aldex2_results.csv: CSV file containing ALDEx2 results. The columns are as follows:<br> * group: metadata category (year, country or host). If the column contains more than one label, e.g., "Denmark - 2020 - Pigs", significance is tested within Danish pig samples from 2020.<br> * rab.all: median clr value for all samples in the feature<br> * rab.win.conditionA: median clr value for the condition A of samples<br> * rab.win.conditionB: median clr value for the condition B of samples<br> * diff.btw: median difference in clr values between A and B conditions<br> * diff.win: median of the largest difference in clr values within A and B conditions<br> * effect : median effect size: diff.btw / max(diff.win) for all instances<br> * overlap : proportion of effect size that overlaps 0 (i.e. no effect)<br> * we.ep: Expected P value of Welch’s t test<br> * we.eBH: Expected Benjamini-Hochberg corrected P value of Welch’s t test<br> * wi.ep: Expected P value of Wilcoxon rank test<br> * wi.eBH: Expected Benjamini-Hochberg corrected P value of Wilcoxon test<br> * parts: gene name<br> * conditionA: label of condition A that is compared against condition B<br> * conditionB: label of condition B that is compared against condition A<br> * conditions.A.vs.B: label to explain condition A compared against condition B<br> NOTE: see for more explanation of the output of ALDEx2 https://www.bioconductor.org/packages/release/bioc/vignettes/ALDEx2/inst/doc/ALDEx2_vignette.html#5_ALDEx2_outputs</p> <p>SD5: Multi-VCF file containing SNP information on mcr alleles. Can be used to construct consensus sequences.</p> <p>SD6: FASTA file containing all unique consensus sequences reported in the manuscript.</p> <p>SD7: CSV file with an overview of which metagenome contains which unique consensus sequence.</p>
Chicken Immunoglobulin Gene Conversion Full Dataset
<p>Full dataset containing PacBio sequencing data and gene conversion information output from Brepconvert, for the immunoglobulin heavy and light chain of six 3 week old Rhode Island Red chickens studied as part of the publication Diversification of Antibodies by Gene Conversion in the Domestic Chicken (Gallus gallus domesticus). </p>
Raw data for: "Postsynaptic autism spectrum disorder genes and synaptic dysfunction"
<p>Schematic illustration representing postsynaptic proteins associated to ASD. These proteins are involved in different synaptic functions, either directly (ion channels and glutamate receptors), or indirectly, including transmembrane heterophilic (NLGNs) and homophilic (NrCAM) cell-adhesion molecules, and scaffolding proteins (PSD-95, Shank, Homer), that link transmembrane and membrane-associated protein complexes with the underlying actin cytoskeleton. Additional cellular functions may influence synaptic activity in ASD, such as alternative splicing (PTEN, RBFOX1, nSR100/SRRM4), RNA editing (FMR1, FXR1), transcription (FOXP1, FOXP2, TBR1, TSHZ3), translation (FMR1), degradation (UBE3A), and mitochondrial activity (AGC1).</p>
COLONOMICS - predictive models for normal colon gene expression and DNA methylation for TWAS and MWAS
<p>We provide significant SNP prediction models derived from the COLONOMICS data (<a href="https://www.colonomics.org">https://www.colonomics.org</a>). Genotypes were obtained by Affymetrix 6.0 array, imputed to TopMed panel. Gene expression was obtained from Affymetrix U219 array, DNA methylation was obtained with Illuminan 450K array and miRNA expression was obtained by NGS. We provide SNP prediction models for 1,758 genes, 30,530 CpG probes and 38 miRNAs obtained from colon normal biopsy samples. These features can be predicted from SNPs located within ±1Mb, which we assumed they act through cis mechanisms. We include the model’s summary statistics and corresponding SNP weights in SQLite objects. Models were trained using the elastic net procedure employed in the PredictDB pipeline (<a href="https://predictdb.org/">https://predictdb.org</a>), according to which only models with a predictive performance p-value < 0.05 and R<sup>2</sup> > 0.1 are considered significant. We adjusted the models by basic covariates, i.e., sex, age, tissue type and colon anatomic location where biopsies were collected (left and right colon). Genome coordinates refer to GRCh37/hg19.</p>
Gene expression count data from human post-mortem spinal cord
<p>Gene expression data from human post-mortem tissue for three spinal cord sections (cervical, thoracic and lumbar) from amyotrophic lateral sclerosis (ALS) patients and non-neurological disease controls. RNA sequencing performed as part of the New York Genome Center ALS Consortium.</p> <p>Analysis workbooks: <a href="https://jackhump.github.io/ALS_SpinalCord_QTLs/">https://jackhump.github.io/ALS_SpinalCord_QTLs/</a> </p> <p>Preprint describing results: <a href="https://www.medrxiv.org/content/10.1101/2021.08.31.21262682v1">https://www.medrxiv.org/content/10.1101/2021.08.31.21262682v1</a> </p> <p><strong>Sample sizes:</strong></p> <table align="left"> <tbody> <tr> <td> <p><strong>Region</strong></p> </td> <td> <p><strong>Control</strong></p> </td> <td> <p><strong>ALS</strong></p> </td> </tr> <tr> <td> <p>Cervical</p> </td> <td> <p>35</p> </td> <td> <p>139</p> </td> </tr> <tr> <td> <p>Thoracic </p> </td> <td> <p>10</p> </td> <td> <p>42</p> </td> </tr> <tr> <td> <p>Lumbar</p> </td> <td> <p>32</p> </td> <td> <p>122</p> </td> </tr> </tbody> </table> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p><strong>Library preparation</strong></p> <p>RNA was extracted from flash-frozen postmortem tissue using TRIzol (Thermo Fisher Scientific) chloroform, followed by column purification (RNeasy Minikit, QIAGEN). RNA integrity number (RIN) was assessed on a Bioanalyzer (Agilent Technologies). RNA-Seq libraries were prepared from 500ng total RNA using the KAPA Stranded RNA-Seq Kit with RiboErase (KAPA Biosystems) for rRNA depletion and Illumina-compatible indexes (NEXTflex RNA-Seq Barcodes, NOVA-512915, PerkinElmer, and IDT for Illumina TruSeq UD Indexes, 20022370). Pooled libraries (average insert size: 375 bp) passing the quality criteria were sequenced either on an Illumina HiSeq 2500 (125 bp paired end) or an Illumina NovaSeq (100 bp paired-end). The samples had a median sequencing depth of 42 million read pairs, with a range between 16 and 167 million read pairs.</p> <p><strong>Data processing</strong></p> <p>Samples were uniformly processed using RAPiD-nf, an efficient RNA-Seq processing pipeline implemented in the NextFlow framework. Following adapter trimming with Trimmomatic (version 0.36), all samples were aligned to the hg38 build (GRCh38.primary_assembly) of the human reference genome using STAR (2.7.2a), with indexes created from GENCODE, version 30. Gene expression was quantified using RSEM (1.3.1) using GENCODE v30. Quality control was performed using SAMtools and Picard, and the results were collated using MultiQC. Various technical metrics for sequencing quality control are provided in the metadata. Estimated read counts and normalised transcripts per million (TPM) matrices provided for each tissue.</p> <p><strong>Provided data:</strong></p> <p><em>gencode.v30.gene_meta.tsv.gz</em> - tab separated table with columns "genename", the HGNC gene symbol, and "geneid" the Ensembl ID, as set in the GENCODE v30 comprehensive annotation.</p> <p>For {tissue} in Cervical_Spinal_Cord, Thoracic_Spinal_Cord, Lumbar_Spinal_Cord:</p> <p><em>{tissue}_metadata.tsv.gz </em>- metadata describing each sample. Each row describes a sample. Descriptions of each column below.</p> <p><em>{tissue}_gene_tpm.tsv.gz</em> - the normalised TPM values from RSEM for all 58,884 genes in GENCODE v30. Each row describes a gene and each column describes a sample.</p> <p><em>{tissue}_gene_counts.tsv.gz</em> - the estimated read counts from RSEM for all 58,884 genes in GENCODE v30. Each row describes a gene and each column describes a sample.</p> <p><strong>Metadata Column Description</strong></p> <p><em>rna_id</em> - de-identified sample ID for each unique RNA-seq sample</p> <p><em>dna_id</em> - de-identified donor ID for each patient enrolled in the study</p> <p><em>site_id</em> - de-identified site name for each contributing site</p> <p><em>tissue</em> - name of tissue/region</p> <p><em>age_rounded</em> - age at death, rounded to nearest decade</p> <p><em>sex</em> - biological sex of donor</p> <p><em>subject_group</em> - long form disease group</p> <p><em>disease</em> - short form disease group</p> <p><em>site_of_motor_onset </em>- for ALS donors, where did symptoms start?</p> <p><em>disease_duration</em> - for ALS donors, how long did donor live with disease? </p> <p><em>mutations </em>- any known ALS gene mutations</p> <p><em>library_prep</em> - type of library preparation method used</p> <p><em>seq_platform </em>- sequencing platform used for sequencing</p> <p><em>rin</em> - RNA integrity number, 0-10</p> <p><em>c9orf72_repeat_size</em> - estimated C9orf72 repeat expansion size</p> <p><em>gPC1 - gPC5 </em>- principal component of genetic ancestry from whole genome sequencing</p> <p>Remaining metadata columns are from Picard - see here: <a href="http://broadinstitute.github.io/picard/picard-metric-definitions.html#RnaSeqMetrics">http://broadinstitute.github.io/picard/picard-metric-definitions.html#RnaSeqMetrics</a> </p> <p> </p>
Distinguishing between canonical and non-canonical tRNA genes reveals that Thermococcaceae adhere to the standard archaeal tRNA gene set
<p><strong>Abstract</strong></p> <p>Automated genome annotation is an essential tool for extracting biological information from sequence data. The identification and annotation of tRNA genes is frequently performed by the software package tRNAscan-SE, the output of which is listed – for selected genomes – in the Genomic tRNA database (GtRNAdb). Given the central role of tRNA in molecular biology, the accuracy and proper application of tRNAscan-SE is important for both interpretation of the output, and continued improvement of the software. Here, we report a manual annotation of the predicted tRNA gene sets for 20 complete genomes from the archaeal taxon Thermococcaceae. According to GtRNAdb, these 20 genomes contain a number of putative deviations from the standard set of canonical tRNA genes in Archaea. However, manual annotation reveals that only one represents a true divergence; the other instances are either (i) non-canonical tRNA genes resulting from the integration of horizontally transferred genetic elements, or CRISPR-Cas activity, or (ii) attributable to errors in the input DNA sequence. To distinguish between canonical and non-canonical archaeal tRNA genes, we recommend using a combination of automated pseudogene detection by tRNAscan-SE and the tRNAscan-SE isotype score, greatly reducing manual annotation efforts and leading to improved predictions of tRNA gene sets in Archaea.</p> <p> </p> <p><strong>Repository contents</strong></p> <p><strong>01_workflow_tRNAscanSE_predictions_210archaea.html </strong>contains the workflow and graphical output for tRNA gene set predictions in 20 Thermococcaceae genomes and 210 archaeal genomes. Files 03 to 06 below are the files quoted in this workflow.</p> <p><strong>02_workflow_tRNAscanSE_predictions_210archaea.Rmd </strong>contains the markdown file associated with 01_workflow_tRNAscanSE_predictions_210archaea.html above.</p> <p><strong>03_thermo_trnas_GtRNAdb.txt</strong><strong> </strong>contains the predicted tRNA gene sets of 20 Thermococcaceae genomes as listed on GtRNAdb (Data Release 19 (June 2021)).</p> <p><strong>04_Archaea_genome_list.txt </strong>contains the details of all 217 archaeal genomes listed on GtRNAdb (Data Release 19 (June 2021)). The seven genomes for which the NCBI genome sequences were no longer available are indicated by #### preceding the name.</p> <p><strong>05_thermo_tRNAs_genome.txt</strong><strong> </strong>contains the predicted tRNA gene sets of 20 Thermococcaceae genomes as predicted by locally run tRNAscan-SE (version 2.0.6), with standard settings for Archaea (option -A). To display the output, options -H and --detail were added. We note that pseudogene detection is active under these conditions.</p> <p><strong>06_Archaea_210_GtRNAdb_tRNAs.txt </strong>contains the predicted tRNA gene sets of the 210 archaeal genomes as listed on GtRNAdb (Data Release 19 (June 2021)).</p> <p><strong>07_Archaea_210genomes_tRNAs.txt</strong> contains the predicted tRNA gene sets of the 210 archaeal genomes as predicted by locally run tRNAscan-SE (version 2.0.6), with standard settings for Archaea (option -A). To display the output, options -H and --detail were added. We note that pseudogene detection is active under these conditions.</p> <p><strong>08_NCBI_genomes.zip</strong> contains the NCBI GenBank genome sequence files used in this study. These include the 20 Thermococcaceae genomes, the wider 210 archaeal genomes, and several others of interest. </p> <p><strong>09_phylogeny.tar.zip</strong> contains the data used to draw a phylogenetic tree for the 20 Thermococcaceae organisms. The folder includes a file listing the details of all data in the folder (Readme.md), a workflow file (workflow_UndinMarkers_v2.md), and data folders.</p> <p> </p> <p><strong>Notes</strong></p> <p>The extended TIGRFAM database referred to in the phylogenetic tree construction process can be found at <a href="https://zenodo.org/record/3839790#.YjByaVzMI3g">https://zenodo.org/record/3839790#.YjByaVzMI3g</a></p> <p>The perl script used during phylogenetic tree construction, catfasta2phyml.pl, is available in the GitHub repository <a href="https://github.com/nylander/catfasta2phyml">https://github.com/nylander/catfasta2phyml</a></p> <p>tRNAscan-SE is a freely available resource available online (<a href="http://lowelab.ucsc.edu/tRNAscan-SE/">http://lowelab.ucsc.edu/tRNAscan-SE/</a>)</p> <p>GtRNAdb is a publicly accessible resource available online (<a href="http://gtrnadb.ucsc.edu/">http://gtrnadb.ucsc.edu/</a>)</p> <p>NCBI is a publicly accessible resource available online (<a href="https://www.ncbi.nlm.nih.gov/">https://www.ncbi.nlm.nih.gov/</a>)</p> <p>rrnDB is a publicly accessible resource available online (<a href="https://rrndb.umms.med.umich.edu/">https://rrndb.umms.med.umich.edu/</a>)</p> <p>BLAST is a publicly accessible resource available online (<a href="https://blast.ncbi.nlm.nih.gov/Blast.cgi">https://blast.ncbi.nlm.nih.gov/Blast.cgi</a>)</p>
Defining the genes required for survival of Mycobacterium bovis in the bovine host offers novel insights into the genetic basis of survival of pathogenic mycobacteria
<p>Supplementary dataset from "<strong><em>Defining the genes required for survival of Mycobacterium bovis in the bovine host offers novel insights into the genetic basis of survival of pathogenic mycobacteria</em></strong>"</p> <p> </p> <p><strong>Supplementary Figure legends</strong></p> <p><strong>Figure S1. Illustration of the transposon insertions around the <em>M. bovis </em>genome. </strong>Sequencing of the input library showed that transposon insertions were evenly distributed around the genome and 27,419 of the permissible 66,931 thymine–adenine dinucleotide (TA) sites contained an insertion representing an insertion density of ~41%. The outer ring are the genomic coordinates, the blue lines represent transposon insertions and the gray boxes indicate regions of that did not have any insertions. Plot made with Circlize (Gu et al, 2014).</p> <p> </p> <p><strong>Figure S2. Diversity of the output library isolated from lung and thoracic lymph node lesions compared to the input library. </strong>On average, libraries recovered from lung lesions contained 14,456 unique mutants and those recovered from the lymph nodes contained an average of 16,210 unique mutants. Insertion density is represented as a proportion of the TA sites that contained insertions. The numbers on the x-axis refer to the sequencing file from that sample and come from individual animals (Bioproject ID: PRJNA816175, Submission ID: SUB11067380).</p> <p> </p> <p><strong>Figure S3. Volcano plots showing the distribution of log<sub>2</sub> fold-changes and -log<sub>10</sub> of adjusted p-values for representative lung (A) and lymph node (B) samples. </strong>Adjusted p-values (BH-fdr correction) < 0.000001 cluster at the limits of the plot and precision reflects the number of resampling iterations (10,000).</p> <p> </p> <p><strong>Figure S4. Scatterplot of mean log<sub>2</sub> fold change per gene for all lung samples against all thoracic lymph node samples</strong>. Correlation between mean log<sub>2</sub> fold change among genes between the tissues was calculated with Spearman's ranked correlation, = 0.878, p-value < 2.2e-16.</p> <p> </p> <p><strong>Figure S5. Fold-changes caused by transposon insertions in <em>RD1<sup>BCG</sup></em> and <em>RD1<sup>MIC</sup> </em>in the lungs and lymph nodes of infected cattle. </strong>Boxplot for log<sub>2 </sub>fold-changes in genes of the RD1<sup>BCG</sup> region. Samples with adjusted p-values (BH-fdr corrected) <0.05 are indicated with purple points. Gene names highlighted in magenta have fewer than 5 TA sites located in the gene; too few to determine the statistical significance of changes in insertion levels with this method.</p> <p> </p> <p><strong>Supplementary Tables </strong></p> <p><strong>Table S1. Sequencing statistics of the input and output transposon libraries. </strong>The numbers in the column labelled “filename” refers to the sequencing file from that sample and come from individual animals (Bioproject ID: PRJNA816175, Submission ID: SUB11067380).</p> <p> </p> <p><strong>Table S2. Tissues collected and scored for gross pathology. </strong>Tissues from head and neck lymph nodes (from the right and left sub-mandibular lymph nodes, the right and left medial retropharyngeal lymph nodes), thoracic lymph nodes (the right and left bronchial lymph nodes, the cranial tracheobronchial lymph nodes, the cranial and caudal mediastinal lymph nodes) and from lung lesions, were collected and scored.</p> <p> </p> <p><strong>Table S3. Log<sub>2</sub> fold-changes for insertions across the entire genome of <em>M. bovis</em> AF2122/97. </strong>Cells are coloured according to log<sub>2</sub> fold-change. Refer to the text for the gene groups in individual tabs.</p> <p> </p> <p><strong>Table S4. </strong>Custom transposon sequencing primers and adaptors used in sequencing of the transposon libraries.</p> <p> </p>
Gene annotation files for Fraxinus excelsior (European ash) genome assembly BATG-0.5
<p>Gene annotation files for <em>Fraxinus excelsior</em> genome assembly v. BATG-0.5, published in Nature (doi:10.1038/nature20786). These files were previously hosted on the Ash Tree Genomes website (http://www.ashgenome.org/transcriptomes) and first made available for download via that site on 2016-02-08.</p> <p>The following annotation files are available:</p> <p>### GFF file of all gene models (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3</p> <p>### FASTA file of all cDNA sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.cdna.fa</p> <p>### FASTA file of all CDS DNA sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.cds.fa</p> <p>### FASTA file of all peptide sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.pep.fa</p> <p>### Functional annotation for each gene model (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.functional_annotation.tsv</p> <p>### GFF file of all gene models (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3</p> <p>### FASTA file of all cDNA sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.cdna.fa</p> <p>### FASTA file of all CDS DNA sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.cds.fa</p> <p>### FASTA file of all peptide sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.pep.fa</p> <p>### GFF file for gene models identified as probable transposable element related sequences (excluded from the other files)<br> Fraxinus_excelsior_38873_TGAC_v2.transposable_elements.gff3</p> <p><br> NB: The annotation files include preliminary annotations for genes within the organellar scaffolds (gene models FRAEX38873_v2_000400370-FRAEX38873_v2_000401330), which were not reported in the publication of the BATG0.5 assembly (doi:10.1038/nature20786).</p>
SEESAW quantification data for temporal gene expression across osteoblastogenesis (B6xCAST), n=9
<p>Osteoblast cells mature from a mesenchymal stem cell pool to become cells capable of forming bone matrix and mineralizing this matrix. The goal of this study was to characterize temporal changes in the transcriptome across osteoblast maturation, starting with committed mesenchymal stem cell/ early pre-osteoblast stage through to mature osteoblasts capable of matrix mineralization. Methods: Enriched populations of pre-osteoblast-like cells were obtained from neonatal calvaria from B6xCAST mice expressing CFP under the control of the Col3.6 promoter. These cells were placed into culture for 4 days, removed from culture and subjected to FACS sorting based on the presence/absence of CFP expression. Cells expressing CFP were returned to culture, subjected to an osteoblast differentiation cocktail and RNA was collected at 2, 4, 6, 8, 10, 12, 14, 16 and 18 days post differentiation. Methods II: mRNA profiles for each time point were generated by next generation RNA sequencing, using an Illumina HiSeq 2000. Three technical replicates per sample were sequenced. Overall design: Gene expression in calvarial osteoblasts from neonatal B6xCAST-Col3.6 CFP mice at 9 time points post differentiation.</p> <p>File description: the R data files (.rda) provide outputs of the scripts in the mikelove/osteoblast-quant GitHub repo (July 2022, commit 01d96490), having run the fishpond package function importAllelicCounts() followed by minimal filtering. The `_counts.rda` files contain SummarizedExperiment objects with estimated count, TPM abundance, and effective length, but do not contain inferential replicates (bootstrap counts), although the transcript-level allelic counts object contains bootstrap mean and variances for every isoform, sample, and allele. The other two `.rda` files are summarized to gene level.</p> <p>The `_quant_dirs.tgz` files contain all the Salmon quantification data including bootstraps for the 9 time points. They are grouped into sets of three for convenience. The `CAST_EiJ.diploid.fa.gz` file provides the transcript sequences that were used for Salmon quantification.</p> <p>The `B6xCAST_discordant_global_AI.csv` file contains the same information as presented in Table S1 of Wu et al (2022). These are TSS-level results for 134 genes showing significant and discordant patterns within gene.</p> <p>The other 6 CSV files provide global and dynamic AI testing results at three levels of resolution: gene level, isoform level (txp), and TSS level where TSS within 50bp are combined into a single TSS-group. The significance cutoff is a q-value of 0.05. The code used for generating these results is provided in the GitHub repo: FennecFish/osteoblast-test.</p>
Cytochrome P450 Genes Expressed in Phasmatodea Midguts
<p>In June 2021, representative sequences for insect cytochrome P450s were downloaded from NCBI, limiting the search to those in the UniProtKB database. The resulting 111 sequences were used as a query to mine the above transcriptomes using tblastn with an expect value threshold of e-10. These were manually annotated by removing truncated sequences, using the ExPASy online translation tool to obtain the complete amino acid sequences, removing duplicates using the sRNAtoolbox webserver, and confirming that the sequences were cytochrome P450s by identifying them using blastp against the NCBI database. The resulting sequences were combined with the representatives from NCBI, aligned using the Clustal W program built into the software MEGA version X. Any sequences missing the heme-binding domain FXXGXXXCXG/A, which is a signature motif for CYPs, were deleted. The sequences were also checked for the presence and absence of other four signature motifs from insect CYPs: helix C (WxxxR), helix I (GxE/DTT/S), helix K (ExLR), and PERF (PxxFxPE/DRE/F).</p>
Collation and orthology-based identification of hormone-related genes in bread wheat
<p>Plant hormones coordinate a plethora of developmental processes in plants, including responses to abiotic and biotic stressors. Here, we collate the findings of previous studies identifying bread wheat (<em>Triticum aestivum</em>) genes related to hormonal processes (<strong>biosynthesis</strong>, <strong>transport</strong>, <strong>signalling</strong>, and <strong>catabolism</strong>) and collect wheat orthologues from hundreds of additional hormone-related genes utilising the Ensembl Plants Compara database. We have initially conducted this procedure for <strong>abscisic acid</strong>, <strong>auxins </strong>(IAA and IBA), <strong>brassinosteroids</strong>, <strong>cytokinins</strong>, <strong>ethylene</strong>, <strong>gibberellins</strong>, and <strong>strigolactone</strong>, yielding a total of over 1,700 putative wheat orthologues. We aim to provide a community resource to aid gene annotation and subsequent analyses. We warmly welcome feedback from the community.</p> <p>Please refer to the file <strong>README.pdf</strong> for further details, including methods and references.</p>
GEnes and LAnguages TOgether
<p>Cite the source of the dataset as:</p> <blockquote> <p>Barbieri et al. 2022. A global analysis of matches and mismatches between human genetic and linguistic histories. PNAS. DOI: 10.1073/pnas.2122084119</p> </blockquote>
Data: multimodal cell tracking from systemic administration to tumour growth by combining gold nanorods and reporter genes
<p>This data set includes multispectral optoacoustic tomography images supporting an article on cell tracking (preprint: bioRxiv 199836; https://doi.org/10.1101/199836). The corresponding bioluminescence results are included too, as well as the spectra used for the multispectral processing. </p>
Interactive map of distribution of gene fragments indicative of cyanotoxin biosynthesis and cyanotoxins in the European Alps
<p><span>Distribution of cyanotoxins and cyanotoxin biosynthesis genes in Alpine region determined by LC-MS/MS and (q)PCR. Cyanotoxins and cyanotoxin genes are mapped on separate layers, and two basemaps are available (simple and relief). Results can be filtered by location, sample type, water body type, cyanotoxins and cyanotoxin genes. Note that cyanotoxin analyses were not performed on all sampling points.</span></p>
What is epigenetics? and how our choices affect our genes
<p>The first video of the EPIBOOST Project, which integrates researchers Cecília Guerra and Maria José Loureiro from CIDTFF, and David Oliveira from DigiMedia, in their multidisciplinary team, is now available. As experts in Didactics and Educational Technology, researchers from CIDTFF and DigiMedia are responsible for producing scientific dissemination videos related to epigenetics, targeting society and to be disseminated on online platforms such as Educast (PT) and YouTube (for global accessibility). The first video explains "What is Epigenetics?" and how our choices affect our genes, with scientific review by Joana Luísa Pereira and Guilherme Jeremias, is already available on the project's page and on YouTube (<a href="https://youtu.be/1Tu1JA1_ICY">https://youtu.be/1Tu1JA1_ICY</a>).</p> <p>Date of the description: May, 2, 2024</p>
Cladonema radiatum Alr and IgSF genes assembly and domain predictions
<p><span>This dataset is related to the </span><span>submitted paper</span><span> " A single gene determines allorecognition in hydrozoan jellyfish <em>Cladonema radiatum</em> inbred lines ".</span></p> <h1><span>Abstract:</span></h1> <p><strong><span> </span></strong></p> <p><span>Allorecognition—the ability of an organism to discriminate between self and non-self—is crucial to colonial marine animals to avoid invasion by other individuals in the same habitat. The cnidarian hydroid <em>Hydractinia</em> has long been a major research model in studying invertebrate allorecognition, establishing a rich knowledge foundation. In this study, we introduce a new cnidarian model <em>Cladonema radiatum</em> (<em>C. radiatum</em>). <em>C. radiatum</em> is a hydroid jellyfish which also forms polyp colonies interconnected with stolons. Allorecognition responses, fusion or regression of stolons, are observed when stolons encounter each other. By transmission electron microscopy, we observe rapid tissue remodelling contributing to gastrovascular system connection in fusion. Rejection responses are regulated by reconstruction of the chitinous exoskeleton perisarc, and induction of necrotic and autophagic cellular responses at cells in contact with the opponent. Genetic analysis identifies allorecognition genes: six<em> Alr </em>genes located on the putative Allorecognition Complex (ARC) and four immunoglobulin superfamily genes on a separate genome region. C. radiatum allorecognition genes show notable conservation with the <em>Hydractinia Alr</em> family. Remarkedly, stolon encounter assays of inbred lines reveal that genotypes of Alr1 solely determine allorecognition outcomes in <em>C. radiatum</em>.</span></p>
Human intestinal Bacteria Collection (HiBC): 16S rRNA gene sequences
<p>The <a href="https://hibc.rwth-aachen.de/" target="_blank" rel="noopener">Human intestinal Bacteria Collection (HiBC)</a> is a collection of bacterial strains, isolated from the human gut for which 16S rRNA gene sequences, genome sequences and culture conditions are made available to the research community. In addition to previously described bacteria, we include strains that represent novel species which have been taxonomically described and validly named, or will be in the future. This collection will be updated regularly.</p> <p>This dataset includes the sequences of the 16S rRNA gene sequences of the isolates in the FASTA nucleotide format. Sequences ending in Sanger were obtained using the Sanger dideoxy sequencing technology. Sequences ending in Genome were obtained from the genome sequence using barrnap.</p>
Row sequcenes data for assessing the risks of potential pathogens and antibiotic resistance genes among heterogeneous habitats in a temperate estuary wetland
<p>The study included 118 usable samples within three different habitats (water, soil, and sediment) across the Liaohe River basin to the Red Beach wetland collected from seven papers, and all of the sequence files were uploaded for availability.</p>
Dataset for the stimulator of interferon genes protein (STING1) antibody screening study
<p><strong>This antibody characterization dataset is related to the F1000 research article openly available at F1000Research.</strong></p> <p><em>This dataset contains the following underlying raw data for a study which characterized sixteen antibodies for the stimulator of interferon genes protein (STING1) in western blot, immunoprecipitation and immunofluorescence. The corresponding study is accessible on the YCharOS community on Zenodo (<a href="https://doi.org/10.5281/zenodo.11582350">https://doi.org/10.5281/zenodo.11582350</a>).</em></p> <p><em>The Dataset is in the format of a zip file. Once downloaded, please expand the zip file to access the folders containing the underlying data for Western blot (Wb), immunoprecipitation (IP) and immunofluorescence (IF).</em></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.