Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
48,977
datasets available to search
ShareScore release 0.7.1
Dataset results
48,977 results for “genes”
Regulation of Dye-decolorizing Peroxidases Gene Expression in Pleurotus ostreatus Grown on Glycerol as the Carbon Source
<p>This dataset contains the raw data and code necessary to reproduce the results of: Regulation of dye peroxidas gene expression in Pleurotus ostreatus grown on glycerol as the carbon source.</p> <p> </p> <p>These data are also available at github: <a href="https://github.com/JLuisCuamatzi/Pleurotus_ostreatus_CarbonSources">JLuisCuamatzi/Pleurotus_ostreatus_CarbonSources: Data and scripts to reproduce the analysis performed at Regulation of dye peroxidas gene expression in Pleurotus ostreatus grown on glycerol as the carbon source (github.com)</a></p>
The North Pacific Eukaryotic Gene Catalog: Raw assemblies from Gradients 1, 2 and 3
<p>The North Pacific Eukaryotic Gene Catalog consolidates eukaryotic metatranscriptome data from three latitudinal transects of the North Pacific transition zone and one cruise in the subtropical gyre. Metatranscriptomes were gathered from latitudinally-resolved surface samples, and diel-resolved temporal studies, with samples taken in triplicate or duplicate and collected on 0.2-100 μm, 0.2-3 μm, and 3 μm-100 or 200 μm size fractions. These metatranscriptome data were <em>de novo</em> assembled into 175 independent assemblies, totalling 182 million clustered nucleotide contigs. Assemblies were annotated by taxonomy and function. This catalog provides assembled environmental contigs, their translated peptide sequences, and their taxonomic and functional annotations with the aim of facilitating continued discoveries about the molecular ecology of microbial eukaryotes in the North Pacific.<br><br>A full description of this data is published in Scientific Data, available here: <a href="https://www.nature.com/articles/s41597-024-04005-5" target="_blank" rel="noopener">The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations</a>. Please cite this publication if your research uses this data:<br><br>Groussman, R. D., Coesel, S. N., Durham, B. P., Schatz, M. J., & Armbrust, E. V. (2024). The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations. <em>Scientific Data</em>, <em>11</em>(1), 1161.</p> <div> <p>This dataset repository is associated with a codebase and documentation repository:<br><a href="https://github.com/armbrustlab/NPac_euk_gene_catalog" target="_blank" rel="noopener">https://github.com/armbrustlab/NPac_euk_gene_catalog</a><br>Please see this code repository for additional data and project updates<br><br>Translated and processed protein sequences and their annotations are available in this repository: <br><a href="../doi/10.5281/zenodo.10472589">https://zenodo.org/doi/10.5281/zenodo.10472589</a><br><br>99% identity clustered nucleotide sequences and kallisto enumerations are available here:<br><a href="../doi/10.5281/zenodo.10570448">https://zenodo.org/doi/10.5281/zenodo.10570448</a></p> </div> <div> <p>File contents: this repository contains five .tar.gz compressed tarballs with raw de novo Trinity assemblies of poly-A selected metatranscriptomes from the Gradients 1 through 3 cruises, and a plain-text file with the custom spike-in mRNA standards (CustomStandardSequences.txt)</p> </div> <div> <p><strong><br>Gradients1.KOK1606.PA.assemblies.tar.gz</strong><br>- Link to <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/tree/main/projects/G1PA" target="_blank" rel="noopener">G1PA project github page</a><br>- Simons CMAP cruise page and datasets: <a href="https://simonscmap.com/catalog/cruises/KOK1606" target="_blank" rel="noopener">https://simonscmap.com/catalog/cruises/KOK1606</a><br>- Short read processing code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G1PA.process_short_reads.sh" target="_blank" rel="noopener">G1PA.process_short_reads.sh</a><br>- Trinity assembly code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G1PA.trinity_assemblies.sh" target="_blank" rel="noopener">G1PA.trinity_assemblies.sh</a></p> </div> <div> <p><strong><br>Gradients2.MGL1704.PA.assemblies.tar.gz</strong><br>- Link to <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/tree/main/projects/G2PA" target="_blank" rel="noopener">G2PA project github page</a><br>- Simons CMAP cruise page and datasets: <a href="https://simonscmap.com/catalog/cruises/MGL1704" target="_blank" rel="noopener">https://simonscmap.com/catalog/cruises/MGL1704</a><br>- Short read processing code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G2PA.process_short_reads.sh" target="_blank" rel="noopener">G2PA.process_short_reads.sh</a><br>- Trinity assembly code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G2PA.trinity_assemblies.sh" target="_blank" rel="noopener">G2PA.trinity_assemblies.sh</a></p> </div> <div> <p><strong><br>Gradients3.KM1906.PA.assemblies.tar.gz</strong><br>- Link go <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/tree/main/projects/G3PA" target="_blank" rel="noopener">G3PA project github page</a><br>- Simons CMAP cruise page and datasets: <a href="https://simonscmap.com/catalog/cruises/KM1906" target="_blank" rel="noopener">https://simonscmap.com/catalog/cruises/KM1906</a><br>- Short read processing code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G3PA_UW.process_short_reads.sh" target="_blank" rel="noopener">G3PA_UW.process_short_reads.sh</a><br>- Trinity assembly code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G3PA_UW.trinity_assemblies.sh" target="_blank" rel="noopener">G3PA_UW.trinity_assemblies.sh</a></p> </div> <div> <p><strong><br>G3_diel.KM1906.PA.assemblies.tar.gz</strong><br>- Link go <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/tree/main/projects/G3PA" target="_blank" rel="noopener">G3PA project github page</a><br>- Simons CMAP cruise page and datasets: <a href="https://simonscmap.com/catalog/cruises/KM1906" target="_blank" rel="noopener">https://simonscmap.com/catalog/cruises/KM1906</a><br>- Short read processing code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G3PA_diel.process_short_reads.sh" target="_blank" rel="noopener">G3PA_diel.process_short_reads.sh</a><br>- Trinity assembly code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G3PA_diel.trinity_assemblies.sh" target="_blank" rel="noopener">G3PA_diel.trinity_assemblies.sh</a></p> </div> <div> <p><strong><br>CustomStandardSequences.txt<br></strong>- Plain-text FASTA file with the spike-in standards used during mRNA extraction and sequencing prep<br>- Link to publication of spike-in standards methods: <a href="https://www.nature.com/articles/s41564-019-0507-5" target="_blank" rel="noopener">https://www.nature.com/articles/s41564-019-0507-5</a></p> </div> <div> <p>The 2015 SCOPE Diel metatranscriptome raw assemblies have been released in a previous Zenodo repository, and are not included again in this deposition. We provide the links to the Diel1 resources here:<br>- Diel1 raw metatranscriptome assembly Zenodo repository: <a href="../records/5009803" target="_blank" rel="noopener">https://zenodo.org/records/5009803</a><br>- Dataset DOI: <a href="https://doi.org/10.5281/zenodo.5009803" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.5009803</a><br>- Associated publication: <a href="https://www.frontiersin.org/articles/10.3389/fmicb.2021.682651/full" target="_blank" rel="noopener">https://www.frontiersin.org/articles/10.3389/fmicb.2021.682651/full</a><br>- Codebase: <a href="https://github.com/armbrustlab/diel_eukaryotes" target="_blank" rel="noopener">https://github.com/armbrustlab/diel_eukaryotes</a><br>- Simons CMAP cruise page and datasets: <a href="https://simonscmap.com/catalog/cruises/KM1513" target="_blank" rel="noopener">https://simonscmap.com/catalog/cruises/KM1513</a><br>- Short read processing code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/D1PA.process_short_reads.sh" target="_blank" rel="noopener">D1PA.process_short_reads.sh</a><br>- Trinity assembly code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/D1PA.trinity_assemblies.sh" target="_blank" rel="noopener">D1PA.trinity_assemblies.sh</a></p> </div> <p><br><br></p>
Gene expression count matrix for 4 T cell subtypes from ROSMAP participants
<p><span>Peripheral blood mononuclear cells (PBMCs) from participants in the Rush Religious Orders Study/Memory and Aging Project (ROSMAP) were isolated by Ficoll gradient centrifugation, then sorted by high-speed flow cytometry into the following T cell subtypes:<span> </span>CD4+CD45RO-, CD4+CD45RO+, CD8+CD45RO-, and CD8+CD45RO+.<span> </span>Total RNA was extracted using buffer TCL (Qiagen), then RNA-seq libraries were prepared according to the Single Cell RNA Barcoding and Sequencing method originally developed for single-cell RNA-seq</span><span>, adapted for extracted total RNA.<span> </span>RNA libraries were collected on a single 384-well plate and sequenced on the Illumina HiSeq </span><span>using the High-throughput 3<span>’</span> Digital Gene Expression (DGE) library</span><span>.<span> The "RNA count matrix" file is the raw counts from the 384-well plate, while the "ROSMAP_Tcell_DGE_PlateMap" file contains metadata for the wells on the plate, by well position.</span></span></p>
Predicted genes from the Amblyomma americanum draft genome assembly
<p>Data for pub "Predicted genes from the <em>Amblyomma americanum </em>draft genome assembly."</p> <ul> <li>Amblyomma_americanum_filtered_assembly.fasta: Decontaminated A. americanum genome with bacterial contigs removed</li> <li>Amblyomma_americanum_bacterial_contigs_info.tsv: Information about contigs classified as bacteria that were removed</li> <li>Amblyomma_americanum_annotation_data.tar.gz: Directory of annotation data produced by EvidenceModeler as part of the nf-core/genomeannotator workflow. Includes files in FASTA format (predicted genes and proteins), set of proteins clustered at 99% identity in FASTA format, and annotations in both GFF3 and GTF formats. GTF file produced from the GFF3 file with AGAT.</li> <li>Amblyomma_americanum_transcriptome_assembly_data.tar.gz: Directory of data generated for the transcriptome assembly that was used for gene prediction</li> </ul>
Supplementary datasets for manuscript titled: Seasonal tissue-specific gene expression in wild crown-of-thorns starfish reveals reproductive and stress-related transcriptional systems
<p>Supplementary datasets for manuscript titled: Seasonal tissue-specific gene expression in wild crown-of-thorns starfish reveals reproductive and stress-related transcriptional systems</p>
MAGE: Multi-ancestry Analysis of Gene Expression
<p>MAGE comprises RNA-seq data from lymphoblastoid cell lines derived from 731 individuals from the <a href="https://doi.org/10.1038/nature15393" rel="nofollow">1000 Genomes Project (1KGP)</a>, representing 26 globally-distributed populations across five continental groups. These data offer a large, geographically diverse, open access resource to facilitate studies of the distribution, genetic underpinnings, and evolution of variation in human transcriptomes and include data from several ancestry groups that were poorly represented in previous studies.</p> <p>Briefly, this repo contains the following data:</p> <ol> <li>Sample metadata and sequencing metrics</li> <li>Gene expression and splicing matrices used for e/sQTL mapping and analyses of global trends of expression/splicing diversity</li> <li>cis-e/sQTL mapping results, including aFC estimates for cis-eQTLs</li> <li>Functional annotations of cis-e/sQTLs</li> <li>Results of colocalization analysis between MAGE e/sQTLs and complex trait GWAS from the <a href="https://doi.org/10.1038/s41586-019-1310-4" rel="nofollow">PAGE</a> study</li> <li>Results of analyses of global trends of expression/splicing diversity</li> <li>Jointly-generated top genotype PCs for samples in MAGE and other resources with paired WGS/RNA-seq data (Geuvadis, GTEx, AFGR)</li> </ol> <p>READMEs are provided for all data in the repo.</p>
scGraph2Vec: a deep generative model for gene embedding augmented by Graph Neural Network and single-cell omics data
<p>This repository contains the training data and source code to reproduce the results of our paper:<br>scGraph2Vec: a deep generative model for gene embedding augmented by Graph Neural Network and single-cell omics data</p> <p>More description can be also found in GitHub (https://github.com/LPH-BIG/scGraph2Vec).</p>
Aqueous geochemical measurements and speciation calculations with concurrent copper resistance gene counts from sediment metagenomes over a seasonal cycle from 2015 to 2016 on Silver Bow Creek and Blacktail Creek near Butte, MT
<p>This dataset contains information from concurrently gathered geochemical and metagenomic samples collected from Silver Bow Creek and Blacktail Creek near Butte, MT (SBC/BC) during 2015 and 2016. SBC/BC is recovering from metal contamination related to extensive mining in the area. Full geochemical measurements, geochemical speciation calculations, and gene counts of sequences mapping to copper resistance genes using MG-RAST are included. </p>
Genomic Typing, Antimicrobial Resistance Gene, Virulence Factor and Plasmid Replicon Dataset for the Important Pathogenic Bacteria Klebsiella pneumoniae
<p>The infections caused by various bacterial pathogens both in clinical and community settings represent a significant threat to public healthcare worldwide. The growing resistance to antimicrobial drugs acquired by bacterial species causing healthcare-associated infections has already become a life-threatening danger noticed by the World Health Organization. Several groups or lineages of bacterial isolates usually called 'the clones of high risk' often drive the spread of resistance within particular species. </p> <p>Thus, it is vitally important to reveal and track the spread of such clones and the mechanisms by which they acquire antibiotic resistance and enhance their survival skills. Currently, the analysis of whole genome sequences for bacterial isolates of interest is increasingly used for these purposes, including epidemiological surveillance and developing of spread prevention measures. However, the availability and uniformity of the data derived from the genomic sequences often represents a bottleneck for such investigations. </p> <p>In this dataset, we present the results of a genomic epidemiology analysis of 61,857 genomes of a dangerous bacterial pathogen <em>Klebsiella pneumoniae</em> obtained from NCBI Genbank database. Important typing information including multilocus sequence typing (MLST)-based sequence types (STs), capsular (KL) and oligosaccharide (OL) types, CRISPR-Cas systems, and cgMLST profiles are presented, as well as the assignment of particular isolates to clonal groups (CG). The presence of antimicrobial resistance and virulence genes, as well as plasmid replicons, within the genomes is also reported. </p> <p>These data will be useful for researchers in the field of <em>K. pneumoniae</em> genomic epidemiology, resistance analysis and prevention measure development.</p>
Efficient and accurate framework for genome-wide gene-environment interaction analysis in large-scale biobanks
<p>Gene-environment interaction (GxE) analysis elucidates the interplay between genetic predispositions and environmental influences, offering significant potential for precision medicine. With the increasing use of electronic health records (EHR) linked to genetic data in large-scale biobanks, genome-wide association studies (GWAS) have expanded to encompass complex traits with intricate structures, such as time-to-event and ordinal categorical traits. Although these complex traits convey more phenotypic information, most existing scalable genome-wide GxE analysis approaches only focus on quantitative or binary traits. In this work, we propose a scalable and accurate analysis framework, SPAGxE<sub>CCT</sub>, that is applicable to a wide variety of trait types. We extend SPAGxE to SPAGxE+, which can account for sample relatedness. In addition, we extend SPAGxE<sub>CCT</sub> to SPAGxEmix<sub>CCT</sub>, which accounts for population stratification and is applicable to include individuals from multiple ancestries or admixed populations. We applied SPAGxE<sub>CCT</sub>, SPAGxE+, and SPAGxEmix<sub>CCT</sub> to analyze time-to-event traits in UK Biobank. For the SPAGxE<sub>CCT</sub> analyses, 281,149 White British individuals were included. For the SPAGxE+ analyses, 337,367 WB individuals with sample relatedness were included. For the SPAGxEmix<sub>CCT</sub> analyses, 338,044 individuals from all ancestries were included. SPAGxE<sub>CCT</sub>, SPAGxE+, and SPAGxEmix<sub>CCT</sub> are computationally efficient to analyze large datasets with hundreds of thousands of individuals, can accurately control type I error rates while remaining powerful to identify novel GxE findings.</p>
Gene tagging and gene deletion resources for Leishmania mexicana MNYC/BZ/62/M379 Cas9/T7 strain
<p><em>primers_barcodes.csv</em>: List of primer sequences necessary for N and C terminus gene tagging, as well as gene deletion in the Leishmania mexicana MNYC/BZ/62/M379 Cas9/T7 strain. Each row contains the gene name and the DF (downstream forward), DR (downstream reverse), DSG (downstream guide sRNA), UF (upstream forward), UFB (upstream forward including a gene-unique 17nt barcode sequence), UR (upstream reverse), USG (upstream guide sRNA), VF (verification forward) and VR (verification reverse) primer sequences. Empty cells indicate that it was not possible to design this primer for this gene. Guide sRNA perfect match and off-target counts are included as well. The primer sequences were designed using LeishGEdit (http://www.leishgedit.net). Barcode sequences and assigned IDs for the unique identification of knock-out or tagged strains are included as separate columns. For recommended methods for endogenous tagging or gene deletion see Beneke <em>et al., </em>R. Soc. Open Sci.4170095 (2017), for generating barcoded deletion mutants see Beneke and Gluenz, Mol. Biochem. Parasitol. 239 (2020).</p> <p><em>genome.gff</em>: Annotated genome of the <em>L. mexicana</em> MNYC/BZ/62/M379 strain, genetically modified to express T7 RNA polymerase and Cas9. The annotation is provided in a combined GFF3 / FASTA format that also includes the sequences of the chromosomes and small contigs. The annotation also specifies polyadenylation sites (PAS features) and splice leader acceptor sites (SLAS features) which were used to refine the boundaries of protein-coding sequences as well as 3' and 5' untranslated regions over the reference genome of <em>L. mexicana</em> MNYC/BZ/62/M379<em>.</em></p> <p><em>c9t7_sequences.fasta</em>: Raw chromosome and contig sequences in FASTA format.</p> <p><em>c9t7_transcripts.fasta</em>: mRNA transcript sequences in FASTA format (includes 5' and 3' UTRs).</p> <p><em>c9t7_transcript_CDSs.fasta</em>: Coding sequences in FASTA format.</p> <p><em>c9t7_predicted_protein_sequences.fasta</em>: Predicted protein amino acid sequences in FASTA format.</p> <p>Note: This version provides an update for <em>c9t7_transcript_CDSs.fasta, c9t7_predicted_protein_sequences.fasta</em> and <em>genome.gff</em>, correcting an off-by-one sequence coordinate in 48 of the the protein-coding genes.</p>
Isoprenoid Gene Network in Arabidopsis Thaliana
<p>To gain more insight into the cross-talk between both pathways at the transcriptional level, gene-expression patterns were monitored under various experimental conditions using 118 GeneChip (Affymetrix) microarrays. To construct a genetic regulatory network in Wille et al. (2004), the interventional dataset focusses on 40 genes, 16 of which were assigned to the cytosolic pathway, 19 to the plastidal pathway and five encode proteins located in the mitochondrion.</p> <p> </p> <p><strong>Task: </strong>The dataset can be used to study causal discovery algorithms.</p> <p> </p> <p><strong>Summary: </strong></p> <ul> <li><strong>Size of dataset</strong>: 118 x 41</li> <li><strong>Task:</strong> Causal Discovery Problem</li> <li><strong>Data Type:</strong> Continuous Data</li> <li><strong>Dataset Scope:</strong> Standalone Dataset</li> <li><strong>Ground Truth:</strong> Known Graph</li> <li><strong>Temporal Structure:</strong> Static Data</li> <li><strong>License:</strong> CC BY 2.0 (see https://genomebiology.biomedcentral.com/articles/10.1186/gb-2004-5-11-r92#MOESM1)</li> <li><strong>Missing Values:</strong> No Missing Values</li> </ul> <p> </p> <p><strong>Missingness Statement: </strong>There are no missing values.<strong><br></strong></p> <p> </p> <p><strong>Samples:</strong></p> <p>Experimental procedures involved seedlings, leaves and roots. For experiments involving seedlings or leaves, plants were grown in growth chambers at 70% humidity and daily cycles of 16 h light at 21◦C and 8 h darkness at 21◦C. Plant material from three independent experiments (replications not in parallel) for each experiment group respectively was pooled prior to RNA extraction. The 118 different experimental conditions are denoted by c1-c118:</p> <ul> <li><strong>c1-c2: </strong>Experiment with wild-type and era mutant seedlings. (growth stage 1.0; tissue: whole seedlings)</li> <li><strong>c3-c5:</strong> Arabidopsis tissue culture, leaf and seedling in a baseline experiment (growth stage: - ; tissue: tissue culture, seedling, adult leaf).</li> <li><strong>c6-c14: </strong>RNA was extracted from seedlings and adult leaves of wild-type and prenylation mutant plants grown under standard conditions (growth stage 1.0; tissue: whole seedlings and adult leaves).</li> <li><strong>c15-c22: </strong>RNA was extracted from wild-type and several transgenic seedlings (growth stage: 1.0; tissue: whole seedlings).</li> <li><strong>c23-c30:</strong> RNA was extracted from a root inducible system exposed to hormonal treatments. (growth stage: 1.0; tissue: lateral roots).</li> <li><strong>c31-c56: </strong>Arabidopsis seedlings were exposed to light and dark conditions in a time-course experiment (0, 10 min, 1h, 5h, 2d, 5d). (growth stage: 1.0; tissue: whole seedlings).</li> <li><strong>c70-c92:</strong> Experiment to assess the effect of inhibitors of the MVA pathway (lovastatin) and the MVA-independent pathway (fosmidomycin) on the expression of genes involved in isoprenoid biosynthesis. (growth stage: 1.0 and 3.90; tissues: whole seedlings and adult leaves).</li> <li><strong>c93-c118:</strong> Arabidopsis seedlings and adult leaves were exposed to ozone for several periods of time. (growth stage: 1.0 and 3.90; tissues: whole seedlings, cauline leaves and adult leaves).</li> </ul> <p> </p> <p><strong>Features: </strong></p> <ul> <li><strong>c[...]: </strong>This row indicates the experiment as descried above</li> <li> <p><strong>AACT1, AACT2, CMK, DPPS1, ...:</strong> Gene expressions level. The column names indicate the gene name.</p> </li> </ul> <div> <p> </p> <p><strong>Note: </strong>Although the gene UPPS2 was analyzed in Willie et al. (2004), it is not provided in its supporting material and in this dataset.</p> </div>
The Minimum Information about a Biosynthetic Gene Cluster (MIBiG) data repository
<p>This dataset was originally published alongside the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) data standard publication(s).</p> <p>It contains JSON files following the MIBiG data standard. Additional information on proteins/genes associated to biosynthetic gene clusters described by MIBiG can be found in the GenBank (gbk) and fasta files.</p> <p>This dataset was uploaded with permission from the corresponding author(s).</p> <p>For more information, see https://mibig.secondarymetabolites.org/.</p>
Most bacterial gene families are biased toward specific chromosomal positions
<p><span>The arrangement of genes along bacterial chromosomes influences their expression through growth rate-dependent gene copy number changes during DNA replication. While translation and transcription genes often cluster near the origin of replication, the extent of positional biases across gene families remains unclear. We hypothesized that natural selection broadly favors specific chromosomal positions to optimize growth rate-dependent expression. Analyzing 910 bacterial species and proteomics data from <em>Escherichia coli</em> and <em>Bacillus subtilis</em>, we find that about two-thirds of bacterial gene families are positionally biased, mainly near the origin or terminus of replication, with the strongest natural selection in fast-growing species. Our findings reveal chromosomal positioning as a fundamental mechanism for coordinating gene expression with growth rate, highlighting evolutionary constraints on bacterial genome architecture.</span></p> <p> </p>
Deep learning models predicting gene functions and pathways using public DRKG knowledge graph and graph neural network
<p>The attached dataset contains pretrained link prediction models, as described in our paper 'Morphological Map of Under- and Over-Expression of Genes in Human Cells'.</p>
Supplementary material for Targeted gene knock-in reduces variation between transformants in the mushroom-forming fungus Schizophyllum commune
<p>Supplementary data for "Targeted gene knock-in reduces variation between transformants in the mushroom-forming fungus <em>Schizophyllum commune</em>"</p> <p>Dataset consists of fluorescent images of <em>S. commune</em> strains with an ectopic or targeted integration of <em>dTomato </em>under the control of the <em>tubulin </em>promoter and <em>hom2 </em>terminator and the obtained fluorescent intensity of each strain. For thesholding the mean intensity of all pixels above 14 (range 0, 255) was calculated.</p> <p>Files are names according to strain (E1-E12 for ectopic integrations and TI1-TI6 for targeted integrations and WT for wildtype) and replicate.</p>
FASTA file containing the MYB encoding gene An2-like and Ant1 coding sequences corresponding to wild and cultivated tomato accessions
<p>The coding sequence (CDS) of the MYB encoding genes <em>Ant1</em> and <em>An2-like</em>. Sequences were retrieved from regions corresponding to the<em> Aft</em> locus from <em>Solanum galapagense </em>accession LA1141, <em>S. lycopersicum</em> variety OH8245, and 84 tomato accessions published as part of The 100 Tomato Genome Sequencing Consortium (The 100 Tomato Genome Sequencing Consortium et al., 2014). Sequences were compared to available CDS available from the Sol genomics network (SGN) and the National Center for Biotechnology Information. The CDS was retrieved from <em>S. lycopersicum</em> variety Indigo Rose [MN433087 (Yan et al., 2020)], <em>S. lycopersicum</em> accession LA1996 [MN242011.1, EF433417.1( Sapir et al., 2008; Colanero et al., 2020)], and <em>S. chilense </em>accession LA1930 [MN242012.1 (Colanero et al., 2020)], The orthologous CDS corresponding to the <em>Aft </em>MYB encoding genes from <em>Solanum tuberosum</em> L. Group Phureja clone DM1-3 genome (PGSC DM v4.03 Pseudomolecules) was retrieved from the Potato Genome Sequence Consortium (PGSC: Potato Genome Sequencing Consortium et al., 2011), and the Capsicum annum cv. CM334 genome was retrieved from <em>Capsicum annuum </em>cv CM334 genome chromosome release 1.55 (Hulse-Kemp et al. 2018). These CDS were obtained using the Basic Local Alignment Search Tool (BLAST) tool available from the Sol Genomics Network (SGN) (available at https://solgenomics.net/tools/blast/). Comparison of syntenic chromosomal regions using known positions of tomato, potato, and pepper markers with comparative map viewer from SGN: (available at https://solgenomics.net/cview) on chromosome 10, was used as a quality check for S.<em> tuberosom</em> and <em>C. annuum.</em> Orthologous CDS corresponding to <em>Salvia miltiorrhiza, Arabidopsis thaliana</em>, [NM_105308.2, NM_105310.4 (Teng et al., 2005, Cominelli et al., 2008; Beradini et al., 2015)] were chosen based on tomato <em>Aft</em> sequence homology and gene annotations of positive R2R3 MYB regulation of anthocyanin. The CDS corresponding to the <em>Aft</em> genes were retrieved from the CDS reference genomes available from the Sol Genomics Network SGN: Tomato Genome CDS (ITAG release 4.0), Potato PGSC DM v3.4 CDS sequences, <em>Capsicum annuum </em>cv CM334 Genome CDS (release 1.55), or from the National Center for Biotechnology Information (NCBI: https://www.ncbi.nlm.nih.gov) reference sequences (RefSeq) section of the Genbank records. When accessed from Genank records, the CDS sequence was extracted from the “features” section and exported as a FASTA file.</p>
Metagenomes: gene function and family annotations
<p>Functional annotations of genes for all contigs in 1,782 metagenomes.</p> <p>Genes were annotated to three sources: (1) COGs, (2) Pfams, and (3) <em>de novo</em> families from reference sequences. These gene annotations are used to train and run PlasX.</p>
Targeted Re-sequencing Identifies Candidate Fusiform Rust Resistance Genes in Loblolly Pine
<p>A fasta file containing the subset of the v2.01 Pita genome in addition to the novel NLR genes that were targeted by hybridization probes. </p> <p>A bed file describing the intervals targeted by the hybridization probes.</p> <p>Trinity assemblies of the 30 RNAseq libraries along with predictions by transdecoder of CDS and peptide sequences from those trinity assemblies. </p>
TBGA: A Large-Scale Gene-Disease Association Dataset for Biomedical Relation Extraction
<p>This repository contains the TBGA dataset. TBGA is a large-scale, semi-automatically annotated dataset for Gene-Disease Association (GDA) extraction. The dataset consists of three text files, corresponding to train, validation, and test sets, plus an additional JSON file containing the mapping between relation names and IDs. Each record in train, validation, or test files corresponds to a single GDA extracted from a sentence. Records are represented as JSON objects with the following structure:</p> <ul> <li><strong>text:</strong> sentence from which the GDA was extracted.</li> <li><strong>relation:</strong> relation name associated with the given GDA.</li> <li><strong>h: </strong>JSON object representing the gene entity, composed of: <ul> <li><strong>id: </strong>NCBI Entrez ID associated with the gene entity.</li> <li><strong>name:</strong> NCBI official gene symbol associated with the gene entity.</li> <li><strong>pos: </strong>list consisting of starting position and length of the gene mention within text.</li> </ul> </li> <li><strong>t:</strong> JSON object representing the disease entity, composed of: <ul> <li><strong>id: </strong>UMLS CUI associated with the disease entity.</li> <li><strong>name:</strong> UMLS preferred term associated with the disease entity.</li> <li><strong>pos:</strong> list consisting of starting position and length of the disease mention within text.</li> </ul> </li> </ul> <p>TBGA contains over 200,000 instances and 100,000 bags.<br> The zip file consists of one folder, named TBGA, containing the files corresponding to the dataset.</p> <p>If you use or extend our work, please cite the following: https://bmcbioinformatics.biomedcentral.com/articles/10.1186/s12859-022-04646-6#citeas<br> TBGA paper can be found at: <a href="https://rdcu.be/cKkY2">https://rdcu.be/cKkY2</a><br> TBGA code is available at: https://github.com/GDAMining/gda-extraction</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.