Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
48,977
datasets available to search
ShareScore release 0.7.1
Dataset results
48,977 results for “Genes”
Plant regeneration in leaf culture of Centaurium erythraea Rafn. Part 3: de novo transcriptome assembly and validation of housekeeping genes for studies of in vitro morphogenesis
<p>Six centaury transcriptomes (embryogenic calli, globular somatic embryos, cotyledonary somatic embryos, adventitious buds, leaves and roots of <em>in vitro</em> grown plants) were sequenced and <em>de novo</em> assembled using <a href="https://github.com/trinityrnaseq/trinityrnaseq/wiki">Trinity</a> .</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/CE_Assembly.tar.gz">CE_Assembly.tar.gz</a> - Centaury referent transcriptome comprises of 160.839 Trinity transcripts grouped in 105.726 Trinity genes.</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/CE_Assembly_fpkm.tar.gz">CE_Assembly_fpkm.tar.gz</a> - fpkm normalized read counts of the assembled transcripts in the six sequenced centaury tissues.</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/nt.db_CE_assembly.tar.gz">nt.db_CE_assembly.tar.gz</a> - annotation of assembled transcripts by mapping them against NCBI nucleotide (NT) database using BLASTn . The obtained results were filtered with E-value E ≤ 10<sup>-3</sup>.</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/swissprot.db_CE_assembly.tar.gz">swissprot.db_CE_assembly.tar.gz</a> - annotation of assembled transcripts by mapping them against NCBI nucleotide (<a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/swissprot.db_CE_assembly.tar.gz">s</a>wissprot) database using BLASTx . The obtained results were filtered with E-value E ≤ 10<sup>-3</sup>.</p> <p><a href="https://zenodo.org/api/files/a0546879-e382-4cf9-8185-f188d1a0c5f0/pfam30.db_CE_assembly.tar.gz">pfam30.db_CE_assembly.tar.gz</a> - annotation of assembled transcripts by mapping them against Pfam30 domain database using hmmer3. The obtained results were filtered with independent E-value E ≤ 10<sup>-3</sup>.</p>
Predicting gene expression using morphological cell responses to nanotopography
<p>This dataset contains the raw files, results files and R workspace files (.RData) associated with the paper:</p> <p>Predicting gene expression using morphological cell responses to nanotopography</p> <p>Please note that this dataset is separated according to the Figure presented in the published and peer-reviewed version of the manuscript. Particular folders contain its own README file to facilitate reproduction/replication of results and figures. </p>
Data for "Gut microbial genes are associated with neurocognition and brain development in healthy children"
<p><strong>Datasets accompanying<em> Gut microbial genes are associated with neurocognition and brain development in healthy children</em>, submitted to Nature Microbiology.</strong></p> <p><strong>Contents:</strong></p> <ul> <li> fecal_samples_master.csv <ul> <li>Metadata for all fecal samples processed by the Klepac-Ceraj Lab at Wellesley College</li> </ul> </li> <li>filemakerdb.csv <ul> <li>Initial export and parsing (long form) of deidentified patient metadata from internal filemnaker pro database</li> </ul> </li> <li>gbm.txt <ul> <li>Info about potentially neuroactive gene sets</li> <li>This was acquired as Supplementary Dataset 1 from <a href="https://doi.org/10.1038/s41564-018-0337-x">https://doi.org/10.1038/s41564-018-0337-x</a></li> </ul> </li> <li>batchXXX_analysis_noknead.tar.gz <ul> <li>Sequencing batches 001-012 (see fecal_samples_master.csv for metadata about samples contained in each batch)</li> <li>Each tarball contains: <ul> <li><strong>cluster.yaml</strong>: configuration file for snakemake pipeline (<a href="https://github.com/Klepac-Ceraj-Lab/snakemake_workflows">repo link</a>)</li> <li><strong>config.yaml</strong>: run configuration for snakemake pipeline</li> <li><strong>.snakemake/</strong>: metadata about snakemake pipeline runs on engaging cluster at MIT</li> <li><strong>output/</strong>: outputs from metaphlan2 and humann2 analysis runs. Note: kneaddata sequence files were not included, but will be uploaded to SRA (link to come)</li> </ul> </li> </ul> </li> <li>All <a href="https://www.uniprot.org/">uniprot</a> searches were performed 2019-09-19 <ul> <li>uniprot-abxr.tsv <ul> <li>search term: "keyword:\"Antibiotic resistance [KW-0046]\" AND reviewed:yes"</li> </ul> </li> <li>uniprot-carbohydrate.tsv <ul> <li>search term: "keyword:\"Carbohydrate metabolism [KW-0119]\" AND reviewed:yes"</li> </ul> </li> <li>uniprot-fa.tsv <ul> <li>search term: (keyword:\"Fatty acid biosynthesis [KW-0275]\" OR keyword:\"Fatty acid metabolism [KW-0276]\") AND reviewed:yes"</li> </ul> </li> </ul> </li> </ul>
Supplementary Data to "Disparate regulation of Smad3 phosphorylation and collagen gene transcription by full-length IL-33"
<p>These are Supplementary Figures for the article "Disparate regulation of Smad3 phosphorylation and collagen transcription by full-length IL-33"</p>
Acute Respiratory Distress Syndrome-Database of Genes (ARDS-DB)
<p>To better understand the gene level associations that are most relevant to Acute Respiratory Distress Syndrome (ARDS), a comprehensive resource is needed. There’s currently no freely available database dedicated to ARDS that provides comprehensive gene lists from experimentally verifiable studies, gene function, gene location, and additional metadata for tracking related link out resources. The need for such a database is only accentuated by the steep rise in ARDS cases due to the 2020 Coronavirus pandemic, in which infected patients admitted to the ICU develop ARDS at a rate of 67% to 85%, calling for an increase into ARDS research. </p> <p>Our goal was to develop such a resource for use by the scientific community to enhance our studies of ARDS and associated genes. Our first step was to perform data mining and curation of scientific literature through a robust review process. Subsequent steps enabled us to refine our data by capturing specific metadata and incorporating these into our database. The version 1 of the database will provide users with access to the database flat file with current genes, gene location, chromosomal information, and more in a freely accessible and downloadable format. Future project goals are to develop a standalone web portal that will integrate the gene level information with network analysis, and other visualizations for users. </p> <p> </p>
Dataset for genes linked to gastroschisis along with bioinformatics analysis
<p>The dataset consists of records from genes linked to gastroschisis. Genes displaying statistical significance with gastroschisis (excluding those genes undergoing adjusted calculations with covariates) were selected and manually curated for further bioinformatics analysis.</p> <p>Figure 1 illustrates the systematic review of the literature search strategy and selection criteria, publishing crude (unadjusted) genes linked to gastroschisis from January 1, 1990 to August 2, 2020. </p> <p>The full list of tables is described in the file READ ME and remains available in CSV files.</p>
Coevolving plasmids drive gene flow and genome plasticity in host-associated intracellular bacteria
<p>Comparative genomics and modeling of plasmids of the obligate host-associated intracellular phylum chlamydiae. </p>
Transcriptome analysis of the effect of over-expressing H2A.J mutants in proliferating WI38 fibroblasts for the paper entitled: The H2A.J histone variant contributes to Interferon-Stimulated Gene expression in senescence by its weak interaction with H1 and the derepression of repeated DNA sequences
<p>Abstract for overall study:</p> <p>The histone variant H2A.J was previously shown to accumulate in senescent human fibroblasts with persistent DNA damage to promote inflammatory gene expression, but its mechanism of action was unknown. We show that H2A.J accumulation contributes to weakening the association of histone H1 to chromatin and increasing its turnover. Decreased H1 in senescence is correlated with increased expression of some repeated DNA sequences, increased expression of STAT/IRF transcription factors, and transcriptional activation of Interferon-Stimulated Genes (ISGs). The H2A.J-specific Val-11 moderates the transcriptional activity of H2A.J, and H2A.J-specific Ser-123 can be phosphorylated in response to DNA damage with potentiation of its transcriptional activity by the phospho-mimetic S123E mutation. Our work demonstrates the functional importance of H2A.J-specific residues and potential mechanisms for its function in promoting inflammatory gene expression in senescence.</p> <p>Specific description for this dataset:</p> <p>H2A.J differs from canonical H2A only by a valine at position 11 instead of alanine, and the 7 C-terminal amino acids containing a potential minimal phosphorylation site SQ for DNA-damage response kinases. To test the functional importance of these H2A.J-specific sequences, we mutated Val-11 to Ala as is found in all canonical H2A sequences, and we mutated Ser-123 to either Glu to mimic a phospho-serine residue or to Ala to prevent phosphorylation. We also substituted the C-terminus of H2A.J with the C-terminus of H2A. These mutants, WT-H2A.J and canonical H2A-type1 were ectopically expressed in proliferating fibroblasts, and their microarray transcriptomes were compared to that of proliferating and senescent fibroblasts without ectopic histone expression. Genome-wide transcriptome analysis indicated that senescent fibroblasts clustered distinctly from proliferating fibroblasts, and proliferating fibroblasts expressing the H2A.J-V11A and H2A.J-S123E mutants clustered distinctly from fibroblasts expressing the other H2A.J mutants, WT-H2A.J, and H2A. Hallmark gene set enrichment analysis of the transcriptomes of fibroblasts expressing H2A.J-V11A or H2A.J-S123E versus control proliferating fibroblasts indicated that they showed the same highly significant enrichment for the Epithelial-Mesenchyme Transition, TNF-Alpha Signaling Via NF-kB, and Inflammatory Response gene sets. Notable inflammatory genes including IL1A, IL1B, IL6, CXCL8, and CCL2 are contained in these gene sets and are often induced in senescence as part of the senescence-associated secretory phenotype. Heat maps showed that the H2A.J-V11A and H2A.J-S123E mutants were particularly apt at activating the expression of these inflammatory genes in proliferating fibroblasts</p>
EoRNA, a barley gene and transcript abundance database
<p>A high-quality, barley gene reference transcript dataset (BaRTv1.0, Rapazote-Flores et al. 2019), was used to quantify gene and transcript abundances from 22 RNA-seq experiments, covering 843 separate samples. Using the abundance data we developed a Barley Expression Database (EoRNA* – Expression of RNA) to underpin a visualisation tool that displays comparative gene and transcript abundance data on demand as transcripts per million (TPM) across all samples and all the genes. EoRNA provides gene and transcript models for all of the transcripts contained in BaRTV1.0, and these can be conveniently identified through either BaRT or HORVU gene names, or by direct BLAST of query sequences. Browsing the quantification data reveals cultivar, tissue and condition specific gene expression and shows changes in the proportions of individual transcripts that have arisen via alternative splicing. TPM values can be easily extracted to allow users to determine the statistical significance of observed transcript abundance variation among samples or perform meta analyses on multiple RNA-seq experiments. * Eòrna is the Scottish Gaelic word for Barley</p>
A comprehensive evaluation of binning methods to recover human gut microbial species from a non-redundant reference gene catalog - Supporting Data
<p><strong>Description </strong></p> <p>The following files are available : </p> <ul> <li>Simulated non-redundant Gene Catalog (SGC) composed of 128267 genes;</li> <li>Gene abundance profiles across 40 samples: raw read counts, gene length normalized base counts, depth file computed by the jgi_summarize_bam_contig_depth script provided by MetaBAT;</li> <li>Gold Standard (GS) and Gold Standard Single Assignment (GS_SA) binning results;</li> <li>Binning results obtained on the SGC with nine binning methods: MSPminer, MGS-canopy, DAS Tool, MaxBin2, MetaBAT2, SolidBin, CONCOCT, COCACOLA and MyCC.</li> </ul> <p><strong>License</strong></p> <p>These files are licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p>
Phylogenomics of Gesneriaceae using targeted capture of nuclear genes
<p>Gesneriaceae (ca. 3400 species) is a pantropical plant family with a wide range of growth form and floral morphology that are associated with repeated adaptations to different environments and pollinators. Although Gesneriaceae systematics has been largely improved by the use of Sanger sequencing data, our understanding of the evolutionary history of the group is still far from complete due to the limited number of informative characters provided by this type of data. To overcome this limitation, we developed here a Gesneriaceae-specific gene capture kit targeting 830 single-copy loci (776,754 bp in total), including 279 genes from the Universal Angiosperm-353 kit. With an average of 557,600 reads and 87.8% gene recovery, our target capture was successful across the family Gesneriaceae and also in other families of Lamiales. From our bait set, we selected the most informative 418 loci to resolve phylogenetic relationships across the entire Gesneriaceae family using maximum likelihood and coalescent-based methods. Upon testing the phylogenetic performance of our baits on 78 taxa representing 20 out of 24 subtribes within the family, we showed that our data provided high support for the phylogenetic relationships among the major lineages, and were able to provide high resolution within more recent radiations. Overall, the molecular resources we developed here open new perspectives for the study of Gesneriaceae phylogeny at different taxonomical levels and the identification of the factors underlying the diversification of this plant group. </p>
Raw and processed GO term data to support running GCEA analyses using ensemble-based nulls, as described in the manuscript, 'Overcoming bias in gene category enrichment analyses of brain-wide transcriptomic data'.
<p>Data to support a toolbox for performing gene category enrichment analyses, including against ensembles of null phenotypes.</p> <p>Descriptions of how these data files can be used for this purpose are in the documentation for the toolbox, at https://github.com/benfulcher/GCEA_FalsePositives</p>
G-quadruplex in the gene of the large subunit of plant RNA polymerase II: billion years old story
<p><strong>Supplementary material to the journal article</strong></p> <p>Consist of:</p> <p>Supplementary material S1: Analyzed <em>RPB1 </em>sequences in 40 plant species together with detailed characteristics and G-quadruplex prediction using four different computational approaches.</p> <p>Supplementary material S2: G4 locus is the most conserved within the <em>RPB1</em> gene (40 bp long potential G4 locus is the most conserved site in the whole ~ 6000 bp long <em>RPB1</em> gene. See the histogram below the alignment: the position of the G4 locus is depicted, together with the horizontal red dashed line indicating relative nucleotide conservation among aligned sequences of<em> RPB1</em>)</p> <p>Supplementary material S3: Multiple alignment of G4 locus of <em>RPB1</em> paralogs in <em>Arabidopsis thaliana </em>centered to G4 locus of <em>RPB1 </em></p> <p>Supplementary material S4: Modelled 3D structure of G4 from <em>Bathycoccus prasinos</em> in PDB format</p> <p>Supplementary material S5: Gel electrophoresis and ThT staining of the selected G4-forming sequences</p> <p>Supplementary material S6: All analyzed <em>RPB1</em> sequences in FASTA format</p> <p>Supplementary material S7: Aligned <em>RPB1</em> sequences in FASTA format</p> <p>Supplementary material S8: <em>RPB1</em> paralogs in <em>Arabidopsis thaliana</em></p> <p>Supplementary material S9: Spectral composition of light used in the UV experiment. Analysis of emitted light was performed by Ocean Optics (HR4000CG-UV-NIR, USA) device.</p> <p>Supplementary material S10: Difference CD spectra - comparison without and with previous UV treatment</p>
Supplement 1: Full list of ICD10 codes and number of gene-disease links (tab-separated-value file); Supplement 2: Mapping (tab-separated-value file)
<p>Supplements to BioMedBridges deliverable 10.2 A prototype linking ICD10/SNOMED CT concepts to Ensembl gene identifiers:</p> <p><strong>Supplement 1</strong>: Full list of ICD10 codes and number of gene-disease links: table_icd10_gene_count_descr.tsv</p> <p><strong>Supplement 2</strong>: Mapping of disease terms: <em>ICD10_to_doid.tsv</em></p>
List of human genes and their probabilities of being intolerant to heterozygous protein truncating variants.
<p>Recalculation of the Supplementary Table 2 (doi:10.1371/journal.pcbi.1004647.s002) of the journal article "The Characteristics of Heterozygous Protein Truncating Variants in the Human Genome" by Bartha and Rausell published in PLoS Computational Biology (http://dx.doi.org/10.1371/journal.pcbi.1004647). Probabilities in this dataset were computed using human variation data from the Exome Aggregation Consortium (http://exac.broadinstitute.org/).</p> <p>Methods described in that article is relevant for this dataset. All author and affiliation information in that article is relevant for this dataset.</p> <p>Credit for the original human variation data is for the Exome Aggregation Consortium (http://exac.broadinstitute.org/, doi:10.1038/nature19057).</p>
Gene co-ordinates, expression levels; SNP identifiers and functions for Drosophila melanogaster (Sussex LHM population)
<p>Data for SNP context information to add to GWAS results. Specifically, SNP functions, sex-bias in gene expression, official SNP idenfiers from NCBI dbSNP, and gene positions and names (from UCSC Genome Browswer). Most of the input files are on-line and their URLs are stated in the code (make_dmel_accessory_data.sh). Also includes code, logs, and exploratory graphs.</p>
Adaptive Introgression in Modern Human Circadian Rhythm Genes Datasets
<p><strong>README:</strong></p> <p>Modern human genetic data with evidence of adaptive introgression from Neanderthals or Denisovans within circadian rhythm genes. The data was generated from the phased gnomAD 1KGP + HGDP callset (Koenig <em>et al</em>., 2024) and introgressed segments were identified by SPrime (Browning <em>et al</em>., 2018). Genes of interest were downloaded from the Circadian Genome Database (CGDB) (Li <em>et al</em>., 2017). Additional variants, haplotypes, and genes that have been previously reported to influence circadian rhythm or chronotype that are thought to be derived from Neanderthals and Denisovans were compiled from Dannemann & Kelso (2017), McArthur et al. (2021), Dannemann et al. (2022), and Velazquez-Arcelay et al. (2023).</p> <p><strong>SPrime ND_Match Files</strong></p> <p>Raw SPrime identified files that we used for our entire analysis. These were modified to include the archaic allele, archaic allele frequency, and average introgressed segment allele frequency. Note that these have been lifted over (Hinrichs <em>et</em> <em>al</em>., 2006) from GRCh38 (hg38) to GRCh37 (hg19) coordinates to match the genome builds of the archaic samples used in our study. As such, any manually generated variant IDs (chromosome:position:ReferenceAllele_AlternativeAllele naming convention) may no longer match the position they are currently sitting on as they were generated with hg38 coordinates. However, all of these were subsequently filtered out of our final results and any proper SNP IDs (dbSNP labels) will be accurate.</p> <p><strong>Supplementary Tables</strong></p> <p>All supplementary tables have an associated README as the first sheet that explains in detail the contents.</p> <p><strong>NEXUS Files</strong></p> <p>NEXUS files were used to generate haplotype networks in PopArt (Leigh & Bryant, 2015). There is a larger, master haplotype file and a smaller subset file. The larger file contains 668 haplotypes from all populations generated in the phased gnomAD 1KGP + HGDP callset (Koenig <em>et al</em>., 2024) for the <em>SUSD1 </em>core haplotype. The smaller subset file is the top 50 haplotypes and ties based on frequency, all Oceanic haplotypes with frequencies of at least 2, and the Neanderthal and Denisovan haplotypes for <em>SUSD1</em>. </p> <p><strong>TRAITS file</strong></p> <p>Accompanies the NEXUS files to create pie graphs for the haplotype network and contains frequency counts of number of haplotypes per region.</p>
SeMRA Gene Mappings Database
<p>Analyze the landscape of gene nomenclature resources, species-agnostic. See instructions for reproduction and usage in the attached README.md.</p>
Gene family expansions underpin context-dependency of the oldest mycorrhizal symbiosis
<p>This Zenodo archive is associated with the manuscript:</p> <p>Hernandez, D.J., Pohlmann, G.B., Afkhami, M.E. (2025) <span>Gene family expansions provide molecular flexibility required for context-dependent species interactions.</span> Ecology Letters.</p> <p>Abstract:</p> <p>As environments worldwide change at unprecedented rates during the Anthropocene, understanding context-dependency – how species regulate interactions to match changing environments – is crucial. However, generalizable molecular mechanisms underpinning context-dependency remain elusive. Combining comparative genomics across 42 angiosperms with transcriptomics, genome-wide association mapping, and gene duplication origin analyses, we show for the first time that gene family expansions undergird context-dependent regulation of species interactions. Gene families expanded in mycorrhizal fungi-associating plants display up to 200% more context-dependent gene expression and double the genetic variation associated with mycorrhizal benefits to plant fitness. Moreover, we discover these gene family expansions arise primarily from tandem duplications with >2-times more tandem duplications genome-wide, indicating gene family expansions continuously supply genetic variation throughout plant evolution allowing fine-tuning of context-dependency in species interactions.</p>
Gene family data from the PhyloGenes (release version 5.0, phylogenes.org)
<p>The data files were generated from the PhyloGenes 5.0release (see release notes <a href="https://conf.arabidopsis.org/display/PHGSUP/About+PhyloGenes">here</a>).</p> <p>About the two zip files: </p> <p>1. phyloXML.zip</p> <p>PhyloGenes gene family trees in PhyloXML format, one file per family (e.g. <family_ID>.xml).</p> <p>The following information is provided for each node of a tree:<br>1) leaf node:<br>branch length<br>name <gene_id><br>taxonomy scientific_name<br>sequence accession <UniProt ID></p> <p>2) non-leaf node:<br>branch length<br>events <duplication or speciation></p> <p><br>2. panther_csv.zip</p> <p>Functional information of family members in CSV format, one file per family (e.g. <family_ID>.csv). </p> <p>A CSV file includes the following columns:<br>Uniprot ID<br>Gene <Gene name. If none then Gene ID><br>Gene ID<br>Gene name<br>Organism<br>Subfamily name</p> <p>The columns displayed after 'Subfamily name', if any, are GO annotations. Each column is a GO molecular function or biological process term that is annotated to at least one member of the gene family AND the annotation is supported by an experimental evidence (indicated by 'EXP') or phylogenetic inference (indicated by 'IBA'). A '0' indicates absence of either annotations.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.