Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
528
datasets available to search
ShareScore release 0.9.0
Dataset results
528 results for “gene prediction”
Predicting gene expression using morphological cell responses to nanotopography
<p>This dataset contains the raw files, results files and R workspace files (.RData) associated with the paper:</p> <p>Predicting gene expression using morphological cell responses to nanotopography</p> <p>Please note that this dataset is separated according to the Figure presented in the published and peer-reviewed version of the manuscript. Particular folders contain its own README file to facilitate reproduction/replication of results and figures. </p>
Predicted genes from the Amblyomma americanum draft genome assembly
<p>Data for pub "Predicted genes from the <em>Amblyomma americanum </em>draft genome assembly."</p> <ul> <li>Amblyomma_americanum_filtered_assembly.fasta: Decontaminated A. americanum genome with bacterial contigs removed</li> <li>Amblyomma_americanum_bacterial_contigs_info.tsv: Information about contigs classified as bacteria that were removed</li> <li>Amblyomma_americanum_annotation_data.tar.gz: Directory of annotation data produced by EvidenceModeler as part of the nf-core/genomeannotator workflow. Includes files in FASTA format (predicted genes and proteins), set of proteins clustered at 99% identity in FASTA format, and annotations in both GFF3 and GTF formats. GTF file produced from the GFF3 file with AGAT.</li> <li>Amblyomma_americanum_transcriptome_assembly_data.tar.gz: Directory of data generated for the transcriptome assembly that was used for gene prediction</li> </ul>
Deep learning models predicting gene functions and pathways using public DRKG knowledge graph and graph neural network
<p>The attached dataset contains pretrained link prediction models, as described in our paper 'Morphological Map of Under- and Over-Expression of Genes in Human Cells'.</p>
COLONOMICS - predictive models for normal colon gene expression and DNA methylation for TWAS and MWAS
<p>We provide significant SNP prediction models derived from the COLONOMICS data (<a href="https://www.colonomics.org">https://www.colonomics.org</a>). Genotypes were obtained by Affymetrix 6.0 array, imputed to TopMed panel. Gene expression was obtained from Affymetrix U219 array, DNA methylation was obtained with Illuminan 450K array and miRNA expression was obtained by NGS. We provide SNP prediction models for 1,758 genes, 30,530 CpG probes and 38 miRNAs obtained from colon normal biopsy samples. These features can be predicted from SNPs located within ±1Mb, which we assumed they act through cis mechanisms. We include the model’s summary statistics and corresponding SNP weights in SQLite objects. Models were trained using the elastic net procedure employed in the PredictDB pipeline (<a href="https://predictdb.org/">https://predictdb.org</a>), according to which only models with a predictive performance p-value < 0.05 and R<sup>2</sup> > 0.1 are considered significant. We adjusted the models by basic covariates, i.e., sex, age, tissue type and colon anatomic location where biopsies were collected (left and right colon). Genome coordinates refer to GRCh37/hg19.</p>
Cladonema radiatum Alr and IgSF genes assembly and domain predictions
<p><span>This dataset is related to the </span><span>submitted paper</span><span> " A single gene determines allorecognition in hydrozoan jellyfish <em>Cladonema radiatum</em> inbred lines ".</span></p> <h1><span>Abstract:</span></h1> <p><strong><span> </span></strong></p> <p><span>Allorecognition—the ability of an organism to discriminate between self and non-self—is crucial to colonial marine animals to avoid invasion by other individuals in the same habitat. The cnidarian hydroid <em>Hydractinia</em> has long been a major research model in studying invertebrate allorecognition, establishing a rich knowledge foundation. In this study, we introduce a new cnidarian model <em>Cladonema radiatum</em> (<em>C. radiatum</em>). <em>C. radiatum</em> is a hydroid jellyfish which also forms polyp colonies interconnected with stolons. Allorecognition responses, fusion or regression of stolons, are observed when stolons encounter each other. By transmission electron microscopy, we observe rapid tissue remodelling contributing to gastrovascular system connection in fusion. Rejection responses are regulated by reconstruction of the chitinous exoskeleton perisarc, and induction of necrotic and autophagic cellular responses at cells in contact with the opponent. Genetic analysis identifies allorecognition genes: six<em> Alr </em>genes located on the putative Allorecognition Complex (ARC) and four immunoglobulin superfamily genes on a separate genome region. C. radiatum allorecognition genes show notable conservation with the <em>Hydractinia Alr</em> family. Remarkedly, stolon encounter assays of inbred lines reveal that genotypes of Alr1 solely determine allorecognition outcomes in <em>C. radiatum</em>.</span></p>
Ontology based text mining of gene-phenotype associations: application to candidate gene prediction
<p>Gene-phenotype associations play an important role in understanding<br> the disease mechanisms which is a requirement for treatment<br> development. A portion of gene-phenotype associations are observed<br> mainly experimentally and made publicly available through several<br> standard resources such as MGI. However, there is still a vast<br> amount of gene--phenotype associations buried in the biomedical<br> literature. Given the large amount of literature data, we need<br> automated text mining tools to alleviate the burden in manual<br> curation of gene-phenotype associations and to develop<br> comprehensive resources. We developed an ontology based<br> approach in combination with statistical methods to text mine<br> gene-phenotype associations from literature. Our method achieved<br> AUC values of 0.90 and 0.75 in recovering known gene-phenotype<br> associations from HPO and MGI respectively. We posit that candidate<br> genes and their relevant diseases should be expressed with similar<br> phenotypes in publications. Thus, we demonstrate the utility of our<br> approach by predicting disease candidate genes based on the semantic<br> similarities of phenotypes associated with genes and diseases. We evaluated our disease candidate prediction model on<br> the gene-disease associations from MGI. Our model achieved AUC<br> values of 0.90 and 0.87 on OMIM (human) and MGI (mouse) datasets of<br> gene-disease associations respectively. Our manual analysis on the<br> text mined data revealed that, our method can accurately extract<br> gene-phenotype associations which are not currently covered by the<br> existing public gene-phenotype resources. Overall, results indicate<br> that our method can precisely extract known as well as new<br> gene-phenotype associations from literature. This released dataset at Zenodo covers our gene-phenotype extracts from the literature. All the methods used to extract the data are available at https://github.com/bio-ontology-research-group/genepheno.</p>
Data from: External validation of prognostic and predictive gene signatures in 1097 European head and neck squamous cell carcinoma patients
<p><span>Anonymized data containing survival endpoints and gene signature scores for head and neck cancer patients.</span></p> <p><span>File <strong>data_os_gs.csv</strong> : data linking overall survival and gene signature scores</span></p> <p><span>File <strong>data_dfs_gs.csv</strong> : data linking disease-free survival and gene signature scores</span></p> <p><span><strong>Variables</strong>:</span></p> <ul> <li><span><em>supertreat_id</em>: patient ID</span></li> <li><span><em>GS_score_172GS</em>: gene signature score for the <em>172-GS</em> signature. The score is Z-score normalized with a mean of 0 and SD of 1. </span></li> <li><span><em>GS_score_3clustersHPV</em>: gene signature score for the <em>3 clusters HPV</em> signature. The score is Z-score normalized with a mean of 0 and SD of 1. </span></li> <li><span><em>GS_score_RSI</em>: gene signature score for the <em>radiosenstivity index (RSI) </em>signature. The score is Z-score normalized with a mean of 0 and SD of 1. </span></li> <li><span><em>GS_score_pancancerCisplatin</em>: gene signature score for the <em>pancancer-cisplatin</em> signature. The score is Z-score normalized with a mean of 0 and SD of 1. </span></li> <li><span><em>GS_score_cl3Hypoxia</em>: gene signature score for the <em>Cl3-hypoxia</em> signature. The score is Z-score normalized with a mean of 0 and SD of 1. </span></li> <li><span>Variables only available in <strong>data_os_gs.csv: </strong></span> <ul> <li><span><em>overall_survival_days_2years</em>: Overall survival censored at 2 years since diagnosis. Number of days from diagnosis to death or censoring.</span></li> <li><span><em>overall_survival_days_5years</em>: Overall survival censored at 5 years since diagnosis. Number of days from diagnosis to death or censoring.</span></li> <li><span><em>overall_survival_status_2years</em>: Overall survival status when censored at 2 years since diagnosis. Coded as 0 if censored, and 1 if dead. </span></li> <li><span><em>overall_survival_status_5years</em>: Overall survival status when censored at 5 years since diagnosis. Coded as 0 if censored, and 1 if dead. </span></li> </ul> </li> </ul> <ul> <li><span>Variables only available in <strong>data_dfs_gs.csv:</strong></span> <ul> <li><span><em>disease_free_survival_days_2years</em>: Disease-free survival censored at 2 years since diagnosis. Number of days from diagnosis to an event (death or cancer recurrence) or censoring.</span></li> <li><span><em>disease_free_survival_days_5years</em>: Disease-free survival censored at 5 years since diagnosis. Number of days from diagnosis to an event (death or cancer recurrence) or censoring.</span></li> <li><span><em>disease_free_survival_status_2years</em>: Disease-free survival status when censored at 2 years since diagnosis. Coded as 0 if censored, and 1 if an event (death or recurrence). </span></li> <li><span><em>disease_free_survival_status_5years</em>: Disease-free survival status when censored at 5 years since diagnosis. Coded as 0 if censored, and 1 if an event (death or recurrence). </span></li> </ul> </li> </ul>
ZIRFs: zero-inflated random forests for estimating gene regulatory networks from single cell RNA-seq data (assessment of predictive accuracy and VIM stability)
<p>We developed a zero-inflated random forests (ZIRFs) algorithm to produce a metric of connection strength between regulator genes and target genes. This file contains SCENIC results for the aorta and diaphragm tissue data sets from the Tabula Muris Consortium results. SCENIC is a genetic regulatory network analysis published by Aibar et al. (2017). The purpose of the data sets and R source code are described by README files in each directory.</p>
Gene Regulatory Network inference in long lived C.elegans reveals modular properties that are predictive of novel ageing genes - Supplementary Tables
<p>This repository contains the Supplementary Tables for Suriyalaksh et al. Gene Regulatory Network inference in long lived C.elegans reveals modular properties that are predictive of novel ageing genes.</p> <p>The list of table files can be found in <a href="https://zenodo.org/api/files/431155b3-c2b2-4fff-92bb-e93cad8bf5c2/Supplementary%20table%20guide.pdf">Supplementary Tables guide.pdf</a></p> <p>Tables S1, S2 and S3 corresponding to physical gene-gene interaction data are in a separate repository doi:10.5281/zenodo.4382337</p> <p>Details about some of the Supplementary tables:</p> <p>TableS4_inferred_networks.csv - list of inferred GRNs for specified input combinations (set of input regulators, length of the time sequence, NI tool and prior used).</p> <p><a href="https://zenodo.org/api/files/431155b3-c2b2-4fff-92bb-e93cad8bf5c2/TableS5_consensus_network_member.xlsx">TableS5_consensus_network_member.xlsx</a> - list of groups of topologically similar GRNs (from Table S4)</p> <p>Table S6: edge lists (source,target) for each one of the three consensus networks selected according to the GS validation metrics: middle PFE/AUFE, max AUFE, max PFE.<br> TableS6a_max_AUFE_GRN.txt - max AUFE; largest network - this is the one we used in the main analysis and discussion<br> TableS6b_max_PFE_GRN.xt - max PFE<br> TableS6c_middle_AUFE_PFE_GRN.txt - middle PFE/AUFE</p> <p><a href="https://zenodo.org/api/files/431155b3-c2b2-4fff-92bb-e93cad8bf5c2/TableS7_qRTPCR_ddCt_network_accuracy.csv">TableS7_qRTPCR_ddCt_network_accuracy.csv</a> - gene expression count differences for RNAi knockdown GRN validation experiments. </p> <p>Table S8: Group membership for each one of the nodes in each one of the selected networks according to the SBM that best describes the observed network topology. Each column shows the group membership for each level in a SBM block hierarchy. Our analysis is in the second most coarse-grained level (level 1).</p> <p>TableS8a_max_AUFE_SBM.csv<br> TableS8b_max_PFE_SBM.csv<br> TableS8c_middle_AUFE_PFE_SBM.csv</p> <p><a href="https://zenodo.org/api/files/431155b3-c2b2-4fff-92bb-e93cad8bf5c2/TableS9_glp_gs_datasets.pdf">TableS9_glp_gs_datasets.pdf</a> - list of datasets used for defining functional clusters.</p> <p>TableS14a_glp_l1_vs_fem_l1_lifespan_assay.xlsx - Day13 survival of fem-3(q20)ts vs day 19 survival of glp-1(e2144)ts;rrf-3(pk1426) RNAi from L1</p> <p>TableS14b_glp_l1_vs_glp_l4_lifespan_assay.xlsx - Day 19 survival of glp-1(e2144)ts;rrf-3(pk1426) RNAi from L4 vs day 19 survival of glp-1(e2144)ts;rrf-3(pk1426) RNAi from L1</p> <p>TableS15a_glp1_in_vivo_fluorescence_data.xlsx - in vivo fluorescent reporter data of glp-1(e2144)ts;rrf-3(pk1426)</p> <p>TableS15b_fem3_in_vivo_fluorescence_data.xlsx - in vivo fluorescent reporter data of fem-3(q20)ts</p> <p>TableS17_input_regulators_annotated.csv - list input regulators used as input for Network Inference Tools annotated by source type (2nd column): GenAge, known transcription factors (TF) and gene with high variability in the gene expression time series (HV). The third column lists whether that regulator has an orthologue in human (y) according to WormBase (v 278).</p> <p>TableS20_epistasis_lifespan_data.xlsx - Epistasis lifespan data of glp-1(e2144)ts</p> <p>All the image (TIF) files represent representative images in the following genetic backgrounds (below) that have been treated </p> <p>with empty vector (EV) or RNAi against the gene highlighted in the title of the image. See methods section for details. </p> <p><strong>femliu1: </strong></p> <p><em>fem-3(q20)ts.; dhs-3p::dhs-3::gfp</em></p> <p><strong>femsod3:</strong></p> <p><em>fem-3(q20)ts.; sod-3p::gfp</em></p> <p><strong>glp1lgg1:</strong></p> <p><em>glp-1(e2144); lgg-1p:lgg-1:gfp</em></p>
Xpresso: Predicting gene expression levels from genomic sequences
<p>Xpresso: Predicting gene expression levels from genomic sequences<br> <br> More info at:<br> Publication: https://doi.org/10.1016/j.celrep.2020.107663<br> Website: https://xpresso.gs.washington.edu/<br> Github: https://github.com/vagarwal87/Xpresso</p>
Predicting placenta transcriptional regulatory interactions based on spatial gene expression data and convolutional neural network
<p><strong>Aims:</strong> The dysfunction of placenta development is correlated to the defects of pregnancy and fetal growth. The detailed molecular mechanism of placenta development is not identified in human due to the lack of material in vivo. Image-based reconstructions of GRN are still very underdeveloped.</p> <p><strong>Methods and Results:</strong> In this study, first-trimester chorionic villus and decidua tissues were collected. Next, we present a machine-learning system to infer gene interaction networks of the human placenta from immunofluorescence images of trophoblast specific transcription factors obtained by a high-resolution scanner.</p> <p><strong>Conclusions:</strong> The experimental results show that deep learning models reveal regulatory roles that have not yet been fully recognized. The spatial expression data reveal new regulatory relationships that traditional experiments have failed to recognize, and has allowed the development of gene regulation networks based on the spatial distribution of gene expression. We demonstrate the effectiveness of this approach in building networks using high-resolution images of the human placenta. Our analysis is of certain significance for further exploration of the development of the placenta and the occurrence of pregnancy-related diseases in the future. The datasets and analysis provide a useful source for the researchers in the field of the maternal-fetal interface and the establishment of pregnancy.</p>
Gene Regulatory Network inference in long lived C.elegans reveals modular properties that are predictive of novel ageing genes - Database of Physical gene-gene Interactions in young adult C.elegans.
<p>This repository contains Supplementary Information for manuscript Suriyalaksh et al Gene Regulatory Network inference in long lived C.elegans reveals modular properties that are predictive of novel ageing genes corresponding to the curation of physical gene-gene interactions for young adult C elegans worms </p> <p>We manually curated 239,001 regulatory interactions from 289 young adult wild-type (WT) C.elegans datasets, consisting of 126 genes and 495 unique transcription factors (see TableS1_datasets_for_prior.csv for references). </p> <p>This repository contains 3 different files:</p> <p>TableS1_datasets_for_prior.csv - contains datasets used as sources for physical gene-gene or TF-gene interactions</p> <p>TableS2_physical_priors.xlsx - contains three tabs:<br> ChIPATAC - contains physical TF-gene interactions from 115 L4 or young-adult ChIP-seq datasets from modERN (Kudron et al., 2018) + ChIP-seq datasets (GSE28350, GSE81521) from (Hochbaum et. al, 2011, Li et. al, 2016).</p> <p>eY1HATAC- contains 3,501 TF-gene interactions from eY1H assay by Fuxman Bass et al. (2016).</p> <p>motifATAC - contains 202 unique TF DNA recognition motifs using “direct evidence” option from CiS-BP motif database (Weirauch et al., 2014), obtained through RTFBSDB R package (Wang et al., 2016) - see TableS1</p> <p>TableS3_WT_functional_priors.csv - contains functional knockdown data that we use as gold standard to validate inferred networks in Suriyalaksh et al. (see TableS1_datasets_for_prior.csv for sources)</p> <p>---</p> <p>Description of methodology to obtain regulatory interactions in TableS2:</p> <p>Regulatory sequences for each gene were acquired from ENSEMBL (Aken et al., 2017), obtained using biomaRt R package (accessed on 31st Oct 2017). This study used WBcel235/ce11 version of the C. elegans genome, and WormBase WS260 genome annotations.</p> <p>For motifs, TFs whose motifs overlapped with an open ATAC-seq region by at least one base pair were kept. For ChIP-seq, TF binding sites that overlapped with an open ATAC-seq region by at least one base pair were kept using bedtools intersect and bedtools merge commands.</p> <p>An interaction from a TF to a gene was inferred by aligning transcription start sites (TSS) using bedtools window commands with 1000 bp window size to the TF-binding locations from ChIP-seq and motifs.</p> <p>For eY1H data, an interaction is included if the TSS site of the target gene overlaps with an open ATAC-seq region by at least one base pair.</p> <p>For gene-gene interactions, of the 298 studies compiled in WormExp v1.0 database (Yang et al, 2016, updated 27/07/16), 98 studies were included in the database spanning 126 different genes (see Table S1 in this repository).</p>
Fecal-bbu-genes-quantification-predicts-L-carnitine-mediated-TMAO-production-and-serves-as-a-biomarker-for-precision-nutrition-code-20231201
<p>Custom code related to the original research article "Fecal bbu genes quantification predicts L-carnitine-mediated TMAO production and serves as a biomarker for precision nutrition"</p>
JSON files containing parameters of training gene models for ab-initio prediction software
<p>These are the JSON files containing parameters of training gene models for ab-initio prediction software. These training datasets are Phytophthora specific and can be further utilized for the gene prediction and annotation of other related Phytophthora strains.</p>
# Single-cell network biology characterizes cell type gene regulation for drug repurposing and phenotype prediction in Alzheimer's disease
<p>Dysregulation of gene expression in Alzheimer’s disease (AD) remains elusive, especially at the cell type level. Gene regulatory network, a key molecular mechanism linking transcription factors (TFs) and regulatory elements to govern target gene expression, can change across cell types in the human brain and thus serve as a model for studying gene dysregulation in AD. However, it is still challenging to understand how cell type networks work abnormally under AD. To address this, we integrated single-cell multi-omics data and predicted the gene regulatory networks in AD and control for four major cell types, excitatory and inhibitory neurons, microglia and oligodendrocytes. Importantly, we applied network biology approaches to analyze the changes of network characteristics across these cell types, and between AD and control. For instance, many hub TFs target different genes between AD and control (rewiring). Also, these networks show strong hierarchical structures in which top TFs (master regulators) are largely common across cell types, whereas different TFs operate at the middle levels in some cell types (e.g., microglia). The regulatory logics of enriched network motifs (e.g., feed-forward loops) further uncover cell type-specific TF-TF cooperativities in gene regulation. The cell type networks are highly modular and several network modules with cell-type-specific expression changes in AD pathology are enriched with AD-risk genes and putative targets of approved and pending AD drugs, suggesting possible cell-type genomic medicine in AD. Finally, using the cell type gene regulatory networks, we developed machine learning models to classify and prioritize additional AD genes. We found that top prioritized genes predict clinical phenotypes (e.g., cognitive impairment) with reasonable accuracy. Overall, this single-cell network biology analysis provides a comprehensive map linking genes, regulatory networks, cell types and drug targets and reveals dysregulated cell type gene dysregulatory mechanisms in AD.</p>
Predicted gene expression in ancestrally diverse populations leads to discovery of susceptibility loci for lifestyle and cardiometabolic traits
<p>Full summary statistics for the publication "Predicted gene expression in ancestrally diverse populations leads to discovery of susceptibility loci for lifestyle and cardiometabolic traits". </p> <p>The files, bmi.UKBBsummary.txt and height.UKBBsummary.txt, contain tissue specific associations with body mass index (BMI) and height respectively. The suffix UKBB450k indicates results from all ~450,000 European ancestry individuals in UK Biobank. The suffix UKBB50k corresponds to results from a subset of 50,000 Europeans in the UK Biobank. The suffix PAGE corresponds to results from ~50,000 individuals in the Population Architecture using Genomics and Epidemiology (PAGE) study. </p> <p>The file PAGE_PrediXcan_associations.txt includes trait~tissue specific GReX associations for 25 traits. The first field specifies the tissue.trait.gene of the association results.</p>
Aberrant gene expression prediction benchmark based on GTEx v8
<p>This repository contains the aberrant gene expression prediction benchmark data as well as the necessary expected gene expression across tissues and tissue-specific isoform contribution scores for AbExp prediction.<br> </p> <p>The aberrant gene expression prediction benchmark data (aberrant_expression_prediction_benchmark.parquet) contains the following columns:</p> <ul> <li>individual: GTEx individual</li> <li>gene: Ensembl gene identifier</li> <li>tissue: GTEx tissue</li> <li>tissue_type: GTEx tissue type</li> <li>mu: OUTRIDER-estimated expected gene expression</li> <li>theta: OUTRIDER-estimated gene dispersion</li> <li>counts: Raw gene expression count</li> <li>normalized_counts: OUTRIDER-normalized gene expression count</li> <li>l2fc: log2 fold change between observed and expected gene expression count</li> <li>zscore: z-score of gene expression, obtained by quantile-mapping the OUTRIDER-estimated distribution to the standard normal distribution</li> <li>nominal_pvalue: OUTRIDER-estimated <em>p</em>-value of being an expression outlier</li> <li>FDR: FDR-adjusted <em>p</em>-value of being an expression outlier</li> <li>is_in_benchmark: Whether this observation is part of the aberrant gene expression prediction benchmark</li> <li>is_underexpressed_outlier: Whether this observation is an underexpression outlier at FDR < 5%. This is the benchmark prediction label.</li> </ul> <p><br>The isoform proportions table (gtex_v8_isoform_proportions.tsv) contains the following columns:</p> <ul> <li>gene: Ensembl gene identifier</li> <li>tissue_type: GTEx tissue type</li> <li>tissue: GTEx tissue</li> <li>transcript: Ensembl transcript identifier</li> <li>mean_transcript_proportions: mean transcript proportions across individuals in GTEx v8</li> <li>median_transcript_proportions: median transcript proportions across individuals in GTEx v8</li> <li>sd_transcript_proportions: standard deviation of transcript proportions across individuals in GTEx v8</li> </ul> <p><br>The expected gene expression table (gtex_v8_expected_expression.tsv) contains the following columns:</p> <ul> <li>gene: Ensembl gene identifier</li> <li>tissue_type: GTEx tissue type</li> <li>tissue: GTEx tissue</li> <li>gene_is_expressed: Whether the gene is expressed in the tissue</li> <li>median_expression: median OUTRIDER-estimated expected gene expression (mu) across individuals</li> <li>expression_dispersion: OUTRIDER-estimated gene dispersion (theta)</li> </ul>
DeepARG: a deep learning approach for predicting antibiotic resistance genes from metagenomic data
<p>Growing concerns about increasing rates of antibiotic resistance call for expanded and comprehensive global monitoring. Advancing methods for monitoring of environmental media (e.g., wastewater, agricultural waste, food, and water) is especially needed for identifying potential resources of novel antibiotic resistance genes (ARGs), hot spots for gene exchange, and as pathways for the spread of ARGs and human exposure. Next-generation sequencing now enables direct access and profiling of the total metagenomic DNA pool, where ARGs are typically identified or predicted based on the “best hits” of sequence searches against existing databases. Unfortunately, this approach produces a high rate of false negatives. To address such limitations, we propose here a deep learning approach, taking into account a dissimilarity matrix created using all known categories of ARGs. Two deep learning models, DeepARG-SS and DeepARG-LS, were constructed for short read sequences and full gene length sequences, respectively. Evaluation of the deep learning models over 30 antibiotic resistance categories demonstrates that the DeepARG models can predict ARGs with both high precision (> 0.97) and recall (> 0.90). The models displayed an advantage over the typical best hit approach, yielding consistently lower false negative rates and thus higher overall recall (> 0.9). As more data become available for under-represented ARG categories, the DeepARG models’ performance can be expected to be further enhanced due to the nature of the underlying neural networks. Our newly developed ARG database, DeepARG-DB, encompasses ARGs predicted with a high degree of confidence and extensive manual inspection, greatly expanding current ARG repositories. The deep learning models developed here offer more accurate antimicrobial resistance annotation relative to current bioinformatics practice. DeepARG does not require strict cutoffs, which enables identification of a much broader diversity of ARGs. The DeepARG models and database are available as a command line version and as a Web service at <a href="http://bench.cs.vt.edu/deeparg">http://bench.cs.vt.edu/deeparg</a>.</p>
Exploring the utility of regulatory network-based machine learning for gene expression prediction in maize
<p>Relevant Data and Code for <em>Exploring the utility of regulatory network-based machine learning for gene expression prediction in maize </em>by Taylor Ferebee and Edward Buckler.</p> <p><strong>Input Data</strong></p> <p>The inputs of the models are enclosed in <em>Input_data-2022-001.zip</em></p> <p><strong>Output Data</strong></p> <p>The outputs of the models are enclosed in <em>Output_Results-2022-001.zip</em></p> <p><strong>Relevant Code </strong></p> <p>The code for all analyses is enclosed in <em>Code_Archive.zip </em></p>
Data for: "Generative prediction of causal gene sets responsible for complex traits"
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.