Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
48,977
datasets available to search
ShareScore release 0.7.1
Dataset results
48,977 results for “Genes”
Stimulating Wnt signaling reveals context-dependent genetic effects on gene regulation in primary human neural progenitors
<p>Summary statistics for chromatin accessibility and gene expression quantitative trait loci (ca/eQTLs) from Matoba, N., Le, B.D., Valone, J.M. <em>et al.</em> Stimulating Wnt signaling reveals context-dependent genetic effects on gene regulation in primary human neural progenitors. <em>Nat Neurosci</em> (2024). https://doi.org/10.1038/s41593-024-01773-6</p>
Processed Data for "Improving Gene Regulatory Network Inference using Dropout Augmentation"
<p>Here are the processed dataset that are used in the manuscript "Improving Gene Regulatory Network Inference using Dropout Augmentation"</p>
Data from: External validation of prognostic and predictive gene signatures in 1097 European head and neck squamous cell carcinoma patients
<p><span>Anonymized data containing survival endpoints and gene signature scores for head and neck cancer patients.</span></p> <p><span>File <strong>data_os_gs.csv</strong> : data linking overall survival and gene signature scores</span></p> <p><span>File <strong>data_dfs_gs.csv</strong> : data linking disease-free survival and gene signature scores</span></p> <p><span><strong>Variables</strong>:</span></p> <ul> <li><span><em>supertreat_id</em>: patient ID</span></li> <li><span><em>GS_score_172GS</em>: gene signature score for the <em>172-GS</em> signature. The score is Z-score normalized with a mean of 0 and SD of 1. </span></li> <li><span><em>GS_score_3clustersHPV</em>: gene signature score for the <em>3 clusters HPV</em> signature. The score is Z-score normalized with a mean of 0 and SD of 1. </span></li> <li><span><em>GS_score_RSI</em>: gene signature score for the <em>radiosenstivity index (RSI) </em>signature. The score is Z-score normalized with a mean of 0 and SD of 1. </span></li> <li><span><em>GS_score_pancancerCisplatin</em>: gene signature score for the <em>pancancer-cisplatin</em> signature. The score is Z-score normalized with a mean of 0 and SD of 1. </span></li> <li><span><em>GS_score_cl3Hypoxia</em>: gene signature score for the <em>Cl3-hypoxia</em> signature. The score is Z-score normalized with a mean of 0 and SD of 1. </span></li> <li><span>Variables only available in <strong>data_os_gs.csv: </strong></span> <ul> <li><span><em>overall_survival_days_2years</em>: Overall survival censored at 2 years since diagnosis. Number of days from diagnosis to death or censoring.</span></li> <li><span><em>overall_survival_days_5years</em>: Overall survival censored at 5 years since diagnosis. Number of days from diagnosis to death or censoring.</span></li> <li><span><em>overall_survival_status_2years</em>: Overall survival status when censored at 2 years since diagnosis. Coded as 0 if censored, and 1 if dead. </span></li> <li><span><em>overall_survival_status_5years</em>: Overall survival status when censored at 5 years since diagnosis. Coded as 0 if censored, and 1 if dead. </span></li> </ul> </li> </ul> <ul> <li><span>Variables only available in <strong>data_dfs_gs.csv:</strong></span> <ul> <li><span><em>disease_free_survival_days_2years</em>: Disease-free survival censored at 2 years since diagnosis. Number of days from diagnosis to an event (death or cancer recurrence) or censoring.</span></li> <li><span><em>disease_free_survival_days_5years</em>: Disease-free survival censored at 5 years since diagnosis. Number of days from diagnosis to an event (death or cancer recurrence) or censoring.</span></li> <li><span><em>disease_free_survival_status_2years</em>: Disease-free survival status when censored at 2 years since diagnosis. Coded as 0 if censored, and 1 if an event (death or recurrence). </span></li> <li><span><em>disease_free_survival_status_5years</em>: Disease-free survival status when censored at 5 years since diagnosis. Coded as 0 if censored, and 1 if an event (death or recurrence). </span></li> </ul> </li> </ul>
Data from: Polymorphic tandem repeats shape single-cell gene expression across the immune landscape
<p>This dataset contains the association summary statistics (v0.1) for genome-wide tandem repeat (TR) expression quantitative trait (eQTL) analysis of TenK10K Phase 1 (https://doi.org/10.1101/2024.11.02.621562). </p> <p>Please access the README for a detailed description of file contents. </p> <p> </p>
Gene Annotations of 49 Bacillariophyta Genome Assemblies (Individual gff3 files)
<div>Contact: katharina.hoff@uni-greifswald.de.</div> <div> </div> <div> <h2>Manuscript</h2> <p>The data hosted here is associated with the preprint <a href="https://doi.org/10.48550/arXiv.2410.05467">https://doi.org/10.48550/arXiv.2410.05467</a> . It is a copy of the data hostet at <a href="https://zenodo.org/records/13933292">https://zenodo.org/records/13933292</a> , but instead of storing one archive will all gff3 files included, the gff3 files are here hosted, individually. This copy was made upon request from the RDA Working Group "FAIRification of Genomic Annotations – metadata harmonisation at scale".</p> <div> <h2>Files</h2> <div>The following gzip-compressed gff3-files with structural and functional genome annotation are included:</div> <div> </div> <div>Asterionella_formosa.gff3.gz<br>Asterionellopsis_glacialis.gff3.gz<br>Bacterosira_constricta.gff3.gz<br>Chaetoceros_muellerii.gff3.gz<br>concatenated_output.gff3.gz<br>Conticribra_guillardii.gff3.gz<br>Conticribra_weissflogii.gff3.gz<br>Craspedostauros_australis.gff3.gz<br>Cyclostephanos_invisitatus.gff3.gz<br>Cyclostephanos_tholiformis.gff3.gz<br>Cyclotella_atomus.gff3.gz<br>Cyclotella_baltica.gff3.gz<br>Cyclotella_choctawhatcheeana.gff3.gz<br>Cyclotella_cryptica.gff3.gz<br>Cylindrotheca_fusiformis.gff3.gz<br>Detonula_confervacea.gff3.gz<br>Discostella_pseudostelligera.gff3.gz<br>Discostella_stelligera.gff3.gz<br>Discostella_stelligeroides.gff3.gz<br>Epithemia_pelagica.gff3.gz<br>Fistulifera_pelliculosa.gff3.gz<br>Fistulifera_solaris.gff3.gz<br>Fragilaria_radians.gff3.gz<br>Fragilariopsis_cylindrus.gff3.gz<br>Licmophora_abbreviata.gff3.gz<br>Mediolabrus_comicus.gff3.gz<br>Nitzschia_palea.gff3.gz<br>Nitzschia_putrida.gff3.gz<br>Porosira_glacialis.gff3.gz<br>Psammoneis_japonica.gff3.gz<br>Pseudo-nitzschia_multiseries.gff3.gz<br>Pseudo-nitzschia_pungens.gff3.gz<br>Skeletonema_costatum.gff3.gz<br>Skeletonema_marinoi.gff3.gz<br>Skeletonema_menzelii.gff3.gz<br>Skeletonema_potamos.gff3.gz<br>Skeletonema_tropicum.gff3.gz<br>Stephanocyclus_meneghinianus.gff3.gz<br>Stephanodiscus_minutulus.gff3.gz<br>Stephanodiscus_triporus.gff3.gz<br>Thalassiosira_allenii.gff3.gz<br>Thalassiosira_delicatula.gff3.gz<br>Thalassiosira_exigua.gff3.gz<br>Thalassiosira_gravida.gff3.gz<br>Thalassiosira_livingstoniorum.gff3.gz<br>Thalassiosira_mediterranea.gff3.gz<br>Thalassiosira_oceanica.gff3.gz<br>Thalassiosira_ordinaria.gff3.gz<br>Thalassiosira_pacifica.gff3.gz<br>Thalassiosira_profunda.gff3.gz</div> <div> </div> <div>To extract individual files after download execute the following command:</div> <div> </div> <div><code>gunzip *.gff3.gz</code></div> <h2>Genome Assemblies</h2> <p> </p> <div>The files in this folder attain to genome assemblies are publicly available at NCBI datasets (https://www.ncbi.nlm.nih.gov/datasets/). We used the following versions:</div> <p> </p> <div>Asterionella formosa GCA_002256025.1</div> <div>Asterionellopsis glacialis GCA_014885115.2</div> <div>Bacterosira constricta GCA_037356235.1</div> <div>Chaetoceros muellerii GCA_019693545.1</div> <div>Conticribra guillardii GCA_036939335.1</div> <div>Conticribra weissflogii GCA_036940025.1</div> <div>Craspedostauros australis GCA_026770025.1</div> <div>Cyclostephanos invisitatus GCA_036939675.1</div> <div>Cyclostephanos tholiformis GCA_036939975.1</div> <div>Cyclotella atomus GCA_036939935.1</div> <div>Cyclotella baltica GCA_036939635.1</div> <div>Cyclotella choctawhatcheeana GCA_036939855.1</div> <div>Cyclotella cryptica GCA_013187285.1</div> <div>Cylindrotheca fusiformis GCA_019693525.1</div> <div>Detonula confervacea GCA_036939415.1</div> <div>Discostella pseudostelligera GCA_036940085.1</div> <div>Discostella stelligera GCA_036939735.1</div> <div>Discostella stelligeroides GCA_036939555.1</div> <div>Epithemia pelagica GCA_946965045.2</div> <div>Fistulifera pelliculosa GCA_026008555.1</div> <div>Fistulifera solaris GCA_030295235.1</div> <div>Fragilaria radians GCA_900642245.1</div> <div>Fragilariopsis cylindrus GCA_900095095.1</div> <div>Licmophora abbreviata GCA_900291995.1</div> <div>Mediolabrus comicus GCA_036940125.1</div> <div>Nitzschia palea GCA_019593585.1</div> <div>Nitzschia putrida GCA_016586335.1</div> <div>Porosira glacialis GCA_036939395.1</div> <div>Psammoneis japonica GCA_008632985.1</div> <div>Pseudo-nitzschia multiseries GCA_037355745.1</div> <div>Pseudo-nitzschia pungens GCA_037355855.1</div> <div>Skeletonema costatum GCA_018806925.1</div> <div>Skeletonema marinoi GCA_030544225.1</div> <div>Skeletonema menzelii GCA_036940005.1</div> <div>Skeletonema potamos GCA_036940105.1</div> <div>Skeletonema tropicum GCA_037178625.1</div> <div>Stephanocyclus meneghinianus GCA_036940045.1</div> <div>Stephanodiscus minutulus GCA_036939435.1</div> <div>Stephanodiscus triporus GCA_036939755.1</div> <div>Thalassiosira allenii GCA_036939655.1</div> <div>Thalassiosira delicatula GCA_036939835.1</div> <div>Thalassiosira exigua GCA_036939895.1</div> <div>Thalassiosira gravida GCA_037356215.1</div> <div>Thalassiosira livingstoniorum GCA_036939595.1</div> <div>Thalassiosira mediterranea GCA_036939795.1</div> <div>Thalassiosira oceanica GCA_019693575.1</div> <div>Thalassiosira ordinaria GCA_036939695.1</div> <div>Thalassiosira pacifica GCA_036939875.1</div> <div>Thalassiosira profunda GCA_036939355.1</div> <p> </p> <h2>Converting to Protein FASTA and Coding Sequences FASTA</h2> <p> </p> <div>To save storage place at Zenodo, we did not upload the protein FASTA and coding sequence FASTA files. They can easily be generated from the genome FASTA file in combination with the respective GFF3 file. To do this, you can use the following commands:</div> <p> </p> <div><code># assume that genome.fa ist you respective genome FASTA file downloaded from NCBI datasets</code></div> <div><code>sed '/^>/ s/ .*//' genome.fasta > genome_short_headers.fasta</code></div> <div><code># assume that file.gff is the respective GFF3 file</code></div> <div><code>getAnnoFastaFromJoingenes.py -g genome_short_headers.fasta -3 file.gff -o nameStem</code></div> <p> </p> <div>This will produce the following files: nameStem.aa (protein FASTA file) and nameStem.codingseq (coding sequence FASTA file).</div> <p> </p> <div>The getAnnoFastaFromJoingenes.py script is available at https://raw.githubusercontent.com/Gaius-Augustus/Augustus/master/scripts/getAnnoFastaFromJoingenes.py . It is part of the AUGUSTUS software package.</div> <h2>Release notes</h2> <p>The submission and release was made upon request of the RDA working group "FAIRification of Genomic Annotations – metadata harmonisation at scale". The contained data is identical to <a href="https://zenodo.org/records/13933292">https://zenodo.org/records/13933292</a></p> <h2>License</h2> <p> </p> <div>The genome annotation files are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.</div> <p> </p> </div> </div>
Identification of Genes Regulating Dexamethasone Resistance and Prognostic Model Development in Acute Lymphoblastic Leukemia
<p>This study investigates the mechanisms of dexamethasone resistance in acute lymphoblastic leukemia (ALL) and presents a prognostic model to predict patient outcomes and immunotherapy responses. By analyzing gene expression data, we identified autophagy-related genes associated with dexamethasone resistance, particularly focusing on STK38L’s role in modulating autophagy via ULK1. Our results reveal that high STK38L expression enhances dexamethasone resistance by promoting autophagy markers LC3II/LC3I and beclin-1. This study provides valuable insights into the molecular basis of dexamethasone resistance and highlights STK38L as a potential biomarker and therapeutic target for improving ALL treatment strategies.</p>
Direct molecular evidence for an ancient, conserved developmental toolkit controlling post-transcriptional gene regulation in land plants
<p>In plants, miRNA production is orchestrated by a suite of proteins that control transcription of the pri-miRNA gene, post-transcriptional processing and nuclear export of the mature miRNA. Post-transcriptional processing of miRNAs is controlled by a pair of physically-interacting proteins, HYL1 and DCL1. However, the evolutionary history and structural basis of the HYL1-DCL1 interaction is unknown. Here we use ancestral sequence reconstruction and functional characterization of ancestral HYL1 <em>in vitro</em> and in <em>Arabidopsis thaliana </em>to better understand the origin and evolution of the HYL1-DCL1 interaction and its impact on miRNA production and plant development. We found the ancestral plant HYL1 evolved high affinity for both double-stranded RNA (dsRNA) and its DCL1 partner before the divergence of mosses from seed plants (~500 Ma), and these high-affinity interactions remained largely conserved throughout plant evolutionary history. Structural modeling and molecular binding experiments suggest that the second of two double-stranded RNA-binding motifs (DSRMs) in HYL1 may interact tightly with the first of two C-terminal DCL1 DSRMs to mediate the HYL1-DCL1 physical interaction necessary for efficient miRNA production. Transgenic expression of the nearly 200 Ma-old ancestral flowering-plant HYL1 in <em>A. thaliana</em> was sufficient to rescue many key aspects of plant development disrupted by HYL1<sup>-</sup> knockout and restored near-native miRNA production, suggesting that the functional partnership of HYL1-DCL1 originated very early in and was strongly conserved throughout the evolutionary history of terrestrial plants. Overall, our results are consistent with a model in which miRNA-based gene regulation evolved as part of a conserved plant ‘developmental toolkit’.</p>
Tissue heterogeneity is prevalent in gene expression studies
<p>This archive contains results associated with the publication</p> <p><em>Tissue heterogeneity is prevalent in gene expression studies. Gregor Sturm, Markus List and Jitao David Zhang.</em></p> <p> </p> <ul> <li>expr.tissuemark.affy.roche.symbols.gmt: The tissue signatures from the BioQC publication used in this study</li> <li>gtex_v6_gini_solid.gmt: The cross-platform cross-species validated tissue signatures produced in this study</li> <li>heterogeneity_results.tsv.gz: Signature scores and heterogeneity calls for each tested signature</li> <li>heterogeneity_fractions.tsv: Fraction of heterogeneous and severely heterogeneous samples per tissue</li> </ul>
GWAS to single cell: Intersecting single-cell transcriptomics and genome wide association studies identifies crucial cell-populations and candidate genes for atherosclerosis.
<p><strong>Background</strong></p> <p>Genome-wide association studies (GWAS) have discovered hundreds of common genetic variants for atherosclerotic disease and cardiovascular risk factors. The translation of susceptibility loci into biological mechanisms and targets for drug discovery remains challenging. Intersecting genetic and gene expression data has led to identification of candidate genes. However, the assayed tissues are often non-diseased and heterogeneous in cell composition confounding the candidate prioritization. We collected single-cell transcriptomics (scRNA-seq) from atherosclerotic plaques and aimed to identify cell-type-specific expression of disease-associated genes. </p> <p> </p> <p><strong>Methods and Results</strong></p> <p>To identify disease-associated candidate genes, we applied gene-based analyses using GWAS summary statistics from 46 atherosclerotic, cardiometabolic, and other traits. Next we intersected these candidates with single-cell transcriptomics (scRNA-seq) to identify those genes that are specifically expressed in individual cell (sub)populations of atherosclerotic plaques. We derive an enrichment score and show that loci that associated with coronary artery disease demonstrated a prominent substrate in plaque smooth muscle cells (<em>SKI</em>, <em>KANK2</em>, <em>SORT1</em>), endothelial cells (<em>SLC44A1</em>, <em>ATP2B1</em>), and macrophages (<em>APOE</em>, <em>HNRNPUL1</em>). Further sub clustering of SMC-subtypes revealed genes in risk loci for coronary calcification specifically enriched in a synthetic cluster of SMCs. To verify the robustness of our approach, we used liver-derived scRNAseq-data and showed enrichment of circulating lipids-associated loci in hepatocytes.</p> <p><br> <strong>Conclusion</strong></p> <p>We confirm known gene-cell pairs relevant for atherosclerotic disease, and discovered novel pairs pointing to new biological mechanisms amenable for therapy. We present an intuitive single-cell transcriptomics driven workflow rooted in human large-scale genetic studies to identify putative candidate genes and affected cells associated with cardiovascular traits.</p> <p> </p>
ZIRFs: zero-inflated random forests for estimating gene regulatory networks from single cell RNA-seq data (assessment of predictive accuracy and VIM stability)
<p>We developed a zero-inflated random forests (ZIRFs) algorithm to produce a metric of connection strength between regulator genes and target genes. This file contains SCENIC results for the aorta and diaphragm tissue data sets from the Tabula Muris Consortium results. SCENIC is a genetic regulatory network analysis published by Aibar et al. (2017). The purpose of the data sets and R source code are described by README files in each directory.</p>
Building Large-Scale Gene-Disease Association Datasets for Biomedical Relation Extraction
<p>This repository contains the GDAb and GDAt datasets. GDAb and GDAt are large-scale, distantly supervised, and manually enhanced datasets for Gene-Disease Association (GDA) extraction. Each dataset consists of three text files, corresponding to train, validation, and test sets, plus an additional JSON file containing the mapping between relation names and IDs. Each record in train, validation, or test files corresponds to a single GDA extracted from a sentence. Records are represented as JSON objects with the following structure:</p> <ul> <li><strong>text:</strong> sentence from which the GDA was extracted.</li> <li><strong>relation:</strong> relation name associated to the given GDA.</li> <li><strong>h: </strong>JSON object representing the gene entity, composed of: <ul> <li><strong>id: </strong>UMLS CUI associated to the gene entity.</li> <li><strong>name:</strong> UMLS preferred name associated to the gene entity.</li> <li><strong>pos: </strong>list consisting of starting position and length of the gene mention within text.</li> </ul> </li> <li><strong>t:</strong> JSON object representing the disease entity, composed of: <ul> <li><strong>id: </strong>UMLS CUI associated to the disease entity.</li> <li><strong>name:</strong> UMLS preferred name associated to the disease entity.</li> <li><strong>pos:</strong> list consisting of starting position and length of the disease mention within text.</li> </ul> </li> </ul> <p>Both datasets contain over 2,500,000 sentences and 500,000 bags.<br> The zip file consists of two folders, GDAb and GDAt, containing the files corresponding to the two datasets, respectively.</p> <p> </p>
Riboswitch-inspired toehold riboregulators for gene regulation in Escherichia coli
<p>This dataset comprises flow cytometry data accompanying a publication on the development of synthetic riboregulators.</p> <p>These riboregulators were inspired by the architecture of naturally occurring riboswitches and toehold-mediated strand displacement. Specifically, we adopt the toehold switch hairpin and inserted regulatory sequences within the loop region of which accessibility can be controlled by toehold-mediated strand displacement. We utilized this design principle to develop toehold translation repressor and toehold transcriptional repressor, which regulate mCherry expression in <em>E. coli </em>in translational and transcriptional levels with certain ON/OFF ratios. Furthermore, we combined these two riboregulators and developed them into a NOR gate switch that can regulate downstream GFP expression in <em>E. coli </em>with different input conditions of trigger RNA. We used flow cytometry to quantify the expression level of the NOR gate switch under different inputs.</p>
EAW FASTQ files for bioinformatic courses (16S rRNA genes, 2018)
<p>Selection of 6 samples, with 6 Forward (R1) and 6 Reverse (R2) files, including primers. R1 and R2 reads are ca. 300 bp long, and were obtained from Illumina MiSeq technologies, at the FEM facility sequencing platform. The files refer to the16S rRNA gene reads obtained from the analyses carried out on the samples collected and filtered (Sterivex<sup>TM</sup> 0.22 µm) in different areas and depths of Lake Garda on September, 2018 (see EAW_2018_FASTQ_16S_description.docx).</p> <p>Sampling and analyses were carried out in the framework of the project Eco-AlpsWater (ASP569), funded by the Interreg Alpine Space program.</p> <p>A bioinformatic protocol for analyzing these data using DADA2 is available in Zenodo:</p> <pre>https://doi.org/10.5281/zenodo.5232772 </pre>
EAW FASTQ files for bioinformatic courses (18S rRNA genes, 2018)
<p>Selection of 6 samples, with 6 Forward (R1) and 6 Reverse (R2) files, including primers. R1 and R2 reads are ca. 300 bp long, and were obtained from Illumina MiSeq technologies, at the FEM facility sequencing platform. The files refer to the18S rRNA gene reads obtained from the analyses carried out on the samples collected and filtered (Sterivex<sup>TM</sup> 0.22 µm) in different areas and depths of Lake Garda on September, 2018 (see EAW_2018_FASTQ_18S_description.docx).</p> <p>Sampling and analyses were carried out in the framework of the project Eco-AlpsWater (ASP569), funded by the Interreg Alpine Space program.</p> <p>A bioinformatic protocol for analyzing these data using DADA2 is available in Zenodo:</p> <pre>https://doi.org/10.5281/zenodo.5233527</pre> <p>The corresponding 16S rRNA gene reads, obtained from the same eDNA extracts, are saved in <a href="https://doi.org/10.5281/zenodo.5215815">https://doi.org/10.5281/zenodo.5215815</a></p>
Gene Regulatory Network inference in long lived C.elegans reveals modular properties that are predictive of novel ageing genes - Supplementary Tables
<p>This repository contains the Supplementary Tables for Suriyalaksh et al. Gene Regulatory Network inference in long lived C.elegans reveals modular properties that are predictive of novel ageing genes.</p> <p>The list of table files can be found in <a href="https://zenodo.org/api/files/431155b3-c2b2-4fff-92bb-e93cad8bf5c2/Supplementary%20table%20guide.pdf">Supplementary Tables guide.pdf</a></p> <p>Tables S1, S2 and S3 corresponding to physical gene-gene interaction data are in a separate repository doi:10.5281/zenodo.4382337</p> <p>Details about some of the Supplementary tables:</p> <p>TableS4_inferred_networks.csv - list of inferred GRNs for specified input combinations (set of input regulators, length of the time sequence, NI tool and prior used).</p> <p><a href="https://zenodo.org/api/files/431155b3-c2b2-4fff-92bb-e93cad8bf5c2/TableS5_consensus_network_member.xlsx">TableS5_consensus_network_member.xlsx</a> - list of groups of topologically similar GRNs (from Table S4)</p> <p>Table S6: edge lists (source,target) for each one of the three consensus networks selected according to the GS validation metrics: middle PFE/AUFE, max AUFE, max PFE.<br> TableS6a_max_AUFE_GRN.txt - max AUFE; largest network - this is the one we used in the main analysis and discussion<br> TableS6b_max_PFE_GRN.xt - max PFE<br> TableS6c_middle_AUFE_PFE_GRN.txt - middle PFE/AUFE</p> <p><a href="https://zenodo.org/api/files/431155b3-c2b2-4fff-92bb-e93cad8bf5c2/TableS7_qRTPCR_ddCt_network_accuracy.csv">TableS7_qRTPCR_ddCt_network_accuracy.csv</a> - gene expression count differences for RNAi knockdown GRN validation experiments. </p> <p>Table S8: Group membership for each one of the nodes in each one of the selected networks according to the SBM that best describes the observed network topology. Each column shows the group membership for each level in a SBM block hierarchy. Our analysis is in the second most coarse-grained level (level 1).</p> <p>TableS8a_max_AUFE_SBM.csv<br> TableS8b_max_PFE_SBM.csv<br> TableS8c_middle_AUFE_PFE_SBM.csv</p> <p><a href="https://zenodo.org/api/files/431155b3-c2b2-4fff-92bb-e93cad8bf5c2/TableS9_glp_gs_datasets.pdf">TableS9_glp_gs_datasets.pdf</a> - list of datasets used for defining functional clusters.</p> <p>TableS14a_glp_l1_vs_fem_l1_lifespan_assay.xlsx - Day13 survival of fem-3(q20)ts vs day 19 survival of glp-1(e2144)ts;rrf-3(pk1426) RNAi from L1</p> <p>TableS14b_glp_l1_vs_glp_l4_lifespan_assay.xlsx - Day 19 survival of glp-1(e2144)ts;rrf-3(pk1426) RNAi from L4 vs day 19 survival of glp-1(e2144)ts;rrf-3(pk1426) RNAi from L1</p> <p>TableS15a_glp1_in_vivo_fluorescence_data.xlsx - in vivo fluorescent reporter data of glp-1(e2144)ts;rrf-3(pk1426)</p> <p>TableS15b_fem3_in_vivo_fluorescence_data.xlsx - in vivo fluorescent reporter data of fem-3(q20)ts</p> <p>TableS17_input_regulators_annotated.csv - list input regulators used as input for Network Inference Tools annotated by source type (2nd column): GenAge, known transcription factors (TF) and gene with high variability in the gene expression time series (HV). The third column lists whether that regulator has an orthologue in human (y) according to WormBase (v 278).</p> <p>TableS20_epistasis_lifespan_data.xlsx - Epistasis lifespan data of glp-1(e2144)ts</p> <p>All the image (TIF) files represent representative images in the following genetic backgrounds (below) that have been treated </p> <p>with empty vector (EV) or RNAi against the gene highlighted in the title of the image. See methods section for details. </p> <p><strong>femliu1: </strong></p> <p><em>fem-3(q20)ts.; dhs-3p::dhs-3::gfp</em></p> <p><strong>femsod3:</strong></p> <p><em>fem-3(q20)ts.; sod-3p::gfp</em></p> <p><strong>glp1lgg1:</strong></p> <p><em>glp-1(e2144); lgg-1p:lgg-1:gfp</em></p>
Chromatin activity identifies differential gene regulation across human ancestries
<p>This repository contains data related to:</p> <p>Chromatin activity identifies differential gene regulation across human ancestries</p> <p>Kade P. Pettie, Maxwell Mumbach, Amanda J. Lea, Julien Ayroles, Howard Y. Chang, Maya Kasowski, Hunter B. Fraser</p> <p> </p>
16S rRNA sequencing gene datasets for CRC data
<p>Used datasets: </p> <table> <thead> <tr> <th scope="col"> <table> <thead> <tr> <th>Dataset</th> <th>16S rRNA Region</th> <th>Control (n)</th> <th>Adenoma (n)</th> <th>CRC (n)</th> <th>Available metadata</th> </tr> </thead> <tbody> <tr> <td><a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4823848/">Baxter</a></td> <td>V4</td> <td>171</td> <td>198</td> <td>120</td> <td>Gender, age, weight, height, BMI, country, race</td> </tr> <tr> <td><a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4221363/">Zackular</a></td> <td>V4</td> <td>30</td> <td>30</td> <td>30</td> <td>Gender, age, weight, height, BMI, country, race, FOBT, medication</td> </tr> <tr> <td><a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4299606/">Zeller</a></td> <td>V4</td> <td>50</td> <td>38</td> <td>41</td> <td>Gender, age, BMI, country, FOBT</td> </tr> <tr> <td><strong>TOTAL</strong></td> <td>V4</td> <td>251</td> <td>266</td> <td>191</td> <td><em>All of the above</em></td> </tr> </tbody> </table> </th> </tr> </thead> <tbody> <tr> <td> </td> </tr> </tbody> </table> <p>Data processing & sharing</p> <p>All datasets were processed using <a href="https://docs.qiime2.org/2021.11/">qiime2</a> pipeline with <a href="https://benjjneb.github.io/dada2/">DADA2</a> for Sequence quality control and feature table construction and <a href="https://www.arb-silva.de/">SILVA</a> database for taxonomic assignment, and then a <em>phyloseq </em>object was constructed.</p> <ul> <li>Abundance table at genus level is in file <em>genus.csv</em> (Sample counts with NO filtering).</li> <li>Clean metadata is in <em>metadata.csv</em> file (Countries: CA - Canada. USA - United States of America. FRA - France.)</li> <li>Phyloseq object is in file <em>physeq.RDS</em> (Saved as an RDS object in R)</li> </ul> <p>More information is <a href="https://hackmd.io/nbsLqCLlSNSRFc5RBX9c5Q?view">here</a>. </p> <p> </p>
Vitamin_B12_related_genes_and_FASTA_sequences
<p>These two files are related to a submitted research paper on the study of a 7-year metagenomic time series carried monthly in a northwestern Mediterranean coastal site (SOLA, Banyuls-sur-Mer, FRANCE) :</p> <p><em>Seasonal succession of different vitamin B12 biosynthesis pathways and producers in coastal marine microbial communities. </em></p> <p>This study focuses on prokaryotes involved in vitamin B12 metabolism (biosynthesis, transport, remodeling) and on the metabolic pathways involved in these processes.</p> <p>The dataset in <strong>.xlsx</strong> format represents the abundance table of all the genes associated with 82 KEGGs involved in vitamin B12 metabolism, and the file in <strong>.fasta</strong> format corresponds to the FASTA sequences associated with these genes.</p>
Dataset for 53 gene families from 16 eukaryotes
<p>NEXUS file representing Guigo et al.'s (1996) dataset for 53 gene families from 16 eukaryotes, used in a number of gene tree reconciliation studies. Original data from Guigo et al., this NEXUS encoding by Roderic Page.</p>
Vertebrate gene family trees
<p>This is a dataset of phylogenies for 118 vertebrate gene families, used in two papers published by James Cotton and Roderic Page in 2002. The folder named "genes.zip" contains FASTA and NEXUS sequence files, and Newick-format tree files for each gene family. There is a PDF file "suppl.pdf" listing the names of each family. The file "final_dataset.gml" contains a graph of the taxonomic overlap in the gene trees. All 118 gene trees have been combined into the file "final_dataset.gtr" which is a NEXUS file with custom blocks recognised by GeneTree.</p> <table> <tbody> <tr> <th>HOVERGEN FAMILY CODE</th> <th>GENE FAMILY NAME</th> </tr> <tr> <td>FAM000030A</td> <td>wnt 5</td> </tr> <tr> <td>FAM000030B</td> <td>wnt 7</td> </tr> <tr> <td>FAM000030C</td> <td>wnt 11</td> </tr> <tr> <td>FAM000030D</td> <td>wnt/int 1</td> </tr> <tr> <td>FAM000030E</td> <td>wnt 4</td> </tr> <tr> <td>FAM000030F</td> <td>wnt 10/12</td> </tr> <tr> <td>FAM000030G</td> <td>wnt 3</td> </tr> <tr> <td>FAM000030H</td> <td>wnt 2</td> </tr> <tr> <td>FAM000030I</td> <td>wnt 8</td> </tr> <tr> <td>FAM000105</td> <td>rhodopsin</td> </tr> <tr> <td>FAM000214</td> <td>beta-B globin</td> </tr> <tr> <td>FAM000033</td> <td>protamine 1</td> </tr> <tr> <td>FAM000370</td> <td>PrP a prion-protein</td> </tr> <tr> <td>FAM000014</td> <td>growth hormone</td> </tr> <tr> <td>FAM000556</td> <td>Rag-1 recombination activation gene</td> </tr> <tr> <td>FAM001493</td> <td>c-mos proto-oncogene</td> </tr> <tr> <td>FAM001462</td> <td>tyrosine kinase / yes / fyn / src / lck</td> </tr> <tr> <td>FAM001041</td> <td>metallothionein</td> </tr> <tr> <td>FAM000364</td> <td>Ldh-2 lactate dehydrogenase-B (EC</td> </tr> <tr> <td>FAM000215</td> <td>alpha globin</td> </tr> <tr> <td>FAM000016</td> <td>placental lactogen - prolactin</td> </tr> <tr> <td>FAM000008</td> <td>insulin</td> </tr> <tr> <td>FAM000824</td> <td>phosphoglycerate kinase</td> </tr> <tr> <td>FAM000550</td> <td>neurotrophin-4 (NT-4)</td> </tr> <tr> <td>FAM001478</td> <td>tyrosine kinase receptor, c-fms oncogene</td> </tr> <tr> <td>FAM000173A</td> <td>guanine nucleotide-binding protein</td> </tr> <tr> <td>FAM000173B</td> <td>transducin alpha</td> </tr> <tr> <td>FAM000502</td> <td>cytochrome P-450 aromatase</td> </tr> <tr> <td>FAM000192</td> <td>alpha-fetoprotein / serum albumin</td> </tr> <tr> <td>FAM000627</td> <td>neurone-specific enolase</td> </tr> <tr> <td>FAM001232</td> <td>preprotrypsin (ta)</td> </tr> <tr> <td>FAM000664</td> <td>complement component 3 (C3)</td> </tr> <tr> <td>FAM000175</td> <td>Ras</td> </tr> <tr> <td>FAM000248</td> <td>alpha B-crystallin</td> </tr> <tr> <td>FAM000006</td> <td>insulin-like growth factor II</td> </tr> <tr> <td>FAM000639</td> <td>transthyretin (prealbumin)</td> </tr> <tr> <td>FAM001303</td> <td>butylcholinesterase (BCHE)</td> </tr> <tr> <td>FAM000242</td> <td>connexin / gap junction protein</td> </tr> <tr> <td>FAM000058</td> <td>dopamine D1 receptor</td> </tr> <tr> <td>FAM000055</td> <td>beta-3-adrenergic receptor .</td> </tr> <tr> <td>FAM000131</td> <td>ATPase (Na+K+, H+K+)</td> </tr> <tr> <td>FAM003983</td> <td>preprogastrin</td> </tr> <tr> <td>FAM000353</td> <td>vasopressin</td> </tr> <tr> <td>FAM000152</td> <td>acetylcholine receptor</td> </tr> <tr> <td>FAM001327</td> <td>peripherin, desmin, vimentin, GFAP</td> </tr> <tr> <td>FAM003199</td> <td>(C57BL/6J)</td> </tr> <tr> <td>FAM002881</td> <td>cytochrome P-450, 17a-hydroxylase (CYP17)</td> </tr> <tr> <td>FAM002789</td> <td>(MUAHRB-1) Ah-receptor (Ah)</td> </tr> <tr> <td>FAM001461</td> <td>tropomyosin</td> </tr> <tr> <td>FAM000385</td> <td>Y3 peptide supply factor</td> </tr> <tr> <td>FAM000378</td> <td>pancreatic polypeptide, neuropeptide Y</td> </tr> <tr> <td>FAM001329</td> <td>cytokeratin</td> </tr> <tr> <td>FAM001619</td> <td>amelogenin (enamel-specific protein)</td> </tr> <tr> <td>FAM000286</td> <td>glucagon</td> </tr> <tr> <td>FAM000330</td> <td>tissue inhibitor of</td> </tr> <tr> <td>FAM000475</td> <td>lipophilin</td> </tr> <tr> <td>FAM001060</td> <td>ornithine carbamoyltransferase</td> </tr> <tr> <td>FAM001328</td> <td>neurofilament</td> </tr> <tr> <td>FAM001370</td> <td>peroxisome proliferator</td> </tr> <tr> <td>FAM000371</td> <td>somatostatin</td> </tr> <tr> <td>FAM001607</td> <td>Wilms tumor assocated protein (WT1)</td> </tr> <tr> <td>FAM000495</td> <td>aldolase A, B, C</td> </tr> <tr> <td>FAM001365</td> <td>liver receptor homologous protein (LRH-1)</td> </tr> <tr> <td>FAM000672</td> <td>ribosomal protein S4,</td> </tr> <tr> <td>FAM000135</td> <td>Na, K-ATPase beta-1 subunit</td> </tr> <tr> <td>FAM000271</td> <td>enkephalin : 1 2.</td> </tr> <tr> <td>FAM000799</td> <td>RING10</td> </tr> <tr> <td>FAM001664</td> <td>glutamate decarboxylase</td> </tr> <tr> <td>FAM000617</td> <td>creatine kinase</td> </tr> <tr> <td>FAM001337</td> <td>amyloid beta protein precursor</td> </tr> <tr> <td>FAM000274</td> <td>basic fibroblast growth factor (bFGF)</td> </tr> <tr> <td>FAM001108</td> <td>Sl-d mutant allele kit ligand (KL)</td> </tr> <tr> <td>FAM000504</td> <td>anion exchange protein 3</td> </tr> <tr> <td>FAM001239</td> <td>prothrombin</td> </tr> <tr> <td>FAM001360</td> <td>high mobility group proteins HMG1 and HMG2</td> </tr> <tr> <td>FAM000534</td> <td>glutamine synthetase</td> </tr> <tr> <td>FAM001053</td> <td>nucleoside diphosphate kinase</td> </tr> <tr> <td>FAM001390</td> <td>low density lipoprotein receptor LDLR</td> </tr> <tr> <td>FAM001464</td> <td>Cek6 receptor tyrosine kinase</td> </tr> <tr> <td>FAM000801</td> <td>manganese-containing superoxide</td> </tr> <tr> <td>FAM003946</td> <td>Six2 / Six1 mRNA</td> </tr> <tr> <td>FAM001606</td> <td>Ikaros binding protein (Ikaros)</td> </tr> <tr> <td>FAM000350</td> <td>SPARC protein</td> </tr> <tr> <td>FAM001339</td> <td>calcium-binding protein</td> </tr> <tr> <td>FAM000553</td> <td>pyruvate kinase</td> </tr> <tr> <td>FAM000604</td> <td>t complex polypeptide 1 (Tcp-1-a)</td> </tr> <tr> <td>FAM002988</td> <td>mSlo</td> </tr> <tr> <td>FAM001266</td> <td>transcription factor / hepatocyte nuclear factor</td> </tr> <tr> <td>FAM000453</td> <td>terminal deoxynucleotidyltransferase</td> </tr> <tr> <td>FAM001642</td> <td>transformation associated protein p53</td> </tr> <tr> <td>FAM000300A</td> <td>glucose-regulated protein 78 / HSP70 PART A</td> </tr> <tr> <td>FAM000300B</td> <td>glucose-regulated protein 78 / HSP70 PART B</td> </tr> <tr> <td>FAM000492A</td> <td>alpha actin etc</td> </tr> <tr> <td>FAM000492B</td> <td>beta actin etc</td> </tr> <tr> <td>FAM000649</td> <td>Myelin Basic Protein</td> </tr> <tr> <td>FAM000843</td> <td>ribosomal protein S4</td> </tr> <tr> <td>FAM000170</td> <td>atrial natriuretic protein</td> </tr> <tr> <td>FAM000871A</td> <td>tyrosinase</td> </tr> <tr> <td>FAM000871B</td> <td>tyrosinase related protein 1</td> </tr> <tr> <td>FAM001605</td> <td>ZFX put. transcription activator</td> </tr> <tr> <td>FAM001479</td> <td>fibroblast growth factor</td> </tr> <tr> <td>FAM000160</td> <td>pro-opiomelanocortin (POMC)</td> </tr> <tr> <td>FAM001644</td> <td>fibrinogen alpha subunit</td> </tr> <tr> <td>FAM000564</td> <td>SNAP-25</td> </tr> <tr> <td>FAM000266</td> <td>nitric oxide synthase</td> </tr> <tr> <td>FAM004159</td> <td>chondroitin-6 sulfotransferase</td> </tr> <tr> <td>FAM000800</td> <td>Lmp-2 (LMPq) proteasome subunit</td> </tr> <tr> <td>FAM001595</td> <td>factor B</td> </tr> <tr> <td>FAM001134</td> <td>zona pellucida (ZP)</td> </tr> <tr> <td>FAM001632</td> <td>stromelysin-3</td> </tr> <tr> <td>FAM001366</td> <td>steroid receptor (TR2-9)</td> </tr> <tr> <td>FAM000526</td> <td>c-ski protein</td> </tr> <tr> <td>FAM002463</td> <td>thymosin beta 4 peptide</td> </tr> <tr> <td>FAM000904</td> <td>sequence-specific DNA-binding protein (AP-2)</td> </tr> <tr> <td>FAM001465</td> <td>T-cell specific tyrosine kinase (ltk)</td> </tr> <tr> <td>FAM000567</td> <td>triosephosphate isomerase</td> </tr> <tr> <td>FAM006113</td> <td>DNA-dependent RNA polymerase III, large subunit</td> </tr> <tr> <td>FAM001733</td> <td>DNA-dependent RNA polymerase II</td> </tr> </tbody> </table>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.