Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
368
datasets available to search
ShareScore release 0.7.1
Dataset results
368 results for “Consensus”
Consensus-seeking and conflict-resolving: an fMRI study on college couples’ shopping interaction
Open the record for dataset details and reuse information.
Global consensus map of human transcription factor footprints
<p>Vierstra, J. <em>et al.</em> <strong>Global reference mapping of human transcription factor footprints.</strong> <em>Nature</em><strong> </strong>583, 729–736 (2020). <a href="https://doi.org/10.1038/s41586-020-2528-x">https://doi.org/10.1038/s41586-020-2528-x</a></p> <p>Preprint @ bioRxiv: <a href="https://doi.org/10.1101/2020.01.31.927798">https://doi.org/10.1101/2020.01.31.927798</a></p> <p><strong>Contact:</strong> Jeff Vierstra (<a href="mailto:jvierstra@altius.org?subject=Consensus%20DNase%20I%20footprints">jvierstra@altius.org</a>)</p> <p>Genomic DNase I footprinting enables quantitative, nucleotide-resolution delineation of sites of transcription factor occupancy within native chromatin. We combined sampling of >67 billion uniquely mapping DNase I cleavages from >240 human cell types and states to index, with unprecedented accuracy and resolution, human genomic footprints and thereby the sequence elements that encode transcription factor recognition sites.</p> <p>Please see <a href="http://vierstra.org/resources/dgf">http://vierstra.org/resources/dgf </a>for additional information and a complete set of raw DNase I data for individual datasets. Additionally, raw data can also be accessed via the ENCODE data portal (<a href="http://encodeproject.org">http://encodeproject.org</a>) using the dataset accessions found in Supplementary Table 1.</p> <p>Code for footprint analysis and tutorials on how to access and manipulate digital genomic footprint data can be found at <a href="https://footprint-tools.readthedocs.io/en/latest/">https://footprint-tools.readthedocs.io/en/latest/</a>.</p> <p>All files herein correspond to human genome build version GRCh38 (UCSC hg38).</p> <p><strong>Dataset contents:</strong></p> <ul> <li><strong>Biosample metadata</strong> – Supplementary_Table_1.xlsx</li> <li><strong>Motif clustering metadata </strong>– Supplementary_Table_2.xlsx</li> <li><strong>ChIP-seq validation metadata </strong>–<strong> </strong>Supplementary_Table_3.xlsx</li> <li><strong>Consensus footprint coordinates and assigned motif archetypes</strong><br> TSV file (BED-format) with consensus footprint (posterior probability>0.99) coordinates and overlaps with matches to motif model clusters. The legend file contains column definitions in detail. <ul> <li>consensus_footprints_and_motifs_hg38.bed.gz</li> <li>consensus_footprints_and_motifs_legend.txt</li> </ul> </li> <li><strong>Motif archetype matches overlapping consensus footprints</strong><br> TSV file (BED-format) containing the coordinates for clustered motif model matches that overlap consensus footprints <ul> <li>collapsed_motifs_overlaping_consensus_footprints.bed.gz</li> <li>collapsed_motifs_overlaping_consensus_footprints_legend.txt</li> </ul> </li> <li><strong>Footprint occupancy matrix of consensus footprints</strong><br> Rows are same order as the consensus footprint file and columns are same order as in the metadata files. <ul> <li>consensus_index_matrix_full_hg38.txt.gz (Values are –log(1-posterior))</li> <li>consensus_index_matrix_binary_hg38.txt.gz (binary occupancy matrix, where footprints with posterior footprint probability >0.99 are considered occupied)</li> </ul> </li> <li><strong>Single nucleotide variants tested for allelic imbalance </strong><br> The legend file contains column definitions in detail. <ul> <li>genotypes.vcf.gz - Genotyping and allelic read depth for each biosample (see header for more information)</li> <li>tested_snvs_padj.bed.gz - SNVs tested for imbalance (TSV, BED-format)</li> <li>tested_snvs_padj_legend.txt</li> </ul> </li> </ul>
Dataset - Terminology of e-Oral Health: Consensus Report of the IADR's e-Oral Health Network Terminology Task Force.
<p>README<br>====================<br>This repository contains the data and documentation for a research project. It includes the dataset,<br>which is provided in CSV format and the original PDF with the survey answers.</p> <p>Research Information<br>====================<br>Terminology of e-Oral Health: Consensus Report of the IADR’s e-Oral Health Network Terminology<br>Task Force. Authors reported multiple definitions of e-oral health and related terms, and used several definitions<br>interchangeably, like mhealth, teledentistry, teleoral medicine and telehealth. The International<br>Association of Dental Research e-Oral Health Network (e-OHN) aimed to establish a consensus on<br>terminology related to digital technologies used in oral healthcare.</p> <p>This dataset contains data from a survey about digital oral health. The survey asked participants to provide their definition of various terms related to digital oral health, as well as their agreement with the provided definitions. The dataset also includes three figures that the participants were asked to review.</p> <p>The purpose of this dataset is to collect data on the public's understanding of digital oral health terms and to identify areas where there may be confusion or misinterpretation. The data from this dataset could be used to develop educational materials or to improve the way that digital oral health information is communicated to the public.</p> <p>Additional notes<br>====================<br>The data is not currently cleaned or preprocessed.</p> <p>Dataset<br>====================<br>The dataset file, named "dataset.csv," is in this repository. It contains the raw anonymized data<br>collected from the participants in a structured format. Each row represents a respondent, and the<br>columns correspond to different variables.</p> <p>Codebook<br>====================<br>The codebook file, named "codebook.pdf," is also included in this repository. It provides a<br>comprehensive description of the variables present in the dataset. The codebook outlines each<br>variable's meaning, type, and possible values, allowing users to understand and analyze the data<br>effectively.</p> <p>Metadata<br>====================<br>No metadata is provided</p> <p>Files<br>====================<br>01_readme.txt this readme file<br>02_codebook.pdf The codebook of the dataset<br>03_dataset.csv The dataset in csv format<br>04_e-OHN Delphi (2023-02-03).pdf The output from the survey</p> <p>Usage<br>====================<br>To work with the dataset, you can download the "dataset.csv" file and import it into your preferred<br>software or programming language for analysis. The codebook provides valuable information about<br>the variables, allowing you to understand the data structure and make informed decisions during your<br>analysis.<br>Please note that while every effort has been made to ensure the accuracy and quality of the data, it is<br>important to review the codebook and understand the context of the research before concluding the<br>dataset.</p> <p>License<br>====================<br>The data and documentation in this repository are provided under the CC BY-SA.<br>This license enables reusers to distribute, remix, adapt, and build upon the material in any medium or format, so long as attribution is given to the creator. The license allows for commercial use. If you remix, adapt, or build upon the material, you must license the modified material under identical terms. CC BY-SA includes the following elements:</p> <p> BY: credit must be given to the creator.<br> SA: Adaptations must be shared under the same terms.<br> <br>Please refer to the license file for further details on how the data can be used and shared.</p> <p>Contact Information<br>====================<br>For any questions, clarifications, or inquiries related to the dataset or research project, please contact<br>Assoc Prof Dr Sergio Uribe, sergio.uribe@rsu.lv</p>
A consensus compound/bioactivity dataset for data-driven drug design and chemogenomics
<p>This is the <strong>updated version</strong> of the dataset from <strong>10.5281/zenodo.6320761</strong></p> <p><strong>Information</strong></p> <p>The diverse publicly available compound/bioactivity databases constitute a key resource for data-driven applications in chemogenomics and drug design. Analysis of their coverage of compound entries and biological targets revealed considerable differences, however, suggesting benefit of a consensus dataset. Therefore, we have combined and curated information from five esteemed databases (ChEMBL, PubChem, BindingDB, IUPHAR/BPS and Probes&Drugs) to assemble a consensus compound/bioactivity dataset comprising <a href="tel:1144648">1144648</a> compounds with 10915362 bioactivities on 5613 targets (including defined macromolecular targets as well as cell-lines and phenotypic readouts). It also provides simplified information on assay types underlying the bioactivity data and on bioactivity confidence by comparing data from different sources. We have unified the source databases, brought them into a common format and combined them, enabling an ease for generic uses in multiple applications such as chemogenomics and data-driven drug design.</p> <p>The consensus dataset provides increased target coverage and contains a higher number of molecules compared to the source databases which is also evident from a larger number of scaffolds. These features render the consensus dataset a valuable tool for machine learning and other data-driven applications in (de novo) drug design and bioactivity prediction. The increased chemical and bioactivity coverage of the consensus dataset may improve robustness of such models compared to the single source databases. In addition, semi-automated structure and bioactivity annotation checks with flags for divergent data from different sources may help data selection and further accurate curation.</p> <p>This dataset belongs to the publication: <a href="https://doi.org/10.3390/molecules27082513">https://doi.org/10.3390/molecules27082513</a><br> </p> <p><strong>Structure and content of the dataset</strong></p> <table align="left"> <caption><strong>Dataset structure</strong></caption> <thead> <tr> <th scope="col"> <p>ChEMBL</p> <p>ID</p> </th> <th scope="col"> <p>PubChem</p> <p>ID</p> </th> <th scope="col"> <p>IUPHAR</p> <p>ID</p> </th> <th scope="col">Target</th> <th scope="col"> <p>Activity</p> <p>type</p> </th> <th scope="col">Assay type</th> <th scope="col">Unit</th> <th scope="col">Mean C (0)</th> <th scope="col">...</th> <th scope="col">Mean PC (0)</th> <th scope="col">...</th> <th scope="col">Mean B (0)</th> <th scope="col">...</th> <th scope="col">Mean I (0)</th> <th scope="col">...</th> <th scope="col">Mean PD (0)</th> <th scope="col">...</th> <th scope="col">Activity check annotation</th> <th scope="col">Ligand names</th> <th scope="col">Canonical SMILES C</th> <th scope="col">...</th> <th scope="col">Structure check (Tanimoto)</th> <th scope="col">Source</th> </tr> </thead> <tbody> <tr> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> <td> </td> </tr> </tbody> </table> <p>The dataset was created using the Konstanz Information Miner (KNIME) (https://www.knime.com/) and was exported as a CSV-file and a compressed CSV-file.</p> <p>Except for the canonical SMILES columns, all columns are filled with the datatype ‘string’. The datatype for the canonical SMILES columns is the smiles-format. We recommend the <strong>File Reader</strong> node for using the dataset in KNIME. With the help of this node the data types of the columns can be adjusted exactly. In addition, only this node can read the compressed format.</p> <p>Column content:</p> <ul> <li>ChEMBL ID, PubChem ID, IUPHAR ID: chemical identifier of the databases</li> <li>Target: biological target of the molecule expressed as the HGNC gene symbol</li> <li>Activity type: for example, pIC<sub>50</sub></li> <li>Assay type: Simplification/Classification of the assay into cell-free, cellular, functional and unspecified</li> <li>Unit: unit of bioactivity measurement</li> <li>Mean columns of the databases: mean of bioactivity values or activity comments denoted with the frequency of their occurrence in the database, e.g. Mean C = 7.5 *(15) -> the value for this compound-target pair occurs 15 times in ChEMBL database</li> <li>Activity check annotation: a bioactivity check was performed by comparing values from the different sources and adding an activity check annotation to provide automated activity validation for additional confidence <ul> <li>no comment: bioactivity values are within one log unit;</li> <li>check activity data: bioactivity values are not within one log unit;</li> <li>only one data point: only one value was available, no comparison and no range calculated;</li> <li>no activity value: no precise numeric activity value was available;</li> <li>no log-value could be calculated: no negative decadic logarithm could be calculated, e.g., because the reported unit was not a compound concentration</li> </ul> </li> <li>Ligand names: all unique names contained in the five source databases are listed</li> <li>Canonical SMILES columns: Molecular structure of the compound from each database</li> <li>Structure check (Tanimoto): To denote matching or differing compound structures in different source databases <ul> <li>match: molecule structures are the same between different sources;</li> <li>no match: the structures differ. We calculated the Jaccard-Tanimoto similarity coefficient from Morgan Fingerprints to reveal true differences between sources and reported the minimum value;</li> <li>1 structure: no structure comparison is possible, because there was only one structure available;</li> <li>no structure: no structure comparison is possible, because there was no structure available.</li> </ul> </li> <li>Source: From which databases the data come from</li> </ul> <p> </p> <p> </p>
Alphafold and ColabFold models of E. coli and consensus Bcs complexes
<p>ColabFold and AlphaFold 3 models used for structure modeling and interpretation in Anso et al. 'Structural basis for synthase activation and cellulose modification in the <em>E. coli</em> Type II Bcs secretion system'. </p>
Consensus models to predict oral rat acute toxicity and validation on a dataset coming from the industrial context
<p>We report predictive models of acute oral systemic toxicity representing a follow-up of our previous work in the framework of the NICEATM project. It includes the update of original models through the addition of new data and an external validation of the models using a dataset relevant for the chemical industry context. A regression model for LD50 and classification model for toxicity classes according to the Global Harmonized System categories were prepared. ISIDA descriptors were used to encode molecular structures. Machine learning algorithms included Support Vector Machine (SVM), Random Forest (RF) and Naïve Bayesian. Selected individual models were combined in consensus.</p> <p>The different datasets were compared using the Generative Topographic Mapping approach. It appeared that the NICEATM datasets were lacking some relevant chemotypes for chemical industry. The new models trained on enlarged data sets have applicability domain (AD) sufficiently large to accommodate industrial compounds. The fraction of compounds inside the models’ AD increased from 58 % (NICEATM model) to 94 % (new model). Yet, the increase of training sets only slightly improved of the models’ prediction performance: RMSE values decreased from 0.56 to 0.47 and balanced accuracies increased from 0.69 to 0.71 for NICEATM and new models, respectively.</p>
Suplementary material: Towards uncoding hepatotoxicity of approved drugs through navigation of multiverse and consensus chemical spaces
<p>Supplementary material: "Towards uncoding hepatotoxicity of approved drugs through navigation of multiverse and consensus chemical spaces"</p>
2018_earthenv_Consensus_Landuse _percentage
<p>This dataset has been downloaded from earthen.org and then windowed to E4warning study area. This dataset is available in TIF format. </p>
Fig. 1. Bayesian majority rule consensus tree reconstructed for 90 in Phylogenetic analysis and systematic position of two new species of the ant genus Crematogaster (Hymenoptera, Formicidae) from Southeast Asia
Fig. 1. Bayesian majority rule consensus tree reconstructed for 90 taxa using five genes (ArgK, CAD, LWRh, Top1, Wg) in a MrBayes analysis. Above node numbers indicate posterior probability. Data were partitioned by PartitionFinder v.1.1.1 and analyzed using a best fit model for each gene and codon position, with 10 million generations and a burn-in of 25 %. Area enclosed by dashed lines is enlarged on Fig. 2.
Consensus QSAR models estimating acute aquatic toxicity for three trophic levels organisms: Algae, Daphnia and Fish
<p>We report new consensus models estimating acute toxicity for algae, daphnia and fish endpoints. We assembled a large collection of 3680 public unique compounds annotated by, at least, one experimental value for the given endpoint. Support Vector Machine models were internally and externally validated following the OECD principles. Reasonable predictive performances were achieved (RMSE<sub>ext</sub> = 0.56 – 0.78) which are in line with those of state-of-the-art models. The known structural alerts are compared with analysis of the atomic contributions to these models obtained using the ISIDA/<em>ColorAtom</em> utility. A benchmarking against existing tools has been carried out on a set of compounds considered more representative and relevant for the chemical space of the current chemical industry. Our model scored one of the best accuracies and data coverage.</p> <p>Nevertheless, industrial data performances were noticeably lower than those on public data, indicating that existing models fail to meet the industrial needs. Thus, final models were updated with the inclusion of new industrial compounds, extending applicability domain and relevance for application in an industrial context. Generate models and collected public data are made freely available.</p> <p><strong>Available fields in the SDF file:</strong></p> <ul> <li>SMILES_Canonical: canonical SMILES code</li> <li>DB: source of the data, "Litterature set" means that the data is originated from an article (see the companion article of the dataset for details).</li> <li>endpoint: organism for which endpoint is available</li> <li>CASRN: CAS registration number</li> <li>98-81-7</li> <li>pEC50 - DAPHNIA: Daphnia, mortality, which is evaluated by the immobilization of the invertebrate is recorded at 48 hours and expressed as the log median effective concentration (pEC50)</li> <li>mg/L - DAPHNIA: Daphnia, mortality, which is evaluated by the immobilization of the invertebrate is recorded at 48 hours and expressed as the median effective concentration (EC50)</li> <li>pLC50 - FISH: Fish, the log median lethal concentration measured at 96 hours is considered (pLC50)</li> <li>mg/L - FISH: Fish, the log median lethal concentration measured at 96 hours is considered (LC50)</li> <li>pEC50 - ALGA: Algae, the purpose is to determine the substance’s growth inhibition effect, expressed as the log median effective concentration (pEC50) measured at 72 hours</li> <li>mg/L - ALGA: Algae, the purpose is to determine the substance’s growth inhibition effect, expressed as the median effective concentration (EC50) measured at 72 hours</li> </ul>
CyclomicsSeq: Accurate detection of circulating tumor DNA using nanopore consensus sequencing
<p>CyclomicsSeq is a protocol designed to produce and sequence long DNA concatemers with a linear repetition to acquire high accuracy consensus reads. In this dataset, we used CyclomicsSeq for sequencing TP53 in cell-free DNA of healthy individuals and of head and neck cancer patients and for sequencing synthetic TP53 DNA sequences that mimic the length of cell-free DNA. This dataset contains data (mainly base calls of the backbone and the insert) of 32 nanopore sequencing runs. <br> </p>
Figure 3. Consensus tree for the cytochrome b in Four New Bat Species (Rhinolophus hildebrandtii Complex) Reflect Plio-Pleistocene Divergence of Dwarfs and Giants across an Afromontane Archipelago
Figure 3. Consensus tree for the cytochrome b dataset for representative genotyped specimens of the Rhinolophus hildebrandtii complex. The topology represents the consensus topology from a 20 million MCMC run implemented in BEAST. Estimates of divergence times (million years ago; Mya) are indicated adjacent to nodes or above branches and grey bars indicate 95% HPD values. The split between the Hipposideridae and Rhinolophidae was used as the calibration point. Taxa names include museum/field numbers which correspond to Appendix S1 or GenBank accession numbers and abbreviations are: RcfH - R. cf. hildebrandtiiı RD - R. darlingiı RE - R. eloquensı RF - R. fumigatusı RH - R. hildebrandtii s.l.ı RL - R. landeri and RR - R. ruwenzorii. Localitiesı where availableı are providedı abbreviations include SA - South Africaı MZ - Mozambiqueı and ZW - Zimbabweı and the numbers in parentheses correspond with place names in Table S1 and Fig. 2 for Clade 1 and 2 individuals. doi:10.1371/journal.pone.0041744.g003
FIG. 5. — Strict consensus tree from eight most parsimonious trees recovered for Molossus E. Geoffroy, 1805 in Diversity, morphological phylogeny, and distribution of bats of the genus Molossus E. Geoffroy, 1805 (Chiroptera, Molossidae) in Brazil
FIG. 5. — Strict consensus tree from eight most parsimonious trees recovered for Molossus E. Geoffroy, 1805. Numbers above the branches indicate Bootstrap values and bottom numbers indicate Bremer support values.
Consensus between pipelines in structural brain networks
<p>Weighted brain network matrices for all subjects reconstructed using each pipeline. Filename letters denote the combination of pipeline stages used to reconstruct the matrix, whereas the number represents the subject. Pipeline stages are described in filename_key.txt. The brain region names and parcellation labels corresponding to the network node row/column are in X_region_names.txt and X_parcellation_labels.txt, where X is the atlas abbreviation described in filename_key.txt.<br /> </p>
Fig. 2. Bayesian consensus tree generated from partial 28S in Relationships Of The Heteronchocleidids (Heteronchocleidus, Eutrianchoratus And Trianchoratus) As Inferred From Ribosomal Dna Nucleotide Sequence Data
Fig. 2. Bayesian consensus tree generated from partial 28S rDNA sequences (D1 domain) with Diplectanum spp. and Gyrodactylus spp. as outgroups. Values shown at each node refer to Bayesian (BI) posterior probabilities/maximum likelihood (ML) percentages of the bootstrap values with 100 replicates. Bootstrap values lower than 50 are given as dashes (-).
Fig. 51. Consensus tree for 13 in Polybia, Paraphyly, and Polistine Phylogeny
Fig. 51. Consensus tree for 13 cladograms resulting from simultaneous analysis of the combined character data in tables 1–3.
Fig. 50. Consensus tree for 79 in Polybia, Paraphyly, and Polistine Phylogeny
Fig. 50. Consensus tree for 79 strictly supported cladograms resulting from analysis of the larval character matrix in table 2, excluding the three terminals that have all missing values.
FIGURE 3 Majority-rule consensus tree from a in Evolutionary history of species of the fireFly subgenus Hotaria (Coleoptera, Lampyridae, Luciolinae, Luciola) inferred from DNA barcoding data
FIGURE 3 Majority-rule consensus tree from a Bayesian analysis (BI) of 128 samples of 14 morphospecies based on COI barcode sequences. The numbers at each node indicate Downloadedposteriorfrom Brill. probabilities com 12. /12/ Weakly 2023 03 sup-:05:57PM ported nodes (posterior via probabilityOpen below Access. 0.95) Thisareis an shownopenin red. access article distributed under the terms of the CC-BY 4.0 License. https://creativecommons.org/licenses/by/4.0/
Figure 23. Strict consensus cladograms constructed using a in Identification of fossil worm tubes from Phanerozoic hydrothermal vents and cold seeps
Figure 23. Strict consensus cladograms constructed using a total of 64 modern and fossil annelid taxa and 48 mostly morphological tube characters. Analyses were performed using implied character weighting, with the concavity constant set as default (k = 3; A), and also set to downweight homoplastic characters less (k = 4; B). Numbers on nodes represent groups present/contradicted support values. Modern taxa are coloured according to taxonomic groups; fossil taxa are in grey. A, consensus of 271 most parsimonious trees (best score = 15.387, consistency index = 0.195, retention index = 0.264); B, consensus of 60 most parsimonious trees (best score = 13.568, consistency index = 0.232, retention index = 0.569). Symbols/colours indicate taxonomic affinities.
Consensus molecular environment of schizophrenia risk genes in co-expression networks shifting across age and brain regions
<p>This is the online data repository accompanying the following manuscript:<br><strong>Consensus molecular environment of schizophrenia risk genes in coexpression networks shifting across age and brain regions</strong></p> <p><em>Giulio Pergola<sup>1,2,3,*</sup>, Madhur Parihar<sup>1</sup>, Leonardo Sportelli<sup>1,2</sup>, Rahul Bharadwaj<sup>1</sup>, Christopher Borcuk<sup>2</sup>, Eugenia Radulescu<sup>1</sup>, Loredana Bellantuono<sup>2,5</sup>, Giuseppe Blasi<sup>2,4</sup>, Qiang Chen<sup>1</sup>, Joel E. Kleinman<sup>1,3</sup>, Yanhong Wang<sup>1</sup>, Srinidhi Rao Sripathy<sup>1</sup>, Brady J. Maher<sup>1,3,7</sup>, Alfonso Monaco<sup>5,9</sup>, Fabiana Rossi<sup>1,2</sup>, Joo Heon Shin<sup>1</sup>, Thomas M. Hyde<sup>1,3,6</sup>, Alessandro Bertolino<sup>2,4,*</sup>, Daniel R. Weinberger<sup>1,7,8,*</sup></em></p> <p> </p> <p><strong>Affiliations:</strong></p> <p><em>1)Lieber Institute for Brain Development, Johns Hopkins Medical Campus, Baltimore, MD (USA)<br>2)Group of Psychiatric Neuroscience, Department of Translational Biomedicine and Neuroscience, University of Bari Aldo Moro, Bari, Italy<br>3)Department of Psychiatry and Behavioral Sciences, Johns Hopkins University School of Medicine, Baltimore, Maryland<br>4)Azienda Ospedaliero-Universitaria Consorziale Policlinico, Bari, Italy<br>5)Istituto Nazionale di Fisica Nucleare (INFN), Bari, Italy<br>6)Department of Neurology, Johns Hopkins University School of Medicine, Baltimore, Maryland<br>7)Department of Neuroscience, Johns Hopkins University School of Medicine, Baltimore, Maryland<br>8)Department of Genetic Medicine, Johns Hopkins University School of Medicine, Baltimore, Maryland<br>9)Dipartimento Interateneo di fisica, Università degli Studi di Bari Aldo Moro, Bari, Italy</em></p> <p> </p> <p><strong>Abstract:</strong></p> <p><em>Schizophrenia is a neurodevelopmental brain disorder whose genetic risk is associated with shifting clinical phenomena across the life span. We investigated the convergence of putative schizophrenia risk genes in brain coexpression networks in postmortem human prefrontal cortex (DLPFC), hippocampus, caudate nucleus, and dentate gyrus granule cells, parsed by specific age periods (total N = 833). The results support an early prefrontal involvement in the biology underlying schizophrenia and reveal a dynamic interplay of regions in which age parsing explains more variance in schizophrenia risk compared to lumping all age periods together. Across multiple data sources and publications, we identify 28 genes that are the most consistently found partners in modules enriched for schizophrenia risk genes in DLPFC; twenty-three are previously unidentified associations with schizophrenia. In iPSC-derived neurons, the relationship of these genes with schizophrenia risk genes is maintained. The genetic architecture of schizophrenia is embedded in shifting coexpression patterns across brain regions and time, potentially underwriting its shifting clinical presentation.</em></p> <p> </p> <p><strong>Citation:</strong> <em>Giulio Pergola et al. ,Consensus molecular environment of schizophrenia risk genes in coexpression networks shifting across age and brain regions.Sci. Adv.9, eade2812(2023).DOI:10.1126/sciadv.ade2812</em></p> <p> </p> <p><strong>Data Files:<br>DLPFC hit.genes_kb_200__online.version.zip: </strong><br>Interactive Sankey plot for age-parsed DLPFC networks with SCZ genes (200 kbp list) only. For Sankey plots, hover mouse over the links to see the list of genes. Also supports zoom, drag and selection.<br><strong>DLPFC hit.genes_kb_200__paper.version.zip:</strong><br>Interactive Sankey plot for age-parsed DLPFC networks with SCZ genes (200 kbp list) only. For paper version of the figure, smaller modules are merged into a macro-module (lightgrey color)<br><strong>DLPFC all.genes_kb_200__online.version.zip:</strong><br>Interactive Sankey plot for age-parsed DLPFC networks with all genes<br><strong>DLPFC all.genes_kb_200__paper.version.zip:</strong><br>Interactive Sankey plot for age-parsed DLPFC networks with all genes. For paper version of the figure, smaller modules are merged into a macro-module (lightgrey color)<br><strong>HP hit.genes_kb_200__online.version.zip:</strong><br>Interactive Sankey plot for age-parsed Hippocampus networks with SCZ genes (200 kbp list) only<br><strong>HP hit.genes_kb_200__paper.version.zip:</strong><br>Interactive Sankey plot for age-parsed Hippocampus networks with SCZ genes (200 kbp list) only. For paper version of the figure, smaller modules are merged into a macro-module (lightgrey color)<br><strong>HP all.genes_kb_200__online.version.zip:</strong><br>Interactive Sankey plot for age-parsed Hippocampus networks with all genes<br><strong>HP all.genes_kb_200__paper.version.zip:</strong><br>Interactive Sankey plot for age-parsed Hippocampus networks with all genes. For paper version of the figure, smaller modules are merged into a macro-module (lightgrey color)<br><strong>Modulewise SCZ enrichment(1.0).xlsx:</strong><br>Excel file contains module level SCZ enrichment results for all networks<br><strong>wide_form_test_slidingwindow_NC_SchizoNew(v1.4)_final.xlsx:</strong><br>Excel file contains WGCNA output for sliding window networks<br><strong>wide_form_WGCNA(v3.7.1)_final.xlsx:</strong><br>Excel file contains WGCNA output for our generated networks and from previously published networks<br><strong>libdnetworks(NC).preprocessed.exp.RData: </strong><br>Preprocessed ranknormalised expression assay for age-parsed/nonparsed NC networks (DLPFC, HP, CAUDATE, DENTATE). For fixed window and sliding window study.<br><strong>libdnetworks(SCZ).preprocessed.exp.RData: </strong><br>Preprocessed ranknormalised expression assay for nonparsed SCZ networks (DLPFC, HP, CAUDATE, DENTATE). For the sliding window study.<br><strong>sample_matched_HP_DG_qsva(NC).preprocessed.exp.RData:</strong><br>Preprocessed ranknormalised expression assay for the sample-matched HP-DG. QSVA removed pipeline. For Cell population enrichment study.<br><strong>sample_matched_HP_DG_noqsva(NC).preprocessed.exp.RData:</strong><br>Preprocessed ranknormalised expression assay for the sample-matched HP-DG. No QSVA removed pipeline. For Cell population enrichment study.<br><strong>stemcell.preprocessed.exp.RData:</strong><br>Preprocessed ranknormalised expression assay for the iPSC network. For replication in human iPSC data study. Neuronal samples averaged for each “RealGenome”.<br><strong>SCZ.ref.list.sciadv.ade2812.rds</strong>: List of All Biotypes/ Protein Coding Schizophrenia reference genelist for following bins: PGC3, 0 kbp, 20 kbp, 50 kbp, 100 kbp, 150 kbp, 200 kbp, 250 kbp, 500 kbp.</p> <p> </p> <p>Accompanying code can be found at: <a href="https://github.com/LieberInstitute/Brain_WGCNA">https://github.com/LieberInstitute/Brain_WGCNA</a><br>Data from this repository is also available at: <a href="https://nets.libd.org/age_wgcna/">https://nets.libd.org/age_wgcna/</a></p> <p> </p> <p>For any data inquiries please contact:<br><strong>Giulio Pergola: </strong><a href="mailto:Giulio.Pergola@libd.org"><strong>Giulio.Pergola@libd.org</strong></a></p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.