Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,324
datasets available to search
ShareScore release 0.7.1
Dataset results
7,324 results for “pathways”
Post-War Pathways of Foreign Fighters from the Croatian War (1991--1995) Dataset
<p>This dataset is a part of project 798392 - MPP funded by Marie-Sklodowska Curie Actions Individual Fellowship, a scholarly grant awarded by the European Commission to Dr Milos Popovic as the principal investigator(PI). In total, the dataset includes information on 25 individuals who fought in the Croatian war (1991–1995)as foreign volunteers based on in-person interviews conducted between August 2020 and August 2021 in Croatia and online. It is by no means the universe of cases or a representative sample. The overarching aim of this research is to collect war memories for scientific and archival purposes. All the interviews were set up and carried out by Tomislav Šulj a historian at the Croatian Homeland War Memorial Documentation Centre (https://centardomovinskograta.hr/). The interviewer adhered to a code of conduct approved by Leiden University and the European Commission. First, any participation in the interview was completely voluntary and the participants had the right to refuse participation or any question at any stage or withdraw at any time without any consequence whatsoever while preserving their anonymity. Second, potential interviewees were given a detailed study guidebook beforehand, including the description, aim and anonymity/confidentiality clauses. Third, the interviewees were asked to sign their consent to the use of interview data for scientific purposes, including the production of this dataset. Interviewees were also informed that their anonymity would be broken only under extraordinary circumstances when there is a serious risk of harm or danger to either the interviewer, interviewee or another individual (e.g. physical, emotional or sexual abuse, concerns for child protection, rape, self-harm, suicidal intent or criminal activity) or if a crime was committed. All the conversations were audio-recorded upon the interviewee’s explicit consent. Every signed consent form was scanned and placed together with Dr Popovic’s original audio recording of the interview. There are two copies of the interview notes and audio recordings. One copy is stored in an encrypted folder on Dr Popovic’s personal and office computer at the Institute of Security and Global Affairs in the Hague until after Dr Popovic’s project has ended on September 1, 2021. The second copy is stored and archived at the CroatianHomeland War Memorial Documentation Centre and, following the authorization of the interview transcript, used for research on the Croatian Homeland War. Under the provisions of the General Data Protection Regulation (EU) 2016/679 (GDPR), every interviewee is entitled to access the information they have provided at any time. </p>
Post-War Pathways of Foreign Fighters from the Bosnian War (1992--1995)
<p>This dataset is a part of project 798392 - MPP funded by Marie-Sklodowska Curie Actions Individual Fellowship, a scholarly grant awarded by the European Commission to Dr. Milos Popovic as the principal investigator (PI). In total, the dataset includes information on 95 individuals who fought in the Bosnian war (1991--1995) as foreign volunteers based on available online resources. It is by no means the universe of cases or a representative sample. The overarching aim of this research is to collect war memories for scientific and archival purposes. </p> <p>\noindent The dataset is the first to date to include information on the Greek and Russian foreign fighters who fought on the side of the Army of the Serb Republic of Bosnia-Herzegovina. The dataset does not feature information on the former mujahideen fighters because there is little information on the individual fighters. The key resource used to gather information on the Greek volunteers is the massive investigative work of the XYZ contagion investigative work, which can be found at https://xyzcontagion.wordpress.com/2015/06/17/maria-lefteris-ethnos-tagmata-matosan-srebrenica/. Information on Russian foreign fighters comes mostly from two books published by former foreign fighters Mikhail Polikarpov and Oleg Valetskiy:</p> <p>Валецкий, О. В. (1993). Волки белые. <em>Сербский дневник русского добровольца (1993--1999</em>), Грифон М;<br> Поликарпов, М. А. (2007). <em>Сербский закат</em>. Москва: Эксмо.</p> <p>For data triangulation purposes, I also used a few online sources from Serbian forums as well as Soldier of Fortune issues (in Russian) for the period 1992-1995. In addition, Aziz Tafro's book <em>Russian and Greek Henchmen in the war in Bosnia-Herzegovina</em> provided useful information on the deceased Russian volunteers. Finally, I relied on numerous articles by former foreign fighters as well as interviews that they gave to various news portals.</p>
Adverse Outcome Pathway Wiki RDF
<p>This dataset is the RDF generated from the AOP-Wiki data release (<a href="https://aopwiki.org/downloads">aopwiki.org/downloads</a>). It was generated using a Jupyter notebook that is available on GitHub (<a href="https://github.com/marvinm2/AOPWikiRDF">github.com/marvinm2/AOPWikiRDF</a>), and the process and additional description of the RDF have been published (<a href="https://doi.org/10.1089/aivt.2021.0010">doi.org/10.1089/aivt.2021.0010</a>).</p>
Meta-analysis on necessary investment shifts to reach net zero pathways in Europe
<p>This is the code and the data necessary to reproduce the six main figures and the t-test presented in the supplementary information of the publication "Meta-analysis on necessary investment shifts to reach net zero pathways in Europe". DOI: 10.1038/s41558-022-01549-5</p>
Anypodetus dichotomous, pathway key in SDD-format
<p>Dichotomous, pathway key to species of genus Anypodetus (Diptera: Asilidae) developed with Lucid Builder v4 in XML Structure of Descriptive Data (SDD) format.</p>
MicroRNA-target pathways in acute myeloid leukaemia
<p>Pathway map of 17 microRNAs (miRs) in acute myeloid leukaemia (AML). Details over- and under-expressed miRs, the impact on relevant targets, interaction of targets with other proteins and/or pathways, and the overall impact on AML onset, progression and/or maintenance. All miR targets identified and verified via miRTarBase, KEGG and relevant literature. </p>
RDF dataset produced in the work "Exploring Adverse Outcome Pathways for Nanomaterials with semantic web technologies"
<p>Adverse Outcome Pathways (AOPs) have been proposed to facilitate mechanistic understanding of interactions of chemicals/materials with biological systems. Each AOP starts with a molecular initiating event (MIE) and possibly ends with adverse outcome(s) (AOs) via a series of key events (KEs). So far, the interaction of engineered nanomaterials (ENMs) with biomolecules, biomembranes, cells, and biological structures, in general, is not yet fully elucidated. There is also a huge lack of information on which AOPs are ENMs-relevant or -specific, despite numerous published data on toxicological endpoints they trigger, such as oxidative stress and inflammation. We propose to integrate related data and knowledge recently collected. Our approach combines the annotation of nanomaterials and their MIEs with ontology annotation to demonstrate how we can then query AOPs and biological pathway information for these materials. We conclude that a FAIR (Findable, Accessible, Interoperable, Reusable) representation of the ENM-MIE knowledge simplifies integration with other knowledge.</p>
Supplementary datasets for "ARBRE: Computational resource to predict pathways towards industrially important aromatic compounds"
<p>Supplementary datasets accompanying the manuscript "ARBRE: Computational resource to predict pathways towards industrially important aromatic compounds" published in the Metabolic Engineering Journal (<a href="https://doi.org/10.1016/j.ymben.2022.03.013">https://doi.org/10.1016/j.ymben.2022.03.013). </a>In line with the standards of open science, the ARBRE toolbox is freely available to the scientific community on gitHub (<a href="https://github.com/EPFL-LCSB/ARBRE">https://github.com/EPFL-LCSB/ARBRE</a>) and we also provide the web-version at <a href="http://lcsb-databases.epfl.ch/arbre/">http://lcsb-databases.epfl.ch/arbre/</a></p> <p>ARBRE: Aromatic compounds RetroBiosynthesis Repository and Explorer is a new computational resource consisting of a comprehensive biochemical reaction network centered around aromatic amino acid biosynthesis and a computational toolbox for navigating this network. ARBRE encompasses over 33′000 known and 390′000 novel reactions predicted with generalized enzymatic reactions rules and over 74′000 compounds, of which 19′000 are known to biochemical databases and 55′000 only to PubChem. Over 1′000 molecules that were solely part of the PubChem database before and were previously impossible to integrate into a biochemical network are included in the ARBRE reaction network by assigning enzymatic reactions. ARBRE can be applied for pathway search, enzyme annotation, pathway ranking, visualization, and network expansion around known biochemical pathways and products of lignin degradation to predict valuable compound derivations.</p> <p>Supplementary files are organized as follows:</p> <p>- 1-s2.0-S1096717622000490-mmc4.docx contains Supplementary Figures 1-4 and Tables 1, 2, and 4.</p> <p>- 1-s2.0-S1096717622000490-mmc2.xlsx contains Supplementary Table 3.</p> <p>- 1-s2.0-S1096717622000490-mmc1.xlsx contains Supplementary Table 5</p> <p>- 1-s2.0-S1096717622000490-mmc3.xlsx contains Supplementary Table 6</p> <p> </p> <p> </p> <p> </p> <p> </p>
Simulated spatially explicit dataset (300 m) on future forest cover changes in Southeast Asia projected under the baseline shared socioeconomic pathways
<p>This is a simulated spatially explicit dataset on future forest cover changes in Southeast Asia projected under the baseline shared socioeconomic pathways. It includes six raster maps at a spatial resolution of 300 m: (1) 2015 baseline forest and non-forest map; (2) SSP1 2050 projected net forest gain map; (3) SSP2 2050 projected net forest gain map; (4) SSP3 2050 projected net forest loss map; (5) SSP4 2050 projected net forest gain map; and SSP5 2050 projected net forest loss map. This dataset is the result of a study published in Nature Communications (2019) (https://doi.org/10.1038/s41467-019-09646-4).</p>
A Complement Atlas identifies interleukin 6 dependent alternative pathway dysregulation as a key druggable feature of COVID-19.
<p>Improvements in COVID-19 treatments, especially for the critically ill, require deeper understanding of the mechanisms driving disease pathology. The complement system is a crucial component of innate host defense, but can also contribute to tissue injury. Although all complement pathways have been implicated in COVID-19 pathogenesis, the upstream drivers and downstream effects on tissue injury remain poorly defined. We demonstrate that complement activation is primarily mediated by the alternative pathway, and we provide a comprehensive atlas of the complement alterations around the time of respiratory deterioration. Proteomic and single-cell sequencing mapping across cell types and tissues reveals a division of labor between lung epithelial, stromal, and myeloid cells in complement production, in addition to liver-derived factors. We identify IL-6 and STAT1/3 signaling as an upstream driver of complement responses, linking complement dysregulation to approved COVID-19 therapies. Furthermore, an exploratory proteomic study indicates that inhibition of complement C5 decreases epithelial damage and markers of disease severity. Collectively, these results support complement dysregulation as a key druggable feature of COVID-19.</p>
Data deposit accompanying Accurate Energy Barriers for Catalytic Reaction Pathways: An Automatic Training Protocol for Machine Learning Force Fields
<p>Dataset accompanying the paper: <em>"Accurate Energy Barriers for Catalytic Reaction Pathways: An Automatic Training Protocol for Machine Learning Force Fields"</em>. Contains the training sets curated during active learning as well as .xyz files used for creating the Figures. <br> <br> The paper highlights that the computational efficiency of ML force fields not only results in decreased computational costs for routine catalytic investigations but also facilitates more comprehensive exploration of catalytic pathways.</p> <p><strong>Published in NPJ Computational Materials</strong>: <a href="https://www.nature.com/articles/s41524-023-01124-2">https://www.nature.com/articles/s41524-023-01124-2</a><br> Formerly on Arxiv: <a href="https://arxiv.org/abs/2301.09931">https://arxiv.org/abs/2301.09931</a></p>
Data for: Sparse subalpine forest recovery pathways, plant communities, and carbon stocks 34 years after stand-replacing fire (Greater Yellowstone Ecosystem, Wyoming, USA; 2022)
We assessed postfire forest recovery pathways, stem densities, understory plant communities, and carbon stocks across 55 plots in areas exhibiting sparse and reduced forest recovery 34 years after the 1988 Yellowstone Fires in the Greater Yellowstone Ecosystem, Wyoming, USA. Recovery pathways were identified using plot-level frequency distributions of tree ages and correlated with potentially important biotic and abiotic variables (e.g., elevation, seed source distance). Species- and age-specific stem densities were similarly regressed across environmental factors to determine variability in forest recovery across the sampled landscape. Understory plant communities were sampled in 0.25m-square quadrats and environmental drivers of individual species occurrence and whole compositional shifts were determined. Finally, carbon stock sizes were derived from field measures of tree characteristics, understory cover, and soil combined with regionally derived allometric equations. Data collection is complete and is part of a forthcoming manuscript at Ecological Monographs.
Age structure, developmental pathways, and fire regime characterization of Douglas-fir/western hemlock forests in the central western Cascades of Oregon
These data are the raw forest stand- and age-structure data from 124 stands in the central western Cascades of Oregon used to construct a conceptual model of stand development under the mixed-severity fire regime that has operated extensively in this region.
Robustness and applicability of transcription factor and pathway analysis tools on single-cell RNA-seq data
<p>Data used to test the robustness and applicability of transcription factor and pathway analysis tools on single-cell RNA-seq data, described in <a href="https://doi.org/10.1186/s13059-020-1949-z">Holland et al. 2020</a>.</p> <p>The folder <em>data </em>contains<em> </em>raw data and the folder <em>output</em> contains intermediate and final results of all analyses. </p> <p>The associated analyses code and more information are available on <a href="https://github.com/saezlab/FootprintMethods_on_scRNAseq">GitHub</a>.</p> <p> </p> <p><strong>Abstract</strong></p> <p><strong>Background</strong></p> <p>Many functional analysis tools have been developed to extract functional and mechanistic insight from bulk transcriptome data. With the advent of single-cell RNA sequencing (scRNA-seq), it is in principle possible to do such an analysis for single cells. However, scRNA-seq data has characteristics such as drop-out events and low library sizes. It is thus not clear if functional TF and pathway analysis tools established for bulk sequencing can be applied to scRNA-seq in a meaningful way.</p> <p><strong>Results</strong></p> <p>To address this question, we perform benchmark studies on simulated and real scRNA-seq data. We include the bulk-RNA tools PROGENy, GO enrichment, and DoRothEA that estimate pathway and transcription factor (TF) activities, respectively, and compare them against the tools SCENIC/AUCell and metaVIPER, designed for scRNA-seq. For the in silico study, we simulate single cells from TF/pathway perturbation bulk RNA-seq experiments. We complement the simulated data with real scRNA-seq data upon CRISPR-mediated knock-out. Our benchmarks on simulated and real data reveal comparable performance to the original bulk data. Additionally, we show that the TF and pathway activities preserve cell type-specific variability by analyzing a mixture sample sequenced with 13 scRNA-seq protocols. We also provide the benchmark data for further use by the community.</p> <p><strong>Conclusions</strong></p> <p>Our analyses suggest that bulk-based functional analysis tools that use manually curated footprint gene sets can be applied to scRNA-seq data, partially outperforming dedicated single-cell tools. Furthermore, we find that the performance of functional analysis tools is more sensitive to the gene sets than to the statistic used.</p> <p> </p> <p>For questions related to the data please write an email to christian.holland@bioquant.uni-heidelberg.de or use the <a href="https://github.com/saezlab/FootprintMethods_on_scRNAseq/issues">GitHub issue system</a>.</p>
Data set for "Pathway-, layer- and cell-type-specific thalamic input to mouse barrel cortex"
<p>Data set for: Sermet BS, Truschow P, Feyerabend M, Mayrhofer JM, Oram TB, Yizhar O, Staiger JF, Petersen CCH (2019) Pathway-, layer- and cell-type-specific thalamic input to mouse barrel cortex. eLife 8: e52665. https://doi.org/10.7554/eLife.52665</p> <p>There are 2 files in this upload:</p> <p>1. The file named "2019_Sermet_eLife.pdf" is the Open Access pdf file of the manuscript published in eLife.</p> <p>2. The file named "Sermet_data_code.zip" (~5 GB) is a zipped version of a folder "Sermet_data_code" (~5 GB), which contains the data analysed in the study along with the Matlab code used to generate the published figures. When unzipped, the folder contains 8 Matlab '.m' files with analysis code and one '.mat' data file. In order to run the analysis of the data set, you need to execute 'PopPlot.m'.</p>
data for "A disordered encounter complex is central to the yeast Abp1p SH3 domain binding pathway"
<p>Protein-protein interactions are involved in a wide range of cellular processes. These interactions often involve intrinsically disordered proteins (IDPs) and protein binding domains. However, the details of IDP binding pathways are hard to characterize using experimental approaches, which can rarely capture intermediate states present at low populations. SH3 domains are common protein interaction domains that typically bind proline-rich disordered segments and are involved in cell signaling, regulation, and assembly. We hypothesized, given the flexibility of SH3 binding peptides, that their binding pathways include multiple steps important for function. Molecular dynamics simulations were used to characterize the steps of binding between the yeast Abp1p SH3 domain (AbpSH3) and a proline-rich IDP, ArkA. Before binding, the N-terminal segment 1 of ArkA is pre-structured and adopts a polyproline II helix, while segment 2 of ArkA (C-terminal) adopts a 310 helix, but is far less structured than segment 1. As segment 2 interacts with AbpSH3, it becomes more structured, but retains flexibility even in the fully engaged state. Binding simulations reveal that ArkA enters a flexible encounter complex before forming the fully engaged bound complex. In the encounter complex, transient nonspecific hydrophobic and long- range electrostatic contacts form between ArkA and the binding surface of SH3. The encounter complex ensemble includes conformations with segment 1 in both the forward and reverse orientation, suggesting that segment 2 may play a role in stabilizing the correct binding orientation. While the encounter complex forms quickly, the slow step of binding is the transition from the disordered encounter ensemble to the fully engaged state. In this transition, ArkA makes specific contacts with AbpSH3 and buries more hydrophobic surface. Simulating the binding between ApbSH3 and ArkA provides insight into the role of encounter complex intermediates and nonnative hydrophobic interactions for other SH3 domains and IDPs in general.</p>
Data from: GPCR genes as activators of surface colonization pathways in a model marine diatom
<p>Surface colonization allows diatoms, a dominant group of phytoplankton in oceans, to adapt to harsh marine environments while mediating biofoulings to human-made underwater facilities. The regulatory pathways underlying diatom surface colonization, which involves morphotype switching in some species, remain mostly unknown. Here, we describe the identifications of 61 signaling genes, including G-protein-coupled receptors (GPCRs) and protein kinases, that are differentially regulated during surface colonization in the model diatom species, <em>Phaeodactylum tricornutum</em>. We show that the transformation of <em>P. tricornutum</em> with constructs expressing individual GPCR genes induces cells to adopt the surface colonization morphology. <em>P. tricornutum</em> cells transformed to express GPCR1A display 30% more resistance to UV light exposure than their non-biofouling wild type counterparts, consistent with increased silicification of cell walls associated with the oval-biofouling morphotype. Our results provide a mechanistic definition of morphological shifts during surface colonization and identify candidate target proteins for the screening of eco-friendly, anti-biofouling molecules.</p>
Efficient pathways performances
<p>The dataset contains values of design indicators for all the efficient pathways identified by the Decision Analytic Framework in the DAFNE Project.</p> <p><strong>Zambezi River Basin</strong> (file <em>zrb_efficient_pathways.txt</em>)</p> <ul> <li>time horizon: 2020-2060</li> <li>design indicators: <ul> <li><em>j_environment</em>: Environmental flow deficit</li> <li><em>j_hydropower</em>: Hydropower production deficit</li> <li><em>j_irrigation_deficit</em>: Normalized irrigation deficit</li> <li><em>j_cost</em>: Total discounted cost</li> </ul> </li> </ul> <p><strong>Omo-Turkana Basin</strong> (file <em>otb_efficient_pathways.zip</em>)</p> <ul> <li>time horizon: 2002-2016</li> <li>design indicators: <ul> <li><em>j_Env</em>: Environmental flow deficit</li> <li><em>j_Hyd: </em>Hydropower production</li> <li><em>j_Irr</em>: Normalized irrigation deficit for large scale irrigation district</li> <li><em>j_Rec</em>: Deficit with respect to the target flood requirement in the Omo Delta for recession agricolture</li> <li><em>j_Fish:</em> Deficit of Fish biomass production in the Lake Turkana with respect to natural condition</li> </ul> </li> <li>pathways are related to four different configuration and labelled progressive within each group: <ul> <li>P0: Baseline, with no infrastructure development</li> <li>P1: Koysha, Baseline + Koysha dam</li> <li>P2: Irrigation, Baseline + Irrigation development</li> <li>P3: Irrigation and Koysha, Baseline + Irrigation development + Koysha dam</li> </ul> </li> </ul> <p>Indicators formulation and details on the efficient pathways generation can be found on DAFNE project Deliverables D5.2 and D5.4.</p> <p>These dataset have been used to populate the Multi-Perspective-Visual-Analytics tool and to perform the screening exercise during the second NSL meeting in both case studies (see also DAFNE Deliverables D7.4).</p>
Data to initialize a TAD_Pathways Analysis
<p>Dataset is required for a TAD_Pathways analysis (see https://github.com/greenelab/tad_pathways_pipeline).</p> <p>Archived folder includes a TAD based gene index file, curated SNPs from the NHGRI-EBI GWAS catalog, and TAD based genes and SNPs for each GWAS.</p>
10 eADAGE models used for the PAO1 KEGG pathways case study in PathCORE
<p>ensemble Analysis using Denoising Autoencoders for Gene Expression<strong> </strong>(<strong>eADAGE</strong>) is an unsupervised feature construction algorithm developed by Tan et al. that uses an ensemble of neural networks (an ensemble of ADAGE models) to capture biological signatures embedded in the expression compendium. By initializing eADAGE with different random seeds, Tan et al. produced 10 eADAGE models that each extracted k=300 features from the compendium of genome-scale <em>P. aeruginosa</em> data.</p> <p>eADAGE is described in Tan et al.'s "System-wide automatic extraction of functional signatures in <em>Pseudomonas aeruginosa</em> with eADAGE" (https://doi.org/10.1101/078659). The code to construct these 10 models is available in this repository: https://bitbucket.org/greenelab/eadage (see eADAGE_construction.sh). </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.