Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,634
datasets available to search
ShareScore release 0.9.0
Dataset results
1,634 results for “Data integration”
Subtyping of common complex diseases and disorders by integrating heterogeneous data. Identifying clusters among women with lower urinary tract symptoms in the LURN study
<p>We present a methodology for subtyping of persons with a common clinical symptom complex by integrating heterogeneous continuous and categorical data. We illustrate it by clustering women with lower urinary tract symptoms (LUTS), who represent a heterogeneous cohort with overlapping symptoms and multifactorial etiology. Data collected in the Symptoms of Lower Urinary Tract Dysfunction Research Network (LURN), a multi-center observational study, included self-reported urinary and non-urinary symptoms, bladder diaries, and physical examination data for 545 women. Heterogeneity in these multidimensional data required thorough and non-trivial preprocessing, including scaling by controls and weighting to mitigate data redundancy, while the various data types (continuous and categorical) required novel methodology using a weighted Tanimoto indices approach. Data domains only available on a subset of the cohort were integrated using a semi-supervised clustering approach. Novel contrast criterion for determination of the optimal number of clusters in consensus clustering was introduced and compared with existing criteria. Distinctiveness of the clusters was confirmed by using multiple criteria for cluster quality, and by testing for significantly different variables in pairwise comparisons of the clusters. Cluster dynamics were explored by analyzing longitudinal data at 3- and 12-month follow-up. Five clusters of women with LUTS were identified using the developed methodology. None of the clusters could be characterized by a single symptom, but rather by a distinct combination of symptoms with various levels of severity. Targeted proteomics of serum samples demonstrated that differentially abundant proteins and affected pathways are different across the clusters. The clinical relevance of the identified clusters is discussed and compared with the current conventional approaches to the evaluation of LUTS patients. The rationale and thought process are described for the selection of procedures for data preprocessing, clustering, and cluster evaluation. Suggestions are provided for minimum reporting requirements in publications utilizing clustering methodology with multiple heterogeneous data domains.</p>
Data from: Integrating effects of neighbor interactions for pollination
<p>Animal-pollinated plants interact with neighbors for both abiotic resources and pollination, with consequences for reproduction and yield. Yet few studies have compared the relative magnitude of these effects, particularly in agroecosystems. In vertically stratified communities, such as agroforests, neighbor effects may be stratum-dependent. Understanding the net effects of neighbors on crop yield is important for managing multifunctional agroecosystems to simultaneously support production and biodiversity. This study evaluated the effects of neighboring plants on pollen deposition, fertilization, and yield in Coffea arabica in a shaded organic coffee farm with high non-crop plant abundance and diversity in Chiapas, Mexico. We assessed the impact of 1) floral resources at three vertical strata (herbs, coffee bushes, and canopy trees) on stigma pollen load (a measure of interaction for pollination), and 2) floral density and canopy cover (proxies for competition for abiotic resources) on yield (final fruit set and per-fruit weight), using structural equation modeling to evaluate the relative effect of each interaction type. Coffee competed for pollination with neighbors (conspecifics and heterospecifics) across strata. Pollen load influenced final fruit set, but the effect of neighbor competition for pollination was weaker than effects mediated by interaction for abiotic resources. Effects of interactions for abiotic resources were heterogeneous across strata, with negligible effects of herb-layer or coffee flower density but net positive effects of canopy trees on final fruit set. Overall effects of neighbors on coffee yield were weak, suggesting that coffee agroecosystems can be managed to maintain high plant density and diversity without sacrificing yield.</p>
Data for empirical example in: An effect size for comparing the strength of morphological integration across studies
<p>Understanding how and why phenotypic traits covary is a major interest in evolutionary biology. Biologists have long sought to characterize the extent of morphological integration in organisms, but comparing levels of integration for a set of traits across taxa has been hampered by the lack of a reliable summary measure and testing procedure. Here we propose a standardized effect size for this purpose, calculated from the relative eigenvalue variance, Vrel. First we evaluate several eigenvalue dispersion indices under various conditions, and show that only Vrel remains stable across samples size and the number of variables. We then demonstrate that Vrel accurately characterizes input patterns of covariation, so long as redundant dimensions are excluded from the calculations. However, we also show that the variance of the sampling distribution of Vrel depends on input levels of trait covariation, making Vrel unsuitable for direct comparisons. As a solution, we propose transforming Vrel to a standardized effect size (Z-score) for representing the magnitude of integration for a set of traits. We also propose a two-sample test for comparing the strength of integration between taxa, and show that this test displays appropriate statistical properties. We provide software for implementing the procedure, and an empirical example illustrates its use.</p>
Integrated Statistical Indicators from Scottish Linked Open Government Data
<p>Integrated statistical indicators that were retrieved from the official Scottish data portal in order to facilitate the exploitation of Machine Learning methods in Open Government Data. Data include 60 statistical indicators from seven categories such as health and social care, housing, and crime and justice. The indicators refer to the 6,976 “2011 data zones” of Scotland, while the year of reference is 2015. Data are ready to be used by the research community, students, policy makers, and journalists and give rise to plenty of social, business, and research scenarios that can be solved using Machine Learning technologies and methods.</p>
U.S. cities increasingly integrate justice into climate planning and create policy tools for climate justice (Diezmartínez & Short Gianotti, 2022) - Data and code
<p>This repository contains datasets and coding corresponding to the journal article titled "U.S. cities increasingly integrate justice into climate planning and create policy tools for climate justice". We include:</p> <ul> <li>DataRegressionAnalysis.csv <ul> <li>CSV file with data used for regression analysis. This file can be used directly to run R code provided in this repository.</li> </ul> </li> <li>QualitativeCodingResults.nvp <ul> <li>NVivo project with all results for the qualitative coding of urban climate action plans.</li> <li>This file also contains all climate action plans analyzed in this research.</li> </ul> </li> <li>QualitativeCodingResults_Summary.xlsx <ul> <li>Excel file with a results summary for the qualitative coding of urban climate action plans.</li> </ul> </li> <li>RegressionAnalysis.Rmd <ul> <li>Rmd file with R code used for regression analysis. </li> </ul> </li> <li>RegressionAnalysis_KnitOutput.html <ul> <li>Knit output from R code with regression analysis results, html format.</li> </ul> </li> <li>RegressionAnalysis_KnitOutput.pdf <ul> <li>Knit output from R code with regression analysis results, PDF format.</li> </ul> </li> </ul>
NAPS Integrated Data
<p>Scripts used to download and clean NAPS Integrated data</p> <p> </p>
Data from: Identification of a minority population of LMO2+ breast cancer cells that integrate into the vasculature and initiate metastasis.
<p>Metastasis is responsible for the majority of breast cancer-related deaths, however, identifying the cellular determinants of metastasis has remained challenging. Here, we identified a minority population of immature THY1+/VEGFA+ tumor epithelial cells in human breast tumor biopsies that display angiogenic features and are marked by the expression of the oncogene, LMO2. Higher abundance of LMO2+ basal cells correlated with tumor endothelial content and predicted poor distant recurrence-free survival in patients. Using MMTV-PyMT/Lmo2CreERT2 mice, we demonstrated that Lmo2 lineage-traced cells integrate into the vasculature and have a higher propensity to metastasize. LMO2 knockdown in human breast tumors reduced lung metastasis by impairing intravasation, leading to a reduced frequency of circulating tumor cells. Mechanistically, we find that LMO2 binds to STAT3 and is required for STAT3 activation by TNFα and IL6. Collectively, our study identifies a population of metastasis-initiating cells with angiogenic features and establishes the LMO2-STAT3 signaling axis as a therapeutic target in breast cancer metastasis.</p>
Data: Applying stochastic and Bayesian integral projection modeling to amphibian population viability analysis
<p>Integral projection models (IPMs) can estimate the population dynamics of species for which both discrete life stages and continuous variables influence demographic rates. Stochastic IPMs for imperiled species, in turn, can facilitate population viability analyses (PVAs) to guide conservation decision-making. Biphasic amphibians are globally distributed, often highly imperiled, and ecologically well-suited to the IPM approach. Herein, we present the first stochastic size- and stage-structured IPM for a biphasic amphibian, the U.S. federally threatened California tiger salamander (<em>Ambystoma</em> <em>californiense</em>; CTS). This Bayesian model reveals that CTS population dynamics show the greatest elasticity to changes in juvenile and metamorph growth and that populations are likely to experience rapid growth at low density. We integrated this IPM with climatic drivers of CTS demography to develop a PVA and examined CTS extinction risk under the primary threats of habitat loss and climate change. The PVA indicates that long-term viability is possible with surprisingly high (20–50%) terrestrial mortality, but simultaneously identified likely minimum terrestrial buffer requirements of 600–1000 m while accounting for numerous parameter uncertainties through the Bayesian framework. These analyses underscore the value of stochastic and Bayesian IPMs for understanding both climate-dependent taxa and those with cryptic life histories (e.g., biphasic amphibians) in service of ecological discovery and biodiversity conservation. In addition to providing guidance for CTS recovery, the contributed IPM and PVA supply a framework for applying these tools to investigations of ecologically-similar species.</p>
MSME T2 data from the Ferret Interactive Integrated Neurodevelopment Atlas
<p>The first days after birth in ferrets provide a unique view of the development of a complex brain. Unlike mice, ferrets develop a rich pattern of deep neocortical folds and cortico-cortical connections. Unlike humans and other primates, whose brains are well differentiated and folded at birth, ferrets are born with a very immature and completely smooth neocortex: folds, neocortical regionalisation and cortico-cortical connectivity develop in ferrets during the first days after birth. After a period of fast neocortical expansion, during which brain volume increases by up to a factor of 4 in 2 weeks, the ferret brain reaches its adult volume at about 6 weeks of age. This dataset contains brain MRI T2 data from 28 ferrets from P0 to Adults. It can be visualised at http://brainbox.pasteur.fr/project/FIIND.</p>
Figure 1. Schema for a data-integration solution-A Proposed Data Driven Architecture for Cardiology Network Application
<p>Data integration has favored loosening the coupling between data. This may involve<br> providing a uniform query interface over a mediated schema (see figure 1), thus transforming<br> a query into specialized queries over the original databases. One can also term this process<br> "view-based query-answering" because each of the data sources functions as a view over the<br> (nonexistent) mediated schema.</p>
Figure 3. Data Integration Sample-A Proposed Data Driven Architecture for Cardiology Network Application
<p>Pentaho Data Integration has implemented a metadata-driven approach where you<br> only specify the data you want integrated, but you do not specify the way you want it done.<br> One of the most important advantages of Pentaho is that one can create complex<br> transformations and jobs in a graphical, drag-and-drop environment without having to create<br> proprietary custom code that will work only with some proprietary application.</p>
DrugComb - an integrative cancer drug combination data portal (v1.4)
<p>Drug combination therapy has the potential to enhance efficacy, reduce dose-dependent toxicity and prevent the emergence of drug resistance. However, discovery of synergistic and effective drug combinations has been a laborious and often serendipitous process. In recent years, identification of combination therapies has been accelerated due to the advances in high-throughput drug screening, but informatics approaches for systems-level data management and analysis are needed. To contribute toward this goal, we created an open-access data portal called DrugComb (<a href="https://drugcomb.org/">https://drugcomb.org</a>) where the results of drug combination screening studies are accumulated, standardized and harmonized. Through the data portal, we provided a web server to analyze and visualize users' own drug combination screening data. The users can also effectively participate a crowdsourcing data curation effect by depositing their data at DrugComb. To initiate the data repository, we collected 437 932 drug combinations tested on a variety of cancer cell lines. We showed that linear regression approaches, when considering chemical fingerprints as predictors, have the potential to achieve high accuracy of predicting the sensitivity of drug combinations. All the data and informatics tools are freely available in DrugComb to enable a more efficient utilization of data resources for future drug combination discovery.</p> <p><strong>Citations:</strong> </p> <p>[1] Nucleic Acids Res. 2021, 49(W1):W174-W184. doi: 10.1093/nar/gkab438</p> <p><a href="https://pubmed.ncbi.nlm.nih.gov/34060634/" target="_blank" rel="noopener">https://pubmed.ncbi.nlm.nih.gov/34060634/ </a></p> <p>[2] Nucleic Acids Res. 2019, 47(W1):W43-W51. doi: 10.1093/nar/gkz337<br><a href="https://pubmed.ncbi.nlm.nih.gov/31066443/" target="_blank" rel="noopener">https://pubmed.ncbi.nlm.nih.gov/31066443/</a></p>
Data and codes for "Decay-protected superconducting qubit with fast control enabled by integrated on-chip filters"
<p>Data and codes for "Decay-protected superconducting qubit with fast control enabled by integrated on-chip filters".</p>
Integrated Machine Learning model in Early Urban Flooding Warning System - Data
<p>AI_DATA.npy - Inundation data (mm) generated from MIKE+ model that has been converted to numpy array</p> <p>INDEX.npy - The index where inundation is > 0 </p> <p>source.tif - Source tif image for creating map from ML models</p>
Dataset for 'A Resource Hub For Interoperability And Data Integration In Heritage Research: The H-Setis Database'
<p> Data and scripts for charts and maps published in "A Resource Hub For Interoperability And Data Integration In Heritage Research: The H-Setis Database".</p>
Improved Detection of Drug-Induced Liver Injury by Integrating Predicted in vivo and in vitro Data
<p><span>This repository provides datasets for the study: https://broad.io/DILIPredictor</span></p> <p><span>Full Paper: https://www.biorxiv.org/content/10.1101/2024.01.10.575128v1</span></p> <p><span>This work is on enhancing the early detection of Drug-Induced Liver Injury (DILI) through the integration of predicted in vivo and in vitro data. This project utilizes advanced machine learning models and chemical informatics to predict the likelihood of DILI for various compounds. </span></p> <p><span>For code see: <a href="https://github.com/srijitseal/DILI">https://github.com/srijitseal/DILI</a><br><br>Drug-induced liver injury (DILI) has been significant challenge in drug discovery, often leading to clinical trial failures and necessitating drug withdrawals. The existing suite of in vitro proxy-DILI assays is generally effective at identifying compounds with hepatotoxicity. However, there is considerable interest in enhancing in silico prediction of DILI because it allows for the evaluation of large sets of compounds more quickly and cost-effectively, particularly in the early stages of projects. In this study, we aim to study ML models for DILI prediction that first predicts nine proxy-DILI labels and then uses them as features in addition to chemical structural features to predict DILI. The features include <em>in vitro</em> (e.g., mitochondrial toxicity, bile salt export pump inhibition) data, <em>in vivo</em> (e.g., preclinical rat hepatotoxicity studies) data, pharmacokinetic parameters of maximum concentration, structural fingerprints, and physicochemical parameters. We trained DILI-prediction models on 888 compounds from the DILIst dataset and tested on a held-out external test set of 223 compounds from DILIst dataset. The best model, DILIPredictor, attained an AUC-ROC of 0.79. This model enabled the detection of top 25 toxic compounds compared to models using only structural features (2.68 LR+ score). Using feature interpretation from DILIPredictor, we were able to identify the chemical substructures causing DILI as well as differentiate cases DILI is caused by compounds in animals but not in humans. For example, DILIPredictor correctly recognized 2-butoxyethanol as non-toxic in humans despite its hepatotoxicity in mice models. Overall, the DILIPredictor model improves the detection of compounds causing DILI with an improved differentiation between animal and human sensitivity as well as the potential for mechanism evaluation. DILIPredictor is publicly available at </span><a href="https://broad.io/DILIPredictor">https://broad.io/DILIPredictor</a> <span>for use <em>via</em> web interface and with all code available for download and local implementation via </span><a href="https://pypi.org/project/dilipred/"><span>https://pypi.org/project/dilipred/</span></a><span>.</span></p>
FCH and FS Datasets for the paper "Integrating Multi-Source Remote Sensing Data for Mapping Boreal Forest Canopy Height and Species in interior Alaska in Support of Radar Modeling"
<p>This dataset provides forest canopy height and forest species in Delta Junction, interior Alaska in 2017. This dataset was produced based on the multi-source remote sensing datasets (AirMOSS, UAVSAR, Sentinel-1, Sentinel-2, topography), using a XGBoost approach.</p>
Fig. 6 in Polyclinum constellatum (Tunicata, Ascidiacea), an emerging non-indigenous species of the Mediterranean Sea: integrated taxonomy and the importance of reliable DNA barcode data Abstract
Fig. 6: ML phylogenetic tree of the genus Polyclinum (sequences abbreviation: Pln) based on COI nucleotide sequences (1560 aligned nucleotide sites; best-fit substitution model GTR+I+G; bootstrap on 100 replicates). Eudistoma and Pseudodistoma species were used as outgroups. The sequence list and species abbreviations are reported in Supplementary table S1. Black dots: bootstrap values ≥ 70 %; red: P. constellatum sequences; blue: P. indicum sequences; yellow background: our sequences.
Fig. 4 in Polyclinum constellatum (Tunicata, Ascidiacea), an emerging non-indigenous species of the Mediterranean Sea: integrated taxonomy and the importance of reliable DNA barcode data Abstract
Fig. 4: A, C) Colonies of Polyclinum constellatum with different colours photographed and collected in the Heraklion marina (Crete) (A: colony K11 and C: colony K12); B) Transversal section of the colonies, joined only at the surface layer (upper white arrow); D) Zooid extracted from the red-orange colony (K11), with magnification of the 6-lobed anus; E) Zooid extracted from the dark blue colony (K12) with magnification of the 6-lobed anus. Both K11 and K12 have the same COI haplotype (sequence AC number: MT873559).
Fig. 5 in Polyclinum constellatum (Tunicata, Ascidiacea), an emerging non-indigenous species of the Mediterranean Sea: integrated taxonomy and the importance of reliable DNA barcode data Abstract
Fig. 5: A) Larva of P. constellatum, showing the ocellus, four long narrow ampullae, three adhesive papillae and a group of a few small ventral vesicles (red arrow). am, ampullae; ap, adhesive papillae; oc, ocellus; B) Larva of P. constellatum, red arrow pointing out the calcite crystal in the middle of the body.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.