Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,634
datasets available to search
ShareScore release 0.9.0
Dataset results
1,634 results for “Data integration”
Data from: Identification of a minority population of LMO2+ breast cancer cells that integrate into the vasculature and initiate metastasis.
Open the record for dataset details and reuse information.
Data for: Parasite prevalence depends on female preference: Integrating parasite-mediated sexual selection and infectious disease dynamics
Open the record for dataset details and reuse information.
Data from: Integrated SDM database: Enhancing the relevance and utility of species distribution models in conservation management
Open the record for dataset details and reuse information.
Data for empirical example in: An effect size for comparing the strength of morphological integration across studies
Open the record for dataset details and reuse information.
Scripts and data for: Integrating different facets of diversity into food web models: how adaptation among and within functional groups shape ecosystem functioning
Open the record for dataset details and reuse information.
Data from: Integrating effects of neighbor interactions for pollination
Open the record for dataset details and reuse information.
Data from: Displacement experiments provide evidence for path integration in Drosophila
Open the record for dataset details and reuse information.
Data from: Evaluating the importance of individual heterogeneity in reproduction to Weddell seal population dynamics using integral projection models
Open the record for dataset details and reuse information.
Data from: Dispersal in a house sparrow metapopulation: an integrative case study of genetic assignment calibrated with ecological data and pedigree information
Open the record for dataset details and reuse information.
Data for: An integrated population model and population viability assessment for the southern population of a data-poor species
Open the record for dataset details and reuse information.
GWAS summary statistics imputation support data and integration with PrediXcan MASHR
<p># GWAS summary statistics imputation, integration with PrediXcan MASHR-M</p> <p> </p> <p>The file `sample_data.tar` contains all necessary files to perform imputation of GWAS summary statistics to the GTEx v8 QTL data set.</p> <p>It includes 1000 Genomes individuals' genotypes as reference panel.</p> <p>The `.tar` archive, upon uncompression, contains the following folder structure:</p> <p>```</p> <p>data<br> |-- coordinate_map<br> |-- gwas<br> |-- liftover<br> |-- models<br> | |-- eqtl<br> | | `-- mashr<br> | `-- sqtl<br> | `-- mashr<br> |-- reference_panel_1000G<br> `-- ucsc</p> <p>```</p> <p> </p> <p>`data/eur_ld.bed.gz` contains definitions of approximately independent LD-regions in hg38 (Berisa-Pickrell regions, lifted over)</p> <p>`data/gtex_v8_eur_filtered_maf0.01_monoallelic_variants.txt.gz` is a snp annotation file, listing all GTEx v8 variants with MAF>0.01 in europeans.</p> <p>`data/coordinate_map` contains precomputed mapping tables that MetaXcan tools can use to convert GWAS' genomic coordinates in GWAS between genome assemblies.</p> <p>`data/gwas` contains a sample GWAS file for the purposes of a tutorial (data obtained from Nikpay et al (Nat Gen 2016) https://www.ncbi.nlm.nih.gov/pubmed/26343387</p> <p>`data/liftover` contains Liftover chains to map coordinates between human genome assemblies (used by full harmonization tools)</p> <p>`data/models` contains PrediXcan MASHR-M models, and cross-tissue S-MultiXcan LD compilation, from eQTL and sQTL.</p> <p>`data/reference_panel_1000G` contains 1000G hg38 genotypes, in parquet format, to be used by imputation tools.</p> <p>`data/ucsc` contains genomic coordinates of rsids in hg17, hg18 and hg19. You can use these to add chromosome and start position information to a GWAS based on its rsids. (column `end` is not used)</p> <p> </p>
Data from: Assessing species boundaries in the open sea: an integrative taxonomic approach to the pteropod genus Diacavolinia
To track changes in pelagic biodiversity in response to climate change, it is essential to accurately define species boundaries. Shelled pteropods are a group of holoplanktonic gastropods that have been proposed as bio-indicators because of their vulnerability to ocean acidification. A particularly suitable, yet challenging group for integrative taxonomy is the pteropod genus Diacavolinia, which has a circumglobal distribution and is the most species-rich pteropod genus, with 24 described species. We assessed species boundaries in this genus, with inferences based on geometric morphometric analyses of shell-shape variation, genetic (cytochrome c oxidase subunit I, 28S rDNA sequences) and geographic data. We found support for a total of 13 species worldwide, with observations of 706 museum and 263 freshly collected specimens across a global collection of material, including holo‐ and paratype specimens for 14 species. In the Atlantic Ocean, two species are well supported, in contrast to the eight currently described, and in the Indo‐Pacific we found a maximum of 11 species, partially merging 13 of the described species. Distributions of these revised species are congruent with well-known biogeographic provinces. Combining varied datasets in an integrative framework may be suitable for many diverse taxa and is an important first step to predicting species-specific responses to global change.
Data from: Estimating fish population abundance by integrating quantitative data on environmental DNA and hydrodynamic modeling
<p>Molecular analysis of DNA left in the environment, known as environmental DNA (eDNA), has proven to be a powerful and cost-effective approach to infer occurrence of species. Nonetheless, relating measurements of eDNA concentration to population abundance remains difficult because detailed knowledge on the processes that govern spatial and temporal distribution of eDNA should be integrated to reconstruct the underlying distribution and abundance of a target species. In this study, we propose a general framework of abundance estimation for aquatic systems on the basis of spatially replicated measurements of eDNA. The proposed method explicitly accounts for production, transport, and degradation of eDNA by utilizing numerical hydrodynamic models that can simulate the distribution of eDNA concentrations within an aquatic area. It turns out that, under certain assumptions, population abundance can be estimated via a Bayesian inference of a generalized linear model. Application to a Japanese jack mackerel (<em>Trachurus japonicus</em>) population in Maizuru Bay revealed that the proposed method gives an estimate of population abundance comparable to that of a quantitative echo sounder method. Furthermore, the method successfully identified a source of exogenous input of eDNA (a fish market), which may render a quantitative application of eDNA difficult to interpret unless its effect is taken into account. These findings indicate the ability of eDNA to reliably reflect population abundance of aquatic macroorganisms; when the "ecology of eDNA" is adequately accounted for, population abundance can be quantified on the basis of measurements of eDNA concentration.</p>
Data from: Integrative genomic analysis in African American children with asthma finds 3 novel loci associated with lung function
<p>Bronchodilator drugs are commonly prescribed for treatment and management of obstructive lung function present with diseases such as asthma. Administration of bronchodilator medication can partially or fully restore lung function as measured by pulmonary function tests. The genetics of baseline lung function measures taken prior to bronchodilator medication has been extensively studied, and the genetics of the bronchodilator response itself has received some attention. However, few studies have focused on the genetics of post-bronchodilator lung function. To address this gap, we analyzed lung function phenotypes in 1,103 subjects from the Study of African Americans, Asthma, Genes, and Environment (SAGE), a pediatric asthma case-control cohort, using an integrative genomic analysis approach that combined genotype, locus-specific genetic ancestry, and functional annotation information. We integrated genome-wide association study (GWAS) results with an admixture mapping scan of three pulmonary function tests (FEV1, FVC, and FEV1/FVC) taken before and after albuterol bronchodilator administration on the same subjects, yielding six traits. We identified 18 GWAS loci, and 5 additional loci from admixture mapping, spanning several known and novel lung function candidate genes. Most loci identified via admixture mapping exhibited wide variation in minor allele frequency across genotyped global populations. Functional fine-mapping revealed an enrichment of epigenetic annotations from peripheral blood mononuclear cells, fetal lung tissue, and lung fibroblasts. Our results point to three novel potential genetic drivers of pre- and post-bronchodilator lung function: ADAMTS1, RAD54B, and EGLN3. </p>
Data from: Defining a spectrum of integrative trait-based vegetation canopy structural types
Vegetation canopy structure is a fundamental characteristic of terrestrial ecosystems that defines vegetation types and drives ecosystem functioning. We use the multivariate structural trait composition of vegetation canopies to classify ecosystems within a global canopy structure spectrum. Across the temperate forest subset of this spectrum we assess gradients in canopy structural traits, characterize canopy structural types (CST), and evaluate drivers and functional consequences of canopy structural variation. We derive CSTs from multivariate canopy structure data, illustrating variation along three primary structural axes and resolution into six largely distinct and functionally relevant CSTs. Our results illustrate that within-ecosystem successional processes and disturbance legacies can produce variation in canopy structure similar to that associated with sub-continental variation in forest types and ecoclimatic zones. The potential to classify ecosystems into CSTs based on suites of structural traits represents an important advance in understanding and modeling structure-function relationships in vegetated ecosystems.
Data from: Integrating population genetics to define conservation units from the core to the edge of Rhinolophus ferrumequinum western range
The greater horseshoe bat (<i>Rhinolophus ferrumequinum</i>) is among the most widespread bat species in Europe but it has experienced severe declines, especially in Northern Europe. This species is listed Near Threatened in the European IUCN Red List of Threatened Animals and it is considered to be highly sensitive to human activities and particularly to habitat fragmentation. Therefore, understanding the population boundaries and demographic history of populations of this species is of primary importance to assess relevant conservation strategies. In this study, we used 17 microsatellite markers to assess the genetic diversity, the genetic structure and the demographic history of <i>R. ferrumequinum</i> colonies in the western part of its distribution. We identified one large population showing high levels of genetic diversity and large population size. Lower estimates were found in England and northern France. Analyses of clustering and isolation by distance suggested that the Channel and the Mediterranean seas could impede <i>R. ferrumequinum</i> gene flow. These results provide important information to improve the delineation of <i>R. ferrumequinum</i> management units. We suggest that a large management unit corresponding to the population ranging from Spanish Basque country to northern France must be considered. Particular attention should be given to mating territories as they seem to play a key role in maintaining the high levels of genetic mixing between colonies. Smaller management units corresponding to English and northern France colonies must also be implemented. These insular or peripheral colonies could be at higher risk of extinction in a near future.
Data from: Integration and harmonization of trait data from plant individuals across heterogeneous sources
<p>Trait data represent the basis for ecological and evolutionary research and have relevance for biodiversity conservation, ecosystem management and earth system modelling. The collection and mobilization of trait data has strongly increased over the last decade, but many trait databases still provide only species-level, aggregated trait values (e.g. ranges, means) and lack the direct observations on which those data are based. Thus, the vast majority of trait data measured directly from individuals remains hidden and highly heterogeneous, impeding their discoverability, semantic interoperability, digital accessibility and (re-)use. Here, we integrate quantitative measurements of verbatim trait information from plant individuals (e.g. lengths, widths, counts and angles of stems, leaves, fruits and inflorescence parts) from multiple sources such as field observations and herbarium collections. We develop a workflow to harmonize heterogeneous trait measurements (e.g. trait names and their values and units) as well as additional information related to taxonomy, measurement or fact and occurrence. This data integration and harmonization builds on vocabularies and terminology from existing metadata standards and ontologies such as the Ecological Trait-data Standard (ETS), the Darwin Core (DwC), the Thesaurus Of Plant characteristics (TOP) and the Plant Trait Ontology (TO). A metadata form filled out by data providers enables the automated integration of trait information from heterogeneous datasets. We illustrate our tools with data from palms (family Arecaceae), a globally distributed (pantropical), diverse plant family that is considered a good model system for understanding the ecology and evolution of tropical rainforests. We mobilize nearly 140,000 individual palm trait measurements in an interoperable format, identify semantic gaps in existing plant trait terminology and provide suggestions for the future development of a thesaurus of plant characteristics. Our work thereby promotes the semantic integration of plant trait data in a machine-readable way and shows how large amounts of small trait data sets and their metadata can be integrated into standardized data products.</p>
Predicting amphibian intraspecific diversity with machine learning: Challenges and prospects for integrating traits, geography, and genetic data
<p>The growing availability of genetic datasets, in combination with machine learning frameworks, offer great potential to answer long-standing questions in ecology and evolution. One such question has intrigued population geneticists, biogeographers, and conservation biologists: What factors determine intraspecific genetic diversity? This question is challenging to answer because many factors may influence genetic variation, including life history traits, historical influences, and geography, and the relative importance of these factors varies across taxonomic and geographic scales. Furthermore, interpreting the influence of numerous, potentially correlated variables is difficult with traditional statistical approaches. To address these challenges, we analyzed repurposed data using machine learning and investigated predictors of genetic diversity, focusing on Nearctic amphibians as a case study. We aggregated species traits, range characteristics, and >42,000 genetic sequences for 299 species using open-access scripts and various databases. After identifying important predictors of nucleotide diversity with random forest regression, we conducted follow-up analyses to examine the roles of phylogenetic history, geography, and demographic processes on intraspecific diversity. Although life history traits were not important predictors for this dataset, we found significant phylogenetic signal in genetic diversity within amphibians. We also found that salamander species at northern latitudes contain lower genetic diversity. Data repurposing and machine learning provide valuable tools for detecting patterns with relevance for conservation, but concerted efforts are needed to compile meaningful datasets with greater utility for understanding global biodiversity.</p>
Dataset accompanying paper submission for "Toward data-driven generation and evaluation of model structure for integrated representations of human behavior in water resources systems"
<p>This data set accompanies code archived at DOI: <a href="https://doi.org/10.5281/zenodo.3833186">10.5281/zenodo.3833186</a>, which was used in the experiments for the paper submission "Toward data-driven generation and evaluation of model structure for integrated representations of human behavior in water resources systems"</p>
Input data of manuscript "CACTUS: integrating clonal architecture with genomic clustering and transcriptome profiling of single tumor cells"
<p>This is the directory containing input data necessary to reproduce analyses presented in the manuscript:</p> <blockquote> <p><strong>CACTUS: integrating clonal architecture with genomic clustering and transcriptome profiling of single tumor cells</strong><br> Shadi Darvish Shafighi, Szymon M Kiełbasa, Julieta Sepúlveda Yáñez, Ramin Monajemi, Davy Cats, Hailiang Mei, Roberta Menafra, Susan Kloet, Hendrik Veelken, Cornelis A.M. van Bergen, Ewa Szczurek</p> </blockquote>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.