Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,634

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,634 results for “Data integration”

Learn how ShareScore rates datasets ↗
zenodo40/100

Figure 5a. from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063

Figure 5a. - Dashboard charts summarizing content from 37 open access articles published in Zootaxa and five articles published in Biodiversity Data Journal containing treatments on spiders (Suppl. materials 8, 9). These charts illustrate interoperability of data from XML-based publishing and subsequently marked up legacy literature.Figure 5a.All treatments regardless of taxonomic rank.Figure 5b.Species-rank treatments. <br> All treatments regardless of taxonomic rank.

opencc-by-4.0Feb 2017View details →
zenodo40/100

Figure 4b. from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063

Figure 4b. - Dashboard charts summarizing content from five articles published in Biodiversity Data Journal containing treatments on spiders (Suppl. materials 6, 7). Data elements were XML encoded as part of the routine publication process.Figure 4a.All treatments regardless of taxonomic rank.Figure 4b.Species-rank treatments. <br> Species-rank treatments.

opencc-by-4.0Feb 2017View details →
zenodo40/100

Figure 3a. from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063

Figure 3a. - Legacy literature dashboard: charts summarizing content from 37 open access articles published in Zootaxa containing treatments on spiders (Suppl. materials 4, 5). Data elements were XML encoded using GoldenGATE.Figure 3a.All treatments regardless of taxonomic rank.Figure 3b.Species-rank treatments. <br> All treatments regardless of taxonomic rank.

opencc-by-4.0Feb 2017View details →
zenodo40/100

Figure 2b. from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063

Figure 2b. - Number of publications (a; Suppl. material 2) and treatments (b; Suppl. material 3) contributing to spider taxonomy sorted by number of treatments per publication source (e.g., journal, book publisher). Source: the World Spider Catalog, accessed 14 October 2014.Figure 2a.12,377 publications listed in the 2014 World Spider Catalog. Zootaxa is the top ranking venue with 509 titles representing just over 4% of spider taxonomy.Figure 2b.The 2014 World Spider Catalog refers to 126,621 treatments. With 3314 treatments (2.6%), Zootaxa is the third ranking all time venue behind two museum monograph series: Bulletin of the American Museum of Natural History (4537 treatments, 3.6%) and Harvard's Bulletin of the Museum of Comparative Zoology (3761 treatments, 3.0%). <br> The 2014 World Spider Catalog refers to 126,621 treatments. With 3314 treatments (2.6%), Zootaxa is the third ranking all time venue behind two museum monograph series: Bulletin of the American Museum of Natural History (4537 treatments, 3.6%) and Harvard's Bulletin of the Museum of Comparative Zoology (3761 treatments, 3.0%).

opencc-by-4.0Feb 2017View details →
zenodo40/100

Figure 2a. from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063

Figure 2a. - Number of publications (a; Suppl. material 2) and treatments (b; Suppl. material 3) contributing to spider taxonomy sorted by number of treatments per publication source (e.g., journal, book publisher). Source: the World Spider Catalog, accessed 14 October 2014.Figure 2a.12,377 publications listed in the 2014 World Spider Catalog. Zootaxa is the top ranking venue with 509 titles representing just over 4% of spider taxonomy.Figure 2b.The 2014 World Spider Catalog refers to 126,621 treatments. With 3314 treatments (2.6%), Zootaxa is the third ranking all time venue behind two museum monograph series: Bulletin of the American Museum of Natural History (4537 treatments, 3.6%) and Harvard's Bulletin of the Museum of Comparative Zoology (3761 treatments, 3.0%). <br> 12,377 publications listed in the 2014 World Spider Catalog. Zootaxa is the top ranking venue with 509 titles representing just over 4% of spider taxonomy.

opencc-by-4.0Feb 2017View details →
zenodo40/100

Figure 1a. from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063

Figure 1a. - GBIF records proportioned by selected taxonomic groups (Suppl. material 1). Inner ring shows vertebrates, insects, arachnids, other animals, plants, and other kingdoms; outer ring separates birds from other vertebrates, Hymenoptera from other insects, and spiders from other arachnids.Figure 1a.All records in GBIF (n = 517,325,595).Figure 1b.Specimen-based records in GBIF (n = 98,144,242). <br> All records in GBIF (n = 517,325,595).

opencc-by-4.0Feb 2017View details →
zenodo40/100

Figure 3b. from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063

Figure 3b. - Legacy literature dashboard: charts summarizing content from 37 open access articles published in Zootaxa containing treatments on spiders (Suppl. materials 4, 5). Data elements were XML encoded using GoldenGATE.Figure 3a.All treatments regardless of taxonomic rank.Figure 3b.Species-rank treatments. <br> Species-rank treatments.

opencc-by-4.0Feb 2017View details →
zenodo40/100

Figure 1b. from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063

Figure 1b. - GBIF records proportioned by selected taxonomic groups (Suppl. material 1). Inner ring shows vertebrates, insects, arachnids, other animals, plants, and other kingdoms; outer ring separates birds from other vertebrates, Hymenoptera from other insects, and spiders from other arachnids.Figure 1a.All records in GBIF (n = 517,325,595).Figure 1b.Specimen-based records in GBIF (n = 98,144,242). <br> Specimen-based records in GBIF (n = 98,144,242).

opencc-by-4.0Feb 2017View details →
zenodo40/100

Supplementary material 15: Species dashboard: Tenuiphantes tenuis from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063

Dashboard charts showing only specimens of the linyphiid spider Tenuiphantes tenuis. When viewed using a browser (such as Google Chrome) with an internet connection, this page sends a series of queries to Plazi and integrates the results with the Google Charts API to produce 37 interactive dashboard charts.

opencc-by-4.0Feb 2017View details →
zenodo40/100

Supplementary material 7: Prospective publishing dashboard: species-rank treatments from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063

Dashboard charts summarizing content from 5 articles published in Biodiversity Data Journal containing treatments on spiders. This page shows data from species-rank treatments. When viewed using a browser (such as Google Chrome) with an internet connection, this page sends a series of queries to Plazi and integrates the results with the Google Charts API to produce 37 dashboard charts.

opencc-by-4.0Feb 2017View details →
zenodo40/100

Supplementary material 14: Treatment dashboard: content from Pardosa zyuzini treatment in Kronestedt and Marusik (2011) from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063

Dashboard charts showing content from one treatment: Pardosa zyuzini in Kronestedt and Marusik (2011). When viewed using a browser (such as Google Chrome) with an internet connection, this page sends a series of queries to Plazi and integrates the results with the Google Charts API to produce 37 interactive dashboard charts.

opencc-by-4.0Feb 2017View details →
zenodo40/100

Supplementary material 13: Article dashboard: content from Kronestedt and Marusik (2011) from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063

Dashboard charts showing content from one article, Kronestedt and Marusik 2011. This page shows data from all treatments. When viewed using a browser (such as Google Chrome) with an internet connection, this page sends a series of queries to Plazi and integrates the results with the Google Charts API to produce 37 interactive dashboard charts.

opencc-by-4.0Feb 2017View details →
zenodo40/100

Supplementary material 6: Prospective publishing dashboard: all treatments from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063

Dashboard charts summarizing content from 5 articles published in Biodiversity Data Journal containing treatments on spiders. This page shows data from all treatments. When viewed using a browser (such as Google Chrome) with an internet connection, this page sends a series of queries to Plazi and integrates the results with the Google Charts API to produce 37 dashboard charts.

opencc-by-4.0Feb 2017View details →
zenodo40/100

Figure 4a. from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063

Figure 4a. - Dashboard charts summarizing content from five articles published in Biodiversity Data Journal containing treatments on spiders (Suppl. materials 6, 7). Data elements were XML encoded as part of the routine publication process.Figure 4a.All treatments regardless of taxonomic rank.Figure 4b.Species-rank treatments. <br> All treatments regardless of taxonomic rank.

opencc-by-4.0Feb 2017View details →
zenodo40/100

Integrated field-aligned radar data and analysis results

<p>Dataset used in &quot;A statistical survey of heat input parameters into the cusp thermosphere&quot; J. Geophys. Res. 2017, doi:10.1002/2016JA023594.</p>

opencc-by-4.0Apr 2017View details →
zenodo40/100

SleepEEGpy: a Python-based software integration package to organize preprocessing, analysis, and visualization of sleep EEG data

<p>This dataset includes three high-density sleep EEG recordings of healthy participants, downsampled to 250 Hz and stored in FIF format:</p> <ol> <li>Nap recording of a young adult participant</li> <li>Overnight recording of a young adult participant</li> <li>Overnight recording of an older adult participant</li> </ol> <p>Additionally, the dataset includes three text files for each recording:</p> <ul> <li>bad_channels.txt: Indexes of noisy channels</li> <li>annotations.txt: Onset and duration of noisy temporal intervals</li> <li>staging.txt: Sleep staging vector</li> </ul> <p>The corresponding package can be found&nbsp;on <a href="https://github.com/NirLab-TAU/sleepeegpy">GitHub.</a></p> <p>For citation, please use:<br>Falach, R., G. Belonosov, J. F. Schmidig, M. Aderka, V. Zhelezniakov, R. Shani-Hershkovich, E. Bar, and Y. Nir. "SleepEEGpy: a Python-based software integration package to organize preprocessing, analysis, and visualization of sleep EEG data." Computers in Biology and Medicine 192 (2025): 110232.<br><a href="https://doi.org/10.1016/j.compbiomed.2025.110232" rel="nofollow">https://doi.org/10.1016/j.compbiomed.2025.110232</a></p>

opencc-by-4.0Dec 2023View details →
dryad40/100

Data from: Integrated species distribution models to account for sampling biases and improve range wide occurrence predictions

<p><strong><span>Aim</span></strong></p> <p><span>Species distribution models (SDMs) that integrate presence-only and presence-absence data offer a promising avenue to improve information on species' geographic distributions. The use of such 'integrated SDMs' on a species range-wide extent has been constrained by the often-limited presence-absence data and by the heterogeneous sampling of the presence-only data. Here, we evaluate integrated SDMs for studying species ranges with a novel expert range map-based evaluation. We build a new understanding about how integrated SDMs address issues of estimation accuracy and data deficiency and thereby offer advantages over traditional SDMs.</span></p> <p><strong><span>Location</span></strong></p> <p><span>South and Central America.</span></p> <p><strong><span>Time period</span></strong></p> <p><span>1979-2017.</span></p> <p><strong><span>Major taxa studied</span></strong></p> <p><span>Hummingbirds.</span></p> <p><strong><span>Methods</span></strong></p> <p><span>We build integrated SDMs by linking two observation models – one for each data type – to the same underlying spatial process.</span> <span>We validate SDMs with two schemes: i) cross-validation with presence-absence data and ii) comparison with respect to the species' whole range as defined with IUCN range maps. We also compare models relative to the estimated response curves and compute the association between the benefit of the data integration and the number of presence records in each data set.</span></p> <p><strong><span>Results</span></strong></p> <p><span>The integrated SDM accounting for the spatially varying sampling intensity of the presence-only data was one of the top-performing models in both model validation schemes. Presence-only data alleviated overly large niche estimates, and data integration was beneficial compared to modelling solely presence-only data for species that had few presence points when predicting the species' whole range. On the community level, integrated models improved the species richness prediction.</span></p> <p><strong><span>Main conclusions</span></strong></p> <p><span>Integrated SDMs combining presence-only and presence-absence data are successfully able to borrow strengths from both data types and offer improved predictions of species' ranges. Integrated SDMs can potentially alleviate the impacts of taxonomically and geographically uneven sampling and to leverage the detailed sampling information in presence-absence data.</span></p>

opencc-zeroNov 2023View details →
dryad40/100

Data from: harnessing the power of regional baselines for broad-scale genetic stock identification: a multistage, integrated, and cost-effective approach

<p>In mixed-stock fishery analyses, genetic stock identification (GSI) estimates the contribution of each population to a mixture and is typically conducted at a regional scale using genetic baselines specific to the stocks expected in that region. Often these regional baselines cannot be combined to produce broader geographical baselines due to non-overlapping populations and genetic markers. In cases where the mixture contains stocks spanning across a wide area, a broad-scale baseline is created, but often at the cost of resolution. Here, we introduce a new GSI method to harness the resolution capabilities of baselines developed for regional applications in the analysis of mixtures containing individuals from a broad geographic range. This method employs a multistage framework that allows disparate baselines to be used in a single integrated process that produces estimates along with the propagated errors from each stage. All individuals in the mixture sample are required to be genotyped for all genetic markers in the baselines used by this model, but the baselines do not require overlap in genetic markers or populations representing the broad-scale or regional baselines.</p> <p>We demonstrate our integrated multistage GSI model using a synthesized data set made up of Chinook salmon, <em>Oncorhynchus tshawytscha</em>, from the North Bering Sea of Alaska. The data set is designed to be run using R package, Ms.GSI, and it does not represent the composition of the real fishery. The results show an improved accuracy for estimates using an integrated multistage framework, compared to the conventional framework of using separate hierarchical steps. The integrated multistage framework allows GSI of a wide geographic area without first developing a large scale, high-resolution genetic baseline or dividing a mixture sample into smaller regions beforehand. This approach is more cost-effective than updating range-wide baselines with all regionally important markers.</p>

opencc-zeroDec 2023View details →
zenodo40/100

From Pixels to Phenotypes: Integrating Image-Based Profiling with Cell Health Data Improves Interpretability

<p>Code: https://github.com/srijitseal/BioMorph_Space<br> <br> Cell Painting assays generate morphological profiles that are versatile descriptors of biological systems and have been used to predict <em>in vitro</em> and <em>in vivo</em> drug effects. However, Cell Painting features are based on image statistics, and are, therefore, often not readily biologically interpretable. In this study, we introduce an approach that maps specific Cell Painting features into the BioMorph space using readouts from comprehensive Cell Health assays. We validated that the resulting BioMorph space effectively connected compounds not only with the morphological features associated with their bioactivity but with deeper insights into phenotypic characteristics and cellular processes associated with the given bioactivity. The BioMorph space revealed the mechanism of action for individual compounds, including dual-acting compounds such as emetine, an inhibitor of both protein synthesis and DNA replication. In summary, BioMorph space offers a more biologically relevant way to interpret cell morphological features from the Cell Painting assays and to generate hypotheses for experimental validation.</p> <p>&nbsp;</p> <p>The following datasets are released:<br> &nbsp;</p> <p>Cell_Health_median_357_profiles_70_labels.csv :<br> The Cell Heath dataset for CRISPR perturbations.&nbsp;Contains&nbsp;median consensus signatures for the 357 consensus profiles (119 CRISPR perturbations &times; 3 cell lines) Ref: Way et al.</p> <p>Cell_Painitng_CRISPR_Perturbations_357_profiles_827_features_scaled.csv:<br> The Cell Painting dataset for CRISPR perturbations.&nbsp;Contains 827 morphology features (and metadata annotation) for 357 consensus profiles (119 CRISPR perturbations &times; 3 cell lines).&nbsp;Ref: Way et al.</p> <p>Cell_Painting_data_658_compounds_827_Features_scaled.csv<br> The Cell Painting dataset for compound perturbations.&nbsp;Contains 658 structurally unique compounds with 827 Cell Painting features. Ref: Bray et al</p> <p>Endpoints_9_Mitotox_biological_activities_658_compounds.csv<br> The biological assay activity labels&nbsp;for compound perturbations.&nbsp;Contains 658 structurally unique compounds with&nbsp;9 biological activity consensus hit calls.&nbsp;Ref: ToxCast/MoleculeNet</p> <p>BioMoprh_pvalue_658_compunds_398_BioMorph_terms.csv:<br> The dataset of standardised BioMorph term p-values. Contains&nbsp;398 BioMorph terms for the 658 compounds in the biological activity dataset.&nbsp;<br> <br> References:&nbsp;<br> Way et al. Predicting cell health phenotypes using image-based morphology profiling. Mol Biol Cell. 2021;32(9):995-1005.<br> Bray et al. A dataset of images and morphological profiles of 30 000 small-molecule treatments using the Cell Painting assay. Gigascience. 2017;6(12):1-5.&nbsp;<br> MoleculeNet: Wu&nbsp;et al. MoleculeNet: A benchmark for molecular machine learning. Chem Sci. 2018;9(2):513-530.&nbsp;<br> ToxCast:&nbsp;Exploring ToxCast Data | US EPA https://www.epa.gov/chemical-research/exploring-toxcast-data (accessed Jul 9, 2023).</p>

opencc-by-4.0Jan 2024View details →
dryad40/100

Data from: Exponential history integration with diverse temporal scales in retrosplenial cortex supports hyperbolic behavior

<p>Animals rely on their experience to guide their next choice. In foraging-type tasks guided by history-dependent value, these experiences are typically integrated such that the weights of past events initially decay quickly over time but show a longer tail than expected by exponential decay. Rather, such integration is better described by a hyperbolic function. Hyperbolic integration affords sensitivity to both recent environmental dynamics and long-term trends, however the mechanism by which the brain implements this hyperbolic integration is unknown. We trained mice on a history-dependent, value-based decision task and found that the mice indeed showed hyperbolic decay on their weighting of past experience. However, the activity of history-encoding cortical neurons showed weighting with exponential decay. In resolving this apparent mismatch, we observed that cortical neurons encode history information heterogeneously across a wide variety of exponential time-constants, with the retrosplenial cortex (RSC) overrepresenting longer time-constants compared to other areas. A model that combines these diverse timescales of exponential history integration can recreate the heavy-tailed, hyperbolic history integration observed in behavior. In particular, time-constants of RSC neurons best matched the behavior, and optogenetic inactivation of RSC uniquely reduced the use of history information. These results indicate that behavior-relevant history information is maintained in neurons across multiple timescales in parallel, and suggest that the neural population in RSC is a critical reservoir of this information guiding decision-making.</p>

opencc-zeroJan 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record