Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

9

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

9 results for “macroecological modelling”

Learn how ShareScore rates datasets ↗
dryad36/100

Data from: Influence of different data cleaning solutions of point-occurrence records on downstream macroecological diversity models

<p><span>Digital point-occurrence records from the Global Biodiversity Information Facility (GBIF) and other data providers enable a wide range of research in macroecology and biogeography. However, data errors may hamper immediate use. Manual data cleaning is time-consuming and often unfeasible, given that the databases may contain thousands or millions of records. Automated data cleaning pipelines are therefore of high importance. This study examined the extent to which cleaned data from six pipelines using data cleaning tools (e.g., the GBIF web application, different R packages) affect downstream species distribution models. In addition, we assessed how the pipeline data differ from expert data. From 13,889 North American <i>Ephedra</i> observations in GBIF, the pipelines removed 31.7% to 62.7% false-positives, invalid coordinates, and duplicates, leading to data sets that included between 9,484 (GBIF application) and 5,196 records (manual-guided filtering). The expert data consisted of 703 thoroughly handpicked records, comparable to data from field studies. Although differences in the record numbers were relatively large, stacked species distribution models (sSDM) from the pipelines and the expert data were strongly related (mean Pearson's <i>r</i> across the pipelines: 0.9986, versus the expert data: 0.9173). The ever-stronger correlations resulted from occurrence information that became increasingly condensed in the course of the workflow (from individual occurrences to collectivized occurrences in grid cells to predicted probabilities in the sSDMs). In sum, our results suggest that the <i>R</i> package-based pipelines reliably identified invalid coordinates. In contrast, the GBIF-filtered data still contained both spatial and taxonomic errors. However, major drawbacks emerge from the fact that no pipeline fully discovered misidentified specimens without the assistance of expert taxonomic knowledge. We conclude that application-filtered GBIF data will still need additional review to achieve higher spatial data quality. Achieving high-quality taxonomic data will require extra effort, probably by thoroughly analyzing the data for misidentified taxa, supported by experts.</span></p>

opencc-zeroJul 2022View details →
dryad36/100

Data from: Influence of different data cleaning solutions of point-occurrence records on downstream macroecological diversity models

Open the record for dataset details and reuse information.

publicJul 2022View details →
dryad32/100

Data from: Predicting spatial patterns of plant species richness: a comparison of direct macroecological and species stacking modelling approaches

PLEASE NOTE, THESE DATA ARE ALSO REFERRED TO IN TWO OTHER PUBLICATIONS. PLEASE SEE http://dx.doi.org/10.1111/j.1365-2486.2008.01766.x AND http://dx.doi.org/10.1111/2041-210X.12222 FOR MORE INFORMATION. Aim: This study compares the direct, macroecological approach (MEM) for modelling species richness (SR) with the more recent approach of stacking predictions from individual species distributions (S-SDM). We implemented both approaches on the same dataset and discuss their respective theoretical assumptions, strengths and drawbacks. We also tested how both approaches performed in reproducing observed patterns of SR along an elevational gradient. Location: Two study areas in the Alps of Switzerland. Methods: We implemented MEM by relating the species counts to environmental predictors with statistical models, assuming a Poisson distribution. S-SDM was implemented by modelling each species distribution individually and then stacking the obtained prediction maps in three different ways – summing binary predictions, summing random draws of binomial trials and summing predicted probabilities – to obtain a final species count. Results: The direct MEM approach yields nearly unbiased predictions centred around the observed mean values, but with a lower correlation between predictions and observations, than that achieved by the S-SDM approaches. This method also cannot provide any information on species identity and, thus, community composition. It does, however, accurately reproduce the hump-shaped pattern of SR observed along the elevational gradient. The S-SDM approach summing binary maps can predict individual species and thus communities, but tends to overpredict SR. The two other S-SDM approaches – the summed binomial trials based on predicted probabilities and summed predicted probabilities – do not overpredict richness, but they predict many competing end points of assembly or they lose the individual species predictions, respectively. Furthermore, all S-SDM approaches fail to appropriately reproduce the observed hump-shaped patterns of SR along the elevational gradient. Main conclusions: Macroecological approach and S-SDM have complementary strengths. We suggest that both could be used in combination to obtain better SR predictions by following the suggestion of constraining S-SDM by MEM predictions.

opencc-zeroDec 2013View details →
zenodo32/100

FIGURE 3 in Molecules meet macroecology—combining Species Distribution Models and phylogeographic studies

FIGURE 3. Potential distribution of Arthroleptis xenodactyloides under current climatic and two proposed Last Glacial Maximum palaeoclimatic scenarios (CCSM, MIROC) showing mean values obtained from 10 models computed with randomly selected 30 % of the 46 records (black dots) for model evaluation and the remaining 70 % for model training and corresponding standard deviations (SD).

opennotspecifiedDec 2010View details →
zenodo32/100

FIGURE 2 in Molecules meet macroecology—combining Species Distribution Models and phylogeographic studies

FIGURE 2. (A) Model performance and (B) presence/absence thresholds obtained from 10 models computed with randomly selected 30 % of the 46 records for model evaluation and the remaining 70 % for model training; as Maxent values the logistic model output is chosen.

opennotspecifiedDec 2010View details →
zenodo32/100

FIGURE 1 in Molecules meet macroecology—combining Species Distribution Models and phylogeographic studies

FIGURE 1. Elevation map of part of coastal eastern Africa showing records of Arthroleptis xenodactyloides (dots) processed in this study; those of Blackburn &amp; Measey (2009) are indicated by name.

opennotspecifiedDec 2010View details →
dryad32/100

Data from: Predicting spatial patterns of plant species richness: a comparison of direct macroecological and species stacking modelling approaches

Open the record for dataset details and reuse information.

publicJul 2014View details →
dryad28/100

Data from: Comparing process-based and constraint-based approaches for modeling macroecological patterns

Ecological patterns arise from the interplay of many different processes, and yet the emergence of consistent phenomena across a diverse range of ecological systems suggests that many patterns may in part be determined by statistical or numerical constraints. Differentiating the extent to which patterns in a given system are determined statistically, and where it requires explicit ecological processes, has been difficult. We tackled this challenge by directly comparing models from a constraint-based theory, the Maximum Entropy Theory of Ecology (METE) and models from a process-based theory, the size-structured neutral theory (SSNT). Models from both theories were capable of characterizing the distribution of individuals among species and the distribution of body size among individuals across 76 forest communities. However, the SSNT models consistently yielded higher overall likelihood, as well as more realistic characterizations of the relationship between species abundance and average body size of conspecific individuals. This suggests that the details of the biological processes contain additional information for understanding community structure that are not fully captured by the METE constraints in these systems. Our approach provides a first step towards differentiating between process- and constraint-based models of ecological systems and a general methodology for comparing ecological models that make predictions for multiple patterns.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Comparing process-based and constraint-based approaches for modeling macroecological patterns

Open the record for dataset details and reuse information.

publicDec 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record