Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
403
datasets available to search
ShareScore release 0.9.0
Dataset results
403 results for “Occurrence Data”
Data from: Do traits of plant species predict the efficacy of species distribution models for finding new occurrences?
Open the record for dataset details and reuse information.
Data from: Predicting species occurrences with habitat network models
Open the record for dataset details and reuse information.
Data from: On the occurrence of three non-native cichlid species including the first record of a feral population of Pelmatolapia (Tilapia) mariae (Boulenger, 1899) in Europe
Open the record for dataset details and reuse information.
Data from: Habitat monitoring and projections for threatened Canada lynx: linking the Landsat archive with carnivore occurrence and prey density
Open the record for dataset details and reuse information.
Data from: Species interactions limit the occurrence of urban-adapted birds in cities
Open the record for dataset details and reuse information.
Data from: SpeciesGeoCoder: fast categorization of species occurrences for analyses of biodiversity, biogeography, ecology and evolution
Open the record for dataset details and reuse information.
Data from: Hide and seek in vegetation: time-to-detection is an efficient design for estimating detectability and occurrence
Open the record for dataset details and reuse information.
Data from: Modeling and mapping the probability of occurrence of invasive wild pigs across the contiguous United States
Open the record for dataset details and reuse information.
Moose occurrence data in Bohemian Forest Ecosystem
Open the record for dataset details and reuse information.
Data from: Recognizing pulses of extinction from clusters of last occurrences
Open the record for dataset details and reuse information.
Filling gaps of occurrence data for Colombia: case of birds
<p>Biological databases from collections, museums, and inventories present gaps and biases as some species and geographical areas remain poorly known. Biological expeditions are thus crucial to increase the knowledge of biodiversity and should be planned carefully. Using 1-km2 spatial resolution, this dataset provides which locations of Colombia that have been surveyed at least one time for the bird group.</p> <p><br> The first two columns describe the latitudinal (lat) and longitudinal (long) coordinate of the locations. Meanwhile, the last column (observed) describes whether the location has been surveyed at least one time, e.g., if the location has at least a point record, its value is 1; otherwise, its value is 0.</p>
Assigning occurrence data to cryptic taxa improves climatic niche assessments: biodecrypt, a new tool tested on European butterflies
<p><b><span>Aim</span></b><br> <span>Occurrence data are fundamental to macroecology, but accuracy is often compromised when multiple units are lumped together (e.g. in recently separated cryptic species or citizen science records). Using amalgamated data leads to inaccuracy in species mapping, to biased beta-diversity assessments and to potentially erroneous</span><span>ly</span><span> predicted responses to climate change. We provide a set of R functions (biodecrypt) to objectively attribute undetermined occurrences to the most probable taxon based on a subset of identified records.</span></p> <p><b><span>Innovation</span></b><br> <span>Biodecrypt assumes </span><span>that unknown occurrences can only be attributed at certain distances from </span><span>areas of </span><span>sympatry. </span><span>The </span><span>function draws concave hulls based on the subset of identified records; subsequently, based on hull geometry, it attributes (or not) unknown records to a given taxon. Concavity can be imposed with an alpha value and sea or land areas can be excluded. A cross-validation function tests attribution reliability and another function optimizes the parameters (alpha, buffer, distance ratio between hulls). We applied the procedure to 16 European butterfly complexes recently separated into 33 cryptic species for which most records were amalgamated. We compared niche similarity and divergence between cryptic taxa, and we re-calculated and </span><span>contributed </span><span>updated </span><span>CLIMBER variables for climatic preferences</span><span>.</span></p> <p><b><span>Main conclusions</span></b><br> Biodecrypt showed a cross-validated correct attribution of known records always ≥98% and attributed more than 80% of unknown records to the most likely taxon in parapatric species. The functions determined where records can be assigned even for largely sympatric species, and highlighted areas where further sampling is required. All the cryptic taxa <span>showed significantly diverging climatic niches, </span>reflected in different values of mean temperature and precipitation compared to the values originally provided in the CLIMBER database. The substantial fraction of cryptic taxa existing across different taxonomic groups and their divergence in climatic niches highlights the importance of using reliably assigned occurrence data in macroecology.</p>
Data from: The role of human outdoor recreation in shaping patterns of grizzly bear-black bear co-occurrence
Species distributions are influenced by a combination of landscape variables and biotic interactions with other species, including people. Grizzly bears and black bears are sympatric, competing omnivores that also share habitats with human recreationists. By adapting models for multi-species occupancy analysis, we analyzed trail camera data from 192 trail camera locations in and around Jasper National Park, Canada to estimate grizzly bear and black bear occurrence and intensity of trail use. We documented (a) occurrence of grizzly bears and black bears relative to habitat variables (b) occurrence and intensity of use relative to competing bear species and motorised and non-motorised recreational activity, and (c) temporal overlap in activity patterns among the two bear species and recreationists. Grizzly bears were spatially separated from black bears, selecting higher elevations and locations farther from roads. Both species co-occurred with motorised and non-motorised recreation, however, grizzly bears reduced their intensity of use of sites with motorised recreation present. Black bears showed higher temporal activity overlap with recreational activity than grizzly bears, however differences in bear daily activity patterns between sites with and without motorised and non-motorised recreation were not significant. Reduced intensity of use by grizzly bears of sites where motorised recreation was present is a concern given off-road recreation is becoming increasingly popular in North America, and can negatively influence grizzly bear recovery by reducing foraging opportunities near or on trails. Camera traps and multi-species occurrence models offer non-invasive methods for identifying how habitat use by animals changes relative to sympatric species, including humans. These conclusions emphasise the need for integrated land-use planning, access management, and grizzly bear conservation efforts to consider the implications of continued access for motorised recreation in areas occupied by grizzly bears.
Data from: Inferring new relations between medical entities using literature curated term co-occurrences
ABSTRACT Objectives Identifying new relations between medical entities, such as drugs, diseases, and side-effects, is typically a resource-intensive task, involving experimentation and clinical trials. The increased availability of related data and curated knowledge enables a computational approach to this task, notably by training models to predict likely relations. Such models rely on meaningful representations of the medical entities being studied. We propose a generic features vector representation that leverages co-occurrences of medical terms, linked with PubMed citations. Materials and Methods We demonstrate the usefulness of the proposed representation by inferring two types of relations: a drug causes a side effect, and a drug treats an indication. To predict these relations and assess their effectiveness, we applied two modeling approaches: multi-task modeling using neural networks, and single-task modeling based on gradient-boosting machines and logistic regression. Results These trained models, which predict either side effects or indications, obtained significantly better results than baseline models that use a single direct co-occurrence feature. The results demonstrate the advantage of a comprehensive representation. Discussion Selecting the appropriate representation has an immense impact on the predictive performance of machine learning models. Our proposed representation is powerful, as it spans multiple medical domains and can be used to predict a wide range of relation types. Conclusion The discovery of new relations between various medical entities can be translated into meaningful insights, for example, related to drug development or disease understanding. Our representation of medical entities can be used to train models that predict such relations, thus accelerating healthcare-related discoveries.
Data from: Phylogenomic analysis of transcriptome data elucidates co-occurrence of a paleopolyploid event and the origin of bimodal karyotypes in Agavoideae (Asparagaceae)
PREMISE OF THE STUDY: The stability of the bimodal karyotype found in Agave and closely related species has long interested botanists. The origin of the bimodal karyotype has been attributed to allopolyploidy, but this hypothesis has not been tested. Next Generation transcriptome sequence data were used to test whether a paleopolyploid event occurred on the same branch of the Agavoideae phylogenetic tree as the origin of the Yucca-Agave bimodal karyotype. METHODS: Illumina RNAseq data were generated for phylogenetically strategic species in Agavoideae. Paleopolyploidy was inferred in analyses of frequency plots for synonymous substitutions per synonymous site (Ks) between Hosta, Agave and Chlorophytum paralogous and orthologous gene pairs. Phylogenies of gene families including paralogous genes for these species and outgroup species were estimated in order to place inferred paleopolyploid events on a species tree. KEY RESULTS: Ks frequency plots suggested paleopolyploid events in the history of the genera Agave, Hosta and Chlorophytum. Phylogenetic analyses of gene families estimated from transcriptome data revealed two polyploid events: one predating the last common ancestor of Agave and Hosta and one within the lineage leading to Chlorophytum. CONCLUSIONS: We found that allopolyoidy and the origin of the Yucca-Agave bimodal karyotype co-occur on the same lineage consistent with the hypothesis that the bimodal karyotype is a consequence of allopolyploidy. We discuss this and alternative mechanisms for the formation of the Yucca-Agave bimodal karyotype. More generally, we illustrate how the use of next generation sequencing technology is a cost-efficient means for assessing genome evolution in non-model species.
Data from: Occurrence, costs and heritability of delayed selfing in a free-living flatworm
Evolutionary theory predicts that in the absence of outcrossing opportunities, simultaneously hermaphroditic organisms should eventually switch to self-fertilization as a form of reproductive assurance. Here we report the existence of facultative self-fertilization in the free-living flatworm Macrostomum hystrix, a species in which outcrossing occurs via hypodermic insemination of sperm into the parenchyma of the mating partner. First, we show that isolated individuals significantly delay the onset of reproduction compared to individuals with outcrossing opportunities ("delayed selfing") as predicted by theory. Second, consistent with the idea of M. hystrix being a preferential outcrosser under natural conditions, we report likely costs of selfing manifested via reduced hatchling production and offspring survival. Third, we demonstrate that selfing propensity has a genetic basis in this species, with a heritability estimated at 0.43 ± 0.11. Variation in selfing propensity could arise due to differing costs of inbreeding among families, though despite marked inter-family variation in apparent costs of inbreeding we found no evidence for such a link. Alternatively, selfing propensity might differ across families because of heritable variation in reproductive traits that determine the likelihood of selfing. We speculate that adaptations to hypodermic insemination under outcrossing, most notably a highly modified copulatory stylet (male copulatory organ) and reduced sperm complexity, could also facilitate facultative selfing in this species.
Data from: Bayesian estimation of speciation and extinction from incomplete fossil occurrence data
The temporal dynamics of species diversity are shaped by variations in the rates of speciation and extinction, and there is a long history of inferring these rates using first and last appearances of taxa in the fossil record. Understanding diversity dynamics critically depends on unbiased estimates of the unobserved times of speciation and extinction for all lineages, but the inference of these parameters is challenging due to the complex nature of the available data. Here, we present a new probabilistic framework to jointly estimate species-specific times of speciation and extinction and the rates of the underlying birth-death process based on the fossil record. The rates are allowed to vary through time independently of each other, and the probability of preservation and sampling is explicitly incorporated in the model to estimate the true lifespan of each lineage. We implement a Bayesian algorithm to assess the presence of rate shifts by exploring alternative diversification models. Tests on a range of simulated data sets reveal the accuracy and robustness of our approach against violations of the underlying assumptions and various degrees of data incompleteness. Finally, we demonstrate the application of our method with the diversification of the mammal family Rhinocerotidae and reveal a complex history of repeated and independent temporal shifts of both speciation and extinction rates, leading to the expansion and subsequent decline of the group. The estimated parameters of the birth-death process implemented here are directly comparable with those obtained from dated molecular phylogenies. Thus, our model represents a step towards integrating phylogenetic and fossil information to infer macroevolutionary processes.
Data from: Taxon abundance, diversity, co-occurrence and network analysis of the ruminal microbiota in response to dietary changes in dairy cows
The effects of sunflower oil (SO) (0 or 50 g/kg diet dry matter), supplemented to diets contrasting in the proportion of forage and concentrate (FC) (65:35 vs 35:65), were evaluated for their influence on rumen microbiome. Four multiparous Nordic Red dairy cows fitted with rumen cannulae were used in a 4 × 4 Latin square with a 2 × 2 factorial arrangement of treatments and four 35-d periods. Ruminal digesta samples were collected on d 22 and d 24 of each experimental period and DNA was extracted from a combined sample. Diet effect on rumen microbial community was explored by qPCR, T-RFLP and metabarcoding sequencing. QPCR analysis showed that the total amounts of bacteria, archaea or ciliate protozoa were not significantly altered either by FC ratio or addition of SO. Only fungi were reduced by half in high concentrate (H) compared to high forage (L) diets (P=0.03). Further significant reduction of fungi was observed due to SO in both HSO and LSO diets but the effect was stronger in HSO (H vs HSO by 10.5x, P=0.03; L vs LSO by 1.9x, P=0.04). Metabarcoding sequencing analysis showed that SO affected bacterial, archaeal, ciliate protozoa and fungal community structure and diversity and the effect was FC ratio dependent. As expected, Simpson's index of diversity was higher in diets containing higher proportion of forage. These diets were dominated by Firmicutes, while Bacteroidetes and Proteobacteria were more abundant in H diets. Methanobrevibacter ruminantium and Methanobrevibacter gottschalkii dominated archaea community but they were in negative relation to each other. Methanobrevibacter gottschalkii was more abundant in L, while Methanobrevibacter ruminantium in H diet. The strongest diet effect was observed on fungal community, represented by both well classified and novel fungal groups. Both, increase in concentrate and supplementation of SO significantly reduced fungal diversity. We explored microbial interactions by building taxa co-occurrence networks. Our results suggest that studying entire rumen microbiome simultaneously is needed aiming to better understand how diet induced changes within microbial community are associated with microbial function, subsequently leading to better understanding of rumen fermentation and methanogenesis.
Data from: Acceptance threshold theory can explain occurrence of homosexual behaviour
Same-sex sexual behaviour (SSB) has been documented in a wide range of animals, but its evolutionary causes are not well understood. Here, we investigated SSB in the light of Reeve's acceptance threshold theory. When recognition is not error-proof, the acceptance threshold used by males to recognize potential mating partners should be flexibly adjusted to maximize the fitness pay-off between the costs of erroneously accepting males and the benefits of accepting females. By manipulating male burying beetles' search time for females and their reproductive potential, we influenced their perceived costs of making an acceptance or rejection error. As predicted, when the costs of rejecting females increased, males exhibited more permissive discrimination decisions and showed high levels of SSB; when the costs of accepting males increased, males were more restrictive and showed low levels of SSB. Our results support the idea that in animal species, in which the recognition cues of females and males overlap to a certain degree, SSB is a consequence of an adaptive discrimination strategy to avoid the costs of making rejection errors.
Data from: A multi-state dynamic occupancy model to estimate local colonization-extinction rates and patterns of co-occurrence between two or more interacting species
1. Although ecology is rife with theory that explores how multiple species co-occur through space and time, the field lacks robust statistical models to parameterize this theory with empirical data, particularly when species are detected imperfectly and data are collected as a time-series. 2. We address this need by developing an occupancy model that estimates local colonization and extinction rates for two or more interacting species when data are collected across multiple sampling occasions. This model estimates how community composition at a site may change across sampling occasions by assuming the latent occupancy state is a categorical random variable. We used a multinomial-logit model to parameterize species-specific parameters and pairwise interactions between species, both of which can be made a function of covariates. These transition probabilities between community states can then be converted to occupancy or co-occurrence probabilities to determine how community composition varies along an environmental gradient or through time. 3. As an example, we estimate patterns of co-occurrence between coyote (Canis latrans), Virginia opossum (Didelphis virginiana), and raccoon (Procyon lotor) in Chicago, Illinois, USA with data from a multi-year camera trapping study. Models with pairwise interactions between species greatly out performed models that assumed independence between species. Opossum and raccoon, for example, were far less likely to go extinct in habitat patches where coyotes were present. 4. Community composition at a site depends on species interactions and the local environment. Our model can separate such effects by estimating the underlying processes that define species occurrence patterns. As a result, our model can more explicitly quantify a wide range of ecological dynamics and therefore be used to empirically test ecological theory, such as estimating priority effects at a site or turnover rates between species, both of which can be made to vary as a function of covariates.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.