Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,940
datasets available to search
ShareScore release 0.9.0
Dataset results
1,940 results for “data sample”
Data from: Assessing the usefulness of Citizen Science Data for habitat suitability modelling: opportunistic reporting versus sampling based on a systematic protocol
<p><strong>Aim:</strong> To evaluate the potential of models based on opportunistic reporting (OR) compared to models based on data from a systematic protocol (SP) for modelling species distributions. We compared model performance for eight forest bird species with contrasting spatial distributions, habitat requirements, and rarity. Differences in the reporting of species were also assessed. Finally, we tested potential improvement of models when inferring high quality absences from OR based on questionnaires sent to observers.</p> <p><strong>Location:</strong> Both datasets cover the same large area (Sweden) and time period (2000 -2013).</p> <p><strong>Methods:</strong> Species distributions were modelled using logistic regression. Predictive performance of OR models to predict SP data were assessed based on AUC. We quantified the congruence in spatial predictions using Spearman's rank correlation coefficient. We related these results to species characteristics and reporting behaviour of observers. We also assessed the gain in predictive performance of OR models by adding inferred absences. Finally, we investigated the potential impact of sampling bias in OR.</p> <p><strong>Results:</strong> For all species, and despite the sampling biases, results from OR overall agreed well with those of SP, for the nationwide spatial congruence of habitat suitability maps and the selection and directions of species-environment relationships. The OR models also performed well in predicting the SP data. The predictive performance of the OR models increased with species rarity and even outperformed the SP model for the rarest species. No significant impact of observer behaviour was found.</p> <p><strong>Main Conclusions:</strong> Relatively simple analyses with inferred absences could produce reliable spatial predictions of habitat suitability. This was especially true for rare species. OR data should be seen as a complement to SP, as the weakness of one is the strength of the other, and OR may be especially useful at large spatial scales or where no systematic data collection protocols exist.</p>
Sample locations and structural data
<p>Available data compile sample locations and structural measurements from several field campains between 2016 and 2018. The area of interest is in the westernmost part of the Mediterranean region, in the Betic Cordillera (Andalucia, Spain).</p> <p> </p> <p>The full dataset is linked to a publication in <em>Tectonics</em> Journal (Bessière et al., 2021):</p> <p>" Exhumation of the Ronda peridotite during hyper-extension: New structural and thermal constraints from the Nieves Unit (western Betic Cordillera, Spain)"</p>
Tree species richness differentially affects the chemical composition of leaves, roots and root exudates in four subtropical tree species - Sampling Raw Data
<p>Sampling Raw Data for the manuscript "<strong>Tree species richness differentially affects the chemical composition of leaves, roots and root exudates in four subtropical tree species </strong>" </p> <p>R Code for producing the sunburst plots from the data obtained by classyFire</p> <p> </p>
Fig. 23 in Phylogenetic Studies On Didelphid Marsupials Ii. Nonmolecular Data And New Irbp Sequences: Separate And Combined Analyses Of Didelphine Relationships With Denser Taxon Sampling
Fig. 23. Skull of Tlacuatzin canescens, a composite drawing based on USNM 125659 and 511261.
Data for the detection of the boreal chorus frog (Pseudacris maculata) using environmental DNA and call surveys at 180 ponds sampled in 2017-2018 in southeastern Québec, Canada
<p>The boreal chorus frog (<em>Pseudacris maculata</em>) is at risk of extinction in parts of its range in Canada. Our objectives were to quantify the influence of local and landscape characteristics on the occurrence of the species in wetlands in southern Québec. We hypothesized that site occupancy depends on local characteristics and landscape characteristics contributing to site connectivity. We developed an environmental DNA (eDNA) method to detect the species and compared the detection probability of this method to traditional call surveys. We collected water samples at a total of 180 sites (90 in 2017, 110 in 2018), whereas we surveyed a subset of 63 sites using both eDNA and call surveys in 2018. Site occupancy varied across years, but was higher in sites where the species had been previously detected during the last 12 years by other studies. Site occupancy did not vary with other local and landscape characteristics, in part due to an apparent decrease in the number of sites occupied by the species since the last 12 years. Detection probability via eDNA (0.81; 95% CI: [0.31; 0.98]) did not differ from that of call surveys (0.62; 95% CI: [0.25; 0.89]). To identify the optimal sampling period for the boreal chorus frog, future studies should estimate the detection probability of eDNA during the breeding season and the larval development period of the species.</p>
Diet DNA metabarcoding data from spiders (Heteropoda venatoria) from Palmyra Atoll (2015-2017) with both individual samples that have and have not been surface sterilized
<p>These are data and code from a study examining the potential for surface contamination to influence diet DNA metabarcoding datasets when DNA is sequenced from full body parts (in this case, the opisthosomas of spider individuals). These datasets include the raw sequencing data, all downstream datasets, and taxonomic assignments collected from database searches on BOLD and GenBank (accessed 2019). The code includes code to reproduce all bioinformatics (merge, filter, match to taxonomies, rarefy, sort) as well as all statistics and figures generated from analyses. Raw data are from DNA extractions of predator gut regions (opisthosomas) and amplification of the CO1 gene using PCR. The predator species is <em>Heteropoda venatoria </em>collected individually with sterilized implements and either the diet sequences from their natural diets were extracted or their diets following feeding spiders in a feeding trial. </p>
Sample locations, micas pictures and Ar data tables
<p>Sample locations, pictures of some dated micas and Ar data tables from step heating ages.</p> <p>Data linked to a publication in Tectonics Journal: "<strong><sup>40</sup></strong><strong>Ar/<sup>39</sup>Ar age constraints</strong> <strong>on H<em>P</em>/L<em>T</em> metamorphism in extensively overprinted units: the example of the Alpujárride subduction Complex (Betic Cordillera, Spain)</strong>" - Bessière et al., 2021.</p>
pyDeltaRCM sample data -- Golf Delta
<p>This model run was created to generate sample data.<br> Model was run on 10/14/2021, at the University of Texas at Austin.</p> <p>Run was computed with pyDeltaRCM v2.1.0. See log file for complete information<br> on system and model configuration.</p> <p>Data available at Zenodo, version 1.1: 10.5281/zenodo.5570962.</p> <p>Version history:<br> v1.1: 10.5281/zenodo.5570962<br> v1.0: 10.5281/zenodo.4456144<br> </p>
Digital Forestry Toolbox - Sample data
<p>This repository contains airborne laser scanning data samples acquired by the states of Geneva, Solothurn and Zurich (in Switzerland). They are used in the tutorials of the <a href="https://github.com/mparkan/Digital-Forestry-Toolbox">Digital Forestry Toolbox for Matlab/Octave</a>.</p> <p>The full datasets (covering the complete state extents) are available from here: </p> <ul> <li><a href="https://maps.zh.ch/">https://maps.zh.ch/</a></li> <li><a href="https://geoweb.so.ch/map/lidar">https://geoweb.so.ch/map/lidar</a></li> <li><a href="https://ge.ch/sitg/donnees">https://ge.ch/sitg/donnees</a></li> </ul> <p><strong>Sources</strong> <strong>and usage conditions</strong>:</p> <ul> <li>Canton de Genève, Département de l'aménagement, du logement et de l'énergie (DALE), Système d'information du territoire à Genève (SITG). Dataset extracted on December 6, 2018. <a href="https://ge.ch/sitg/media/sitg/files/documents/conditions_generales_dutilisation_des_donnees_et_produits_du_sitg_en_libre_acces.pdf">See usage conditions</a>.</li> <li>Kanton Zürich, <a href="https://are.zh.ch/internet/baudirektion/are/de/aktuell.html">Baudirektion, Amt für Raumentwicklung</a>. Dataset extracted on December 6, 2018. <a href="https://are.zh.ch/internet/baudirektion/are/de/geoinformation/geodaten_uebersicht/Open_Data_Kanton_Zuerich.html#datenbezug">See usage conditions</a>.</li> <li>Kanton Solothurn, <a href="https://www.so.ch/verwaltung/bau-und-justizdepartement/">Bau- und Justizdepartement, Amt für Geoinformation</a>. Dataset extracted on December 6, 2018. <a href="https://geoweb.so.ch/geodaten/index.php?action=nutzung&UID=&USR=&user_id=&lang=de&menue=&aktuell=&sogis_zip_ie=">See usage conditions</a>.</li> </ul>
This is sample of the Data used to analyse AMV impacts on the global monsoon [Monerie et al. GRL]
<p>These data have been generated using MetUM-GOML2, and are used to analyse the impact of an anomaly cool/warm North Atlantic Ocean on tropical precipitation. These shared data could be used to reproduce the results shown in Monerie et al.</p> <p>More data will be available through the EU PRIMAVERA project via BADC.</p> <p>For more details about the metadata/having more data, please contact me at p.monerie@reading.ac.uk</p>
Sampling time-dependent artifacts in single-cell genomics studies: scRNA-seq data
<p>Robust protocols and automation now enable large-scale single-cell RNA and ATAC sequencing experiments and their application on biobank and clinical cohorts. However, technical biases introduced during sample acquisition can hinder solid, reproducible results, and a systematic benchmarking is required before entering large-scale data production. Here, we report the existence and extent of gene expression and chromatin accessibility artifacts introduced during sampling and identify experimental and computational solutions for their prevention.</p> <p>This repository contains the expression matrices and Seurat objects associated with the scRNA-seq data of the manuscript: "Sampling time-dependent artifacts in single-cell genomics studies" published in Genome Biology in 2020. The purpose of this repo is to share processed files and metadata for immediate access and reproducibility. The code to analyze it is thoroughly documented at the associated Github repository (https://github.com/massonix/sampling_artifacts).</p>
Simulated data set for SAXS tensor tomography, sample "M"
<p>This is a simulated data set intended for testing SAXS tensor tomographic reconstructions, including the original field from which the data was generated.</p>
Northern Nevada wildlife and topography: Camera trapping data set for 14 mammal species collected from 100 sampling sites in northwestern Nevada
<p>Camera traps are one of the most common field techniques for surverying terrestrial mammal communities and thus, much work has gone into understanding how different factors influence species detection at camera trap locations. However, the effect of fine-scale topography, such as terrain slope and position, on wildlife detection has not been explicitly quantified despite strong effects of topography on animal movement in mountainous regions. This data set contains weekly detection non-detection data for 14 mammal species from 100 camera traps sites monitored for 28 months (June 2018 - September 2020) in northwestern Nevada, U.S.A. This sampling extent was split into 3 month sampling seasons, exclusive of winter (Dec, Jan, Feb) and spring 2020, when data were sparse. In addition to species detection data, that dataset includes topographic variables at cameras sites: 1) terrain slope, calculated in R package raster from a 10m digital elevation model and 2) Topographic position index averaged across three buffer sizes around points 270m, 810m, and 2430m. The land cover variables proportion mixed conifer and proportion pinyon-juniper woodland within a 5000m buffer of sites are also included. Both are derived from the USDA/US DOI Landfire 2016 dataset. The luring variable indicates whether attractant was applied at a site during a given week, the effect of which was assumed to last for a month after the last application. </p>
Data of "A workflow to study the microbiota profile of piglet's umbilical cord blood: from sampling to data analysis".
<p>The present study proposes a workflow – from the sampling method to DNA extraction, bioinformatics and data analysis – that characterises the bacterial profile of umbilical cord blood samples, taking into account the contaminants found throughout the procedure of bacterial DNA extraction and amplification.</p> <p>Ps_umbilical.rds: A phyloseq object file of data containing the amplicon sequences variants (ASVs) of thirteen umbilical cord samples and two negative control samples, created by DADA2.</p> <p>R-script.doc: A word document containing the scripts used to characterize the taxonomical composition of the fifteen umbilical cord samples and two negative control samples before and after the application of Decontam R-package (Davis et al., 2018).</p> <p>metadata.docx: meta data for R-script.doc</p>
The impacts of fine-tuning, phylogenetic distance, and sample size on big-data bioacoustics
<p>Vocalizations in animals, particularly birds, are critically important behaviors that influence their reproductive fitness. While recordings of bioacoustic data have been captured and stored in collections for decades, the automated extraction of data from these recordings has only recently been facilitated by artificial intelligence methods. These have yet to be evaluated with respect to accuracy of different automation strategies and features. Here, we use a recently published machine learning framework to extract syllables from ten bird species ranging in their phylogenetic relatedness from 1 to 85 million years, to compare how phylogenetic relatedness influences accuracy. We also evaluate the utility of applying trained models to novel species. Our results indicate that model performance is best on conspecifics, with accuracy progressively decreasing as phylogenetic distance increases between taxa. However, we also find that the application of models trained on multiple distantly related species can improve the overall accuracy to levels near that of training and analyzing a model on the same species. When planning big-data bioacoustics studies, care must be taken in sample design to maximize sample size and minimize human labor without sacrificing accuracy.</p>
Sample based prevalence data complementing the European Union One Health 2021 Zoonoses Report - Croatia
<p>This dataset contains monitoring data on zoonoses and zoonotic agents under the Directive 2003/99/EC. This Directive requires Member Sates (MSs) to collect, evaluate and report data on zoonoses and zoonotic agents. MSs can also report monitoring data and information on some other pathogenic microbiological agents in foodstuffs. Relevant EU legislation: Commission Regulation (EC) No 2073/2005,Commission Regulation (EC) No 1441/2007, Commission Regulation (EU) No 1086/2011, Commission Regulation (EU) No 209/2013, Commission Regulation(EU) No 217/2014.</p>
Sample based prevalence data complementing the European Union One Health 2021 Zoonoses Report - the United Kingdom (Northern Ireland)
<p>This dataset contains monitoring data on zoonoses and zoonotic agents under the Directive 2003/99/EC. This Directive requires Member Sates (MSs) to collect, evaluate and report data on zoonoses and zoonotic agents. MSs can also report monitoring data and information on some other pathogenic microbiological agents in foodstuffs. Relevant EU legislation: Commission Regulation (EC) No 2073/2005,Commission Regulation (EC) No 1441/2007, Commission Regulation (EU) No 1086/2011, Commission Regulation (EU) No 209/2013, Commission Regulation(EU) No 217/2014.</p>
Sample based prevalence data complementing the European Union One Health 2021 Zoonoses Report - Finland
<p>This dataset contains monitoring data on zoonoses and zoonotic agents under the Directive 2003/99/EC. This Directive requires Member Sates (MSs) to collect, evaluate and report data on zoonoses and zoonotic agents. MSs can also report monitoring data and information on some other pathogenic microbiological agents in foodstuffs. Relevant EU legislation: Commission Regulation (EC) No 2073/2005,Commission Regulation (EC) No 1441/2007, Commission Regulation (EU) No 1086/2011, Commission Regulation (EU) No 209/2013, Commission Regulation(EU) No 217/2014.</p>
Sample based prevalence data complementing the European Union One Health 2021 Zoonoses Report - Luxembourg
<p>This dataset contains monitoring data on zoonoses and zoonotic agents under the Directive 2003/99/EC. This Directive requires Member Sates (MSs) to collect, evaluate and report data on zoonoses and zoonotic agents. MSs can also report monitoring data and information on some other pathogenic microbiological agents in foodstuffs. Relevant EU legislation: Commission Regulation (EC) No 2073/2005,Commission Regulation (EC) No 1441/2007, Commission Regulation (EU) No 1086/2011, Commission Regulation (EU) No 209/2013, Commission Regulation(EU) No 217/2014.</p>
Sample based prevalence data complementing the European Union One Health 2021 Zoonoses Report - Ireland
<p>This dataset contains monitoring data on zoonoses and zoonotic agents under the Directive 2003/99/EC. This Directive requires Member Sates (MSs) to collect, evaluate and report data on zoonoses and zoonotic agents. MSs can also report monitoring data and information on some other pathogenic microbiological agents in foodstuffs. Relevant EU legislation: Commission Regulation (EC) No 2073/2005,Commission Regulation (EC) No 1441/2007, Commission Regulation (EU) No 1086/2011, Commission Regulation (EU) No 209/2013, Commission Regulation(EU) No 217/2014.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.