Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,940
datasets available to search
ShareScore release 0.9.0
Dataset results
1,940 results for “data sample”
Raw data on sample languages
<p>Raw data on sample languages</p>
PyHawk_sample_data
Open the record for dataset details and reuse information.
Non-freely available datasets for the following work: "Comparison of consumption data and phenotypical antimicro-bial resistance in E. coli isolates of human urinary samples and of weaning and fattening pigs from surveillance and monitor-ing systems in Germany."
<p>Human and animal databases used for the creation of the work "Comparison of consumption data and phenotypical antimicrobial resistance in E. coli isolates of human urinary samples and of weaning and fattening pigs from surveillance and monitoring systems in Germany."</p>
Photizo: an open-source library for cross-sample analysis of FTIR spectroscopy data
<p>With continually improved instrumentation, Fourier transform infrared (FTIR) microspectroscopy can now be used to capture thousands of high-resolution spectra for chemical characterisation of a sample. The spatially resolved nature of this method lends itself well to histological characterisation of complex biological specimens. However, commercial software currently available can make joint analysis of multiple samples challenging and, for large datasets, computationally infeasible. In order to overcome these limitations, we have developed Photizo - an open-source Python library for spectral analysis which includes functions for pre-processing, visualisation and downstream analysis, including principal component analysis, clustering, macromolecular quantification and biochemical mapping. This library can be used for analysis of spectroscopy data without a spatial component, as well as spatially-resolved data, such as data obtained via infrared (IR) microspectroscopy in scanning mode and IR imaging by focal plane array (FPA) detector. The data set made available here was FTIR microspectroscopy spatially resolved data used for demonstrating cross-sample analysis using Photizo including example metadata. </p>
Sample Weather Observation Website Data (6 Sites, 1 Year each)
<p><strong>Contains public sector information licensed under the Open Government Licence v3.0</strong></p> <p>An example of data that can be downloaded from <a href="https://wow.metoffice.gov.uk/">WOW</a>, used to demonstrate data surrounding the Met Office in a <a href="https://github.com/informatics-lab/spaceapps-2022-mo-bootcamp">data tooling and analysis tutorial</a>. File names are adapted from the source by giving the WOW site id and the year of the downloaded data.</p>
Data from: Is Hydroides dianthus (Verrill, 1873) really a Mediterranean native? Increased sampling in the eastern United States reveals enhanced genetic diversity
<p>The introduction of non-indigenous species (NIS) is a significant threat to marine biodiversity, facilitated by vectors such as shipping and aquaculture. <em>Hydroides dianthus,</em> a tubicolous polychaete worm, is known for its biofouling capabilities, impacting both shipping and aquaculture. Traditionally, the east coast of the United States has been considered the native range of <em>H. dianthus</em>. However, previous studies have suggested the Mediterranean region as the species' true native range based on higher genetic diversity. This study aims to re-evaluate the genetic diversity patterns of <em>H. dianthus</em> on the east coast of the United States by expanding the cytochrome c oxidase I (COI) dataset currently available for the species. Samples were collected from various locations on the east coast and analyzed using DNA barcoding. The results revealed a three-fold increase in haplotype diversity on the east coast compared to previous findings. A hierarchical AMOVA indicated significant genetic structuring between the Mediterranean and U.S. populations (ϕST = 0.51, P < 0.05). Despite a higher genetic diversity in the Mediterranean, this study highlights the variability of genetic diversity estimates and the challenges in using such metrics to delineate native ranges. Factors such as multiple introductions, genetic drift, and sampling bias can significantly alter genetic variability within populations. The findings suggest that the east coast's genetic diversity is likely underestimated and that more comprehensive data, including high-throughput genomic analyses and ecological studies, are needed to determine the native range of <em>H. dianthus</em> conclusively. This study underscores the complexity of using genetic data to trace the biogeography and invasion pathways of marine species.</p>
Noise samples for Costantino et al., Seismic Source Characterization From GNSS Data Using Deep Learning (2023)
Open the record for dataset details and reuse information.
Nanostring GeoMx DSP data for TNBC Samples_ M-712Baylor
<div> <div> <div> <div> <p>Gene expression analysis using the GeoMX DSP system (NanoString) was conducted to profile 57 FFPE TNBC tissues from self-reported 26 BA and 31 WA patients. We examined the transcriptome of these samples using the GeoMX Human Whole Transcriptome Atlas (WTA), which measures approximately 18,000 protein-coding genes. The GeoMX DSP analysis involved staining the slides with two morphology markers—Pan Cytokeratin (PanCK) to identify tumor cells, and CD45 to identify immune cells—along with a nuclear stain (DAPI).</p> </div> </div> </div> </div> <div> <div> <div> </div> </div> </div>
Supplementary Data for article: Robust characterization of forest structure from airborne laser scanning – a systematic assessment and sample workflow for ecologists
<p>This is a collection of scripts and research data to assess the robustness of forest structure characterization from airborne laser scanning (ALS). It replicates the main analysis in the article <em>Robust characterization of forest structure from airborne laser scanning – a systematic assessment and sample workflow for ecologists</em> and accompanies the main research data set (<a href="https://doi.org/10.5281/zenodo.10878070" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.10878070</a>).</p> <p>In this replication study, we assess the derivation of canopy height models (CHMs) from point cloud data, how sensitive CHM algorithms are to pulse density variation and how uncertainties and biases propagate to commonly used forest structure metrics.</p> <p>The main data source for this study are ALS point clouds from nine U.S. sites, acquired by the National Ecological Observatory Network (NEON, 6 sites, 3 km x 3 km) and by the United States Geological Survey's 3DEP program (3 sites, also 3 km x 3 km). The underlying data can be found here: https://data.neonscience.org/data-products/DP3.30024.001 (NEON) and here: https://apps.nationalmap.gov/lidar-explorer (3DEP)</p> <p>The different data layers are:</p> <p><strong>replicate.US.R</strong> </p> <p> is a single R script that contains all the code necessary to reproduce the analyses, including point cloud manipulations and derivation of CHMs from the raw data as well as the overall robustness analysis. To replicate the processing of the raw point clouds step by step, this script should be located in a folder called "rscripts".</p> <p><strong>pointclouds_original.zip</strong></p> <p> is the set of original point clouds (3 km x 3 km in extent) used for the replication test, separated into 9 subfolders/sites. Can be used to reproduce the original workflow by placing them in a folder called "data/original". The script will then automatically produce derived point clouds at pulse densities of 2 and 16 per squaremetre and process them into digital terrain models (DTMs), digital surface models (DSMs) and CHMs. Note that, for convienence, these derived products are also included in a separate .zip file (cf. below).</p> <p><strong>processed_foranalysis.zip</strong></p> <p> is the set of derived products (DTMs, DSMs, CHMs), separated into 9 subfolders/sites, i.e. the result of processing the original point clouds. To use these layers directly with the provided script, they should be put into a folder called "processed_foranalysis".</p> <p><strong>summaries.zip</strong></p> <p><strong> </strong>is a set of summary statistics (as .csv files) that were used to generate the main analysis tables in the replication study. To use these summary statistics directly with the provided script, they should be put into a folder called "summaries".</p> <p><strong>figures.zip</strong></p> <p> is a set of figures displayed in the Supplementary Material of the paper <em>Robust characterization of forest structure from airborne laser scanning – a systematic assessment and sample workflow for ecologists</em>.</p>
Sample data from trench sediments and rocks
<p>The data of glycerol dialkyl glycerol tetraether (GDGT) membrane lipids in sediments and rocks of the Mariana and Yap Trench, northwest Pacific Ocean.</p>
Water sampling data from Outokumpu area 2024
<h2>Abstract</h2> <p>Report on the water sampling campaign performed in Outokumpu during summer 2024</p> <p>This depositry contains data generated within the European S34 project. </p> <h2>Metadata Information</h2> <table> <tbody> <tr> <td> <p><strong>Identification</strong></p> </td> </tr> <tr> <td> <p>Full Title</p> </td> <td> <p>Water analysis Outokumpu 2024</p> </td> </tr> <tr> <td> <p>Abstract</p> </td> <td> <p>Report on the water sampling campaign performed in Outokumpu during summer 2024</p> </td> </tr> <tr> <td> <p>Keywords</p> </td> <td> <p>AMD, water analysis</p> </td> </tr> <tr> <td> <p>Pilot area</p> </td> <td> <p>Keretti-Outokumpu</p> </td> </tr> <tr> <td> <p>Associated resources</p> </td> <td> <p>/</p> </td> </tr> <tr> <td> <p>Language</p> </td> <td> <p>English</p> </td> </tr> <tr> <td> <p>URL</p> </td> <td> <p>/</p> </td> </tr> <tr> <td> <p>Categories</p> </td> <td> <p>Other</p> </td> </tr> <tr> <td> <p><strong>Temporal reference</strong></p> </td> </tr> <tr> <td> <p>Creation date (dd.mm.yyyy)</p> </td> <td> <p>28.08.2024</p> </td> </tr> <tr> <td> <p>Revision date (dd.mm.yyyy)</p> </td> <td> <p>28.08.2024</p> </td> </tr> <tr> <td> <p><strong>Quality and validity</strong></p> </td> </tr> <tr> <td> <p>Representation type</p> </td> <td> <p>Other</p> </td> </tr> <tr> <td> <p>Format</p> </td> <td> <p>PDF</p> </td> </tr> <tr> <td> <p>Lineage</p> </td> <td> <p>/</p> </td> </tr> <tr> <td> <p>Spatial resolution</p> </td> <td> <p>0,20m</p> </td> </tr> <tr> <td> <p>Positional accuracy</p> </td> <td> <p>0,1</p> </td> </tr> <tr> <td> <p>Maintenance information</p> </td> <td> <p>/</p> </td> </tr> <tr> <td> <p>Coordinate system</p> </td> <td> <p>EPSG 32635</p> </td> </tr> <tr> <td> <p><strong>Constranits related to access and use</strong></p> </td> </tr> <tr> <td> <p>Use limitation</p> </td> <td> <p>None- refer to GTK as source of information</p> </td> </tr> <tr> <td> <p>Access constraint</p> </td> <td> <p>/</p> </td> </tr> <tr> <td> <p>Public/Private</p> </td> <td> <p>Public</p> </td> </tr> <tr> <td> <p><strong>Responsible organisation</strong></p> </td> </tr> <tr> <td> <p>Responsible Contact</p> </td> <td> <p>Nike Luodes (nike.luodes@gtk.fi)</p> </td> </tr> <tr> <td> <p>Responsible Party</p> </td> <td> <p>GTK</p> </td> </tr> <tr> <td> <p><strong>Metadata on metadata</strong></p> </td> </tr> <tr> <td> <p>Contact</p> </td> <td> <p>Nike Luodes (nike.luodes@gtk.fi)</p> </td> </tr> <tr> <td> <p>Metadata language</p> </td> <td> <p>English</p> </td> </tr> </tbody> </table>
PP-recyclates characterization after different sampling and recycling strategies - ATR-FTIR data
<p>The purpose of this analysis is to evaluate the efficiency of different recycling procedures and the quality of the resulting recyclates.</p> <p>This dataset contains raw parallel plate rheology data of PP-recyclates from different recycling strategies. The content is:</p> <ul> <li>One Excel file containing ATR-FTIR data of PP recyclates after scCO2 recycling, reference samples and measurement protocol</li> <li>One Excel file containing ATR-FTIR of PP recyclates after solvent-based recycling, reference samples and measurement protocol</li> <li>One Excel file containing ATR-FTIR data of PP recyclates after upcycling, reference samples and measurement protocol</li> <li>One Readme file containing further information about the methodology and nomenclature</li> </ul> <p>This dataset was generated in the framework of PRecycling Horizon Europe project (101058670)</p>
PP-recyclates characterization after different sampling and recycling strategies – XRF data
<p><span>The purpose of this analysis is the development of an efficient sampling protocol for plastic waste streams</span><span> after </span><span>upcycling or </span><span>the treatment with scCO2</span><span> <span>and solvent-based recycling.</span></span></p> <p><span>Taking into consideration the morphology of the recyclates after upcycling, the samples were prepared based on the model waste (MW) approaches. </span><span>The</span><span> recyclates after upcycling are in the form of pellets, therefore the MW approaches </span><span>were </span><span>selected </span><span>in which</span><span> extrusion processes were used.</span></p> <p><span>Taking into consideration the morphology of the recyclates after scCO2, the samples were prepared based on the model waste (MW) approaches. </span><span>The</span><span> recyclates after scCO2 treatment are in the form of flakes, however due to the high heterogeneity of the material and possible low accuracy of analytical evaluation, an extra step of grinding was ad</span><span>d</span><span>ed to increase their representativeness. </span></p> <p><span>Taking into consideration the morphology of the recyclates after </span><span>solvent-based recycling</span><span>, the samples were prepared based on the model waste (MW) approaches. </span><span>The</span><span> recyclates after </span><span>solvent-based recycling</span><span> treatment are in the form of powder, therefore the MW approaches </span><span>were </span><span>selected </span><span>in which</span><span> </span><span>cryogenic grinding processes were used.</span></p> <p><span>The sampling approaches from the (MW) were applied to the recyclates, and each </span><span>was assessed via various analytical technique</span><span>s.</span></p> <p><span>This dataset contains raw XRF data of the r</span><span>ecyclate </span><span>from the different sampling approaches</span><span>, such as:</span></p> <p><span><span>•<span> </span></span></span><span>Excel file</span><span>s</span><span> containing XRF data of the r</span><span>ecyclate</span><span>s, wherein the approach was based on upcycling (compounding)</span><span>,</span><span> <span>supercritical CO2 treatment and solvent-based recycling treatment</span></span><span> with a </span><span>measurement protocol</span></p> <ul> <li><span>One Readme file containing further information about the methodology and nomenclature </span></li> </ul> <p><span>This dataset was generated in the framework of PRecycling Horizon Europe project (101058670)</span></p>
Data gathered from grabbed samples on DWSPs (drinking water service public - houses of water) for CS#2
<p>Data gathered from grabbed samples on DWSPs used to compare data obtained from online sensor installed on DWSPs</p>
Data from: Studying the microbiota of bats: accuracy of direct and indirect samplings
Given the recurrent bat-associated disease outbreaks in humans and recent advances in metagenomics sequencing, the microbiota of bats is increasingly being studied. However, obtaining biological samples directly from wild individuals may represent a challenge and thus, indirect passive sampling (without capturing bats) is sometimes used as an alternative. Currently, it is not known whether the bacterial community assessed using this approach provides an accurate representation of the bat microbiota. This study was designed to compare the use of direct sampling (based on bat capture and handling) and indirect sampling (collection of bat's excretions under bat colonies) in assessing bacterial communities in bats. Using high-throughput 16S rRNA sequencing of urine and faeces samples from Rousettus aegyptiacus, a cave-dwelling fruit bat species, we found evidence of niche specialization among different excreta samples, independent of the sampling approach. However, sampling approach influenced both the alpha-and beta-diversity of urinary and faecal microbiotas. In particular, increased alpha-diversity and more overlapping composition between urine and faeces samples was seen when direct sampling was used, suggesting that cross-contamination may occur when collecting samples directly from bats in hand. In contrast, results from indirect sampling in the cave may be biased by environmental contamination. Our methodological comparison suggested some influence of the sampling approach on the bat-associated microbiota, but both approaches were able to capture differences among excreta samples. Assessment of these techniques opens an avenue to use more indirect sampling, in order to explore microbial community dynamics in bats.
sample data
<p>This is a sample data file to check how zenodo works</p>
Microsatellite data of 21 EST-SSR loci from all 469 Emmenopterys henryi samples
<p>Several hypotheses are available to predict change in genetic diversity at expanding range margins. However, empirical evidence to test predictions of the central-marginal hypothesis (CMH) at contracting range limits is scarce. To address this issue, we assessed spatial patterns of genetic variation, effective population size, and contemporary and historical gene flow in a widespread, Tertiary relict tree species from subtropical China, <i>Emmenopterys henryi</i>.</p>
Data for "Improving sampling of crystallographic disorder in Ensemble Refinement"
<p>Program outputs and data for "Improving sampling of crystallographic disorder in Ensemble Refinement"</p>
SAXS data collection and analysis of Q23 HTT and HTT-HAP40 samples - 2018/10/08
<p><strong>Project</strong> - Huntingtin structure-function open lab notebook. </p> <p><strong>Experiment</strong> - Small angle X-ray scattering (SAXS) of huntingtin-HAP40 complex samples with Q23 polyQ lengths.</p>
BSSRDF-data for 12 translucent samples
<p>This dataset contains BSSRDF-data in MATLAB format (.mat) for 12 translucent samples used in the Optics Express article "Primary facility for traceable measurement of the BSSRDF", see <a href="https://doi.org/10.1364/OE.439108">https://doi.org/10.1364/OE.439108</a>. Also, a description of the files and the format is found in the included WORD file (.docx).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.