Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,940

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,940 results for “data sample”

Learn how ShareScore rates datasets ↗
zenodo40/100

Rt-Cloud Sample Project 'AmygActivation' DICOM Data

<p>This upload contains&nbsp;the same data as published in <a href="https://doi.org/10.5281/zenodo.3677090">our previous zenodo dataset upload</a>. Unlike our previous upload,&nbsp;this version contains data&nbsp;after&nbsp;transferring the DICOMs&nbsp;directly from the Siemens Skyra 3T to our Linux machine (as done in real-time experiments). The purpose of this separate upload is to serve as sample data for our <a href="https://github.com/brainiak/rt-cloud">real-time cloud software</a>, for a <a href="https://github.com/amennen/amygActivation">specific sample project</a>. The brain data are contributed by author S.A.N. and are authorized for non-anonymized distribution.</p>

opencc-by-4.0May 2020View details →
zenodo40/100

The Set Increment with Limited Views Encoding Ratio (SILVER) Method for Optimizing Radial Sampling of Dynamic MRI: Supporting Data

<p>This data was created to perform the first experiments with the SILVER method of optimizing radial MRI acquisition. The files are mainly .mat files containing the numerical data used to assess the SILVER method using MATLAB (Version R2018b). Instructions on how to use the data and how to reproduce the experiments are available on https://github.com/SophieSchau/SILVER</p>

opencc-by-4.0Jun 2020View details →
dryad40/100

Data from: Effects of taxon sampling and tree reconstruction methods on phylodiversity metrics

1. The amount and patterns of phylodiversity in a community are often used to draw inferences about the local and historical factors affecting community assembly and can be used to prioritize communities and locations for conservation. Because measures of phylodiversity are based on the topology and branch lengths of phylogenetic trees, which are affected by the number and diversity of taxa in the tree, these analyses may be sensitive to changes in taxon sampling and tree reconstruction methods. 2. To investigate the effects of taxon sampling and tree reconstruction methods on measures of phylodiversity, we investigated the community phylogenetics of the Ordway-Swisher Biological Station (Florida), which is home to over 600 species of vascular plants. We studied the effects of 1) the number of taxa included in the regional phylogeny; 2) random vs. targeted sampling of species to assemble the regional species pool; 3) including only species from specific clades rather than broad sampling; 4) using trees reconstructed directly for the taxa under study compared to trees pruned from a larger reconstructed tree; and 5) using phylograms compared to chronograms. 3. We found that including more taxa in a study increases the likelihood of observing significantly non-random phylogenetic patterns. However, there were no consistent trends in the phylodiversity patterns based on random taxon sampling compared to targeted sampling, or within individual clades compared to the complete dataset. Using pruned and reconstructed phylogenies resulted in similar patterns of phylodiversity, while chronograms in some cases led to significantly different results from phylograms. 4. The methods commonly used in community phylogenetic studies can significantly impact the results, potentially influencing both inferences of community assembly and conservation decisions. We highlight the need for both careful selection of methods in community phylogenetic studies and appropriate interpretation of results, depending on the specific questions to be addressed.

opencc-zeroJun 2020View details →
zenodo40/100

Survey data, models and dated samples of the Pliocene shorelines of Camarones, Argentina (Ver 1.1).

<p>The dataset cosists of a spreadsheet containing data on GPS surveys, dynamic topography extracted from published models (gplates.org), Shell preservation scoring, Strontium Isotopic Stratigraphy ages, and Global mean Sea Level calculations.</p> <p>Version 1.1 contains fixes to small errors and formulas.</p>

opencc-by-4.0Jul 2020View details →
zenodo40/100

Netflow data without sampling for test (D2)

<p>NetFlow traffic generated using&nbsp;<strong>DOROTHEA</strong>&nbsp;(<strong>DO</strong>cker-based f<strong>R</strong>amework f<strong>O</strong>r ga<strong>TH</strong>ering n<strong>E</strong>tflow tr<strong>A</strong>ffic)</p> <p>NetFlow is a network protocol developed by Cisco for the collection and monitoring of network traffic flow data generated. A flow is defined as a unidirectional sequence of packets with some common properties that pass through a network device.</p> <p>NetFlow flows have been captured without sampling at the packet level. A sampling means that 1 out of every X packets is selected to be flow while the rest of the packets are not valued.</p> <p>The version of NetFlow used to build the datasets is 5.</p> <p>In the construction of the datasets, different percentages of flows considered attacks and flows considered normal traffic have been used.</p> <p>These datasets have been used to test&nbsp;machine learning models.</p>

opencc-by-4.0Oct 2020View details →
zenodo40/100

Netflow data without sampling for training (D1)

<p>NetFlow traffic generated using&nbsp;<strong>DOROTHEA</strong>&nbsp;(<strong>DO</strong>cker-based f<strong>R</strong>amework f<strong>O</strong>r ga<strong>TH</strong>ering n<strong>E</strong>tflow tr<strong>A</strong>ffic)</p> <p>NetFlow is a network protocol developed by Cisco for the collection and monitoring of network traffic flow data generated. A flow is defined as a unidirectional sequence of packets with some common properties that pass through a network device.</p> <p>NetFlow flows have been captured without sampling at the packet level. A sampling means that 1 out of every X packets is selected to be flow while the rest of the packets are not valued.</p> <p>The version of NetFlow used to build the datasets is 5.</p> <p>In the construction of the datasets, different percentages of flows considered attacks and flows considered normal traffic have been used.</p> <p>These datasets have been used to train machine learning models.</p>

opencc-by-4.0Oct 2020View details →
zenodo40/100

Application of contact-resonance AFM methods to polymer samples (Raw Data)

<p>Raw data and figures of the article &quot;Application of contact-resonance AFM methods to polymer samples&quot;, published in Beilstein Journal of Nanotechnology on 12 Nov 2020</p> <p>&nbsp;</p> <p>The raw data can be opened with the software &quot;Igor Pro&quot;</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Data: Flower visiting insects of kiwifruit within New Zealand commercial orchard blocks sampled over two years in the Bay of Plenty, New Zealand

<p>These data are total counts of individual bee and non&ndash;bee insects observed visiting the flowers of kiwifruit (<em>Actinidia chinensis</em> var.deliciosa) (&lsquo;Hayward&rsquo;) vines in three commercial orchards located in the Bay of Plenty Region of New Zealand (37&deg; 46&#39; 56&quot; S; 176&deg; 19&#39; 10&quot; E). Each block was located on a different farm and each separated by a distance of at least two kilometres and surveyed twice in two consecutive years. A total of 1181 insects were observed, 741 in the 2014 season and 460 in the 2015 season. Insects from four orders were recorded. The most abundant species were honey bees <em>Apis mellifera</em> (n= 1068; 90.4%), flower longhorn beetles <em>Zorion guttigerum</em> (n= 52; 4.4%), the native bee <em>Lasioglossum</em> <em>sordidum</em>/c<em>ognatum</em> (n= 12; 1.0%) and the hover fly <em>Melanostoma fasciatum</em> (n= 11; 0.9%)&nbsp; Others insects represented 3.2% of individuals observed (n=38). We present a table of counts of the insects observed.</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Replication Data for: Continuously moving table MRI with golden angle radial sampling

<p>Continuously moving table (CMT) MRI is a high throughput technique that has multiple applications in whole-body imaging. In this work, CMT MRI based on golden angle (GA, 111.246&deg; azimuthal step) radial sampling is developed at 3 Tesla (T), with the goal of increased flexibility in image reconstruction using arbitrary profile groupings.</p>

openmit-licenseOct 2014View details →
zenodo40/100

USPTO patent data: 250k random sample and NPEs' patents

<p>This package includes:</p> <ol> <li>Two Stata .dta files consisting of information on patents assigned by the United States Patent and Trademark Office between 1976 and 2014: a random sample of 250 000 US patents, and data on patent owned by Intellectual Ventures, RPX, and several other companies. The variables for example include: grant date, application date, forward and backward citations, renewals, claims&nbsp;and others.</li> <li>Source codes and methods used in generating and analyzing the two data files.</li> </ol> <p>A bachelor thesis with further information will be linked here.</p>

opencc-zeroMay 2016View details →
zenodo40/100

Summarised contextual data about metabarcoding Tara Oceans samples (2009-2013)

<p>Tab-separated values table describing the metabarcoding samples from the expedition Tara Oceans (2009-2013).</p> <p>Information such as depth, time, geographic position, size fraction, collected from <a href="https://pangaea.de/">Pangaea</a>, are listed in context_general tables. In context_stat tables, you will find a selection of physico-chemical parameters. Tara_Oceans_Pangaea_context.rds gathers all the data collected from Pangaea in a single R object.</p> <p>These tables have been built using the code here: <a href="https://gitlab.com/tara-and-friends-euk-metab/tara-oceans-metab-context/-/tree/v1.1.1" target="_blank" rel="noopener">https://gitlab.com/tara-and-friends-euk-metab/tara-oceans-metab-context/-/tree/v1.1.2</a> (v1.1.2).</p> <p>In this version 16S metabarcoding samples missing in previous versions were added in context_general.* and context_sats.*</p>

opencc-by-4.0Oct 2022View details →
dryad40/100

Data from: Integrated species distribution models to account for sampling biases and improve range wide occurrence predictions

<p><strong><span>Aim</span></strong></p> <p><span>Species distribution models (SDMs) that integrate presence-only and presence-absence data offer a promising avenue to improve information on species' geographic distributions. The use of such 'integrated SDMs' on a species range-wide extent has been constrained by the often-limited presence-absence data and by the heterogeneous sampling of the presence-only data. Here, we evaluate integrated SDMs for studying species ranges with a novel expert range map-based evaluation. We build a new understanding about how integrated SDMs address issues of estimation accuracy and data deficiency and thereby offer advantages over traditional SDMs.</span></p> <p><strong><span>Location</span></strong></p> <p><span>South and Central America.</span></p> <p><strong><span>Time period</span></strong></p> <p><span>1979-2017.</span></p> <p><strong><span>Major taxa studied</span></strong></p> <p><span>Hummingbirds.</span></p> <p><strong><span>Methods</span></strong></p> <p><span>We build integrated SDMs by linking two observation models – one for each data type – to the same underlying spatial process.</span> <span>We validate SDMs with two schemes: i) cross-validation with presence-absence data and ii) comparison with respect to the species' whole range as defined with IUCN range maps. We also compare models relative to the estimated response curves and compute the association between the benefit of the data integration and the number of presence records in each data set.</span></p> <p><strong><span>Results</span></strong></p> <p><span>The integrated SDM accounting for the spatially varying sampling intensity of the presence-only data was one of the top-performing models in both model validation schemes. Presence-only data alleviated overly large niche estimates, and data integration was beneficial compared to modelling solely presence-only data for species that had few presence points when predicting the species' whole range. On the community level, integrated models improved the species richness prediction.</span></p> <p><strong><span>Main conclusions</span></strong></p> <p><span>Integrated SDMs combining presence-only and presence-absence data are successfully able to borrow strengths from both data types and offer improved predictions of species' ranges. Integrated SDMs can potentially alleviate the impacts of taxonomically and geographically uneven sampling and to leverage the detailed sampling information in presence-absence data.</span></p>

opencc-zeroNov 2023View details →
dryad40/100

Data from: Wildlife fecal microbiota exhibit community stability across a semi-controlled longitudinal non-invasive sampling experiment

<p>Wildlife microbiome studies are being used to assess microbial links with animal health and habitat. The gold standard of sampling microbiomes directly from captured animals is ideal for limiting potential abiotic influences on microbiome composition, yet fails to leverage the many benefits of non-invasive sampling. Application of microbiome-based monitoring for rare, endangered, or elusive species creates a need to non-invasively collect scat samples shed into the environment. Since controlling sample age is not always possible, the potential influence of time-associated abiotic factors was assessed. To accomplish this, we analyzed partial 16S rRNA genes of fecal metagenomic DNA sampled non-invasively from Rocky Mountain elk (<em>Cervus canadensis</em>) near Yellowstone National Park. We sampled pellet piles from four different elk, then aged them in a natural forest plot for 1, 3, 7, and 14 days, with triplicate samples at each time point (i.e., a blocked, repeat measures (longitudinal) study design). We compared microbiomes of each elk through time with point estimates of diversity, bootstrapped hierarchical clustering of samples, and a version of ANOVA–simultaneous components analysis (ASCA) with PCA (LiMM-PCA) to assess the variance contributions of time, individual and sample replication. Our results showed community stability through days 0, 1, 3 and 7, with a modest but detectable change in abundance in only 2 genera (<em>Bacteroides</em> and <em>Sporobacter</em>) at day 14. The total variance explained by time in our LiMM-PCA model across the entire 2-week period was not statistically significant (p&gt;0.195) and the overall effect size was small (&lt;10% variance) compared to the variance explained by the individual animal (p&lt;0.0005; 21% var.). We conclude that non-invasive sampling of elk scat collected within one week during winter/early spring provides a reliable approach to characterize microbiome composition in a 16S rDNA survey and that sampled individuals can be directly compared across unknown time points with minimal bias. Further, point estimates of microbiome diversity were not mechanistically affected by sample age. Our assessment of samples using bootstrap hierarchical clustering produced clustering by animal (branches) but not by sample age (nodes). These results support greater use of non-invasive microbiome sampling to assess ecological patterns in animal systems.</p>

opencc-zeroNov 2023View details →
zenodo40/100

IW-NET sample data-set: River Weser IW network

<p>This data-set provides a sample of the data that was utilized for the analysis performed in the context of the IW-NET research project. The data have been collected via publicly available sources and are offered in this package as a sample. The use case the data refer to is River Weser, in northern Germany. The following table explains the contents of each of the uploaded files.</p> <table> <tbody> <tr> <td><strong>File</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td><a href="../api/records/10391858/draft/files/RiverWeserWaterway.json/content" target="_blank" rel="noopener noreferrer">RiverWeserWaterway.json</a></td> <td>OpenStreetMap data describing the Weser region</td> </tr> <tr> <td><a href="../api/records/10391858/draft/files/RiverWeserRelationsWaysNodes.json/content" target="_blank" rel="noopener noreferrer">RiverWeserRelationsWaysNodes.json</a></td> <td>OpenStreetMap data describing the Weser region</td> </tr> <tr> <td> <div><a href="../api/records/10391858/draft/files/IWTWeather.json/content" target="_blank" rel="noopener noreferrer">IWTWeather.json</a></div> </td> <td>Weather reports from 6 stations in the Weser region</td> </tr> <tr> <td><a href="../api/records/10391858/draft/files/unCitiesDE.json/content" target="_blank" rel="noopener noreferrer">unCitiesDE.json</a></td> <td>UN/LOCODE data in json format.</td> </tr> <tr> <td> <div><a href="../api/records/10391858/draft/files/MMSI.xlsx/content" target="_blank" rel="noopener">MMSI.xlsx</a></div> </td> <td>Correspondance of MMSI codes to Vessel registration country</td> </tr> </tbody> </table> <p>These resources were combined with AIS data logs from vessels active in the area - which, for legal reasons, cannot be made publicly available to provide powerful insights into the logistics operations and their intricacies.</p>

opencc-by-4.0Dec 2023View details →
dryad40/100

Data From: what mandrills leave behind: using fecal samples to characterize the major histocompatibility complex in a threatened primate

<p>The major histocompatibility complex (MHC) can be useful in guiding conservation planning because of its influence on immunity, fitness, and reproductive ecology in vertebrates. The mandrill (<em>Mandrillus sphinx</em>) is a threatened primate endemic to central Africa. Considerable research in this species has shown that the MHC is important for disease resistance, mate choice, and reproductive success. However, all previous MHC research in mandrills has focused on an inbred semi-captive population, so their genetic diversity may have been underestimated. Here we expand our current knowledge of mandrill MHC variation by performing next-generation sequencing of non-invasively collected fecal samples from a large wild horde in central Gabon. We observe MHC lineages and alleles shared with other primates, and we uncover 45 putative new class II MHC DRB alleles, including representatives of the DRB9 pseudogene, which has not previously been identified in mandrills. We also document methodological challenges associated with fecal samples in NGS-based MHC research. Even with high read depth, the replicability of alleles from fecal samples was lower than that of tissue samples, and allele assignments are inconsistent between sample types. Further, the common assumption that variants with very high read depth should represent true alleles does not appear to be reliable for fecal samples. Nevertheless, the use of degraded DNA in the present study still enabled significant progress in quantifying immunogenetic diversity and its evolution in wild primates.</p>

opencc-zeroJan 2024View details →
zenodo40/100

Improving Algorithm-Selectors and Performance-Predictors via Learning Discriminating Training Samples - Code and Data

<p>This repository contains the code and data for reproducibility of the paper 'Improving Algorithm-Selectors and Performance-Predictors via Learning Discriminating Training Samples'.&nbsp;</p> <p>The following files are included:</p> <ul> <li>Plots: Additional plots not in the paper;</li> <li>Code: Python scripts to generate trajectories and perform classification/regression;</li> <li>best_algo.csv : Labels for the classification;</li> <li>performances.csv : Performances used for the regression;</li> <li>SA_parameters.csv : SA parameters for all machine learning tasks;</li> <li>irace_scenario.txt : scenario used for the tuning.</li> </ul>

opencc-by-4.0Jan 2024View details →
zenodo40/100

Data from: Transcriptome analysis of apical meristem enriched bud samples for size dependent flowering commitment in Crocus sativus reveal role of sugar and auxin signalling

<p><strong>Background</strong></p> <p>Cultivation of <em>Crocus sativus</em> (saffron) faces challenges due to inconsistent flowering patterns and variations in yield. Flowering takes place in a graded way with smaller corms unable to produce flowers. Enhancing the productivity requires a comprehensive understanding of the underlying genetic mechanisms that govern this size based flowering initiation and commitment. Therefore, samples enriched with non-flowering and flowering apical buds from small (&lt;6g) and large (&gt;14g) corms were sequenced.&nbsp;</p> <p><strong>Methods and Results</strong></p> <p>Apical bud enriched samples from small and large corms were collected immediately after break of dormancy in July. RNA sequencing was performed using Illumina Novaseq 6000. <em>De-novo</em> transcriptome assembly and analysis using flowering committed buds from large corms at post-dormancy and their comparison with vegetative shoot primordia from small corms pointed out the major role of Auxin and ABA hormonal regulation. Many genes with known dual responses in flowering development and circadian rhythm like Flowering locus T and Cryptochrome 1 along with a transcript showing homology with small auxin upregulated RNA (SAUR) exhibited induced expression in flowering buds. Thorough prediction of&nbsp;<em>Crocus sativus</em> non-coding RNA repertoire has been carried out for the first time. Enolase was found to be acting as a major hub with protein-protein interaction analysis using Arabidopsis counterparts.</p> <p><strong>Conclusion</strong></p> <p>Transcripts belong to key pathways including phenylpropanoid biosynthesis, hormone signaling and carbon metabolism were found significantly modulated. KEGG assessment and protein-protein interaction analysis confirm the expression data. Findings unravel the genetic determinants driving the size-dependent&nbsp;flowering in <em>Crocus sativus</em>.</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Chemical analysis dataset for contaminants of emerging concern and bioanalytical data for samples from a low mountain stream in Central Germany

<p>In 2022, river-water samples were collected at six sampling sites along the Holtemme River in Central Germany using large-volume solid phase extraction. The extracts were analysed by target chemical analysis for contaminants of emerging concern. In addition, the extracts were analysed in a bioanalytical test battery using effect-based tools. The battery included assays for cytotoxicity (neutral red retention assay), oxidative stress (Nrf2-CALUX&reg;), endocrine disruption (ER-, AR-, anti-ER-, anti-AR-, GR- and PR-CALUX&reg;) and the fish embryotoxicity test with zebrafish (<em>Danio rerio</em>). The data obtained are included in the .csv files in this repository.</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Data set for BA and BQ samples

<p><span>Collection of data for the manuscript entitled</span></p> <p><span>Carbonized apples and quinces stillage for electromagnetic shielding</span></p> <p><span>Data for BA and BQ is a collection of data - file is manuscript, file type .pdf</span></p> <p><span>FTPO_TGA_BA.txt &ndash; thermogravimetric data of carbon-based nanomaterial produced from biomass apple, file type .txt</span></p> <p><span>VINCA_Raman_BA.txt &ndash; Raman spectra data of BA, file type .txt</span></p> <p><span>FTPO_TGA_BQ.txt - thermogravimetric data of carbon-based nanomaterial produced from biomass quince, file type .txt</span></p> <p><span>UniOldenburg_VNA_BA.txt &ndash; Vector network analysis of shielding efficiency BA, file type .txt</span></p> <p><span>UniOldenburg_VNA_BQ.txt &ndash; Vector network analysis of shielding efficiency BQ, file type .txt</span></p> <p><span>VINCA_Raman_BQ.txt &ndash; Raman spectra data of BQ, file type .txt</span></p> <p>&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

nanoCT data for the samples published in the article doi 10.1002/jemt.24746

<p>nanoCT data for the samples published in the article "Dealing with Missing Angular Sections in NanoCT Reconstructions of Low Contrast Polymeric Samples Employing a Mechanical In Situ Loading Stage" -&nbsp; Rafaela Debastiani, Chantal Miriam Kurpiers, Enrico Domenico Lemma, Ben Breitung, Martin Bastmeyer, Ruth Schwaiger, Peter Gumbsch</p> <p><a href="https://doi.org/10.1002/jemt.24746">https://doi.org/10.1002/jemt.24746</a></p> <p>Samples:<br>Sample A_180&deg;_phase<br>Sample A_140&deg;_pos0&deg;_phase<br>Sample A_140&deg;_pos0&deg;_absorption<br>Sample A_140&deg;_pos45&deg;_phase<br>Sample A_140&deg;_pos45&deg;_absorption<br>Sample A_140&deg;_pos90&deg;_phase<br>Sample A_140&deg;_pos90&deg;_absorption</p> <p>Sample B_140&deg;_pos0&deg;_phase<br>Sample B_140&deg;_pos90&deg;_absorption<br>Sample B_140&deg;_pos90&deg;_phase<br>Sample B_coated_140&deg;_absorption<br>Sample B_coated_140&deg;_phase</p> <p>&nbsp;</p> <p>Phase = Zernike phase contrast</p> <p>Absorption = absorption contrast</p>

opencc-by-4.0Nov 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record