Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,940
datasets available to search
ShareScore release 0.9.0
Dataset results
1,940 results for “data sample”
Denoising Autoencoders for Phenotype Stratification (DAPS) Sample Trained Patient Data
<p>Data for: https://github.com/greenelab/DAPS</p>
Biotea sample data
<p>This is the RDF loaded at http://biotea.linkeddata.es/sparql and described at http://biotea.github.io</p>
VAPOR Sample Data
<p>VAPOR is the Visualization and Analysis Platform for Ocean, Atmosphere, and Solar Researchers. VAPOR provides an interactive 3D visualization environment that can also produce animations and still frame images. VAPOR runs on most UNIX and Windows systems equipped with modern 3D graphics cards.</p> <p>VAPOR is a product of the National Center for Atmospheric Research's Computational and Information Systems Lab. Support for VAPOR is provided by the U.S. National Science Foundation and by the Korea Institute of Science and Technology Information</p> <p>This dataset contains sample files of model outputs from numerical simulations that VAPOR is capable of directly reading. They are not related to each other aside from being sample data for VAPOR.</p> <p>To unpack the tar.gz files on Linux/OSX, issue the command `tar -xzvf [myFile].tar.gz` on the file you've downloaded. On Windows, a program like 7-zip can perform that operation. Once unpacked, the files can be directly imported into VAPOR, or converted to VDC. For more information see the "Getting Data Into VAPOR" Related Link below.</p>
Flask sampling data obtained during the CoMet campaign in summer 2018 over Silesia on DLR Cessna
Open the record for dataset details and reuse information.
Article data of The Impact of MDR-1 Gene Polymorphism (rs1128503) on Response to Imatinib and nilotinib Treatment in A Sample of Iraqi Chronic Myeloid Leukemia-Chronic Phase Patients
Open the record for dataset details and reuse information.
Article data of The Impact of MDR-1 Gene Polymorphism (rs1128503) on Response to Imatinib and nilotinib Treatment in A Sample of Iraqi Chronic Myeloid Leukemia-Chronic Phase Patients
Open the record for dataset details and reuse information.
Data and Scripts for Anonymous et al. "How to optimally allocate sampling effort in experimental ecology"
Open the record for dataset details and reuse information.
Freezing nucleus spectra for hailstone samples in China from droplet freezing experiments—Data
<p>INP concentrations in hailstone samples from droplet freezing experiments.</p>
Writing Sample Data Sheet
Open the record for dataset details and reuse information.
Supporting publication for 'Prevalence sample-based guidance for reporting 2023 data'
<p>The record is aimed at helping the reporting countries to submit their 2023 sample-based level data to the EFSA Data Collection Framework. We include here two excel files and one XML file, and we give below specific information on their use.</p> <p>The two Excel documents help in mapping terms from the matrix catalogue ZOO_CAT_MATRIX used in the aggregated prevalence data model to FoodEx2 codes, and offer examples on how prevalence data can be reported using SSD2 and how data are aggregated afterwards. The XML file is the same example as in the Excel file with similar title but in the XML format that allows for it be uploaded in the Data Collection Framework.</p>
PyHawk_Sample_Data
Open the record for dataset details and reuse information.
A detailed ultrastructural examination of lung cryobiopsy samples from a COVID-19 patient case series – Data set 04
<p>We investigated six cryobiopsy samples from six deceased patients (patients C03 to C08 from Barisione et al. 2020 <a href="https://doi.org/10.1007/s00428-020-02934-1">doi.org/10.1007/s00428-020-02934-1</a>) by using thin section electron microscopy (Cortese et al. 2022 <a href="http://doi.org/10.1007/s00428-022-03308-5">doi.org/10.1007/s00428-022-03308-5</a>). A detailed description of the methods and the data set is provided in the download container.</p> <p>Data set 04 contains a stitched image montage of the first and of the last semithin section from the analysis of patient C06, which was acquired by bright-field light microscopy.</p>
A detailed ultrastructural examination of lung cryobiopsy samples from a COVID-19 patient case series – Data set 16
<p>We investigated six cryobiopsy samples from six deceased patients (patients C03 to C08 from Barisione et al. 2020 <a href="https://doi.org/10.1007/s00428-020-02934-1">doi.org/10.1007/s00428-020-02934-1</a>) by using thin section electron microscopy (Cortese et al. 2022 <a href="http://doi.org/10.1007/s00428-022-03308-5">doi.org/10.1007/s00428-022-03308-5</a>). A detailed description of the methods and the data set is provided in the download container.</p> <p>Data set 16 contains stitched image montages of a thin section through the lung of patient C06 which were acquired by scanning electron microscopy. The file “Data_set_16.tif” contains a montage of the entire thin section while the other files contain selected areas of the section recorded at higher resolution.</p>
Sample site coordinates, environmental data, number of copies of target DNA/ul for each sample and limit of detection plot
<p>Human activities in coastal areas are accelerating ecosystem changes at an unprecedented pace, resulting in habitat loss, hydrological modifications, and predatory species declines. Understanding how these changes potentially cascade across marine and freshwater ecosystems requires knowing how mobile euryhaline species link these seemingly-disparate systems. As upper trophic level predators, bull sharks (<i>Carcharhinus leucas</i>) play a crucial role in marine and freshwater ecosystem health. Telemetry studies in Mobile Bay, Alabama suggest that bull sharks extensively use the northern portions of the bay, an estuarine-freshwater interface known as the Mobile-Tensaw Delta. To assess whether bull sharks use freshwater habitats in this region, environmental DNA surveys were conducted during the dry summer and wet winter seasons in 2018. In each season, 5 x<span> 1</span> L water samples were collected at each of 21 sites: five sites in Mobile Bay, six sites in the Mobile-Tensaw Delta, and ten sites throughout the Mobile-Tombigbee and Tensaw-Alabama Rivers. Water samples were vacuum-filtered, DNA extractions were performed on the particulate, and DNA extracts were analyzed with Droplet Digital™ Polymerase Chain Reaction using species-specific primers and an internal probe to amplify a 237-base pair fragment of the mitochondrial NADH dehydrogenase subunit 2 gene in bull sharks. One water sample collected during the summer in the Alabama River met the criteria for a positive detection, thereby confirming the presence of bull shark DNA. While preliminary, this finding suggests that bull sharks use less urbanized, riverine habitats up to 120 km upriver during Alabama's dry summer season.</p>
Netflow data with sampling 1000 for test (D6)
<p>NetFlow traffic generated using DOROTHEA (DOcker-based fRamework fOr gaTHering nEtflow trAffic) NetFlow is a network protocol developed by Cisco for the collection and monitoring of network traffic flow data generated. A flow is defined as a unidirectional sequence of packets with some common properties that pass through a network device.</p> <p>NetFlow flows have been captured with sampling 1000 at the packet level. A sampling means that 1 out of every X packets is selected to be flow while the rest of the packets are not valued.</p> <p>The version of NetFlow used to build the datasets is 5.</p>
Audio samples from generative models trained on the TIMIT speech data.
<p>The snippets include samples and reconstructions. All samples are completely unconditional and utilise only the prior internal representations learned by the models. Reconstructions are computed from a given test audio snippet by first encoding it to a learned representation and then decoding that to a reconstruction of the audio.</p> <p>All models are trained on the TIMIT speech dataset (<a href="https://catalog.ldc.upenn.edu/LDC93s1">https://catalog.ldc.upenn.edu/LDC93s1</a>). Some snippets are from models trained at different temporal resolutions denoted by `s1` and `s64`. We refer to the paper for details.</p> <p>The files include:</p> <ul> <li>`clockwork-vae-s64-reconstruction-*` <ul> <li>Four reconstructions using a two-layered Clockwork VAE trained with temporal resolution s=64.</li> </ul> </li> <li>`clockwork-vae-s64-sample-*` <ul> <li>Four samples from the prior of a Clockwork VAE trained with temporal resolution s=64.</li> </ul> </li> <li>`original-*` <ul> <li>Four original samples from TIMIT corresponding in pairs to the reconstructions.</li> </ul> </li> <li>`vrnn-s64-sample-*` <ul> <li>Two samples from the prior of a VRNN trained with temporal resolution s=64.</li> </ul> </li> <li>`vrnn-s1-sample-*` <ul> <li>Two samples from the prior of a VRNN trained with temporal resolution s=1.</li> </ul> </li> <li>`srnn-s64-sample-*` <ul> <li>Two samples from the prior of a SRNN trained with temporal resolution s=64.</li> </ul> </li> <li>`srnn-s1-sample-*` <ul> <li>Two samples from the prior of a SRNN trained with temporal resolution s=1.</li> </ul> </li> <li>`wavenet-s64-sample-*` <ul> <li>Four samples from a WaveNet trained with temporal resolution s=1.</li> </ul> </li> <li>`wavenet-s1-sample-*` <ul> <li>Two samples from a WaveNet trained with temporal resolution s=64.</li> </ul> </li> </ul>
NetFlow data collected with different packet sampling rates
<p>NetFlow traffic generated using DOROTHEA (DOcker-based fRamework fOr gaTHering nEtflow trAffic) NetFlow is a network protocol developed by Cisco for the collection and monitoring of network traffic flow data generated. A flow is defined as a unidirectional sequence of packets with some common properties that pass through a network device.</p> <p>NetFlow flows have been captured with different sampling at the packet level. A sampling means that 1 out of every X packets is selected to be flow while the rest of the packets are not valued.</p> <p>The version of NetFlow used to build the datasets is 5.</p>
sample data to reproduce https://github.com/OpenMS/OpenMS/issues/5842
<p>sample data to reproduce https://github.com/OpenMS/OpenMS/issues/5842</p>
Data from: Estimating bee abundance: Can mark-recapture methods validate common sampling protocols?
<p>Wild bees can be essential pollinators in natural, agricultural, and urban systems, but populations of some species have declined. Efforts to assess the status of wild bees are hindered by uncertainty in common sampling methods, such as pan traps and aerial netting, which may or may not provide a valid index of abundance across species and habitats. Mark-recapture methods are a common and effective means of estimating population size, widely used in vertebrates but rarely applied to bees. Here we review existing mark-recapture studies of wild bees and present a new case study comparing mark-recapture population estimates to pan trap and net capture for four taxa in a wild bee community. Net, but not trap, capture was correlated with abundance estimates across sites and taxa. Logistical limitations ensure that mark-recapture studies will not fully replace other bee sampling methods, but they do provide a feasible way to monitor selected species and measure the performance of other sampling methods.</p>
Nuclear Magnetic Resonance Data of Sandstones Samples
<ul> <li> <p>Longitudinal (T1) and transverse (T2) NMR data of 18 sandstone samples</p> </li> <li> <p>Performance comparison for two different NMR devices</p> </li> <li> <p>for details see the readme file</p> </li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.