Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

126

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

126 results for “Extraction Method”

Learn how ShareScore rates datasets ↗
zenodo44/100

A Data Set of 255,000 Randomly Selected and Manually Classified Extracted Ion Chromatograms for Evaluation of Peak Detection Methods

<p>Non-targeted mass spectrometry (MS) has become an important method over the last years in the fields of metabolomics and environmental research. While more and more algorithms and workflows become available to process a large number of data sets nontargeted, there still exist few manually evaluated universal test data sets for refining and evaluating these methods. The first step of non-targeted screening, peak detection (and refinement of it) is arguably the most important step for non-targeted screening. However, the absence of a model data set makes it harder for researchers to evaluate peak detection methods. In this Data Descriptor, we provide a manually checked data set consisting of 255,000 EICs (5000 peaks randomly sampled from across 51 samples) for the evaluation on peak detection and gap filling algorithms. The data set was created from a previous real-world study, of which a subset was used to extract and manually classify ion chromatograms by three mass spectrometry experts. The data set consists of:</p> <ul> <li>51 converted mass spectral files in mzML format</li> <li>An .RData-file containing the extracted ion chromtograms (EICs)</li> <li>The randomly selected subset and the original output table of MZmine in .csv-format</li> <li>Example .xlsx files for the classification</li> <li>2 central classification tables</li> <li>Several tables with additional information about the sampling, chemical analysis and expert jugdement on EICs</li> </ul> <p>For a full description of the experiment and the data set, please read the related Data Descriptor with the title &quot;A data set of 255000 randomly selected and manually classified extracted ion chromatograms for evaluation of peak detection methods&quot; in Metabolites (https://www.mdpi.com/journal/metabolites; DOI: https://doi.org/10.3390/metabo10040162).</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

Formal Methods in Railways: a Systematic Mapping Study - List of Primary Studies and Data Extraction

<p>This Excel file includes the list of papers analyzed in the systematic mapping study titled &quot;Formal Methods in Railways: a Systematic Mapping Study&quot;. The study has been submitted for publication, and its preprint is also included in this repository.&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Burial mounds dataset associated with the paper Geomorphometric methods for tumuli recognition and extraction from high resolution LiDAR DEMs

<p>This dataset consist of the burial (tumuli) mounds delineation and the associated data produced for the article Geomorphometric methods for tumuli recognition and extraction from high resolution LiDAR DEMs, submitted to Sensors. The dataset work in conjunction with the script (http://doi.org/10.5281/zenodo.3628805) to make the work reproductible. The DEM is available only by request to mihai.niculita@uaic.ro.</p>

opencc-by-4.0Jan 2020View details →
zenodo40/100

BioPropaPhenKG on Multi-Relation Extraction Methods on Online Newspapers

<p>The coronavirus disease (COVID-19) spread rampantly around the world at the beginning of 2020 before the governments of each country could prevent it by making decisions based on medical data analysis. With proper formalization, the terabytes of new textual data available online every day could have been used for the early description and detection of cases of this virus. Since then, the number of Event-Based Surveillance (EBS) applications has increased exponentially. These applications aim to mine channels of unstructured data to detect signs of possible public health events. However, one problem with such systems is the need for expert intervention to define which event will be captured, which relevant terms should be used in the search, and to analyze the events to modify the search procedure constantly. Another problem is that many of these applications do not consider both spatial and temporal characteristics. Addressing such limitations, this article presents a novel approach. We propose the biomedical domain specialization of the Core Propagation Phenomenon Ontology (PropaPhen) to capture spatiotemporal characteristics of the propagation of health-related phenomena. We also propose the Description-Detection-Framework (DDF), which leverages PropaPhen, UMLS, and OpenStreetMaps to detect new medical events automatically. Finally, we demonstrate a use case with experiments on extracts from online newspapers about COVID-19. The results show that DDF can be useful for detecting clusters of suspicious cases of possible emerging health-related phenomena.</p> <p>BioPropaPhenKG, its ontology and other useful information can be found in&nbsp;<a href="../records/10911980">https://zenodo.org/records/10911980</a>. The code used for this use case can be found in <a href="https://github.com/Gabriel382/DDPF-Health-Risks">https://github.com/Gabriel382/DDPF-Health-Risks</a> . Finally, the datasets used where UMLS MetamorphoSys, OpenStreetMaps, Wikidata, <a href="https://aylien.com/blog/free-coronavirus-news-dataset">Aylien</a> (data only from November of 2019).</p> <p>&nbsp;</p> <p>To read, you just need to load it with Neo4j:4.4.3. Alternatively, you can open it with docker using the following command:&nbsp;</p> <p>docker run --interactive --tty --rm \<br>&nbsp; &nbsp; --publish=7474:7474 --publish=7687:7687 \<br>&nbsp; &nbsp; --volume=/path-to-data-folder:/data --user="$(id -u):$(id -g)"\<br>&nbsp; &nbsp; neo4j:4.4.3 \<br>neo4j-admin load --from=/data/BioPropaPhenKG-Journal-MultiRE.dump --database "neo4j" --force</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

All raw data for Fukuzawa, T et al. "Environmental DNA extraction method for a high and stable DNA yield"

<p>The all raw data of quantitative PCR for environmental DNA in&nbsp;Fukuzawa, T et al. &quot;Environmental DNA extraction method for a high and stable DNA yield&quot;.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Figure 6. (a1), (a2), (a3), (a4), (a5), (a6), (a7) and (a8) watermarked image is degraded respectively through JPEG2000 compression, JPEG compression, median filtering, adding Salt&Pepper noise, rotating, center cropping, surrounding cropping and scaling. (b1), (b2), (b3), (b4), (b5), (b6), (b7) and (b8) The corresponding extracted watermarks.-Discrete Wavelet Transform Method: A New Optimized Robust Digital Image Watermarking Scheme

<p>This paper has described a scheme for digital watermarking of still images based on discrete<br> wavelet transform. In the proposed method, the embedded logo watermark can be extracted without<br> access to the original image. It has been confirmed that the proposed watermarking method is able<br> to extract the embedded logo watermark from the watermarked images that have degraded through<br> compression, filtering, cropping and scaling. Although this algorithm is not robust against rotation,<br> it can completely extract the watermark from watermarked images that lose about 35% of their<br> areas by cropping attack.</p>

opencc-by-4.0Jun 2012View details →
zenodo40/100

Figure 1. (a) Original watermark (b) extracted watermarks after compression(c) merged watermark-Discrete Wavelet Transform Method: A New Optimized Robust Digital Image Watermarking Scheme

<p>Therefore, each bit of the logo watermark is stored in one coefficient of a sub-block to keep<br> the capacity of watermarking fixed.<br> When a region of the watermarked image is destroyed; the whole watermark can be<br> extracted using other regions of the watermarked image by merging extracted watermarks. Figure 1<br> shows result of merging logo watermarks that were extracted from a compressed (with JPEG2000<br> algorithm) watermarked image.</p>

opencc-by-4.0Jun 2012View details →
zenodo40/100

Figures 2–6 in A reliable and efficient BioPulverizer method in preparing and grinding nematodes for nucleic acid extraction and molecular identification

Figures 2–6 PCR amplification product with primers: (2) PCR products from the seminested primer pairs (NemF and 18Sr2b; NF1 and 18Sr2b) obtained from soil nematode samples with BioPulverizer grinding; (3) PCR amplification with the first cycle of the primer pair (NemF and 18Sr2b) from soil samples without BioPulverizer grinding; (4–5) Two amplification bands, 181 bp with the species-specific primer pair GlyF1/rDNA2 (4) and 477 bp with SCNF1/SCNR1 (5), were amplified for Heterodera glycines; (6) PCR products with the universal primer pair 194F/195R and the species-specific primer pair (Meloidogyne incognita) from potato tuber samples. All molecular markers (M) are 100 bp DNA ladders. The concentrations of agarose gel are 1.8% in Figs 2–3 and 1% in Figs 4–6.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figure 1 in A reliable and efficient BioPulverizer method in preparing and grinding nematodes for nucleic acid extraction and molecular identification

Figure 1 Equipment used for nematode preparation and grinding. The names of all the equipment are listed above or under each respective one.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figure 1 in A Damage-Limiting Method for Extracting Bark and Ambrosia Beetles (Coleoptera: Curculionidae: Scolytinae) from Their Tunnels in Host Plants of Conservation Concern

Figure 1. Use of sticky cockroach traps to extract the Hawaiian endemic bark beetle Xyleborus mauiensis Perkins, 1900, from a Hawaiian endemic olapa (Cheirodendron trigynum) tree on the island of Lanai. a: Section of sticky trap ready for action; note the glue pushed to one side of the strip to form a globule; b: Beetle stuck to the glue and extracted from the wood.

opencc-by-4.0Dec 2020View details →
zenodo40/100

Figure 1. DNA extraction with two different protocols from different noninvasive samples. Lines 1, 3, 5 in Evaluation of methods for molecular sex-typing of three heron species from different DNA sources

Figure 1. DNA extraction with two different protocols from different noninvasive samples. Lines 1, 3, 5, and 7: DNA extraction with commercial kit; Lines 2, 4, 6, and 8: DNA extracted with modified standard protocol. Lines 1–2: eggshells (Grey Heron); lines 3–4: eggshell swabs (Grey Heron); lines 5–6: pin feathers (Purple Heron); lines 7–8: contour feathers (Great Egret); 9: negative control; M: molecular marker.

opencc-by-4.0Jan 2017View details →
zenodo40/100

Quantification of Fatty Acids in Hemp Seeds (Cannabis sativa L.) and Yield Prediction Using Machine Learning for Soxhlet and Ultrasound Extraction Methods

<p>This study focuses on the quantification of fatty acids present in hemp seeds (Cannabis sativa L.) cultivated in the Ecuadorian Andes using Soxhlet and ultrasound extraction methods. The aim is to evaluate and compare the extraction efficiency of these two techniques. Furthermore, machine learning models are applied to predict extraction yields based on experimental conditions. Using locally cultivated seeds provides valuable insights into the influence of regional agro-climatic conditions on the chemical composition. The integration of predictive algorithms offers a novel approach to optimizing the extraction process, enhancing both precision and efficiency. The findings could contribute to developing sustainable extraction methods for high-value bioactive compounds in the food and pharmaceutical industries.</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Retinal status analysis method based on feature extraction and quantitative grading in OCT images

<p>The raw database includes 200 retinal OCT&nbsp;images judged as normal by ophthalmologists and 100 images with various abnormalities. The software includes the main steps for retinal status analysis. The software was carried out in Matlab.</p>

opencc-zeroJun 2016View details →
zenodo36/100

Application of high-throughput sequencing (HTS) metabarcoding to diatom biomonitoring: Do DNA extraction methods matter?

<p>The 8 benthic samples from Mainland France (stream Edian, stream Aire, lake Geneva), Sweden (stream Dåmman, Agricultural stream, lake Båtkåjaure) and Mayotte (stream Dapani, stream Majimbini) were collected by scraping material from the surface of stones, following the French standard (AFNOR 2007) used in routine biomonitoring programs.DNA was extracted from each sample (2 replicates) using five DNA extraction methods, followed by the amplification of a short rbcL DNA barcode (312bp) specific to diatoms. PCR products were then sequenced in one random direction using the Ion Torrent™ Personal Genome Machine® (PGM) System according to the manufacturer’s instructions. The data file contains one fastq file per library sequenced with the raw DNA reads, as provided by the sequencing platform (demultiplexing performed by the sequencing platform). An excel file is also provided to make the link between the fastq file number and the sample information (sample origin, DNA extraction method used, number of raw reads).</p>

opencc-by-4.0Nov 2016View details →
zenodo36/100

Methods for Extracting and Characterizing RNA from Urine: for downstream PCR and RNAseq Analysis

<p>Readily accessible samples such as urine or blood are seemingly ideal for differentiating and stratifying patients, however, it has proven a daunting task to identify reliable biomarkers in such samples. Noncoding RNA holds great promise as a source of biomarkers distinguishing physiologic wellbeing or illness.</p>

opencc-by-4.0Jun 2017View details →
dryad36/100

Data for: A new threshold selection method for species distribution models with presence-only data: extracting the mutation point of the P/E curve by threshold regression

<p>Selecting thresholds to convert continuous predictions of species distribution models proves critical for many real-world applications and model assessments. Prevalent threshold selection methods for presence-only data require unproven pseudo-absence data or subjective researchers' decisions. This study proposes a new method, Boyce-Threshold Quantile Regression (BTQR), to determine thresholds objectively without pseudo-absence data. We summarize that the mutation point is a typical shape feature of the predicted-to-expected (P/E) curve after reviewing relevant articles. Analysis based on source-sink theory suggests that this mutation point may represent a transition in habitat types and serve as an appropriate threshold. Threshold regression is introduced to accurately locate the mutation point.</p> <p>To validate the effectiveness of BTQR, we used four virtual species of varying prevalence and a real species with reliable distribution data. Six different species distribution models were employed to generate continuous suitability predictions. BTQR and nine other traditional methods transformed these continuous outputs into binary results. Comparative experiments show that BTQR has advantages in terms of accuracy, applicability, and consistency over the existing methods.</p>

opencc-zeroMar 2024View details →
zenodo36/100

Supplementary Table S1 (raw data) of "Filtration extraction method using microfluidic channel for measuring environmental DNA "

<p>Supplementary Table S1 (all&nbsp;raw data)&nbsp;of &quot;Filtration extraction method using microfluidic channel for measuring environmental DNA &quot;. Each data of the validation experiment; Experiment 1-4 was located in different sheets..</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Data Set used in "Full backward and forward dependencies through regional hypothetical extraction method"

<p>This set of data was obtained from&nbsp;EUREGIO database, developed by the Tinbergen Institute, which&nbsp;is a set of global IO tables with regional and sectoral disaggregation. The EUREGIO database collects the productive structure and commercial relations of the WIOD in the period 2000-2010. The table is broken down into 249 administrative regions at the NUTS2 level, from 24 EU countries, 16 non-EU countries, and a block that brings together countries from the rest of the world, making a total of 266 regions. The statistical information is organised in 11 IO tables, one for each year.</p> <p>The data base that we provid in this repository is used in our study with the aim&nbsp;to determine the key regions of the Spanish economy. In order to address this objective, IO tables of smaller dimensions are built, through an aggregation and disaggregation procedure. First, the 14 industries are grouped, then the 4 sectors of final demand and, lastly, the 4 components of value added. Below, the 266 EUREGIO regions are grouped into 21 regions. Of these, 19 regions correspond to Spain&nbsp;<a href="#_ftn1">[1]</a>, one region includes the rest of the NUTS2 in the EU and another region covers the rest of the world.</p> <p><a href="#_ftnref1">[1]</a> The 17 Spanish regions and the two autonomous cities of Ceuta and Melilla.</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Formal Methods for NFA Equivalence: QBFs, Witness Extraction, and Encoding Verification

<p>Supplemental material to the paper.</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Legal Case Document Summarization: Extractive and Abstractive Methods and their Evaluation

<p>This repository contains the following 3&nbsp;datasets for legal document summarization :</p> <p>- IN-Abs : Indian Supreme Court case documents &amp; their `abstractive&#39; summaries, obtained from http://www.liiofindia.org/in/cases/cen/INSC/<br> - IN-Ext : Indian Supreme Court case documents &amp; their `extractive&#39; summaries, written by two law experts (A1, A2).<br> - UK-Abs : United Kingdom (U.K.) Supreme Court case documents &amp; their `abstractive&#39; summaries, obtained from https://www.supremecourt.uk/decided-cases/</p> <p>Please refer to the paper and the README file for more details.</p>

opencc-by-4.0Nov 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record