Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
355
datasets available to search
ShareScore release 0.9.0
Dataset results
355 results for “data extraction”
Marcell Experimental Forest peat core extraction chemical analysis data (DOC, Fe, Ca, Mg, K, P, Al)
This data set reports iron (Fe), dissolved organic carbon (DOC), calcium (Ca), magnesium (Mg), potassium (K), phosphorus (P), and aluminum (Al) measured in extractions of soil cores sampled from two boreal peatlands, the S1 and S2 bogs, in the Marcell Experimental Forest (MEF) in Itasca County, Minnesota. The soil cores were sampled on September 2, 2017. Elements were quantified in extractions with hydrochloric acid, sodium dithionite, sodium sulfate, and sodium dithionite plus hydrochloric acid to examine how iron influences carbon and nutrient cycling in peatlands. The S1 and S2 sites are research catchments instrumented for hydrologic monitoring. The S1 bog is also the location of the Spruce and Peatland Responses Under Changing Environments (SPRUCE) experiment. These data are used, analyzed, and reported in Curtinrich et al. (2021, Ecosystems).
Data for "A Dataset for Multi-lingual Epidemiological Event Extraction"
<p>This is the data for the LREC 2020 paper "<a href="https://zenodo.org/record/3693647">A Dataset for Multi-lingual Epidemiological Event Extraction</a>". If you use this resource, please cite the paper:</p> <pre><code>@inproceedings{mutuvi2020dataset, title = "A Dataset for Multi-lingual Epidemiological Event Extraction", author = {Mutuvi, Stephen and Doucet, Antoine, and Lejeune, Gaël and Odeo, Moses}, booktitle = "Proceedings of the Eighth International Conference on Language Resources and Evaluation ({LREC}'20)", year = "2020", location = "Marseille" }</code></pre> <p> </p> <p>This work has been supported by the European Union Horizon 2020 research and innovation programme under grants 825153 (Embeddia) and 770299 (NewsEye).</p>
Data from: Optimisation of biogenic synthesis of silver nanoparticles from flavonoid-rich Clinacanthus nutans leaf and stem aqueous extracts
<p>BACKGROUND: Silver nanoparticles (AgNPs) are widely used in food industries, biomedical, dentistry, catalysis, diagnostic biological probes, and sensors. The use of plant extract for AgNPs synthesis eliminates the process of maintaining cell culture and the process could be scaled up under a non-aseptic environment. The purpose of this study is to determine the classes of phytochemicals, to biosynthesise and characterise the AgNPs using Clinacanthus nutans leaf and stem extracts. In this study, AgNPs was synthesised from the aqueous extracts of C. nutans leaves and stems through a non-toxic, cost effective and eco-friendly method.</p> <p>RESULTS: The formation of AgNPs was confirmed by UV-Vis spectroscopy, and the size of AgNP-L (leaf) and AgNP-S (stem) were 114 and 129 nm, respectively. Transmission electron microscopy (TEM) analysis showed spherical nanoparticles with AgNP-L and AgNP-S ranging from 10-300 nm and 10-180 nm; with zeta potentials of AgNP-L and AgNP-S at -42.8 and -43.9 mV, respectively. XRD analysis matched the face-centred cubic structure of silver and was capped with bioactive compounds. FTIR analysis revealed the presence of few functional groups of phenolic and flavonoid compounds. These functional groups act as reducing agents in AgNPs synthesis.</p> <p>CONCLUSION: This result showed that the biogenically synthesised nanoparticles reduced silver ions to silver nanoparticles in aqueous condition and the AgNPs formed were stable and less toxic.</p>
Table S1: Data extraction table. A summary of the 166 candidate biomarkers that met the inclusion/exclusion criteria.
<p>Supplementary material to Translating Biomarkers of Cholangiocarcinoma for Theranosis: A Systematic Review. Version 2. </p>
Source code for models of floral initiation in pea and gene expression data extracted from published sources
<p>The dataset contains the source code for computational models of a gene network controlling transition to flowering in pea (<em>Pisum sativum</em>). The models were based on ordinary differential equations (ODE) or neural networks. It also includes data on the expression dynamics of genes involved in the network, which was used for model fitting. The expression data was extracted from the following papers: </p> <p>Hecht, V., Laurie, R. E., Schoor, K. Vander, Ridge, S., Knowles, C. L., Liew, L. C., Sussmilch, F. C., et al. (2011). The Pea GIGAS Gene Is a FLOWERING LOCUS T Homolog Necessary for Graft-Transmissible Specification of Flowering but Not for Responsiveness to Photoperiod. 23, 147–161. doi:10.1105/tpc.110.081042</p> <p>Sussmilch, F. C., Berbel, A., Hecht, V., Schoor, K. Vander, Ferrándiz, C., Madueño, F., et al. (2015). Pea VEGETATIVE2 Is an FD Homolog That Is Essential for Flowering and Compound In fl orescence Development. 27, 1046–1060. doi:10.1105/tpc.115.136150</p> <p>The source code of the DEEP software used for parameter optimization in the model fitting can be found in the Gitlab repository (https://gitlab.com/mackoel/deepmethod/-/tree/master).</p> <p>The files are the supplement to the following manuscript, submitted to Frontiers in Genetics:</p> <p>"Dynamical Modeling of the Core Gene Network Controlling Transition to Flowering in <em>Pisum sativum</em>" by Polina Pavlinova, Maria G. Samsonova, and Vitaly V. Gursky.</p> <p>All possible questions can be sent to: Polina Pavlinova (polina.pavlina1004@gmail.com), Vitaly Gursky (gursky@math.ioffe.ru).</p>
Europe Road Network extracted from OpenStreetMap data
<p>The data extracts of Europe region downloaded on 04/07/2020 was used to create this dataset. From this the <strong>highways</strong> tagged as <strong>motorway, trunk, primary, secondary, tertiary, unclassified</strong> and <strong>residential </strong>are selected and the information was saved as line strings. CRS: WGS84 (EPSG:4326)<br> Files available are in parquet and csv format. Please feel free to convert the files in to desired file formats.</p>
Data from Nicolle et al. LC-HRMS study of Streptomyces sp. AgN23 Culture Media Extract. Study of AgN23 exometabolome and analysis of Arabidopsis metabolomic responses to the bacteria
<p>This archive compiles several datasets related to studies of <i>Streptomyces</i> sp. AgN23 interaction with <i>Arabidopsis thaliana</i>. Ultra-high-performance liquid chromatography-high-resolution MS (UHPLC-HRMS) analyses were performed on a Q Exactive Plus quadrupole (Orbitrap) mass spectrometer, equipped with a heated electrospray probe (HESI II) coupled to a U-HPLC Ultimate 3000 RSLC system (Thermo Fisher Scientific, Hemel Hempstead, United Kigdom). For each biological sample, the RAW file obtained in ESI+ and ESI- mode were retrieved from the Xcalibur version 4.4 software and are deposited in separate sub-folders termed "RawPos" and "RawNeg". Each experimental cohort is grouped in a folder where the biological repeats can be retrieved, as well as QC (Quality Check, pool of all samples from the cohort), Blank samples and eventual alternative control such as Bennett, the mock control media of <i>Streptomyces</i> sp. AgN23. The details regarding samples preparation, analytic parameters and mass spectrometry, statistical treatment and visualization of the data will be made available in the publication relating to this archive. The folder " AgN23-WT_AgN23-pSC004" contains chromatograms related to metabolomic study of Wild-type and pSC004-1, pSC004-10, pSC004-16 and pSC004-22 mutants of <i>Streptomyces</i> sp. AgN23. The folder " Col-0_AgN23" contains chromatograms related to metabolomic study of <i>Arabidopsis thaliana</i> Col-0 responses to colonization by <i>Streptomyces</i> sp. AgN23-WT. The folder " Col-0_pad3-1_AgN23" contains chromatograms related to metabolomic study of <i>Arabidopsis thaliana</i> Col-0 and the <i>Arabidopsis</i> pad3-1 mutant responses to colonization by <i>Streptomyces</i> sp. AgN23-WT. The folder " Col-0_pSC004" contains chromatograms related to metabolomic study of <i>Arabidopsis thaliana</i> Col-0 responses to colonization by <i>Streptomyces</i> sp. AgN23-WT and the AgN23 mutants pSC004-10 and pSC004-22. It should be noted that in the publication associated with this archive, the pSC004-1, pSC004-10, pSC004-16 and pSC004-22 mutants are referred to as ΔgbnB-1, ΔgbnB-2, ΔgbnB-3 and ΔgbnB-4, respectively.</p>
Data for: A new threshold selection method for species distribution models with presence-only data: extracting the mutation point of the P/E curve by threshold regression
<p>Selecting thresholds to convert continuous predictions of species distribution models proves critical for many real-world applications and model assessments. Prevalent threshold selection methods for presence-only data require unproven pseudo-absence data or subjective researchers' decisions. This study proposes a new method, Boyce-Threshold Quantile Regression (BTQR), to determine thresholds objectively without pseudo-absence data. We summarize that the mutation point is a typical shape feature of the predicted-to-expected (P/E) curve after reviewing relevant articles. Analysis based on source-sink theory suggests that this mutation point may represent a transition in habitat types and serve as an appropriate threshold. Threshold regression is introduced to accurately locate the mutation point.</p> <p>To validate the effectiveness of BTQR, we used four virtual species of varying prevalence and a real species with reliable distribution data. Six different species distribution models were employed to generate continuous suitability predictions. BTQR and nine other traditional methods transformed these continuous outputs into binary results. Comparative experiments show that BTQR has advantages in terms of accuracy, applicability, and consistency over the existing methods.</p>
Automatic User Story Generation: A Comprehensive Systematic Literature Review - Data Extraction
<p>This document presents the data extraction performed for the Systematic Literature Review in Automatic User Story Generation.</p>
Data from: Extraction and identification of a wide range of microplastic polymers in soil and compost
<p>Microplastic (MP) pollution is globally widespread, however their presence in soil systems is poorly understood due to complexity of soil and lack of standardised extraction methods. Datasets provided contain data from optimisation (recoveries) of MPs extraction protocol from soil and compost based on olive oil and density separation using zinc chloride, in case of low-density polyethylene and polyethylene terephthalate. Density separation was further used to extract five microplastic polymers (PET, PS, PE, PP and PVC), added to soil and compost at different concentrations, which is also included in the dataset. Additionally, identification data of extracted MPs from spiked soil and compost samples.</p>
Assessment of animal diseases caused by bacteria resistant to antimicrobials: Cattle- Appendix B: Excel file with all data extracted
<p>Information on all the full-text studies that were assessed, including the reason for exclusion for those that were excluded at the full-text screening and the data extracted from the included studies, can be consulted here. </p> <p>The extensive literature review was carried out by the University of Copenhagen under the contract OC/EFSA/ALPHA/2020/02 – LOT 1 (https://ted.europa.eu/udl?uri=TED:NOTICE:457654-2020:TEXT:EN:HTML)</p>
Supplementary Table S1 (raw data) of "Filtration extraction method using microfluidic channel for measuring environmental DNA "
<p>Supplementary Table S1 (all raw data) of "Filtration extraction method using microfluidic channel for measuring environmental DNA ". Each data of the validation experiment; Experiment 1-4 was located in different sheets..</p>
Wilcoxon Rank Sum Test and Keyphrase Extraction Data Cited in "What Everyone Says: Public Perceptions of the Humanities in the Media"
<p>This repository contains Wilcoxon rank sum test and keyphrase extraction data cited in the WhatEvery1Says (WE1S) Project's article "What Everyone Says: Public Perceptions of the Humanities in the Media". The organization of the materials is discussed below.</p> <p><strong>Wilcoxon Rank Sum Test</strong></p> <p>All data and results from Wilcoxon rank sum testing can be found in the <code>wilcoxon-tests</code> folder of the extracted zip file <code>we1s_about_the_humanities.zip</code>. The Wilcoxon rank sum test identifies specific words that appear significantly more in one group of documents as compared to another, thus providing researchers with an understanding of what words are “distinctive” to each group. Further information on WE1S's use of Wilcoxon rank sum testing can be found at <a href="https://we1s.ucsb.edu/wp-content/uploads/M-15-Wilcoxon-Test.pdf">https://we1s.ucsb.edu/wp-content/uploads/M-15-Wilcoxon-Test.pdf</a>.</p> <p>Each subdirectory in the <code>wilcoxon-test</code> folder contains the data and results of a particular comparison experiment based on a metadata category such as whether the data contained articles published by public or private institutions. Each data file is a <code>.txt</code> file representing a sample of the overall data from the collection. The <code>README</code> file provides information on the collection used, the sample size, and the nature of the comparison. The results for the test are in a file called <code>results.csv</code>.</p> <p>The <code>results.csv</code> file for each test includes a row for each term included in the test. Each row displays the term, the term's raw count in each category compared (count 1 and count 2), the difference between the 2 counts (count 1 minus count 2), the percentage change in the counts, the Wilcoxon statistic, and the Wilcoxon p-value. Sorting the csv by the Wilcoxon stat from greatest to least will cause the terms most strongly associated with category 1 to come to the top (category 1 is the category listed first in the title field of the README.md file for each test), while sorting it by the Wilcoxon stat from least to greatest will cause the terms most strongly associated with category 2 to come to the top (category 2 is the category listed second). The p-value column provides you with information about how confident you can be about each comparison's significance.</p> <p><strong>Keyphrase Extraction</strong></p> <p>All data and results from Wilcoxon rank sum testing can be found in the <code>keyphrase-extraction</code> folder of the extracted zip file <code>we1s_about_the_humanities.zip</code>. Keyphrase extraction generates a list of the most significant words or phrases (1-6 words long) within individual documents. WE1S takes the top ten keyphrases in each document and ranks them according to their frequency across the collection. WE1S uses the SGRank algorithm for keyphrase extraction, and because this algorithm is computationally intensive, WE1S limits keyphrases to lemmatized nouns and proper nouns within a window of 70 words to either side of candidate keyphrases. Further information on WE1S's use of keyphrase extaction can be found at <a href="https://we1s.ucsb.edu/wp-content/uploads/M-14-Keyphrase-Extraction.pdf">https://we1s.ucsb.edu/wp-content/uploads/M-14-Keyphrase-Extraction.pdf</a>.</p> <p>Each subdirectory in the <code>keyphrase-extraction</code> folder contains the data and results of keyphrase extraction on a particular collection. Details of the collection and resulting files can be found in each subdirectory. Each list of keyphrases is in a file called <code>SGRank.csv</code>, which lists the keyphrases and their number of occurrences in the collection. The article additionally cites keyphrases that are shared with the terms in the public topic model produced by Andrew Goldstone and Ted Underwood, “The Quiet Transformations of Literary Studies: What Thirteen Thousand Scholars Could Tell Us,” <em>New Literary History</em> 45, no. 3 (2014): 359–84, <a href="https://doi.org/10.1353/nlh.2014.0025">https://doi.org/10.1353/nlh.2014.0025</a>. The list of terms is derived from the public visualization at <a href="https://www.sas.rutgers.edu/virtual/ag978/quiet/#/words">https://www.sas.rutgers.edu/virtual/ag978/quiet/#/words</a>. Keyphrases extracted from WE1S data were split into single-word terms and compared with the list of vocabulary in Goldstone and Underwood's word list (<code>quiet_transformations_wordlist.txt</code>) to compile lists of shared vocabulary. These lists are given in files called <code>shared_terms.txt</code>.</p> <p>Note that keyphrases were extracted for corpora produced using the Python <a href="https://textacy.readthedocs.io/en/latest/index.html">Textacy</a> library. Because these corpora contain the full text of articles with intellectual property restrictions they cannot be reproduced here.</p>
rime fraction training data set extracted from BAECC
<p>training data set used in Vogl et al. (https://amt.copernicus.org/preprints/amt-2021-137/) to derive rime mass fraction from Doppler cloud radar observations. Extracted from the BAECC data set.</p> <p>rime mass fraction retrieved from PIP data</p> <p>cloud radar observations at Ka- and W-band (ARM KAZR and MWACR)</p> <p>attenuation estimated using the Passive and Active Microwave Remote Sensing Tool (PAMTRA)</p>
Assessment of animal diseases caused by bacteria resistant to antimicrobials: Fishes - Appendix B: Excel file with all data extracted
<p>Information on all the full-text studies that were assessed, including the reason for exclusion for those that were excluded at the full-text screening and the data extracted from the included studies, can be consulted here. </p> <p>The extensive literature review was carried out by the University of Copenhagen under the contract OC/EFSA/ALPHA/2020/02 – LOT 1 (https://ted.europa.eu/udl?uri=TED:NOTICE:457654-2020:TEXT:EN:HTML)</p>
Assessment of animal diseases caused by bacteria resistant to antimicrobials: Rabbits - Appendix B: Excel file with all data extracted
<p>Information on all the full-text studies that were assessed, including the reason for exclusion for those that were excluded at the full-text screening and the data extracted from the included studies, can be consulted here. </p> <p>The extensive literature review was carried out by the University of Copenhagen under the contract OC/EFSA/ALPHA/2020/02 – LOT 1 (https://ted.europa.eu/udl?uri=TED:NOTICE:457654-2020:TEXT:EN:HTML)</p>
Nanopore MinION Run Metrics and genomic DNA fragment size analysis data from automated phenol-chloroform extractions (RBI LabDroid Maholo)
<p>Nanopore MinION run MinKNOW statistical metrics output, Agilent Femto Pulse and Tape Station gDNA fragment size analysis reports of genomic DNA isolated from automated RBI LabDroid Maholo organic extractions.</p>
Data set for "Token-Level Multilingual Epidemic Dataset for Event Extraction"
<p>This is the data for the TPDL 2021 paper "<a href="https://zenodo.org/record/5780020">Token-Level Multilingual Epidemic Dataset for Event Extraction</a>". If you use this resource, please cite the paper:</p> <pre><code>@inproceedings{mutuvi2021dataset, title = "Token-level Multilingual Epidemic Dataset for Event Extraction", author = {Mutuvi, Stephen and Boros, Emanuela and Doucet, Antoine, and Lejeune, Gaël and Jatowt, Adam and Odeo, Moses}, booktitle = "Proceedings of the 25th International Conference on Theory and Practice of Digital Libraries, September 13–17, 2021, TPDL 2021", year = "2021", location = "Online" }</code></pre> <p> </p> <p>This work has been supported by the European Union Horizon 2020 research and innovation programme under grants 825153 (Embeddia) and 770299 (NewsEye).</p>
Labeled data for citation field extraction
<p>Citations are an important part of scientific papers, and the proper handling of them is indispensable for the science of science. Citation field extraction is the task of parsing citations: given a citation string, extract authors, title, venue, doi etc. Since the number of citations is counted by hundreds millions, efficient computer based methods for this task are very important.</p> <p>The development of machine learning methods for citation field extraction requires ground truth: a large corpus of labeled citations. This dataset provides a very large (41M) corpus of labeled data obtained by the reverse process: we took structured citation lists and used BibTeX to generate labeled citation strings.</p>
Data Set used in "Full backward and forward dependencies through regional hypothetical extraction method"
<p>This set of data was obtained from EUREGIO database, developed by the Tinbergen Institute, which is a set of global IO tables with regional and sectoral disaggregation. The EUREGIO database collects the productive structure and commercial relations of the WIOD in the period 2000-2010. The table is broken down into 249 administrative regions at the NUTS2 level, from 24 EU countries, 16 non-EU countries, and a block that brings together countries from the rest of the world, making a total of 266 regions. The statistical information is organised in 11 IO tables, one for each year.</p> <p>The data base that we provid in this repository is used in our study with the aim to determine the key regions of the Spanish economy. In order to address this objective, IO tables of smaller dimensions are built, through an aggregation and disaggregation procedure. First, the 14 industries are grouped, then the 4 sectors of final demand and, lastly, the 4 components of value added. Below, the 266 EUREGIO regions are grouped into 21 regions. Of these, 19 regions correspond to Spain <a href="#_ftn1">[1]</a>, one region includes the rest of the NUTS2 in the EU and another region covers the rest of the world.</p> <p><a href="#_ftnref1">[1]</a> The 17 Spanish regions and the two autonomous cities of Ceuta and Melilla.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.