Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
355
datasets available to search
ShareScore release 0.9.0
Dataset results
355 results for “data extraction”
Some original, intermediate, and result data in the papar entited "Highway marking extraction and degradation analysis by using MLS point clouds"
<p>Some original, intermediate, and result data in the papar entited "Highway marking extraction and degradation analysis by using MLS point clouds"</p>
Lakkasuo carbon isotope and AWEN extraction data for Yasso-C13 model development
<p>This package contains carbon isotope and AWEN extraction data from litterbag experiments (5 year) at Lakkasuo, a raised bog complex near Hyytiälä weather station in Finland. Additionally present are driving data and parameter values needed to run soil carbon model Yasso. The given data is used to implement and calibrate carbon-13 related soil organic matter decomposition in the Yasso model. The dataset also contains calibration results, scripts to run the results anew and to produce plots and images. The updated dataset uses new Yasso20 parameter values.</p>
Regeneration data of bryophyte fragments extracted from feces of Chloephaga picta and Attagis malouinus in Navarino Island, sub-Antarctic Chile
<p class="Normal1"><span><span><span><span><span><span><span><span><span><span><span>Birds are known to act as potential vectors for the exogenous dispersal of bryophyte diaspores. Given the totipotency of vegetative tissue of many bryophytes, birds could also contribute to endozoochorous bryophyte dispersal. Research has shown that fecal samples of the upland goose (<i>Chloephaga picta</i>) and white-bellied seedsnipe (<i>Attagis malouinus</i>) contain bryophyte fragments. Although few fragments from bird feces have been known to regenerate, the evidence for the viability of diaspores following passage through the bird intestinal tract remains ambiguous. We evaluated the role of endozoochory in these same herbivorous and sympatric bird species in sub-Antarctic Chile. We hypothesized that fragments of bryophyte gametophytes retrieved from their feces are viable and capable of regenerating new plant tissue. Eleven feces samples containing undetermined moss fragments from <i>C. picta</i> and <i>A. malouinus</i>, six moss fragment samples from wild collected mosses (<i>Conostomum </i><i>tetragonum</i>,<i> Syntrichia </i><i>robusta</i>,<i> </i>and <i>Polytrichum </i><i>strictum</i>), and one spore sample from <i>C. tetragonum </i>were grown ex situ in peat soil and in vitro<i> </i>using a Gamborg (agar) medium. After 91 days, 20% of fragments from <i>A. malouinus </i>feces, 50% of fragments from <i>C. picta </i>feces, and 57% of propagules from wild mosses produced new growth. The fact that moss diaspores remained viable and can regenerate under experimental conditions following the passage through the intestinal tracts of these robust fliers and altitudinal and latitudinal migrants, suggests that sub-Antarctic birds may play a </span></span></span></span></span></span></span></span></span></span></span><span><span><span><span><span><span><span><span><span><span><span>critical, role in bryophyte dispersal. This relationship may have important implications in the way bryophytes disperse and colonize facing climate change.</span></span></span></span></span></span></span></span></span></span></span></p>
Collection of Extracted Data as part of the Publication "(Why) Do Taint Analyzers Fail in Practice? A Critical Literature Review"
<p>This document provides an extensive list of the data extraction process for the systematic literature review. This anonymous version is review-friendly for double-blinded peer reviews.</p>
A novel technique to simulate and characterize a yarn's mechanical behavior based on a geometrical fiber model extracted from micro-CT imaging: geometry and simulation data
<p>This dataset contains the original µCT scan data, the scripts and intermediate results for the generation of the geometrical fiber model, as well as the structural simulation files and their experimental validation data described in the paper <a href="https://journals.sagepub.com/doi/10.1177/00405175221137009">"A novel technique to simulate and characterize a yarn's mechanical behavior based on a geometrical fiber model extracted from micro-CT imaging"</a>, published in Textile Research Journal.</p>
SSP2017 - Experiment Data - Towards Extracting Realistic User Behavior Models
<p>This package contains the monitoring data, the ideal behavior models, computed clustering results for behavior models based on the monitoring data, and the interpretation of the analysis results. We processed these data with the following two tooling sources:</p> <p>Experiment setup: https://doi.org/10.5281/zenodo.883069</p> <p>Analysis software: https://doi.org/10.5281/zenodo.883061</p>
Extracted data for meta-analysis: Distal versus proximal radial access in coronary angiography
<p>Data base underlying quantitative meta-analysis in the manuscript titled "Distal versus proximal radial access in coronary angiography: A meta-analysis" by Lueg, Schulze, Stöhr, & Leistner.</p> <p>Data allows meta-analysis of primary endpoint (RAO), secondary endpoints, meta-regression for moderator analysis, and publication bias analysis.</p>
Data set for the figures in the manuscript "Real-Time Identification of Aerosol-Phase Carboxylic Acid Production Using Extractive Electrospray Ionization Mass Spectrometry"
Open the record for dataset details and reuse information.
Data for: Environmental DNA storage and extraction method affects detectability for multiple aquatic invasive species
<p>Environmental DNA (eDNA) refers to genetic material released by organisms into their surrounding environment. Collecting and identifying eDNA has gained popularity for monitoring and surveillance of aquatic invasive species. Invasive species management is most successful when an invasion is identified early while population size is likely to be low, highlighting the importance of eDNA detection sensitivity. Various factors influence DNA yield recovered from environmental samples. Environmental DNA storage and extraction methods, for example, can be adjusted to maximize DNA yield, thereby improving detectability. In this study, we compared the performance of two eDNA storage and extraction methods in detecting three common aquatic invasive species (<em>Bythotrephes longimanus</em>, <em>Dreissena polymorpha</em>, and <em>Faxonius rusticus</em>) across five natural ecosystems of Minnesota, United States. One method involved storing filters in 95% ethanol (EtOH) and extracting DNA using a DNeasy PowerSoil Pro Kit (Qiagen, Hilden, Germany), whereas the other method used cetyl trimethylammonium bromide (CTAB) for storage and a phenol–chloroform–isoamyl (PCI) procedure for DNA extraction. We also investigated the effect of DNA extract volume (1 μL relative to 3 μL) in qPCR reactions on eDNA detections for the commercial kit method. The CTAB‐PCI method yielded significantly more positive detections, across all three species, compared to the EtOH‐Qiagen method. Moreover, we found that using 1 μL of DNA extract in qPCR reactions was equally effective as using 3 μL. To improve detections of aquatic invasive species, we recommend that researchers store eDNA sample filters in CTAB or a similar lysis buffer such as Longmire's solution and extract with PCI when feasible, but note that lower extract volumes might be used without negative effect when either increasing technical replicates or repurposing samples for the detection of multiple species.</p>
NED data for the paper Into the Single Cell Multiverse: an End-to-End Dataset for Procedural Knowledge Extraction in Biomedical Texts
<p>This data repository contains NED data from the paper, <em><a href="https://arxiv.org/abs/2309.01812">Into the Single Cell Multiverse: an End-to-End Dataset for Procedural Knowledge Extraction in Biomedical Texts.</a></em></p> <p>Additional data for the NER classification task can be found here: <a href="../records/10050681">zenodo</a></p> <p> </p> <p> </p>
Data from: Minimally destructive hDNA extraction method for retrospective genetics of pinned historical Lepidoptera specimens
<p>The millions of specimens stored in entomological collections provide a unique opportunity to study historical insect diversity. Current technologies allow to sequence entire genomes of historical specimens and estimate past genetic diversity of present-day endangered species, advancing our understanding of anthropogenic impact on genetic diversity and enabling the implementation of conservation strategies. A limiting challenge is the extraction of historical DNA (hDNA) of adequate quality for sequencing platforms. We tested four hDNA extraction protocols on five body parts of pinned false heath fritillary butterflies, <em>Melitaea diamina</em>, aiming to minimise specimen damage, preserve their scientific value to the collections, and maximise DNA quality and yield for whole-genome re-sequencing. We developed a very effective approach that successfully recovers hDNA appropriate for short-read sequencing from a single leg of pinned specimens using silica-based DNA extraction columns and an extraction buffer that includes SDS, Tris, Proteinase K, EDTA, NaCl, PTB, and DTT. We observed substantial variation in the ratio of nuclear to mitochondrial DNA in extractions from different tissues, indicating that optimal tissue choice depends on project aims and anticipated downstream analyses. We found that sufficient DNA for whole genome re-sequencing can reliably be extracted from a single leg, opening the possibility to monitor changes in genetic diversity maintaining the scientific value of specimens while supporting current and future conservation strategies.</p>
Extracted experimental data for research paper titled "Micro-thermomechanical Modeling of Rocks with Temperature-dependent Friction and Damage Laws"
<p>This repository contains the experimental data on stress-strain curves extracted from the following original publications for constitutive model validation in our manuscript.</p> <p>[1] Jinping marble: Zhong, Y. Y. (2017). Research on mechanical properties of marble and the effects on rock burst under thermal-mechanical coupling (in Chinese) (Master’s thesis, Chengdu University of Technology). doi: 10.26986/d.cnki.gcdlc.2017.000109.</p> <p>[2] Beibei sandstone: Long, L. J. (2021). Study on mechanical and seepage properties of sandstone under the coupling of temperature-seepage-stress (in Chinese) (Doctoral dissertation, Chongqing University). doi: 10.27670/d.cnki.gcqdu.2021.001009.</p> <p>[3] Gongjue granite: Zhou, H. Y., Liu, Z. B., Shen, W. Q., Feng, T., & Zhang, G. Z. (2022). Mechanical property and thermal degradation mechanism of granite in thermal-mechanical coupled triaxial compression. International Journal of Rock Mechanics and Mining Sciences, 160, 105270. doi: 10.1016/j.ijrmms.2022.105270.</p>
Forest Fire Dataset for Peninsular Malaysia (2001-2023) Extracted from Multiple-Source Remote Sensing Data using Google Earth Engine
<ul> <li>Dataset: Forest Fire data</li> <li>Time Period: 2001 to 2023</li> <li>Location: Peninsular Malaysia</li> <li>Historical Fire Source: MCD64A1 and FIRMS Hotspots</li> <li>Fire Factors Extracted: Global Remote Sensing Data from GEE</li> </ul> <p>The framework extraction process can be reffered from the following publication:</p> <ul> <li>Framework to Create Inventory Dataset for Disaster Behavior Analysis Using Google Earth Engine: A Case Study in Peninsular Malaysia for Historical Forest Fire Behavior Analysis</li> <li>Journal: <em>Forests</em> <strong>2024</strong>, <em>15</em>(6), 923;</li> <li><a href="https://doi.org/10.3390/f15060923">https://doi.org/10.3390/f15060923</a></li> <li>The variables name such as AET (actual evapotranspiration) can be found from the article.</li> </ul> <p>Access the framework code from: </p> <ul> <li><a href="https://github.com/chewyeejian/GEE_FrameworkForestFireDataset">https://github.com/chewyeejian/GEE_FrameworkForestFireDataset</a></li> </ul> <p>The time sequence in the variable indicate whether it's a monthly data / yearly accumulated data / seasonal data, example:</p> <ul> <li>200101_aet (Year 2001, Month 01, value for aet (actual evapotranspiration)</li> <li>2001_aet_DJF (Average of December, January, February)</li> <li>2001_aet_MAM (Seasonal Average of March, April, May)</li> <li>2001_aet_JJA (Seasonal Average of June, July, August)</li> <li>2001_aet_SON (Seasonal Average of September, October, November)</li> <li>2001_aet_annual (Annual average of 2001)</li> </ul>
Data set for ODI cricket matches from 1987 to 2023 (extracted from ESPN Cricinfo) and code (R) used for a statistical study
<p>Here I present the data and code that has been used to study the statistical evolution of ODI cricket. The preprint for this research is available at: </p> <div> <div> <div> <table> <tbody> <tr> <td><a href="https://doi.org/10.48550/arXiv.2406.11652">https://doi.org/10.48550/arXiv.2406.11652</a> <div><span>Focus to learn more</span></div> </td> </tr> </tbody> </table> </div> </div> </div> <div> </div>
[Data augmentation in a TTL] - Fictive dataset (27.5M) with up to 5k reactions per template // (13'953 template extracted from USPTO-FULL IBM version)
<p>Full generated fictive dataset, containing 27.5M reactions with up to 5000 reactions per radius 1 reaction template (13'953 reaction templates from USPTO-full, IBM version).</p> <p>Title of the manuscript:</p> <p>"Data augmentation in a Triple Transformer Loop retrosynthesis model"</p> <p>Abstract: </p> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div dir="auto"> <div>Reactions in the US Patent Office (USPTO) are biased towards a few over-represented reaction types, which potentially limits its usefulness for computer-assisted synthesis planning (CASP). To obtain an equilibrated dataset, we applied retrosynthesis templates to USPTO molecules as products (P) to generate starting materials (SM). We then used transformer T2 from our recently reported triple transformer loop (TTL) retrosynthesis model to predict reagents (R) for the SM®P reaction. Finally, we validated the prediction by requesting a high confidence prediction (>95%) for the prediction of P from SM+R by TTL transformer T3. We generated up to 5,000 reactions per template, resulting in 27.5 million validated fictive reactions covering the chemical space of the original UPSTO dataset. To exemplify the use of this dataset, we show that a single-step retrosynthesis transformer model trained with a template equilibrated subset of 1,097,374 fictive reactions outperforms the corresponding model trained on USPTO reactions only.</div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> <div></div>
Manual Dependency Annotation of Three German Text Extracts from the Project hermA (Gold Standard Data)
<p>This dataset was created in the digital humanities project hermA (www.herma.uni-hamburg.de) and comprises annotated extracts of the following three texts:</p> <ul> <li>Modern literature (Lit2009): novel <em>Corpus Delicti: Ein Prozess </em>by German author Juli Zeh, published in Frankfurt/Main in 2009.</li> <li>Non-contemporary literature (Lit1850):<em> Eine Frauenfahrt um die Welt </em>('A woman’s journey around the world') by Austrian author Ida Pfeiffer (1850). Full text available at Deutsches Textarchiv: http://www.deutschestextarchiv.de/pfeiffer_frauenfahrt01_1850/6.</li> <li>Modern academic writing (Aca2009): <em>Stand, Möglichkeiten und Grenzen der Telemedizin in Deutschland</em> ('Telemedicine in Germany: status, chances and limits') by Rüdiger Klar and Ernst Pelikan, published in Bundesgesundheitsblatt ('Federal Health Gazette') in 2009. DOI 10.1007/s00103-009-0787-7.</li> </ul> <p>The texts are annotated for part-of-speech and dependency syntax and are made available in CoNLL file format. We describe the annotation process and report inter-annotator agreements in:</p> <p>Adelmann, Benedikt, Melanie Andresen, Wolfgang Menzel & Heike Zinsmeister. 2018. Evaluation of Out-Of Domain Dependency Parsing for its Application in a Digital Humanities Project. <em>Proceedings of the 14th Conference on Natural Language Processing (KONVENS 2018)</em>. Vienna, Austria.</p>
Expression data of the flowering time genes in chickpea, extracted from Ridge et al. (2017). Plant Physiology 175, 802-815.
<p>This data is supplementary to the following paper: Gursky, V.V., Kozlov, K.N., Nuzhdin, S.V., and Samsonova, M.G. (2018) Dynamical Modeling of the Core Gene Network Controlling Flowering Suggests Cumulative Activation from the <em>FLOWERING LOCUS T </em>Gene Homologs in Chickpea. <em>Frontiers in Genetics</em>. 9:547. doi: 10.3389/fgene.2018.00547</p> <p>The data was obtained by digitizing Figure 5 of the following paper: Ridge, S., Deokar, A., Lee, R., Daba, K., Macknight, R. C., Weller, J. L., and Tar'an, B. (2017). The chickpea Early flowering 1 (Efl1) locus is an ortholog of arabidopsis ELF3. <em>Plant Physiology </em>175, 802-815. doi:10.1104/pp.17.00082</p> <p>The archive contains files (in csv format) with the expression data of each of the following ten genes: <em>FTa1</em>, <em>FTa2</em>, <em>FTa3</em>, <em>FTb</em>, <em>FTc</em>, <em>AP1</em>, <em>FD</em>, <em>TFL1a</em>, <em>TFL1c</em>, and <em>LFY</em>, for the cultivars CDC Frontier and ICCV 96029 and for two growth conditions (long day, LD, and short day, SD). Each file is named according to the following scheme: <Gene name>_<Cultivar name>_<Growth conditions>.csv. Each file contains values in the following three columns (separated by commas): time (in days after sowing), relative transcription level (%ACTIN), and standard error. In the case of the genes <em>AP1</em>, <em>FD</em>, <em>TFL1a</em>, <em>TFL1c</em>, and <em>LFY</em>, the standard error was assumed equal to the size of the points in the figure when the actual error range was smaller than that size (and, thus, not visible in the figure). In the case of the genes <em>FTa1</em>, <em>FTa2</em>, <em>FTa3</em>, <em>FTb</em>, and <em>FTc</em>, the standard error was recorded as 0 for such points (the error for these genes was not used in the study).</p> <p>The data was extracted with the help of the web-based tool <em>WebPlotDigitizer</em> (https://automeris.io/WebPlotDigitizer).</p>
Scoping review of the term 'genetic identity' - data extraction table
<p>This dataset comes from a systmatic scoping reivew of the term ‘genetic identity’ within different academic discourses. This review systematically identifies all uses of the term ‘genetic identity’ within the academic literature. Content analysis is used to code and categorize its use, and the emerging discourses are related to disciplinary perspective, year of publication, and geographical setting.</p>
Data for scenario extraction ESR 12
<p>Data for scenario extraction by Mobileye camera.</p>
Data and Scripts for the Article "Structural Descriptors and Information Extraction from X-ray Emission Spectra: Aqueous Sulfuric Acid"
<p>Data and scripts for the article titled "Structural Descriptors and Information Extraction from X-ray Emission Spectra: Aqueous Sulfuric Acid".</p> <p>For further details on the contents, see the "readme.md"-file.</p> <p>Article available at <a href="https://doi.org/10.1039/D4CP02454K">10.1039/D4CP02454K</a>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.