Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

355

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

355 results for “data extraction”

Learn how ShareScore rates datasets ↗
zenodo32/100

Guidelines for Software Product Line Experiments: SMS extracted data and guidelines evaluation

<p>Guidelines for Softwrae Product Line Experiments: SMS extracted data and guidelines evaluation</p>

opencc-by-4.0Jun 2020View details →
dryad32/100

Data from: Concordance in wetland physicochemical conditions, vegetation, and surrounding land cover is robust to data extraction approach

Concordance among wetland physicochemical conditions, vegetation, and surrounding land cover may result from the influence of land cover on the sources of plant propagules, on physicochemical conditions, and their subsequent determination of growing conditions. Alternatively, concordance may result if differences in climate, soils, and species pools are spatially confounded with differences in human population density and land conversion. Further, we expect that land cover within catchment boundaries will be more predictive than land cover in symmetrical buffers if runoff is a major pathway. We measured concordance between land cover, wetland vegetation and physicochemical conditions in 48 prairie pothole wetlands, controlling for inter-wetland distance. We contrasted land-cover data collected over a four-year period by multiple extraction approaches including topographically-delineated catchments and nested 30 m to 5,000 m radius buffers. After factoring out inter-wetland distance, physiochemical conditions were significantly concordant with land cover. Vegetation was not significantly concordant with land cover, though it was strongly and significantly concordant with physicochemical conditions. More, concordance was as strong when land cover was extracted from buffers &lt;500 m in radius as from catchments, indicating the mechanism responsible is not topographically constrained. We conclude that local landscape structure does not directly influence wetland vegetation composition, but rather that vegetation depends on physicochemical conditions in the wetland (which are affected by surrounding land cover) and on regional factors such as the vegetation species pool and geographic gradients in climate, soil type, and land use.

opencc-zeroDec 2018View details →
dryad32/100

Data from: Comparative analysis of DNA extraction methods to study the body surface microbiota of insects: a case study with ant cuticular bacteria

High-throughput sequencing of the 16S rRNA gene has considerably helped revealing the essential role of bacteria living on insect cuticles in the ecophysiology and behavior of their hosts. However, our understanding of host-cuticular microbiota feedbacks remains hampered by the difficulties to working with low bacterial DNA quantities as in individual insect cuticle samples, which are more prone to molecular biases and contaminations. Herein, we conducted a methodological benchmark on the cuticular bacterial loads retrieved from two Neotropical ant species of different body size and ecology: Atta cephalotes (~15 mm) and Pseudomyrmex penetrator (~5 mm). We evaluated the richness and composition of the cuticular microbiota, as well as the amount of biases and contamination produced by four DNA extraction protocols. We also addressed how bacterial communities' characteristics would be affected by the number of individuals or individual body size used for DNA extraction. Most extraction methods yielded similar results in term of bacterial diversity and composition for A. cephalotes (~15 mm). In contrast, greater amounts of artifactual sequences and contaminations, as well as noticeable differences in bacterial communities' characteristics were observed between the extraction methods for P. penetrator (~5 mm). We also found that large (~15 mm) and small (~5 mm) A. cephalotes individuals harbor different bacterial communities. Our benchmark hence suggests that cuticular microbiota of single insect individuals can be reliably retrieved provided that blank controls, appropriate data cleaning, and standardization of individual body size are considered in the experiment.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Can differential nutrient extraction explain property variations in a predatory trap?

Predators exhibit flexible foraging to facilitate taking prey that offer important nutrients. Because trap-building predators have limited control over the prey they encounter, differential nutrient extraction and trap architectural flexibility may be used as a means of prey selection. Here, we tested whether differential nutrient extraction induces flexibility in architecture and stickiness of a spider's web by feeding Nephila pilipes live crickets (CC), live flies (FF), dead crickets with the web stimulated by flies (CD) or dead flies with the web stimulated by crickets (FD). Spiders in the CD group consumed less protein per mass of lipid or carbohydrate, and spiders in the FF group consumed less carbohydrates per mass of protein. Spiders from the CD group built stickier webs that used less silk, whereas spiders in the FF group built webs with more radii, greater catching areas and more silk, compared with other treatments. Our results suggest that differential nutrient extraction is a likely explanation for prey-induced spider web architecture and stickiness variations.

opencc-zeroDec 2014View details →
dryad32/100

Data from: High-throughput sequencing of nematode communities from total soil DNA extractions

Background: Nematodes are extremely diverse and numbers of species are predicted to be more than a million. Studies on nematode diversity are difficult and laborious using standard methods such as identification based on morphology and therefore high-throughput sequencing is an attractive alternative. Generally, primers that have been used for generating amplicons for sequencing are not nematode specific and also amplify other groups such as fungi and plantae. Thus a nematode enrichment step must be included that may introduce biases. Results: An amplification strategy, including a new primer, which selectively amplifies nematodes and other metazoans was developed. When this strategy was tested on DNA templates from a set of 22 agricultural soils, we obtained 64.4 % sequences of nematode origin in total, whereas the remaining sequences were almost entirely metazoan. The nematode sequences were derived from a broad taxonomic range and most sequences were from nematode taxa that have previously been found to be abundant in soil such as Tylenchida, Rhabditida, Dorylaimida, Triplonchida and Araeolaimida. Conclusions: This amplification and sequencing strategy for assessing nematode diversity was demonstrated to be able to collect a broad taxonomy of nematodes without prior enrichment and thus the method will be highly valuable in ecological studies of nematodes. Keywords: nematode, community, next-generation sequencing, SSU, diversity, 18S, rDNA

opencc-zeroDec 2014View details →
zenodo32/100

Data from: Decentralizing cell-free RNA sensing with the use of low-cost cell extracts

<p>Data and DNA sequences for the publication &quot;Decentralizing cell-free RNA sensing with the use of low-cost cell extracts&quot; (https://www.biorxiv.org/content/10.1101/2021.05.29.446205v1)</p>

opencc-by-4.0Jun 2021View details →
dryad32/100

Data from: Extracting DNA from 'jaws': high yield and quality from archived tiger shark (Galeocerdo cuvier) skeletal material

Archived specimens are highly valuable sources of DNA for retrospective genetic/genomic analysis. However, often limited effort has been made to evaluate and optimize extraction methods, which may be crucial for downstream applications. Here, we assessed and optimized the usefulness of abundant archived skeletal material from sharks as a source of DNA for temporal genomic studies. Six different methods for DNA extraction, encompassing two different commercial kits and three different protocols, were applied to material, so-called bio-swarf, from contemporary and archived jaws and vertebrae of tiger sharks (Galeocerdo cuvier). Protocols were compared for DNA yield and quality using a qPCR approach. For jaw swarf, all methods provided relatively high DNA yield and quality, while large differences in yield between protocols were observed for vertebrae. Similar results were obtained from samples of white shark (Carcharodon carcharias). Application of the optimized methods to 38 museum and private angler trophy specimens dating back to 1912 yielded sufficient DNA for downstream genomic analysis for 68% of the samples. No clear relationships between age of samples, DNA quality and quantity were observed, likely reflecting different preparation and storage methods for the trophies. Trial sequencing of DNA capture genomic libraries using 20 000 baits revealed that a significant proportion of captured sequences were derived from tiger sharks. This study demonstrates that archived shark jaws and vertebrae are potential high-yield sources of DNA for genomic-scale analysis. It also highlights that even for similar tissue types, a careful evaluation of extraction protocols can vastly improve DNA yield.

opencc-zeroDec 2015View details →
dryad32/100

Data from: A proteomic method to extract, concentrate, digest, and enrich peptides from fossils with colored (humic) substances for mass spectrometry analyses

Humic substances are break-down products of decaying organic matter that co-extract with proteins from fossils. These substances are difficult to separate from proteins in solution, and interfere with analyses of fossil proteomes. We introduce a method combining multiple recent advances in extraction protocols to both concentrate proteins from fossil specimens with high humic content, and remove humics, producing clean samples easily analyzed by mass spectrometry (MS). This method includes: 1) a non-demineralizing extraction buffer that eliminates protein loss during the demineralization step in routine methods; 2) filter-aided sample preparation (FASP) of peptides, which concentrates and digests extracts in one filter, allowing the separation of large humics after digestion; 3) centrifugal stage-tipping, which further clarifies and concentrates samples in a uniform process performed simultaneously on multiple samples. We apply this method to a moa fossil (~800¬–1000 yr) dark with humic content, generating colorless samples and enabling the detection of more proteins with greater sequence coverage than previous MS analyses on this same specimen. This workflow allows analyses of low-abundance proteins in fossils containing humics, and thus may widen the range of extinct organisms and regions of their proteomes we can explore with MS.

opencc-zeroJul 2019View details →
dryad32/100

Data from: Improving PCR detection of prey in molecular diet studies: importance of group-specific primer set selection and extraction protocol performances

While morphological identification of prey remains in feces of predators is the method most commonly used to study trophic interactions, many studies indicate that this method does not detect all consumed prey. Polymerase Chain Reaction based methods are increasingly used to detect prey DNA in the predator food bolus and have proved themselves efficient, with high accuracy. When studying complex diet samples, the extraction of total DNA is a critical step, as PCR inhibitors may be co-extracted. Another critical step consist in carefully select suitable group-specific primer sets that should only amplify prey DNA from the targeted taxon. In this study, the food boluses of five Rattus rattus and seven Rattus exulans were analyzed using both morphological and molecular methods. We tested a panel of 30 PCR specific primer sets targeting Bird, Invertebrate and Plant sequences and four were finally selected to be use as group-specific primer pairs in PCR protocols. The performances of four DNA extraction protocols (QIAamp DNA stool mini kit, DNeasy mericon food kit and two CTAB-based methods) were compared using four variables: DNA concentration, A260/A280 absorbance ratio, food compartment analyzed (stomach or fecal contents), total number of prey specific PCR amplification per sample. Our results clearly indicate that the A260/A280 absorbance ratio, which varies between extraction protocols, is positively correlated to the number of PCR amplifications of each prey taxon. We recommend using the DNeasy mericon food kit (Qiagen), which yielded results very similar to those achieved with the morphological approach.

opencc-zeroDec 2011View details →
dryad32/100

Data from: Reconciling timber extraction with biodiversity conservation in tropical forests using reduced-impact logging

1. Over 20% of the world's tropical forests have been selectively logged, and large expanses are allocated for future timber extraction. Reduced-impact logging (RIL) is being promoted as best practice forestry that increases sustainability and lowers CO2 emissions from logging, by reducing collateral damage associated with timber extraction. RIL is also expected to minimize the impacts of selective logging on biodiversity, although this is yet to be thoroughly tested. 2. We undertake the most comprehensive study to date to investigate the biodiversity impacts of RIL across multiple taxonomic groups. We quantified birds, bats and large mammal assemblage structures, using a before-after control-impact (BACI) design across 20 sample sites over a 5-year period. Faunal surveys utilized point counts, mist nets and line transects and yielded &gt;250 species. We examined assemblage responses to logging, as well as partitions of feeding guild and strata (understorey vs. canopy), and then tested for relationships with logging intensity to assess the primary determinants of community composition. 3. Community analysis revealed little effect of RIL on overall assemblages, as structure and composition were similar before and after logging, and between logging and control sites. Variation in bird assemblages was explained by natural rates of change over time, and not logging intensity. However, when partitioned by feeding guild and strata, the frugivorous and canopy bird ensembles changed as a result of RIL, although the latter was also associated with change over time. Bats exhibited variable changes post-logging that were not related to logging, whereas large mammals showed no change at all. 4. Indicator species analysis and correlations with logging intensities revealed that some species exhibited idiosyncratic responses to RIL, whilst abundance change of most others was associated with time. 5. Synthesis and applications. Our study demonstrates the relatively benign effect of reduced-impact logging (RIL) on birds, bats and large mammals in a neotropical forest context, and therefore, we propose that forest managers should improve timber extraction techniques more widely. If RIL is extensively adopted, forestry concessions could represent sizeable and important additions to the global conservation estate – over 4 million km2.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Increase in extraction of I-123 iomazenil in patients with chronic cerebral ischemia

Background: Cerebral extraction of diffusively distributed substances like oxygen has been suggested to change according to the cerebral blood flow (CBF) and status of the microvasculature. The relationships between the cerebral extraction of diffusively distributed lipophilic tracers and the severity of cerebral ischemia has not yet been clarified. In the present study, we attempted to elucidate the association between the extraction fraction of the lipophilic tracer I-123 iomazenil (IMZ) (IMZ-EF) and the oxygen extraction fraction (OEF) derived from O-15 PET in patients with chronic steno-occlusive disease of internal carotid artery (ICA) or middle cerebral artery (MCA). Methods: Seven patients with unilateral chronic severe stenosis or occlusion of the middle cerebral/internal cerebral artery were prospectively recruited for this study. All the patients underwent both O-15 PET and quantitative I-123 IMZ SPECT. Parametric images derived from the PET and SPECT scans were anatomically normalized and evaluated by automated image analysis based on the volume-of-interest template. Results: The asymmetry index (AI) of IMZ-EF was shown to significantly correlated with the AI of OEF (r = 0.562, P &lt; 0.001) in the internal carotid artery perfusion area. Strong and significant correlation between the AI of the influx rate constant K1 of IMZ and the AI of the cerebral metabolic rate of oxygen (r = 0.552, P = 0.001) was clarified. Conclusions: Our results suggested that the transportation efficiency of I-123 IMZ into the brain tissue was an indicator for evaluating severity of cerebral ischemia in patients with chronic steno-occlusive disease of ICA or MCA. Cerebral metabolic state can possibly be estimated by I-123 IMZ SPECT without cyclotron.

opencc-zeroDec 2017View details →
zenodo32/100

Weak evidence base for bee protective pesticide mitigation measures: Data Behind Systematic Review- Raw, Extracted and Tabulated

<p>An excel file with the raw exported data from Web of Science, the data extracted from it, and the summary statistics.&nbsp;</p> <p>Data associated with https://doi.org/10.1093/jee/toad118</p> <p>Edward A Straw, Dara A Stanley, Weak evidence base for bee protective pesticide mitigation measures, <em>Journal of Economic Entomology</em>, Volume 116, Issue 5, October 2023, Pages 1604&ndash;1612, <a href="https://doi.org/10.1093/jee/toad118">https://doi.org/10.1093/jee/toad118</a></p> <p>&nbsp;</p> <p>Straw and Stanley, 2023.&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2023View details →
zenodo32/100

Reduced Gun Violence Frame Corpus data set for the Text2Story 2024 article: "Evaluating the Ability of Computationally Extracted Narrative Maps to Encode Media Framing"

<p><strong>Title: </strong>Simplified Gun Violence Frame Corpus (GVFC) Subset</p> <p><strong>Description:</strong><br>This data set is a simplified subset of the Gun Violence Frame Corpus (GVFC) from Liu et al. (2019). The original GVFC consists of 1300 news articles in English from multiple U.S. based sources extracted during the year 2018, focusing on media frames commonly used when reporting the issue of Gun Violence. The original data set has 9 types of frames, including both issue-specific and generic frames. Due to high computational costs in our analysis methods, we decreased the data set size from 1300 articles to 131 articles using stratified sampling, maintaining the original distribution of the frame labels. We also manually searched for the original sources of each article based on its headline and added the missing temporal information and news source to the data set, as it was required by our algorithms.</p> <p>To further reduce the complexity of the framing model and account for the smaller data set size, we grouped the original nine frames into three higher-level frames:</p> <p>1. Frame 1: Political Issues - Combining the first, second, and third frames, which focus on political issues mostly related to gun control.<br>2. Frame 2: Public Services - Combining the fourth and fifth frames, which focus on mental healthcare issues, as well as school and public safety.<br>3. Frame 3: Cultural and Societal Issues - Combining the last four frames, which are oriented towards cultural or societal issues, including discussions around race and ethnicity, public opinion, and economic consequences.</p> <p>The resulting simplified data set contains 131 news articles, each labeled with one of the three higher-level frames, along with the necessary temporal information and news source for the narrative extraction process.</p> <p>If you use this data set, please make sure to cite the original GVFC paper and our workshop paper please.&nbsp;</p> <p><strong>References:</strong></p> <ol> <li>Liu et al. (2019) "Detecting Frames in News Headlines and Its Application to Analyzing News Framing Trends Surrounding US Gun Violence", 23rd Conference on Computational Natural Language Learning (CoNLL 2019).</li> <li>Concha Mac&iacute;as, Sebasti&aacute;n and Keith Norambuena, Brian (2024). "Evaluating the Ability of Computationally Extracted Narrative Maps to Encode Media Framing", Text2Story 2024 Workshop, ECIR 2024.</li> </ol>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Programs and data used for Extracting latent variables from forecast ensembles and advancements in similarity metric utilizing optimal transport

<p>This is compiled from the program and output data using in Nishizawa (2024).</p> <p>&nbsp;</p> <p>Nishizawa, 2024: Extracting latent variables from forecast ensembles and advancements in similarity metric utilizing optimal transport. submitted to JGR: Machine Learning and Computation.</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

Iodine deficiency data extraction sheet

<p><span>The contents of the Excel file reporting the studies assessing the prevalence of iodine deficiency and associated factors among school-age children in Ethiopia, 2023 are: <span>name of the authors, year of publication, study year, country, region, study design, study setting, quality of the paper, total population, iodine deficient population, prevalence of iodine deficiency, standard error of the prevalence and factors such as age, sex, goitrogenic food consumption, salt container used to store the iodized salt, salt adding time and maternal educational status and their derivatives.</span></span></p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Metadata of the extracted data assessing iodine deficiency and associated factors among school-age children in Ethiopia, 2023.

<p>This is Meta-data of the extracted data assessing the prevalence of iodine deficiency and associated factors among school-age children in Ethiopia, 2023</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Data extraction sheet: Knowledge, Attitudes, and Practices of Women and Men Towards Infertility: A Scoping Review

<p>This is the data extraction tool that includes all the studies reviewed and inclided in: Knowledge, Attitudes, and Practices of Women and<br>Men Towards Infertility: A Scoping Review</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

MOFSimplify: Machine Learning Models with Extracted Stability Data of Three Thousand Metal-Organic Frameworks

<p>Solvent removal stability and thermal stability associated with structurally characterized metal organic frameworks.</p>

opencc-by-4.0Sep 2021View details →
zenodo32/100

Data extraction form - A Systematic Literature Review on Prioritizing Software Test Cases using Markov Chains

<p>A data extraction form was created to gather all relevant data from the identified studies and manage the selection process in this systematic literature review. Some of the main information presented in this form was followed by a protocol, identifier (id) for each returned study, bibliographic reference, and answers to research questions. This catalog helps us in the data extraction and synthesis procedures and may be used by potentially interested, for example, for updating or replication.&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

Marikina Valley Fault Creeping Segment (Philippines) - Groundwater Extraction Data

<p>This dataset contain information on groundwater extraction for different well locations along the creeping segment of the Marikina Valley Fault System (Philippines).&nbsp;</p>

opencc-by-4.0Dec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record