Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

355

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

355 results for “data extraction”

Learn how ShareScore rates datasets ↗
dryad28/100

Data from: Toward synthesizing our knowledge of morphology: using ontologies and machine reasoning to extract presence/absence evolutionary phenotypes across studies

The reality of larger and larger molecular databases and the need to integrate data scalably have presented a major challenge for the use of phenotypic data. Morphology is currently primarily described in discrete publications, entrenched in noncomputer readable text, and requires enormous investments of time and resources to integrate across large numbers of taxa and studies. Here we present a new methodology, using ontology-based reasoning systems working with the Phenoscape Knowledgebase (KB; kb.phenoscape.org), to automatically integrate large amounts of evolutionary character state descriptions into a synthetic character matrix of neomorphic (presence/absence) data. Using the KB, which includes more than 55 studies of sarcopterygian taxa, we generated a synthetic supermatrix of 639 variable characters scored for 1051 taxa, resulting in over 145,000 populated cells. Of these characters, over 76% were made variable through the addition of inferred presence/absence states derived by machine reasoning over the formal semantics of the source ontologies. Inferred data reduced the missing data in the variable character-subset from 98.5% to 78.2%. Machine reasoning also enables the isolation of conflicts in the data, that is, cells where both presence and absence are indicated; reports regarding conflicting data provenance can be generated automatically. Further, reasoning enables quantification and new visualizations of the data, here for example, allowing identification of character space that has been undersampled across the fin-to-limb transition. The approach and methods demonstrated here to compute synthetic presence/absence supermatrices are applicable to any taxonomic and phenotypic slice across the tree of life, providing the data are semantically annotated. Because such data can also be linked to model organism genetics through computational scoring of phenotypic similarity, they open a rich set of future research questions into phenotype-to-genome relationships.

opencc-zeroDec 2014View details →
dryad28/100

Data from: DNA extraction method affects the detection of a fungal pathogen in formalin-fixed specimens using qPCR

Museum collections provide indispensable repositories for obtaining information about the historical presence of disease in wildlife populations. The pathogenic amphibian chytrid fungus Batrachochytrium dendrobatidis (Bd) has played a significant role in global amphibian declines, and examining preserved specimens for Bd can improve our understanding of its emergence and spread. Quantitative PCR (qPCR) enables Bd detection with minimal disturbance to amphibian skin and is significantly more sensitive to detecting Bd than histology; therefore, developing effective qPCR methodologies for detecting Bd DNA in formalin-fixed specimens can provide an efficient and effective approach to examining historical Bd emergence and prevalence. Techniques for detecting Bd in museum specimens have not been evaluated for their effectiveness in control specimens that mimic the conditions of animals most likely to be encountered in museums, including those with low pathogen loads. We used American bullfrogs (Lithobates catesbeianus) of known infection status to evaluate the success of qPCR to detect Bd in formalin-fixed specimens after three years of ethanol storage. Our objectives were to compare the most commonly used DNA extraction method for Bd (PrepMan, PM) to Macherey-Nagel DNA FFPE (MN), test optimizations for Bd detection with PM, and provide recommendations for maximizing Bd detection. We found that successful detection is relatively high (80–90%) when Bd loads before formalin fixation are high, regardless of the extraction method used; however, at lower infection levels, detection probabilities were significantly reduced. The MN DNA extraction method increased Bd detection by as much as 50% at moderate infection levels. Our results indicate that, for animals characterized by lower pathogen loads (i.e., those most commonly encountered in museum collections), current methods may underestimate the proportion of Bd-infected amphibians. Those extracting DNA from archived museum specimens should ensure that the techniques they are using are known to provide high-quality throughput DNA for later analysis.

opencc-zeroDec 2014View details →
dryad28/100

Data from: More than skin and bones: comparing extraction methods and alternative sources of DNA from avian museum specimens

Next-generation sequencing has greatly expanded the utility and value of museum collections by revealing specimens as genomic resources. As the field of museum genomics grows, so does the need for extraction methods that maximize DNA yields. For avian museum specimens, the established method of extracting DNA from toe pads works well for most specimens. However, for some specimens, especially those of birds that are very small or very large, toe pads can be a poor source of DNA. In this study, we apply two DNA extraction methods (phenol-chloroform and silica column) to three different sources of DNA (toe pad, skin punch, and bone) from ten historical avian museum specimens. We show that a modified phenol-chloroform protocol yielded significantly more DNA than a silica column protocol (e.g., Qiagen DNeasy Blood & Tissue Kit) across all tissue types. However, extractions using the silica column protocol contained longer fragments on average than those using the phenol-chloroform protocol, likely a result of loss of small fragments through the silica column. While toe pads yielded more DNA than skin punches and bone fragments, skin punches proved to be a reliable alternative source of DNA and might be especially appealing when toe pad extractions are impractical. Overall, we found that historical bird museum specimens contain substantial amounts of DNA for genomic studies under most extraction scenarios, but that a phenol-chloroform protocol consistently provides the high quantities of DNA required for most current genomic protocols.

opencc-zeroJul 2019View details →
dryad28/100

Data from: Efficient and accurate extraction of in vivo calcium signals from microendoscopic video data

In vivo calcium imaging through microendoscopic lenses enables imaging of previously inaccessible neuronal populations deep within the brains of freely moving animals. However, it is computationally challenging to extract single-neuronal activity from microendoscopic data, because of the very large background fluctuations and high spatial overlaps intrinsic to this recording modality. Here, we describe a new constrained matrix factorization approach to accurately separate the background and then demix and denoise the neuronal signals of interest. We compared the proposed method against previous independent components analysis and constrained nonnegative matrix factorization approaches. On both simulated and experimental data recorded from mice, our method substantially improved the quality of extracted cellular signals and detected more well-isolated neural signals, especially in noisy data regimes. These advances can in turn significantly enhance the statistical power of downstream analyses, and ultimately improve scientific conclusions derived from microendoscopic data.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Photocatalytic activity of graphene oxide–TiO2 thin films sensitized by natural dyes extracted from Bactris guineensis

This study synthesized and characterized composites of graphene oxide and TiO2 (GO-TiO2). GO-TiO2 thin films were deposited using the Doctor Blade technique. Subsequently, the thin films were sensitized with a natural dye extracted from a Colombian source (Bactris guineensis). Thermogravimetric analysis, X-ray diffraction, Raman spectroscopy, scanning electron microscopy (SEM), X-ray photoelectron spectroscopy (XPS), and diffuse reflectance measurements were used for physicochemical characterization. All the samples were polycrystalline in nature, and the diffraction signals corresponded to the TiO2 anatase crystalline phase. Raman spectroscopy and FT-IR verified the synthesis of composite thin films, and the SEM analysis confirmed the TiO2 films morphological modification after the process of GO incorporation and sensitization. XPS results suggested a possibility of appearance of Titanium (III) through the formation of oxygen vacancies (Ov). Furthermore, the optical results indicated that the presence of the natural sensitizer and GO improved the optical properties of TiO2 in the visible range. Finally, the photocatalytic degradation of Methylene Blue (MB) was studied under visible irradiation in aqueous solution, and pseudo-first order model was used to obtain kinetic information about photocatalytic degradation. These results indicated that the presence of GO has an important synergistic effect in conjunction with the natural sensitizer, reaching a photocatalytic yield of 33%.

opencc-zeroDec 2018View details →
dryad28/100

Data from: Preservation-induced morphological change in salamanders and failed DNA extraction from a decades-old museum specimen: implications for Plethodon ainsworthi

<p>Natural history collections are important data repositories, but different chemical treatments of specimens can influence morphological measurements and DNA extraction, complicating taxonomic and conservation decisions dependent upon these data. One such example is the Bay Springs Salamander (<i>Plethodon ainsworthi</i>), the only United States amphibian categorized as Extinct by the IUCN.<i> </i>Recent research has proposed that <i>P. ainsworthi </i>is an invalid taxon, arguing that the 55-year-old type specimens' morphological distinctiveness from syntopic <i>P. mississippi</i> is a preservation artifact. To address this controversy, we tested for morphological changes across five experimental treatments in proxy <i>P. shermani</i> specimens, and we re-examined the datasets used to support the invalidity of <i>P. ainsworthi</i>. We also tested recently developed DNA extraction techniques on the putatively formalin-fixed <i>P. ainsworthi</i> holotype. We used Bayesian models to demonstrate that preservation method can differentially bias morphological measurements, with most methods causing lower estimates of mass and modestly higher estimates of snout-vent-length:head width ratio. These results are broadly consistent with previous studies of other vertebrates, but inconsistent with the hypothesis that <i>P. ainsworthi </i>type specimens are actually poorly preserved <i>P. mississippi</i>. Attempts to extract DNA from the <i>P. ainsworthi</i> holotype unfortunately proved unsuccessful, preventing conclusive resolution of its status and emphasizing the limitations of promising new methods. Nonetheless, we tentatively recommend continued recognition of <i>P. ainsworthi</i> as a valid but possibly extinct taxon. More generally, we invite all authors who study preserved specimens to recognize and report how certain chemical treatments might impact their results.</p>

opencc-zeroNov 2019View details →
dryad28/100

Data from: Membrane-assisted extraction of monoterpenes: from in-silico solvent screening towards biotechnological process application

This work focuses on the process development of membrane-assisted solvent extraction of hydrophobic compounds such as monoterpenes. Beginning with the choice of suitable solvents, quantum chemical calculations with the simulation tool COSMO-RS were carried out to predict the partition coefficient (logP) of (S)-(+)-carvone and terpinen-4-ol in various solvent-water systems and validated afterwards with experimental data. COSMO-RS results show good prediction accuracy for nonpolar solvents like n-hexane, ethyl acetate and n-heptane even in the presence of salts and glycerol in aqueous medium. Based on the high logP value, n-heptane was chosen for the extraction of (S)-(+)-carvone in a lab-scale hollow-fiber membrane contactor. Two operation modes are investigated where experimental and theoretical mass transfer values, based on their related partition coefficients were compared. In addition, the process is evaluated in terms of extraction efficiency and overall product recovery, and its biotechnological application potential discussed. Our work demonstrates that the combination of in-silico prediction by COSMO-RS with membrane-assisted extraction is a promising approach for the recovery of hydrophobic compounds from aqueous solutions.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Ginkgo biloba extract for prevention of acute mountain sickness: a systematic review and meta-analysis of randomized controlled trials

Study objective: Trials of ginkgo biloba extract (GBE) for the prevention of acute mountain sickness (AMS) have been published since 1996. Because of their conflicting results, the efficacy of GBE remains unclear. We performed a systematic review and meta-analysis to assess whether GBE prevents acute mountain sickness. Methods: The Cochrane Library, EMBASE, Google Scholar, and PubMed databases were searched for articles published up to May 20, 2017. Only randomized controlled trials were included. AMS defined as acute mountain sickness–cerebral(AMS-C) score≧0.7 or Lake Louise Score (LLS)≧3 with headache. The main outcome measures were the relative risks of AMS in participants receiving GBE for prophylaxis. Meta-analyses were conducted using random-effects models. Sensitivity analyses, subgroup analyses and tests for publication bias were conducted. Results: Six published articles with a total of 451 participants met all eligibility criteria. In the primary meta-analysis of all 7 study groups, GBE showed trend of AMS prophylaxis, but it is not statistically significant (RR =0.68; 95% CI: 0.45 to 1.04; p-value=0.08) (Figure 2). The I2 statistic was 58.7% (p-value=0.02), indicating substantial heterogeneity. The results of subgroup analyses of studies with low risk of bias, low starting altitude (&lt;2500 m), number of treatment days before ascending and dosage of GBE were similar. Conclusions: The currently available data suggest that although GBE may tend toward AMS prophylaxis, there are not enough data to show the statistically significant effect of GBE for preventing AMS. Further large randomized control studies are warranted.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Robust extraction of quantitative structural information from high-variance histological images of livers from necropsied Soay sheep

Quantitative information is essential to the empirical analysis of biological systems. In many such systems, spatial relations between anatomical structures is of interest, making imaging a valuable data acquisition tool. However, image data can be difficult to analyse quantitatively. Many image processing algorithms are highly sensitive to variations in the image, limiting their current application to fields where sample and image quality may be very high. Here, we develop robust image processing algorithms for extracting structural information from a dataset of high-variance histological images of inflamed liver tissue obtained during necropsies of wild Soay sheep. We demonstrate that features of the data can be measured in a fully automated manner, providing quantitative information which can be readily used in statistical analysis. We show that these methods provide measures that correlate well with a manual, expert operator-led analysis of the same images, that they provide advantages in terms of sampling a wider range of information and that information can be extracted far more quickly than in manual analysis.

opencc-zeroDec 2016View details →
dryad28/100

Data from: Parser extraction of triples in unstructured text

The web contains vast repositories of unstructured text. We investigate the opportunity for building a knowledge graph from these text sources. We generate a set of triples which can be used in knowledge gathering and integration. We define the architecture of a language compiler for processing subject-predicate-object triples using the OpenNLP parser. We implement a depth-first search traversal on the POS tagged syntactic tree appending predicate and object information. A parser enables higher precision and higher recall extractions of syntactic relationships across conjunction boundaries. We are able to extract 2-2.5 times the correct extractions of ReVerb. The extractions are used in a variety of semantic web applications and question answering. We verify extraction of 50,000 triples on the ClueWeb dataset.

opencc-zeroDec 2018View details →
zenodo28/100

Data extracted from acoustic device.

<p>Data extracted from acoustic device.</p>

opencc-by-4.0Oct 2016View details →
zenodo28/100

Musical characteristics and experience: data extraction for our systematic review

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo28/100

Protocol data extraction - Literature Review creating the Integrated List of Agile Practices

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →
dryad28/100

Data from: Analysis of phytochemicals, antioxidants, and antimicrobial properties in non-polar extracts of Magnolia virginiana L. flowers from Saudi Arabia

<div> <p><em><span>Magnolia virginiana</span></em><span> (<em>M. virginiana</em>) L., a native North American plant, is globally cultivated for shade and ornamental purposes, including in Saudi Arabia. This study analyzed the chemical diversity and biological activity of non-polar extracts (n-hexane and diethyl ether) from <em>M. virginiana</em> flowers. The major components identified by </span><span>gas chromatography-mass spectroscopy (GC-MS) analysis were aromatic and aliphatic esters, triterpenes, steroids, and phenolic acids. The total phenolic content (TPC) of n-hexane and diethyl ether extracts was determined to be 29.66 and 29.44 mGAE/g, respectively. The extracts showed strong antioxidant activity, with the diethyl ether extract having more reducing power than the <em>n</em>-hexane extract. The diethyl ether extract also showed greater Trolox equivalent values in total antioxidant capacity (TAC) and ferric reducing antioxidant power (FRAP) assays, but the n-hexane extract exhibited higher metal chelating activity (MCA) and free radical scavenging activity (DPPH-SA) levels. The diethyl ether extract displayed stronger antimicrobial potential than the <em>n</em>-hexane extract, particularly against <em>Staphylococcus saprophyticus</em>, with a zone of inhibition diameter (ZID) of 20.0 ± 0.3 mm. The minimum inhibitory concentration (MIC), minimum biocidal concentration (MBC), minimum biofilm inhibitory concentration (MBIC), and minimum biofilm eradication concentration (MBEC) values were 0.78, 1.56, 1.56, and 3.125 mg/mL, respectively. </span></p> </div>

opencc-zeroNov 2023View details →
zenodo28/100

Data Extraction Sheet_SLR

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo28/100

Figure 3 from: Mesa-Varona O, Plaza-Rodríguez C, Valentin L, Filter M (2024) WarenstromInfo: a tool for the easy extraction and visualisation of trade flow data. Research Ideas and Outcomes 10: e112227. https://doi.org/10.3897/rio.10.e112227

Figure 3 WI User Interfaces: UI3. In the UI3, countries involved in the query can be selected (I and J). Input data from UI3 are included in the query by clicking "Next" (K).

opencc-by-4.0Feb 2024View details →
zenodo28/100

Figure 5 from: Mesa-Varona O, Plaza-Rodríguez C, Valentin L, Filter M (2024) WarenstromInfo: a tool for the easy extraction and visualisation of trade flow data. Research Ideas and Outcomes 10: e112227. https://doi.org/10.3897/rio.10.e112227

Figure 5 A screenshot of the BACI SQLite database builder workflow. The workflow presented here is organised into two main sections, with interconnected KNIME nodes displayed in each section. The yellow-boxed nodes are responsible for creating and loading the database where the data are stored. The green-boxed section shows: Locate BACI csv files (BACI files must have been previously downloaded and stored in a folder that is pointed in the "List Files/Folders" KNIME node). Split the files according to the HS data provided and Load csv files, filtering and storage in SQLite database (carried out in each consecutive metanode). In this last section, the user can decide not to filter the data or customise the filter criteria of the BACI database, including more types of data apart from those agrifood data that are included in the default configuration of the workflow. Variable flow connections (red lines) can be removed in the second green-boxed section, if an HPC (High Performance Computing) service is available. This will allow the workflow to run much faster as the nodes will not be executed one after the other, but all at the same time.

opencc-by-4.0Feb 2024View details →
zenodo28/100

Figure 4 from: Mesa-Varona O, Plaza-Rodríguez C, Valentin L, Filter M (2024) WarenstromInfo: a tool for the easy extraction and visualisation of trade flow data. Research Ideas and Outcomes 10: e112227. https://doi.org/10.3897/rio.10.e112227

Figure 4 WI User Interfaces: UI4. The UI4 provides an overview of the final WI outputs, displaying a list with the initial input variables (L), three maps displaying trade flows (a world map, a world map focused on Europe and an European map) (M), download options (N), slide filter bars for value and weight preview (O) and trade flow data in the preview table (P).

opencc-by-4.0Feb 2024View details →
zenodo28/100

Figure 1 from: Mesa-Varona O, Plaza-Rodríguez C, Valentin L, Filter M (2024) WarenstromInfo: a tool for the easy extraction and visualisation of trade flow data. Research Ideas and Outcomes 10: e112227. https://doi.org/10.3897/rio.10.e112227

Figure 1 WI User Interfaces: UI1. In the UI1, users can select the desired database (EUROSTAT/BACI) (A). The database selection is concluded by clicking "Next" (B), which will lead the user to the following UI.

opencc-by-4.0Feb 2024View details →
zenodo28/100

Figure 2 from: Mesa-Varona O, Plaza-Rodríguez C, Valentin L, Filter M (2024) WarenstromInfo: a tool for the easy extraction and visualisation of trade flow data. Research Ideas and Outcomes 10: e112227. https://doi.org/10.3897/rio.10.e112227

Figure 2 User Interfaces: UI2. In the UI2, the following input options can be selected: standard selection (C), time range of the search (D), the email option (E), the pre-filter option (F) and the table with the embedded live filter (G). Input data from UI2 are processed by clicking "Next" (H).

opencc-by-4.0Feb 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record