Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
107
datasets available to search
ShareScore release 0.9.0
Dataset results
107 results for “automated analysis”
Flow virometry for water-quality assessment: Protocol optimization for a model virus and automation of data analysis
Open the record for dataset details and reuse information.
Automated analysis of bird head motion in unconstrained settings: A foundational study on semicircular canal evolution in archosaurs
Open the record for dataset details and reuse information.
Fine-grained automated visual analysis of herbarium specimens for phenological data extraction: an annotated dataset of reproductive organs in Strepanthus herbarium specimens
<p>This dataset contains annotations of 31 herbarium specimens of <em>Streptanhus tortuosus Kellogg</em> for which we have we carefully and manually drew and annotated the contours of four reproductive organs: “bud”, “flower”, “immature fruit” and “mature fruit”.</p> <p>The dataset can be used to assess the ability of automated methods to count and detect precisely the shapes of these reproductive organs, with a view to conducting phenological studies.</p> <p>The annotations are formatted in accordance with the COCO data format, a usual format for object detection tasks in the field of Computer Vision. The annotations are divided into two files:</p> <ul> <li>train_21_full_masks.json contains the mask coordinates and labels of 21 herbarium sheets that can be used for training models</li> <li>test_10_full_masks.json contains the mask coordinates and labels of 10 other herbarium that can be used as a groundtruth file for evaluating the predictions, typically with the COCO evaluation scripts (<a href="https://github.com/cocodataset/cocoapi">https://github.com/cocodataset/cocoapi</a>)</li> </ul> <p>Please refer to the following publication for a first assessment of this dataset with a Mask-RCNN approach:</p> <p><em>H. Goëau, A. Mora-Fallas, J. Champ, N. Love, S. Mazer, E. Mata-Montero, A. Joly, P. Bonnet. </em>2020. New fine-grained method for automated visual analysis of herbarium specimens: a case study for phenological data extraction. <em>Applications in Plant Sciences </em></p> <p> </p> <p> </p> <p> </p> <p> </p>
Automated analysis of scanning electron microscopic images for assessment of hair surface damage
<p>Mechanical damage of hair can serve as an indicator of health status and its assessment relies on the measurement of morphological features via microscopic analysis, yet few studies have categorized the extent of damage sustained, and instead, have depended on qualitative profiling based on the presence or absence of specific features. We describe the development and application of a novel quantitative measure for scoring hair surface damage in scanning electron microscopic (SEM) images without predefined features, and automation of image analysis for characterization of morphological hair damage after exposure to an explosive blast. Application of an automated normalization procedure for SEM images revealed features indicative of contact with materials in an explosive device and characteristic of heat damage, though many were similar to features from physical and chemical weathering. Assessment of hair damage with tailing factor, a measure of asymmetry in pixel brightness histograms and proxy for surface roughness, yielded 81% classification accuracy to an existing damage classification system, indicating good agreement between the two metrics. Further ability of tailing factor to score features of hair damage reflecting explosion conditions demonstrates the broad applicability of the metric to assess damage to hairs containing a diverse set of morphological features. </p>
Data from: Influence of parameter settings in automated scoring of AFLPs on population genetic analysis
The use of procedures for the automated scoring of AFLP fragments has recently increased. Corresponding software does not only automatically score the presence or absence of AFLP fragments, but also allows an evaluation of how different settings of scoring parameters influence subsequent population genetic analyses. In this study, we used the automated scoring package RAWGENO to evaluate how five scoring parameters influence the number of polymorphic bins and estimates of pairwise genetic differentiation between populations (Fst). Steps were implemented in R to automatically run the scoring process in RAWGENO for a set of different parameter combinations. While we found the scoring parameters minimum bin width and minimum number of samples per bin to have only weak influence on pairwise Fst values, maximum bin width and bin reproducibility had much stronger effects. The minimum average bin fluorescence scoring parameter affected Fst values in an only moderate way. At a range of scoring parameters around the default settings of RAWGENO, the number of polymorphic bins as well as pairwise Fst values stayed rather constant. This study thus shows the particularities of AFLP scoring, be it either manual or automatical, can have profound effects on subsequent population genetic analysis.
Data from: FEATHER: automated analysis of force spectroscopy unbinding and unfolding data via a Bayesian algorithm
Single-molecule force spectroscopy (SMFS) provides a powerful tool to explore the dynamics and energetics of individual proteins, protein-ligand interactions, and nucleic acid structures. In the canonical assay, a force probe is retracted at constant velocity to induce a mechanical unfolding/unbinding event. Next, two energy landscape parameters, the zero-force dissociation rate constant (ko) and the distance to the transition state (Δx‡), are deduced by analyzing the most probable rupture force as a function of the loading rate, the rate of change in force. Analyzing the shape of the rupture force distribution reveals additional biophysical information, such as the height of the energy barrier (ΔG‡). Accurately quantifying such distributions requires high-precision characterization of the unfolding events and significantly larger data sets. Yet, identifying events in SMFS data is often done in a manual or semiautomated manner and is obscured by the presence of noise. Here, we introduce, to our knowledge, a new algorithm, FEATHER (force extension analysis using a testable hypothesis for event recognition), to automatically identify the locations of unfolding/unbinding events in SMFS records and thereby deduce the corresponding rupture force and loading rate. FEATHER requires no knowledge of the system under study, does not bias data interpretation toward the dominant behavior of the data, and has two easy-to-interpret, user-defined parameters. Moreover, it is a linear algorithm, so it scales well for large data sets. When analyzing a data set from a polyprotein containing both mechanically labile and robust domains, FEATHER featured a 30-fold improvement in event location precision, an eightfold improvement in a measure of the accuracy of the loading rate and rupture force distributions, and a threefold reduction of false positives in comparison to two representative reference algorithms. We anticipate FEATHER being leveraged in more complex analysis schemes, such as the segmentation of complex force-extension curves for fitting to worm-like chain models and extended in future work to data sets containing both unfolding and refolding transitions.
Comparative analysis of metabolic models of microbial communities reconstructed from automated tools and consensus approaches
<p>Generated draft and consensus reconstructions for the manuscript "Comparative analysis of metabolic models of microbial communities reconstructed from automated tools and consensus approaches" (Hsieh, Tandon, Verbruggen, & Nikoloski).</p>
Analysis of maize growth under drought in an automated plant phenotyping platform
<p>Supplemental data accompanying the PhD thesis of Lennart Verbraeken.</p>
Constellation analysis of automated local public transport shuttles in the north western development area of Berlin
<h2><strong><span>Konstellationsanalyse: Autonome ÖPNV-Shuttles im Entwicklungsband Nordwest </span></strong></h2> <p>Das Dokument beinhalten stellt eine deutschsprachige Ergebnisdokumentation einer Konstellationsanalyse aus dem Projekt „NOWEL4 – Berliner NordWestraum Level 4" dar . Die Konstellationsanalyse ist ein Brückenkonzept für inter- und transdisziplinäre Zusammenarbeit, das u.a. in der Nachhaltigkeits- und Technikforschung eingesetzt werden kann. Die Ergebnisse zeigen eine gegenwärtige Konstellation und eine Zielkonstellation der Einbindung autonomer Shuttles in den ÖPNV im Berliner Entwicklungsband NordWest. </p> <h2> </h2> <p> </p>
Empirical Review of Automated Analysis Tools on 47,587 Ethereum Smart Contracts
<p>This dataset contains the full output of the execution of 9 state-of-the-art automated analysis tools on 47,518 Solidity contracts.</p> <p>The data structure is as follows</p> <pre><code>├─ results │ └─ <tool_name> │ └─ <dataset_name> │ └─ <contract_address> │ ├─ <result.log> # stdout of the analysis │ └─ <result.json> # parsable output analysis</code></pre> <p> </p>
Supplementary Material - Dataset for "Automating Quantum Software Maintenance: Flakiness Detection and Root Cause Analysis"
<h2>README</h2> <p>The dataset consists of the following components:<br> <br>- `<strong>prompts.txt</strong>`: This file contains the prompts used for large language models.<br> <br>- `<strong>Dataset</strong>` directory: includes general information about the dataset. Specifically, the `dataset.xlsx` file lists flaky and non-flaky tests, along with their root causes and fix types.<br> <br>- `<strong>Full</strong>` directory contains two subdirectories: `Flaky` and `Non-flaky`. Each of these directories is organized by individual GitHub organization projects, with each project having its list of repository subdirectories. These subdirectories are further divided into “issues” and “pull requests” (PRs).</p> <p><br>- `<strong>Method</strong>` level subdirectory has a similar structure but contains extracted code snippets at the method level instead of full code listings. The `code.diff` file is copied over and left unaltered. </p> <p><br>- <strong>Issue Directories (IRs):</strong> Named with an `issueID` template, each issue directory contains a `log.issue` file that includes the extracted description, comments, and metadata.<br> <br>- <strong>PR Directories (PRs)</strong>: Named using the `prID` template, each PR directory contains the text, comments, and metadata in the `pr.log` file. The text of the associated issue is stored in the `log.issue` file. Code listings are stored in a file with the `.bug` suffix, while the corresponding fixed version is in a `.fix` file. The `code.diff` file contains the patch that transforms the `.bug` version into the `.fix` version.</p> <p><br><strong>Additional notes:</strong><br>Issues with associated pull requests in `dataset.xlsx` are combined into the pull request directory template. If two pull requests are listed for a row, a PR directory is created for each. Due to updates in the extended dataset, some repositories have been renamed or archived, meaning the current repository directory names in `Dataset` will include both the previous and new names if it has been changed (e.g., a repository previously saved as Qiskit/qiskit-terra may now be saved as Qiskit/qiskit following the renaming from qiskit-terra to qiskit).</p> <h2>Directory Structure:</h2> <p><br>├── prompts.txt<br>├── Dataset/<br> └── dataset.xlsx<br>├── Full/<br> ├── Flaky/<br> └── <Organization>/<Repository>/...<br> ├── Non-Flaky/<br> └── <Organization>/<Repository>/...<br>├── Method/<br> ├── Flaky/<br> └── <Organization>/<Repository>/...<br> ├── Non-flaky/<br> └── <Organization>/<Repository>/...</p> <p> </p>
Towards the conservation of Brazilian legumes: Summary, land cover and use analytics, and a comparative analysis of locations count methods (curated vs. automated [buffer-dissolution])
<p>This dataset presents key findings from the study "<strong>Automating and Enhancing Species Extinction Risk Assessments with Historical Land Use and Land Cover Data</strong>."</p> <p>The file `<em>summary-threatened-legume-species.ods</em>` provides a summary of all threatened species of the Leguminosae family native to Brazil, including IUCN category and criteria, location counts, and the year of the latest assessment, sourced from official records. Additionally, it includes calculated values for Area of Occupancy (AOO) and Extent of Occurrence (EOO), along with trend data indicating natural area change rates (decline [positive number, red] or growth [negative number, green]) within both AOO and EOO. Location counts are further detailed across various buffer radii (1-5 km), utilizing a buffer-dissolution method for automated location counting. Each buffer radius is analyzed to assess AOO and EOO decline, where '1' indicates a species is threatened and '0' indicates it is not. This data provides an efficient method for screening threatened species under criterion B of the IUCN Red List guidelines.</p> <p>The file `<em>overlay-analysis.ods</em>` contains overlay analysis results for AOO and EOO of each species using MapBiomas land use and land cover (LULC) data (specifically MapBiomas Brazil, collection 7.1) from 1985 to 2021, covering all threatened legume species. This file provides both absolute area in square kilometers and percentages for each LULC class. The overlay analysis results support estimates of growth and decline trends for each LULC class.</p> <p>The file `<em>trend-analysis.ods</em>` presents results of annual rate estimates from trend analysis across LULC classes, and including both natural and anthropic groupings. A complete JSON database with p-values and R² values is provided in `<em>trend-analysis.json</em>`.</p> <p>This approach, combining all results for each species in a comprehensive, merged dataset, allows for effective filtering and ranking of the most threatened species as well as identification of their primary threats.</p> <p>We recommend opening the ODS files with LibreOffice, as Microsoft Excel may experience issues parsing decimal formats accurately.</p> <p>More detailed maps and graphs are available at <a title="LULC-MapBiomas-Leguminosae" href="https://github.com/lsbjordao/LULC-MapBiomas-Leguminosae" target="_blank" rel="noopener">https://github.com/lsbjordao/LULC-MapBiomas-Leguminosae</a>.</p>
DATASET - Automated grain sizing from UAV imagery of a gravel-bed river: benchmarking of three object-based methods and analysis of particle-size clustering
<p>Dataset used to compute the grain size distributions from in-field line sampling and digitally on orthoimages with automated methodologies and by manual labelling. It also contains the data used to produce spatial statistics.</p>
Dataset for Thesis "Automated Consistency of Legal and Software Architecture System Specifications for Data Protection Analysis"
<p>Dataset for Thesis "Automated Consistency of Legal and Software Architecture System Specifications for Data Protection Analysis"</p>
Prediction of Extubation Readiness in Extreme Preterm Infants by the Automated Analysis of CardioRespiratory Behavior
ClinicalTrials.gov study NCT01909947. IPD Sharing: Not stated. Countries: 2. Publications: 2.
Automated Quantitative Ulcer Analysis Study
ClinicalTrials.gov study NCT04420962. IPD Sharing: NO. Countries: 2. Publications: 8.
Pain Detection Through Automated Video Analysis
ClinicalTrials.gov study NCT04011189. IPD Sharing: NO. Countries: 1. Publications: 0.
Automated Vision Assessment and Impairment Detection Through Gaze Analysis in Wet AMD Patients
ClinicalTrials.gov study NCT06518512. IPD Sharing: NO. Countries: 1. Publications: 10.
Automated Analysis of EIT Data for PEEP Setting
ClinicalTrials.gov study NCT03653806. IPD Sharing: NO. Countries: 1. Publications: 3.
User-centric Study of Patients' Receptiveness Towards the Web-based Automated Vision Impairment Gaze-tracking Analysis Systems
ClinicalTrials.gov study NCT07338513. IPD Sharing: Not stated. Countries: 1. Publications: 11.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.