Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
355
datasets available to search
ShareScore release 0.9.0
Dataset results
355 results for “data extraction”
Data from: How many specimens make a sufficient training set for automated three dimensional feature extraction?
<p>Deep learning has emerged as a robust tool for automating feature extraction from 3D images, offering an efficient alternative to labour-intensive and potentially biased manual image segmentation methods. However, there has been limited exploration into the optimal training set sizes, including assessing whether artificial expansion by data augmentation can achieve consistent results in less time and how consistent these benefits are across different types of traits. In this study, we manually segmented 50 planktonic foraminifera specimens from the genus Menardella to determine the minimum number of training images required to produce accurate volumetric and shape data from internal and external structures. The results reveal unsurprisingly that deep learning models improve with a larger number of training images with eight specimens being required to achieve 95% accuracy. Furthermore, data augmentation can enhance network accuracy by up to 8.0%. Notably, predicting both volumetric and shape measurements for the internal structure poses a greater challenge compared to the external structure, due to low contrast differences between different materials and increased geometric complexity. These results provide novel insight into optimal training set sizes for precise image segmentation of diverse traits and highlight the potential of data augmentation for enhancing multivariate feature extraction from 3D images. </p>
COVID-19 daily situation reports - partial data extracted to .csv
<p><span>Government of Nepal Ministry of Health and Population (MoHP) and World Health Organization’s Country Office in Nepal published daily situation reports monitoring the pandemic activity on a national level. Daily situation reports were published in a PDF format including up-to-date figures on the number of COVID-19 cases, the number of PCR-tests performed at each laboratory as well as the respective number of positive test results. The provided data contains parts of the data from these .pdf reports extracted to a .csv file. The data was extracted during a joint project between MoHP, WHO Country Office Nepal, WHO South East Asia Regional Office, Polytechnique Montréal, University of Amsterdam, and Karlsruhe Institute of Technology. </span></p>
Extracting Ridge and Valley Lines in Mountainous Areas from Airborne Lidar Data by Utilizing Line Feature Strength
<p><strong><span>Background</span></strong><strong><span>:</span></strong><span> </span><span>DEMs (digital elevation models) are very important in many fields, such as in Geomatics and in water conservation of mountainous areas etc. Geomorphic feature lines are necessary data for the topography interpolation and computation from DEMs.</span></p> <p><strong><span>Methods</span></strong><strong><span>:</span></strong><span> </span><span>Instead of the parameter space, we propose a novel automatic extraction of Geomorphic feature lines in the feature space from discrete airborne LiDAR (Light detection and ranging) data by TVM (tensor voting method) developed originally for image data in this article. A tensor field for discrete airborne LiDAR points is first established and then utilizing the TVM, a new geometric feature metric of data, the line feature strength, was captured. A practical line growing method based on the local maximum line feature strength is proposed in the article.</span></p> <p><strong><span>Results</span></strong><strong><span>:</span></strong><span> </span><span>Compared with the general line growing that is based on a certain threshold, our line growing method is quite effective, in particular for the extraction of primary and minor ridge and valley lines in mountainous areas.</span></p> <p><strong><span>Conclusions</span><span>:</span></strong><span> </span><span>The method presented in this paper is fast and automated and can furnish operators with a wealth of detailed information about minor line features. This will enable the extraction of ridge and valley lines tailored to specific requirements. It is no doubt that the method developed here can be generalized to a large amount of Lidar data.</span></p>
Compound profiling matrices extracted from screening data
<p>Compound profiling matrices record assay results for compound libraries tested against panels of targets. In addition to their relevance for exploring structure-activity relationships, such matrices are of considerable interest for chemoinformatic and chemogenomic applications. For example, profiling matrices provide a valuable data resource for the development and evaluation of machine learning approaches for multi-task activity prediction. However, experimental compound profiling matrices are rare in the public domain. Although they are generated in pharmaceutical settings, they are typically not disclosed. Herein, we present an algorithm for the generation of large profiling matrices, for example, containing more than 100,000 compounds exhaustively tested against 50 to 100 targets. The new methodology is a variant of bi-clustering algorithms originally introduced for large-scale analysis of genomics data. Our approach is applied here to assays from the PubChem BioAssay database and generates profiling matrices of increasing assay or compound coverage by iterative removal of entities that limit coverage. Weight settings control final matrix size by preferentially retaining assays or compounds. In addition, the methodology can also be applied to generate matrices enriched with active entries representing above-average assay hit rates.</p>
A Genetic Algorithm Approach to Regenerate Image from a Reduce Scaled Image Using Bit Data Count-Figure 9. Resized image and extracted data for a 100*100 size image
<p>For the extraction, our goal is to divide an image into smaller blocks and keep the row and column data for these blocks. But for our experiment we used a single block, which means taking the full image as a single block. For bigger image we should always divide the image in separate blocks and work on them par rally. As in figure 8, after extracting the data we can add the row and column bits information in the resized image or saved in a separate file. For proof of concept we saved it in a text file. And later that file is used to feed GA to make the fitness function, in figure 9 an extraction has been shown. In upper and side textbox containing the information which later is saved in a text file.</p>
A Genetic Algorithm Approach to Regenerate Image from a Reduce Scaled Image Using Bit Data Count-Figure 8. Sample Data Extraction for a 20*20 size image
<p>As in figure 7 we are storing the extra data which is look like figure 8. Where a 20*20 size image of alphabet ‘A’ data has been stored. When we regenerate image, we are using these data.</p>
The development and evaluation of an online application to assist in the extraction of data from graphs for use in systematic reviews (data repository)
<p>These are the data we generated in our evaluation of the graphical user interface.</p> <p>Please see our publication on Wellcome Open Research for information about the evaluations.</p>
Extended Data for Publication "Testing the Spectroscopic Extraction of Suppression of Convective Blueshift"
<p>Efforts to detect low-mass exoplanets using stellar radial velocities (RVs) are currently limited by magnetic photospheric activity. Suppression of convective blueshift is the dominant magnetic contribution to RV variability in low-activity Sun-like stars. Due to convective plasma motions, the magnitude of RV contributions from the suppression of convective blueshift is roughly correlated with the depth of formation of photospheric spectral lines used to compute the RV time series. Meunier et al. (2017), used this relation to demonstrate a method for spectroscopic extraction of the suppression of convective blueshift in order to isolate RV contributions, including planetary RVs, that contribute equally to the timeseries for each spectral line. In this publication, we extract disk-integrated solar RVs from observations over a 2.5 year time span made with the solar telescope integrated with the HARPS-N spectrograph at the Telescopio Nazionale Galileo (La Palma, Canary Islands, Spain). We apply the methods outlined by Meunier et al. (2017) - as part of this analysis, we fit Gaussian line profiles to 765 iron lines measured over 457 exposures.</p> <p>Here, we provide the complete line list (Table 1; Table1_LineList.csv) used in our analysis, and the resulting RVs time series in their entirety (Table 3; Table3_TimeSeries.csv). Wavelengths are given in Angstroms, and RVs in m/s.</p> <p>We also include 4 CSV files with the line fit parameters of each line profile: Each row corresponds to a single exposure time (corresponding to the JDs in Table 3) and each column corresponds to a specific spectral line (with wavelength specified in Table 1).</p> <p>We fit each spectral line to a Gaussian of the form:</p> <p><span class="math-tex">\(f(\lambda) = p_1 - p_2 \exp \left[- {1 \over 2} \left({{\lambda - p_3} \over p_4}\right)^2 \right]\)</span></p> <p>p<sub>1</sub> is the continuum level in arbitrary units (LineProfiles_Continuum.csv)<br> p<sub>2</sub> is the line strength in arbitrary units (LineProfiles_Amplitude.csv)<br> p<sub>3</sub> is the line shift in Angstroms (LineProfiles_Shift.csv)<br> p<sub>4</sub> is the line width in Angstroms (LineProfiles_Width.csv)</p>
Systematic Mapping Extraction Data
<p><strong>Context:</strong> Agile methodologies use continuous delivery and adaptability to develop software that meets the needs of its users. However, such methods are prone to accumulate technical debt (TD). Agile teams must balance the benefits and risks of incurring debt by managing TD items. Knowing how TD management is conducted<br>in agile software teams can help agile practitioners increase their ability to handle TD items. <strong>Aims:</strong> To investigate, based on the state of the art, how agile software teams have managed existing TD items in their projects. <strong>Method: </strong>We carried out a systematic mapping study covering articles from 2010 to 2023. <strong>Results:</strong> Among the agile methodologies, Scrum was the most used for TD management. Regarding practices, agile teams employ mostly user stories and sprint backlogs to identify TD items. Sprints and sprint backlogs are used to monitor debt and refactoring to prevent and repay TD items. <strong>Conclusion: </strong>This study maps current knowledge on TD management in agile methodologies, serving as a starting point for new investigations in the area.</p>
Ice Anatomy: A Benchmark Dataset and Methodology for Automatic Ice Boundary Extraction from Radio-Echo Sounding Data
<p>The measurement of ice thickness is of great importance for the accurate estimation of glacier volume and the delineation of their bedrock topography. In particular, this is a crucial factor in forecasting the future evolution of glaciers in the context of a changing climate. In order to derive the ice thickness, the travel time of electromagnetic waves in radargrams acquired by radio-echo sounding (RES) systems is analyzed. This can only be achieved by identifying the ice surface and underlying ice bottom in corresponding radargrams. Manually identifying these two reflection horizons in RES data is a laborious and time-consuming process. Consequently, scientists are attempting to automate this task through the use of techniques such as deep learning. Such automation can significantly reduce the time between a field campaign and the calculation of the glacier's ice thickness distribution. In this paper, we present the first benchmark dataset for delineating the ice surface and bottom boundaries in RES data, to facilitate straightforward comparisons of deep learning models in the future. The ``IceAnatomy'' dataset comprises radargrams and the corresponding manual picks, amounting to a total of over 45,000km of observations. The RES data originates from three sources: FAU, CReSIS, and AWI. The dataset comprises different RES systems as well as different pre-processing methods. In addition, the data was acquired over a large range of geographical and glaciological settings, featuring different thermal regimes present in Antarctica and the Southern Patagonian Icefield. This diversity ensures that the models' behaviors can be analyzed in different scenarios. We define a standardized train-test split for each source in the dataset. This allows us to introduce not only a baseline model trained on the entire training set (the ``omni'' model), but also three source-specific baseline models. The source-specific models are trained exclusively on the subset of the training data acquired by the specified source. The baseline models provide an initial benchmark against which subsequent models can be compared. The source-specific models demonstrate more accurate results than the omni model. For the FAU, CReSIS, and AWI test sets, the source-specific models achieve Mean Meter Errors of 2.1m, 23.1m, and 4.9m for the ice surface and 9.1m, 78.2m, and 29.3m for the ice bottom. In relation to the mean measured ice thickness, these errors equate to 1.2%, 3.1%, and 0.3% for the ice surface and 4.9%, 10.4%, and 1.5% for the ice bottom.</p> <p> </p> <p> For more information, please read the following paper:</p> <p>[Coming soon. Currently under review.]</p> <p>Please also cite this paper if you plan on using the dataset.</p> <p> </p> <p>For the implementation of a baseline model please visit:</p> <p>[Coming soon]</p> <p> </p>
Average daily air temperature, precipitation and relative sunshine duration for Vallon de Nant catchment, extracted from gridded MeteoSwiss data (1961-2020)
<p>This excel file contains time series of daily temperature, precipitation and relative sunshine duration obtained as the spatial average of gridded data sets. The underlying original gridded data sets produced by MeteoSwiss are known as RhiresD, TabsD and SrelD. All meta data are included in the excel file.</p> <p>The data can e.g. be used for hydrological modelling. Comparison to local station data is not included.</p> <p><strong>This data set as well as the original data set should be cited</strong>.</p>
UPLC-QTOF data from a river sample extracted with various solid phase extraction phases - mz5 files with scans in centroid spectrum format
<p>This dataset has been acquired from a river sample (Marne River, France) collected as part of the Screenatm'eau project (Observatoire des Sciences de l’Univers - Enveloppes FLUides de la Ville à l’Exobiologie - OSU-EFLUVE, Université Paris-Est Créteil). The sample was processed by solid-phase extraction using multiple cartridges and phases, in triplicates (see the list of samples in the samplemetadata.csv file).</p> <p>Data was acquired with a SYNAPT HDMS QTOF (Waters) coupled with a Nano ACQUITY UPLC System (Waters), in ESI positive mode. Raw data was converted using Proteowizard MSConvert version 3.0.21288, using zlib compression, with the CWT peakpicking algorithm (snr=0, peakSpace=0) to transform profile to centroid data, and the zeroSamples option (removeExtra 1-). The chosen output formats were .mzML and .mz5.</p> <p>A list of markers was obtained after peak picking and alignment across all samples using the PatRoon R package with the OpenMS algorithm, and exported as a markertable.csv file. Each detected marker has m/z and retention time values (see the variablemetadata.csv file), and intensity values in all samples (markertable.csv).</p> <p>For file naming explanation and details about each SPE phase, see the readme.txt file.</p>
Systematic Literature Mapping - Selected Articles Data Extraction
<p>- CSV with the raw data extracted from abstract + full-text review from the selected articles.</p> <p>- CSV with normalized data.</p> <p>- CSV with the normalized data after being a feature engineering step in which some columns have been deleted.</p> <p>- CSV with formal context to extract the concepts lattice.</p>
List of non-naturalized plant species present in France extracted from Pl@ntNet data (exotic ornamental and cultivated plants in particular).
<p>This dataset contains the list of plant species that have been observed on the French territory using the <a href="https://plantnet.org/">Pl@ntNet </a>application and that are NOT know as being either native or naturalized according to Kew's Plants of the World Online repository (<a href="https://powo.science.kew.org/">POWO</a>). Such species are typically exotic species managed by humans in anthropized environments such as guardens, houses or cultivated areas. This includes commercialized plants for various usage such as ornemental plants, eatable plants, phytotherapy, etc. The list contains 5,589 species, each associated with its scientific name and the number of valid Pl@ntNet observations of that species geo-localized in the metropolitan French territory. </p>
Data for cyanobacteria, nutrients and color from 588 lakes in Finland Norway Sweden and UK extracted from the WISER database
<p>Data on phosphorus, color, cyanobacteria biovolume, total phytoplankton biovolume and cyanobacteria proportion of the total phytoplankton biovolume used for GAMM</p>
Systematic Literature Review - Selected Articles - Data Extraction
<p>- CSV with the raw data extracted from abstract + full-text review from the selected articles.</p> <p>- CSV with normalized data.</p> <p>- CSV with the normalized data after being a feature engineering step in which some columns have been deleted.</p> <p>- CSV with formal context to extract the concepts lattice.</p>
Data from: How many specimens make a sufficient training set for automated three dimensional feature extraction?
Open the record for dataset details and reuse information.
Extracting abundance information from DNA-based data
Open the record for dataset details and reuse information.
Data from: A comprehensive suite for extracting neuron signals across multiple sessions in one-photon calcium imaging
Open the record for dataset details and reuse information.
Using convolutional neural networks to efficiently extract immense phenological data from community science images
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.