Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
23
datasets available to search
ShareScore release 0.9.0
Dataset results
23 results for “Code Extraction”
Deep learning to extract the meteorological by-catch of wildlife cameras: Supporting data, models and code
<p>This repository contains the data, models and code to train and deploy deep learning models related to the paper "Deep learning to extract the meteorological by-catch of wildlife cameras" published in the journal Global Change Biology (<a href="https://doi.org/10.1111/gcb.17078"><strong>https://doi.org/10.1111/gcb.17078</strong></a>).</p>
Data and Code for "Extracting reproductive parameters from GPS tracking data for a nesting raptor in Europe"
<p>Understanding population dynamics requires estimation of demographic parameters. We build on existing approaches to develop a new tool that uses GPS tracking data to estimate breeding propensity and breeding success, and show that this tool yielded accurate predictions for two red kite populations in Central Europe. The tool is available as an R package at <a href="https://github.com/Vogelwarte/NestTool">https://github.com/Vogelwarte/NestTool</a> and will facilitate the estimation of demographic parameters from tracking data to inform population assessments. The files in this repository contain the data and analytical code to replicate the results of the publication in the Journal of Avian Biology (DOI: 10.1111/jav.03246). The version contained in this repository does not include updates and improvements that occurred after the 29 August 2024.</p>
TrainTicket microservice testbench extracted information for our work: Evaluating ChatGPT's Proficiency in Understanding and Answering Microservice Architecture Queries Using Source Code Insights
<p>It contains the CSV file output of our tool implemented in the paper: "Evaluating ChatGPT’s Proficiency in Understanding and Answering Microservice Architecture Queries Using Source Code Insights." applied to the TrainTicket microservice testbench. The information in this CSV was used for In-Context-Learning for ChatGPT.</p>
Source data and code for: Existing fossil fuel extraction would warm the world beyond 1.5°C
<p>Source data and code for the study, "Existing fossil fuel extraction would warm the world beyond 1.5°C." Datasets 1-4 include mine-level data collected for China (Dataset 1), India (Dataset 2), and five other countries (Dataset 3) that are among the world's top nine coal producers - the United States, Indonesia, Australia, South Africa, and Poland. Dataset 4 includes global and country-level output data from the 1,000-run Monte Carlo simulation. <Committed_Reserves_Monte_Carlo_Input_Data.zip> includes data and code to replicate the Monte Carlo simulation.</p>
Train and Evaluation Code, Road Classification Models and Test set of the paper "Insights into the Effects of Image Overlap and Image Size on Semantic Segmentation Models Trained for Road Surface Area Extraction from Aerial Orthophotography"
<p>This repository contains the Python scripts built for training and evaluation of the implementation, together with the test data and the resulting road segmentation models corresponding to the paper "Insights into the Effects of Image Overlap and Image Size on Semantic Segmentation Models Trained for Road Surface Area Extraction from Aerial Orthophotography". The scripts make use of the Tensorflow with Keras framework and their additional required dependencies.</p> <p>The training and validation set is based on the binary SROADEX dataset (<a href="../records/6482346">https://zenodo.org/records/6482346</a>) that was re-split into tiles that feature the image resolutions (256 x 256, 512 x 512, and 1024 x 1024 pixels) and image overlaps (0% and 12.5%) considered in this study. The data have been generated using scripts developed in Python using Open Source libraries (GDAL/OGR and MapScript) for rasterization of vector cartography that represents the axes of the different types of roads (urban, interurban and rural). This binary road data contains information from 16 full orthoimages (28.5 km * 18.5 km) with spatial resolution of 0.5 m/pixel from the insular and peninsular Spanish territory. Due to the size on disk of approximately 492 gigabytes, this training and validation data is only available upon request from the corresponding author. The test set has been generated from a novel area from Palencia (Spain) and features 18 million pixels labelled with the positive "Road" class. The test sets are provided in the repository for each resolution (with no overlap), so that additional DL models can be evaluated on the same data and compared with the results achieved in this study.</p> <p>The structure of the information shared in this repository is as follows:<br>The scripts have been grouped by tile resolution (256, 512 and 1024). First, the test set and the evaluation script can be found. For each tile resolution, there are two subfolders (corresponding to the "no overlap" and "12.5% overlap"). In each case, the Python scripts for training the models in the three repetitions are shared, and the trained models (H5 format) are shared in compressed form. Finally, for each resolution we also share the testing dataset which consists of two folders.</p> <p>The material is distributed under a CC-BY 4.0 license.</p>
Source code for models of floral initiation in pea and gene expression data extracted from published sources
<p>The dataset contains the source code for computational models of a gene network controlling transition to flowering in pea (<em>Pisum sativum</em>). The models were based on ordinary differential equations (ODE) or neural networks. It also includes data on the expression dynamics of genes involved in the network, which was used for model fitting. The expression data was extracted from the following papers: </p> <p>Hecht, V., Laurie, R. E., Schoor, K. Vander, Ridge, S., Knowles, C. L., Liew, L. C., Sussmilch, F. C., et al. (2011). The Pea GIGAS Gene Is a FLOWERING LOCUS T Homolog Necessary for Graft-Transmissible Specification of Flowering but Not for Responsiveness to Photoperiod. 23, 147–161. doi:10.1105/tpc.110.081042</p> <p>Sussmilch, F. C., Berbel, A., Hecht, V., Schoor, K. Vander, Ferrándiz, C., Madueño, F., et al. (2015). Pea VEGETATIVE2 Is an FD Homolog That Is Essential for Flowering and Compound In fl orescence Development. 27, 1046–1060. doi:10.1105/tpc.115.136150</p> <p>The source code of the DEEP software used for parameter optimization in the model fitting can be found in the Gitlab repository (https://gitlab.com/mackoel/deepmethod/-/tree/master).</p> <p>The files are the supplement to the following manuscript, submitted to Frontiers in Genetics:</p> <p>"Dynamical Modeling of the Core Gene Network Controlling Transition to Flowering in <em>Pisum sativum</em>" by Polina Pavlinova, Maria G. Samsonova, and Vitaly V. Gursky.</p> <p>All possible questions can be sent to: Polina Pavlinova (polina.pavlina1004@gmail.com), Vitaly Gursky (gursky@math.ioffe.ru).</p>
Data set for ODI cricket matches from 1987 to 2023 (extracted from ESPN Cricinfo) and code (R) used for a statistical study
<p>Here I present the data and code that has been used to study the statistical evolution of ODI cricket. The preprint for this research is available at: </p> <div> <div> <div> <table> <tbody> <tr> <td><a href="https://doi.org/10.48550/arXiv.2406.11652">https://doi.org/10.48550/arXiv.2406.11652</a> <div><span>Focus to learn more</span></div> </td> </tr> </tbody> </table> </div> </div> </div> <div> </div>
The role of non-native plant species in modulating riverbank erosion: a systematic review - extracted coded articles dataset
<p>Data extraction sheet of code typologies for the review titled: The role of non-native plant species in modulating riverbank erosion: a systematic review. The uploaded csv file contains all the extracted values for each typology. </p>
Yields, Cannabinoids Quantification, and Predictive Programming Codes Using Machine Learning for Non-Psychoactive Cannabis Flowers and Extracts (Cannabis sativa L.) Cultivated in Ecuador.
<p>This publication presents data from various extraction methods, including maceration, Soxhlet, and supercritical fluids, performed on different cannabis flower varieties (Cannabis sativa L.) under varying operating conditions. We quantified the amounts of CBD, THC, CBG, and CBN in the extracts produced by each method using High-Performance Liquid Chromatography (HPLC). Using this data, we developed a machine learning algorithm in RStudio to make predictions and determine the best conditions and yields for each extraction method. The analysis focuses on different varieties of non-psychoactive cannabis cultivated in Ecuador at altitudes over 2,450 m.a.s.l.</p>
Extracting interaction potentials from trajectories: code and a set of 30 simulated contest trajectories
<p>Dataset and code provided as supporting materials for the manuscript "Modeling animal contests based on spatio-temporal dynamics". See the Supporting Information of the manuscript for detailed instructions of use.</p>
Extracting Educational Code Scenarios from Python Textbooks
<p><strong>Dataset Overview:</strong> This dataset complements the research project titled "Extracting Learning Scenarios from Python Textbooks." It consists of 1,017 chapter titles collected from 76 Python textbooks.</p><p><strong>Research Findings:</strong> Our analysis revealed that learning scenarios (referred to as "Scenarios") are a prevalent theme, constituting approximately 39.5% of the total chapter titles in comparison to other content categories. We further categorized these scenarios into four types, including Application Programming Interfaces, Data and Processing, Graphical User Interfaces, and other scenarios. Additionally, we identified a list of 19 Python modules commonly used within these scenarios.</p><p><strong>Purpose:</strong> We envision that this work and its insights can serve as stepping stones and lay the groundwork for further extraction and the effective application of how Python can be utilized for its diverse audience.</p>
Dataset for "Extracting Traceability Links between Javadoc Sentences and JUnit Test Code Lines"
<p>Dataset for "Extracting Traceability Links between Javadoc Sentences and JUnit Test Code Lines"</p>
The public part of the interview extractions as thematic codings: used for the consumer behavior study
<p>Extractions from interview transcripts created by the thematic analysis, this data is part of the interview study done on consumer energy behavior during price hikes,</p>
Code and data for "ActivityGen: Extracting Enabled Activities from Screenshots"
<p>The code and data for the paper "ActivityGen: Extracting Enabled Activities from Screenshots" is provided here. </p> <p><strong>Abstract</strong></p> <p>Many tasks in organizations are performed in a desktop environment. It is possible to record users' interactions in a desktop environment by taking screenshots when an action happens. The result is an interaction log. By considering the associated images of a record, it is possible to detect which activity was performed and which activities were enabled. This information can be extracted, resulting in a translucent event log. Such a translucent event log is valuable and can be used as input for dedicated process-mining techniques. The results can be used to analyze human-computer interactions or create bots for robotic process automation. However, current techniques for extracting information on enabled activities rely on template matching, which is rigid and sensitive to variations. To solve this issue, we present our modular framework, ActivityGen. ActivityGen detects and labels graphical user interface elements by also considering additional information. ActivityGen uses more advanced techniques to overcome the limitations of previous approaches and can extract information without a user's input. Furthermore, it can be adjusted to a user's needs. It detects graphical user interface elements more accurately than state-of-the-art techniques and labels them faster, more robust, and more domain-oriented than state-of-the-art techniques.</p> <p><strong>Data</strong></p> <p>ReDraw_CLS and ReDraw_ViSM are specified in the work. </p> <p>The basis for ReDraw_CLS is the ReDraw dataset. We focus on the following components: Button, CheckBox, EditText, Image, ImageButton (which we refer to as icon), RadioButton, and Switch. We noticed that the examples of ImageView and ImageButton are similar, primarily consisting of icon images. Therefore, we removed the ImageView class and introduced an Image class instead. The Image class contains images from the validation set of the Coco validation set 2017 and the YouTube Thumbnails dataset, enabling the detection of general website images.</p> <p>ReDraw_ViSM iterates add 6,000 synthetically created buttons to the former dataset by distributing them in the same ratio into train, test, and validation sets.</p> <p>lm_basic and lm_extended contain the text training for the language models. </p> <p><strong>Code</strong></p> <p>The code allows for the execution of ActivityGen. Moreover, we provide our evaluation scripts. However, the models do not have to be trained. The models' weights are provided in the model folder.</p>
A Mechanism for Automatically Extracting Reusable and Maintainable Code Idioms from Software Repositories
<p>The provided dataset contains the data used by "A Mechanism for Automatically Extracting Reusable and Maintainable Code Idioms from Software Repositories", in order to extract maintainable and reusable code idioms from the most popular GitHub repositories. The dataset includes the code snippets, the Abstract Syntax Trees and repositories information about Control Flow Statements (if-blocks, for-blocks, enhanced-for-blocks, try-blocks, while-blocks, switch-blocks and do-blocks) data coming from the top GitHub projects.</p>
Identifiers in source code extracted from 13,000,000 public GitHub repositories (October 2016)
<p>101 files, indexed LZO archives for Hadoop/Spark.<br> Each file is text lines with the format:</p> <p>(‘<GitHub repo name>', [(‘<name>', <count>),(‘<name>', <count>),(‘<name>', <count>)])</p> <p> </p>
Dataset and code for 'AI-based Knowledge Extraction from the Bioprinting Literature for identifying technology trends'
<p>Zip file containing the dataset and code for the paper 'AI-based Knowledge Extraction from the Bioprinting Literature for identifying technology trends'. The dataset is composed of:</p> <ul> <li>A train_data.csv file, containing all annotated keywords used for classifier training.</li> <li>A filt_ls.pkl file, containing the sentences used to train the embeddings model.</li> <li>A train.py file, to train the composite keyword annotation model.</li> </ul> <p>The authors acknowledge the supported by the European Union’s Horizon 2020 research and innovation program under the project GIOTTO: “Giotto: Active ageing and osteoporosis: The next challenge for smart nanobiomaterials and 3D technologies,” grant agreement no. 814410.</p>
Data and Codes of "A Pharmacological Representation-based LSTM Network for Drug–Drug Interaction Extraction"
<p>The datasets and source codes used and/or analyzed in study "A Pharmacological Representation-based LSTM Network for Drug–Drug Interaction Extraction".</p>
Code for: <em>Triolena</em> anisophylly data extraction
Open the record for dataset details and reuse information.
Codes used to extract coeliac disease, type I diabetes, rotavirus vaccination and DTP vaccination from CPRD Aurum
<p>Codes used in medcode and prodcode fields of CPRD Aurum data</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.