Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,773
datasets available to search
ShareScore release 0.9.0
Dataset results
1,773 results for “mines”
FIGURES 93 – 97 in Leaf-mining Nepticulidae (Lepidoptera) from record high altitudes: documenting an entire new fauna in the Andean páramo and puna
FIGURES 93 – 97. Stigmella rigida Diškus & Stonis, sp. nov. 93, male adult, holotype; left side; 94, same, right side; 95, 96, male genitalia, holotype, genitalia slide no. AD 625, capsule without phallus; 97, same, phallus (ZMUC).
FIGURES 98 – 102 in Leaf-mining Nepticulidae (Lepidoptera) from record high altitudes: documenting an entire new fauna in the Andean páramo and puna
FIGURES 98 – 102. Stigmella altiplanica Diškus & Stonis, sp. nov. 98, male adult, holotype; 99, same, left side; 100, male genitalia, holotype, genitalia slide no. AD 647, capsule without phallus; 101, same, phallus; 102, same, dorsal view of capsule (ZMUC).
FIGURES 32 – 33 in Leaf-mining Nepticulidae (Lepidoptera) from record high altitudes: documenting an entire new fauna in the Andean páramo and puna
FIGURES 32 – 33. Details of morphology and bionomics. 32, abdominal apex of female of Stigmella calceolarifoliae sp. n.; 33, leaf-mines of S. paramica sp. n. in Pentacalia leaf, and cocoon.
FIGURES 118 – 122. Stigmella schoorli Puplesis & Robinson. 118, 119 in Leaf-mining Nepticulidae (Lepidoptera) from record high altitudes: documenting an entire new fauna in the Andean páramo and puna
FIGURES 118 – 122. Stigmella schoorli Puplesis & Robinson. 118, 119, male adult; 120, male genitalia, capsule without phallus, paratype, genitalia slide no. Diškus 201; 121, same, holotype, genitalia slide no. Diškus 200 (after Puplesis & Robinson 2000); 122, phallus, paratype, genitalia slide no. Diškus 201 (ZMUC).
Data: Evolution of anatomical concept usage over time: Mining 200 years of biodiversity literature
<p>This data package contains data and results corresponding to the paper titled "Evolution of anatomical concept usage over time: Mining 200 years of biodiversity literature". </p>
Energy Mining
<p><strong>Overview of Data</strong></p> <p>Excessive energy consumption in mobile apps can be a consequence of energy greedy hardware, bad programming practices, or particular API usage patterns. We present the largest to date quantitative and qualitative empirical investigation into the categories of API calls and usage patterns that—in the context of the Android development framework—exhibit particularly high energy consumption profiles. By using a hardware power monitor, we measure energy consumption of method calls when executing typical usage scenarios in 55 mobile apps from different domains. Based on the collected data, we mine and analyze energy-greedy APIs and usage patterns. We zoom in and discuss the cases where either the anomalous energy consumption is unavoidable or where it is due to suboptimal usage or choice of APIs. Finally, we synthesize our findings into actionable knowledge and recipes for developers on how to reduce energy consumption while using certain categories of Android APIs and patterns</p> <p><strong>Abstract</strong></p> <p>Energy consumption of mobile applications is nowadays a hot topic, given the widespread use of mobile devices. The high demand for features and improved user experience, given the available powerful hardware, tend to increase the apps’ energy consumption. However, excessive energy consumption in mobile apps could also be a consequence of energy greedy hardware, bad programming practices, or particular API usage patterns. We present the largest to date quantitative and qualitative empirical investigation into the categories of API calls and usage patterns that—in the context of the Android development framework—exhibit particularly high energy consumption profiles. By using a hardware power monitor, we measure energy consumption of method calls when executing typical usage scenarios in 55 mobile apps from different domains. Based on the collected data, we mine and analyze energy-greedy APIs and usage patterns. We zoom in and discuss the cases where either the anomalous energy consumption is unavoidable or where it is due to suboptimal usage or choice of APIs. Finally, we synthesize our findings into actionable knowledge and recipes for developers on how to reduce energy consumption while using certain categories of Android APIs and patterns.</p>
Map of the Roman mining hydraulic system - Las Médulas, León, Spain
<p>This resource is related with the Roman hydraulic system linked to Las Médulas gold mining complex in northwest Iberia. The dataset includes a detailed digital cartography of the hydraulic network, which extends over 1100km. It identifies 41 canals distributed between La Cabrera and El Bierzo regions, (33 and 8, respectively), with 14 canals supplying water to Las Médulas. Additionally, the dataset includes the locations of the Roman mining sites.</p> <p><a title="Map of the Roman mining hydraulic system - Las Médulas, León, Spain" href="https://earth.google.com/earth/d/1argDbqGaEMdlE7Uomb4z70LGYce9pCBb?usp=sharing">View on Google Earth (3D)</a></p> <p><a title="Map of the Roman mining hydraulic system - Las Médulas, León, Spain" href="https://www.google.com/maps/d/edit?mid=12yHtyAN3qJcmrt-P_3abv3rd6PaI_ZQ&usp=sharing">View on Google Maps (2D)</a></p> <p> </p>
Unraveling community adaption and survival strategy of soil microbiome under vanadium stress in nationwide mining environments
<p class="Abstract"><span><span>The vanadium (V) smelters soil harbor wide ranges of microorganisms, whose survival relies on their metabolic activities under stress.</span><span> Nonetheless, the characteristics and functions of soil microbiome in V mining environments have not been recognized at a continental scale. This study investigates microbial diversity, community assembly and metabolic traits of soil microbiome across 90 V smelters in China. A decrease in alpha diversity is observed, along with community variation, which is also jointly explained by other environmental, climatic and geographic factors. Null model shows that V promotes homogeneous selection. V also mediates co-occurrence patterns, with increased positive interspecific associations under higher V concentrations (</span><span>></span><span>559.6 mg/kg)</span><span>, e.g., <em>f_Gemmatimonadaceae</em>, <em>Nocardioides</em>, <em>Micromonospora</em>, <em>Rubrobacter</em>.</span><span> In addition, 67 metagenome assembled genomes are retrieved via metagenomic analysis. The metabolic pathways of keystone taxa are disentangled to reveal their putative involvement in the V(V) reduction process. Nitrate and nitrite reductase (<em>nirK</em>, <em>narG</em>), and <em>mtrABC</em> are found to be taxonomically affiliated with <em>Micromonospora</em>. sp, <em>FEN-1250</em>. sp, <em>Nocardioides</em>. sp, etc. Additionally, reverse citric acid cycle (rTCA) serves the main carbon fixation pathway, synthetizing alternative energy for putative V reducers, highlighting a synergistic relationship between autotrophic and heterotrophic processes to support the microbial survival. Our findings comprehensively reveal the driving forces for soil community variation under V stress, suggesting the robust strategies adopted by indigenous microorganisms to alleviate V impact, which can be exploited for bioremediation application.</span></span></p>
Dataset: Behavior of Participants in Hands-on Cybersecurity Training Suitable for Process Mining
<p>This repository contains supplementary materials for the following journal paper:</p> <p>Radek Ošlejšek, Martin Macák, Karolína Dočkalová Burská.<br><em>Hands-on cybersecurity training behavior data for process mining.</em><br>In Elsevier Data in Brief. 2023.<br>Available as open-access article on <a href="https://doi.org/10.1016/j.dib.2023.109956">https://doi.org/10.1016/j.dib.2023.109956</a></p> <p><strong>Contents</strong></p> <p>Datasets store event logs of trainees participating in hands-on cybersecurity exercises organized in the <a href="https://www.kypo.cz">KYPO Cyber Range</a>. The data includes training scenarios (expected behavior), raw event logs in the JSON format, and aggregated behavioral data suitable for process mining analysis.</p> <ol> <li><strong>Data1:</strong> A dataset of 52 trainees participating in the <a href="https://gitlab.ics.muni.cz/muni-kypo-trainings/games/locust-3302">Locust 3302</a> exercise adapted an insider attack scenario. No time restrictions were posed on playtime. The data file is structured as follows: <ul> <li>training_definition.json: The exercise content – cybersecurity tasks and hints. The training is based on the <a href="https://gitlab.ics.muni.cz/muni-kypo-trainings/games/locust-3302">Locust 3302</a> game adapted to an insider attack scenario.</li> <li>training_events: Recorded progress of trainees within the exercise, i.e., the status of completing tasks.</li> <li>command_histories: Recorded commands executed on network hosts.</li> <li>process_mining.csv: Complete PM-ready dataset suitable for process discovery or conformance analysis.</li> <li>process_mining_simplified.csv : Reduced PM-ready dataset with semantically identical events being removed.</li> </ul> </li> <li><strong>Data2:</strong> A dataset of 48 trainees participating in the original <a href="https://gitlab.ics.muni.cz/muni-kypo-trainings/games/locust-3302">Locust 3302</a> exercise. Three supervised training sessions were restricted to two hours of playtime. The structure follows the structure of Data1.</li> <li><strong>Tool:</strong> A Java application used to aggregate raw JSON data and transform them into a CSV format suitable for process mining techniques.</li> </ol> <p><strong>How to cite</strong></p> <p>If you use or build upon the materials, please use the BibTeX entry below to cite the original work.</p> <pre><code>@article{Oslejsek2023dataset, author = {Radek O\v{s}lej\v{s}ek and Martin Mac\'{a}k and Karol\'{i}na {Do\v{c}kalov\'{a} Bursk\'{a}}}, title = {Hands-on cybersecurity training behavior data for process mining}, journal = {{Data in Brief}}, publisher = {Elsevier}, issn = {2352-3409}, year = {2023}, volume = {52}, doi = {10.1016/j.dib.2023.109956}, url = {https://www.sciencedirect.com/science/article/pii/S2352340923009873} }</code></pre>
Mining the health disparities and minority health bibliome: A computational scoping review and gap analysis of 200,000+ articles
<p>Without comprehensive examination of available literature on health disparities and minority health (HDMH), the field is left vulnerable to disproportionately focus on specific populations or conditions, curtailing our ability to fully advance health equity. Using scalable open-source methods, we conducted a computational scoping review of more than 200,000 articles to investigate major populations, conditions, and themes in the literature as well as notable gaps. We also compared trends in studied conditions to their relative prevalence in the general population using insurance claims (42 million Americans). HDMH publications represent 1% of articles in MEDLINE. Most studies are observational in nature, though randomized trial reporting has increased five-fold in the last twenty years. Half of all HDMH articles concentrate on only three disease groups (cancer, mental health, endocrine/metabolic disorders), while hearing, vision, and skin-related conditions are among the least well-represented despite substantial prevalence. To support further investigation, we also present HDMH Monitor, an interactive dashboard and repository generated from the HDMH bibliome.</p>
Figure 2 in Differences in the soil seed bank of a mining area and its surroundings: a case study inserted in the Cerrado domain
Figure 2. Curves of species accumulation for the seed bank of Vazante city, northwest of the state of Minas Gerais, Brazil. A. Accumulation curve for the seed bank and confidence interval in a mining area and surrounding area separated; B. Accumulation curve for the seed bank, confidence interval and first-order Jackknife richness estimator in a mining area and surrounding area together.
Figure 3 in Differences in the soil seed bank of a mining area and its surroundings: a case study inserted in the Cerrado domain
Figure 3. Two-dimensional ordination diagram based on the nonmetric multidimensional scaling (nMDS) for species density of the seed bank from a mining pit and surrounding areas in Vazante municipality, northwest of the state of Minas Gerais, Brazil.
Global ML-ready dataset for mining areas in satellite images
<p>This dataset is a global resource for machine learning applications in mining area detection and semantic segmentation on satellite imagery. It contains Sentinel-2 satellite images and corresponding mining area masks + bounding boxes for 1,210 sites worldwide. Ground-truth masks are derived from <a href="https://doi.org/10.1594/PANGAEA.942325" target="_blank" rel="noopener">Maus et al. (2022)</a> and <a href="https://doi.org/10.5281/zenodo.6806817" target="_blank" rel="noopener">Tang et al. (2023)</a>, and validated through manual verification to ensure accurate alignment with Sentinel-2 imagery from specific timestamps. </p> <p>The dataset includes three mask variants:</p> <ul> <li>Masks exclusively from Maus et al. (n=1,090)</li> <li>Masks exclusively from Tang et al. (n=817)</li> <li>A preferred mask selected from either Maus or Tang based on alignment quality determined during manual review (n=1,210).</li> </ul> <p>Each tile corresponds to a 2048x2048 pixel Sentinel-2 image, with metadata on mine type (surface, placer, underground, brine & evaporation) and scale (artisanal, industrial). For convenience, the preferred mask dataset is already split into training (75%), validation (15%), and test (10%) sets. </p> <p>Furthermore, dataset quality was validated by re-validating test set tiles manually and correcting any mismatches between mining polygons and visually observed true mining area in the images, resulting in the following estimated quality metrics: </p> <table> <tbody> <tr> <td> </td> <td>Combined</td> <td>Maus</td> <td>Tang</td> </tr> <tr> <td>Accuracy</td> <td>99.78</td> <td>99.74</td> <td>99.83</td> </tr> <tr> <td>Precision</td> <td>99.22</td> <td>99.20</td> <td>99.24</td> </tr> <tr> <td>Recall</td> <td>95.71</td> <td>96.34</td> <td>95.10</td> </tr> </tbody> </table> <p>Note that the dataset does not contain the Sentinel-2 images themselves but contains a reference to specific Sentinel-2 images. Thus, for any ML applications, the images must be persisted first. For example, Sentinel-2 imagery is available from Microsoft's Planetary Computer and filterable via STAC API: <a href="https://planetarycomputer.microsoft.com/dataset/sentinel-2-l2a" target="_blank" rel="noopener">https://planetarycomputer.microsoft.com/dataset/sentinel-2-l2a</a>. Additionally, the temporal specificity of the data allows integration with other imagery sources from the indicated timestamp, such as Landsat or other high-resolution imagery.</p> <p>Source code used to generate this dataset and to use it for ML model training is available at <a href="https://github.com/SimonJasansky/mine-segmentation" target="_blank" rel="noopener">https://github.com/SimonJasansky/mine-segmentation</a>. It includes useful Python scripts, e.g. to <a href="https://github.com/SimonJasansky/mine-segmentation/blob/main/src/data/05_persist_pixels_masks.py">download Sentinel-2 images via STAC API</a>, or to <a href="https://github.com/SimonJasansky/mine-segmentation/blob/main/src/data/06_make_chips.py">divide tile images (2048x2048px) into smaller chips (e.g. 512x512px)</a>. </p> <p>A database schema, a schematic depiction of the dataset generation process, and a map of the global distribution of tiles are provided in the accompanying images. </p>
Comprehensive Ethereum Execution Data for Object-Centric Process Mining of Decentralized Applications (DApps)
<p>The dataset pertains to the collection and analysis of blockchain execution data, particularly from Ethereum-based Decentralized Applications (DApps). This data includes transactions, transaction receipts, and detailed transaction traces, documenting the execution steps performed by the Ethereum Virtual Machine (EVM). Such traces are essential for understanding the interaction between smart contracts and accounts, including Contract Accounts (CAs) and Externally Owned Accounts (EOAs).</p> <p>A blockchain is an append-only ledger that chronologically records data in blocks. Each block contains transactions that signify state transitions, and transaction receipts that provide a hashed result of these transitions to ensure uniform results across different executions. The dataset includes a classification of Ethereum accounts, detailing the functions and interactions between EOAs and CAs, where CAs deploy and execute smart contract code.</p> <p>The dataset captures the granular operational data of blockchain transactions, such as function calls, contract creations, and log entries generated by smart contracts. These details are crucial for creating object-centric event logs, aiding in process mining and analysis to bridge the gap between theoretical process models and actual execution.</p> <p>Contract creations and function calls are fundamental components of the dataset. The former documents the deployment of smart contracts, including the mechanics of contract updates and additions through various design patterns. Function calls between accounts are also extensively logged, providing insights into the flow of Ethereum's native token, Ether, and other transactional data within the blockchain.</p> <p>Delegated calls and log entries represent more specialized interactions within Ethereum, where delegated calls allow contracts to use code from other contracts to manipulate their own state, supporting upgradeable contract designs. Log entries, specified within smart contract code, facilitate the communication of contract execution details to external systems.</p> <p>To handle the diverse and dynamic nature of blockchain data, the dataset employs the Object-Centric Event Log (OCEL) format. This format accommodates multiple object types in a single log, addressing issues such as event divergence and convergence, typical of traditional single-case logs. The latest version, OCEL 2.0, supports documenting dynamic object roles and relationships, improving the fidelity of logs in capturing blockchain operations.</p> <p>In summary, the dataset is structured to support a comprehensive analysis of blockchain behaviors, particularly focusing on Ethereum DApps. It is tailored to assist researchers and practitioners in understanding and analyzing the decentralized execution of smart contracts and the associated data flows within the blockchain environment.</p>
Mining and Extractivism Records Data for Bibliometric Analysis (Scopus database 1992-2020)
<p>The dataset file export from scopus database and the dataset file export as bibliometrix file on excel format from biblioshiny.</p>
BioVAE: a pre-trained latent variable language model for biomedical text mining
<p>We release BioVAE, the first large-scale pre-trained latent variable language model for the biomedical domain, which uses the OPTIMUS framework to train on large volumes of biomedical text.</p> <p>This version contains the pre-trained models for text mining tasks such as named entity recognition or relation extraction, and text generation task.</p> <p>Explanation of each file: (lt32: latent_size = 32, beta05: beta=0.5)</p> <ul> <li>pm-full-lt32-beta00</li> <li>pm-full-lt32-beta05</li> <li>pm-full-lt768-beta00</li> <li>pm-full-lt768-beta05</li> <li>pm-full-generation</li> </ul>
Text-mining-derived cancer associations - PMED
<p>JSON file containing gene::cancer, variant::cancer and gene/variant::cancer:drug-response associations derived from an in-house text-mining pipeline utilizing full-text articles from PMC.</p>
Fatiando a Terra Data: Osborne Mine, Australia - Airborne total-field magnetic anomaly
<p>This is a section of a survey acquired in 1990 by the Queensland Government, Australia. The data are good quality with approximately 80 m terrain clearance and 200 m line spacing. The anomalies are very visible and present interesting processing and modelling challenges, as well as plenty of literature about their geology.</p> <p><strong>Note:</strong> This is a processed and formatted version of the source dataset below. It's meant for use in documentation and tutorials of the <a href="https://www.fatiando.org">Fatiando a Terra</a> project. Please <strong>cite the original authors</strong> when using this dataset.</p> <p><strong>Changes made: </strong>Change the horizontal datum from GDA94 to WGS84. Convert terrain clearance to flight height using an SRTM grid. Keep only the coordinates, AWAGS leveled magnetic anomaly, and flight line ID. Cut to a smaller region containing only the 2 anomalies of interest.</p> <p><strong>Source: </strong>Geophysical Acquisition & Processing Section 2019. MIM Data from Mt Isa Inlier, QLD (P1029), magnetic line data, AWAGS levelled. Geoscience Australia, Canberra. <a href="http://pid.geoscience.gov.au/dataset/ga/142419">http://pid.geoscience.gov.au/dataset/ga/142419</a></p> <p><strong>Source license: </strong><a href="http://pid.geoscience.gov.au/dataset/ga/142419">CC-BY</a></p> <p><strong>Repository: </strong><a href="https://github.com/fatiando-data/osborne-magnetic">https://github.com/fatiando-data/osborne-magnetic</a></p>
Mining folded proteomes in the era of accurate structure prediction
<p>Supplementary data to accompany the manuscript “Mining folded proteomes in the era of accurate structure prediction”. Contains three zip files with fold matching search results to support results in the main text.</p>
Requirements (enhancements) of 64 Mozilla projects mined from Bugzilla
<p>The dataset consists of 4200 enhancements that are in some form of dependency with the others (such as blocks, depends_on etc.) This data spans from 08/05/2001 to 09/08/2019.</p> <p>This dataset is gathered using Bugzilla's REST API for 64 projects. We have id, summary, priority, severity, type, version, target_milesotne, product, depends_on, blocks fields information for each one of these enhancements.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.