Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
318
datasets available to search
ShareScore release 0.7.1
Dataset results
318 results for “Data mining”
Data mining tool to discover DevOps trends from public repositories: Predicting Release Candidates with gthbmining.rc
<p>Public repositories have been performing an essential role in bring- ing software and services to technical communities and general users. Most of the cases, public repositories have a DevOps tool, with a live and historical database behind it, to support delivering and all steps this software or service should adopt before going to production. This paper introduces gthbmining, a data mining set of tools to discover DevOps trends from public repositories, and presents the module gthbmining.rc. Considering the premise of a GitHub public repository, the main contribution here is pre- dicting release candidates, an important label a software release has. The methodology, architecture, components and interfaces are explained, as well as potential users. The results show a reliable and flexible tool, as classifiers metrics and graphics are provided, along with the possibility to add new data mining algorithms in the open source module presented. Related works are also supplied, and a conclusion shows the outcomes gthbmining.rc can provide.</p>
Data from: Evidence-based restoration of freshwater biodiversity after mining: Experience from Central European spoil heaps
<p>Supplementary datasets for <strong>Kolar et al. (in press) Evidence-based restoration of freshwater biodiversity after mining: Experience from Central European spoil heaps. <em>Journal of Applied Ecology.</em></strong></p> <p>All related information can be found in the cited paper.</p> <p>When using the dataset for anything, cite the Kolar et al. <em>Journal of Applied Ecology </em>paper.</p> <p>For additional information, refer the paper or write to robert.tropek@gmail.com and kolarvojta@seznam.cz </p>
Data supplementing the conference paper "Who you gonna call? Analyzing web requests in Android applications", 14th International Conference on Mining Software Repositories 2017.
<p>This repository contains the data supplementing the paper:</p> <p>M. Rapoport, P. Suter, E. Wittern, O. Lhótak, J. Dolby, "Who you gonna call? Analyzing web requests in Android applications", MSR 2017.</p> <p>A detailed description of the data is included in the archive in README.md.</p>
Data mining-based machine learning methods for improving hydrological data: a case study of salinity field in the Western Arctic Ocean
<p><span><span>Salinity variations in Arctic Ocean determine the strength of stratification,</span> <span>ocean circulation, and biogeochemical cycles. Therefore, a<span>ccurate </span>salinity product is of great significance for our study of the Arctic Ocean. The mean density structure and wind-driven surface circulation of the Arctic Ocean are largely dominated by the anti-cyclonic Beaufort Gyre in the Canadian Basin, along with the Transpolar Drift<span> (Hall </span><span>et al.,2022)</span>. We focus on the salinity in Western Arctic Ocean. Multiple machine learning methods were used to reconstruct annual salinity product in the Western Arctic Ocean temporal for the period 2003-2022.</span></span></p>
Data from: Abiotic legacies mediate plant-soil feedback during early vegetation succession on rare earth element mine tailings
<p>An increasing number of studies have shown how feedback interactions between plants and soil can influence primary and secondary succession. However, very little is known about the patterns and mechanisms of such plant-soil feedbacks on stressed mine tailings ecosystem, which can be severely contaminated by a range of toxic elements. </p> <p>In a two-phase plant-soil feedback experiment based on the rare earth element (REE) mine tailing soil, we investigated biotic (changes in bacterial and fungal community) and abiotic legacies (changes in chemical properties) of three pioneer grass species, and examined feedback effects of three grasses, two legumes and two woody plants with different root traits.</p> <p>Positive plant-soil feedback was found in Miscanthus sinensis, Paspalum thunbergii and Tephrosia candida, and neutral feedback was observed in other four plants. These effects corresponded with an increase of nutrients and total organic carbon, as well as a decrease of acidity and extractable aluminum and REEs. There were less signs of biotic changes in the conditioned tailings. </p> <p>The correlation analysis suggested a relationship between responses to soil legacies and root traits, as well as root economics spectrum. On the mine tailings, acquisitive species with higher specific root length appeared to have greater potential for positive feedback. </p> <p>Synthesis and application: Our study shows that early succession on contaminated REE mine tailings may lead to more positive plant-soil feedback than predicted based on results of non-contaminated soils, mainly due to the alleviation of abiotic stress in tailings. Therefore, the improvement of specific abiotic soil stress and the trait-based selection of acquisitive plants should be preferentially considered to promote the primary restoration of degraded land.</p>
Digitization Workflow for Data Mining in Production Technology applied to a Feed Axis of a CNC Milling Machine
<p>Dataset accompanying the publication "Digitization Workflow for Data Mining in Production Technology<br>applied to a Feed Axis of a CNC Milling Machine" (<a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.procs.2024.01.017" target="_blank" rel="noreferrer noopener"><span>https://doi.org/10.1016/j.procs.2024.01.017</span></a>).</p>
Data from: Do metal mines and their runoff affect plumage color? A regional scale study of streak-backed orioles in south-central Mexico
<p>Metal mining causes serious ecological disturbance, due partly to heavy metal (HM) pollution that can accumulate at mining sites themselves and be dispersed downstream as runoff. Plumage coloration is important in birds' social and ecological interactions and sensitive to environmental stressors, and several local-scale studies have found decreased carotenoid-based plumage and/or increased melanin-based plumage in wild birds exposed to HM pollution. We investigated regional-scale effects of proximity to mines and their downstream rivers as a proxy of exposure to HM-contaminated mining waste on plumage coloration in streak-backed orioles (<em>Icterus pustulatus</em>) in south-central Mexico. We measured the plumage color of museum skins using reflectance spectrometry and digital photography, then used geographic information systems to estimate each specimen's distance from the nearest mining concession and river and determine whether that river's watershed contained mines. Proximity to mines and their downstream rivers was related to ventral (but not dorsal) carotenoid-based coloration; birds collected farther from mines had more vivid yellow-orange breast plumage, and belly plumage was more vivid and redder with increasing distance from rivers with upstream mines. Breast background reflectance unexpectedly decreased with mine distance and was higher among birds whose nearest river had mines upstream. The area (but not reflectance) of melanin-based plumage was also related to mines. The area of dark back streaks decreased with mine distance, while the bib patch was smaller among birds presumably more exposed to mining waste. While some of these results are consistent with predicted effects of HM pollution on plumage, most were not straightforward, and effects differed among plumage patches and variables. Further investigation is needed to understand the direct (e.g., toxicity, oxidative stress) and/or indirect (e.g., decreased availability of carotenoid-rich food) mechanisms responsible and their individual, population, and community-level implications. </p>
Data from: Freshwater ecological quality assessment of the gold mining Mashcon watershed, Cajamarca - Peru
<p>These data were generated to investigate the aquatic community gradients across different types of anthropogenic impacts (reference, mining, rural and urban) and Andean environmental gradients (headwaters, midstream and downstream) in the Mashcon watershed. </p> <p>The macroinvertebrates' taxa abundance and abundance of traits modalities serve as biological values. The latter in combination with abiotic measurements of the freshwater habitats (i.e. physicochemical water quality and hydromorphology) constitute the ecological data.</p> <p>The values were obtained from river sediments (macroinvertebrates collection), water samples for laboratory analyses, field protocols and in-situ water quality measurements at 40 sites, wherein 6 were downstream of the gold mine's artificially recharged headwaters, 8 sites at near-pristine headwaters tributary streams, 14 sites at midstream rural areas and 12 sites at downstream urban areas (sample grouping information is shown in the Field_protocol.xlsx file).</p>
Text Mining of Archaeological Reports for Urban farming (data and code)
<p>This release is created to create a DOI in Zenodo for the data related to Fischer, AD, van Londen, H, Blonk-van den Bercken, AL, Visser, RM and Renes, J. 2021. Urban farming and ruralisation in the Netherlands (1250 up tot the nineteenth century), unravelling farming practice and the use of (open) space by synthesising archaeological reports using text mining. Nederlandse Archeologische Rapporten 68. Amersfoort: Rijksdienst voor het Cultureel Erfgoed. <a href="https://www.cultureelerfgoed.nl/publicaties/publicaties/2021/01/01/urban-farming-and-ruralisation-in-the-netherlands">https://www.cultureelerfgoed.nl/publicaties/publicaties/2021/01/01/urban-farming-and-ruralisation-in-the-netherlands</a></p>
Accompanying data for the paper "Making Sense of Wildlife Habitat Use on Active Oil Sands Mines: Quasi-experiments, Occupancy Models, Trends Assessments, and Upland Habitat Reclamation"
<p>This data set contains both the raw species detection records and the derived occupancy model data used to assess usage patterns for the nine species of wildlife. Data have been anonymized by using non-identifying company and lease names. These attributes are not required to reproduce the results in this paper and was done per contractual requirements between LGL Limited and its clients.</p> <p>Data is currently being reviewed by the client and will be shared publicly once final approval has been received.</p>
Data showcase papers published in the Mining Software Repositories (MSR) conference
<p>Data regarding data showcase papers published in the Mining Software Repositories (MSR) conference.</p> <p>The following data files are included.</p> <p>citation-table.csv: SWEBOK areas of citing studies<br> citations.bib: Bibliographic details of citing studies<br> citing_dp_dois_citations.txt: Citations of citing studies<br> data_papers.bib: MSR data papers<br> dp_dois_citations.txt: Citations of data papers<br> false_citations.bib: Citing studies that don't actuall use data papers<br> msr-all: Bibliographic details of all MSR papers<br> ndp_dois_citations.txt: Citations of non-data papers<br> ndp_rand_dois_citations.txt: Citations of a randomly chose non-data paper weighted sample<br> self-citations.csv: Data papers citations by their authors</p> <p> </p>
The coincidence degree between the geochemical behavior of elements and the periodic variation of elements based on geochemical data of C2 coal seam in the Fengfeng mining area of the Handan Coalfield in Hebei, China
<p>In this data-set, based on the geochemical data of C2 coal seam in the Fengfeng mining area of the Handan Coalfield in Hebei (China), where provided ideal coal samples changing continuously from low-rank metamorphic coal to high-rank metamorphic coal, the coincidence degree (or similarity degree) between the geochemical behavior of 57 elements and the periodic variation of elements during the thermal metamorphism process is calculated.</p>
Identification of processes in Cu-ore heap leaching using Cu isotopes and leachate chemistry at Tschudi mine, northern Namibia - Supplementary data
<p>This is a supplementary dataset to the paper:</p> <p>Sracek O., Ettler V., Mihaljevič M., Kříbek B., Mapani B., Penížek V., Zádorová T., Vaněk A. (2024): Identification of processes in Cu-ore heap leaching using Cu isotopes and leachate chemistry at Tschudi mine, northern Namibia. <em>Hydrometallurgy</em> <strong>228</strong>, 106356.</p> <p>This research was supported by the Johannes Amos Comenius Programme (OP JAC), project No. CZ.02.01.01/00/22_008/0004605, Natural and anthropogenic georisks. The dataset is published under the Creative Commons Attribution 4.0 International License (CC-BY-4.0). This license allows others to distribute, remix, adapt, and build upon the dataset for any purpose, even commercially, as long as they give appropriate credit to the original creator(s).</p>
Data mining antibody sequences for database searching in bottom-up proteomics
<p>Mass spectrometry (MS)-based proteomics is a powerful method for identifying and quantifying antibodies. Among the various MS approaches, bottom-up proteomics is especially effective for analyzing thousands of antibodies in complex mixtures. In this method, proteins are enzymatically digested into smaller peptides, typically using the protease trypsin, which are then analyzed via mass spectrometry. These peptides are matched to sequences in standard databases like UniProt or NCBI-RefSeq for identification.</p> <p>However, a major limitation of this approach is the absence of comprehensive disease-specific antibody databases. Current databases, such as UniProt, include only a fraction of the antibody sequences present in the human body. For instance, as of January 2024, UniProt contains just 38,800 immunoglobulin sequences, far short of the billions of antibodies the human immune system can produce. As a result, relying on such limited databases can lead to under-detection of antibodies, particularly those associated with specific diseases. Expanding antibody databases with disease-specific sequences is crucial for improving the accuracy of MS-based proteomics in identifying antibodies relevant to human health.</p> <p>Recently, through next-generation sequencing of antibody gene repertoires, it has become possible to obtain billions of antibody sequences (in amino acid format) by annotating, translating, and numbering antibody gene sequences. These large numbers of sequences are now available in public databases such as the <a href="https://opig.stats.ox.ac.uk/webapps/oas/" rel="nofollow">Observed Antibody Space</a>. We hypothesize that using these theoretical antibody sequences as new databases for bottom-up proteomics could address the current lack of antibody coverage in standard databases.</p> <p>We developed a workflow to create disease-specific antibody peptide databases for bottom-up proteomics. The workflow details are available on <a href="https://github.com/trinhxt/SDU_Immunoinformatics">GitHub</a>. The database and metadata files generated by this workflow are stored in this Zenodo dataset, and they are used in DAT-DB — a web application that allows researchers to obtain FASTA files of disease-specific antibody peptides for direct use in bottom-up proteomics (see <a href="https://trinhxt.shinyapps.io/DAT-DB/">Demo version</a>).</p> <p>Each database file in this dataset is in <em>.duckdb</em> format and contains tables with 10 columns: Sequence, Filename, Patient, BSource, BType, Isotype, N_patient, N_antibody, Length_aa, and CDR3. The "<strong>Sequence</strong>" column contains tryptic peptides. "<strong>Filename</strong>" is the file where the data was collected. "<strong>Patient</strong>" refers to the patient number as listed in <em>metadata2.csv</em>. "<strong>BSource</strong>" refers to the B-cells' source, and "<strong>BType</strong>" refers to the type of B-cells. "<strong>Isotype</strong>" specifies the antibody isotype (IgA, IgD, IgE, IgG, IgM, or Bulk). "<strong>N_patient</strong>" indicates the number of patients having this peptide, and "<strong>N_antibody</strong>" specifies the number of antibodies containing this peptide. "<strong>Length_aa</strong>" indicates the number of amino acids in the peptide, while "<strong>CDR3</strong>" shows whether the peptide is found in the CDR3 region.</p> <p>The file <em>metadata1.csv</em> contains information about each database file, while <em>metadata2.csv</em> provides details about the sources of the collected antibodies.</p>
Table 5. Analytical data for miscellaneously attributed samples from Broken Hill, including the Broken Hill Consols Mine. a in Compositions of silver halides from the Broken Hill district, New South Wales
<p><b>Table</b> 5. Analytical data for miscellaneously attributed samples from Broken Hill, including the Broken Hill Consols Mine.a</p><table><tbody><tr><th>No.</th><th><b>Analyses</b></th></tr></tbody><tbody><tr><th>D46338b</th><td>Cl</td><td>60</td><td>57</td><td>57</td><td>53</td><td>56</td><td>57</td><td>58</td><td>68</td><td>56</td><td>66</td><td></td></tr><tr><th></th><td>Br</td><td>32</td><td>35</td><td>34</td><td>37</td><td>35</td><td>34</td><td>32</td><td>31</td><td>35</td><td>33</td><td></td></tr><tr><th></th><td>I</td><td>8</td><td>8</td><td>9</td><td>10</td><td>9</td><td>9</td><td>10</td><td>1</td><td>9</td><td>1</td><td></td></tr><tr><th>D26168c</th><td>Cl</td><td>72</td><td>76</td><td>63</td><td>65</td><td>60</td><td>77</td><td>77</td><td>72</td><td>66</td><td>65</td><td></td></tr><tr><th></th><td>Br</td><td>28</td><td>24</td><td>37</td><td>35</td><td>40</td><td>22</td><td>23</td><td>28</td><td>34</td><td>35</td><td></td></tr><tr><th>D28184d</th><td>I Cl</td><td>75</td><td>68</td><td>86</td><td>70</td><td>70</td><td>1 70</td><td>78</td><td>70</td><td>84</td><td></td><td></td></tr><tr><th></th><td>Br</td><td>25</td><td>30</td><td>14</td><td>28</td><td>30</td><td>29</td><td>21</td><td>30</td><td>16</td><td></td><td></td></tr><tr><th>D28185d</th><td>I Cl</td><td>55</td><td>2 56</td><td>55</td><td>2 56</td><td>55</td><td>155</td><td>1 60</td><td>50</td><td>45</td><td>50</td><td>56</td></tr><tr><th></th><td>Br</td><td>43</td><td>43</td><td>44</td><td>43</td><td>44</td><td>43</td><td>40</td><td>46</td><td>45</td><td>45</td><td>41</td></tr><tr><th></th><td>I</td><td>2</td><td>1</td><td>1</td><td>1</td><td>1</td><td>2</td><td></td><td>4</td><td>10</td><td>5</td><td>3</td></tr></tbody></table><p>a Analyses reported as in Table I. b Trace S detected. C Traces Pb, Cu, S, As detected. d Consols Mine; traces Pb, As, S, Fe detected.</p>
Data from: Limited biomass recovery from gold mining in Amazonian forests
<ol> <li>Gold mining has rapidly increased across the Amazon Basin in recent years, especially in the Guiana shield, where it is responsible for >90% of total deforestation. However, the ability of forests to recover from gold mining activities remains largely unquantified. </li> <li>Forest inventory plots were installed on recently abandoned mines in two major mining regions in Guyana, and re-censused 18 months later, to provide the first ground-based quantification of gold mining impacts on Amazon forest biomass recovery. </li> <li>We found that woody biomass recovery rates on abandoned mining pits and tailing ponds are amongst the lowest ever recorded for tropical forests, with close to no woody biomass recovery after 3-4 years. </li> <li>On the overburden sites (i.e. areas not mined but where excavated soil is deposited), however, aboveground biomass recovery rates (0.4 - 3.5 Mg ha<sup>-1</sup> yr<sup>-1</sup>) were within the range of those recorded in other secondary forests across the Neotropics following abandonment of pastures and agricultural lands. </li> <li>Our results suggest that forest recovery is more strongly limited by severe mining-induced depletion of soil nutrients, especially nitrogen, than by mercury contamination, due to slowing of growth in nutrient-stripped soils. </li> <li>We estimate that the slow recovery rates in mining pits and ponds currently reduce carbon sequestration across Amazonian secondary forests by ~21,000 t C yr<sup>-1</sup>, compared to the carbon that would have accumulated following more traditional land uses such as agriculture or pasture. </li> <li> <i>Synthesis and applications</i>. To achieve large-scale restoration targets, Guyana and other Amazonian countries will be challenged to remediate previously mined lands. The recovery process is highly dependent on nitrogen availability rather than mercury contamination, affecting woody biomass regrowth. The significant recovery in overburden zones indicates that one potential active remediation strategy to promote biomass recovery may be to back-fill mining pits and ponds with excavated soil. </li> </ol>
Dataset: Process Mining for Reliability Modeling of Smart Manufacturing Systems with Reduced Data Requirements
<p>Operational state logs from the Industry 4.0 Lab, University of Southern Denmark.</p> <p>"I4.0Lab_state_log.csv" -> without failures</p> <p>I4.0Lab_state_log_failures.csv -> with failures</p>
Appendix-C: Text Data and Mining Licensing Conditions
<p>Appendix C is associated with <em><strong>Chapter 11: Text Data and Mining Ethics</strong></em> of the book -- Manika Lamba and Margam Madhusudhan (2021) Text Mining for Information Professionals: An Uncharted Territory, SpringerNature.</p>
Evaluation of restoration on post–mining areas using a game theory - data
<p>Supplementary data to the article - numerical values from graphs (Fig 1-2, 5-12) + input data for calculations. The resulting NE probability values can be verified at equsis.com (outside of Fig 8, the limit is 32 cases).</p>
A novel big data mining framework for reconstructing large-scale daily MAIAC AOD across China from 2000 to 2020
<p>This dataset contains the mean reconstructed MAIAC AOD and mean percentage of coverage of reconstructed MAIAC AOD across China from 2000 to 2020. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.