Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
318
datasets available to search
ShareScore release 0.7.1
Dataset results
318 results for “Data mining”
Research exceptions in copyright laws around the world in Legal reform to enhance global text and data mining research.
Research exceptions in copyright laws around the world
Data from: Multiple facets of biodiversity are threatened by mining-induced land-use change in the Brazilian Amazon
<p><strong>Aim</strong> </p> <p>Mining is increasingly pressuring areas of critical importance for biodiversity conservation, such as the Brazilian Amazon. Biodiversity data are limited in the tropics, restricting the scope for risks to be appropriately estimated before mineral licencing decisions are made. As the distributions and range sizes of other taxa differ markedly from those of vertebrates – the common proxy for analysis of risk to biodiversity from mining – whether mining threatens lesser-studied taxonomic groups differentially at a regional scale is unclear.</p> <p><strong>Location </strong></p> <p>Brazilian Amazon</p> <p><strong>Methods </strong></p> <p>We assess risks to several facets of biodiversity from industrial mining by comparing mining areas (within 70km of an active mining lease) and areas unaffected by mining, employing species richness, species endemism, phylogenetic diversity, and phylogenetic endemism metrics calculated for angiosperms, arthropods, and vertebrates.</p> <p><strong>Results </strong></p> <p>Mining areas contained higher densities of species occurrence records than the unaffected landscape, and we accounted for this sampling bias in our analyses. None of the four biodiversity metrics differed between mining and non-mining areas for vertebrates. For arthropods, species endemism was greater in mined areas. Mined areas also had greater angiosperm species richness, phylogenetic diversity, and phylogenetic endemism, although lower species endemism than unmined areas.</p> <p><strong>Main Conclusions </strong></p> <p>Unlike for vertebrates, facets of angiosperm and arthropod diversity are relatively higher in areas of mining activity, underscoring the need to consider multiple taxonomic groups and biodiversity facets when assessing risk and evaluating management options for mining threats. Particularly concerning is the proximity of mining to areas supporting deep evolutionary history, which may be impossible to recover or replace. As pressures to expand mining in the Amazon grow, impact assessments with broader taxonomic reach and metric focus will be vital to conserving biodiversity in mining regions.</p>
Acupuncture for ADHD: Acupoint Data Mining, Clinical Effectiveness, and Interviews to Explore Treatment Outcomes.
ClinicalTrials.gov study NCT06860763. IPD Sharing: NO. Countries: 1. Publications: 19.
Data and code from: Investigating the Yanomami malaria outbreak: Gold mining and malaria
Open the record for dataset details and reuse information.
Data from: Ecological resilience in a primate community affected by gold mining in Suriname
Open the record for dataset details and reuse information.
Data from: Do metal mines and their runoff affect plumage color? A regional scale study of streak-backed orioles in south-central Mexico
Open the record for dataset details and reuse information.
Data from: Why do youths initiate to smoke? A data mining analysis on tobacco advertising, peer, and family factors for Indonesian youths
Open the record for dataset details and reuse information.
Data from: Variations in root functional traits facilitate the adaptation of pioneer plants to rare earth element mine tailings
Open the record for dataset details and reuse information.
Characterizing and classifying neuroendocrine neoplasms through microRNA sequencing and data mining
Open the record for dataset details and reuse information.
Data from: Phenotypic plasticity accounts for changes in plant phosphorus-acquisition strategies from mining to scavenging along a gradient of soil phosphorus availability in South American Campos grasslands
Open the record for dataset details and reuse information.
Data from: Dark ophiuroid biodiversity in a prospective abyssal mine field
Open the record for dataset details and reuse information.
Data from: Abiotic legacies mediate plant-soil feedback during early vegetation succession on rare earth element mine tailings
Open the record for dataset details and reuse information.
Data from: Limited biomass recovery from gold mining in Amazonian forests
Open the record for dataset details and reuse information.
Data from: Feature sequence-based genome mining uncovers the hidden diversity of bacterial siderophore pathways
Open the record for dataset details and reuse information.
Data from: Freshwater ecological quality assessment of the gold mining Mashcon watershed, Cajamarca - Peru
Open the record for dataset details and reuse information.
Data from: Horses in the Cloud: big data exploration and mining of fossil and extant Equus (Mammalia: Equidae)
Open the record for dataset details and reuse information.
Data from: Multiple facets of biodiversity are threatened by mining-induced land-use change in the Brazilian Amazon
Open the record for dataset details and reuse information.
[Model outputs] Identifying major hydrologic change drivers in a highly managed transboundary endorheic basin: integrating hydro‐ecological models and time‐series data mining techniques
Open the record for dataset details and reuse information.
Data and material for: "Mining file histories: should we consider branches?"
<p>This repository is the online appendix of our paper:</p> <p>Vladimir Kovalenko, Fabio Palomba, and Alberto Bacchelli. 2018. Mining file histories: should we consider branches? In Proceedings of the 33rd ACM/IEEE International Conference on Automated Software Engineering (ASE 2018). Association for Computing Machinery, New York, NY, USA, 202–213. <a href="https://doi.org/10.1145/3238147.3238169">DOI</a></p> <p>The <code>results</code> folder contains full results per project for RQ1 and for the reviewer recommendation part of RQ2, as well as lists of projects used for the evaluation of defect prediction and change recommendation algorithms. The rest of the results are provided in the paper.</p> <p>The code that we have used to download and process the data (code reviews, repositories, and change recommendation) is located in the <code>processor</code> folder.</p> <p>The <code>processor</code> depends on <code>git2neo</code> -- a tool to load Git metadata to neo4j databases and retrieve the histories, which is described in Section III.B of the paper.<br> This release of <code>git2neo</code> is provided as is to facilitate the evaluation and reproducibility of our work.</p>
Data showcase papers published in the Mining Software Repositories (MSR) conference (v2.2)
<p>Data regarding data showcase papers published in the Mining Software Repositories (MSR) conference.</p> <p>The following data files are included.</p> <ul> <li>citing_dp_dois_citations.txt: Strong and weak citations of (strong and weak) citation papers</li> <li>data_paper_clustering.csv: The clustering process of MSR data papers</li> <li>data_paper_clusters.csv: Clusters of MSR data papers</li> <li>data_papers.bib: Bibliographic details of MSR data papers, along with their assigned clusters (field 'cluster') and strong citations (field 'usedby')</li> <li>dp_dois_citations.txt: Strong and weak citations of MSR data papers</li> <li>msr-all: Bibliographic details of all MSR (data and non-data) papers</li> <li>ndp_dois_citations.txt: Strong and weak citations of MSR non-data papers</li> <li>ndp_rand_dois_citations.txt: Strong and weak citations of a randomly chosen MSR non-data paper weighted sample</li> <li>self-citations.txt: Strong citations of MSR data papers by their authors</li> <li>strong_citation_classification.csv: The classification process of strong citation papers according to the SWEBOK knowledge areas</li> <li>strong_citation_fields.csv: SWEBOK knowledge areas of strong citation papers</li> <li>strong_citations.bib: Bibliographic details of strong citation papers</li> <li>survey_questionnaire.pdf: The final survey questionnaire</li> <li>survey_responses.csv: Anonymized responses of the final survey questionnaire (Email addresses have been excluded for privacy reasons.)</li> <li>weak_citations_notes.bib: Weak citations of MSR data papers and the use they make</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.