Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
318
datasets available to search
ShareScore release 0.7.1
Dataset results
318 results for “Data mining”
FIGURES 19–26 in New species of leaf-mining Nepticulidae (Lepidoptera) from the Neotropical and Ando-Patagonian regions, with new data on host plants
FIGURES 19–26. Stigmella decora Diškus & Stonis, sp. nov. 19, 21, host plant Rhynchotheca spinosa Ruiz & Pav.; 20–24, leaf-mines; 25, cocoon; 26, adult, holotype.
FIGURES 14–18 in New species of leaf-mining Nepticulidae (Lepidoptera) from the Neotropical and Ando-Patagonian regions, with new data on host plants
FIGURES 14–18. Female genitalia of Stigmella apicibrunella Diškus & Stonis, sp. nov., genitalia slide no. AD813 (ZMUC). 14, apophyses; 15, ductus spermathecae; 16, slerite of ductus spermathecae; 17, false signum; 18, corpus bursae.
FIGURES 60–67 in New species of leaf-mining Nepticulidae (Lepidoptera) from the Neotropical and Ando-Patagonian regions, with new data on host plants
FIGURES 60–67. Stigmella patula Remeikis & Stonis, sp. nov., holotype. 60–62, male adult; 63, 64, genitalia slide no. RA282, uncus and gnathos; 65, same, phallus; 66–67, same, capsule with phallus removed (ZMUC).
FIGURES 89–94 in New species of leaf-mining Nepticulidae (Lepidoptera) from the Neotropical and Ando-Patagonian regions, with new data on host plants
FIGURES 89–94. Stigmella monstrata Remeikis & Stonis, sp. nov. 89, male holotype; 90, male genitalia, capsule, slide no. RA283, paratype; 91, phallus, slide no. RA284, holotype; 92, same, slide no. RA283, paratype; 93, capsule, slide no. RA284, holotype; 94, female genitalia, slide no. RA307 (ZMUC).
FIGURES 8–13 in New species of leaf-mining Nepticulidae (Lepidoptera) from the Neotropical and Ando-Patagonian regions, with new data on host plants
FIGURES 8–13. Male genitalia of Stigmella apicibrunella Diškus & Stonis, sp. nov. 8–11, capsule with phallus removed, holotype, genitalia slide no. AD800; 12, phallus, paratype, slide no. AD772; 13, same, holotype, slide no. AD800 (ZMUC).
FIGURES 2–7 in New species of leaf-mining Nepticulidae (Lepidoptera) from the Neotropical and Ando-Patagonian regions, with new data on host plants
FIGURES 2–7. Stigmella apicibrunella Diškus & Stonis, sp. nov. 2—habitat, western slopes of the equatorial Andes (Ecuador: Chimborazo Province) at altitudes about 1800 m; 3—host plant Acalypha padifolia Kunth (Euphorbiaceae); 4, leafmines; 5–7, female paratype adults.
FIGURES 112–115 in New species of leaf-mining Nepticulidae (Lepidoptera) from the Neotropical and Ando-Patagonian regions, with new data on host plants
FIGURES 112–115. Stigmella gallicola van Nieukerken & Nishida, 2016. 112, male adult; 113, 114, male genitalia, slide no. RA617, capsule with phallus removed; 115, same, phallus (USNM, deposited with a label as holotype of "eunigra"; see Remarks on the species).
FIGURES 31–35 in New species of leaf-mining Nepticulidae (Lepidoptera) from the Neotropical and Ando-Patagonian regions, with new data on host plants
FIGURES 31–35. Stigmella decora Diškus & Stonis, sp. nov. 31, 32, male genitalia, paratype, genitalia slide no. AD801; 33, dorsal processes, slide no. AD784 (not type series); 34, paratype, slide no. AD801; 35, female genitalia, holotype, slide no. AD795 (ZMUC).
Data from: Mining the in-use stock of energy-transition materials for closed-loop e-mobility
<p>Material flow analysis dataset for energy-transition materials developed within the Spoke11 - CNMS MOST - WP2:Design for Sustainability</p>
Data from: Phylogeny of gracillariid leaf-mining moths: evolution of larval behaviour inferred from phylogenomic and Sanger data
<p>Gracillariidae is the most taxonomically diverse cosmopolitan leaf-mining moth family, consisting of nearly 2000 named species in 105 described genera, classified into eight extant subfamilies. The majority of gracillariid species are internal plant feeders as larvae, creating mines and galls in plant tissue. Despite their diversity and ecological adaptations, their phylogenetic relationships, especially at the subfamily level, remain largely uncertain. Genomic data (83 taxa and 589 loci) were integrated with Sanger data (130 taxa and 22 loci), to reconstruct a phylogeny of Gracillariidae. Based on analyses of both data sets combined and analyzed separately, the monophyly of Gracillariidae and all its subfamilies, and the monophyly of the clade 'LAMPO' (subfamilies: Lithocolletinae, Acrocercopinae, Marmarinae, Phyllocnistinae, and Oecophyllembiinae) and relationships of its subclade 'AMO' (subfamilies: Acrocercopinae, Marmarinae, and Oecophyllembiinae) were strongly supported. A sister group relationship of Ornixolinae to the remainder of the family, and a monophyletic leaf roller lineage (<i>Callicercops</i> Vári + Parornichinae) + Gracillariinae, as sister to the 'LAMPO' clade were supported by the best hypotheses. Based on these results, a new subfamily, Callicercopinae Li, Ohshima et Kawahara, is established to accommodate the enigmatic genus <i>Callicercops</i>. Dating analyses indicate a mid-Cretaceous (105.3 Ma) origin of the family, followed by a rapid diversification into the nine subfamilies predating the K-Pg extinction. We hypothesize that advanced larval behaviours, such as making keeled or tentiform blotch mines, rolling leaves, and making galls, accelerated the diversification of Gracillariidae by avoiding larval parasitoids.</p>
[Research Data] Mining Relevant Solutions for Programming Tasks from Search Engine Results
<p>[Abstract]</p> <p>Software development is a knowledge-intensive activity. Official documentation for developers may not be sufficient for all developer needs. Searching for information on the Internet is a usual practice, but finding really useful information may be challenging, because the best solutions are not always among the first ranked pages. So, developers have to read and discard irrelevant pages, that is, pages that do not have code examples or that have content with little focus on the desired solution. This work aims at proposing an approach to mine relevant solutions for programming tasks from search engine results that remove irrelevant pages. The approach works as follows: a query related to the programming task is prepared, and given as an input to a search engine. The returned pages pass through an automatic filter to select relevant pages. We evaluated the top-20 pages returned by the Google search engine, for 10 different queries, and observed that only 31\% of the evaluated pages are relevant to developers. Then, we proposed and evaluated three different approaches to mine the relevant pages returned by the search engine. Google’s search engine has been used as a baseline, and our results have shown that Google’s search engine returns a reasonable number of irrelevant pages for developers, and we could find an effective approach to remove irrelevant pages, suggesting that developers could benefit from a customized web search filter for development content.</p> <p>[Contents of Research Data.rar file]</p> <p>The Research Data.rar file has a folder called Research Data that contains 3 folders internally, with the names: “01 – Source Code”, “02 - Data” and “03 – Preprocessing rules”. The folder “01 – Source Code” contains the JAVA source code of the implementations of the proposed approaches. The folder “02 - Data” contains the data of the evaluations carried out in the work, which are in the folders “01 - Evaluation results of pages returned by Google” and “02 - Results of approaches comparisons”. The folder “01 - Evaluation results of pages returned by Google” has the evaluations carried out on the first 20 pages returned by Google, following the criteria defined in the work, for the 10 queries considered in the evaluation. The folder “02 - Results of approaches comparisons” contains the results of the evaluation of the proposed approaches, for the 10 queries considered in the evaluation. In this evaluation, the number of pages given as input for the approaches was increased from 3 to 20 pages, for each number of pages a folder was generated with the results. In addition to the results of the Precision, Recall and F-Measure metrics that are in the file named Results Approaches.txt, other files were generated for analysis. For example, the Instances_without_outliers.txt file shows which pages were filtered out after applying the outlier page removal filter. The Selected Pages Approach 4.txt file, on the other hand, shows which pages were filtered after applying the filters of the GORCUO approach. The folder “03 - Preprocessing rules” has a file called Rules.java. In this file, there is the commented JAVA source code, from the implementation of the rules created in the pre-processing stage of the proposed approach.</p>
Raw data for the samples collected from Hole B of the ICDP DSeis project at the Moab Khotsong Gold Mine in South Africa
<p>This data set is corresponding to the Miyamoto, T. et al. "Characteristics of Seismogenic Fault Rock Related to the 2014 Orkney Earthquake (M5.5) Beneath the Moab Khotsong Gold Mine, South Africa" Geophysical Research Letters, 2022. This data set shows physical property, magnetic susceptibility, mineral assemblage, and element composition of all Hole B subsamples, and frictional properties of 5 Hole B subsamples.</p>
Would you manage a vibrant data mine?
<p>The cryptocurrencies have revolutionized financial data management. Due to this situation, now we have to think: How will we manage all these interactions? How will we preserve digital information? and How will we reduce the environmental footprint of these systems?<br> This new decentralized informational architecture is putting many banking entities in check. It allows, without intermediaries, data transfers at a global level. Which is why this intercontinental database is gaining more and more followers. For this reason, professionals are required capable of managing the frenzied number of algorithms and in turn ensuring the security of the entire informational chain.</p>
Dataset for "Data-Mining of In-Situ TEM Experiments: Towards Understanding Nanoscale Fracture"
<p>Dataset accompanying the publication "Data-Mining of In-Situ TEM Experiments: Towards Understanding Nanoscale Fracture"</p>
The first comprehensive revision of all the species attributed to Melomys led J. I. Menzies in 1996 to resurrect the genus Paramelomys and to redefine its morphologicallimits and species content. Menzies created P. gressitti as a new species belonging to a group displaying morphological similarities and including also P. lorentzii and P. moncktoni. Monotypic Distribution. E New Guinea. Descriptive notes. Head-body 135-162 mm, hindfoot 30-34 mm; no specific data are available for body weight. Gressitt's Mosaic-tailed Rat is a medium-sized Paramelomys with a soft, thick and woolly pelage, a long narrow foot, and a tail with three hairs per scale. It exhibits a medium-sepia dorsal pelage and a gray-buff ventral one. Tail is slightly shorter (99%) than head-body length. The skull has a narrow zygomatic plate. Habitat. Moist tropical mountain forest between 2300 m and 2400 m. Food and Feeding. No information. Breeding. No information. Activity patterns. Gressitt's Mosaic-tailed Rat is terrestrial. Movements, Home range and Social organization. No information. Status and Conservation. Classified as Endangered on The IUCN Red List owing to its small geographic range (less than 3500 km?*) and the destruction ofits habitat by mining and logging activities. The major threat to Gressitt's Mosaic-tailed Rat is ongoing habitat degradation caused by nearby human populations; habitat on Mount Kandy has been destroyed by gold-miners and wood-cutters. Bibliography. Menzies (1996). in Muridae
The first comprehensive revision of all the species attributed to Melomys led J. I. Menzies in 1996 to resurrect the genus Paramelomys and to redefine its morphologicallimits and species content. Menzies created P. gressitti as a new species belonging to a group displaying morphological similarities and including also P. lorentzii and P. moncktoni. Monotypic Distribution. E New Guinea. Descriptive notes. Head-body 135-162 mm, hindfoot 30-34 mm; no specific data are available for body weight. Gressitt's Mosaic-tailed Rat is a medium-sized Paramelomys with a soft, thick and woolly pelage, a long narrow foot, and a tail with three hairs per scale. It exhibits a medium-sepia dorsal pelage and a gray-buff ventral one. Tail is slightly shorter (99%) than head-body length. The skull has a narrow zygomatic plate. Habitat. Moist tropical mountain forest between 2300 m and 2400 m. Food and Feeding. No information. Breeding. No information. Activity patterns. Gressitt's Mosaic-tailed Rat is terrestrial. Movements, Home range and Social organization. No information. Status and Conservation. Classified as Endangered on The IUCN Red List owing to its small geographic range (less than 3500 km?*) and the destruction ofits habitat by mining and logging activities. The major threat to Gressitt's Mosaic-tailed Rat is ongoing habitat degradation caused by nearby human populations; habitat on Mount Kandy has been destroyed by gold-miners and wood-cutters. Bibliography. Menzies (1996).
Automatic ESG Assessment of Companies by Mining and Evaluating Media Coverage Data: NLP Approach and Tool
<p><strong>Replication package for our paper </strong><a href="https://arxiv.org/abs/2212.06540">Automatic ESG Assessment of Companies by Mining and Evaluating Media Coverage Data: NLP Approach and Tool</a></p> <p>It contains the following files:</p> <ol> <li>Our data set of 432,411 news headlines annotated as being environmental-, governance-, or social-related. We encourage fellow researchers to use the corpus as a benchmark for other ESG-relevant NLP tasks.</li> <li>Code/Notebooks that we used for the training and evaluation of our company detection, ESG classification, and sentiment models</li> <li>Full tables detailing the results of all experiments performed </li> </ol>
Learning about Learning: Mining Human Brain Sub-Network Biomarkers from fMRI Data
<p>These are coherence matrices originally described in: Dynamic reconfiguration of human brain networks during learning. Bassett DS, Wymbs NF, Porter MA, Mucha PJ, Carlson JM, Grafton ST. Proc Natl Acad Sci U S A. 2011 May 3;108(18):7641-6. doi:10.1073/pnas.1018985108. Epub 2011 Apr 18. Later the matrices were also studied in Cohesive network reconfiguration accompanies extended training. Telesford QK, Ashourvan A, Wymbs NF, Grafton ST, Vettel JM, BassettDS. Hum Brain Mapp. 2017 Sep;38(9):4744-4759. doi: 10.1002/hbm.23699. Epub 2017 Jun 24.</p>
Data for "Improving semantic video retrieval models by training with a relevance-aware online mining strategy"
<p>This repository contains all the data available for the publication:</p> <p><a href="https://doi.org/10.1016/j.cviu.2024.104035">Alex Falcon, Giuseppe Serra, and Oswald Lanz. <em>Improving semantic video retrieval models by training with a relevance-aware online mining strategy</em>. <strong>Computer Vision and Image Understanding</strong>. 2024.</a></p> <p>Code is available at: <a href="https://github.com/aranciokov/ranp/">https://github.com/aranciokov/ranp/</a></p> <p>The data includes:</p> <ul> <li>pre-extracted features (ordered_feature_*.zip files)</li> <li>annotations, such as pre-extracted semantic graphs, glove checkpoints, class annotations, etc (annotations_*.zip files)</li> <li>train/val/test, when available, split information (public_split_*.zip) files</li> <li>pretrained models for HGR and EAO (details in the github repo)</li> </ul>
Appendix to "Process Mining Pipelines with Controlled Sharing of Data and Algorithms"
<p><strong>Abstract: </strong>Process mining leverages execution traces within an organisation's IT systems to gain insights into its processes. Despite being a mature discipline in academia and industry, setting up process mining pipelines is still a complex task and involves programming, manual steps, and considerations of privacy and intellectual property.</p> <p>This paper introduces a platform based on a distributed architecture that helps define, deploy, and execute process mining pipelines across organisations. The requirements for this distributed architecture and platform are derived from a set of process mining scenarios, whose relevance is validated through a survey.</p> <p>Furthermore, this paper introduces a prototype for an initial version of the platform, demonstrating feasibility and supporting the specified requirements. This development is a major step in advancing process mining, offering simpler and more efficient ways of implementing and managing complex process mining pipelines on a larger scale.</p> <p><strong>Description: </strong>This dataset presents the support for non-functional requirements identified in the paper "Process Mining Pipelines with Controlled Sharing of Data and Algorithms" by existing process mining platforms.</p> <p><strong>Legend:</strong> Green cells indicate complete fulfilment. Yellow indicates partial fulfilment. Blue cells indicate uncertain fulfilment. Red indicates no fulfilmnet.</p>
Raw data article "Surviving adversity: exploring the presence of Lunularia cruciata (L.) Dum. on metal-polluted mining waste"
<p>This is the raw data of the article "Surviving adversity: exploring the presence of Lunularia cruciata (L.) Dum. on metal-polluted mining waste".</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.