Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
12
datasets available to search
ShareScore release 0.9.0
Dataset results
12 results for “web mining”
Unveiling Web Fingerprinting in the Wild Via Code Mining and Machine Learning
<p>Dataset of Javascripts used for training and testing the fingerprinting algorithms described in </p> <p>Rizzo, Valentino, Stefano Traverso, and Marco Mellia. "Unveiling Web Fingerprinting in the Wild Via Code Mining and Machine Learning." <em>Proceedings on Privacy Enhancing Technologies</em> 2021.1 (2021): 43-63.</p>
Data supplementing the conference paper "Who you gonna call? Analyzing web requests in Android applications", 14th International Conference on Mining Software Repositories 2017.
<p>This repository contains the data supplementing the paper:</p> <p>M. Rapoport, P. Suter, E. Wittern, O. Lhótak, J. Dolby, "Who you gonna call? Analyzing web requests in Android applications", MSR 2017.</p> <p>A detailed description of the data is included in the archive in README.md.</p>
WEB MINING TECHNIQUE: IMPLEMENTATION OF STUDENTS BEHAVIOR IN WEB SERVICES
<p>Web mining makes benefit of data mining techniques. It deals with empathetic<br>behavior of students by making use of files. When students interact with the web, the<br>interacted information’s of the user are stored in a special type of log files. User Name,<br>Time Stamp, IP Address, Access Request, variety of Bytes Transferred, URL that<br>Referred, Result Status and User Agent are the information considers in the service files.<br>The service files are managed by the web servers. Our proposal focuses on the use of web<br>mining techniques to classify web pages and web site type according to students visit.<br>Keywords- Student’s behavior, service files, web mining, clustering, classification</p>
Results of Salinicola genome mining using the antiSMASH web tool.
<p>A collection of GenBank files (.gbk) and html files obtained from 14 <em>Salinicola</em> genomes.<br>Biosynthetic Gene Clusters searched and detected using antiSMASH v7.1.</p> <p>List of strains:</p> <p><em>S. tamaricis</em> CSA4-1, <em>S. tamaricis</em> F01, <em>S. endophyticus</em> CPA92, <em>S. aestuarinus</em> CP62, <em>S. halophilus</em> CECT 5903,<em> S. acroporae</em> LMG 28587, <em>S. lusitanus</em> CR50, <em>S. corii </em>L3, <em>S. salarius</em> DSM 18044, <em>S. socius</em> DSM 19940, <em>S. halimionae</em> CPA60, <em>S. halophyticus</em> CR45, <em>S. peritrichatus </em>JCM 18795 and <em>S. rhizosphaerae </em>KCTC 32998.</p> <pre>-. .-. .-. .-. .-. .-. . ||\|||\ /|||\|||\ /|||\|||\ /| |/ \|||\|||/ \|||\|||/ \|||\|| ~ `-~ `-` `-~ `-` `-~ `- </pre> <p><br> </p>
Mining the UK Web Archive for Semantic Change Detection (Dataset)
<p>The dataset that was used and released with the RANLP 2019 paper, titled "Mining the UK Web Archive for Semantic Change Detection" (see <a href="https://github.com/adtsakal/Semantic_Change">https://github.com/adtsakal/Semantic_Change</a>). It contains annual word2vec representations of more than 47K words over the period 2000-2013, along with a list of 65 words with known semantic change over the same time period. </p>
Supplementary web page for the paper "SEAL: Integrating Program Analysis and Repository Mining"
<p>This is an archive of the supplementary material for the paper “SEAL: Integrating Program Analysis and Repository Mining” including the website and dataset. The website can also be viewed here: <a href="https://se-sic.github.io/paper-SEAL/">https://se-sic.github.io/paper-SEAL/</a></p>
Changes in stream food web structure across a gradient of acid mine drainage increases local community stability
<p>Understanding what makes food webs stable has long been a goal of ecologists. Topological structure and the distribution and magnitude of interaction strengths in food webs have been shown to confer important stabilizing properties. However, our understanding of how variable species interactions affect food web structure and stability is still in its infancy. Anthropogenic stress, such as acid mine drainage, is likely to place severe limitations on the food web structures possible due to changes in community composition and body mass distributions. Here, we used mechanistic models to infer food web structure and quantify stability in streams across a gradient of acid mine drainage. Multiple food webs were iterated for each community based on species pairwise interaction probabilities, in order to incorporate the variability of realistic food web structure. We found that food web structure was altered systematically with a 32-fold decrease in the number of links and a 2-fold increase in connectance across the gradient. Stability generally increased 6-fold with increasing acid mine drainage stress, regardless of how interaction strengths were estimated. However, the distribution of the stability measure, s, for some impacted communities separated into clusters of higher and lower magnitude depending on how interaction strengths were estimated. Management and restoration of impacted sites needs to consider their increased stability, as this may have important implications for the re-colonization of desirable species. Furthermore, active species introductions may be required to overcome the internal ecological inertia of affected communities.</p>
Katharina Schmid, WG3: Link mining from web archives
<p>Katharina Schmid, WG3: Link mining from web archives</p> <p>WARCnet Luxembourg meeting 2020</p>
Changes in stream food web structure across a gradient of acid mine drainage increases local community stability
Open the record for dataset details and reuse information.
Social Web Mining for Suicide Prevention
ClinicalTrials.gov study NCT04052477. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.
Client-side Web Mining for Community Formation in Peer-to-Peer Environments
In this paper we present a framework for forming interests-based Peer-to-Peer communities using client-side web browsing history. At the heart of this framework is the use of an order statistics-based approach to build communities with hierarchical structure. We have also carefully considered privacy concerns of the peers and adopted cryptographic protocols to measure similarity between them without disclosing their personal profiles. We evaluated our framework on a distributed data mining platform we have developed. The experimental results show that our framework could effectively build interests-based communities.
Datasets for publication "Post-mining effects on fish communities and food web dynamics"
<p>Raw data and R scripts of all analyses of the paper.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.