Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
54
datasets available to search
ShareScore release 0.9.0
Dataset results
54 results for “Apache”
Problem discovery and resolution activities in the Apache HTTP Server Project (March 2001- March 2013).
<p>This is a dynamic visualization of problem discovery and resolution activities observed in the in the development of the Apache HTTP Server Project during the period March 2001- March 2013. The nodes in the network represent problems (software bugs). Anthropomorphic icons represent participants (software developers). The network edges connect participants to problems. Numerical labels record the internal identification numbers or participants and problems. The visible clusters represent the software modules. The central node is the project core module. Participants move closer to problems that attract their attention. When a participant allocates attention to a problem, an edge emerges connecting the two. The edge is green when a participant opens a bug report (i.e., when he discovers a new problem), red when the participant closes the bug report (i.e., when she solves an existing problem), and yellow when any other action is recorded. Problems (white nodes) are green when they first appear. They turn red immediately before being closed, and are yellow when the corresponding bug report is being modified. The animation advances by 0.05 seconds every day of historical time.</p> <p>The animation is produced using the Gource server control visualization tool developed by Andrew Caldwell (<a href="https://gource.io/">https://gource.io/</a>)</p>
A Queueing Network Model for Performance Prediction of Apache Cassandra
<p>The dataset consists in several csv files containing Cassandra and ScyllaDB performance.</p> <p>The experiments are organized in folders. There are three main folders containing:<br> - Cassandra 4 nodes: the files related to the Cassandra experiments conducted on a cluster composed of four nodes.<br> - ScyllaDB 4 nodes: The files related to the ScyllaDB experiments conducted on a cluster composed of four nodes. <br> - Cassandra QUORUM variant: the simulation data where a different kind of QUORUM is implemented in Cassandra.<br> <br> "Cassandra 4 nodes" and "ScyllaDB 4 nodes" include some subfolders, each one containing the files of the Consistency Level applied for those experiments. Each experiment is composed by three files (data*.csv) with the data reported by Yahoo! Cloud System Benchmark (YCSB) in the end of the experiment execution. Each folder contains also a sim.csv file with the data gathered from the simulation of the model inside Java Modeling Tool.</p> <p>The data*.csv files are composed by:<br> -Number of threads or clients<br> -Overall Throughput<br> -Number of Read requests<br> -Overall Read Response Time<br> -95 percentile Read Response Time<br> -99 percentile Read Response Time<br> -99.9 percentile Read Response Time<br> <br> Differently, the sim.csv files are composed by:<br> -Number of threads or clients<br> -Overall Throughput<br> -Overall Read Response Time</p>
Rediscovery Datasets: Connecting Duplicate Reports of Apache, Eclipse, and KDE
<p>We present three defect rediscovery datasets mined from Bugzilla. The datasets capture data for three groups of open source software projects: Apache, Eclipse, and KDE. The datasets contain information about approximately 914 thousands of defect reports over a period of 18 years (1999-2017) to capture the inter-relationships among duplicate defects. </p> <p><strong>File Descriptions</strong></p> <ul> <li>apache.csv - Apache Defect Rediscovery dataset</li> <li>eclipse.csv - Eclipse Defect Rediscovery dataset</li> <li>kde.csv - KDE Defect Rediscovery dataset</li> </ul> <p> </p> <ul> <li>apache.relations.csv - Inter-relations of rediscovered defects of Apache</li> <li>eclipse.relations.csv - Inter-relations of rediscovered defects of Eclipse</li> <li>kde.relations.csv - Inter-relations of rediscovered defects of KDE</li> </ul> <p> </p> <ul> <li>create_and_populate_neo4j_objects.cypher - Populates Neo4j graphDB by importing all the data from the CSV files. Note that you have to set dbms.import.csv.legacy_quote_escaping configuration setting to false to load the CSV files as per https://neo4j.com/docs/operations-manual/current/reference/configuration-settings/#config_dbms.import.csv.legacy_quote_escaping</li> <li>create_and_populate_mysql_objects.sql - Populates MySQL RDBMS by importing all the data from the CSV files</li> <li>rediscovery_db_mysql.zip - For your convenience, we also provide full backup of the MySQL database</li> </ul> <p> </p> <ul> <li>neo4j_examples.txt - Sample Neo4j queries</li> <li>mysql_examples.txt - Sample MySQL queries</li> <li>rediscovery_eclipse_6325.png - Output of Neo4j example #1</li> </ul> <p> </p> <ul> <li>distinct_attrs.csv - Distinct values of bug_status, resolution, priority, severity for each project</li> </ul>
Apache Jira Issue Tracking Dataset
<p>This dataset contains the Jira Issue Tracking data of the Apache Software Foundation, enriched with topic modeling information using the BERTopic technique.</p> <p>You can use the dataset with the following steps:</p> <ol> <li>Set up a MongoDB instance.</li> <li>Download the data.</li> <li>Navigate to the download folder and use the mongorestore command (<a href="https://docs.mongodb.com/database-tools/mongorestore/" target="_blank" rel="noopener">https://docs.mongodb.com/database-tools/mongorestore/</a>) with the --gzip flag.</li> </ol> <p>Detailed instructions for setting up/reproducing/updating the dataset are also provided in the website <a href="https://authecesofteng.github.io/semantics-jira-dataset/" target="_blank" rel="noopener">https://authecesofteng.github.io/semantics-jira-dataset/</a> (relevant repos <a href="https://github.com/AuthEceSoftEng/jira-apache-downloader" target="_blank" rel="noopener">https://github.com/AuthEceSoftEng/jira-apache-downloader</a> and <a href="https://github.com/AuthEceSoftEng/jira-topic-extractor" target="_blank" rel="noopener">https://github.com/AuthEceSoftEng/jira-topic-extractor</a>)</p> <p><em>Note: you do not need to download the models .rar files if you do not use them, as the issue-topic distribution (along with probabilities) is already computed and stored in the topics Mongo collection.</em></p>
WiDS mortality dataset - APACHE diagnoses enriched
<p>The WiDS mortality dataset was modified, adding the APACHE diagnoses using the original column "apache_3j_diagnosis_code".</p> <p>This dataset is a merge from:</p> <ol> <li><strong>Mortality data</strong>: https://www.kaggle.com/competitions/widsdatathon2020/data</li> <li><strong>APACHE</strong>: https://www.kaggle.com/datasets/danofer/apache-iiij-icu-diagnosis-codes?select=icu-apache-Subdiagnosis-codes-ANZICS.csv</li> </ol>
Fig. 18. Slaterocoris apache, endosoma. A in Revision And Phylogenetic Analysis Of The North American Genus Slaterocoris Wagner With New Synonymy, The Description Of Five New Species And A New Genus From Mexico, And A Review Of The Genus Scalponotatus Kelton (Heteroptera: Miridae: Orthotylinae)
Fig. 18. Slaterocoris apache, endosoma. A. Roosevelt, UT. B. Eagar, AZ. C. Kimball Jct., UT. D, E. 12 mi W of Pueblo, CO.
Fig. 17. Slaterocoris apache, right paramere. A in Revision And Phylogenetic Analysis Of The North American Genus Slaterocoris Wagner With New Synonymy, The Description Of Five New Species And A New Genus From Mexico, And A Review Of The Genus Scalponotatus Kelton (Heteroptera: Miridae: Orthotylinae)
Fig. 17. Slaterocoris apache, right paramere. A. Roosevelt, UT. B. 29 mi SW of Norwood, CO. C. White Rock Overlook, NM. D. Berthoud Pass, CO. E. Taos, NM. F. 33 mi SW of Dulce, NM. G. Connors Pass, NV. H, I. Lehman Caves, NV.
Raw Data and PCA Filtering of Apache Point Observatory NMSU 1m StellaCam Observations of LCROSS
<p>This archive contains the raw data and data products from observations of the 2009-10-09 impact of the Lunar CRater Observation and Sensing Satellite (LCROSS) spacecraft on the Moon by the StellaCam instrument on the Apache Point Observatory NMSU 1m telescope.</p> <p>Full details about the raw data are available in Chanover, N. J. et al. Results from the NMSU-NASA Marshall Space Flight Center LCROSS observational campaign. <em>J. Geophys. Res. (Planets)</em> <strong>116</strong>, E08003 (2011). <a href="https://doi.org/10.1029/2010JE003761">https://doi.org/10.1029/2010JE003761</a></p> <p>We use principal component analysis (PCA) filtering both to coregister the raw time series and to effectively remove a static background signal that is spatially and temporally modified by atmospheric and instrumental effects. We iteratively remove principal components from the data through cumulative sequential elimination (CSE) resulting in a non-detection of the LCROSS ejecta plume signal.</p> <p>Full details are available in the published journal article:</p> <p>Strycker, Paul D., Nancy J. Chanover, Ruth L. Temme, Jonathan M. Schotte, Payton L. Mueller, and Emily L. Karls. 2023. "Time Series Analysis Methods and Detectability Factors for Ground-Based Imaging of the LCROSS Impact Plume" <em>Remote Sensing</em> <strong>15</strong>, no. 1: 37. <a href="https://doi.org/10.3390/rs15010037">https://doi.org/10.3390/rs15010037</a></p> <p>This work was supported by NASA’s Lunar Data Analysis Program through grant number NNX15AP92G.</p>
PCA Filtering of Apache Point Observatory 3.5m Agile Observations of LCROSS
<p>This archive contains data products from observations of the 2009-10-09 impact of the Lunar CRater Observation and Sensing Satellite (LCROSS) spacecraft on the Moon by the Agile instrument on the Apache Point Observatory 3.5m telescope. We use principal component analysis (PCA) filtering both to improve the coregistration of the raw time series and to effectively remove a static background signal that is spatially and temporally modified by atmospheric and instrumental effects. We iteratively remove principal components from the data through cumulative sequential elimination (CSE) to find a maximum signal-to-noise ratio of the LCROSS ejecta plume signal.</p> <p>Full details are available in the published journal article:</p> <p>Strycker, Paul D., Nancy J. Chanover, Ruth L. Temme, Jonathan M. Schotte, Payton L. Mueller, and Emily L. Karls. 2023. "Time Series Analysis Methods and Detectability Factors for Ground-Based Imaging of the LCROSS Impact Plume" <em>Remote Sensing</em> <strong>15</strong>, no. 1: 37. <a href="https://doi.org/10.3390/rs15010037">https://doi.org/10.3390/rs15010037</a></p> <p>This work was supported by NASA’s Lunar Data Analysis Program through grant number NNX15AP92G.</p>
SegInfoSoS 2.0 - Ontologia para Gestão da Segurança da Informação em Sistemas-de-Sistemas com Apache Jena
<p>As intensas transformações ocorridas na sociedade nesta década tornaram os sistemas de informação mais complexos. Tal complexidade se relaciona a uma categoria de sistemas definida como sistemas-de-sistemas (SoS). Embora o SoS ofereça benefícios às organizações, a dificuldade dos gestores de Tecnologia da Informação (TI) em lidar com a segurança da informação nesses sistemas pode deixá-los vulneráveis a ameaças e impactos causados por ataques cibernéticos. As ontologias podem ser utilizadas como solução para esse problema porque elas definem estruturas de conhecimento e promovem um entendimento<br> compartilhado de um domínio, tarefa ou aplicação. Nesse sentido, uma ontologia de domínio foi desenvolvida por meio de OWL API em Java para garantir a gestão do conhecimento em segurança em SoS e o seu entendimento compartilhado de modo que os stakeholders, gestores de TI e suas equipes possam evitar os riscos, as vulnerabilidades e as ameaças em SoS. Cabe também destacar a praticidade para a extração das informações da ontologia para utilização em aplicativos de celular, sites, smartwhatches, smartphones e eletrodomésticos da linha smarts, assistentes virtuais como Alexa (Amazon), Siri (Apple), Cortana<br> (Microsoft), ChatGPT, entre outros. Tal utilidade simula querys de Sistemas de Gerenciamento de Banco de Dados para o uso de pequenas e médias empresas. Neste livro é abordado o tema gestão da segurança da informação em SoS por meio de uma ontologia de domínio com foco na sua aplicabilidade para indústria.</p>
Litter Fall Collection Study in Pinyon-Juniper, Cottowood, and Spruce-Fir-Aspen Forests at the Sevilleta NWR, Bosque del Apache NWR, and the Cibola National Forest, New Mexico (1992-1993)
The litterfall study was designed to assess the quantity of biomass (leaves, twigs, reproductive materials) falling from tree species in different ecosystem types. Three study sites selected were:  (1) the pinyon-juniper woodland site near Cerro Montoso on the Sevilleta NWR; (2) the cottonwood forest LTER site along the Rio Grande at Bosque del Apache NWR; and (2) the old-growth spruce-fir-aspen site near South Baldy in the Magdalena Mountains (Cibola National Forest).  The study was conducted over two years (1992-1993) to compare litterfall rates and quantities among sites, seasons and years.
Apache Water Tus
Apache Water Tus at the Hutchings Museum Institute. 172 images, Canon EOS 80D, 35mm, F/16, ISO 100, Agisoft Metashape, Windows 10. Low Poly Source: Objaverse 1.0 / Sketchfab
apache squire
<p>Overview of Data</p> <p>Contains 3 tables in SQL format:</p> <p>**Table 1 (apache_datasources):** List of Apache Data sources<br> Attributes (Name, Type):<br> `datasource_id` Integer<br> `item_description` Character<br> `date_posted` Date<br> `date_collected` Date<br> `method` Character<br> `item_url` Character<br> `last_updated` Date</p> <p>**Table 2(apache_people_projects):** List of people committing various projects<br> Attributes (Name, Type):<br> `svn_id` Character<br> `real_name` Character<br> `web_site` Character<br> `datasource_id` Integer<br> `project_name` Character<br> `role_on_project` Character<br> `details` Character<br> `email` Character<br> `organization` Character<br> `last_updated` Date</p> <p>**Table 3(apache_unlisted_cla):** List of people with signed CLAs but are not committers<br> Attributes (Name, Type):<br> `real_name` Character<br> `datasource_id` Integer<br> `last_updated` Date</p>
Dataset for Characterizing Distributed Machine Learning Workloads on Apache Spark
<div> <div> <div> <ul> <li> <p><span>YasmineDjebrouni,IsabellyRocha,SaraBouchenak,LydiaChen,PascalFelber,Vania Marangozova, and Valerio Schiavoni. 2023. Characterizing Distributed Machine Learning Workloads on Apache Spark. In Proceedings of the 24th International Middleware Conference (Middleware ’23). Association for Computing Machinery, New York, NY, USA, 151–164. </span></p> </li> </ul> </div> </div> </div>
Case Study - Performance Changes of Apache Tomcat at Code Level
<p>This dataset provides the data of the case study on <a href="https://github.com/apache/tomcat">Apache Tomcat</a> that was conducted as part of the bachelor thesis "Extending Peass to Detect Performance Changes of Apache Tomcat". The data was generated by <a href="https://github.com/DaGeRe/peass">Peass</a> and the <a href="https://github.com/stro18/peass-ant">Peass-Ant plugin</a>.</p> <p><strong>Contents</strong></p> <p>The dataset is divided into case_study_1v-100v.tar.xz and case_study_100v-200v.tar.xz. Both parts contain the following data:</p> <ul> <li>select/ - Results of regression test selection, especially: <ul> <li>results/traceTestSelection_tomcat.json - Tests selected based on static code analysis and trace analysis</li> </ul> </li> <li>3 x measure_p*_100vm_400iter/ - Results of performance measurements, especially: <ul> <li>clean/ - Mean of measurements per VM</li> <li>results/changes.json - Tests containing performance changes, computed with t-test</li> </ul> </li> <li>rca_p3_100vm_400iter/ - Results of root cause analysis, especially: <ul> <li>rca/tree/ - Measurements results per called method</li> <li>results/*/*.html - Visualization of root cause analysis</li> </ul> </li> </ul> <p>Additionally, f1_score.tar.xz is included in this dataset, It contains heatmaps visualizing the F<sub>1</sub>-score that were used to compare different measurement configurations.</p> <p><strong>Extraction of Data</strong></p> <p>Move all .tar.xz files to an empty directory and execute for each file:</p> <pre><code class="language-bash">tar -Jxf <file></code></pre> <p> </p>
Figures 17-19. Laemophloeus apache, n in A review of New World Laemophloeus Dejean (Coleoptera: Laemophloeidae): 3. Nearctic species
Figures 17-19. Laemophloeus apache, n. sp. 17) Head and pronotum. 18) Aedeagus. 19) Parameres.
Figure 8. Laemophloeus apache, n in A review of New World Laemophloeus Dejean (Coleoptera: Laemophloeidae): 3. Nearctic species
Figure 8. Laemophloeus apache, n. sp., male habitus.
Map 3. Localities for Slaterocoris apache and S in Revision And Phylogenetic Analysis Of The North American Genus Slaterocoris Wagner With New Synonymy, The Description Of Five New Species And A New Genus From Mexico, And A Review Of The Genus Scalponotatus Kelton (Heteroptera: Miridae: Orthotylinae)
Map 3. Localities for Slaterocoris apache and S. croceipes.
Apache Point Observatory Lunar Laser-ranging Operation (APOLLO) normal point data: 2006 through 2020
<p>Normal point data from the Apache Point Observatory Lunar Laser-ranging Operation (APOLLO) covering the 15-year span from April 2006 through the end of 2020. APOLLO measures the earth-moon separation by recording the round-trip travel time of photons from the Apache Point Observatory to five retro-reflector arrays on the moon. The APOLLO data set, combined with the 50-year archive of measurements from other lunar laser ranging (LLR) stations, can be used to probe fundamental physics such as gravity and Lorentz symmetry, as well as properties of the moon itself. These normal points have a median nightly accuracy of 1.7mm, which is an order of magnitude better than other LLR stations. Data is provided in both the "JPL" format (.np) and the Consolidated Laser Ranging Data Format (.crd)</p>
Comparative Analysis of APACHE II and P-POSSUM
ClinicalTrials.gov study NCT02471612. IPD Sharing: NO. Countries: 1. Publications: 1.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.