Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
257
datasets available to search
ShareScore release 0.7.1
Dataset results
257 results for “data package”
Data package for NutNet project: Compositional variation in grassland plant communities (60 sites, 2007-2020)
Data associated with a manuscript examining compositional variation in grassland plant communities around the globe. We used a globally distributed experiment to examine variation in species composition within 60 grasslands on 6 continents. Each site had an identical experimental and sampling design: 24 plots x 4 years. We expressed compositional variation within each site—not across sites—using abundance- and incidence-based metrics of the magnitude of dissimilarity (Bray-Curtis and Sorensen, respectively), abundance- and incidence-based measures of the relative importance of replacement (balanced variation and species turnover, respectively), and species richness at two scales (per plot-year (alpha) and per site (gamma)). We assessed species composition separately for each site and then compared patterns among sites, asking: (1) How does small-scale compositional variation differ among grasslands?; (2) Can compositional variation within a site be predicted by its biotic and abiotic context?; and (3) Does a combination of metrics enhance our understanding of the ecological processes at individual sites? This data package includes the site-specific values of each metric and the explanatory variables used to predict differences in compositional variation among sites. Scripts to conduct analyses and create figures and tables are also provided.
Data package for opportunistic constant target matchup study
<p>Opportunistic constant target matching is a new method for satellite<br> intercalibration.</p> <p>It is complementary to the traditional simultaneous nadir overpass<br> (SNO) method because it can provide warm matchups in cases where the<br> SNO method provides only cold matchups.</p> <p>A geostationary infrared sensor (SEVIRI) is used to select constant<br> target matches for two different microwave sensors (NOAA 18 and Metop<br> A). This is the data package for a publication where we discuss the<br> main assumptions and limitations of the new method and explore its<br> statistical properties with a simple Monte Carlo simulation and with<br> real observations from NOAA 18 and Metop A.</p>
Data package from "Regional Mapping and Spatial Distribution Analysis of Canopy Palms in an Amazon Forest Using Deep Learning and VHR Images"
<p>This data package contains the very high resolution maps of canopy palms from the paper "Regional Mapping and Spatial Distribution Analysis of Canopy Palms in an Amazon Forest Using Deep Learning and VHR Images". These maps have been produced with two GeoEye-1 very high resolution images (0.5 m) and a Deep Learning method for image segmentation called U-net, methods and data are fully described in the article. The total size of the decompressed archive is 2.56 Go and is distributed in two shapefiles, one for each GeoEye-1 image. When using this dataset, please cite the original article https://doi.org/10.3390/rs12142225</p>
Replication Package for the Paper: "Will Data Influence the Experiment Results?: A Replication Study of Automatic Identification of Decisions"
<p>This is the replication package for the paper: "Will Data Influence the Experiment Results?: A Replication Study of Automatic Identification of Decisions". It contains the source code and dataset of our experiment for the replication by other researchers. In the meanwhile, we provide brief description of the files in the replication package below.</p> <p><strong>1. main_code folder</strong></p> <ul> <li><em>automatic_approach.py </em>contains the main source code of the automatic approach for identifying decisions in our experiment, which is conducted on MacOs and Python 3.7.9. <strong>Note that you may get slightly</strong> <strong>different experiment results when conducting the experiments on different environment configurations.</strong></li> <li><em>requirement.txt</em> records all the installation packages and their version numbers needed for the current program to run. You can use "<em>pip install -r requirement.txt</em>" to rebuild the project and install all dependencies. <strong>Note that you may get slightly different experiment results when using different packages or versions. </strong></li> </ul> <p><strong>2. dataset folder</strong></p> <ul> <li><em>EASE2020 - 650 decisions.xlsx </em>contains 650 decision sentences from our previous work (EASE2020)</li> <li><em>EASE2020 - 650 non-decisions.xlsx </em>contains 650 non-decision sentences from our previous work (EASE2020)</li> <li><em>Our 844 relabeled decisions.xlsx</em> contains 844 relabeled decisions in this work.</li> <li><em>Our 750 assumptions.xlsx</em> contains 750 assumptions from our previous work (APSEC2019)</li> </ul> <p><strong>3. RQ1 folder</strong></p> <ul> <li><em>experiment_RQ1.py</em> contains the main source code of the experiment for answering RQ1, which is conducted on the same environment configuration as the <em>automatic_approach.py.</em></li> </ul> <p><strong>4. RQ2 folder</strong></p> <ul> <li><em>experiment_RQ2.py</em> contains the main source code of the experiment for answering RQ2, which is conducted on the same environment configuration as the <em>automatic_approach.py.</em></li> </ul> <p><strong>5. RQ3 folder</strong></p> <ul> <li><em>experiment_RQ3.py</em> contains the main source code of the experiment for answering RQ3, which is conducted on the same environment configuration as the <em>automatic_approach.py.</em></li> </ul>
Data from: nlstimedist: an R package for the biologically meaningful quantification of unimodal phenology distributions
Phenological investigation can provide valuable insights into the ecological effects of climate change. Appropriate modelling of the time distribution of phenological events is key to determining the nature of any changes, as well as the driving mechanisms behind those changes. Here we present the nlstimedist R package, a distribution function and modelling framework that describes the temporal dynamics of unimodal phenological events. The distribution function is derived from first principles and generates three biologically interpretable parameters. Using seed germination at different temperatures as an example, we show how the influence of environmental factors on a phenological process can be determined from the quantitative model parameters. The value of this model is its ability to represent various unimodal temporal processes statistically. The three intuitively meaningful parameters of the model can make useful comparisons between different time periods, geographical locations or species' populations, in turn allowing exploration of possible causes.
Data from: An R package and online resource for macroevolutionary studies using the ray-finned fish tree of life
1. Comprehensive, time-scaled phylogenies provide a critical resource for many questions in ecology, evolution, and biodiversity. Methodological advances have increased the breadth of taxonomic coverage in phylogenetic data; however, accessing and reusing these data remain challenging. 2. We introduce the Fish Tree of Life website and associated R package fishtree to provide convenient access to sequences, phylogenies, fossil calibrations, and diversification rate estimates for the most diverse group of vertebrate organisms, the ray-finned fishes. The Fish Tree of Life website presents subsets and visual summaries of phylogenetic and comparative data, and is complemented by the R package, which provides flexible programmatic access to the same underlying data source for advanced users wishing to extend or reanalyze the data. 3. We demonstrate functionality with an overview of the website, and show three examples of advanced usage through the R package. First, we test for the presence of long branch attraction artifacts across the fish tree of life. The second example examines the effects of habitat on diversification rate in the pufferfishes. The final example demonstrates how a community phylogenetic analysis could be conducted with the package. 4. This resource makes a large comparative vertebrate dataset easily accessible via the website, while the R package enables the rapid reuse and reproducibility of research results via its ability to easily integrate with other R packages and software for molecular biology and comparative methods.
GenomicInteractions: an R/Bioconductor package for manipulating and investigating chromatin interaction data.
<p>Files required to regenerate figures in main text of paper, using existing Supplemental R Markdown files. </p>
Data providers package for reporting monitoring results for veterinary medicinal product residues (2017 test phase)
<p>This data providers package provides the data collection configuration and supporting materials for reporting veterinary medicinal product residues (VMPR) results according to Council Directive 96/23/EC of 29 April 1996 on measures to monitor certain substances and residues thereof in live animals and animal products and repealing Directives 85/358/EEC and 86/469/EEC and Decisions 89/187/EEC and 91/664/EEC. These are to be used for the 2017 data reporting test phase.</p> <p>The package includes;</p> <p>Advice on the values to be reported for the mandatory fields specified for this data collection</p> <p>The Standard Sample Description Version 2 XML schema definition for VMPR reporting</p> <p>The STX transformation file which automatically assigns sampEventId and sampAnId when this information is not provided</p> <p>The general and VMPR specific business rules applied for the automatic validation of the submitted datasets</p> <p>The VMPR specific terminologies to be used for reporting analytical methods and the residues included in the scope of the analytical methods</p> <p>An excel tool which can support users in creating the required XML file for submission where automated data collation tools are not available and a manual for the tool</p> <p> </p>
GeoLocator Data Package: South African Woodland Kingfisher
<p>This repository contains the raw data and the trajectory information generated with the GeoPressureR workflow, following the <a href="https://raphaelnussbaumer.com/GeoLocator-DP/">GeoLocator Data Package standard</a>. The more complete code used to generate this datapackage can be found on the Github Repository <a href="https://github.com/Rafnuss/WoodlandKingfisher">Rafnuss/WoodlandKingfisher</a>. </p> <p>It contains 5 tags equipped on Woodland Kingfisher in Mogalakwena, Limpopo, South Africa between 2017-2020.</p> <p> </p>
Data associated with: Recording animal-view videos of the natural world using a novel camera system and software package
<p>Data associated with Vasas V, Lowell MC*, Villa J*, Jamison QD*, Siegle AG*, Katta PVR*, Bhagavathula P*, Kevan PG, Fulton D, Losin N, Kepplinger D, Salehian S, Forkner RE, Hanley D (2023) Recording animal-view videos of the natural world using a novel camera system and software package. PLoS Biology. DOI: 10.1371/journal.pbio.3002444</p>
Replication Package of Deep Learning and Data Augmentation for Detecting Self-Admitted Technical Debt
<p>Self-Admitted Technical Debt (SATD) refers to circumstances where developers use code comments, issues, pull requests, or other textual artifacts to explain why the existing implementation is not optimal. Past research in detecting SATD has focused on either identifying SATD (classifying SATD instances as SATD or not) or categorizing SATD (labeling instances as SATD that pertain to requirements, design, code, test, etc.). However, the performance of such approaches remains suboptimal, particularly when dealing with specific types of SATD, such as test and requirement debt. This is mostly because the used datasets are extremely imbalanced.</p> <p>In this study, we utilize a data augmentation strategy to address the problem of imbalanced data. We also employ a two-step approach to identify and categorize SATD on various datasets derived from different artifacts. Based on earlier research, a deep learning architecture called BiLSTM is utilized for the binary identification of SATD. The BERT architecture is then utilized to categorize different types of SATD. We provide the dataset of balanced classes as a contribution for future SATD researchers, and we also show that the performance of SATD identification and categorization using deep learning and our two-step approach is significantly better than baseline approaches.</p> <p>Therefore, to showcase the effectiveness of our approach, we compared it against several existing approaches:</p> <ol> <li>Natural Language Processing (NLP) and Matches task Annotation Tags (MAT) [<a href="https://github.com/Naplues/MAT" target="_blank" rel="noopener">Github</a>]</li> <li>eXtreme Gradient Boosting+Synthetic Minority Oversampling Technique (XGBoost+SMOTE) [<a href="https://figshare.com/s/87a4b5002c7488822e60" target="_blank" rel="noopener">Figshare</a>]</li> <li>eXtreme Gradient Boosting+Easy Data Augmentation (XGBoost+EDA) [<a href="https://github.com/shenyuanduanzui/xgboost_satd" target="_blank" rel="noopener">Github</a>]</li> <li>MT-Text-CNN [<a href="https://github.com/yikun-li/satd-different-sources-data" target="_blank" rel="noopener">Github</a>]</li> </ol> <p> </p> <div> <div><strong>Structure of the Replication Package:</strong></div> </div> <p>In accordance with the original dataset, the dataset comprises four distinct CSV files delineated by the artifacts under consideration in this study. Each CSV file encompasses a text column and a class, which indicate classifications denoting specific types of SATD, namely code/design debt (C/D), documentation debt (DOC), test debt (TES), and requirement debt (REQ) or Not-SATD.</p> <div> <div><code>├── SATD Keywords</code></div> <div><code>│ ├── Keywords based on Source of Artifacts</code></div> <div><code>│ │ ├── Code comment.txt</code></div> <div><code>│ │ ├── Commit message.txt</code></div> <div><code>│ │ ├── Issue section.txt</code></div> <div><code>│ │ └── Pull section.txt</code></div> <div><code>│ ├── Keywords based on Types of SATD</code></div> <div><code>│ │ ├── code-design debt.txt</code></div> <div><code>│ │ ├── documentation debt.txt</code></div> <div><code>│ │ ├── requirement debt.txt</code></div> <div><code>│ │ └── test debt.txt</code></div> <div><code>├── src</code></div> <div><code>│ ├── bert.py</code></div> <div><code>│ ├── bilstm.py</code></div> <div><code>│ └── preprocessing.py</code></div> <div><code>├── data-augmentation-code_comments.csv</code></div> <div><code>├── data-augmentation-commit_messages.csv</code></div> <div><code>├── data-augmentation-issues.csv</code></div> <div><code>├── data-augmentation-pull_requests.csv</code></div> <div> <div><code>└── Supplementary Material.docx</code></div> <div> </div> </div> <div> </div> </div> <p><strong>Requirements:</strong></p> <div> <div><a href="https://nlp.stanford.edu/projects/glove/" target="_blank" rel="noopener">glove</a></div> <div>nltk</div> <div>transformers</div> <div>torch</div> <div>tensorflow</div> <div>keras</div> <div>langdetect</div> <div>inflect</div> <div>inflection</div> </div> <div> </div> <div> </div> <div> <div><strong>Project sources for each artifact are as follows:</strong></div> </div> <div> </div> <div> <table> <tbody> <tr> <td><strong>Source code comment</strong></td> <td><strong>Issue section</strong></td> <td><strong>Pull section</strong></td> <td><strong>Commit message</strong></td> </tr> <tr> <td>ant<br>argouml<br>columba<br>emf<br>hibernate<br>jedit<br>jfreechart<br>jmeter<br>jruby<br>squirrel</td> <td> <div>camel</div> <div>chromium</div> <div>gerrit</div> <div>hadoop</div> <div>hbase</div> <div>impala</div> <div>thrift</div> </td> <td> <div>accumulo</div> <div>activemq</div> <div>activemq-artemis</div> <div>airflow</div> <div>ambari</div> <div>apisix</div> <div>apisix-dashboard</div> <div>arrow</div> <div>attic-apex-core</div> <div>attic-apex-malhar</div> <div>attic-stratos</div> <div>avro</div> <div>beam</div> <div>bigtop</div> <div>bookkeeper</div> <div>brooklyn-server</div> <div>calcite</div> <div>camel</div> <div>camel-k</div> <div>camel-quarkus</div> <div>camel-website</div> <div>carbondata</div> <div>cassandra</div> <div>cloudstack</div> <div>commons-lang</div> <div>couchdb</div> <div>cxf</div> <div>daffodil</div> <div>drill</div> <div>druid</div> <div>dubbo</div> <div>echarts</div> <div>fineract</div> <div>flink</div> <div>fluo</div> <div>geode</div> <div>geode-native</div> <div>gobblin</div> <div>griffin</div> <div>groovy</div> <div>guacamole-client</div> <div>hadoop</div> <div>hawq</div> <div>hbase</div> <div>helix</div> <div>hive</div> <div>hudi</div> <div>iceberg</div> <div>ignite</div> <div>incubator-brooklyn</div> <div>incubator-dolphinscheduler</div> <div>incubator-doris</div> <div>incubator-heron</div> <div>incubator-hop</div> <div>incubator-mxnet</div> <div>incubator-pagespeed-ngx</div> <div>incubator-pinot</div> <div>incubator-weex</div> <div>infrastructure-puppet</div> <div>jena</div> <div>jmeter</div> <div>kafka</div> <div>karaf</div> <div>kylin</div> <div>lucene-solr</div> <div>madlib</div> <div>myfaces-tobago</div> <div>netbeans</div> <div>netbeans-website</div> <div>nifi</div> <div>nifi-minifi-cpp</div> <div>nutch</div> <div>openwhisk</div> <div>openwhisk-wskdeploy</div> <div>orc</div> <div>ozone</div> <div>parquet-mr</div> <div>phoenix</div> <div>pulsar</div> <div>qpid-dispatch</div> <div>reef</div> <div>rocketmq</div> <div>samza</div> <div>servicecomb-java-chassis</div> <div>shardingsphere</div> <div>shardingsphere-elasticjob</div> <div>skywalking</div> <div>spark</div> <div>storm</div> <div>streams</div> <div>superset</div> <div>systemds</div> <div>tajo</div> <div>thrift</div> <div>tinkerpop</div> <div>tomee</div> <div>trafficcontrol</div> <div>trafficserver</div> <div>trafodion</div> <div>tvm</div> <div>usergrid</div> <div>zeppelin</div> <div>zookeeper</div> </td> <td> <div>accumulo</div> <div>activemq</div> <div>activemq-artemis</div> <div>airflow</div> <div>ambari</div> <div>apisix</div> <div>apisix-dashboard</div> <div>arrow</div> <div>attic-apex-core</div> <div>attic-apex-malhar</div> <div>attic-stratos</div> <div>avro</div> <div>beam</div> <div>bigtop</div> <div>bookkeeper</div> <div>brooklyn-server</div> <div>calcite</div> <div>camel</div> <div>camel-k</div> <div>camel-quarkus</div> <div>camel-website</div> <div>carbondata</div> <div>cassandra</div> <div>cloudstack</div> <div>commons-lang</div> <div>couchdb</div> <div>cxf</div> <div>daffodil</div> <div>drill</div> <div>druid</div> <div>dubbo</div> <div>echarts</div> <div>fineract</div> <div>flink</div> <div>fluo</div> <div>geode</div> <div>geode-native</div> <div>gobblin</div> <div>griffin</div> <div>groovy</div> <div>guacamole-client</div> <div>hadoop</div> <div>hawq</div> <div>hbase</div> <div>helix</div> <div>hive</div> <div>hudi</div> <div>iceberg</div> <div>ignite</div> <div>incubator-brooklyn</div> <div>incubator-dolphinscheduler</div> <div>incubator-doris</div> <div>incubator-heron</div> <div>incubator-hop</div> <div>incubator-mxnet</div> <div>incubator-pagespeed-ngx</div> <div>incubator-pinot</div> <div>incubator-weex</div> <div>infrastructure-puppet</div> <div>jena</div> <div>jmeter</div> <div>kafka</div> <div>karaf</div> <div>kylin</div> <div>lucene-solr</div> <div>madlib</div> <div>myfaces-tobago</div> <div>netbeans</div> <div>netbeans-website</div> <div>nifi</div> <div>nifi-minifi-cpp</div> <div>nutch</div> <div>openwhisk</div> <div>openwhisk-wskdeploy</div> <div>orc</div> <div>ozone</div> <div>parquet-mr</div> <div>phoenix</div> <div>pulsar</div> <div>qpid-dispatch</div> <div>reef</div> <div>rocketmq</div> <div>samza</div> <div>servicecomb-java-chassis</div> <div>shardingsphere</div> <div>shardingsphere-elasticjob</div> <div>skywalking</div> <div>spark</div> <div>storm</div> <div>streams</div> <div>superset</div> <div>systemds</div> <div>tajo</div> <div>thrift</div> <div>tinkerpop</div> <div>tomee</div> <div>trafficcontrol</div> <div>trafficserver</div> <div>trafodion</div> <div>tvm</div> <div>usergrid</div> <div>zeppelin</div> <div>zookeeper</div> </td> </tr> </tbody> </table> </div> <div> <div> </div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <p>This dataset has undergone a data augmentation process using the <a href="https://arxiv.org/abs/2302.13007" target="_blank" rel="noopener">AugGPT</a> technique. Meanwhile, the original dataset can be downloaded via the following link: <a href="https://github.com/yikun-li/satd-different-sources-data">https://github.com/yikun-li/satd-different-sources-data</a></p> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> <div> </div> </div> </div> </div> </div> </div>
TWIN SEEDS Work Package 2 data
<p>Data collected within Work Package 2 of the Horizon Europe project TWIN SEEDS (Grant agreement ID: 101056793).</p> <p>The WP2 report "Emerging trends of Global Value Chains and Multinational Enterprises in the pandemic time", using as inputs these data is publicly accessible here: https://twinseeds.eu/wp-content/uploads/2023/12/WP2-Report.pdf</p>
Data Package for "A Platform-Agnostic Approach for Automatically Identifying Real-Life Performance Issue Reports with Heuristic Linguistic Patterns"
<p>This Zenodo repository contains the data supporting the findings of the journal paper, titled "A Platform-Agnostic Approach for Automatically Identifying Real-Life Performance Issue Reports with Heuristic Linguistic Patterns", published on IEEE Transactions on Software Engineering, including:</p> <ol> <li><strong>Heuristic Linguistic Pattern Set</strong>: <span>we listed the 80 HLP we derived from </span><span>Apache's JIRA issue tracking system</span><span>. Column "</span><span>Category" </span><span>lists the type of each pattern. Namely, LEX represents lexical pattern, STR represents structural pattern, SEM represents semantic pattern, and PRF represents profiling pattern. Column "Name" is a descriptive name we give to each pattern. Column "Definition" defines the detailed content in each pattern.</span></li> <li><strong>Manual Tagging Results</strong>: manual_tagging.xlsx spreadsheet <span>comprises both sentence-level and issue-level manually tagging results for three datasets: 'Dataset-1: Apache Jira's Homologous Evaluation', '</span><span>Dataset-</span><span>2: Apache Jira's Heterologous Evaluation', and '</span><span>Dataset-</span><span>3: Other Platform's Evaluation'. The tagging results are segmented into sentence-level tabs ("Dataset-1 Sen", "Dataset-2 Sen", "Dataset-3 Sen") and issue-level tabs ("Dataset-1 Issue", "Dataset-2 Issue", "Dataset-3 Issue").</span></li> <li><span><strong>RQ Findings</strong>: </span> <p><span>This section contains detailed data findings from six research questions (RQ1 to RQ6).</span></p> <ul> <li> <p><span>The RQ1 tab provides an evaluation of our HLP-based approach, showing the precision, recall, and F1-Score of eight classifiers. These results are juxtaposed with the corresponding values from baseline methods, at both sentence and issue levels for automatic tagging.</span></p> </li> <li> <p><span>The RQ2 tab illustrates the precision, recall, and F1-Score of eight classifiers under two training conditions: a balanced training dataset (BT+HLP) and an imbalanced training dataset (UBT+HLP). These outcomes are contrasted with the equivalent values from baseline methods, also trained under balanced (BT+BLM) and imbalanced (UBT+BLM) conditions. The results are shown at both sentence and issue levels for automatic tagging.</span></p> </li> <li> <p><span>The RQ3 tab evaluates the dataset transferability of our HLP-based approach in comparison to baseline methods. It achieves this by analyzing the precision, recall, and F1-Score metrics for eight classifiers under two different "training/testing" dataset conditions, i.e., 'D1/D1' and 'D1/D3'. These conditions allow for a direct comparison of performance when applied to the same dataset ('D1/D1') versus when transferred to a different dataset ('D1/D3'). Additionally, the tab includes an 'Avg Change' and 'p-value' section, summarizing the statistical change in performance metrics between the two dataset conditions. </span></p> </li> <li> <p><span>The RQ4 tab presents a direct comparison between strict and fuzzy HLP matching approaches, assessed through precision, recall, and F1-Score metrics across eight issue classifiers.</span></p> </li> <li> <p><span>The RQ5 tab examines the influence of sentence order on the accuracy of eight classifiers within our approach. It shows the change in precision, recall, and F1-Score when the sentence order feature is taken into consideration versus when it is not.</span></p> </li> <li> <p><span>The RQ6 tab explores the impact of feature selection algorithms on both issue and sentence-level tagging accuracy. This tab presents the average precision, recall, and F1-Score for three experiments: Boruta, Recursive Feature Elimination (RFE), and the usage of all 80 features. </span></p> </li> </ul> </li> <li><strong>Qualitative Analysis</strong>: <p><span>This spreadsheet offers a comprehensive examination of the data supporting Section 6.1, which focuses on Qualitative Analysis. It is organized into several tabs, each dedicated to specific research questions (RQs) as outlined below:</span></p> <ul> <li> <p><span>Tab "RQ-1" showcases performance issue reports accurately detected by our High-Level Performance (HLP) approach's top model, XGBoost, which were not identified by the benchmark method's leading model, BERT. This highlights the comparative advantage of our approach in identifying nuanced performance issues.</span></p> </li> <li> <p><span>Tab "RQ-2" continues the exploration of performance issue reports, presenting cases with specific details (to be added).</span></p> </li> <li> <p><span>Tab "RQ-3" delves into the unique capabilities of XGBoost, the leading model in our HLP approach, showcasing its ability to detect performance issues missed by the baseline's top model, BERT. This comparison is drawn under distinct conditions: with pre-training (Dataset 1) and without pre-training (Dataset 3), illustrating the robustness and adaptability of our model.</span></p> </li> <li> <p><span>Tab "RQ-4" focuses on performance issue reports uniquely identified through the implementation of Fuzzy HLP Matching within our HLP approach. This method underscores the innovative matching techniques that enhance issue detection.</span></p> </li> <li> <p><span>Tab "RQ-5" presents performance issue reports pinpointed exclusively by applying the Issue HLP Matrix within our approach. This tab demonstrates the effectiveness of our matrix-based analysis in isolating and identifying specific performance concerns.</span></p> </li> <li> <p><span>Tab "RQ-6" is dedicated to performance issue reports uniquely detected by incorporating feature selection techniques into our HLP approach. This illustrates the value of advanced feature selection in improving the precision of performance issue identification.</span></p> </li> </ul> </li> <li><strong>LLM Experiment Data</strong>: presents the tagging outcomes of Large Language Models (LLMs), specifically ChatGPT-3.5 and ChatGPT-4, across three distinct datasets: 'Dataset-1: Apache Jira's Homologous Evaluation', 'Dataset-2: Apache Jira's Heterologous Evaluation', and 'Dataset-3: Evaluation on Other Platforms'. The results are organized into three separate tabs: 'Dataset-1 Issue', 'Dataset-2 Issue', and 'Dataset-3 Issue'.</li> <li><strong>ChatGPT Operation Python Script</strong>: crafted for automating the evaluation and tagging of issue reports in Excel using Large Language Models (LLMs) like ChatGPT-3.5 and ChatGPT-4. It underscores the importance of administrative rights for file modifications and outlines procedures for reading from and writing responses to Excel files. Key functions include querying LLMs with issue descriptions, processing their responses, and updating the spreadsheet with 'Yes' or 'No' labels and explanatory reasons, thereby facilitating an organized review of LLM performance across different datasets.</li> </ol>
TWIN SEEDS Work Package 5 data
<p><span>Data collected within Work Package 5 of the Horizon Europe project TWIN SEEDS (Grant agreement ID: 101056793).</span></p> <p><span>The WP5 report "Recent and emerging impact of GVCs and MNEs on employment and inequalities", using as inputs these data is publicly accessible here: <a href="https://twinseeds.eu/projects-outputs/reports/">https://twinseeds.eu/projects-outputs/reports/</a> </span></p>
TWIN SEEDS Work Package 4 data
<p><span>Data collected within Work Package 4 of the Horizon Europe project TWIN SEEDS (Grant agreement ID: 101056793).</span></p> <p><span>The WP4 report "Recent and emerging impact of GVCs and MNEs on employment and inequalities", using as inputs these data is publicly accessible here: <a href="https://twinseeds.eu/projects-outputs/reports/">https://twinseeds.eu/projects-outputs/reports/</a></span></p>
TWIN SEEDS Work Package 3 data
<p><span>Data collected within Work Package 3 of the Horizon Europe project TWIN SEEDS (Grant agreement ID: 101056793).</span></p> <p><span>The WP3 report "Recent and emerging impact of GVCs and MNEs on employment and inequalities", using as inputs these data is publicly accessible here: <a href="https://twinseeds.eu/projects-outputs/reports/">https://twinseeds.eu/projects-outputs/reports/</a> </span></p>
Data and code to replicate: Diet analysis using generalized linear models derived from foraging processes using R package mvtweedie
<p>Diet analysis integrates a wide variety of visual, chemical and biological identification of prey. Samples are often treated as compositional data, where each prey is analyzed as a continuous percentage of the total. However, analyzing compositional data results in analytical challenges, e.g., highly parameterized models or prior transformation of data. Here, we present a novel approximation involving a Tweedie generalized linear model (GLM). We first review how this approximation emerges from considering predator foraging as a thinned and marked point process (with marks representing prey species and individual prey size). This derivation can motivate future theoretical and applied developments. We then provide a practical tutorial for the Tweedie GLM using new package <i>mvtweedie</i> that extends capabilities of widely used packages in R (<i>mgcv</i> and <i>ggplot2</i>) by transforming output to calculate prey compositions. We demonstrate this approach and software using two examples. Tufted puffins (<i>Fratercula cirrhata</i>) provisioning their chicks on a colony in the northern Gulf of Alaska show decadal prey switching among sand lance and prowfish (1980-2000) and then Pacific herring and capelin (2000-2020), while wolves (<i>Canis lupus ligoni</i>) in Southeast Alaska forage on mountain goats and marmots in northern uplands and marine mammals in seaward island coastlines. </p>
Replication Package for the Paper: Transfer Learning with Time Series Data: A Systematic Mapping Study
<p>This is a replication package for the paper "Transfer Learning with Time Series Data: A Systematic Mapping Study".</p> <p>It provides</p> <ul> <li>a documentation of the conducted electronic literature search,</li> <li>exports of the search results from each literature database,</li> <li>and an excel file on the included literature and extracted data.</li> </ul>
Data package for Pozzolanic activity quantification of hollow glass microspheres
<p>Raw data of calengorimetry, TGA, UCS and XRD tests developed as part of the research published in Cement and Concrete Composites. The dataset is deposited in a .zip file in .txt or .xls(.xlsx) file format.</p> <p>Linked publication:</p> <p>C.M. Martín, N.B. Scarponi, Y.A. Villagrán, D.G. Manzanal, T.M. Piqué,<br> Pozzolanic activity quantification of hollow glass microspheres,<br> Cement and Concrete Composites,<br> Volume 118,<br> 2021,<br> 103981,<br> ISSN 0958-9465,<br> https://doi.org/10.1016/j.cemconcomp.2021.103981.<br> (https://www.sciencedirect.com/science/article/pii/S0958946521000500)</p>
seshatdb (Equinox Packaged Data)
<p>This is the first official release of Equinox Data on GitHub / Zenodo.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.