Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

997

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

997 results for “AWARENESS”

Learn how ShareScore rates datasets ↗
zenodo44/100

Replication Package for the paper "Conversing with business process-aware Large Language Models: the BPLLM framework"

<p>Replication Package for the research paper "<em>Conversing with business process-aware Large Language Models: the BPLLM framework</em>".</p> <p>The package includes the process models, the questions (and expected answers), the results of the qualitative evaluation, and the Hugging Face links to the fine-tuned versions of Llama 3.1 8B employed in the quantitative evaluation of the framework.</p> <p>In particular, the process models are:</p> <ul> <li>The natural language Directly-follows graph (DFG) of the Food Delivery process: <em>food_delivery_activities.txt</em> for the definition of the activities and <em>food_delivery_flow.txt</em> for the sequence flow.</li> <li>The BPMN model of the Food Delivery, E-commerce, and Reimbursement processes: <em>ecommerce.bpmn</em>, <em>food_delivery.bpmn</em>, and <em>reimbursement.bpmn</em>.</li> </ul> <p>The datasets with the questions and the expected answers are:</p> <ul> <li><em>1_questions_answers_not_refined_for_DFG.csv</em> ;</li> <li><em>1.1_questions_answers_refined_for_DFG.csv</em> ;</li> <li><em>2_questions_answers_not_refined.csv</em> ;</li> <li><em>3_questions_answers_refined.csv</em> ;</li> <li><em>4_questions_answers_different_processes.csv</em> ;</li> <li><em>5_questions_answers_similar_processes.csv</em> ;</li> <li><em>6_questions_answers_refined_ft.csv</em> .</li> </ul> <p>The complete results of the qualitative evaluation are contained in the file <em>qualitative_experiments_results.pdf</em>.</p> <p>The Hugging Face links to the fine-tuned versions of Llama 3.1 8B are reported in <em>hf_links_finetuned_models.pdf</em>.</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Regional Disparities in AI Awareness

<p>The dataset consists of variables linked to the benefits, risks, and impacts of AI and regional development in a variety of areas. Sources of the data: Office for National Statistics (ONS), UK.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Supporting data for "Raising awareness of potential biases in medical machine learning: Experience from a Datathon"

<p>This archive contains files from a Datathon held virtually in February<br>2024 to introduce clinicians and data scientists to the challenge of<br>reviewing a clinical dataset for potential biases.</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Raw data for Infrastructure and Awareness Landscape Analysis in sub-Saharan Africa

<p>Persistent identifiers that are well-connected are essential for enhancing research, researchers, and research institutions. The comprehensive raw data shared on PIDs infrastructure and awareness landscape analysis in sub-Saharan Africa is taken from service providers and organizations, including Open DOAR, the Registry of Open Access Repositories (ROAR), the Registry of Research Data Repositories (Re3data), UNESCO, Lyrasis (Dspace), Dataverse, Open Journal System (OJS), among others. The data shared here was collected in August 2023. The data shared are secondary data, and the position is strictly based on the primary source data author. The data is restricted to what is available on the internet and does not include locally hosted offline data.</p> <p>Further analysis of the raw data suggested some salient implications for PIDs awareness in the region. Variations were observed across the various data sources, while some interesting correlations emerged from the collected data. There are countries with some PIDs infrastructure, while others are yet to establish a visible presence in PIDs infrastructure. The visibility of PIDs is somewhat related to the awareness level as well as the policy established on open access in the represented countries across the region.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Longitudinal data to explore changes in flood risk awareness and preparedness

<p>This upload includes two different longitudinal datasets. The Panel datasets includes two rounds of surveys where the same individuals were interviewed. The Repeated Cross-Sectional includes two rounds of surveys where different individuals were interviewed in each round. The first survey round is the same in the two datasets. Data were collected in the municipality of Negrar (VR), in North-eastern Italy in February 2019 and in February 2020, following a flash flood which occurred in September 2018.&nbsp;</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

AWARE: Dataset for Aspect-Based Sentiment Analysis of Apps Reviews

<p>&nbsp;</p> <p><em><strong>The&nbsp;peer-reviewed paper of&nbsp;AWARE dataset is published in ASEW 2021, and can be accessed&nbsp;through: <a href="http://doi.org/10.1109/ASEW52652.2021.00049">http://doi.org/10.1109/ASEW52652.2021.00049</a>.&nbsp;Kindly cite this paper when using AWARE dataset.</strong></em></p> <p>&nbsp;</p> <p>Aspect-Based Sentiment Analysis (ABSA) aims to identify the opinion (sentiment) with respect to a specific aspect. Since there is a lack of&nbsp;<em>smartphone apps reviews</em>&nbsp;dataset that is annotated to support the ABSA task, we present AWARE:&nbsp;<strong>A</strong>BSA&nbsp;<strong>W</strong>arehouse of&nbsp;<strong>A</strong>pps&nbsp;<strong>RE</strong>views.</p> <p>AWARE contains apps reviews from three different domains (Productivity, Social Networking, and Games), as each domain has its distinct functionalities and audience. Each sentence is annotated with three labels, as follows:&nbsp;</p> <ul> <li><strong>Aspect Term:&nbsp;</strong>a term that exists in the sentence and describes an aspect of the app that is expressed by the sentiment. A term value of &ldquo;N/A&rdquo; means that the term is not explicitly mentioned in the sentence.</li> <li><strong>Aspect Category:</strong>&nbsp;one of the pre-defined set of domain-specific categories that represent an aspect of the app (e.g., security, usability, etc.).</li> <li><strong>Sentiment:</strong>&nbsp;positive or negative.</li> </ul> <p><em>Note: games domain does not contain aspect terms.</em></p> <p>We provide a comprehensive dataset of 11323 sentences from the three domains, where each sentence is additionally annotated with a Boolean value indicating whether the sentence expresses a positive/negative opinion. In addition, we provide three separate datasets, one for each domain, containing only sentences that express opinions. The file named &ldquo;AWARE_metadata.csv&rdquo; contains a description of the dataset&rsquo;s columns.</p> <p><strong>How AWARE can be used?</strong></p> <p>We designed AWARE such that it can be used to serve various tasks. The tasks can be, but are not limited to:</p> <ul> <li>Sentiment Analysis.</li> <li>Aspect Term Extraction.</li> <li>Aspect Category Classification.</li> <li>Aspect Sentiment Analysis.</li> <li>Explicit/Implicit Aspect Term Classification.</li> <li>Opinion/Not-Opinion Classification.</li> </ul> <p>Furthermore, researchers can experiment with and investigate the effects of different domains on users&#39; feedback.</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

IIIF: raising awareness of the user benefits for scholarly editions - usability testing results

<p>Usability testing results of the bachelor&#39;s thesis titled <a href="https://doc.rero.ch/record/306498/"><em>&quot;The International Image Interoperability Framework (IIIF): raising awareness of the user benefits for scholarly editions&quot;</em></a><a href="https://doc.rero.ch/record/306498/">.</a></p> <p>Remote and in-person usability tests on the <a href="http://universalviewer.io/">Universal Viewer</a> and <a href="http://projectmirador.org/">Mirador</a>, two IIIF-compliant clients, took place between March and May 2017. The tests were conducted with Loop11 (remote testing) and Morae (in-person testing).</p> <p>The dataset is composed of Excel files, screenshots (tasks and heat maps) as well as videos.</p>

opencc-by-4.0Jul 2017View details →
zenodo44/100

Diversity Awareness in Software Engineering Participant Research

<p>This dataset contains the result of a classification of three ICSE venues namely, ICSE 2019, 2020, and 2021 technical tracks, as stated in the methodology of the paper &ldquo;Diversity&nbsp;awareness in software engineering participant studies&rdquo; by Dutta et al. (2023).</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Green Intelligence Awareness, inhabitants attended in Valladolid city

<p>Changes in behavior and human attitudes are fundamental to achieve a more sustainable world, so that, it is very interesting to analyze the potential of an activity or intervention to increase the green intelligence awareness of a population.</p> <p>There is enormous opportunity for nature-based solutions to promote understanding of sustainability in ways that positively influence citizen behavior. There are numerous available resources to learn and understand the fragility of our environmental and the responsibility of humans to protect, preserve and respect the world. Therefore, this KPI aims to reflect how the intervention is used for educational purposes and enhancement of public awareness.&nbsp; &nbsp;&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Green Intelligence Awareness, educational actions in Valladolid city

<p>Changes in behavior and human attitudes are fundamental to achieve a more sustainable world, so that, it is very interesting to analyze the potential of an activity or intervention to increase the green intelligence awareness of a population.</p> <p>There is enormous opportunity for nature-based solutions to promote understanding of sustainability in ways that positively influence citizen behavior. There are numerous available resources to learn and understand the fragility of our environmental and the responsibility of humans to protect, preserve and respect the world. Therefore, this KPI aims to reflect how the intervention is used for educational purposes and enhancement of public awareness.&nbsp; &nbsp;&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

ASPLOS20-AE Artifact Dataset for 'Noise-Aware Dynamical System Compilation for Analog Devices with Legno'

<p>The empirical model database and dataset for the ASPLOS 2020 Paper &#39;Noise-Aware Dynamical System Compilation for Analog Devices with Legno&#39;</p>

opencc-by-4.0Jan 2020View details →
zenodo40/100

Training dataset used in the magazine paper entitled "A Flexible Machine Learning-Aware Architecture for Future WLANs"

<p><a href="https://arxiv.org/pdf/1910.03510.pdf"><strong>A Flexible Machine Learning-Aware Architecture for Future WLANs</strong></a></p> <p><strong>Authors: </strong>Francesc Wilhelmi, Sergio Barrachina-Mu&ntilde;oz, Boris Bellalta, Cristina Cano, Anders Jonsson &amp; Vishnu Ram.</p> <p><strong>Abstract:&nbsp;</strong>Lots of hopes have been placed in Machine Learning (ML) as a key enabler of future wireless networks. By taking advantage of the large volumes of data generated by networks, ML is expected to deal with the ever-increasing complexity of networking problems. Unfortunately, current networking systems are not yet prepared for supporting the ensuing requirements of ML-based applications, especially for enabling procedures related to data collection, processing, and output distribution. This article points out the architectural requirements that are needed to pervasively include ML as part of future wireless networks operation. To this aim, we propose to adopt the International Telecommunications Union (ITU) unified architecture for 5G and beyond. Specifically, we look into Wireless Local Area Networks (WLANs), which, due to their nature, can be found in multiple forms, ranging from cloud-based to edge-computing-like deployments. Based on ITU&#39;s architecture, we provide insights on the main requirements and the major challenges of introducing ML to the multiple modalities of WLANs.</p> <p><strong>Dataset description:&nbsp;</strong>This is the dataset generated for training a Neural Network (NN) in the Access Point (AP) (re)association problem in IEEE 802.11 Wireless Local Area Networks (WLANs).&nbsp;</p> <p>In particular, the NN is meant to output a prediction function of the throughput that a given station (STA) can obtain from a given Access Point (AP) after association. The features included in the dataset are:</p> <ol> <li>Identifier of the AP to which the STA has been associated.</li> <li>RSSI obtained from the AP to which the STA has been associated.</li> <li>Data rate in bits per second (bps) that the STA is allowed to use for the selected AP.</li> <li>Load in packets per second (pkt/s)&nbsp;that the STA generates.</li> <li>Percentage of data that the AP is able to serve before the user association is done.</li> <li>Amount of traffic load in pkt/s handled by the AP before the user association is done.</li> <li>Airtime in % that the AP enjoys before the user association is done.</li> <li>Throughput in pkt/s that the STA receives after the user association is done.</li> </ol> <p>The dataset has been generated through random simulations, based on the model provided in <a href="https://github.com/toniadame/WiFi_AP_Selection_Framework">https://github.com/toniadame/WiFi_AP_Selection_Framework</a>. More details regarding the dataset generation have been provided in&nbsp;<a href="https://github.com/fwilhelmi/machine_learning_aware_architecture_wlans">https://github.com/fwilhelmi/machine_learning_aware_architecture_wlans</a>.</p>

opencc-by-4.0Jan 2020View details →
zenodo40/100

Öffentlichkeitsarbeit Forschungsdatenmanagement (PR-Work) - 'Awareness' zum Forschungsdatenmanagement

<p>Diese Fotos dienen der &#39;Awareness&#39; zum FDM. Sie werden f&uuml;r Brosch&uuml;ren im Sommersemester 2020 benutzt, auf Info.screens, auf der Webseite der Stiftung Universit&auml;t Hildesheim, usw.</p>

opencc-by-4.0May 2020View details →
zenodo40/100

COCOA: Cold Start Aware Capacity Planning for Function-as-a-Service Platforms

<p>This dataset release supports the results presented in the paper &quot;COCOA: Cold Start Aware Capacity Planning for Function-as-a-Service Platforms&quot; by A. U. Gias and G. Casale, accepted in IEEE International Symposium on Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS), 2020.</p> <p>When referring to the dataset please cite the paper above.</p>

opencc-by-4.0Sep 2020View details →
zenodo40/100

AWARE characterization factor samples

<p>Files contain 5000 samples of AWARE characterization factors, as well as sampled independent data used in their calculations and selected intermediate results.</p> <p>AWARE is a consensus-based method development to assess water use in LCA. It was developed by the&nbsp;<a href="http://www.wulca-waterlca.org/index.html">WULCA UNEP/SETAC working group</a>. Its characterization factors represent the relative Available WAter REmaining per area in a watershed, after the demand of humans and aquatic ecosystems has been met. It assesses the potential of water deprivation, to either humans or ecosystems, building on the assumption that the less water remaining available per area, the more likely another user will be deprived.</p> <p>The code used to generate the samples can be found here:&nbsp;<a href="https://github.com/PascalLesage/aware_cf_calculator/">https://github.com/PascalLesage/aware_cf_calculator/</a></p> <p>Samples were updated from v1.0 in 2020 to include model uncertainty associated with the choice of&nbsp;WaterGap as the global hydrological model (GHM).</p> <p>The following datasets are supplied:</p> <p><strong>1) AWARE_characterization_factor_samples.zip</strong></p> <p>Actual characterization factors resulting from the Monte Carlo Simulation. Contains 4 zip files:</p> <p>&nbsp; &nbsp; *&nbsp;monthly_cf.zip: contains 116,484 arrays of 5000 monthly characterization factor samples for each of 9707 watershed and for each month, in csv format. Names are cf_&lt;BAS34S_ID&gt;_&lt;MONTH&gt;.csv, where &lt;BAS34S_ID&gt; is the watershed id and &lt;MONTH&gt; is the first three letters of the month (&#39;jan&#39;, &#39;feb&#39;, etc.).</p> <p>&nbsp; &nbsp; *&nbsp;average_agri_cf.zip: contains 9707&nbsp;arrays of 5000 annual average, agricultural use,&nbsp;characterization factor samples for each watershed, in csv format. Names are cf_average_agri_&lt;BAS34S_ID&gt;.csv.</p> <p>&nbsp; &nbsp; *&nbsp;average_non_agri_cf.zip: contains 9707&nbsp;arrays of 5000 annual average, non-agricultural use,&nbsp;characterization factor samples for each watershed, in csv format. Names are cf_average_non_agri_&lt;BAS34S_ID&gt;.csv.</p> <p>&nbsp; &nbsp; *&nbsp;average_unknown_cf.zip: contains 9707&nbsp;arrays of 5000 annual average, unspecified&nbsp;use,&nbsp;characterization factor samples for each watershed, in csv format. Names are cf_average_unknown_&lt;BAS34S_ID&gt;.csv..</p> <p><strong>2) AWARE_base_data.xlsx</strong>&nbsp;</p> <p>Excel file with the deterministic data, per watershed and per month, for each of the independent variables used in the calculation of AWARE characterization factors. Specifically, it includes:</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; Monthly irrigation<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Description: irrigation water, per month, per basin<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Unit: m3/month<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Location in Excel doc: Irrigation<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; File name once imported: irrigation.pickle<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; table shape: (11050, 12)</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; Non-irrigation hwc: electricity, domestic, livestock, manufacturing<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Description: non-irrigation uses of water<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Unit: m3/year<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Location in Excel doc: hwc_non_irrigation<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; File name once imported: electricity.pickle, domestic.pickle,<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; livestock.pickle, manufacturing.pickle<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; table shape: 3 x (11050,)</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; avail_delta<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Description: Difference between &quot;pristine&quot; natural availability<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; reported in PastorXNatAvail and natural availability calculated<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; from &quot;Actual availability as received from WaterGap - after<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; human consumption&quot; (Avail!W:AH) plus HWC.<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; This should be added to calculated water availability to<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; get the water availability used for the calculation of EWR<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Unit: m3/month<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Location in Excel doc: avail_delta<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; File name once imported: avail_delta.pickle<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; table shape: (11050, 12)</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; avail_net<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Description: Actual availability as received from WaterGap - after human consumption<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Unit: m3/month<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Location in Excel doc: avail_net<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; File name once imported: avail_net.pickle<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; table shape: (11050, 12)</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; pastor<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Description: fraction of PRISTINE water availability that should be reserved for environment<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Unit: unitless<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Location in Excel doc: pastor<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; File name once imported: pastor.pickle<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; table shape: (11050, 12)</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; area<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Description: area<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Unit: m2<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Location in Excel doc: area<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; File name once imported: area.pickle<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; table shape: (11050,)<br> It also includes:</p> <p>* information (k values) on the distributions used for each variable (uncertainty tab)</p> <p>* information (k values) on the model uncertainty (model uncertainty tab)</p> <p>* two filters used to exclude watersheds that are either in Greenland (polar filter) or without data from the Pastor et al. (2014) method (122 cells), representing small coastal cells with no direct overlap (pastor filter). (filters tab)</p> <p><strong>3) independent_variable_samples.zip</strong></p> <p>Samples for each of the independent variables used in the calculation of characterization factors. Only random variables are contained. For all watershed or watershed-months without samples, the Monte Carlo simulation used the deterministic values found in the&nbsp;AWARE_base_data.xlsx file.&nbsp;</p> <p>The files are in csv format. The first column contains the watershed id (BAS34S_ID) if the data is annual or the (BAS34S_ID, month) for data with a monthly resolution. the other 5000 columns contain the sampled data.</p> <p>The names of the files are &lt;variable_name.csv&gt;.</p> <p><strong>4) intermediate_variables.zip</strong></p> <p>Contains results of intermediate calculations, used in the calculation of characterization factors. The zip file contains 3&nbsp;zip files:</p> <p>&nbsp; &nbsp; * AMD_world_over_AMD_i.zip: contains&nbsp;116,484 arrays (for each watershed-month) of 5000 calculated values of the ratio&nbsp;between the AMD (Availability Minus Demand) for the watershed-month and AMD_glo, the world weighted AMD average. Format is csv.<br> &nbsp; &nbsp; * AMD_world.zip: contains one array of 5000 calculated values of the world average AMD. Format is csv.</p> <p>&nbsp; &nbsp; * HWC.zip:&nbsp;&nbsp;contains&nbsp;116,484 arrays (for each watershed-month) of 5000 calculated values of the total Human Water Consumption. Format is csv.</p> <p><strong>5) watershedBAS34S_ID.zip</strong></p> <p>Contains the GIS files to link the watershed ids (BAS34S_ID) to actual spatial data.&nbsp;</p>

opencc-by-4.0Sep 2019View details →
zenodo40/100

Communication Characterization and Optimization of Applications Using Topology-Aware Task Mapping on Large Supercomputers (Data)

<p>Auxiliary materials for ICPE 2016 paper titled:<br /> &quot;Communication Characterization and Optimization of Applications Using Topology-Aware Task Mapping on Large Supercomputers&quot;.</p> <p>&nbsp;</p>

opencc-by-sa-4.0Feb 2016View details →
zenodo40/100

An Uncertainty-Aware Approach to Optimal Configuration of Stream Processing Systems

<p>The datasets in this release support the results presented in the paper</p> <blockquote> <p>P. Jamshidi, G. Casale, "An Uncertainty-Aware Approach to Optimal Configuration of Stream Processing Systems", accepted for presentation at MASCOTS 2016.</p> </blockquote> <p>An open access to the paper is available at https://arxiv.org/abs/1606.06543</p> <blockquote> <p>Also open source code is available at https://github.com/dice-project/DICE-Configuration-BO4CO</p> </blockquote> <p>The archive contains 10 comma separated datasets representing performance measurements (throughput and latency) for 3 different stream benchmark applications. These have been experimentally collected on 5 different cloud cluster over the course of 3 months (24/7). Each row in the datasets represents a different configuration setting for the application and the last two columns represent the average performance of the application measured over the course of 10 minutes under that specific configuration setting. The datasets contains a full factorial and exhaustive measurements for all possible settings limited to a predetermined interval for each variable. Each dataset is named in the following format: "<em>benchmark_application-dimensions-cluster_name</em>". For example, "wc-6d-c1" refers to WordCount benchmark application with 6 dimensions (i.e., we varied 6 configuration parameters) and the application was deployed on c1 cluster (OpenNebula, see Appendix). This resulted in a dataset of size 2880, i.e., it has taken 2880*10m=480h=20days for collecting the data!  </p> <p>For more information about the data refer to the appendix of the paper: https://arxiv.org/abs/1606.06543. </p> <p>When referring to the dataset or code please cite the paper above.</p>

openbsd-3-clauseJun 2016View details →
zenodo40/100

Interformer: An Interaction-Aware Model for Protein-Ligand Docking and Affinity Prediction

<p>The code, dataset, and model weights are described in the paper "Interformer: An Interaction-Aware Model for Protein-Ligand Docking and Affinity Prediction."</p> <p>&nbsp;</p> <p><strong>experiment_results.zip:</strong> Contains generated results that can reproduce the result from the reported paper.</p> <p><strong>benchmark.zip:</strong> Contains docking and affinity input data of the interformer. You can use the source code to make predictions and reproduce the number of the reported paper.</p> <p><strong>checkpoints.zip: </strong>Contains one weight for the Energy and four PoseScore and Affinity models.</p> <p><strong>source_code_1.0.zip:</strong> Contains the initial version of the source code.</p> <p><strong>interformer_train.tar.gz:</strong> Contains prepared training data for interformer. poses/ contains all structure need for training, poses/ligand contains the re-docking poses generated by interformer energy, poses/ligand/rcsb contains the conformation of reference ligand, poses/pocket contains all pocket extract by raw PDB from rcsb, poses/uff contains all ligand conformation minimized using UFF from reference ligand, and train/ contains the training csv.</p> <p><strong>baseline_results.tar.gz:</strong>&nbsp; Contains the predictions from three methods: Interformer, DiffDock, and DeepDock. The results align with the exact numbers reported in the paper. For further details, please refer to the <em>eda/ </em>directory.</p> <p>&nbsp;</p> <p>You can also find the newest version of the source code at <a href="https://github.com/tencent-ailab/Interformer" target="_blank" rel="noopener">https://github.com/tencent-ailab/Interformer</a></p> <p>&nbsp;</p>

openapache2.0Mar 2024View details →
zenodo40/100

Supplementary materials for "Improving diffusion-based protein backbone generation with global-geometry-aware latent encoding"

<h1>Info</h1> <p>This dataset contains the supplementary materials for &nbsp;"Improving diffusion-based protein backbone generation with global-geometry-aware latent encoding".&nbsp;</p> <p>For&nbsp;<strong>source code&nbsp;</strong>and&nbsp;<strong>detailed instructions on usage,&nbsp;</strong>please refer to our <a href="https://github.com/meneshail/TopoDiff/tree/main" target="_blank" rel="noopener">github</a> .</p> <h1>Supplementary data</h1> <h2>weights.tar.gz</h2> <p>The trained model weights used in the paper.</p> <h2>dataset.zip</h2> <p>CATH-60 Dataset used in the paper. In the notebook directory of our <a href="https://github.com/meneshail/TopoDiff/tree/main" target="_blank" rel="noopener">github</a> , we provide an example on encoding and visualize it with our trained encoder.</p> <h2>design.zip</h2> <p>The 21 novel mainly-beta designs selected for experiment validation. Along with the generated backbone, we also provide the prediction results from AlphaFold and ESMFold.</p> <h2>benchmark_sample.zip</h2> <p>Sampled backbones used for all benchmark experiment (All methods and variants included).</p> <h2>evaluation.tar.gz</h2> <p>Precomputed CATH reference data for coverage metric computation. Need to be downloaded for using evaluation scripts.&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Diversity-aware Fairness Testing of Machine Learning Classifiers through Hashing-based Sampling

<p>The experimental results of the evaluation of VBT-X.</p> <h2>Abstract</h2> <div> <h3>Context:</h3> <p>There are growing concerns about algorithmic fairness, as some machine learning (ML)-based algorithms have been found to exhibit biases against protected attributes such as gender, race, age and so on. Individual fairness requires an ML classifier to produce similar outputs for similar individuals. Verification Based Testing (<span>Vbt</span>) is a state-of-the-art black-box testing algorithm for individual fairness that leverages constraint solving to generate test cases.</p> </div> <div> <h3>Objective:</h3> <p>Generating diverse test cases is expected to facilitate efficient detection of diverse discriminatory data instances (i.&nbsp;e., cases that violate individual fairness). Hashing-based sampling techniques draw a sample approximately uniformly at random from the set of solutions of given Boolean constraints. We propose <span>Vbt</span>-X, which improves <span>Vbt</span> with hashing-based sampling, aiming to improve its testing performance.</p> </div> <div> <h3>Method:</h3> <p>We realize hashing-based sampling for <span>Vbt</span>. The challenge is that the off-the-shelf hashing-based sampling techniques cannot be integrated in a straightforward manner because the constraints in <span>Vbt</span> are generally not Boolean. Moreover, we propose several enhancement techniques to make <span>Vbt</span>-X more efficient.</p> </div> <div> <h3>Results:</h3> <p>To evaluate our method, we conduct experiments, where <span>Vbt</span>-X is compared to <span>Vbt</span>, <span>Sg</span> and ExpGA (other well-known fairness testing algorithms) over a set of configurations consisting of several datasets, protected attributes, and ML classifiers. The results show that, with each configuration, <span>Vbt</span>-X detects more discriminatory data instances with higher diversity than <span>Vbt</span> and <span>Sg</span>. <span>Vbt</span>-X detects discriminatory data instances with higher diversity than ExpGA, though the number of discriminatory data instances detected by <span>Vbt</span>-X is lesser than ExpGA.</p> </div> <div> <h3>Conclusion:</h3> <p>Our proposed method performs better than other state-of-the-art black-box fairness testing algorithms, particularly in terms of diversity. Our method can serve to efficiently identify flaws in ML classifiers with respect to individual fairness for subsequent improvements of an ML classifier. On the other hand, although our method is specific to individual fairness, it could work for testing other aspects of a software system such as security and counterfactual explanations with some technical adaptations, which remains for future work.</p> </div> <p>&nbsp;</p> <div> <h2>Acknowledgments</h2> <p>This paper is partly based on results obtained from a project, JPNP20006, commissioned by the New Energy and Industrial Technology Development Organization (NEDO). This paper is supported by JST SPRING, Grant Number JPMJSP2131.</p> </div>

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record