Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
725
datasets available to search
ShareScore release 0.9.0
Dataset results
725 results for “recommendation”
Pubmed Journal Recommendation System dataset
<p>Dataset for Journal recommendation, includes title, abstract, keywords, and journal.</p> <p>We extracted the journals and more information of:</p> <p>Jiasheng Sheng. (2022). PubMed-OA-Extraction-dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6330817.</p> <p>Dataset Components:</p> <ul> <li> <p><strong>data_pubmed_all:</strong> This dataset encompasses all articles, each containing the following columns: 'pubmed_id', 'title', 'keywords', 'journal', 'abstract', 'conclusions', 'methods', 'results', 'copyrights', 'doi', 'publication_date', 'authors', 'AKE_pubmed_id', 'AKE_pubmed_title', 'AKE_abstract', 'AKE_keywords', 'File_Name'.</p> </li> <li> <p><strong>data_pubmed:</strong> To focus on recent and relevant publications, we have filtered this dataset to include articles published within the last five years, from January 1, 2018, to December 13, 2022—the latest date in the dataset. Additionally, we have exclusively retained journals with more than 200 published articles, resulting in 262,870 articles from 469 different journals.</p> </li> <li> <p><strong>data_pubmed_train, data_pubmed_val, and data_pubmed_test:</strong> For machine learning and model development purposes, we have partitioned the 'data_pubmed' dataset into three subsets—training, validation, and test—using a random 60/20/20 split ratio. Notably, this division was performed on a per-journal basis, ensuring that each journal's articles are proportionally represented in the training (60%), validation (20%), and test (20%) sets. The resulting partitions consist of 157,540 articles in the training set, 52,571 articles in the validation set, and 52,759 articles in the test set.</p> </li> </ul>
IPBES Data Management Tutorials - Session 4.3: Recommendations and considerations for data backups
<p>The <em>IPBES data management tutorials</em> are short videos to help experts implement the IPBES data management Policy. They cover topics ranging from data management policy, reports, active research data, tools, and examples.</p> <p>The <em>Data management of active research data </em>chapter provides an introduction for IPBES experts on how to manage data while actively being used, analyzed, and produced to fulfill the criteria of the IPBES data management policy.</p> <p>This session,<em> Recommendations and considerations for data backups</em>, reviews the importance of data backups, provides resources for further information, and discusses specific considerations one should keep in mind. </p>
Recommended food alternatives (healthier, eco-friendly, and cost-effective)
<p>It includes recommended food alternatives (healthier, eco-friendly, and cost-effective) for items selected from receipts. This dataset is valuable for research in consumer food science, as it captures the food choices of a small group of consumers over 21 days. It is also useful for machine learning training. All food items are linked to NAct ontology.</p> <p> </p>
Dataset for paper: A Systematic Literature Review and Recommendations for Ontology-based Support of Digital Forensics
<p>PLEASE, READ THE README.TXT FILE</p> <p>This document describes how to interpret the data and metadata files, and it is licensed under Creative Commons CC BY-NC-AS (https://creativecommons.org/licenses).</p> <p>The file "primary_studies_final_set-DATA.csv" is a CSV file format and contains the raw data extracted from our systematic literature review primary studies. Such data were extracted based on the research questions defined for our study.<br> The file "primary_studies_final_set-METADATA.csv" is a CSV file format and contains the following:<br> - the first row contains two pieces of information: the data type, which might be original or reused;<br> - the second row contains the reused data URL/DOI, which should inform the URL or DOI from which the data was reused, or n/a if the data is original;<br> - the third row contains the date of data generation in the format mm/dd/yyyy;<br> - the fourth row contains 11 elements describing each of the fields of the file "primary_studies_final_set-DATA.csv": the study ID, title, objective, six research questions, and an observation field; and<br> - the fifth row describes the data type of each field of the file "primary_studies_final_set-DATA.csv".<br> The .bib files contain the bibtex entry for the final set of studies.<br> The license.txt file describes the Creative Commons license for this material.</p> <p>We hope you have an excellent read!!</p> <p>Cheers!<br> Thiago, Edson, and Avelino</p>
Land Productivity Dynamics (LPD) maps based on approaches recommended by the United Nations Convention to Combat Desertification (UNCCD) for the Brazilian Semiarid Region (BSR).
<p>Maps of the land productivity dynamic (LPD) based on approaches recommended by the United Nations Convention to Combat Desertification (UNCCD) for the Brazilian Semiarid Region for the period 2001-2015. LPD from Trends.Earth (TE), LPD from the Joint Research Center (JRC-LPD), and LPD from the Food and Agriculture Organization – World Overview of Conservation Approaches and Technologies (FAO-WOCAT LPD). In addition, TE-based LPD with climate correction based on Rain Use Efficiency (RUE), Residual Trend Analysis (RESTREND), and Water Use Efficiency (WUE) calculated using the procedures described in the second version of the Good Practice Guidance for SDG Indicator 15.3.1. The annual LCLU maps from the MapBiomas project at 30 m spatial resolution for 2001 and 2015, and the 16-day MOD13Q1 NDVI dataset from 2001 to 2015 were used as inputs.</p> <p>**********************</p> <p>A total of seven GeoTIFF files in Geographic Tagged Image File Format (GeoTIFF) format are provided at 250 m spatial resolution.</p> <p>Coding for the LPDs.</p> <p>Value Meaning</p> <p>-32768 No data</p> <p>1 Declining</p> <p>2 Moderate decline</p> <p>3 Stressed</p> <p>4 Stable</p> <p>5 Increasing</p> <p>Funding: this study was undertaken as part of the Satellite Desertification Monitoring Program in the Brazilian Semiarid Region [Grant Number 403223/2021-0] supported by the CNPq. It also had the support of Capes, through Notice no. 28/2022 – PDPG Social Vulnerability & Human Rights [Grant Number 88881.705050/2022-01].</p>
Culture-Aware Music Recommendation Dataset
<p><strong>LFM-1b dataset extended by acoustic track features and cultural cues describing users</strong></p> <p> </p> <p>This dataset is based on the LFM-1b dataset (cf. <a href="http://www.cp.jku.at/datasets/LFM-1b/">http://www.cp.jku.at/datasets/LFM-1b/</a>), however, adds acoustic features describing the tracks to the original dataset as well as cultural aspects describing users (taken from Hofstede's six dimension model and the World Happiness Report) on the country-level.</p> <p>For the creation of the dataset, we extract all users for which the original dataset contains country information for. We extract the listening events of these users and match the tracks against the Spotify API to subsequently retrieve the acoustic features of these tracks (cf. [Spotify Audio Feature Description](https://developer.spotify.com/documentation/web-api/reference/object-model/#audio-features-object)). The final dataset contains only events of users with country information and tracks with acoustic features, which can be matched with the country-level data of the World Happiness Report and Hofstede's cultural dimensions to add cultural and socio-economic aspects for users.</p> <p>This new dataset contains</p> <ul> <li>55,190 users</li> <li>3,471,884 tracks including acoustic features</li> <li>351,469,333 listening events of those users for tracks we have obtained acoustic features for</li> <li>Hofstede's cultural dimensions for 47 countries</li> <li>World Happiness Report (WHR) data for 164 countries</li> </ul> <p> </p> <p><strong>Files</strong><br> All files are tab-separated, with no quoting of strings. The dataset contains the following files, whose content we describe in more detail in the following parts.</p> <p>* acoustic_features_lfm_id.tsv: acoustic features for all tracks in the dataset, identified by their LFM track identifier<br> * events.tsv: listening events for all users<br> * hofstede.tsv: Hofstede's cultural dimensions<br> * users.tsv: user metadata<br> * world_happiness_report_2018.tsv: World Happiness Report data</p> <p>For further information on the contents of these files, please cf. the Readme file.</p> <p> </p> <p>Please cite the following paper when using the dataset:<br> Zangerle, E., Pichl, M. and Schedl, M., 2020. User Models for Culture-Aware Music Recommendation: Fusing Acoustic and Cultural Cues. <em>Transactions of the International Society for Music Information Retrieval</em>, 3(1), pp.1–16. DOI: <a href="http://doi.org/10.5334/tismir.37">http://doi.org/10.5334/tismir.37</a></p>
GitRec - Github Project Recommender Systems
<p>This dataset contains the data collected using the Google API for the GHTorrent project and which were applied in the doctoral thesis directed to recommending projects on the GitHub platform</p>
Collected recommendations and requirements for FAIR-enabling services
<p>Within FAIRsFAIR task 2.4, we carried out a structured literature review to extract requirements, recommendations and other desiderata for FAIR-enabling services. This document contains the full list of excerpts, structured and annotated. This work has been used as input for the basic framework on FAIRness of services developed by FAIRsFAIR task 2.4 (see https://doi.org/10.5281/zenodo.4292599).</p>
Identifying patterns and recommendations of and for sustainable open data initiatives: a benchmarking-driven analysis of open government data initiatives among European countries
<p>This dataset contains data collected during a study <a href="https://www.sciencedirect.com/science/article/pii/S0740624X23000989"><em><strong>"Identifying patterns and recommendations of and for sustainable open data initiatives: a benchmarking-driven analysis of open government data initiatives among European countries"</strong></em></a> conducted by <em>Martin Lnenicka (University of Pardubice, Pardubice, Czech Republic), Anastasija Nikiforova (University of Tartu, Tartu, Estonia), Mariusz Luterek (University of Warsaw, Warsaw, Poland), Petar Milic (University of Pristina - Kosovska Mitrovica, Kosovska Mitrovica, Serbia), Daniel Rudmark (University of Gothenburg and RISE Research Institutes of Sweden, Gothenburg, Sweden), Sebastian Neumaier (St. Pölten University of Applied Sciences, Austria), Caterina Santoro (KU Leuven, Leuven, Belgium), Cesar Casiano Flores (University of Twente, Twente, the Netherlands), Marijn Janssen (Delft University of Technology, Delft, the Netherlands), Manuel Pedro Rodríguez Bolívar (University of Granada, Granada, Spain).</em></p> <p>It is being made public both to act as supplementary data for "<em>Identifying patterns and recommendations of and for sustainable open data initiatives: a benchmarking-driven analysis of open government data initiatives among European countries</em>", Government Information Quarterly*, and in order for other researchers to use these data in their own work. </p> <p>***Methodology***</p> <p>The paper focuses on benchmarking of open data initiatives over the years and attempts to identify patterns observed among European countries that could lead to disparities in the development, growth, and sustainability of open data ecosystems. </p> <p>This study examines existing benchmarks, indices, and rankings of open (government) data initiatives to find the contexts by which these initiatives are shaped, both of which then outline a protocol to determine the patterns. The composite benchmarks-driven analytical protocol is used as an instrument to examine the understanding, effects, and expert opinions concerning the development patterns and current state of open data ecosystems implemented in eight European countries - Austria, Belgium, Czech Republic, Italy, Latvia, Poland, Serbia, Sweden. 3-round Delphi method is applied to identify, reach a consensus, and validate the observed development patterns and their effects that could lead to disparities and divides. Specifically, this study conducts a comparative analysis of different patterns of open (government) data initiatives and their effects in the eight selected countries using six open data benchmarks, two e-government reports (57 editions in total), and other relevant resources, covering the period of 2013–2022.</p> <p>***Description of the data in this data set***</p> <p>The file "OpenDataIndex_<em>2013_</em>2022" collects an overview of 27 editions of 6 open data indices - for all countries they cover, providing respective ranks and values for these countries. These indices are:</p> <p>1) Global Open Data Index (GODI) (4 editions)</p> <p>2) Open Data Maturity Report (ODMR) (8 editions)</p> <p>3) Open Data Inventory (ODIN) (6 editions)</p> <p>4) Open Data Barometer (ODB) (5 editions)</p> <p>5) Open, Useful and Re-usable data (OURdata) Index (3 editions)</p> <p>6) Open Government Development Index (OGDI) (2 editions)</p> <p>These data shapes the third context - open data indices and rankings. The second sheet of this file covers countries covered by this study, namely, Austria, Belgium, Czech Republic, Italy, Latvia, Poland, Serbia, Sweden. It serves the basis for Section 4.2 of the paper.</p> <p>Based on the analysis of selected countries, incl. the analysis of their specifics and performance over the years in the indices and benchmarks, covering 57 editions of OGD-oriented reports and indices and e-government-related reports (2013-2022) that shaped a protocol (see paper, Annex 1), 102 patterns that may lead to disparities and divides in the development and benchmarking of ODEs were identified, which after the assessment by expert panel were reduced to a final number of 94 patterns representing four contexts, from which the recommendations defined in the paper were obtained. These patterns are available in the file "OGDdevelopmentPatterns". The first sheet contains the list of patterns, while the second sheet - the list of patterns and their effect as assessed by expert panel.</p> <p>***Format of the file***<br>.xls, .csv (for the first spreadsheet only)</p> <p>***Licenses or restrictions***<br>CC-BY</p> <p> </p> <p>For more info, see README.txt<br> </p>
zbMATHOpenRec: A Gold Standard Dataset for Recommending Scientific Documents with Mathematical Content
<p> </p> <p>Here we include the first gold standard dataset for recommending scientific documents with mathematical content. </p> <p><strong>Contents: </strong></p> <p>As of Feb-2023, there are 421 recommendation pairs with 80 seed documents.</p> <ol> <li>All recommendation pairs are available: recommendationPairs.csv</li> <li>Each document's contents, such as title, abstract/review/summary, authors, MSC codes, Full-text link, references, etc. are available in: documentContents.csv</li> </ol> <p><strong>Dataset construction process</strong>:</p> <p>This is the first gold standard content-based RS dataset, consisting of 421 scientific research entry recommendation pairs with mathematical content. The purpose is to enable math in scientific documents for document recommendations, meaning if two documents have similar math content, one could be recommended to the other. </p> <p>To create this dataset, we analyzed 4.5 million research entires from zbMATH Open (https://zbmath.org/) and performed the following steps to obtain the final dataset:</p> <ol> <li>We selected 80 seeds that capture the most word and math tokens in zbMATH Open using statistical measures.</li> <li>Three experts, one with several years of experience reviewing research entries in mathematics, curated the recommendations for 80 seeds.</li> </ol> <p>Using this dataset, researchers can accelerate the development and testing of recommendation approaches for scientific literature with mathematical content, improving recommendations for the STEM fields where mathematical content is currently being ignored</p> <p>## License </p> <p>Legal restrictions and copyright: The zbMATH Open data is subject to the Terms and Conditions for the zbMATH Open API Service of FIZ Karlsruhe – Leibniz-Institut für Informationsinfrastruktur GmbH. Content generated by zbMATH Open, such as reviews, classifications, software, or author disambiguation data, are distributed under CC-BY-SA 4.0. This defines the license for the whole dataset, which also contains non-copyrighted bibliographic metadata and reference data derived from I4OSC (CC0).</p>
GERONTE H2020 project - GERDAT004 - Dataset of self-management recommendations
<p><strong>The present document is a dataset generated as part of Deliverable D1.1. of the GERONTE project, which has received funding from the European Union’s Horizon 2020 Programme under Grant Agreement N°945218. It aims to provide the geriatric oncology professional community with a dataset of self-management recommendations that can be used in the care for older patients with cancer and multimorbidity.</strong></p> <p>GERONTE is a 5-year research and innovation project (April 2021 to Mars 2026) funded by the European Union within the framework of the H2020 Research and Innovation programme, in response to the health societal challenge topic SC1-BHC-24-2020 “Healthcare interventions for the management of the elderly multimorbid patient”. The overall aim of GERONTE is to improve quality of life - defined as well-being on three levels: global health status, physical functioning and social functioning- for older multimorbid patients, while reducing overall costs of care. To this end, GERONTE will co-design, test, and prepare for deployment an innovative cost-effective patient-centred holistic health management system, hereafter referred to as the GERONTE intervention. GERONTE intervention will rely on an ICT based application for real-time collection and integration of standardised clinical and home patient-reported data. GERONTE intervention will be demonstrated in the context of care of multimorbid patients having cancer as a dominant morbidity, and be adaptable to any other combination of morbidities.</p> <p>Patient empowerment by supporting self-management is an important component of the Geronte care pathway. As patients will be monitoring themselves at home, to register side-effects of treatment, decompensation of comorbidities and signs of functional decline, they will also be faced with questions about how to deal with the issues that they are having. While one important component of the care pathway is early signalling of complications to allow for early intervention by health care professionals, there is also a lot that patients can do for themselves at home to enhance their life-style, decrease burden of signs and symptoms or to improve outcomes.</p> <p>This led to the composition of a dataset to be included in the GERONTE care pathway, which is presented here.</p>
Community Established Best Practice Recommendations for Tephra Studies-from Collection through Analysis
<p>Tephra is a unique volcanic product with an unparalleled role in understanding past eruptions, long-term behavior of volcanoes, and the effects of volcanism on climate and the environment. Tephra deposits also provide spatially widespread, extremely high-resolution time-stratigraphic markers across a range of sedimentary settings and are used in a range of disciplines (e.g., volcanology, climate science, archaeology, ecology, and impact assessment). Nonetheless, the study of tephra deposits is challenged by a lack of standardization that often inhibits data integration across geographic regions and across disciplines.</p> <p>Here we present comprehensive recommendations for tephra data gathering and reporting that were developed by the tephra science community to serve as guidelines for future investigators and to ensure that sufficient data are gathered for transparency and interoperability. Recommendations include standardized field and laboratory data collection along with reporting and correlation guidance. These are organized as tabulated lists of key metadata with their definition and purpose. They are system independent and usable for template, tool, and database development. This new standardized framework promotes consistent tephra documentation and archiving, fosters interdisciplinary communication, and improves effectiveness of data sharing among diverse communities of researchers. Wider adoption will help to expand the applicability and usability of tephra data and facilitate scientific collaboration and data reuse.</p> <p>For additional details, see the accompanying manuscript:</p> <p>Wallace, K.*, Bursik, M. Kuehn, S., Kurbatov, A., Abbott, P., Bonadonna, C., Cashman, K., Davies, S., Jensen, B., Lane, C., Plunkett, G., Smith, V. Tomlinson, E., Thordarsson, T., and Walker, D. Community established best practice recommendations for tephra studies—from collection through analysis. <em>Sci Data</em> <strong>9, </strong>447 (2022). <a href="https://doi.org/10.1038/s41597-022-01515-y">https://doi.org/10.1038/s41597-022-01515-y</a></p> <p>*corresponding author: Kristi Wallace, <a href="mailto:kwallace@usgs.gov">kwallace@usgs.gov</a></p> <p>Open access article is available online here <a href="https://doi.org/10.1038/s41597-022-01515-y">https://doi.org/10.1038/s41597-022-01515-y</a> or as a PDF here <a href="https://gcc02.safelinks.protection.outlook.com/?url=https%3A%2F%2Fwww.nature.com%2Farticles%2Fs41597-022-01515-y.pdf&data=05%7C01%7Ckwallace%40usgs.gov%7C673f9f39fd3e4122dd9b08da6f3b9667%7C0693b5ba4b184d7b9341f32f400a5494%7C0%7C0%7C637944598967375940%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C3000%7C%7C%7C&sdata=%2BKVfwK2FbUKAoJf2gMerCmBMEQE1rvMDkS6xIk3DGKY%3D&reserved=0">https://www.nature.com/articles/s41597-022-01515-y.pdf</a>.</p>
Libraries & Recommended Citations for using PLAsTiCC Models
<p>Text file libraries for transient and variable source models used in the "Photometric LSST Astronomical Time-Series Classification Challenge" (PLAsTiCC). The original challenge (Sep 28, 2018 - Dec 17, 2018) was hosted at https://www.kaggle.com/c/PLAsTiCC-2018. See AAA_README.pdf for more information.</p>
Dataset: "Balancing consumer and business value of recommender systems: A simulation-based analysis"
<p>The data files in this directory contain to the results of the simulations reported in the paper: "Balancing Consumer and Business Value of Recommender Systems: A Simulation-based Analysis" published in Electronic Commerce Research and Applications. The paper is available here: <a href="https://doi.org/10.1016/j.elerap.2022.101195">https://doi.org/10.1016/j.elerap.2022.101195</a></p> <p> </p>
Context-based Entity Recommendation on Real-Life Knowledge Work in Context (RLKWiC dataset)
<h2><a href="../records/11059573">RLKWiC</a> Add-on: Benchmarking Dataset for Entity Recommendation</h2> <p>This benchmark, built on top of the Real-Life Knowledge Work in Context (<a href="../records/11059573">RLKWiC</a>) dataset, is designed to evaluate context-based entity recommendation by simulating a scenario where participants receive entities extracted from their activities across their defined contexts. </p> <p>In total, 1850 entity recommendations were generated across 56 contexts. After deduplication, these entities were presented to participants for explicit relevance assessment on a 3-point scale:</p> <ul> <li>0 [Irrelevant]: Signifying a lack of relevance between the recommended entity and the context.</li> <li>1 [Relevant] Denoting a connection between the entity and the context, although it may not fully represent it.</li> <li>2 [Representative]: The entity closely aligns with the context, indicating a high level of relevance where the context can be inferred to be about this entity.</li> </ul> <p>Participants could also suggest additional relevant entities. The resulting dataset comprises 1067 entities with explicit relevance scores, offering a resource for benchmarking entity recommendation in real-life knowledge work.</p> <h3><strong>Paper: </strong><a href="https://dl.acm.org/doi/10.1145/3640457.3688068" target="_blank" rel="noopener">Context-based Entity Recommendation for Knowledge Workers: Establishing a Benchmark on Real-life Data</a></h3>
Experimental result to investigate the influence of user's tweets and diversification on serendipitous research paper recommendations
<p>This is a raw dataset of the experiment result to investigate the influence of user's tweets and diversification on serendipitous research paper recommendations.</p> <p> </p>
Takeout Recommendation Dataset (TRD) from Meituan Takeout app
<p>This is a takeout recommendation dataset (TRD) which contains a vast amount of meta information from Meituan Takeout app. We collect orders from 11 commercial districts in Beijing between March 1st and March 28th, 2021. The first three weeks of orders are as training, while the last week is used for testing to avoid data leakage. We briefly summarize each file as follows and for more details, please refer to README.md.</p> <p>1. users.txt (attributes of all users)<br> 2. pois.txt (attributes of all takeout restaurants)<br> 3. spus.txt (attributes of all food)<br> 4. orders_poi_session.txt (a sequence of restaurants clicked by user before ordering)<br> 5. orders_spu_train.txt (order-food in training set)<br> 6. orders_train.txt (order-restaurant in training set)<br> 7. orders_test.txt (order-restaurant in training set)<br> 8. orders_poi_test_label.txt (test labels of order-restaurant)<br> 9. orders_spu_test_label.txt (test labels of order-food)</p> <p>10 graph.bin (graph in DGL format)</p> <p>graph.bin is build by above *.txt files, there is a vast amount of meta informarion on nodes and edges. just several codes can load this graph with 18,931,400 edges and 408,849 nodes:</p> <pre><code class="language-python">from dgl import load_graphs #should install dgl ds,_ = load_graphs("./graph.bin") g = ds[0] print(g)</code></pre> <p> </p>
Cold rolling mill: Dataset for Recommender system for process optimization
<p>The dataset (pickle formatted with version pickle=4.0) contains a collection of process values obtained from a process line. Each value represents the mean measurement within a window at specific intervals along the distance domain. These intervals were equidistant and sampled under steady-state conditions, ensuring consistent data collection.</p> <p>The process values included in the dataset cover a range of parameters and variables relevant to the process line. These values provide information about the behavior and characteristics of the process at different points along the distance domain.</p> <p> </p> <p>Associate source code is available at: <a href="https://github.com/CuAuPro/opti-rec-sys">Recommender system for process optimization (github.com)</a>.</p>
Recommendation dataset for Cultural Heritage
<p>A set of 24 JSON files that each represents a recommendation category and contains a set of items with all necessary information as extracted by the <a href="https://cuhe.in-two.com/">CUHE</a> platform. For each of the items this file also contains the apiResponse element, which corresponds to the complete API response from <a href="https://www.europeana.eu/">Europeana</a>. The Europeana API description is available <a href="https://pro.europeana.eu/page/intro">here</a>, the apiResponse follows the EDM (Europeana Data Model) specification which is available <a href="https://pro.europeana.eu/page/intro#edm">here</a>.</p> <p>The <a href="https://cuhe.in-two.com/">CUHE project</a> has indirectly received funding from the European Union's Horizon 2020 research and innovation action programme, via the AI4Media Open Call #1 issued and executed under the <a href="https://www.ai4media.eu/">AI4Media project</a> (Grant Agreement no. 951911).</p>
Dataset: Music Industry Professionals' Perspectives on Music Streaming Services and Recommendation
<p><strong>Questionnaire response data set</strong><br> Here, we include the data retrieved from participants at Eurosonic Noorderslag 2023, as described in the paper cited above.<br> When using, analyzing, or publishing this data in any way, please make sure to attribute it to the authors and cite it accordingly.<br> <br> We include the data in .xlsx, .csv format (semicolon-separated, and .tsv format (tab-separated). We suggest using the Excel file, as its layout makes it more easily readable.<br> <br> The complete question list as used in the questionnaire is published separately on <a href="https://doi.org/10.5281/zenodo.8121151">https://doi.org/10.5281/zenodo.8121151</a>.<br> <br> <strong>Paper title</strong><br> Looking at the FAccTs: Exploring Music Industry Professionals’ Perspectives on Music Streaming Services and Recommendations<br> <br> <strong>Paper abstract</strong><br> Music recommender systems, commonly integrated into streaming services, help listeners find music. Previous research on such systems has focused on providing the best possible recommendations for these services' consumers, as well as on fairness for artists who release their music on streaming services. While those insights are imperative, another group of stakeholders has been omitted so far: the many other professionals working in the music industry. They, too, are (in)directly affected by music streaming services. Therefore, this work explores the perspective of music industry professionals. We present a study that addresses the role of streaming services and recommender systems in their jobs. Results indicate this role is significant. Furthermore, participants feel that music recommender systems lack transparency and are insufficiently controllable, for both customers and artists. Finally, participants desire that music streaming services take charge of increasing recommendation diversity, and variety in consumers' listening behavior and taste.</p> <p><strong>Citation</strong><br> Karlijn Dinnissen, Isabella Saccardi, Marloes Vredenborg, and Christine Bauer. 2023. Looking at the FAccTs: Exploring Music Industry Professionals’ Perspectives on Music Streaming Services and Recommendations. In 2nd International Conference of the ACM Greek SIGCHI Chapter (CHIGREECE 2023), September 27–28, 2023, Athens, Greece. ACM, New York, NY, USA, 5 pages. https://doi.org/10.1145/3609987.3610011</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.