Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

48

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

48 results for “recommender systems”

Learn how ShareScore rates datasets ↗
zenodo52/100

Pubmed Journal Recommendation System dataset

<p>Dataset for Journal recommendation, includes title, abstract, keywords, and journal.</p> <p>We extracted the journals and more information of:</p> <p>Jiasheng Sheng. (2022). PubMed-OA-Extraction-dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6330817.</p> <p>Dataset Components:</p> <ul> <li> <p><strong>data_pubmed_all:</strong> This dataset encompasses all articles, each containing the following columns: 'pubmed_id', 'title', 'keywords', 'journal', 'abstract', 'conclusions', 'methods', 'results', 'copyrights', 'doi', 'publication_date', 'authors', 'AKE_pubmed_id', 'AKE_pubmed_title', 'AKE_abstract', 'AKE_keywords', 'File_Name'.</p> </li> <li> <p><strong>data_pubmed:</strong> To focus on recent and relevant publications, we have filtered this dataset to include articles published within the last five years, from January 1, 2018, to December 13, 2022&mdash;the latest date in the dataset. Additionally, we have exclusively retained journals with more than 200 published articles, resulting in 262,870 articles from 469 different journals.</p> </li> <li> <p><strong>data_pubmed_train, data_pubmed_val, and data_pubmed_test:</strong> For machine learning and model development purposes, we have partitioned the 'data_pubmed' dataset into three subsets&mdash;training, validation, and test&mdash;using a random 60/20/20 split ratio. Notably, this division was performed on a per-journal basis, ensuring that each journal's articles are proportionally represented in the training (60%), validation (20%), and test (20%) sets. The resulting partitions consist of 157,540 articles in the training set, 52,571 articles in the validation set, and 52,759 articles in the test set.</p> </li> </ul>

opencc-by-4.0Oct 2023View details →
zenodo44/100

GitRec - Github Project Recommender Systems

<p>This dataset contains the data collected using the Google API for the GHTorrent project and which were applied in the doctoral thesis directed to recommending projects on the GitHub platform</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Dataset: "Balancing consumer and business value of recommender systems: A simulation-based analysis"

<p>The data files in this directory contain to the results of the simulations reported in the paper: &quot;Balancing Consumer and Business Value of Recommender Systems: A Simulation-based Analysis&quot; published in Electronic Commerce Research and Applications. The paper is available here:&nbsp;<a href="https://doi.org/10.1016/j.elerap.2022.101195">https://doi.org/10.1016/j.elerap.2022.101195</a></p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Cold rolling mill: Dataset for Recommender system for process optimization

<p>The dataset (pickle formatted with version pickle=4.0)&nbsp;contains a collection of process values obtained from a process line. Each value represents the mean measurement within a window at specific intervals along the distance domain. These intervals were equidistant and sampled under steady-state conditions, ensuring consistent data collection.</p> <p>The process values included in the dataset cover a range of parameters and variables relevant to the process line. These values provide information about the behavior and characteristics of the process at different points along the distance domain.</p> <p>&nbsp;</p> <p>Associate source code is available at:&nbsp;<a href="https://github.com/CuAuPro/opti-rec-sys">Recommender system for process optimization&nbsp;(github.com)</a>.</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Dataset for Application of Recommender Systems on Police Photo Lineup Assembling task

<p>For more information about the Recommender Systems for Police Photo Lineup project, please visit: http://www.ksi.mff.cuni.cz/~peska/lineup/paper.pdf</p> <p>The dataset consists of two parts:</p> <p><strong>The raw dataset </strong>contains visual and attribute-based desriptors of the list of candidate persons:</p> <p>- personsCB_IDs.csv: ordered list of persons IDs<br> - personsData.csv: raw attribute features of the persons<br> - personsCBVectors.csv: derived binary attribute-features with TF-IDF applied. The ordering of records is the same as in personsCB_IDs.csv<br> - personsVectors.csv: derived visual descriptors of persons' images. Probability layer of VGG-Face CNN was applied for this task. The ordering of records is the same as in personsCB_IDs.csv.</p> <p><strong>The implicit feedback dataset</strong> (feedbackDatasetCB-RSFeatures.csv) contains information received from the user-study on lineup assembling task performed by seven domain experts. This table has following structure:</p> <p>- evaluatorID;<br> - lineupID (id of the suspect);<br> - candidateID (person recommended to this particular suspect);<br> - calculated content-based similarity<br> - selection (1 if evaluator selected this candidate, 0 otherwise);<br> - substraction of CB features of the suspect and the candidate (this can be simply modified by accessing raw dataset to contain visual descriptors or any other combination of features)</p> <p> </p>

opencc-by-4.0Jun 2017View details →
zenodo40/100

Empowering Coffee Farming Using Counterfactual Recommendation based RNN-IoT Integrated Soil Fertility Control System

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Case Base for Fertilizer Recommender System

<p>Dataset that contains recommendations for NPK fertilizer rates for different cases of climatic and soil conditions in coffee crops in the Cauca region in Colombia.</p> <p>The variables of this dataset are described below:</p> <p>Density: Represents the planting density of plants in a coffee crop, that is, the number of plants per hectare.</p> <p>Shadow coverage: Indicates the shade per hectare of the coffee crop, measured on a scale from 0 to 1.</p> <p>Season: Refers to the climatic season to which the year in which fertilization is to be carried out was classified. 'Seca' represents the dry season, 'Lluviosa' represents the rainy season and 'Normal' represents the normal season.</p> <p>Humidity: Represents the percentage of moisture in the soil, measured on scales from 0 to 1.</p> <p>N: It is the level of N in the soil, measured in mg/Kg.</p> <p>P: It is the level of P in the soil, measured in mg/Kg.</p> <p>K: It is the level of K in the soil, measured in mg/Kg.</p> <p>pH: It is the pH level in the soil.</p> <p>Cond_N: It is the rate of nitrogenous fertilizer to recommend.</p> <p>Cond_P: It is the rate of phosphorus fertilizer to recommend.</p> <p>Cond_K: It is the rate of potassium fertilizer to recommend.</p> <p>Note: Fertilizer rates are represented as follows: 1 is low rate, 2 is normal rate, 3 is high rate and 4 is very high rate. The fertilizer rate depends on the sowing density, this is indicated in the corresponding research article.</p> <p>Plus_A and Plus_B are additional recommendations, regarding irrigation and soil and crop care.</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Recommender Systems for Science: A basic Taxonomy

<p>This dataset is accompanying the &quot;<strong>Recommender system for science: A basic taxonomy</strong>&quot; paper published at IRCDL 2022 conference.&nbsp;</p> <p>This study had a Systematic Mapping Approach on the Recommender system for science. In particular, the study aims at responding to four questions on recommender systems in science cases: users and their interests representation, item typologies and their representation, recommendation algorithms, and evaluation, and then providing a taxonomy.&nbsp;</p> <p>This dataset contains&nbsp;<strong>209 papers&nbsp;</strong>of interest that have been published between 2015 and 2022.</p> <p>The dataset has <strong>11</strong> columns which organised as follows:&nbsp;</p> <p>Column&nbsp;<strong>Title:&nbsp;</strong>This column contains the title of the papers.</p> <p>Column&nbsp;<strong>DOI:&nbsp;</strong>This column contains the DOI of the papers.</p> <p>Column&nbsp;<strong>Publication_year</strong>: This column contains the year that the paper is published.</p> <p>Column&nbsp;<strong>DB:&nbsp;</strong>This column contains the repository that the paper is retrieved.</p> <p>Column&nbsp;<strong>Keywords</strong>: This column contains the keywords provided for the paper.</p> <p>Column&nbsp;<strong>Content_type:&nbsp;</strong>This column contains the paper type which can be:&nbsp;<strong>Article,</strong>&nbsp;<strong>Conference</strong>&nbsp;or&nbsp;<strong>Review.</strong></p> <p>Column&nbsp;<strong>Citing_paper_count:&nbsp;</strong>This column contains the citation number of the paper.</p> <p>Column&nbsp;<strong>Recommended_artefact:&nbsp;</strong>This column contains the scientific product that is recommended to users which can be <strong>paper</strong>, <strong>workflow</strong>, <strong>collaborator</strong>, <strong>dataset</strong> or <strong>others</strong>.</p> <p>Column&nbsp;<strong>User_type:&nbsp;</strong>This column contains the type of user who receives the recommendation, which can be&nbsp;an<strong> Individual </strong>user&nbsp;or&nbsp;a<strong> Group</strong>&nbsp;of users.</p> <p>Column&nbsp;<strong>Algorithm</strong><strong>:&nbsp;</strong>This column contains the recommendation algorithm that the paper proposed, which can be:&nbsp;<strong>HB&nbsp;</strong>(Hybrid-based),&nbsp;<strong>CB</strong>&nbsp;(Content-based),&nbsp;<strong>CFB</strong>&nbsp;(Collaborative-filtering-based), or&nbsp;<strong>GB</strong>&nbsp;(Graph-based).</p> <p>Column&nbsp;<strong>Evaluation_method</strong><strong>:&nbsp;</strong>This column contains the method of the algorithm evaluation which can be&nbsp;<strong>OFFLINE</strong>,&nbsp;<strong>ONLINE, BOTH,&nbsp;</strong>or<strong>&nbsp;NO_EVALUATION.</strong></p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

#nowplaying-RS: A New Benchmark Dataset for Building Context-Aware Music Recommender Systems

<p>Music recommender systems can offer users personalized and contextualized recommendation and are therefore important for music information retrieval. An increasing number of datasets have been compiled to facilitate research on different topics, such as content-based, context-based or next-song recommendation. However, these topics are usually addressed separately using different datasets, due to the lack of a unified dataset that contains a large variety of feature types such as item features, user contexts, and timestamps. To address this issue, we propose a large-scale benchmark dataset called #nowplaying-RS, which contains 11.6 million music listening events (LEs) of 139K users and 346K tracks collected from Twitter. The dataset comes with a rich set of item content features and user context features, and the timestamps of the LEs. Moreover, some of the user context features imply the cultural origin of the users, and some others&mdash;like hashtags&mdash;give clues to the emotional state of a user underlying an LE. In this paper, we provide some statistics to give insight into the dataset, and some directions in which the dataset can be used for making music recommendation. We also provide standardized training and test sets for experimentation, and some baseline results obtained by using factorization machines.</p> <p>The dataset contains three files:</p> <ul> <li>user_track_hashtag_timestamp.csv contains basic information about each listening event. For each listening event, we provide an id, the user_id, track_id, hashtag, created_at&nbsp;</li> <li>context_content_features.csv: contains all context and content features. For each listening event, we provide the id of the event, user_id, track_id, artist_id, content features regarding the track mentioned in the event (instrumentalness, liveness, speechiness, danceability, valence, loudness, tempo, acousticness, energy, mode, key) and context features regarding the listening event (coordinates (as geoJSON), place (as geoJSON), geo (as geoJSON), tweet_language, created_at, user_lang, time_zone, entities contained in the tweet).</li> <li>sentiment_values.csv contains sentiment information for hashtags. It contains the hashtag itself and the sentiment values gathered via four different sentiment dictionaries: AFINN, Opinion Lexicon, Sentistrength Lexicon and vader. For each of these dictionaries we list the minimum, maximum, sum and average of all&nbsp;sentiments of the tokens of the hashtag (if available, else we list empty values). However, as most hashtags only consist of a single token, these&nbsp;values are equal in most cases. Please note that the lexica are rather diverse and therefore, are able to resolve very different terms against a score. Hence,&nbsp;the resulting csv is rather sparse. The file contains the following comma-separated values: &lt;hashtag, vader_min, vader_max, vader_sum,vader_avg, &nbsp;afinn_min, afinn_max,&nbsp;afinn_sum, afinn_avg, ol_min, ol_max, ol_sum, ol_avg, ss_min, ss_max, ss_sum, ss_avg &gt;, where we abbreviate all scores gathered over the Opinion Lexicon with the&nbsp;prefix &#39;ol&#39;. Similarly, &#39;ss&#39; stands for SentiStrength.&nbsp;</li> </ul> <p>Please also find the training and test-splits for the dataset in this repo. Also, prototypical implementations of a context-aware recommender system based on the dataset can be found at&nbsp; <a href="https://github.com/asmitapoddar/nowplaying-RS-Music-Reco-FM">https://github.com/asmitapoddar/nowplaying-RS-Music-Reco-FM</a>.</p> <p>If you make use of this dataset, please cite the following paper where we describe and experiment with the dataset:</p> <p>@inproceedings{smc18,<br> title = {#nowplaying-RS: A New Benchmark Dataset for Building Context-Aware Music Recommender Systems},<br> author = {Asmita Poddar and Eva Zangerle and Yi-Hsuan Yang},<br> url = {http://mac.citi.sinica.edu.tw/~yang/pub/poddar18smc.pdf},<br> year = {2018},<br> date = {2018-07-04},<br> booktitle = {Proceedings of the 15th Sound &amp; Music Computing Conference},<br> address = {Limassol, Cyprus},<br> note = {code at https://github.com/asmitapoddar/nowplaying-RS-Music-Reco-FM},<br> tppubtype = {inproceedings}<br> }</p>

opencc-by-4.0Jul 2018View details →
zenodo40/100

Datasets from the KDD 2021 article "A Semi-Personalized System for User Cold Start Recommendation on Music Streaming Apps"

<p>We publicly release&nbsp;the anonymized&nbsp;<em>song_embeddings.parquet&nbsp; user_embeddings.parquet&nbsp; user_features_test.parquet&nbsp; user_features_train.parquet&nbsp; user_features_validation.parquet</em>&nbsp;datasets, with each of the&nbsp;TT-SVD or UT-ALS versions of embeddings, from the music streaming platform Deezer, as described in the&nbsp;article &quot;<em>A Semi-Personalized System for User Cold Start Recommendation on Music Streaming Apps&quot;</em>&nbsp;published in the proceedings of the 27TH ACM SIGKDD conference on knowledge discovery and data mining&nbsp;(<em>KDD 2021</em>). The paper is available&nbsp;<a href="https://arxiv.org/abs/2106.03819">here</a>.</p> <p>These datasets are used in the&nbsp;GitHub repository&nbsp;<a href="https://github.com/deezer/semi_perso_user_cold_start">deezer/semi_perso_user_cold_start</a>&nbsp;to reproduce experiments from the article.</p> <p>Please cite our paper if you use our code or data in your work.</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

GUI evaluation data for an IDE command recommender system

<p>This dataset contains results of the study conducted among the participants of the XP 2016 (a scientific conference with a strong participation of practitioners from the industry). The objective of the study was to evaluate the acceptance and usability of the proposed Graphical User Interface (GUI) for an Integrated Development Environment (IDE) command recommender system (RS). The data was collected by the questionnaire and the interviews. The data is anonymized.</p> <p>Content:</p> <ul> <li>README.txt</li> <li>./Survey answers.csv - the questionnaire answers</li> <li>./Interviews <ul> <li>./interviewXXX.txt - a file with a transcribed interview</li> <li>./mapping-codes-to-primary-documents.csv - a binary table summarizing interviews</li> </ul> </li> </ul>

opencc-by-4.0May 2017View details →
zenodo36/100

Dataset Online grocery shopping recommender systems: common approaches and practices

<p>Standardized Excel form for data extraction. Extraction criteria were defined based on the sub-research questions that are provided in the paper.&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

PODCAST: The Impact of News Recommender Systems on our Personal Identity

<p>This 45-minute podcast is a result of qualitative research which aims to identify if we should be concerned that news recommender systems may have an impact on our personal identities over time. Ind sets out to achieve a number of objectives.</p> <p>1. To understand how news recommender systems influence the way of seeing and experiencing the world in which we anchor our identities<br> 2. To examine the extent to which news recommendations are transforming and shaping our preferences and behaviours<br> 3. To determine if we should be concerned about personal data privacy, as news recommendations technology collects our personal data to create algorithmically generated news recommendations<br> 4. To explore if we think that personalised news recommendations can serve to sharpen our focus and broaden our minds in parallel</p> <p>The Impact of News Recommender Systems on our Personal Identity&nbsp;is licensed under a Creative Commons License.</p>

opencc-by-4.0Aug 2019View details →
zenodo36/100

Datensatz für BA: Entwicklung eines Recommender-Systems für die Zuordnung von Anforderungen zu IT-Services

<ul> <li> <p><strong>afos_03_promise_nrf.csv</strong>: Diese Datei enth&auml;lt funktionale Anforderungen aus dem PROMISE-Datensatz. Die Abk&uuml;rzung &bdquo;nrf&ldquo; k&ouml;nnte sich auf spezifische Notwendigkeiten oder Funktionen innerhalb der PROMISE-Daten beziehen. Sie dient als Quelle f&uuml;r Anforderungen, die im Recommender-System genutzt werden k&ouml;nnen.</p> </li> <li> <p><strong>afos_tcs_sampled.csv</strong>: Diese Datei enth&auml;lt eine Stichprobe von funktionalen Anforderungen aus einem zus&auml;tzlichen Quellenbestand oder einem Datensatz mit technischen Anforderungen. Diese Anforderungen erg&auml;nzen den PROMISE-Datensatz und erweitern die Datenbasis f&uuml;r die Entwicklung und das Training des Recommender-Systems.</p> </li> <li> <p><strong>Goldstandard_256.csv</strong>: Diese Datei bildet den Goldstandard f&uuml;r das Recommender-System und enth&auml;lt 256 funktionale Anforderungen, die manuell den entsprechenden IT-Services zugeordnet wurden. Sie dient als Referenz zur Bewertung der Qualit&auml;t des Recommender-Systems und erm&ouml;glicht die Validierung und Optimierung des Modells.</p> </li> <li> <p><strong>ITSM_Set_with_Descriptions.xlsx</strong>: Diese Datei enth&auml;lt eine Sammlung von IT-Services mit detaillierten Beschreibungen. Die Tabelle umfasst Kategorien und Unterkategorien der IT-Services, die als Basis f&uuml;r die Empfehlungen des Recommender-Systems dienen. Diese Beschreibungen erm&ouml;glichen eine semantische Analyse und unterst&uuml;tzen das Modell bei der Zuordnung zu passenden Anforderungen.</p> </li> <li> <p><strong>Recommender_System.ipynb</strong>: Dies ist ein Jupyter-Notebook, das den Code zur Entwicklung, Training und Evaluierung des Recommender-Systems enth&auml;lt. Das Notebook fasst die Implementierung der Empfehlungslogik zusammen und stellt die Grundlage f&uuml;r die experimentellen Auswertungen dar.</p> </li> <li> <p><strong>requirements.txt</strong>: Diese Datei listet die Python-Bibliotheken und deren Versionen auf, die f&uuml;r das Recommender-System ben&ouml;tigt werden. Sie stellt sicher, dass alle ben&ouml;tigten Abh&auml;ngigkeiten f&uuml;r die erfolgreiche Ausf&uuml;hrung des Codes installiert sind.</p> </li> </ul>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Dataset used for "A Recommender System of Buggy App Checkers for App Store Moderators"

<p>This is the dataset used for paper:&nbsp;&quot;A Recommender System of Buggy App Checkers for App Store Moderators&quot;, published on the <em>International Conference on Mobile Software Engineering and Systems (MOBILESoft)</em>&nbsp;in 2015.<br> <br> <strong>Dataset Collection</strong><br> We built a dataset that consists of a random sample of <em><strong>Android app metadata</strong></em> and <em><strong>user reviews</strong></em> available on the <em>Google Play Store</em> on January and March&nbsp;2014.<br> Since the Google Play Store is continuously evolving (adding, removing and/or updating apps), we updated the dataset twice.<br> The dataset D1 contains available apps in the Google Play Store in January 2014.<br> Then, we created a new snapshot (D2) of the Google Play Store&nbsp;in March 2014.<br> <br> The apps belong to the 27 different categories&nbsp;defined by Google (at the time of writing the paper), and the 4 predefined subcategories (free, paid, new_free, and new_paid). For each category-subcategory pair (e.g. tools-free, tools-paid, sports-new_free, etc.), we collected a maximum of 500 samples, resulting in a&nbsp; median number of 1.978 apps per category.</p> <p>For&nbsp;each app, we retrieved the following metadata: <em>name, package, creator, version code, version name, number of downloads, size, upload date, star rating, star counting</em>, and the set of <em>permission requests</em>.<br> <br> In addition, for each app, we collected up to a maximum of the latest 500 reviews posted by users in the Google Play Store. For each review, we retrieved its metadata:<em> title, description, device,</em> and <em>version</em> of the app. None of these fields were mandatory, thus<br> several reviews lack some of these details.<br> From all the reviews attached to an app, we only considered the reviews associated with the latest version of the app &mdash;i.e., we discarded unversioned and old-versioned reviews. Thus, resulting in a corpus of 1,402,717 reviews (2014 Jan.).</p> <p>&nbsp;</p> <p><strong>Dataset Stats</strong><br> Some stats about the datasets:</p> <p>- <strong>D1</strong> (<em>Jan. 2014</em>) contains 38,781 apps requesting 7,826 different permissions, and 1,402,717 user reviews.</p> <p>- <strong>D2</strong> (<em>Mar. 2014</em>) contains 46,644 apps and 9,319 different permission requests, and 1,361,319 user reviews.</p> <p>Additional stats about the datasets are available&nbsp;<a href="https://sites.google.com/site/androidbuggyappcheckers">here</a>.<br> <br> <br> <strong>Dataset Description</strong><br> To store the dataset, we created a graph database with <a href="https://neo4j.com/">Neo4j</a>. This dataset therefore consists of a graph describing the apps as nodes and edges. &nbsp;We chose a graph database because the graph visualization helps to identify connections among data (e.g.,<br> clusters of apps sharing similar sets of permission requests).<br> <br> In particular, our dataset graph contains six&nbsp;<em>types of nodes</em>:<br> -&nbsp;<em>APP</em> nodes containing metadata of each app,<br> -&nbsp;<em>PERMISSION</em> nodes describing permission types,<br> -&nbsp;<em>CATEGORY</em> nodes describing app categories,<br> -&nbsp;<em>SUBCATEGORY</em> nodes describing app subcategories,<br> - <em>USER_REVIEW</em> nodes storing user reviews.<br> - <em>TOPIC</em> topics mined from user reviews (using LDA).<br> <br> Furthermore, there are five&nbsp;<em>types of relationships</em> between APP nodes and each of the remaining nodes:</p> <p>- <em>USES_PERMISSION</em> relationships between APP and PERMISSION nodes<br> - <em>HAS_REVIEW</em> between APP and USER_REVIEW nodes<br> - <em>HAS_TOPIC</em> between USER_REVIEW and TOPIC nodes<br> -&nbsp;<em>BELONGS_TO_CATEGORY</em> between APP and CATEGORY nodes<br> - <em>BELONGS_TO_SUBCATEGORY</em>&nbsp;between APP and SUBCATEGORY nodes</p> <p><br> <strong>Dataset Files Info</strong></p> <ul> <li><strong>Neo4j 2.0 Databases</strong> <ul> <li><em><strong>googlePlayDB1-Jan2014_neo4j_2_0.rar</strong></em></li> <li><em><strong>googlePlayDB2-Mar2014_neo4j_2_0.rar</strong></em><br> We provide two&nbsp;Neo4j databases containing the 2 snapshots of the Google Play Store (January and March 2014).&nbsp;These are the original databases created for the paper. The databases were created with <strong>Neo4j 2.0. </strong>In particular with the tool version&nbsp;<em>&#39;Neo4j 2.0.0-M06 Community Edition&#39; (latest version available at the time of implementing&nbsp;the paper in 2014).</em><br> &nbsp;</li> </ul> </li> <li><strong>Neo4j 3.5&nbsp;Databases</strong> <ul> <li><strong>googlePlayDB1-Jan2014_neo4j_3_5_28.rar</strong></li> <li><strong>googlePlayDB2-Mar2014_neo4j_3_5_28.rar</strong><br> Currently,&nbsp;the version Neo4j 2.0 is deprecated and it is not available for download in the official <em>Neo4j Download Center</em>. We have migrated the original databases (Neo4j 2.0) to Neo4j 3.5.28.<br> The databases can be opened with the tool version: <em>&#39;Neo4j Community Edition 3.5.28&#39;.<br> The tool can be downloaded from the official </em><a href="https://neo4j.com/download-center/#community">Neo4j Donwload</a>&nbsp;page.<br> <br> In order to open the databases with more recent versions of Neo4j, the databases must be first migrated to the corresponding version. Instructions about the migration process can be found in the&nbsp;<a href="https://neo4j.com/docs/upgrade-migration-guide/current/understanding-upgrades-migration/">Neo4j Migration Guide</a>.<br> <br> First time&nbsp;the Neo4j database is connected, it could request credentials. The username and pasword are: neo4j/neo4j&nbsp;<br> <br> &nbsp;</li> </ul> </li> </ul> <p>&nbsp;</p>

opencc-by-4.0May 2015View details →
dryad32/100

Data from: A clinical decision support system learned from data to personalize treatment recommendations towards preventing breast cancer metastasis

Objective: A Clinical Decision Support System (CDSS) that can amass Electronic Health Record (EHR) and other patient data holds promise to provide accurate classification and guide treatment choices. Our objective is to develop the Decision Support System for Making Personalized Assessments and Recommendations Concerning Breast Cancer Patients (DPAC), which is a CDSS learned from data that recommends the optimal treatment decisions based on a patient's features. Method: We developed a Bayesian network architecture called Causal Modeling with Internal Layers (CAMIL), and an algorithm called Treatment Feature Interactions (TFI), which learns from data the interactions needed in a CAMIL model. Using the TFI algorithm, we learned interactions for six treatments from the Lynn Sage Data Set (LSDS). We created a CAMIL model using these interactions, resulting in a DPAC which recommends treatments towards preventing 5-year breast cancer metastasis. Results: In a 5-fold cross-validation analysis, we compared the probability of being metastasis free in 5 years for patients who made decisions recommended by DPAC to those who did not. These probabilities are (the probability for those making the decisions appears first): chemotherapy (.938, .872); breast/chest wall radiation (.939, .902); nodal field radiation (.940, .784); antihormone (.941, .906); HER2 inhibitors (.934, .880); neadjuvant therapy (.931, .837). In an application of DPAC to the independent METABRIC dataset, the probabilities for chemotherapy were (.845, .788). Discussion: Patients who took the advice of DPAC had, as a group, notably better outcomes than those who did not. We conclude that DPAC is effective at amassing and analyzing data towards treatment recommendations. Some of the findings in DPAC are controversial. For example, DPAC says that chemotherapy increases the chances of metastasis for many node negative patients. This controversy shows the importance of developing a conclusive version of DPAC to ensure we provide patients with the best patient-specific treatment recommendations.

opencc-zeroDec 2018View details →
zenodo32/100

Recommender Systems and AI Techniques in E-commerce: An Analysis of Trends and the Research Agenda

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →
zenodo32/100

Empowering Coffee Farming Using Counterfactual Recommendation based RNN-IoT Integrated Soil Fertility Control System

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo32/100

Game-Shapley recommender system demonstration

<p>Recommender system using Game-Shapley algorithm</p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

Top 250 IMDb Movies Dataset for Recommendation Systems

<p>Dataset obtenido en la pr&aacute;ctica 1 de la asignatura "Tipolog&iacute;a y ciclo de vida de los datos", del M&aacute;ster en ciencia de datos de la UOC. Ha sido obtenido por Ignacio Gimeno Alonso y Morad Kharraz Senhaji.</p> <p>Los datos de este dataset han sido extra&iacute;dos de la lista de las 250 pel&iacute;culas mejor valoradas presente en la web de IMDb (https://www.imdb.com/chart/top/?ref_=nv_mv_250)</p> <p>El dataset contiene los siguientes campos:</p> <p>&middot; &nbsp; &nbsp; &nbsp; ranking: Puesto de la pel&iacute;cula en la lista de las 250 mejor valoradas.</p> <p>&middot; &nbsp; &nbsp; &nbsp; nombre: T&iacute;tulo de la versi&oacute;n espa&ntilde;ola de la pel&iacute;cula.</p> <p>&middot; &nbsp; &nbsp; &nbsp; enlace: P&aacute;gina web de la pel&iacute;cula en <a href="http://www.imdb.com">www.imdb.com</a>.</p> <p>&middot; &nbsp; &nbsp; &nbsp; ano_lanz: A&ntilde;o de estreno de la pel&iacute;cula.</p> <p>&middot; &nbsp; &nbsp; &nbsp; duraci&oacute;n: Duraci&oacute;n de la pel&iacute;cula, en horas y minutos.</p> <p>&middot; &nbsp; &nbsp; edad: Clasificaci&oacute;n de edad. Puede estar en distintos formatos, seg&uacute;n el a&ntilde;o de estreno y el pa&iacute;s de producci&oacute;n (18, A, apta para mayores,...).</p> <p>&middot; &nbsp; &nbsp; &nbsp; rating: Puntuaci&oacute;n media dada por los usuarios de IMDb, de 0 a 10.</p> <p>&middot; &nbsp; &nbsp; &nbsp; num_votos: Cantidad de valoraciones que ha recibido la pel&iacute;cula.</p> <p>&middot;&nbsp; &nbsp; titulo_original: T&iacute;tulo original de la pel&iacute;cula. Si est&aacute; vac&iacute;o, significa que el t&iacute;tulo original coincide con el t&iacute;tulo en la versi&oacute;n espa&ntilde;ola.</p> <p>&middot; &nbsp; sinopsis: Resumen de la pel&iacute;cula en espa&ntilde;ol. Es un resumen corto, de unas pocas frases.</p> <p>&middot; &nbsp; &nbsp; &nbsp; genero: g&eacute;neros en los que se engloba la pel&iacute;cula, en ingl&eacute;s.</p> <p>&middot; &nbsp; &nbsp; &nbsp; direccion: Director o directores de la pel&iacute;cula.</p> <p>&middot; &nbsp; &nbsp; &nbsp; guionistas: Guionistas de la pel&iacute;cula.</p> <p>&middot; &nbsp; &nbsp; &nbsp; elenco: Actores / actrices principales de la pel&iacute;cula.</p> <p>Los datos contenidos en el dataset est&aacute;n referidos a pel&iacute;culas desde 1921 hasta 2024, pero las valoraciones est&aacute;n referidas al momento de recolecci&oacute;n de los datos (octubre-noviembre de 2024).</p> <p>&nbsp;</p>

opencc-by-nc-sa-2.0Nov 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record