Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
48
datasets available to search
ShareScore release 0.7.1
Dataset results
48 results for “recommender systems”
Pubmed Journal Recommendation System dataset
<p>Dataset for Journal recommendation, includes title, abstract, keywords, and journal.</p> <p>We extracted the journals and more information of:</p> <p>Jiasheng Sheng. (2022). PubMed-OA-Extraction-dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6330817.</p> <p>Dataset Components:</p> <ul> <li> <p><strong>data_pubmed_all:</strong> This dataset encompasses all articles, each containing the following columns: 'pubmed_id', 'title', 'keywords', 'journal', 'abstract', 'conclusions', 'methods', 'results', 'copyrights', 'doi', 'publication_date', 'authors', 'AKE_pubmed_id', 'AKE_pubmed_title', 'AKE_abstract', 'AKE_keywords', 'File_Name'.</p> </li> <li> <p><strong>data_pubmed:</strong> To focus on recent and relevant publications, we have filtered this dataset to include articles published within the last five years, from January 1, 2018, to December 13, 2022—the latest date in the dataset. Additionally, we have exclusively retained journals with more than 200 published articles, resulting in 262,870 articles from 469 different journals.</p> </li> <li> <p><strong>data_pubmed_train, data_pubmed_val, and data_pubmed_test:</strong> For machine learning and model development purposes, we have partitioned the 'data_pubmed' dataset into three subsets—training, validation, and test—using a random 60/20/20 split ratio. Notably, this division was performed on a per-journal basis, ensuring that each journal's articles are proportionally represented in the training (60%), validation (20%), and test (20%) sets. The resulting partitions consist of 157,540 articles in the training set, 52,571 articles in the validation set, and 52,759 articles in the test set.</p> </li> </ul>
GitRec - Github Project Recommender Systems
<p>This dataset contains the data collected using the Google API for the GHTorrent project and which were applied in the doctoral thesis directed to recommending projects on the GitHub platform</p>
Dataset: "Balancing consumer and business value of recommender systems: A simulation-based analysis"
<p>The data files in this directory contain to the results of the simulations reported in the paper: "Balancing Consumer and Business Value of Recommender Systems: A Simulation-based Analysis" published in Electronic Commerce Research and Applications. The paper is available here: <a href="https://doi.org/10.1016/j.elerap.2022.101195">https://doi.org/10.1016/j.elerap.2022.101195</a></p> <p> </p>
Cold rolling mill: Dataset for Recommender system for process optimization
<p>The dataset (pickle formatted with version pickle=4.0) contains a collection of process values obtained from a process line. Each value represents the mean measurement within a window at specific intervals along the distance domain. These intervals were equidistant and sampled under steady-state conditions, ensuring consistent data collection.</p> <p>The process values included in the dataset cover a range of parameters and variables relevant to the process line. These values provide information about the behavior and characteristics of the process at different points along the distance domain.</p> <p> </p> <p>Associate source code is available at: <a href="https://github.com/CuAuPro/opti-rec-sys">Recommender system for process optimization (github.com)</a>.</p>
Dataset for Application of Recommender Systems on Police Photo Lineup Assembling task
<p>For more information about the Recommender Systems for Police Photo Lineup project, please visit: http://www.ksi.mff.cuni.cz/~peska/lineup/paper.pdf</p> <p>The dataset consists of two parts:</p> <p><strong>The raw dataset </strong>contains visual and attribute-based desriptors of the list of candidate persons:</p> <p>- personsCB_IDs.csv: ordered list of persons IDs<br> - personsData.csv: raw attribute features of the persons<br> - personsCBVectors.csv: derived binary attribute-features with TF-IDF applied. The ordering of records is the same as in personsCB_IDs.csv<br> - personsVectors.csv: derived visual descriptors of persons' images. Probability layer of VGG-Face CNN was applied for this task. The ordering of records is the same as in personsCB_IDs.csv.</p> <p><strong>The implicit feedback dataset</strong> (feedbackDatasetCB-RSFeatures.csv) contains information received from the user-study on lineup assembling task performed by seven domain experts. This table has following structure:</p> <p>- evaluatorID;<br> - lineupID (id of the suspect);<br> - candidateID (person recommended to this particular suspect);<br> - calculated content-based similarity<br> - selection (1 if evaluator selected this candidate, 0 otherwise);<br> - substraction of CB features of the suspect and the candidate (this can be simply modified by accessing raw dataset to contain visual descriptors or any other combination of features)</p> <p> </p>
Empowering Coffee Farming Using Counterfactual Recommendation based RNN-IoT Integrated Soil Fertility Control System
Open the record for dataset details and reuse information.
Case Base for Fertilizer Recommender System
<p>Dataset that contains recommendations for NPK fertilizer rates for different cases of climatic and soil conditions in coffee crops in the Cauca region in Colombia.</p> <p>The variables of this dataset are described below:</p> <p>Density: Represents the planting density of plants in a coffee crop, that is, the number of plants per hectare.</p> <p>Shadow coverage: Indicates the shade per hectare of the coffee crop, measured on a scale from 0 to 1.</p> <p>Season: Refers to the climatic season to which the year in which fertilization is to be carried out was classified. 'Seca' represents the dry season, 'Lluviosa' represents the rainy season and 'Normal' represents the normal season.</p> <p>Humidity: Represents the percentage of moisture in the soil, measured on scales from 0 to 1.</p> <p>N: It is the level of N in the soil, measured in mg/Kg.</p> <p>P: It is the level of P in the soil, measured in mg/Kg.</p> <p>K: It is the level of K in the soil, measured in mg/Kg.</p> <p>pH: It is the pH level in the soil.</p> <p>Cond_N: It is the rate of nitrogenous fertilizer to recommend.</p> <p>Cond_P: It is the rate of phosphorus fertilizer to recommend.</p> <p>Cond_K: It is the rate of potassium fertilizer to recommend.</p> <p>Note: Fertilizer rates are represented as follows: 1 is low rate, 2 is normal rate, 3 is high rate and 4 is very high rate. The fertilizer rate depends on the sowing density, this is indicated in the corresponding research article.</p> <p>Plus_A and Plus_B are additional recommendations, regarding irrigation and soil and crop care.</p>
Recommender Systems for Science: A basic Taxonomy
<p>This dataset is accompanying the "<strong>Recommender system for science: A basic taxonomy</strong>" paper published at IRCDL 2022 conference. </p> <p>This study had a Systematic Mapping Approach on the Recommender system for science. In particular, the study aims at responding to four questions on recommender systems in science cases: users and their interests representation, item typologies and their representation, recommendation algorithms, and evaluation, and then providing a taxonomy. </p> <p>This dataset contains <strong>209 papers </strong>of interest that have been published between 2015 and 2022.</p> <p>The dataset has <strong>11</strong> columns which organised as follows: </p> <p>Column <strong>Title: </strong>This column contains the title of the papers.</p> <p>Column <strong>DOI: </strong>This column contains the DOI of the papers.</p> <p>Column <strong>Publication_year</strong>: This column contains the year that the paper is published.</p> <p>Column <strong>DB: </strong>This column contains the repository that the paper is retrieved.</p> <p>Column <strong>Keywords</strong>: This column contains the keywords provided for the paper.</p> <p>Column <strong>Content_type: </strong>This column contains the paper type which can be: <strong>Article,</strong> <strong>Conference</strong> or <strong>Review.</strong></p> <p>Column <strong>Citing_paper_count: </strong>This column contains the citation number of the paper.</p> <p>Column <strong>Recommended_artefact: </strong>This column contains the scientific product that is recommended to users which can be <strong>paper</strong>, <strong>workflow</strong>, <strong>collaborator</strong>, <strong>dataset</strong> or <strong>others</strong>.</p> <p>Column <strong>User_type: </strong>This column contains the type of user who receives the recommendation, which can be an<strong> Individual </strong>user or a<strong> Group</strong> of users.</p> <p>Column <strong>Algorithm</strong><strong>: </strong>This column contains the recommendation algorithm that the paper proposed, which can be: <strong>HB </strong>(Hybrid-based), <strong>CB</strong> (Content-based), <strong>CFB</strong> (Collaborative-filtering-based), or <strong>GB</strong> (Graph-based).</p> <p>Column <strong>Evaluation_method</strong><strong>: </strong>This column contains the method of the algorithm evaluation which can be <strong>OFFLINE</strong>, <strong>ONLINE, BOTH, </strong>or<strong> NO_EVALUATION.</strong></p>
#nowplaying-RS: A New Benchmark Dataset for Building Context-Aware Music Recommender Systems
<p>Music recommender systems can offer users personalized and contextualized recommendation and are therefore important for music information retrieval. An increasing number of datasets have been compiled to facilitate research on different topics, such as content-based, context-based or next-song recommendation. However, these topics are usually addressed separately using different datasets, due to the lack of a unified dataset that contains a large variety of feature types such as item features, user contexts, and timestamps. To address this issue, we propose a large-scale benchmark dataset called #nowplaying-RS, which contains 11.6 million music listening events (LEs) of 139K users and 346K tracks collected from Twitter. The dataset comes with a rich set of item content features and user context features, and the timestamps of the LEs. Moreover, some of the user context features imply the cultural origin of the users, and some others—like hashtags—give clues to the emotional state of a user underlying an LE. In this paper, we provide some statistics to give insight into the dataset, and some directions in which the dataset can be used for making music recommendation. We also provide standardized training and test sets for experimentation, and some baseline results obtained by using factorization machines.</p> <p>The dataset contains three files:</p> <ul> <li>user_track_hashtag_timestamp.csv contains basic information about each listening event. For each listening event, we provide an id, the user_id, track_id, hashtag, created_at </li> <li>context_content_features.csv: contains all context and content features. For each listening event, we provide the id of the event, user_id, track_id, artist_id, content features regarding the track mentioned in the event (instrumentalness, liveness, speechiness, danceability, valence, loudness, tempo, acousticness, energy, mode, key) and context features regarding the listening event (coordinates (as geoJSON), place (as geoJSON), geo (as geoJSON), tweet_language, created_at, user_lang, time_zone, entities contained in the tweet).</li> <li>sentiment_values.csv contains sentiment information for hashtags. It contains the hashtag itself and the sentiment values gathered via four different sentiment dictionaries: AFINN, Opinion Lexicon, Sentistrength Lexicon and vader. For each of these dictionaries we list the minimum, maximum, sum and average of all sentiments of the tokens of the hashtag (if available, else we list empty values). However, as most hashtags only consist of a single token, these values are equal in most cases. Please note that the lexica are rather diverse and therefore, are able to resolve very different terms against a score. Hence, the resulting csv is rather sparse. The file contains the following comma-separated values: <hashtag, vader_min, vader_max, vader_sum,vader_avg, afinn_min, afinn_max, afinn_sum, afinn_avg, ol_min, ol_max, ol_sum, ol_avg, ss_min, ss_max, ss_sum, ss_avg >, where we abbreviate all scores gathered over the Opinion Lexicon with the prefix 'ol'. Similarly, 'ss' stands for SentiStrength. </li> </ul> <p>Please also find the training and test-splits for the dataset in this repo. Also, prototypical implementations of a context-aware recommender system based on the dataset can be found at <a href="https://github.com/asmitapoddar/nowplaying-RS-Music-Reco-FM">https://github.com/asmitapoddar/nowplaying-RS-Music-Reco-FM</a>.</p> <p>If you make use of this dataset, please cite the following paper where we describe and experiment with the dataset:</p> <p>@inproceedings{smc18,<br> title = {#nowplaying-RS: A New Benchmark Dataset for Building Context-Aware Music Recommender Systems},<br> author = {Asmita Poddar and Eva Zangerle and Yi-Hsuan Yang},<br> url = {http://mac.citi.sinica.edu.tw/~yang/pub/poddar18smc.pdf},<br> year = {2018},<br> date = {2018-07-04},<br> booktitle = {Proceedings of the 15th Sound & Music Computing Conference},<br> address = {Limassol, Cyprus},<br> note = {code at https://github.com/asmitapoddar/nowplaying-RS-Music-Reco-FM},<br> tppubtype = {inproceedings}<br> }</p>
Datasets from the KDD 2021 article "A Semi-Personalized System for User Cold Start Recommendation on Music Streaming Apps"
<p>We publicly release the anonymized <em>song_embeddings.parquet user_embeddings.parquet user_features_test.parquet user_features_train.parquet user_features_validation.parquet</em> datasets, with each of the TT-SVD or UT-ALS versions of embeddings, from the music streaming platform Deezer, as described in the article "<em>A Semi-Personalized System for User Cold Start Recommendation on Music Streaming Apps"</em> published in the proceedings of the 27TH ACM SIGKDD conference on knowledge discovery and data mining (<em>KDD 2021</em>). The paper is available <a href="https://arxiv.org/abs/2106.03819">here</a>.</p> <p>These datasets are used in the GitHub repository <a href="https://github.com/deezer/semi_perso_user_cold_start">deezer/semi_perso_user_cold_start</a> to reproduce experiments from the article.</p> <p>Please cite our paper if you use our code or data in your work.</p>
GUI evaluation data for an IDE command recommender system
<p>This dataset contains results of the study conducted among the participants of the XP 2016 (a scientific conference with a strong participation of practitioners from the industry). The objective of the study was to evaluate the acceptance and usability of the proposed Graphical User Interface (GUI) for an Integrated Development Environment (IDE) command recommender system (RS). The data was collected by the questionnaire and the interviews. The data is anonymized.</p> <p>Content:</p> <ul> <li>README.txt</li> <li>./Survey answers.csv - the questionnaire answers</li> <li>./Interviews <ul> <li>./interviewXXX.txt - a file with a transcribed interview</li> <li>./mapping-codes-to-primary-documents.csv - a binary table summarizing interviews</li> </ul> </li> </ul>
Dataset Online grocery shopping recommender systems: common approaches and practices
<p>Standardized Excel form for data extraction. Extraction criteria were defined based on the sub-research questions that are provided in the paper. </p>
PODCAST: The Impact of News Recommender Systems on our Personal Identity
<p>This 45-minute podcast is a result of qualitative research which aims to identify if we should be concerned that news recommender systems may have an impact on our personal identities over time. Ind sets out to achieve a number of objectives.</p> <p>1. To understand how news recommender systems influence the way of seeing and experiencing the world in which we anchor our identities<br> 2. To examine the extent to which news recommendations are transforming and shaping our preferences and behaviours<br> 3. To determine if we should be concerned about personal data privacy, as news recommendations technology collects our personal data to create algorithmically generated news recommendations<br> 4. To explore if we think that personalised news recommendations can serve to sharpen our focus and broaden our minds in parallel</p> <p>The Impact of News Recommender Systems on our Personal Identity is licensed under a Creative Commons License.</p>
Datensatz für BA: Entwicklung eines Recommender-Systems für die Zuordnung von Anforderungen zu IT-Services
<ul> <li> <p><strong>afos_03_promise_nrf.csv</strong>: Diese Datei enthält funktionale Anforderungen aus dem PROMISE-Datensatz. Die Abkürzung „nrf“ könnte sich auf spezifische Notwendigkeiten oder Funktionen innerhalb der PROMISE-Daten beziehen. Sie dient als Quelle für Anforderungen, die im Recommender-System genutzt werden können.</p> </li> <li> <p><strong>afos_tcs_sampled.csv</strong>: Diese Datei enthält eine Stichprobe von funktionalen Anforderungen aus einem zusätzlichen Quellenbestand oder einem Datensatz mit technischen Anforderungen. Diese Anforderungen ergänzen den PROMISE-Datensatz und erweitern die Datenbasis für die Entwicklung und das Training des Recommender-Systems.</p> </li> <li> <p><strong>Goldstandard_256.csv</strong>: Diese Datei bildet den Goldstandard für das Recommender-System und enthält 256 funktionale Anforderungen, die manuell den entsprechenden IT-Services zugeordnet wurden. Sie dient als Referenz zur Bewertung der Qualität des Recommender-Systems und ermöglicht die Validierung und Optimierung des Modells.</p> </li> <li> <p><strong>ITSM_Set_with_Descriptions.xlsx</strong>: Diese Datei enthält eine Sammlung von IT-Services mit detaillierten Beschreibungen. Die Tabelle umfasst Kategorien und Unterkategorien der IT-Services, die als Basis für die Empfehlungen des Recommender-Systems dienen. Diese Beschreibungen ermöglichen eine semantische Analyse und unterstützen das Modell bei der Zuordnung zu passenden Anforderungen.</p> </li> <li> <p><strong>Recommender_System.ipynb</strong>: Dies ist ein Jupyter-Notebook, das den Code zur Entwicklung, Training und Evaluierung des Recommender-Systems enthält. Das Notebook fasst die Implementierung der Empfehlungslogik zusammen und stellt die Grundlage für die experimentellen Auswertungen dar.</p> </li> <li> <p><strong>requirements.txt</strong>: Diese Datei listet die Python-Bibliotheken und deren Versionen auf, die für das Recommender-System benötigt werden. Sie stellt sicher, dass alle benötigten Abhängigkeiten für die erfolgreiche Ausführung des Codes installiert sind.</p> </li> </ul>
Dataset used for "A Recommender System of Buggy App Checkers for App Store Moderators"
<p>This is the dataset used for paper: "A Recommender System of Buggy App Checkers for App Store Moderators", published on the <em>International Conference on Mobile Software Engineering and Systems (MOBILESoft)</em> in 2015.<br> <br> <strong>Dataset Collection</strong><br> We built a dataset that consists of a random sample of <em><strong>Android app metadata</strong></em> and <em><strong>user reviews</strong></em> available on the <em>Google Play Store</em> on January and March 2014.<br> Since the Google Play Store is continuously evolving (adding, removing and/or updating apps), we updated the dataset twice.<br> The dataset D1 contains available apps in the Google Play Store in January 2014.<br> Then, we created a new snapshot (D2) of the Google Play Store in March 2014.<br> <br> The apps belong to the 27 different categories defined by Google (at the time of writing the paper), and the 4 predefined subcategories (free, paid, new_free, and new_paid). For each category-subcategory pair (e.g. tools-free, tools-paid, sports-new_free, etc.), we collected a maximum of 500 samples, resulting in a median number of 1.978 apps per category.</p> <p>For each app, we retrieved the following metadata: <em>name, package, creator, version code, version name, number of downloads, size, upload date, star rating, star counting</em>, and the set of <em>permission requests</em>.<br> <br> In addition, for each app, we collected up to a maximum of the latest 500 reviews posted by users in the Google Play Store. For each review, we retrieved its metadata:<em> title, description, device,</em> and <em>version</em> of the app. None of these fields were mandatory, thus<br> several reviews lack some of these details.<br> From all the reviews attached to an app, we only considered the reviews associated with the latest version of the app —i.e., we discarded unversioned and old-versioned reviews. Thus, resulting in a corpus of 1,402,717 reviews (2014 Jan.).</p> <p> </p> <p><strong>Dataset Stats</strong><br> Some stats about the datasets:</p> <p>- <strong>D1</strong> (<em>Jan. 2014</em>) contains 38,781 apps requesting 7,826 different permissions, and 1,402,717 user reviews.</p> <p>- <strong>D2</strong> (<em>Mar. 2014</em>) contains 46,644 apps and 9,319 different permission requests, and 1,361,319 user reviews.</p> <p>Additional stats about the datasets are available <a href="https://sites.google.com/site/androidbuggyappcheckers">here</a>.<br> <br> <br> <strong>Dataset Description</strong><br> To store the dataset, we created a graph database with <a href="https://neo4j.com/">Neo4j</a>. This dataset therefore consists of a graph describing the apps as nodes and edges. We chose a graph database because the graph visualization helps to identify connections among data (e.g.,<br> clusters of apps sharing similar sets of permission requests).<br> <br> In particular, our dataset graph contains six <em>types of nodes</em>:<br> - <em>APP</em> nodes containing metadata of each app,<br> - <em>PERMISSION</em> nodes describing permission types,<br> - <em>CATEGORY</em> nodes describing app categories,<br> - <em>SUBCATEGORY</em> nodes describing app subcategories,<br> - <em>USER_REVIEW</em> nodes storing user reviews.<br> - <em>TOPIC</em> topics mined from user reviews (using LDA).<br> <br> Furthermore, there are five <em>types of relationships</em> between APP nodes and each of the remaining nodes:</p> <p>- <em>USES_PERMISSION</em> relationships between APP and PERMISSION nodes<br> - <em>HAS_REVIEW</em> between APP and USER_REVIEW nodes<br> - <em>HAS_TOPIC</em> between USER_REVIEW and TOPIC nodes<br> - <em>BELONGS_TO_CATEGORY</em> between APP and CATEGORY nodes<br> - <em>BELONGS_TO_SUBCATEGORY</em> between APP and SUBCATEGORY nodes</p> <p><br> <strong>Dataset Files Info</strong></p> <ul> <li><strong>Neo4j 2.0 Databases</strong> <ul> <li><em><strong>googlePlayDB1-Jan2014_neo4j_2_0.rar</strong></em></li> <li><em><strong>googlePlayDB2-Mar2014_neo4j_2_0.rar</strong></em><br> We provide two Neo4j databases containing the 2 snapshots of the Google Play Store (January and March 2014). These are the original databases created for the paper. The databases were created with <strong>Neo4j 2.0. </strong>In particular with the tool version <em>'Neo4j 2.0.0-M06 Community Edition' (latest version available at the time of implementing the paper in 2014).</em><br> </li> </ul> </li> <li><strong>Neo4j 3.5 Databases</strong> <ul> <li><strong>googlePlayDB1-Jan2014_neo4j_3_5_28.rar</strong></li> <li><strong>googlePlayDB2-Mar2014_neo4j_3_5_28.rar</strong><br> Currently, the version Neo4j 2.0 is deprecated and it is not available for download in the official <em>Neo4j Download Center</em>. We have migrated the original databases (Neo4j 2.0) to Neo4j 3.5.28.<br> The databases can be opened with the tool version: <em>'Neo4j Community Edition 3.5.28'.<br> The tool can be downloaded from the official </em><a href="https://neo4j.com/download-center/#community">Neo4j Donwload</a> page.<br> <br> In order to open the databases with more recent versions of Neo4j, the databases must be first migrated to the corresponding version. Instructions about the migration process can be found in the <a href="https://neo4j.com/docs/upgrade-migration-guide/current/understanding-upgrades-migration/">Neo4j Migration Guide</a>.<br> <br> First time the Neo4j database is connected, it could request credentials. The username and pasword are: neo4j/neo4j <br> <br> </li> </ul> </li> </ul> <p> </p>
Data from: A clinical decision support system learned from data to personalize treatment recommendations towards preventing breast cancer metastasis
Objective: A Clinical Decision Support System (CDSS) that can amass Electronic Health Record (EHR) and other patient data holds promise to provide accurate classification and guide treatment choices. Our objective is to develop the Decision Support System for Making Personalized Assessments and Recommendations Concerning Breast Cancer Patients (DPAC), which is a CDSS learned from data that recommends the optimal treatment decisions based on a patient's features. Method: We developed a Bayesian network architecture called Causal Modeling with Internal Layers (CAMIL), and an algorithm called Treatment Feature Interactions (TFI), which learns from data the interactions needed in a CAMIL model. Using the TFI algorithm, we learned interactions for six treatments from the Lynn Sage Data Set (LSDS). We created a CAMIL model using these interactions, resulting in a DPAC which recommends treatments towards preventing 5-year breast cancer metastasis. Results: In a 5-fold cross-validation analysis, we compared the probability of being metastasis free in 5 years for patients who made decisions recommended by DPAC to those who did not. These probabilities are (the probability for those making the decisions appears first): chemotherapy (.938, .872); breast/chest wall radiation (.939, .902); nodal field radiation (.940, .784); antihormone (.941, .906); HER2 inhibitors (.934, .880); neadjuvant therapy (.931, .837). In an application of DPAC to the independent METABRIC dataset, the probabilities for chemotherapy were (.845, .788). Discussion: Patients who took the advice of DPAC had, as a group, notably better outcomes than those who did not. We conclude that DPAC is effective at amassing and analyzing data towards treatment recommendations. Some of the findings in DPAC are controversial. For example, DPAC says that chemotherapy increases the chances of metastasis for many node negative patients. This controversy shows the importance of developing a conclusive version of DPAC to ensure we provide patients with the best patient-specific treatment recommendations.
Recommender Systems and AI Techniques in E-commerce: An Analysis of Trends and the Research Agenda
Open the record for dataset details and reuse information.
Empowering Coffee Farming Using Counterfactual Recommendation based RNN-IoT Integrated Soil Fertility Control System
Open the record for dataset details and reuse information.
Game-Shapley recommender system demonstration
<p>Recommender system using Game-Shapley algorithm</p>
Top 250 IMDb Movies Dataset for Recommendation Systems
<p>Dataset obtenido en la práctica 1 de la asignatura "Tipología y ciclo de vida de los datos", del Máster en ciencia de datos de la UOC. Ha sido obtenido por Ignacio Gimeno Alonso y Morad Kharraz Senhaji.</p> <p>Los datos de este dataset han sido extraídos de la lista de las 250 películas mejor valoradas presente en la web de IMDb (https://www.imdb.com/chart/top/?ref_=nv_mv_250)</p> <p>El dataset contiene los siguientes campos:</p> <p>· ranking: Puesto de la película en la lista de las 250 mejor valoradas.</p> <p>· nombre: Título de la versión española de la película.</p> <p>· enlace: Página web de la película en <a href="http://www.imdb.com">www.imdb.com</a>.</p> <p>· ano_lanz: Año de estreno de la película.</p> <p>· duración: Duración de la película, en horas y minutos.</p> <p>· edad: Clasificación de edad. Puede estar en distintos formatos, según el año de estreno y el país de producción (18, A, apta para mayores,...).</p> <p>· rating: Puntuación media dada por los usuarios de IMDb, de 0 a 10.</p> <p>· num_votos: Cantidad de valoraciones que ha recibido la película.</p> <p>· titulo_original: Título original de la película. Si está vacío, significa que el título original coincide con el título en la versión española.</p> <p>· sinopsis: Resumen de la película en español. Es un resumen corto, de unas pocas frases.</p> <p>· genero: géneros en los que se engloba la película, en inglés.</p> <p>· direccion: Director o directores de la película.</p> <p>· guionistas: Guionistas de la película.</p> <p>· elenco: Actores / actrices principales de la película.</p> <p>Los datos contenidos en el dataset están referidos a películas desde 1921 hasta 2024, pero las valoraciones están referidas al momento de recolección de los datos (octubre-noviembre de 2024).</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.