Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
29
datasets available to search
ShareScore release 0.9.0
Dataset results
29 results for “YouTube videos”
Disinformation on YouTube: A dataset of YouTube comments on videos related to claims made by Trump and Vance on Haitian immigrants
<div> <div> <div> <div> <div> <p>The corpus contains three files. First, the youtube_haitian_disinformation_videos_meta.csv file includes comments and YouTube video metadata. Data is organized around per video information. The columnar values are: </p> </div> </div> </div> <div> <ul> <li> <p>video_id </p> <p> </p> </li> <li> <p>date (video publication date) </p> <p> </p> </li> <li> <p>title </p> <p> </p> </li> <li> <p>description </p> <p> </p> </li> <li> <p>channel_title </p> <p> </p> </li> <li> <p>transcript </p> <p> </p> </li> <li> <p>transcript_str (Video transcript without timestamps) </p> <p> </p> </li> <li> <p>views </p> <p> </p> </li> <li> <p>likes </p> <p> </p> </li> <li> <p>comments (all comments per video) </p> </li> </ul> </div> <div> <div> <div> <p>Second, the youtube_haitian_disinformation_comment_reply_metadata.csv file includes comments, replies, and comment metadata. Each comment occupies its own row in the spreadsheet. The columnar field are: </p> </div> </div> </div> <div> <ul> <li> <p>video_id </p> <p> </p> </li> <li> <p>comment </p> <p> </p> </li> <li> <p>comment_date </p> <p> </p> </li> <li> <p>comment_like_count </p> <p> </p> </li> <li> <p>author </p> <p> </p> </li> <li> <p>comment_id </p> <p> </p> </li> <li> <p>in_reply_to </p> <p> </p> </li> <li> <p>neg, neu, pos, compound (VADER polarity scores) </p> <p> </p> </li> <li> <p>named_entities (spaCy named entity tags with tokens) </p> <p> </p> </li> <li> <p>emoji (spaCy Emojis and token spans) </p> </li> </ul> </div> </div> <div> <div> <div> <div> <p>Comments and associated metadata are represented in individual rows. </p> </div> <div> <p>The third file contains the results of the TFIDF analysis described herein. The TFIDF analysis features the top 5000 terms weights for the comments to each video in a .csv file. </p> </div> </div> </div> <div> <ul> <li> <p>youtube_disinfo_comments_tfidf_results_per_video.csv </p> </li> </ul> </div> </div> </div>
Protests Ukraine Covid 2020-22: YouTube Videos
The collection "YTV Protests Ukraine Covid 2020-22" contains 146 videos (mp4) on protests relating to government measures due to Covid-19. We have downloaded all data in October 2024 and made screenshots (pdf) of websites so that the discussion and comments on the single video posts can be followed. All data is processed in an MS Excel database with metadata. We collect all videos that are 1) event related, 2) show actions of this event, 3) we can find with our search words during a particular period. We strictly aim at a systematic and objective selection and organized storage of protest-related videos. The collection is based on extensive research into Covid-related protest events in Ukraine, which made it possible to identify relevant search words. According to the snowball principle, we then start the collection of videos with the help of these search words and try to download as much relevant content as possible. However, we cannot guarantee the completeness of protest videos on the particular event. We search the videos and include them into the collection until a particular degree of saturation has been reached. Due to copyright restrictions, we are only allowed to give access to the database of the collected video files including the hyperlinks with its metadata and not to the videos themselves. The videos have been posted mainly by TV channels and news outlets. Therefore, the material is only an extract and biased by the perspective of the single creator/creating institution. The collection is part of a larger and ongoing collection of videos on protest events in the post-Soviet region.
Protests Kazakhstan 2022: YouTube Videos
The collection "Protests Kazakhstan 2022" contains 315 videos (mp4) on protests in mainly January 2022 triggered by a sharp increase in gas prices. We have downloaded all data in September 2024 and made screenshots (pdf) of websites so that the discussion and comments on the single video posts can be followed. All data is processed in an MS Excel database with metadata. We collect all videos that are 1) event related AND show actions of this event, 2) downloadable, 3) we can find with our search words during a particular period. We strictly aim at a systematic and objective selection and organized storage of protest-related videos. We identify particular event-related search words after intense research on the event. According to the snowball principle, we then start the collection of videos with the help of these search words and try to download as much relevant content as possible. However, we cannot guarantee the completeness of protest videos on the particular event. We search the videos and include them into the collection until a particular degree of saturation has been reached. Due to copyright restrictions, we are only allowed to give access to the database of the collected video files including the hyperlinks with its metadata and not to the videos themselves. The videos have been posted mainly by the participants of the events. Therefore, the material is only an extract and biased by the perspective of the single creator. The collection is part of a larger and ongoing collection of videos on protest events in the post-Soviet region.
YouNICon: YouTube's CommuNIty of Conspiracy Videos
<p>this repository contains 7 files. </p> <ol> <li>all_videos.csv: contains all videos with the metadata</li> <li>comments_anon.csv: contains comment id, video id, anonymised author id of the comment, Perspective API scores of the comment text. (note: the actual comment text and author id has been removed in this dataset, rehydration using the YouTube Data API is required if the actual comment is needed)</li> <li>conspiracy_label.csv: contains video id and the label if the video contains conspiracy. </li> <li>Keywords_List.csv: contains the keywords used for topic inference. </li> <li>train_final.csv: training set for the model</li> <li>val_final.csv: validation set for the model </li> <li>test_final.csv: test set for the model </li> </ol>
Comments on YouTube videos of top two Indian political parties named Indian National Congress and Bhartiya Janata Party
<p>The datasets are taken from Various YouTube videos of top two Indian political parties named Indian National Congress and Bhartiya Janata Party.</p> <p>Both the datasets are divided into two categories: -</p> <p>Label 1- Positive</p> <p>Label 2- Negative</p> <p>All the labelling has been done manually.</p> <p><br> </p> <p><strong>Indian National Congress dataset:</strong></p> <p>Characteristics: Bivariate</p> <p>Number of instances in dataset: 1998</p> <p>Area of subject: Politics</p> <p>Attribute characteristics of dataset: Real</p> <p>Number of attributes: 2</p> <p>Date donated: March, 2019</p> <p>Task associated: Classification(binary)</p> <p>Missing values: Null</p> <p><br> </p> <p><strong>Bhartiya Janata Party dataset:</strong></p> <p>Characteristics: Bivariate</p> <p>Number of instances in dataset: 1952</p> <p>Area of subject: Politics</p> <p>Attribute characteristics of dataset: Real</p> <p>Number of attributes: 2</p> <p>Date donated: March, 2019</p> <p>Task associated: Classification(binary)</p> <p>Missing values: Null</p> <p><br> </p> <p><br> </p> <p><strong>Bothe datasets contains equal number of positive and negative comments:</strong></p> <p><br> </p> <p>Total number of positive comments present in Bhartiya Janata Party dataset =976</p> <p>Total number of negative comments present in Bhartiya Janata Party dataset =976</p> <p>Total number of positive comments present in Indian National Congress dataset=999</p> <p>Total number of negative comments present in Bhartiya Janata Party dataset=999</p> <p><br> </p> <p><strong>Both datasets contain following attributes:</strong></p> <ul> <li> <p><strong>comment text </strong></p> </li> <li> <p><strong>Labels </strong></p> </li> </ul>
Tour de Canoz - Youtube video
La tour de Canoz. Patrimoine Jurasien. Modèle 3D issu d'une vidéo youtube : https://youtu.be/ubGVF2URG28 96 pictures - Agisoft Photoscan Source: Objaverse 1.0 / Sketchfab
Motivations for citing research in comments to YouTube videos
<p><strong>This dataset includes 300 YouTube comments that link to research publications or preprints. Each comment has been assigned into one category that define why individuals mention scholarly publications in comments to YouTube videos. </strong></p> <p>Each row of the file "Categories_300random.xlsx" represents one distinct comment to YouTube video.<br> The file includes the following columns:</p> <ul> <li><strong>Comment_text </strong>- the full text of a comment,</li> <li><strong>YouTube_link </strong>- URL to the video where the comment was left,</li> <li><strong>Comment_link </strong>- URL to the thread with the comment,</li> <li><strong>Category </strong>- the final category that was assigned to the comment,</li> <li>Columns <strong>Category_old_schema_SS</strong>, <strong>Category_old_schema_IP</strong>, <strong>Category_old_schema_OZ </strong>- categories assigned by different researchers according to old categorization schema and were used for validation reasons,</li> <li>Columns <strong>Category_OZ</strong>, and <strong>Caregory_LB </strong>- categories assigned by different researchers according to the final schema and were used for validation.<br> <br> </li> </ul>
A Labelled Dataset for Sentiment Analysis of Videos on YouTube, TikTok, and other sources about the 2024 outbreak of Measles
<p><strong>Please cite the following paper when using this dataset:</strong></p> <p>N. Thakur, V. Su, M. Shao, K. Patel, H. Jeong, V. Knieling, and A. Bian “A labelled dataset for sentiment analysis of videos on YouTube, TikTok, and other sources about the 2024 outbreak of measles,” Proceedings of the 26th International Conference on Human-Computer Interaction (HCII 2024), Washington, USA, 29 June - 4 July 2024. (Accepted as a Late Breaking Paper, Preprint Available at: <a href="https://doi.org/10.48550/arXiv.2406.07693" rel="nofollow">https://doi.org/10.48550/arXiv.2406.07693</a>)</p> <p><strong>Abstract</strong></p> <p>This dataset contains the data of 4011 videos about the ongoing outbreak of measles published on 264 websites on the internet between January 1, 2024, and May 31, 2024. These websites primarily include YouTube and TikTok, which account for 48.6% and 15.2% of the videos, respectively. The remainder of the websites include Instagram and Facebook as well as the websites of various global and local news organizations. For each of these videos, the URL of the video, title of the post, description of the post, and the date of publication of the video are presented as separate attributes in the dataset. After developing this dataset, sentiment analysis (using VADER), subjectivity analysis (using TextBlob), and fine-grain sentiment analysis (using DistilRoBERTa-base) of the video titles and video descriptions were performed. This included classifying each video title and video description into (i) one of the sentiment classes i.e. positive, negative, or neutral, (ii) one of the subjectivity classes i.e. highly opinionated, neutral opinionated, or least opinionated, and (iii) one of the fine-grain sentiment classes i.e. fear, surprise, joy, sadness, anger, disgust, or neutral. These results are presented as separate attributes in the dataset for the training and testing of machine learning algorithms for performing sentiment analysis or subjectivity analysis in this field as well as for other applications. The paper associated with this dataset (please see the above-mentioned citation) also presents a list of open research questions that may be investigated using this dataset.</p>
[LS2N_IPI_YouTube_UGC] A DATASET FOR UNDERSTANDING OPEN UGC VIDEO DATASETS
<p>User Generated Content (UGC) video streaming is a major application on the Internet. Even small bitrate savings can have large network impacts at this scale. In order to achieve improvements without sacrificing experience, the quality of UGC videos needs to be better understood. In recent years video quality evaluation models designed for the evaluation of UGC videos have received a lot of attention. However, considering that these models are learning-based models, they heavily depend on the training data that has been used. </p> <p>In this paper, a new dataset is introduced that allows studying the differences in characteristics between existing UGC video datasets. It reveals the range of quality that was covered by existing UGC video datasets, and the implication of these quality ranges on training and validation performance of UGC video quality prediction models. Furthermore, this work demonstrates that dataset alignment enables existing UGC models to achieve higher performance.</p>
YA Domain Dataset: Dataset of scholarly bibliographic references on YouTube videos
<p><strong>Abstract</strong></p> <p>Scholarly communication through YouTube videos has been increasing. Although Altmetric (<a href="https://altmetric.com/">https://altmetric.com/</a>) provides the dataset on such references, its coverage is unclear, and it does not contain the original external links in each video. Considering this background, we built and published a dataset of scholarly bibliographic references on YouTube videos by using YouTube Data API v3, targeting six types of domain names: "doi.org," "ncbi.nlm.nih.gov," ieeexplore.ieee.org," "link.springer.com," "onlinelibrary.wiley.com," and "sciencedirect.com." As a result, we identified approximately 480,000 references associated with Crossref DOIs among 230,000 videos published by December 31, 2023, posted on 55,000 channels. Notably, over half of these references were not covered by the Altmetric dataset, resulting in a 150% increase in the number of references when combining the dataset constructed by the proposed method with the Altmetric dataset, compared to the Altmetric dataset alone. Regarding external links, PubMed and DOI links were prominent; however, a substantial number of direct links to publisher platforms were observed. Most channels and videos contained external links to a single platform, scattered across each platform. This dataset is helpful for identifying and analyzing scholarly references on YouTube.<br>As for the original paper related to this dataset, please refer to the references section.</p> <p> </p> <p><strong>Data Records</strong></p> <p>The data format of the dataset is JSON lines, where each line is a single record. The data is split into files by DOI Registration Agencies. A sample of the record is as follows:</p> <table> <tbody> <tr> <td>{<br> "channel_id": "UCEfEi-IMiB87UsxY3765P6w",<br> "video_id": "e7YmyVd4uOE",<br> "is_covered_by_altmetric_com": false,<br> "youtube_data_api_search": [<br> {<br> "query": "doi.org",<br> "uri": "http://dx.doi.org/10.1145/2807442.2814654"<br> }<br> ],<br> "doi": "10.1145/2807442.2814654",<br> "doiRA": "Crossref"<br>}</td> </tr> </tbody> </table> <ul> <li>channel_id (String) -- Channel ID of the YouTube channel that uploaded the video.</li> <li>video_id (String) -- Video ID.</li> <li>is_covered_by_altmetric_com (Boolean) -- Whether this reference is covered by altmetric.com or not.</li> <li>youtube_data_api_search (Array) <ul> <li> query (String) -- The query used in the search:list of YouTube Data API v3. (<a href="https://developers.google.com/youtube/v3/docs/search/list?hl=en">https://developers.google.com/youtube/v3/docs/search/list?hl=en</a>)</li> <li> uri (String)-- The original external links written in the description text or video title in each video.</li> </ul> </li> <li>doi (String) -- DOI corresponding to the bibliographic reference in the video.</li> <li>doiRA (String) -- DOI registration agency for the DOI. We obtained this data using the WhichRA? API (<a href="https://www.doi.org/the-identifier/resources/factsheets/doi-resolution-documentation#4-which-ra">https://www.doi.org/the-identifier/resources/factsheets/doi-resolution-documentation#4-which-ra</a>).</li> </ul> <p>We note that the altmetric dataset obtained from Altmetric Explorer in this study is not included in this dataset.</p> <p><strong>References</strong></p> <ul> <li>Kikkawa, Jiro; Takaku, Masao; Yoshikane, Fuyuki: "Enhancing Identification of Scholarly Reference on YouTube: Method Development and Analysis of External Link Characteristics", <em>Proceedings of the 28th International Conference on Theory and Practice of Digital Libraries (<a href="https://tpdl2024.nuk.si/">TPDL 2024</a>)</em>, Ljubljana, Slovenia, Lecture Notes in Computer Science (LNCS), Vol.15178, 2024.09. (in press).</li> </ul> <p><strong>Fundings</strong></p> <p>JSPS KAKENHI Grant Numbers <a href="https://kaken.nii.ac.jp/en/grant/KAKENHI-PROJECT-22K18147/">JP22K18147</a>, <a href="https://kaken.nii.ac.jp/en/grant/KAKENHI-PROJECT-23K11761">JP23K11761</a>, and <a href="https://kaken.nii.ac.jp/en/grant/KAKENHI-PROJECT-24K15652">JP24K15652</a>.</p>
Commonfare videos on YouTube data (until September 30, 2019)
<p>Views of the PIE News project 16 videos in the "COMMONFARE videos" YouTube channel to September 30, 2019.</p>
Building better conservation media for primates and people: A case study of orangutan rescue and rehabilitation YouTube videos
<p>1. Conservation organizations rely on social/internet media platforms to raise awareness and fundraise. Social media is a double-edged sword: it can be a wide-reaching and effective tool for education and fundraising, but can also have counter-productive impacts on public views toward wildlife and understanding of wildlife conservation.</p> <p>2. For example, depicting humans interacting with wildlife in media may increase video popularity, but animals shown in anthropogenic contexts are also viewed as appealing pets. We are interested in understanding whether this is true for social media posts (YouTube videos) by orangutan rescue and rehabilitation organizations, which rely on social media for fundraising and awareness-raising. Our goal is to provide data and recommendations to guide these organizations in building media with positive conservation impact while minimizing potential negative effects.</p> <p>3. Using YouTube analytics and sentiment analysis of comments on 118 videos, we ask how viewer responses to videos vary with 1) the amount of human-orangutan interaction depicted, 2) the ages of the orangutans featured, and 3) the mention of threats to orangutans.</p> <p>4. Videos with longer human-orangutan interaction time were viewed more, but comments on them were significantly more likely to be negative toward Indonesian/Malaysian people. Comments on orangutan rescue/rehabilitation videos were more likely to be categorized as negative for orangutan conservation compared to videos about orangutans generally, and within these, so were comments on videos featuring infant and juvenile orangutans.</p> <p>5. Based on our findings, we recommend that orangutan rescue and rehabilitation organizations feature adult and mixed age groups of orangutans rather than infants and juveniles, minimize the amount of human-orangutan interaction shown, and talk about conservation threats to orangutans in their videos. We also recommend that, as a precaution, other primate rescue and rehabilitation groups also abide by these suggestions.</p>
Technical Land-Sea Spaces. Impacts of the Port Clusterization Phenomenon on coasts, cities and architectures [YouTube Video Version]
<p>Beatrice Moretti lectures on the phenomenon of spatial stretching that is imposing a profound evolution, both formal and institutional, in the sphere of contemporary port cities and regions, by giving first insights about the research methodology oriented in this phase to the definition of a indicator systems of the cluster dimension. The presentation questions the spatial impacts introduced by port clusters in the field of architectural design.</p> <p>[YouTube Video Version]<br> <br> <a href="https://www.iccaua.com/page/conference-brochure">6th International Conference of Contemporary Affairs in Architecture and Urbanism - ICCAUA2023</a><br> Alanya Hamdullah Emin Paşa University, Istanbul (Turkey)<br> Chairman of the Conference:<strong> </strong>Dr. <a href="https://arch.alanyahep.edu.tr/en/akademik-kadro">Hourakhsh A. Nia</a>, AHEP University, Alanya, Antalya, TR.<br> Special Session "Coastal and Maritime Spaces", proposed by The University IUAV, Venice (IT)<br> Chairs: Paolo De Martino and Fabio Carella (IUAV).</p>
Raw data for manuscript: Use of immunology in news and YouTube videos in the context of COVID-19: politicization and information bubbles
<p>Coding of newsarticles and videos related to immunology and COVID-19 in Italian and English</p>
Building better conservation media for primates and people: A case study of orangutan rescue and rehabilitation YouTube videos
Open the record for dataset details and reuse information.
Video information from the search "Deep learning" recursively on youtube through recommended videos
Open the record for dataset details and reuse information.
YouTube videos
Open the record for dataset details and reuse information.
Reliability and Quality of YouTube Videos Related to Balance Exercise
ClinicalTrials.gov study NCT07117734. IPD Sharing: UNDECIDED. Countries: 1. Publications: 4.
The content and quality of YouTube videos relating to interproximal reduction
Open the record for dataset details and reuse information.
Improving Patient Understanding of the Surgical Hospital Experience: Use of YouTube Video Playlist
ClinicalTrials.gov study NCT02546180. IPD Sharing: Not stated. Countries: 0. Publications: 1.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.