Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

89

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

89 results for “youtube”

Learn how ShareScore rates datasets ↗
zenodo48/100

Disinformation on YouTube: A dataset of YouTube comments on videos related to claims made by Trump and Vance on Haitian immigrants

<div> <div> <div> <div> <div> <p>The corpus&nbsp;contains three files. First, the youtube_haitian_disinformation_videos_meta.csv file includes comments and YouTube video metadata. Data is organized around per video information. The columnar&nbsp;values are:&nbsp;</p> </div> </div> </div> <div> <ul> <li> <p>video_id&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>date (video publication date)&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>title&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>description&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>channel_title&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>transcript&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>transcript_str (Video transcript without timestamps)&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>views&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>likes&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>comments (all comments per video)&nbsp;</p> </li> </ul> </div> <div> <div> <div> <p>Second, the youtube_haitian_disinformation_comment_reply_metadata.csv file includes comments, replies, and comment metadata. Each comment occupies its own row in the spreadsheet. The columnar field are:&nbsp;</p> </div> </div> </div> <div> <ul> <li> <p>video_id&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>comment&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>comment_date&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>comment_like_count&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>author&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>comment_id&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>in_reply_to&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>neg, neu, pos, compound (VADER polarity scores)&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>named_entities (spaCy named entity tags with tokens)&nbsp;</p> <p>&nbsp;</p> </li> <li> <p>emoji (spaCy Emojis and token spans)&nbsp;</p> </li> </ul> </div> </div> <div> <div> <div> <div> <p>Comments and associated metadata are represented in individual rows.&nbsp;</p> </div> <div> <p>The third file contains the results of the TFIDF analysis described herein. The TFIDF analysis features the top 5000 terms weights for the comments to each video in a .csv file.&nbsp;</p> </div> </div> </div> <div> <ul> <li> <p>youtube_disinfo_comments_tfidf_results_per_video.csv&nbsp;</p> </li> </ul> </div> </div> </div>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Youtube comments on Smart Farming

<p>Please cite the following paper, if you are working on or using the above dataset.</p> <p>https://www.mdpi.com/2624-7402/4/2/29</p> <p>Farming Dataset</p> <p>The datasets are taken from the following 16 YouTube videos related to farming:<br> 1. How Japan Is Reshaping Its Agriculture By Harnessing Smart Farming Technology-<br> Science Insider<br> 2. Europe has the best regenerative farmers in the world - Richard Perkins<br> 3. The Futuristic Farms That Will Feed the World - Freethink<br> 4. Vertical farms could take over the world - Freethink<br> 5. India&rsquo;s largest Precision Farm - Simply Fresh Discover Agriculture<br> 6. Solar Panels Plus Farming? Agrivoltaics Explained Undecided with Matt Ferell<br> 7. 7 Israeli Agriculture Technologies Israel<br> 8. The CNH Industrial Autonomous Tractor Concept (Full Version CNH Industrial<br> 9. IoT Smart Plant Monitoring System Viral Science -the home of creativity<br> 10. Singapore&rsquo;s Bold Plan to Build the Farms of the Future Tomorrow&rsquo;s Build<br> 11. Smart Vertical Farms in Sharjah Episode Up<br> 12. Drones, robots, and super sperm &ndash; the future of farming DW Documentary<br> 13. Simply Fresh - India&rsquo;s Largest State Of The Art Precision Farm<br> 14. RIPPA The Farm Robot Exterminates Pests And Weeds- ABC Science<br> 15. Top 10 Agritech Startups Empowering Indian Farmers Backstage With Millionaires<br> 16. This Farm of the Future Uses No Soil and 95% Less Water Stories<br> Total counts: 7334<br> The comments have been divided into the following labels:<br> 1. Praising (970 comments)<br> 2. Opinion (3224 comments),<br> 3. Suggestion (231 comments),<br> 4. Undefined (1386 comments),<br> 5. Queries (1007 comments)<br> 6. Hybrid (516 comments).</p> <p>The file includes:<br> 1. Comment<br> 2. Label</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Spotify and Youtube

<p>This is the statistics for the Top 10 songs of various spotify artists and their YouTube videos. The Creators above generated the data and uploaded it to Kaggle on February 6-7 2023. The license to use this data is "CC0: Public Domain", allowing the data to be copied, modified, distributed, and worked on without having to ask permission. The data is in numerical and textual CSV format as attached.</p><p>This dataset contains the statistics and attributes of the top 10 songs of various artists in the world. As described by the creators above, it includes 26 variables for each of the songs collected from spotify. These variables are briefly described next:</p><ul><li><strong>Track</strong>: name of the song, as visible on the Spotify platform.</li><li><strong>Artist</strong>: name of the artist.</li><li><strong>Url_spotify</strong>: the Url of the artist.</li><li><strong>Album</strong>: the album in wich the song is contained on Spotify.</li><li><strong>Album_type</strong>: indicates if the song is relesead on Spotify as a single or contained in an album.</li><li><strong>Uri</strong>: a spotify link used to find the song through the API.</li><li><strong>Danceability</strong>: describes how suitable a track is for dancing based on a combination of musical elements including tempo, rhythm stability, beat strength, and overall regularity. A value of 0.0 is least danceable and 1.0 is most danceable.</li><li><strong>Energy</strong>: is a measure from 0.0 to 1.0 and represents a perceptual measure of intensity and activity. Typically, energetic tracks feel fast, loud, and noisy. For example, death metal has high energy, while a Bach prelude scores low on the scale. Perceptual features contributing to this attribute include dynamic range, perceived loudness, timbre, onset rate, and general entropy.</li><li><strong>Key</strong>: the key the track is in. Integers map to pitches using standard Pitch Class notation. E.g. 0 = C, 1 = C♯/D♭, 2 = D, and so on. If no key was detected, the value is -1.</li><li><strong>Loudness</strong>: the overall loudness of a track in decibels (dB). Loudness values are averaged across the entire track and are useful for comparing relative loudness of tracks. Loudness is the quality of a sound that is the primary psychological correlate of physical strength (amplitude). Values typically range between -60 and 0 db.</li><li><strong>Speechiness</strong>: detects the presence of spoken words in a track. The more exclusively speech-like the recording (e.g. talk show, audio book, poetry), the closer to 1.0 the attribute value. Values above 0.66 describe tracks that are probably made entirely of spoken words. Values between 0.33 and 0.66 describe tracks that may contain both music and speech, either in sections or layered, including such cases as rap music. Values below 0.33 most likely represent music and other non-speech-like tracks.</li><li><strong>Acousticness</strong>: a confidence measure from 0.0 to 1.0 of whether the track is acoustic. 1.0 represents high confidence the track is acoustic.</li><li><strong>Instrumentalness</strong>: predicts whether a track contains no vocals. "Ooh" and "aah" sounds are treated as instrumental in this context. Rap or spoken word tracks are clearly "vocal". The closer the instrumentalness value is to 1.0, the greater likelihood the track contains no vocal content. Values above 0.5 are intended to represent instrumental tracks, but confidence is higher as the value approaches 1.0.</li><li><strong>Liveness</strong>: detects the presence of an audience in the recording. Higher liveness values represent an increased probability that the track was performed live. A value above 0.8 provides strong likelihood that the track is live.</li><li><strong>Valence</strong>: a measure from 0.0 to 1.0 describing the musical positiveness conveyed by a track. Tracks with high valence sound more positive (e.g. happy, cheerful, euphoric), while tracks with low valence sound more negative (e.g. sad, depressed, angry).</li><li><strong>Tempo</strong>: the overall estimated tempo of a track in beats per minute (BPM). In musical terminology, tempo is the speed or pace of a given piece and derives directly from the average beat duration.</li><li><strong>Duration_ms</strong>: the duration of the track in milliseconds.</li><li><strong>Stream</strong>: number of streams of the song on Spotify.</li><li><strong>Url_youtube</strong>: url of the video linked to the song on Youtube, if it have any.</li><li><strong>Title</strong>: title of the videoclip on youtube.</li><li><strong>Channel</strong>: name of the channel that have published the video.</li><li><strong>Views</strong>: number of views.</li><li><strong>Likes</strong>: number of likes.</li><li><strong>Comments</strong>: number of comments.</li><li><strong>Description</strong>: description of the video on Youtube.</li><li><strong>Licensed</strong>: Indicates whether the video represents licensed content, which means that the content was uploaded to a channel linked to a YouTube content partner and then claimed by that partner.</li><li><strong>official_video</strong>: boolean value that indicates if the video found is the official video of the song.</li></ul><p>The data was last updated on February 7, 2023.</p>

opencc-zeroFeb 2023View details →
zenodo44/100

Net-activism and whistleblowing dataset from Youtube

<p>Files contained in this archive is about net activism and whistleblowing</p> <p>version 1</p> <p>Documents of this dataset have been downloaded from Youtube mainly in 2015 with an update in 2018.</p> <p>Please cite usage of this dataset by</p> <p>Turenne N &nbsp; Net activism and whistleblowing on YouTube: a text mining analysis (2022).<br> &nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

MineDojo Internet Knowledge Base (YouTube)

<p><strong>Project website:</strong>&nbsp;<a href="https://minedojo.org">minedojo.org</a></p> <p><strong>Paper:</strong>&nbsp;<a href="https://arxiv.org/abs/2206.08853">arxiv.org/abs/2206.08853</a></p> <p><strong>GitHub:</strong>&nbsp;<a href="https://github.com/MineDojo/MineDojo">github.com/MineDojo/MineDojo</a></p> <p>Minecraft is among the most streamed games on YouTube. Human players have demonstrated a stunning range of creative activities and sophisticated missions that take hours to complete. We collect 730K+ narrated Minecraft videos, which add up to&nbsp;<strong>33 years of duration and 2.2B words</strong>&nbsp;in English transcripts. The time-aligned transcripts enable the agent to ground free-form natural language in video pixels and learn the semantics of diverse activities without laborious human labeling.</p> <p>There are two files in&nbsp;our YouTube knowledge base.</p> <ul> <li><strong>youtube_tutorial.json</strong> (tutorial videos):&nbsp; <p>Minecraft tutorial videos include step-by-step demonstrations and sometimes detailed verbal explanations. They also serve as a rich source of creative missions that humans find interesting. We harvest thousands of tasks from these videos in our benchmarking suite.&nbsp;</p> </li> <li><strong>youtube_full.json</strong> (general gameplay videos): <p>Unlike tutorials, general gameplay videos do not necessarily provide guidance on particular tasks. Instead, they capture the &ldquo;in-the-wild&rdquo; human experiences that are much larger in quantity, diverse in contents, and rich in learning signals.</p> </li> </ul> <p>Data Structure</p> <pre><code class="language-python">list[     {         "id": str,         # video id         "title": str,      # video title         "link": str,       # video link         "view_count": int  # number of times the video has been viewed         "like_count": int  # number of users who have indicated that they liked the video         "duration": float  # video duration in seconds         "fps": float,      # video FPS     } ]</code></pre> <p>Check out our&nbsp;paper!</p> <pre><code class="language-markdown">@article{fan2022minedojo, title = {MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge}, author = {Linxi Fan and Guanzhi Wang and Yunfan Jiang and Ajay Mandlekar and Yuncong Yang and Haoyi Zhu and Andrew Tang and De-An Huang and Yuke Zhu and Anima Anandkumar}, year = {2022}, journal = {arXiv preprint arXiv: Arxiv-2206.08853} }</code></pre> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Introducing the COVID-19 YouTube (COVYT) speech dataset featuring the same speakers with and without infection

<p>The COVYT dataset contains speech samples from individuals who self-reported their COVID-19 infection on public social media platforms (YouTube, Xiaohongshu). These videos, as well as accompanying videos of the same people prior to infection, were mined in an attempt to gather publicly-available data for COVID-19 research. This release includes the links to the original videos along with the accompanying&nbsp;manual segmentation and diarisation that identifies the utterances of the target individuals. We are additionally releasing features derived from the segmented utterances. Finally, the dataset includes partitioning information&nbsp;according to 4 different cross-validation schemes. See the arxiv pre-print for more details:&nbsp;https://arxiv.org/abs/2206.11045</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

YouTube RAI channel dataset

<p>id, title and youtube segmentation of videos from the official youtube RAI channel (<a href="https://www.youtube.com/@rai" rel="nofollow">https://www.youtube.com/@rai</a>) longer than 5 minutes. For each video the segmentation is a list composed by the start time (in milliseconds) and the title of each chapter. The dataset is already divided in two non-overlapping sets: 614 in "test_yt_over5min.json" and 2460 in "train_yt_over5min.json".</p>

openapache2.0Jun 2024View details →
zenodo44/100

Protests Ukraine Covid 2020-22: YouTube Videos

The collection "YTV Protests Ukraine Covid 2020-22" contains 146 videos (mp4) on protests relating to government measures due to Covid-19. We have downloaded all data in October 2024 and made screenshots (pdf) of websites so that the discussion and comments on the single video posts can be followed. All data is processed in an MS Excel database with metadata. We collect all videos that are 1) event related, 2) show actions of this event, 3) we can find with our search words during a particular period. We strictly aim at a systematic and objective selection and organized storage of protest-related videos. The collection is based on extensive research into Covid-related protest events in Ukraine, which made it possible to identify relevant search words. According to the snowball principle, we then start the collection of videos with the help of these search words and try to download as much relevant content as possible. However, we cannot guarantee the completeness of protest videos on the particular event. We search the videos and include them into the collection until a particular degree of saturation has been reached. Due to copyright restrictions, we are only allowed to give access to the database of the collected video files including the hyperlinks with its metadata and not to the videos themselves. The videos have been posted mainly by TV channels and news outlets. Therefore, the material is only an extract and biased by the perspective of the single creator/creating institution. The collection is part of a larger and ongoing collection of videos on protest events in the post-Soviet region.

openodc-byDec 2023View details →
zenodo44/100

Protests Kazakhstan 2022: YouTube Videos

The collection "Protests Kazakhstan 2022" contains 315 videos (mp4) on protests in mainly January 2022 triggered by a sharp increase in gas prices. We have downloaded all data in September 2024 and made screenshots (pdf) of websites so that the discussion and comments on the single video posts can be followed. All data is processed in an MS Excel database with metadata. We collect all videos that are 1) event related AND show actions of this event, 2) downloadable, 3) we can find with our search words during a particular period. We strictly aim at a systematic and objective selection and organized storage of protest-related videos. We identify particular event-related search words after intense research on the event. According to the snowball principle, we then start the collection of videos with the help of these search words and try to download as much relevant content as possible. However, we cannot guarantee the completeness of protest videos on the particular event. We search the videos and include them into the collection until a particular degree of saturation has been reached. Due to copyright restrictions, we are only allowed to give access to the database of the collected video files including the hyperlinks with its metadata and not to the videos themselves. The videos have been posted mainly by the participants of the events. Therefore, the material is only an extract and biased by the perspective of the single creator. The collection is part of a larger and ongoing collection of videos on protest events in the post-Soviet region.

openodc-byDec 2023View details →
zenodo44/100

SHS-YouTube1300: A YouTube-based Cover Song Dataset (Crema-PCP Features)

<p>These are the CREMA-PCP Features for our SHS-YouTube-1300 dataset and crawl. This crawl is based on the SHS100K dataset and contains YouTube videos of which a subset was annotated by crowd-workers and in-house annotators.</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

CNN YouTube: Titles, Views & Posting Dates Dataset

<p>This dataset was scrapped using one of the python modules called &quot;SiteScraper&quot;. The &quot;SiteScraper&quot; module is build on top of selenium. Here is the link to the module :&nbsp;<code>https://github.com/ibrahim-string/SiteScraper</code></p> <p>Text summarisation can be done on titles and correlation between views and titles can be explored using this dataset.</p> <p>This dataset will be very usefull To find out what kind of content democrats watch.</p> <p>This dataset will be updated every week.</p>

opencc-byJun 2023View details →
zenodo40/100

MALAYALAM LANGUAGE (MIX CODE) Recipe channels Youtube Comments

<p>The dataset used for performing Text Classification is a combination of two different datasets scrapped from the comment section of two Youtube channels namely &quot;Veen&#39;s Curryworld&quot; and &quot;Lekshmi Nair&quot;. The data is extracted using the YouTube API and contains two atrributes namely &ldquo;text&rdquo; and &ldquo;label&rdquo; where the former contains the comments and later contains the corresponding label. The comments are classified into 7 labels:</p> <p>&nbsp;</p> <p>- Label 1: Gratitude</p> <p>- Label 2: About the recipe</p> <p>- Label 3: About the video</p> <p>- Label 4: Praising</p> <p>- Label 5: Hybrid</p> <p>- Label 6: Undefined</p> <p>- Label 7: Suggestions and Queries</p> <p>&nbsp;</p> <p>The number of instances in each category is listed below:</p> <p>&nbsp;</p> <p>- Label 1: 484</p> <p>- Label 2: 396</p> <p>- Label 3: 300</p> <p>- Label 4: 362</p> <p>- Label 5: 249</p> <p>- Label 6: 2062</p> <p>- Label 7: 438</p>

opencc-by-4.0May 2020View details →
zenodo40/100

Canadian Politicians on YouTube

<div> <h4>Description of Columns</h4> </div> <p><code>Honorific title</code>: MPs who are members of the Canadian Privy Council and use the title &ldquo;The Honourable&rdquo; (Hon.)</p> <p><code>First name</code>: MPs&rsquo; first name</p> <p><code>Last name</code>: MPs&rsquo; last name</p> <p><code>Username</code>: MPs&rsquo; YouTube username</p> <p><code>Profile URL</code>: URL for MPs&rsquo; YouTube account</p> <p><code>Status</code>: Whether an MP&rsquo;s YouTube account is&nbsp;<em>Active</em>&nbsp;or&nbsp;<em>Inactive</em></p> <ul> <li> <p><code>Active</code>: At least one short or full length video was posted in 2025</p> </li> <li> <p><code>Inactive</code>: The last short or full length video was posted on December 31, 2024 or earlier</p> </li> </ul> <p><code>Gender</code>: MPs&rsquo; gender, as categorized by the&nbsp;<a href="https://www.ourcommons.ca/Members/en/search" rel="nofollow">House of Commons</a></p> <p><code>Political Affiliation</code>: MPs&rsquo; political party affiliation, as categorized by the&nbsp;<a href="https://www.ourcommons.ca/Members/en/search" rel="nofollow">House of Commons</a></p> <p><code>Constituency</code>: Name of the MPs&rsquo; constituency (as of the 45th Parliament, following redistribution)</p> <p><code>Province/Territory</code>: Where the MPs&rsquo; constituency is located</p>

openmit-licenseSep 2024View details →
zenodo40/100

Data Set Analisis Perilaku dan Interaksi pada YouTuber Gaming berdasarkan Persebaran Gender

<p>Data Set&nbsp;Analisis Perilaku dan Interaksi pada YouTuber Gaming berdasarkan Persebaran Gender</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Youtube cookery channels viewers comments in Hinglish

<p>The data was collected from the famous cookery Youtube channels in India. The major focus was to collect the viewers&#39; comments in Hinglish languages.&nbsp;The datasets are taken from top 2 Indian cooking channel named Nisha Madhulika channel and Kabita&rsquo;s&nbsp; Kitchen channel.</p> <p>Both the datasets comments are divided into seven categories:-</p> <p>Label 1- Gratitude</p> <p>Label 2- About the recipe</p> <p>Label 3- About the video</p> <p>Label 4- Praising</p> <p>Label 5- Hybrid</p> <p>Label 6- Undefined</p> <p>Label 7- Suggestions and queries</p> <p>All the labelling has been done manually.</p> <p>&nbsp;</p> <p><strong>Nisha Madhulika dataset:</strong></p> <p><strong>Dataset characteristics: Multivariate</strong></p> <p><strong>Number of instances: 4900</strong></p> <p><strong>Area: Cooking </strong></p> <p><strong>Attribute characteristics: Real</strong></p> <p><strong>Number of attributes: 3</strong></p> <p><strong>Date donated: March, 2019</strong></p> <p><strong>Associate tasks: Classification</strong></p> <p><strong>Missing values: Null</strong></p> <p>&nbsp;</p> <p><strong>Kabita Kitchen dataset:</strong></p> <p><strong>Dataset characteristics: Multivariate</strong></p> <p><strong>Number of instances: 4900</strong></p> <p><strong>Area: Cooking </strong></p> <p><strong>Attribute characteristics: Real</strong></p> <p><strong>Number of attributes: 3</strong></p> <p><strong>Date donated: March, 2019</strong></p> <p><strong>Associate tasks: Classification</strong></p> <p><strong>Missing values: Null</strong></p> <p>&nbsp;</p> <p>There are two separate datasets file of each channel named as preprocessing and main file .</p> <p>The files with preprocessing names are generated after doing the preprocessing and exploratory data analysis on both the datasets. This file includes:</p> <ul> <li>&nbsp;Id</li> <li>Comment text</li> <li>Labels</li> </ul> <ul> <li>Count of stop-words</li> <li>Uppercase words</li> <li>Hashtags</li> <li>Word count</li> <li>Char count</li> <li>Average words</li> <li>Numeric</li> </ul> <p>&nbsp;</p> <p>The main file includes:</p> <ul> <li>Id</li> <li>comment text</li> <li>Labels</li> </ul> <p>Please cite the paper</p> <p>https://www.mdpi.com/2504-2289/3/3/37</p> <p>&nbsp;</p> <p><strong>MDPI and ACS Style</strong></p> <p>Kaur, G.; Kaushik, A.; Sharma, S. Cooking Is Creating Emotion: A Study on Hinglish Sentiments of Youtube Cookery Channels Using Semi-Supervised Approach. <em>Big Data Cogn. Comput.</em> <strong>2019</strong>, <em>3</em>, 37.</p>

openodc-byMay 2019View details →
zenodo40/100

Universidades españolas en Youtube

<p>Estudio sobre la presencia de las universidades españolas en Youtube realizado por Lydia Gil. Los datos fueron recogidos de sus webs, canales institucionales en Youtube y SocialBlade, entre el 9 y 14 de marzo de 2017.</p> <p>Versión: 16 de marzo de 2017</p> <p> </p>

opencc-by-4.0Sep 2017View details →
zenodo40/100

Sexual Abusive Comments from YouTube by Roma Tre University

<p>1000 sexually abusive comments collected from popular YouTube Videos including cartoons like Peppa Pig</p>

opencc-by-4.0Nov 2017View details →
zenodo40/100

YouNICon: YouTube's CommuNIty of Conspiracy Videos

<p>this repository contains 7&nbsp;files.&nbsp;</p> <ol> <li>all_videos.csv: contains all videos with the metadata</li> <li>comments_anon.csv: contains comment id, video id, anonymised author id of the comment, Perspective API scores of the comment text. (note: the actual comment text and author id has been removed in this dataset, rehydration using the YouTube Data API is required if the actual comment is needed)</li> <li>conspiracy_label.csv: contains video id and the label if the video contains conspiracy.&nbsp;</li> <li>Keywords_List.csv: contains the keywords used for topic inference.&nbsp;</li> <li>train_final.csv: training set for the model</li> <li>val_final.csv: validation set for the model&nbsp;</li> <li>test_final.csv: test set for the model&nbsp;</li> </ol>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Regra de três: teorias da conspiração sobre Covid-19 no YouTube

<p>Material suplementar do artigo &quot;Regra de tr&ecirc;s: teorias da conspira&ccedil;&atilde;o sobre Covid-19 no YouTube&quot;</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Comments on YouTube videos of top two Indian political parties named Indian National Congress and Bhartiya Janata Party

<p>The datasets are taken from Various YouTube videos of top two Indian political parties named Indian National Congress and Bhartiya Janata Party.</p> <p>Both the datasets are divided into two categories: -</p> <p>Label 1- Positive</p> <p>Label 2- Negative</p> <p>All the labelling has been done manually.</p> <p><br> &nbsp;</p> <p><strong>Indian National Congress dataset:</strong></p> <p>Characteristics: Bivariate</p> <p>Number of instances in dataset: 1998</p> <p>Area of subject: Politics</p> <p>Attribute characteristics of dataset: Real</p> <p>Number of attributes: 2</p> <p>Date donated: March, 2019</p> <p>Task associated: Classification(binary)</p> <p>Missing values: Null</p> <p><br> &nbsp;</p> <p><strong>Bhartiya Janata Party dataset:</strong></p> <p>Characteristics: Bivariate</p> <p>Number of instances in dataset: 1952</p> <p>Area of subject: Politics</p> <p>Attribute characteristics of dataset: Real</p> <p>Number of attributes: 2</p> <p>Date donated: March, 2019</p> <p>Task associated: Classification(binary)</p> <p>Missing values: Null</p> <p><br> &nbsp;</p> <p><br> &nbsp;</p> <p><strong>Bothe datasets contains equal number of positive and negative comments:</strong></p> <p><br> &nbsp;</p> <p>Total number of positive comments present in Bhartiya Janata Party dataset =976</p> <p>Total number of negative comments present in Bhartiya Janata Party dataset =976</p> <p>Total number of positive comments present in Indian National Congress dataset=999</p> <p>Total number of negative comments present in Bhartiya Janata Party dataset=999</p> <p><br> &nbsp;</p> <p><strong>Both datasets contain following attributes:</strong></p> <ul> <li> <p><strong>comment text </strong></p> </li> <li> <p><strong>Labels </strong></p> </li> </ul>

opencc-by-4.0May 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record