Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

89

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

89 results for “YouTube”

Learn how ShareScore rates datasets ↗
zenodo36/100

Raw data for manuscript: Use of immunology in news and YouTube videos in the context of COVID-19: politicization and information bubbles

<p>Coding of newsarticles and videos related to immunology and COVID-19 in Italian and English</p>

opencc-by-4.0Sep 2023View details →
dryad36/100

Building better conservation media for primates and people: A case study of orangutan rescue and rehabilitation YouTube videos

Open the record for dataset details and reuse information.

publicSep 2021View details →
dryad36/100

Data from: The viewer doesn’t always seem to care - response to fake animal rescues on YouTube and implications for social media self-policing policies

Open the record for dataset details and reuse information.

publicOct 2022View details →
zenodo32/100

Webis YouTube 8M Augmented 2018

<p>We used the YouTube Data API&nbsp;to augment the&nbsp;<a href="https://research.google.com/youtube8m/">YouTube 8M</a>&nbsp;corpus by crawling a variety of meta data for the videos.</p> <p>First point of interest was the &quot;video resource,&quot;&nbsp;which comprises data about the video, such as the video&rsquo;s title, description, uploader name, tags, view count, and more. Also included in the meta data is whether comments have been left for the video. If so, we downloaded them as well, including information about their authors, likes, dislikes, and responses.</p> <p>There is no property which specifies a video&rsquo;s&nbsp;language, since this information is not mandatory when uploading a video. Also, the API provides only information about the available captions, but not the captions themselves. Only the uploader of a video is given access to its captions via the API; we extracted them using <a href="https://ytdl-org.github.io/youtube-dl/">youtube-dl</a>.&nbsp;For each video, all manually created captions were downloaded, and auto-generated captions in the &quot;default&quot;&nbsp;language and English. The &quot;default&quot;&nbsp;auto-generated caption gives perhaps the only hint at a video&rsquo;s original language.</p> <p>Finally, we downloaded all thumbnails used to advertise a video, which are not available via the API, but only via a canonical URL. Our corpus provides the possibility to recreate the way a video is presented on YouTube (meta data and thumbnail), what the actual content is ((sub)titles and descriptions), and how its viewers reacted (comments).<br> <br> If you use this dataset in your publication, <strong>please cite the dataset as outlined in the right column.</strong></p>

opencc-by-4.0Jul 2018View details →
zenodo32/100

YouTube-ASMR-300K

<p>The YouTube-ASMR dataset contains URLS for over 900 hours of ASMR video clips with stereo/binaural audio produced by various YouTube artists. The following paper contains a detailed description of the dataset and how it was compiled:</p> <p>K. Yang, B. Russell and J. Salamon, &quot;Telling Left from Right: Learning Spatial Correspondence of Sight and Sound&quot;, IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Virtual Conference, June 2020.</p>

opencc-by-4.0Jun 2020View details →
zenodo32/100

Video information from the search "Deep learning" recursively on youtube through recommended videos

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →
zenodo32/100

Replication package for: "Advertising and Content Differentiation: Evidence from YouTube"

<p>Replication Package for: Kerkhof, A. (forthcoming) Advertising and Content Differentiation: Evidence from YouTube, Economic Journal&nbsp;</p> <p><strong>Important note regarding file format:</strong></p> <p><em>There are different instructions for windows and unix (MacOS) based systems.</em></p> <p><strong>Windows:</strong></p> <p>The zip-file was split into three components, using WinZip.&nbsp;</p> <ul> <li>Two of the segments of the split Zip file have the extensions ".z01" and ".z02".</li> <li>The last file has the extension ".zip".</li> <li>To open the split Zip file, open the file with the ".zip" extension. Don't try to open any of the files with the numbered extensions; WinZip won't recognize them as Zip files.</li> <li>Once the split Zip file has been opened, you can work with it much as you would work with a regular Zip file, except you can't add any new files or remove existing files. Some operations such as creating self-extracting Zip files and editing comments are also disabled for split Zip files.</li> <li>The split Zip file format is an extension of the Zip 2.0 specification. Therefore, some Zip utility programs may not be able to open split Zip files. Please see&nbsp;<a href="https://kb.winzip.com/help/HELP_SPLIT_ZIP_INFO.htm">Split Zip file compatibility information</a> for more details.</li> </ul> <p>For further information, see here:&nbsp;</p> <p><a href="https://kb.winzip.com/help/help_splitdlg.htm?alid=165377265.1715071480">Splitting Zip files (winzip.com)</a></p> <p>&nbsp;</p> <p><strong>Unix (MacOS):</strong></p> <p>The following sequence of commands will re-assemble the original zip file:</p> <ol> <li>click&nbsp;<em>download all</em> and unzip the enclosing archive on your computer. You will see 3 files in the archive.</li> <li>execute on your terminal&nbsp;<code>zip -s0 3-replication-package.zip --out package.zip</code></li> <li>execute on your terminal <code>unzip package.zip</code>&nbsp; to extract all files as usual. &nbsp;</li> </ol>

opencc-by-4.0May 2024View details →
zenodo32/100

Youtube

YouTube videos shared with EOL

opennotspecifiedAug 2024View details →
zenodo32/100

YouTube videos

Open the record for dataset details and reuse information.

opencc-by-sa-4.0Apr 2024View details →
zenodo32/100

Dataset for "Multimodal Translation in YouTube Shorts from Spanish Football Teams"

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo32/100

YouTube Content Analysis

<p>The two datasets related to the doctoral thesis of Nathan Richards - Thesis titled: In Three Acts: A Meditation on Black British Diaspora History and Memory in the Digital.&nbsp;</p> <p>&nbsp;</p> <p>The datasets refer to chapter 4 - Appendix 4 and Appendix 4.1&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

A Data-Driven Approach for Finding Requirements Relevant Feedback from TikTok and YouTube

<p>This dataset includes the list of videos from TikTok and YouTube, regarding 20 different products, used in our study on utilizing videos to identify requirements relevant user feedback. We also provide the content and labeling for each video.&nbsp;In addition, we provide the search terms for each of the products that helped us find the videos.&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo32/100

Proposing more ecologically-valid experiment protocol using YouTube platform - Dataset

<p>Uploaded data related to publication &quot;Proposing more ecologically-valid experiment protocol using YouTube platform&quot;.&nbsp;Included in the files are: data analysis project, database, video recordings of testers&#39; behaviors.</p>

opencc-by-4.0Jan 2023View details →
zenodo32/100

Data for publication "Wissenschaft fürs Wohnzimmer" - 2 years of weekly interactive, scientific livestreams on YouTube in Journal Polarforschung

<p>Data from the YouTube channel &quot;Wissenschaft f&uuml;rs Wohnzimmer&quot;</p>

opencc-by-4.0Aug 2023View details →
zenodo32/100

Menghubungkan Data Musik dari Spotify dan YouTube dengan Menggunakan Vocabularies Extended dan Linked Data

<p><strong>Menghubungkan Data Musik dari Spotify dan YouTube dengan Menggunakan Vocabularies dan Linked Data. Dataset diambil dari link&nbsp;</strong><a href="https://www.kaggle.com/datasets/salvatorerastelli/spotify-and-youtube">Spotify and Youtube (kaggle.com)</a></p>

opencc-by-4.0Sep 2023View details →
ClinicalTrials.gov32/100

YouTube-Delivered Physical Activity Intervention

ClinicalTrials.gov study NCT04499547. IPD Sharing: NO. Countries: 1. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov32/100

Reliability and Quality of YouTube Videos Related to Balance Exercise

ClinicalTrials.gov study NCT07117734. IPD Sharing: UNDECIDED. Countries: 1. Publications: 4.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo28/100

x264 performances on the Youtube UGC Dataset

<p>The res_ugc zip contains the results of a performance analysis of x264, for 201 configurations and 1397 videos of the Youtube User General Content Dataset (see https://media.withyoutube.com/).</p> <p>For each video (e.g. Animation_360P-3e40), there is a corresponding file (e.g. Animation_360P-3e40.csv)&nbsp; gathering the measurements for this video. The experimental protocol is detailed in the related paper (to link).</p>

opencc-by-4.0Jul 2020View details →
zenodo28/100

Public Dataset for the Paper "Using Hybrid Bayesian Networks to Detect Audience Behaviour Changes in Youtube"

<p>This is the dataset used in the production of the paper &quot;Using Hybrid Bayesian Networks to Detect Audience Behaviour Changes in Youtube&quot; to be published in the EMSS 2020. Official site:&nbsp;http://www.msc-les.org/conf/emss2020/</p>

opencc-by-4.0Jul 2020View details →
zenodo28/100

Youtube-Dataset for Language Identification in Speech Signals

<p><strong>Youtube-Dataset for Language Identification in Speech Signals</strong></p> <p>- for scientific use only, for questions contact: jakob.abesser@idmt.fraunhofer.de</p> <p><strong>Reference</strong></p> <p>In case you use this dataset for your research, please cite</p> <p>Alexandra Draghici, Jakob Abe&szlig;er &amp; Hanna Lukashevich: A Study on Spoken Language Identification<br> using Deep Neural Networks, Proceedings of the Audio Mostly Conference 2020</p> <p><strong>Dataset</strong></p> <p>The YouTube News Collection is a collection of videos from various<br> Youtube news channels. We gathered data from channels like BBC<br> news, France24, DW News, and Noticias Telemundo.</p> <p>- 135664 npy files (numpy matrices exported from Python)<br> - each npy file includes a mel spectrogram (see below) of an audio file<br> - the subfolders &quot;0&quot; - &quot;5&quot; encode the language id:<br> &nbsp; 0 - English<br> &nbsp; 1 - French<br> &nbsp; 2 - German<br> &nbsp; 3 - Greek<br> &nbsp; 4 - Italian<br> &nbsp; 5 - Spanish</p> <p><strong>Audio Processing</strong></p> <p>- mono, sample rate 22.05 kHz<br> - mel spectrogram (librosa python package)<br> - windows size 512 samples<br> - hopsize 441 samples (20 ms)<br> - 129 mel bands<br> - file-level spectrogram are normalized to maximum of 1<br> &nbsp;</p>

opencc-by-4.0Jul 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record