Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

155

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

155 results for “Popular”

Learn how ShareScore rates datasets ↗
zenodo48/100

Supplementary File: Entertainment interspersed with propaganda: How non-legacy-news accounts deliver explicitly political content to mass audiences on Russia's most popular social network VK

<p>Supplementary file and dataset for the paper "Entertainment interspersed with propaganda: How non-legacy-news accounts deliver explicitly political content to mass audiences on Russia&rsquo;s most popular social network VK"</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Github data for static site generators (SSG) popularity

<p>Number of Github stars, forks, open issues, create and last modified dates for 30 open source static site generators (SSG), including Hugo, Jekyll and Gatsby.</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Top 120+ popular movies 2023 from RottenTomatoes

<p>[ENG] The file contains information about popular movies on the webpage Rotten Tomatoes. We suggest using R or Python to work with the dataset. The dataset has not been cleaned, so spaces, NaN values and unmatched variables types may be present.</p><p>[CAT] El fitxer conté informació sobre pel·lícules populars a la pàgina web Rotten Tomatoes. Suggerim utilitzar R o Python per a treballar amb el conjunt de dades. El conjunt de dades no s'ha netejat, de manera que els espais, els valors NaN i els tipus de variables que no coincideixn poden estar presents.</p><p>[ESP] El fichero contiene información sobre películas populares en la página web Rotten Tomatoes. Sugerimos utilizar R o Python para trabajar con el conjunto de datos. El conjunto de datos no se ha limpiado, de forma que los espacios, los valores NaN y los tipos de variables que no coincidan pueden estar presentes.</p><p>&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Co-occurrences of trending keywords in popular tech media (01.2016-04.2021)

<p>Sources with weights</p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5&nbsp;%</li> <li>IEEE Spectrum 5&nbsp;%</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p>Methodology</p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues have been selected (e.g. &#39;metoo&#39;, &#39;gdpr&#39;)</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Keyword frequencies in popular tech media (01.2016-04.2021)

<p>Sources with weights</p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5&nbsp;%</li> <li>IEEE Spectrum 5&nbsp;%</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p>Methodology</p> <ul> <li>Frequency of appearances for all unigrams and bigrams in the texts</li> <li>Frequency: number of appearances of every term divided by the number of all terms&nbsp;(for every month and source)&nbsp;</li> <li>Several media sources: a representative index is calculated with weighted average (weights as above)</li> <li>Average monthly change in the analised term&#39;s frequency is calculated by OLS regressions</li> <li>The dependent variable of the estimation is the frequency index, while the number of months since the beginning of the analysed period (January 2016) is the independent variable</li> <li>The regression coefficient (referred to as coef) shows by how much on average the analysed expression&rsquo;s frequency changed with every observed month (marginal change of the frequency), revealing which keywords had the biggest monthly growth</li> </ul> <p>Columns</p> <p>freq_months (e.g. freq_2019-04):&nbsp;the average frequency of the term</p> <p>coef:&nbsp;the regression coefficient</p> <p>coef_norm:&nbsp;the regression coefficient divided by the mean frequency of the keyword</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Co-occurrences of trending keywords in popular tech media during the COVID-19 pandemic (01.2020-06.2020)

<p>Sources:&nbsp;</p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe&nbsp;</li> <li>IEEE Spectrum&nbsp;</li> <li>Techforge&nbsp;</li> <li>Fastcompany&nbsp;</li> <li>The Guardian (Tech)&nbsp;</li> <li>Arstechnica&nbsp;</li> <li>Reuters&nbsp;</li> <li>Gizmodo&nbsp;</li> <li>ZDNet&nbsp;</li> <li>The Register&nbsp;</li> <li>The Verge&nbsp;</li> <li>TechCrunch&nbsp;</li> </ul> <p>Methodology</p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues and technologies have been selected (e.g. covid19)</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Keyword frequencies in popular tech media during the COVID-19 pandemic (01.2020-06.2020)

<p>Sources:&nbsp;</p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe&nbsp;</li> <li>IEEE Spectrum&nbsp;</li> <li>Techforge&nbsp;</li> <li>Fastcompany&nbsp;</li> <li>The Guardian (Tech)&nbsp;</li> <li>Arstechnica&nbsp;</li> <li>Reuters&nbsp;</li> <li>Gizmodo&nbsp;</li> <li>ZDNet&nbsp;</li> <li>The Register&nbsp;</li> <li>The Verge&nbsp;</li> <li>TechCrunch&nbsp;</li> </ul> <p>Methodology is modified relative to the regular trend analysis due to the short period of analysis (weekly freqiencies)</p> <ul> <li>Frequency of appearances for all unigrams and bigrams in the texts</li> <li>Frequency: number of appearances of every term divided by the number of all terms&nbsp;(for every week)&nbsp;</li> <li>Several media sources: all articles are treated equally</li> <li>Average monthly change in the analised term&#39;s frequency is calculated by OLS regressions</li> <li>The dependent variable of the estimation is the frequency index, while the number of weeks since the beginning of the analysed period (January 2020) is the independent variable</li> <li>The regression coefficient (referred to as coef) shows by how much on average the analysed expression&rsquo;s frequency changed with every observed week&nbsp;(marginal change of the frequency), revealing which keywords had the biggest weekly&nbsp;growth</li> </ul> <p>Columns</p> <p>freq_2020_weeks&nbsp;(e.g. freq_2020_ww0):&nbsp;the average frequency of the term</p> <p>coef:&nbsp;the regression coefficient</p> <p>coef_norm:&nbsp;the regression coefficient divided by the mean frequency of the keyword</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Data from: Tracking the popularity and outcomes of all bioRxiv preprints

<p>The data used to generate figures in the manuscript titled <a href="https://www.biorxiv.org/content/early/2019/01/13/515643">&quot;Tracking the popularity and outcome of all bioRxiv preprints,&quot;</a> posted to bioRxiv 13 Jan 2019.</p> <ul> <li><strong>22 Mar 2019:</strong> PDFs of each figure from the paper have been added to the repository. In addition, the license has been changed from CC-BY-NC to CC0.</li> </ul>

opencc-zeroJan 2019View details →
zenodo44/100

Co-occurrences of trending keywords in popular tech media

<p><strong>Co-occurrences of trending keywords in the tech media (01.2016-03.2019)</strong></p> <p><strong>Sources</strong></p> <ul> <li>Gigaom 0.5%</li> <li>Euractiv 0.9%</li> <li>The Conversation 1.3%</li> <li>Politico Europe 1.3%</li> <li>IEEE Spectrum 1.8%</li> <li>Techforge 4.3%</li> <li>Fastcompany 4.5%</li> <li>The Guardian (Tech) 9.2%</li> <li>Arstechnica 10.0%</li> <li>Reuters 11%</li> <li>Gizmodo 17.5%</li> <li>ZDNet 18.3%</li> <li>The Register 19.5%</li> </ul> <p><strong>Methodology</strong></p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues have been selected (e.g. &#39;metoo&#39;, &#39;gdpr&#39;)</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul> <p><strong>Files</strong></p> <p>unigram-unigram co-occurrences: cooc11weighted.csv</p> <p>unigram-bigram co-occurrences: cooc12weighted.csv</p> <p>bigram-unigram co-occurrences: cooc21weighted.csv</p> <p>bigram-bigram co-occurrences: cooc22weighted.csv<br> &nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2019View details →
zenodo44/100

Keyword frequency in popular tech media

<p><strong>Keywords trending in the tech media (01.2016-03.2019)</strong></p> <p><strong>Sources</strong></p> <ul> <li>Gigaom 0.5%</li> <li>Euractiv 0.9%</li> <li>The Conversation 1.3%</li> <li>Politico Europe 1.3%</li> <li>IEEE Spectrum 1.8%</li> <li>Techforge 4.3%</li> <li>Fastcompany 4.5%</li> <li>The Guardian (Tech) 9.2%</li> <li>Arstechnica 10.0%</li> <li>Reuters 11%</li> <li>Gizmodo 17.5%</li> <li>ZDNet 18.3%</li> <li>The Register 19.5%</li> </ul> <p><strong>Methodology</strong></p> <ul> <li>Frequency of appearances for all unigrams and bigrams in the texts</li> <li>Frequency: number of appearances of every term divided by the number of published articles (for every month and source)</li> <li>This measure reveals how many times an expression has been mentioned on average per article</li> <li>Several media sources: a representative index is calculated with weighted average</li> <li>Average monthly change in the analised term&#39;s frequency is calculated by OLS regressions</li> <li>The dependent variable of the estimation is the frequency index, while the number of months since the beginning of the analysed period (January 2016) is the independent variable</li> <li>The regression coefficient (referred to as coef) shows by how much on average the analysed expression&rsquo;s frequency changed with every observed month (marginal change of the frequency), revealing which keywords had the biggest monthly growth</li> </ul> <p><strong>Files</strong></p> <ul> <li>unigrams: coefs_1weighted_site.csv</li> <li>bigrams: coefs_2weighted_site.csv</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jan 2019View details →
zenodo40/100

FIG. 5 in Early illustrations of Aepyornis eggs (1851 - 1887): from popular science to Marco Polo's roc bird

FIG. 5. — Comparative dimensions of various bird eggs. From Capus (1885). Aepyornis egg (4) compared with eggs of an ostrich (3), a hen (2) and a hummingbird (1). Photo E. Buffetaut.

opencc-zeroSep 2019View details →
zenodo40/100

FIG. 4 in Early illustrations of Aepyornis eggs (1851 - 1887): from popular science to Marco Polo's roc bird

FIG. 4. — "The Ruc's egg". Colour lithograph, frontispiece of The Book of Ser Marco Polo, the Venetian, Concerning the Kingdoms and Marvels of the East. Volume 2 (Yule 1871). Photo E. Buffetaut.

opencc-zeroSep 2019View details →
zenodo40/100

FIG. 6 in Early illustrations of Aepyornis eggs (1851 - 1887): from popular science to Marco Polo's roc bird

FIG. 6. — "Egg of the Epiornis" [sic]. From Scientific American, March 12, 1887. Photo E. Buffetaut.

opencc-zeroSep 2019View details →
zenodo40/100

FIG. 1 in Early illustrations of Aepyornis eggs (1851 - 1887): from popular science to Marco Polo's roc bird

FIG. 1. – Egg of Aepyornis maximus Geoffroy Saint-Hilaire, 1851, from Rowley (1878). The original plate shows the egg (from Rowley's collection) at its actual size. This appears to be the first illustration of an Aepyornis egg in a scientific paper. Photo E. Buffetaut.

opencc-zeroSep 2019View details →
zenodo40/100

FIG. 3 in Early illustrations of Aepyornis eggs (1851 - 1887): from popular science to Marco Polo's roc bird

FIG. 3. — Egg of Aepiornis [sic] maximus, compared with a hen's egg. From Ward (1866). Photo E. Buffetaut.

opencc-zeroSep 2019View details →
zenodo40/100

Popularity Dataset for Online Stats Training

<p>This is a dataset&nbsp;used for the online stats training website (<a href="https://www.rensvandeschoot.com/tutorials/">https://www.rensvandeschoot.com/tutorials/</a>) and is based on the data used by&nbsp;&nbsp;<a href="https://doi.org/10.1016/j.adolescence.2009.12.004">van de Schoot, van der Velden, Boom, and Brugman (2010)</a>.</p> <p>The dataset is based on a study that investigates an association between popularity status and antisocial behavior from at-risk adolescents (n = 1491), where gender and ethnic background are moderators under the association. The study distinguished subgroups within the popular status group in terms of overt and covert antisocial behavior.For more information on the sample, instruments, methodology, and research context, we refer the interested readers to <a href="https://doi.org/10.1016/j.adolescence.2009.12.004">van de Schoot, van der Velden, Boom, and Brugman (2010)</a>.</p> <p>&nbsp;</p> <p>Variable name&nbsp;&nbsp; Description</p> <p>Respnr =&nbsp; Respondents&rsquo; number</p> <p>Dutch =&nbsp; Respondents&rsquo; ethnic background (0 = Dutch origin, 1 = non-Dutch origin)</p> <p>gender&nbsp; = Respondents&rsquo; gender (0 = boys, 1 = girls)</p> <p>sd =&nbsp;&nbsp;Adolescents&rsquo; socially desirable answering patterns</p> <p>covert =&nbsp;Covert antisocial behavior</p> <p>overt =&nbsp; Overt antisocial behavior</p>

opencc-by-4.0Jul 2020View details →
zenodo40/100

Listening preferences for the different reproduction systems Stereo, Surround, and Wave Field Synthesis in the context of popular music

<p>We did a paired comparison preference test where listeners rated their listening preference for four different pop musical pieces presented by WFS, stereo or surround. The musical pieces were all mixed by the same person in order to try to minimize the influence of the mix on the ratings, but still trying to get the best out of every system, see [1] for details. The mixes are available at https://doi.org/10.14279/depositonce-5173.</p> <p>Here, we provide the results of the 22 listeners that participated in the experiment together with an analysis which calculates a Bradley-Terry-Luce model after Wickelmayer et al. [2].</p> <p>[1] Hold, C., Wierstorf, H., Raake, A. (2016), “The Difference Between Stereophony and Wave Field Synthesis in the Context of Popular Music,” 140th AES Convention, Paper 9533</p> <p>[2] https://cran.r-project.org/web/packages/eba/index.html</p>

opencc-by-4.0Nov 2016View details →
zenodo40/100

An Empirical Study of Activity, Popularity, Size, Testing, and Stability in Continuous Integration

<p>A good understanding of the practices followed by software development projects can positively impact their success --- particularly for attracting talent and on-boarding new members. In this paper, we perform a cluster analysis to classify software projects that follow continuous integration in terms of their activity, popularity, size, testing, and stability. Based on this analysis, we identify and discuss four different groups of repositories that have distinct characteristics that separates them from the other groups.  With this new understanding, we encourage open source projects to acknowledge and advertise their preferences according to these defining characteristics, so that they can recruit developers who share similar values.</p>

opencc-by-4.0May 2017View details →
zenodo40/100

An Exploratory Study of Documentation Strategies for Product Features in Popular GitHub Projects [Replication Package]

<h2>Artefact Summary</h2> <p>This repository contains the replication package for the paper 'An Exploratory Study of Documentation Strategies for Product Features in Popular GitHub Projects,' presented at the <em><a href="https://cyprusconferences.org/icsme2022/" target="_blank" rel="noopener">38th IEEE International Conference on Software Maintenance and Evolution (ICSME'22)</a></em>.</p> <p>The purpose of the package is to facilitate the verification and reproduction of the study results.<br>It provides all computational notebooks used to collect and analyse data, as well as the slides of the conference presentation.</p> <h2>Paper Abstract</h2> <p>[Background] In large open-source software projects, development knowledge is often fragmented across multiple artefacts and contributors such that individual stakeholders are generally unaware of the full breadth of the product features. However, users want to know what the software is capable of, while contributors need to know where to fix, update, and add features. [Objective] This work aims at understanding how feature knowledge is documented in GitHub projects and how it is linked (if at all) to the source code. [Method] We conducted an in-depth qualitative exploratory content analysis of 25 popular GitHub repositories that provided the documentation artefacts recommended by GitHub&rsquo;s Community Standards indicator. We extracted strategies used to document software features in textual artefacts and which strategies were used to link the feature documentation with source code. [Results] We observed feature documentation in all studied projects in artefacts such as READMEs, wikis, and website resource files. However, the features were often described in an unstructured way. Additionally, tracing techniques to connect feature documentation and source code were rarely used. [Conclusions] Our results suggest a lacking (or a low-prioritised) feature documentation in open-source projects, little use of normalised structures, and a rare explicit referencing to source code. As a result, product feature traceability is likely to be very limited, and maintainability to suffer over time.</p> <h2>References</h2> <p>The published paper is available on <a href="https://doi.org/10.1109/ICSME55016.2022.00043" target="_blank" rel="noopener">IEEE Xplore</a> and the preprint on <a href="https://doi.org/10.48550/arXiv.2208.01317" target="_blank" rel="noopener">arXiv</a>.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Representation of crowd accidents in popular media

<p>This repository contains results related to the analysis of a corpus of news reports covering the topic of crowd accidents. To facilitate online visualization and offline analysis, the files are organized by assigning a number to each. The number system and the details of each set of files are described as follows:</p> <ul> <li><strong>Class 0</strong> &ndash; This contains the same files provided in this repository, but they are organized into folders to make analysis easier. If you intend to analyze the data from our lexical analysis, we suggest using this file since it is better organized and can be directly downloaded.</li> <li><strong>Class 1</strong> &ndash; This contains the sources and relevant information for people who are interested in replicating our dataset or accessing the news reports used in our analysis. Please note that due to copyright regulations, the texts cannot be shared. However, you can refer to the links provided in these files to access the news articles and Wikipedia pages. Some links have stopped working during the time we were working on this study, and others may be unreachable in the future.</li> <li><strong>Class 2</strong> &ndash; This contains the results from a lexical analysis of the corpus. The HTML page allows you to visualize each result interactively through the online VOSviewer app (you need to download the file and open it using a browser since Zenodo does not recognize this as a link). It is possible that this service (VOSviewer app) may be discontinued at some point in the future. PNG images of lexical maps are, therefore, available for download through the ZIP archive, although they do not allow interactive access. If you plan to read our results using the offline VOSviewer software or perform a more systematic analysis, JSON files are available for each category (time period, geographical area of the reporting institution, and purpose of gathering). The same files can be also find in the ZIP archive in class 0.</li> <li><strong>Class 3</strong> &ndash; These are the results of the sentiment analysis. For each report, a single result is generated for the title. However, for the body, the text is divided into parts, which are analyzed independently.</li> <li><strong>Class 4</strong> &ndash; These two files contains the corpus of Wikipedia relative to 68 crowd accidents which occurred between 1990 and 2019. The text for all accidents were scraped on October 15th, 2022 (<em>before</em> the tragedy in Itaewon) and on May 25th, 2023 (<em>after</em> the tragedy). Sources relative to the content in Wikipedia are listed in the file contained in Class 1 ("1_list_wiki_report.csv"). More generally, accidents listed on dedicated Wikipedia pages on <a href="https://en.wikipedia.org/wiki/List_of_fatal_crowd_crushes" target="_blank" rel="noopener">https://en.wikipedia.org/wiki/List_of_fatal_crowd_crushes</a> are reported in the corpus provided here (the period 1900-2019 is considered here).</li> </ul> <p>The format of CSV and JSON files should be self-explanatory after reading our publication. For specific questions or queries, please contact one of the authors, and we will try to assist you.</p>

opencc-by-4.0Sep 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record