Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
155
datasets available to search
ShareScore release 0.7.1
Dataset results
155 results for “Popular”
Code to reproduce the figures in the paper 'Listener Preference for Different Reproduction Systems and Mixes in Popular Music'
<p>In this upload you find all the scripts and data you need in order to reproduce<br> the figure from the paper Wierstorf et al., "Listener Preference for Different<br> Reproduction Systems and Mixes in Popular Music" [1].</p> <p>Software Requirements<br> ---------------------</p> <p>For the statistic analysis you will need [python](https://www.python.org) and<br> [R](https://www.r-project.org). I have used python 3.5.2 and R 3.2.3 for<br> published analysis.</p> <p>Under R you need to install the [eba](https://cran.r-project.org/package=eba)<br> package, which implements the Bradley-Terry-Luce model. You can install it in R<br> by running `install.packages("eba")`.</p> <p>Under python you have to install pandas and numpy.</p> <p>Reproduce figures<br> -----------------</p> <p>All figures were plotted using gnuplot 5.0. Every figure folder has an<br> ``figXX.plt`` (replace ``XX`` by the figure number) file that you can execute<br> and you will get the resulting pdf file. For Fig. 5 up to Fig. 9, also a<br> ``figXX.sh`` file is provided, that will rerun the statistical analysis of the<br> data presented in the figures.</p> <p>References<br> ----------</p> <p>[1] H. Wierstorf, C. Hold, A. Raake, "Listener Preference for Different<br> Reproduction Systems and Mixes in Popular Music," J. Audio. Eng. Soc, submitted. <br> </p>
Dataset: Popular, Inc. (BPOP) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Dataset: Popular Capital Trust II PFD GTD 6.125% (BPOPM) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
On the Popularity of Modern Open Source Software
<p><strong>This dataset contains the data analyzed on the paper:</strong></p> <p>Hudson Borges and Marco Tulio Valente. <em>On the Popularity of Modern Open Source Software</em>. Submitted to Journal of Systems and Software (JSS), 2018.</p> <p><strong>Files:</strong></p> <ul> <li><em>cdf.csv</em>: Cumulative distribution function data.</li> <li><em>contributos.csv</em>: List of contributors of the repositories.</li> <li><em>developers_perceptions.csv</em>: Survey of Developers' Perceptions on Growth Patterns.</li> <li><em>factors.[activity,owner,repository].csv</em>: Additional information of the analyzed repositories and their owners.</li> <li><em>growth_patterns.zip</em>: A compressed file containing the output of the KSC algorithm (time series clusters).</li> <li><em>motivations_for_starring.csv</em>: Developers' motivations for starring projects.</li> <li><em>owners.csv</em>: Information of the repositories' owners.</li> <li><em>releases.csv</em>: Releases considered in the study (i.e., major and minor releases only).</li> <li><em>repositories.csv</em>: Information of the analyzed repositories (e.g., stars, forks, owner, creation date, etc.).</li> <li><em>timeseries.json</em>: File containing the number of stars gained by week for each repository since their creation.</li> </ul>
Business popularity in Yelp
<p>Postprocessed Yelp data for business popularity prediction.</p>
Transkripte von elf Video-Ansprachen der Schweizer Regierung vor Volksabstimmungen. Transcripts of Eleven TV Addresses Given by the Swiss Government before Popular Votes
<p>Der Datensatz enthält Transkripte (doc, html, pdf, txt) von elf TV-Ansprachen der Schweizer Regierung vor Volksabstimmungen. Die Ansprachen wurden nach GAT 2 transkribiert. / The dataset contains transcripts (doc, html, pdf, txt) of eleven TV addresses given by the Swiss government before popular votes. The addresses were transcribed according to GAT 2.</p> <p> </p> <p><strong>Quellenangabe der Transkripte / Reference to the Transcripts</strong></p> <p>Schröter, Juliane, Keller, Stefan, 2018. Transkripte von elf Video-Ansprachen der Schweizer Regierung vor Volksabstimmungen. Transcripts of Eleven TV Addresses Given by the Swiss Government before Popular Votes. doi: 10.5281/zenodo.1324476.</p> <p><em>Falls Sie sich nur auf eines oder einige der elf Transkripte beziehen, passen Sie die Quellenangabe bitte entsprechend an. / If you are only referring to one or some of the eleven transcripts, please adopt the reference accordingly. </em></p> <p><em>Disclaimer: Die Mitglieder des Bundesrates haben mündliche Ansprachen gehalten. </em><em>Für den Wortlaut der Transkripte sind sie nicht verantwortlich. / The members of the Federal Council have delivered oral addresses. They are not responsible for the wording of the transcripts.</em></p> <p> </p> <p><strong>Quellenangaben der Videos / References to the Videos</strong></p> <p>Bundesrat, 2017a. [TV-Ansprache zum] Bundesgesetz „Unternehmenssteuerreform III“. Produziert von SRG SSR. <a href="https://www.admin.ch/gov/de/start/dokumentation/abstimmungen/20170212/bundesgesetz-ueber-steuerliche-massnahmen-zur-staerkung-der-wett.html">https://www.admin.ch/gov/de/start/dokumentation/abstimmungen/20170212/bundesgesetz-ueber-steuerliche-massnahmen-zur-staerkung-der-wett.html</a> (Abfrage: 23.05.2018).<br> <br> Bundesrat, 2017b. [TV-Ansprache zum] Energiegesetz. Produziert von SRG SSR. <a href="https://www.admin.ch/gov/de/start/dokumentation/abstimmungen/20170521/Energiegesetz.html">https://www.admin.ch/gov/de/start/dokumentation/abstimmungen/20170521/Energiegesetz.html</a> (Abfrage: 23.05.2018).<br> <br> Bundesrat, 2016. [TV-Ansprache zur] Initiative „Für Ehe und Familie gegen die Heiratsstrafe“. Produziert von SRG SSR. <a href="https://www.srf.ch/play/tv/abstimmungen-teilw--in-gebaerdensprache/video/vorlage-heiratsstrafe-sendung-mit-gebaerdensprache?id=8e482ca4-52e3-4777-b956-529ce64f96d1">https://www.srf.ch/play/tv/abstimmungen-teilw--in-gebaerdensprache/video/vorlage-heiratsstrafe-sendung-mit-gebaerdensprache?id=8e482ca4-52e3-4777-b956-529ce64f96d1</a> (Abfrage: 23.05.2018).<br> <br> Bundesrat, 2015a. [TV-Ansprache zur] Änderung des Bundesgesetzes über Radio und Fernsehen. Produziert von SRG SSR. <a href="https://www.srf.ch/play/tv/abstimmungen/video/br-ansprache-zum-rtvg?id=dd58817b-235b-472a-96a2-a67c721e6775&station=69e8ac16-4327-4af4-b873-fd5cd6e895a7">https://www.srf.ch/play/tv/abstimmungen/video/br-ansprache-zum-rtvg?id=dd58817b-235b-472a-96a2-a67c721e6775&station=69e8ac16-4327-4af4-b873-fd5cd6e895a7</a> (Abfrage: 23.05.2018).<br> <br> Bundesrat, 2015b. [TV-Ansprache zur] Präimplantationsdiagnostik. Produziert von SRG SSR. <a href="https://www.srf.ch/play/tv/ansprachen-bundesrat-in-gebaerdensprache/video/br-ueli-maurer-zum-fortpflanzungsmedizingesetz-fmedg-geb-?id=cbbecc0d-c210-41a9-8190-4d282926c3a8&station=69e8ac16-4327-4af4-b873-fd5cd6e895a7">https://www.srf.ch/play/tv/ansprachen-bundesrat-in-gebaerdensprache/video/br-ueli-maurer-zum-fortpflanzungsmedizingesetz-fmedg-geb-?id=cbbecc0d-c210-41a9-8190-4d282926c3a8&station=69e8ac16-4327-4af4-b873-fd5cd6e895a7</a> (Abfrage: 23.05.2018).<br> <br> Bundesrat, 2014a. [TV-Ansprache zur] Beschaffung des Kampfflugzeuges Gripen. Produziert von SRG SSR. <a href="https://www.srf.ch/play/tv/abstimmungen/video/bundesrat-ueli-maurer-zur-beschaffung-des-kampfflugzeuges-gripen?id=febd8e03-d4c3-42e7-b163-37e9e0d0176c&station=69e8ac16-4327-4af4-b873-fd5cd6e895a7">https://www.srf.ch/play/tv/abstimmungen/video/bundesrat-ueli-maurer-zur-beschaffung-des-kampfflugzeuges-gripen?id=febd8e03-d4c3-42e7-b163-37e9e0d0176c&station=69e8ac16-4327-4af4-b873-fd5cd6e895a7</a> (Abfrage: 23.05.2018).<br> <br> Bundesrat, 2014b. [TV-Ansprache zur] Initiative „Für den Schutz fairer Löhne“. Produziert von SRG SSR. <a href="https://www.srf.ch/play/tv/abstimmungen-teilw--in-gebaerdensprache/video/ansprache-von-bundesrat-johann-schneider-ammann-vom-20-04-2014?id=3ea16bd4-401d-4d68-9d06-b07b83582405">https://www.srf.ch/play/tv/abstimmungen-teilw--in-gebaerdensprache/video/ansprache-von-bundesrat-johann-schneider-ammann-vom-20-04-2014?id=3ea16bd4-401d-4d68-9d06-b07b83582405</a> (Abfrage: 23.05.2018).<br> <br> Bundesrat, 2014c. [TV-Ansprache zur] Initiative „Gegen Masseneinwanderung“. Produziert von SRG SSR. <a href="https://www.youtube.com/watch?v=A91gPrFXFCs">https://www.youtube.com/watch?v=A91gPrFXFCs</a> (Abfrage: 23.05.2018).<br> <br> Bundesrat, 2014d. [TV-Ansprache zur] Initiative „Pädophile sollen nicht mehr mit Kindern arbeiten dürfen“. Produziert von SRG SSR. <a href="https://www.bk.admin.ch/bk/de/home/dokumentation/volksabstimmungen/volksabstimmung-20140518.html">https://www.bk.admin.ch/bk/de/home/dokumentation/volksabstimmungen/volksabstimmung-20140518.html</a> (Abfrage: 18.04.2018).<br> <br> Bundesrat, 2013. [TV-Ansprache zum] Bundesbeschluss über die Familienpolitik. Produziert von SRG SSR. Video bereitgestellt von SRG SSR.<br> <br> Bundesrat, 2010. [TV-Ansprache] Zur Ausschaffungsinitiative und zum Gegenentwurf des Bundesrates. Produziert von SRG SSR. Video bereitgestellt von SRG SSR.</p> <p> </p> <p><strong>Quellenangabe des Transkriptionssystems / Reference to the Conventions of Transcription</strong></p> <p>Selting, Margret, Auer, Peter, Barth-Weingarten, Dagmar et al., 2009. Gesprächsanalytisches Transkriptionssystem 2 (GAT 2). Gesprächsforschung 10, 353-402.</p>
Maven 99 most popular library statical usages
<p>A SQL database containing the static usages of API elements of any version of the 99 most used maven artifact, by any of it client on maven central.</p>
Co-occurrences of trending keywords in popular tech media (01.2016-04.2019)
<p>Sources with weights</p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5 %</li> <li>IEEE Spectrum 5 %</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p>Methodology</p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues have been selected (e.g. 'metoo', 'gdpr')</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul> <p> </p>
Keyword frequencies in popular tech media (01.2016-04.2019)
<p><strong>Sources with weights</strong></p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5 %</li> <li>IEEE Spectrum 5 %</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p><strong>Methodology</strong></p> <ul> <li>Frequency of appearances for all unigrams and bigrams in the texts</li> <li>Frequency: number of appearances of every term divided by the number of published articles (for every month and source)</li> <li>This measure reveals how many times an expression has been mentioned on average per article</li> <li>Several media sources: a representative index is calculated with weighted average (weights as above)</li> <li>Average monthly change in the analised term's frequency is calculated by OLS regressions</li> <li>The dependent variable of the estimation is the frequency index, while the number of months since the beginning of the analysed period (January 2016) is the independent variable</li> <li>The regression coefficient (referred to as coef) shows by how much on average the analysed expression’s frequency changed with every observed month (marginal change of the frequency), revealing which keywords had the biggest monthly growth</li> </ul> <p><strong>Files</strong></p> <p>The dataset contains two files:</p> <p>Unigrams: coefs_1weighted_site.csv</p> <p>Bigrams: coefs_2weighted_site.csv</p> <p><strong>Columns</strong></p> <p>freq_months (e.g. freq_2019-04): the average frequency of the term</p> <p>coef: the regression coefficient</p> <p>coef_norm: the regression coefficient divided by the mean frequency of the keyword</p> <p>coef_norm_max: the regression coefficient divided by the maximum frequency of the keyword</p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p>
How Confidence in Prior Attitudes, Social Tag Popularity, and Source Credibility Shape Confirmation Bias Toward Antidepressants and Psychotherapy in a Representative German Sample: Randomized Controlled Web-Based Study
<p>ABSTRACT</p> <p>Background: In health-related, Web-based information search, people should select information in line with expert (vs nonexpert) information, independent of their prior attitudes and consequent confirmation bias.</p> <p>Objective: This study aimed to investigate confirmation bias in mental health–related information search, particularly (1) if high confidence worsens confirmation bias, (2) if social tags eliminate the influence of prior attitudes, and (3) if people successfully distinguish high and low source credibility.</p> <p>Methods: In total, 520 participants of a representative sample of the German Web-based population were recruited via a panel company. Among them, 48.1% (250/520) participants completed the fully automated study. Participants provided <em>prior attitudes</em> about antidepressants and psychotherapy. We manipulated (1) <em>confidence</em> in prior attitudes when participants searched for blog posts about the treatment of depression, (2) <em>tag popularity</em> —either psychotherapy or antidepressant tags were more popular, and (3) <em>source credibility</em> with banners indicating high or low expertise of the tagging community. We measured <em>tag</em> and <em>blog post</em> selection, and <em>treatment</em><em>efficacy ratings</em> after navigation.</p> <p>Results: Tag popularity predicted the proportion of selected antidepressant tags (beta=.44, SE 0.11; <em>P</em><.001) and blog posts (beta=.46, SE 0.11; <em>P</em><.001). When confidence was low (−1 SD), participants selected more blog posts consistent with prior attitudes (beta=−.26, SE 0.05; <em>P</em><.001). Moreover, when confidence was low (−1 SD) and source credibility was high (+1 SD), the efficacy ratings of attitude-consistent treatments increased (beta=.34, SE 0.13; <em>P</em>=.01).</p> <p>Conclusions: We found correlational support for defense motivation account underlying confirmation bias in the mental health–related search context. That is, participants tended to select information that supported their prior attitudes, which is not in line with the current scientific evidence. Implications for presenting persuasive Web-based information are also discussed.</p> <p>Trial Registration: ClinicalTrials.gov NCT03899168; https://clinicaltrials.gov/ct2/show/NCT03899168 (Archived by WebCite at http://www.webcitation.org/77Nyot3Do)</p> <p>J Med Internet Res 2019;21(4):e11081</p> <p>doi:10.2196/11081</p>
Most popular scholarly works in the English Wikipedia and their transition to open access
<p>Following the release of "The future of OA" by Piwowar, Priem, Orr (2019), interest has grown on how to accelerate the share of scholarly works consultations which meet an open access record.</p> <p>Based on download patterns for over 23 million DOIs in 2017, released by Elbakyan (2018), we found that the 1 million most downloaded DOIs accounted for over 30 % of the total downloads. Of these 1 million DOIs, over 50 thousands (5 %) were previously identified as cited on the English Wikipedia and not open access (Leva 2018). Of these, 2440 DOIs are now open access according to the Unpaywall API as of 2019-10-25: a list of the corresponding OA URL and host type is enclosed, showing that 34 % became OA at the publisher while 66 % were made OA by a repository. The newly OA works were hosted at over 400 domains of which over 300 repositories, but the top 10 repositories accounted for a large portion of the works, with the top 3 repositories accounting for over 40 % of the newly found green open access DOIs.</p> <p>Part of the newly OA works were just false negatives in Unpaywall in 2018, but a small manual sample shows that most are truly new deposits. Works from 2017 can be expected to be over-represented in the sample given that they were probably the most popular downloads of 2017 and could have been under embargo in 2018 when the previous measure of open access status was made.</p>
Figure 1 in Domestication level of the most popular aquarium fish species: is the aquarium trade dependent on wild populations?
Figure 1. – Number of aquarium fish species per domestication level (white: freshwater, n = 50 species; black: marine, n = 50 species).
Metadatos y portadas de los juegos más nuevos y populares en la plataforma Steam
<p><span>Este dataset contiene un ranking de los juegos más destacados en la plataforma Steam por género y subgénero en una fecha en específico. A su vez contiene el listado de géneros y subgéneros, la metainformación de cada juego y la imagen de portada del mismo. El scrapeo se ha realizado el 10 de noviembre de 2024. Es posible ejecutar el código de manera periódica para ir recopilando datos históricos.</span></p>
Información sobre los juegos más nuevos y populares por género de la plataforma Steam
<p><span>El conjunto de datos obtenido en el presente proyecto se compone de un dataset principal, dos datasets complementarios, y una carpeta que contiene imágenes. </span></p> <p> </p> <ul> <li><span>El dataset principal contiene información detallada de una variedad de videojuegos como el nombre, la fecha de lanzamiento, y el precio, entre otros.</span></li> <li><span>El primero de los datasets complementarios contiene un listado de los géneros y subgéneros que Steam considera como más destacados. Entre los que se encuentran acción, aventuras, y estrategia, entre otros. </span></li> <li><span>El segundo dataset complementario contiene un listado de los juegos más recientes y populares por cada género y subgénero. </span></li> <li><span>Por último, la carpeta de imágenes contiene la portada de cada uno de los videojuegos. </span></li> </ul>
Características de películas y programas de televisión más populares en la base de datos de Imdb
<p>Se ha realizado una extracción de datos a través de técnicas de web scraping en la web Imdb, con las películas y programas de televisión más populares distribuidos por género.</p> <p>El dataset cuenta con información referente a las películas y programas de televisión más populares según la comunidad cinéfila de Imdb. Esta información se puede utilizar para clasificar estas películas entre las más votadas, las mejores valoradas, las que más actores aparecen, las se pueden enmarcar en más tipos de géneros o incluso saber el género que presenta las películas peor valoradas. Además, el dataset se ha construido con solo las primeras cincuenta películas más populares de cada género ya que la web contiene más de 2 millones de títulos y no nos interesa tener un dataset tan grande para su posterior tratamiento.</p> <p>Toda la información que se ha recogido se presenta en un fichero CSV para facilitar su posterior limpieza y análisis en la siguiente práctica.</p>
Datasets to Evaluate Accuracy, Miscalibration and Popularity Lift in Recommendations
<p>This repository contains three datasets for evaluating accuracy, miscalibration and popularity lift in recommender systems. All datasets contain genre/category information in addition to different user group splits:</p> <ol> <li>Last.fm (lfm.zip), based on the LFM-1b dataset of JKU Linz (http://www.cp.jku.at/datasets/LFM-1b/)</li> <li>MovieLens (ml.zip), based on MovieLens-1M dataset (https://grouplens.org/datasets/movielens/1m/)</li> <li>MyAnimeList (anime.zip), based on the MyAnimeList dataset of Kaggle (https://www.kaggle.com/CooperUnion/anime-recommendations-database)</li> </ol> <p>'user_events_cats.txt' contains the users' rating/interaction data along with a list of genres/categories assigend to the rated items. The list of categories is given in 'categories.txt'. Additionally, assignments to three user groups that differ in their inclination to popular/mainstream items are provided: LowPop in 'low_main_users.txt', MedPop in 'med_main_users.txt', and HighPop in 'high_main_users.txt'.</p> <p>The format of the three user files are "user,mainstreaminess"</p> <p>The format of the user-events files are "user,item,preference,cats", where different categories are separated by '|'</p> <p>The format of the categories files are "category-name,index", where index refers to the category-id in the user-events files</p> <p>Example Python-code for analyzing the datasets as well as empirical results on calibration, popularity lift and accuracy can be found on GitHub: https://github.com/domkowald/FairRecSys</p>
DFRWS EU '23: Hamming Distributions of Popular Perceptual Hashing Techniques - DATASET
<p><strong>Dataset Purpose and Citation</strong></p> <p>This repository contains raw data and plots for the experimental work in the paper:</p> <p>McKeown, S., Buchanan, WJ. (2023 - In Press). Hamming Distributions of Popular Perceptual Hashing Techniques. DFRWS EU 2023, Bonn, Germany.</p> <p>Six perceptual hashes are evaluated on the Flickr 1 Million dataset against a variety of content-preserving attacks in order to better understand their overall behaviour.</p> <p><strong>Approach</strong></p> <p>Hashes for each algorithm and modification (listed below) are created and compared using the Normalised Hamming distance. Three main comparisons are done:</p> <ul> <li> <p>Inter-score originals (original unrelated images in Flickr 1 Million)</p> </li> <li> <p>Inter-score modified (modified unrelated versions of the Flickr images compared to each other within class, e.g. cropped to cropped)</p> </li> <li> <p>Intra-score (original to modified version of the same image)</p> </li> </ul> <p><strong>Repository Contents</strong></p> <p>The top-level of the repository contains the SHA256 hashes of the Flickr 1 Million dataset used for exact file deduplication.</p> <p>There are then six sub-folders, one for each perceptual hashing algorithm, containing:</p> <ul> <li> <p>Zip files of the raw hashes generated by the algorithm for the original Flickr 1 Million dataset</p> </li> <li> <p>A separate Inter and Intra score sub-directory containing:</p> </li> </ul> <p>-- Zipped CSV files containing the Normalised Hamming distance for comparison files</p> <p>-- An analysis folder containing plotted histograms and derived statistical information (e.g. percentiles, mean)</p> <p>The analysis folders were largely used to generate the Figures and Tables in the paper, however more of them are present here than was possible to include in the conference paper.</p> <p>It should be noted that 50 million comparison samples are taken for the inter-score original comparisons (out of a possible 500 billion), while only 250k modifications were generated, resulting in a smaller pool of comparisons for the modified inter-scores and original to modified intra-scores.</p> <p><strong>Tested Algorithms</strong></p> <ul> <li> <p>Blockhash - <a href="https://github.com/commonsmachinery/blockhash">Commons Machinery</a></p> </li> <li> <p>ColourHash - from the <a href="https://github.com/JohannesBuchner/imagehash">Python ImageHash Libary</a></p> </li> <li> <p>NeuralHash - Apple's models extracted using <a href="https://github.com/AsuharietYgvar/AppleNeuralHash2ONNX">A. Ygvar's method</a></p> </li> <li> <p>PDQ - from <a href="https://github.com/facebook/ThreatExchange/tree/main/pdq/python">Facebook's ThreatExchange</a></p> </li> <li> <p>Phash - from the <a href="https://github.com/JohannesBuchner/imagehash">Python ImageHash Libary</a></p> </li> <li> <p>Wavehash - from the <a href="https://github.com/JohannesBuchner/imagehash">Python ImageHash Libary</a></p> </li> </ul> <p><strong>Tested Modifications/Attacks</strong></p> <ul> <li> <p>Border (30 pixel, black)</p> </li> <li> <p>Compression (quality level 30 JPEG)</p> </li> <li> <p>Crop (5% around all edges)</p> </li> <li> <p>Mirror (x-axis)</p> </li> <li> <p>Scaling (1.5x)</p> </li> <li> <p>Thumbnails (Windows 96 pixel, see <a href="https://commons.erau.edu/jdfsl/vol14/iss3/1/">McKeown et al. 2019</a>)</p> </li> <li> <p>Watermarking (Image added to bottom-right corner, 10% of image height or minimum of 40 pixels)</p> </li> </ul>
Dataset for the paper "Website Fingerprinting: Attacking Popular Privacy Enhancing Technologies with the Multinomial Naïve-Bayes Classifier"
<p>This dataset contains website fingerprints of 775 websites analyzed in the paper "Website Fingerprinting: Attacking Popular Privacy Enhancing Technologies with the Multinomial Naïve-Bayes Classifier" published in the Proceedings of the 2009 ACM workshop on Cloud computing security (CCSW 2009, DOI: 10.1145/1655008.1655013).</p>
Spacial and temporal genetic pattern of Semaprochilodus insignis (Prochilodontidae), the most popular fish from the Amazon basin
<p>Dataset of 8 genotyped microsatellite loci for 180 individuals of Semaprochilodus insignis from 11 locations in the Amazon basin.</p>
Co-occurrences of trending keywords in popular tech media (01.2016-02.2020)
<p><strong>Sources with weights</strong></p> <pre> Arstechnica: 1/8, Euractiv: 1/8, Fastcompany: 1/8, The Register: 1/8, Techcrunch: 1/8, The Guardian: 1/8, Venturebeat: 1/8, The Verge: 1/8</pre> <p><strong>Methodology</strong></p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues and technologies have been selected (e.g. 'gdpr', '5G')</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.