Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
47
datasets available to search
ShareScore release 0.9.0
Dataset results
47 results for “instagram”
Dataframe - Posts-Instagram-Torrevieja-Colores dominantes
<p>Una base de datos con extracción de colores dominantes de los post de instagram en torrevijea</p>
A set of generated Instagram Data Download Packages (DDPs) to investigate their structure and content
<p><strong>Instagram data-download example dataset</strong></p> <p>In this repository you can find a data-set consisting of 11 personal Instagram archives, or Data-Download Packages (DDPs).</p> <p> </p> <p><strong>How the data was generated</strong></p> <p>These Instagram accounts were all new and generated by a group of researchers who were interested to figure out in detail<br> the structure and variety in structure of these Instagram DDPs. The participants user the Instagram account extensively for approximately a week. The participants also intensively communicated with each other so that the data can be used as an example of a network. </p> <p>The data was primarily generated to evaluate the performance of de-identification software. Therefore, the text in the DDPs particularly contain many randomly chosen (Dutch) first names, phone numbers, e-mail addresses and URLS. In addition, the images in the DDPs contain many faces and text as well. The DDPs contain faces and text (usernames) of third parties. However, only content of so-called `professional accounts' are shared, such as accounts of famous individuals or institutions who self-consciously and actively seek publicity, and these sources are easily publicly available. Furthermore, the DDPs do not contain sensitive personal data of these individuals. </p> <p><br> <strong>Obtaining your Instagram DDP</strong></p> <p>After using the Instagram accounts intensively for approximately a week, the participants requested their personal Instagram DDPs by using the following steps. You can follow these steps yourself if you are interested in your personal Instagram DDP. </p> <p>1. Go to www.instagram.com and log in<br> 2. Click on your profile picture, go to *Settings* and *Privacy and Security*<br> 3. Scroll to *Data download* and click *Request download*<br> 4. Enter your email adress and click *Next*<br> 5. Enter your password and click *Request download*</p> <p>Instagram then delivered the data in a compressed zip folder with the format **username_YYYYMMDD.zip** (i.e., Instagram handle and date of download) to the participant, and the participants shared these DDPs with us.</p> <p> </p> <p><strong>Data cleaning</strong></p> <p>To comply with the Instagram user agreement, participants shared their full name, phone number and e-mail address. In addition, Instagram logged the i.p. addresses the participant used during their active period on Instagram. After colleting the DDPs, we manually replaced such information with random replacements such that the DDps shared here do not contain any personal data of the participants.</p> <p> </p> <p><strong>How this data-set can be used</strong></p> <p>This data-set was generated with the intention to evaluate the performance of the de-identification software. We invite other researchers to use this data-set for example to investigate what type of data can be found in Instagram DDPs or to investigate the structure of Instagram DDPs. The packages can also be used for example data-analyses, although no substantive research questions can be answered using this data as the data does not reflect how research subjects behave `in the wild'. </p> <p><br> <strong>Authors</strong></p> <p>The data collection is executed by Laura Boeschoten, Ruben van den Goorbergh and Daniel Oberski of Utrecht University. For questions, please contact l.boeschoten@uu.nl. </p> <p> </p> <p><strong>Acknowledgments</strong></p> <p>The researchers would like to thank everyone who participated in this data-generation project.</p>
Engagement and Purchase Intention in storydoing and storytelling for Instagram ads
<p>This is a database with which we have worked on the engagement and purchase intention generated by storydoing and storytelling advertising.</p>
Sponsored and Unsponsored Posts on Instagram
<p>The dataset release comprises 2 Comma Separated Values (CSV) files, one for sponsored posts and another for unsponsored posts collected between 4/15/20 and 5/15/20. The first file "sponsored_posts.csv" contains 29,425 rows with each row representing a post. The second file "unsponsored_posts.csv" contains 28,779 rows with each row representing a post. For each post, the following attributes are collected as columns in each file:</p> <ol> <li>user_id: The Instagram ID of the user</li> <li>user_followers:The number of out-degree nodes for the user</li> <li>likes: The number of likes for the post</li> <li>comments: The number of comments for the post</li> <li>id: The Instagram ID of the Post</li> </ol>
Instagram: #erasurepoetry query
<p><strong>Tags: #erasurepoetry vs. #blackoutpoetry on Instagram (Nov. 19, 2020) </strong></p> <p>As of November 19, 2020, searching on Instagram for the tags “#erasurepoetry” and “#blackoutpoetry” shows some overlapping of techniques in both tags.</p>
Instagram data set BalanceTonPorc
<p>The #MeToo campaign had different moments and expressions online and offline in France. The original #BalanceTonPorc continued to spread on various digital platforms. On Instagram, we identified 7 hashtags derived from #BalanceTonPorc that had at least 250 posts: #BalanceTonBahut, #BalanceTonBar, #BalanceTonHosto, #BalanceTonQuoi, #BalanceTonRappeur, #BalanceTonTiktokeur and #BalanceTonYoutubeur.</p> <p>We downloaded 5,797 posts to synthesize a network of 14,478 hashtags, linked by 267,507 edges that indicate the number of times each pair of hashtags is in the same post. </p> <p>The resulting dataset is a list of adjacencies with hashtags that have co-occurred in the conversation, together with the variable "weight", which indicates the number of times each combination of hashtags has co-occurred. It should be noted that being an undirected network, each combination is given in the table in a unique way, regardless of the position of the values.</p>
Dataset for the Instagram and TikTok problematic use
<p>This dataset supports research on how engagement with social media (Instagram and TikTok) was related to problematic social media use (PSMU) and mental well-being. There are three different files. The SPSS and Excel spreadsheet files include the same dataset but in a different format. The SPSS output presents the data analysis in regard to the difference between Instagram and TikTok users.</p>
The Olympic gold medalists and Instagram - A longitudinal study on user characteristics
<p>This dataset includes Instagram user characteristics of those Olympic athletes who won gold medals in the individual events of Rio2016. The name of all these gold medalists of individual events are in the dataset (226 athletes), however only 144 athletes (83 men and 61 women) had a publicly available Instagram account in all of the observations during the 4 months period of data gathering. The first round of data gathering (first observation, i.e. OlympicAthletesData_1) took place 9-Aug-2019 to 12-Aug-2019, the second round of data gathering (second observation, i.e. OlympicAthletesData_2) took place 9-Sep-2019 to 12-Sep-2019, the third round of data gathering (third observation, i.e. OlympicAthletesData_3) took place 9-Oct-2019 to 12-Oct-2019, the fourth round of data gathering (fourth observation, i.e. OlympicAthletesData_4) took place 9-Nov-2019 to 12-Nov-2019. The data gathered for each user (in each observation) consists of:</p> <p>1- Name of the individual event </p> <p>2- Country</p> <p>3- Name</p> <p>4- Gender</p> <p>5- Instagram ID</p> <p>6- Number of Posts</p> <p>7- Number of followers</p> <p>8- Number of followings</p> <p>9- Maximum Number of likes (in the last 10 photo posts)</p> <p>10- Number of comments for the post with Maximum Number of likes (in the last 10 photo posts)</p> <p>11- Number of self-presenting posts in the last 10 photo posts (those posts in which the athlete is present)</p> <p>12- Number of pure self-presenting posts in the last 10 photo posts (those posts in which the athlete is the only person who is present)</p> <p>13- Age</p> <p>14- Date of data crawling</p>
Anonymized Instagram network data from Amsterdam and Copenhagen, Pajek format
<p>Networks of reciprocated recognition (mutual liking and/or commenting) among Instagram users in Amsterdam and Copenhagen, on the basis of data collected over a twelve-week period in 2015. User names are hashed to anonymize the data.</p>
Dynamics of Instagram Users
<p>These two data sets are gathered from Instagram users who were chosen randomly.</p> <p>The Main data set encompasses data for 1K users including 500 men and 500 women. The Test data set encompasses data for 100 users including 50 men and 50 women.</p> <p>Data gathered for each user includes :</p> <p>1- number of posts</p> <p>2- number of followers</p> <p>3- number of followings</p> <p>4- number of likes for the tenth previous post</p> <p>5- number of likes for the eleventh previous post</p> <p>6- number of likes for the twelfth previous post</p> <p>7- number of self-presenting posts from nine previous posts</p> <p>8- gender</p>
BullyBlocker Instagram Coding Dataset
<p>Annotated Instagram dataset with detailed labels about key cyberbullying properties, such as the content type, purpose, directionality, and co-occurrence with other phenomena.</p>
Recolección de datos de las publicaciones en Instagram de Operación Triunfo 2024
Open the record for dataset details and reuse information.
Anexo 2. Enlace a dataset de la exportación de metadatos de Instagram zeeschuimer-export-instagram.com-2024-05-Suchard_ES, de la sección 4.2.2
<p><span>Documento de Excel con la exportación de datos y el análisis de los posts de Instagram de las campañas “La vida es” (2023) y “La primera navidad” (2022). </span></p>
Time and Dynamics of Instagram Users
<p>These four datasets are gathered from Instagram users who were chosen randomly.</p> <p>The MainDataset encompasses data for 818 users. The TestDataset encompasses data for 78 users.</p> <p>Data gathered for each user includes :</p> <p>1- number of posts</p> <p>2- number of followers</p> <p>3- number of followings</p> <p>4- number of likes for the tenth previous post</p> <p>5- number of likes for the eleventh previous post</p> <p>6- number of likes for the twelfth previous post</p> <p>7- number of self-presenting posts from nine previous posts</p> <p>8- gender</p> <p><br> The MainDataset_after_150_days and TestDataset_after_150_days encompass data of the users of the Main data set and the Test data set, respectively, for after 150 days. For example, User_1 in the MainDataset has 486 posts and in the MainDataset_after_150_days has 562 posts, which means over the course of 150 days he had published 76 posts.</p>
query suggestion with siri - query suggestions for instagram accounts
<p>Das Datenset enthält Query Suggestions zu 20 Instragram-Accounts bzw. Personen oder Firmen dahinter. (Instagram, Cristiano Ronaldo, Ariana Grande, Selena Gomez, The Rock, Kim Kardashian West, Kylie Jenner, Beyoncé, Taylor Swift, Leo Messi, Neymar jr, Kendall Jenner, Justin Bieber, National Geographic, Barbie, Khloe Kardashian, Jennifer Lopez, Miley Cyrus, Nike, Katy Perry)</p> <p>Die Datenerhebung erfolgte vom 26.05.2019 - 23.06.2019. Enthalten sind Vorschläge der Suchmaschinen Bing, DuckDuckGo, Google und der Suche in Siri (an einem Tablet und in einer XCode Simulation), wobei die Anzahl der zurückgelieferten Vorschläge variiert.</p> <p>Das Datenset enthält folgende Spalten: Plattform, Suchbegriff, Datum, Vorschlag Position</p> <p> </p>
Instagram Characteristics of Olympic gold medalists (Rio2016)
<p>This dataset includes Instagram user characteristics of those Olympic athletes who won gold medals in the individual events of Rio2016. The name of all these gold medalists of individual events are in the dataset (226 athletes), however only 149 athletes (85 men and 64 women) had their Instagram publicly available at the time of data crawling (the whole dataset was crawled from 9-Aug-2019 to 12-Aug-2019). Thus, for some athletes we could not present data in the dataset. The data gathered for each user consists of:</p> <p>1- Name of the individual event </p> <p>2- Country</p> <p>3- Name</p> <p>4- Gender</p> <p>5- Instagram ID</p> <p>6- Number of Posts</p> <p>7- Number of followers</p> <p>8- Number of followings</p> <p>9- Maximum Number of likes (in the last 10 photo posts)</p> <p>10- Number of comments for the post with Maximum Number of likes (in the last 10 photo posts)</p> <p>11- Number of self-presenting posts in the last 10 photo posts (those posts in which the athlete is present)</p> <p>12- Number of pure self-presenting posts in the last 10 photo posts (those posts in which the athlete is the only person who is present)</p> <p>13- Age</p> <p>14- Date of data crawling</p>
Mpox Narrative on Instagram: A Labeled Multilingual Dataset of Instagram Posts on Mpox for Sentiment, Hate Speech, and Anxiety Analysis
<p><strong>Please cite the following paper when using this dataset</strong>:</p> <p>N. Thakur, “Mpox narrative on Instagram: A labeled multilingual dataset of Instagram posts on mpox for sentiment, hate speech, and anxiety analysis,” arXiv [cs.LG], 2024, URL: https://arxiv.org/abs/2409.05292</p> <p><strong>Abstract</strong></p> <p>The world is currently experiencing an outbreak of mpox, which has been declared a Public Health Emergency of International Concern by WHO. During recent virus outbreaks, social media platforms have played a crucial role in keeping the global population informed and updated regarding various aspects of the outbreaks. As a result, in the last few years, researchers from different disciplines have focused on the development of social media datasets focusing on different virus outbreaks. No prior work in this field has focused on the development of a dataset of Instagram posts about the mpox outbreak. The work presented in this paper (stated above) aims to address this research gap. It presents this <strong>multilingual dataset of</strong> <strong>60,127 Instagram posts</strong> about mpox, published between <strong>July 23, 2022, and September 5, 2024</strong>. This dataset contains Instagram posts about mpox in <strong>52 languages</strong>. For each of these posts, the Post ID, Post Description, Date of publication, language, and translated version of the post (translation to English was performed using the Google Translate API) are presented as separate attributes in the dataset.</p> <p>After developing this dataset, sentiment analysis, hate speech detection, and anxiety or stress detection were also performed. This process included classifying each post into</p> <ul> <li>one of the fine-grain sentiment classes, i.e., <strong>fear, surprise, joy, sadness, anger, disgust, or neutral</strong>, </li> <li><strong>hate or not hate</strong></li> <li><strong>anxiety/stress detected or no anxiety/stress detected</strong>.</li> </ul> <p>These results are presented as separate attributes in the dataset for the training and testing of machine learning algorithms for sentiment, hate speech, and anxiety or stress detection, as well as for other applications. </p> <p><strong>The 52 distinct languages in which Instagram posts are present in the dataset </strong><strong>are </strong>English, Portuguese, Indonesian, Spanish, Korean, French, Hindi, Finnish, Turkish, Italian, German, Tamil, Urdu, Thai, Arabic, Persian, Tagalog, Dutch, Catalan, Bengali, Marathi, Malayalam, Swahili, Afrikaans, Panjabi, Gujarati, Somali, Lithuanian, Norwegian, Estonian, Swedish, Telugu, Russian, Danish, Slovak, Japanese, Kannada, Polish, Vietnamese, Hebrew, Romanian, Nepali, Czech, Modern Greek, Albanian, Croatian, Slovenian, Bulgarian, Ukrainian, Welsh, Hungarian, and Latvian. </p> <p>The following table represents the data description for this dataset</p> <table> <tbody> <tr> <td> <p><strong>Attribute Name</strong></p> </td> <td> <p><strong>Attribute Description</strong></p> </td> </tr> <tr> <td> <p>Post ID</p> </td> <td> <p>Unique ID of each Instagram post</p> </td> </tr> <tr> <td> <p>Post Description</p> </td> <td> <p>Complete description of each post in the language in which it was originally published</p> </td> </tr> <tr> <td> <p>Date</p> </td> <td> <p>Date of publication in MM/DD/YYYY format</p> </td> </tr> <tr> <td> <p>Language</p> </td> <td> <p>Language of the post as detected using the Google Translate API</p> </td> </tr> <tr> <td> <p>Translated Post Description</p> </td> <td> <p>Translated version of the post description. All posts which were not in English were translated into English using the Google Translate API. No language translation was performed for English posts.</p> </td> </tr> <tr> <td> <p>Sentiment</p> </td> <td> <p>Results of sentiment analysis (using translated Post Description) where each post was classified into one of the sentiment classes: fear, surprise, joy, sadness, anger, disgust, and neutral</p> </td> </tr> <tr> <td> <p>Hate</p> </td> <td> <p>Results of hate speech detection (using translated Post Description) where each post was classified as hate or not hate</p> </td> </tr> <tr> <td> <p>Anxiety or Stress</p> </td> <td> <p>Results of anxiety or stress detection (using translated Post Description) where each post was classified as stress/anxiety detected or no stress/anxiety detected.</p> </td> </tr> </tbody> </table>
Dataset of Instagram images related to Korean museums
<p>This dataset contains images crawled from Instagram using two hashtags related to two exhibitions held Seoul (South Korea) in 2020 and 2021. One is a generic hashtag for the Museum of Modern and Contemporary Art, MMCA, (#국립현대미술관), the other is for the "instagrammable" exhibition "Yumi's Cell Special Exhibition" (#유미의세포들특별전).</p> <p>The dataset contains more than 20,000 images (Image_all_*.zip).</p> <p>It also includes a selection of images in which there are humans interacting with objects and art installations: 9,409 (Yumi) and 810 (MMCA) (Image_Human_*.zip). The two archives Skeleton_*.zip contain the results of the skeleton analysis of such images, done with OpenPose.</p>
Live no Instagram sobre Psicologia Perinatal
<p>Live realizada no Instagram <a href="https://www.instagram.com/acalanto.perinatal/">@acalanto.perinatal</a>. Uma conversa franca abordando a psicologia nos períodos de gravidez, parto e puerpério com a Psicóloga Miriam Calisman, coordenadora do Espaço Calisman, e com mais de 20 anos de experiência em psicologia clínica.</p> <p>Link da publicação original: <a href="https://www.instagram.com/tv/CFQcamBlo2Y/?igshid=MDJmNzVkMjY%3D">https://www.instagram.com/tv/CFQcamBlo2Y/?igshid=MDJmNzVkMjY%3D</a><br> </p>
Live no Instagram sobre amamentação - Agosto Dourado
<p>Live realizada no Instagram <a href="https://www.instagram.com/acalanto.perinatal/">@acalanto.perinatal</a>. Participação da Prof Abilene Gouveia - Hupe/UERJ, membro do grupo técnico da SES e da Comissão de Aleitamento da SOPERJ. Professora da UVA. Incentivo da família, mitos, bancos de leite, cuidado no regresso ao trabalho, e outros tópicos.</p> <p>Link da publicação original: <a href="https://www.instagram.com/tv/CD4W2IJl0Dt/?igshid=MDJmNzVkMjY%3D">https://www.instagram.com/tv/CD4W2IJl0Dt/?igshid=MDJmNzVkMjY%3D</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.