Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
444
datasets available to search
ShareScore release 0.9.0
Dataset results
444 results for “social network”
Mental Health in Social Networks with Machine Learning Algorihtms
<p>The <strong>DatasetMH.xlsx</strong> excel corresponds to a corpus of mental health in social networks labelled with polarity and stigma. In particular, the corpus consists of 2,287 comments labelled with polarity (positive, negative, neutral) and stigma from comments on Instagram posts about celebrity mental health disclosures:</p> <ol> <li>Polarity: It consists of giving a positive, negative or neutral/undefined value to the comments in response to the disclosure or description of the symptomatology in the post. Positive polarity reflects understanding, encouragement or even admiration of the publication. E.g., “Cheer up, we love you".”. Negative polarity is assigned when the person expresses negative opinions, usually questioning the post with ironic, sarcastic or even mocking and disparaging comments. E.g., “how you show that you don't know what depression or anxiety is, shame on you!”. Neutral or undefined polarity is assigned in cases where no clear opinion is detected or can be interpreted in both directions. E.g., “take medication, it will help you" "and your partner?”</li> <li>Stigma: stigmatising responses to comments are behaviours in which negative beliefs and emotions towards MH problems are expressed. Stigma manifests in a variety of forms including rejection and anger against the person, which may extend to contempt or mockery, belittling their problem. E.g.,"What a desire to draw attention to yourself"; "what you have is a story"; "you're so inconsistent and seeking the limelight". Because socially we know that "stigma is wrong" many rejection comments are made in an ironic or sarcastic way. E.g., and how do you write on insta?"; "better information from someone who doesn't have a current account". Additionally, anger is shown by arguing that such posts "trivialise or commercialise" MH. E.g.,"don't come and tell me your false stories of overcoming, without even knowing what it is to work...". Other times the stigma manifests itself as pity or sorrow for the person. E.g.,“It breaks my heart”; “poor thing”.</li> </ol> <p>The file <strong>DatasetMH_Emotions.xlsx</strong> corresponds to a corpus of mental health in social networks labelled with emotions. In particular, the corpus consists of 2,287 comments labelled with five emotions plus a neutral class from comments on Instagram posts about celebrity mental health disclosures. These emotions are:</p> <ul> <li>Love/admiration: This emotion involves messages where admiration, approval and love are closely related.</li> <li>Gratitude: the messages imply a sincere appreciation for sharing mental health-related content on social networks.</li> <li>Comprehension/empathy/identification: The messages involve interest in and understanding of the message, including self-identification with the situation or context.</li> <li>Sadness: This primary emotion is produced by events that are not pleasant and that denote heaviness. It includes many manifestations of pity for the person.</li> <li>Anger/contempt/mockery: This emotion involves responses of irritation and attacks on the person as ridiculous and superficial.</li> <li>Neutral: This category corresponds to messages without emotions.</li> </ul> <p>The labelling process of both datasets was divided into two phases: an initial phase with a pilot corpus (N = 787 comments) and a second phase focused on the development of the corpus with all the comments of the selected posts (N = 21151). The same methodology was followed in both phases: once the comments were collected, the corpus was cleaned, and then two independent experts were responsible for labelling each category. A third expert then reviewed the comments to resolve discrepancies. In the third and final phase, a final corpus for application to the machine learning algorithms is built from the large corpus (N = 2287).</p> <p>Classification models are a set of machine learning algorithms developed to assess emotional response, i.e. polarity, stigma and emotions in social networks, based on previously developed datasets (<strong>DatasetMH_Emotions.xlsx</strong>, <strong>DatasetMH.xlsx</strong>).</p>
Identity Lexicon and Keyword Counts for "The Life of a Tie: Social Origins of Network Diversity"
<p>Prototype identity lexicon for the categories of occupation (e.g., "reporter" at Boston Globe), familial roles (e.g., proud "father"), political affiliation (life-long "democrat"), and cultural and sports interests (e.g., "hiphop", "NFL").</p> <p>Also includes a CSV file with counts of each identity keywords matched for all 572K users in our dataset. </p> <p> </p>
Medical Concept Normalization in Social Media Posts with Recurrent Neural Networks
<p>Text mining of scientific libraries and social media has already proven itself as a reliable tool for<br> drug repurposing and hypothesis generation. The task of mapping a disease mention to a concept<br> in a controlled vocabulary, typically to the standard thesaurus in the Unified Medical Language<br> System (UMLS), is known as medical concept normalization. This task is challenging due to the<br> differences in medical terminology between health care professionals and social media texts coming<br> from the lay public. To bridge this gap, we use sequence learning with recurrent neural networks<br> and semantic representation of one- or multi-word expressions: we develop end-to-end architectures<br> directly tailored to the task, including bidirectional Long Short-Term Memory and Gated Recurrent<br> Units with an attention mechanism and additional semantic similarity features based on UMLS.<br> Our evaluation over a standard benchmark shows that recurrent neural networks improve results<br> over an effective baseline for classification based on convolutional neural networks. A qualitative<br> examination of mentions discovered in a dataset of user reviews collected from popular online health<br> information platforms as well as quantitative evaluation both show improvements in the semantic<br> representation of health-related expressions in social media.</p>
Interaction-Based Behavioral Analysis in Twitter Social Network
<p>Literature studies usually use data sets consisting of data collected from many different metrics and user counts collected over different time periods. The data set used in this article was formed using completely up-to-date data obtained as a result of metrics measured in terms of scope and efficiency, sufficient and effective user counts, and filtering processes. To classify users correctly and make the classification performance high—in addition to parameters used in the literature such as tweets, account age, follower rank, average retweets and average likes—other parameters such as diameter, density, reciprocity, centralization and modularity were used. These metrics are the parameters that focus on a different area to reveal many aspects in which social network users interact. The data used to create the data set was collected from Twitter. The metric data forming the data set was extracted using Twitter Rest API V1.1 supporting search/tweet endpoints by means of the SocialBlade and Netlytic platforms.</p>
Dataset from: Parental mediation and the use of social networks: A systematic review
<p>Translation of the previous file: aplicación_criterios_elegibilidad_sin_duplicados and modification of an error detected during the review process. </p>
Data for "Information sharing within a social network is key to behavioral flexibility – lessons from mice tested under semi-naturalistic conditions"
<p>Data for "Information sharing within a social network is key to behavioral flexibility – lessons from mice tested under semi-naturalistic conditions", currently under review in Science Advances. </p>
Data from: Linking social and spatial networks to viral community phylogenetics reveals subtype-specific transmission dynamics in African lions
1.Heterogeneity within pathogen species can have important consequences for how pathogens transmit across landscapes; however, discerning different transmission routes is challenging. 2.Here we apply both phylodynamic and phylogenetic community ecology techniques to examine the consequences of pathogen heterogeneity on transmission by assessing subtype specific transmission pathways in a social carnivore. 3.We use comprehensive social and spatial network data to examine transmission pathways for three subtypes of feline immunodeficiency virus (FIVPle) in African lions (Panthera leo) at multiple scales in the Serengeti National Park, Tanzania. We used FIVPle molecular data to examine the role of social organization and lion density in shaping transmission pathways and tested to what extent vertical (i.e., father and/or mother offspring relationships) or horizontal (between unrelated individuals) transmission underpinned these patterns for each subtype. Using the same data, we constructed subtype specific FIVPle co-occurrence networks and assessed what combination of social networks, spatial networks, or co-infection best structured the FIVPle network. 4.While social organization (i.e., pride) was an important component of FIVPle transmission pathways at all scales, we find that FIVPle subtypes exhibited different transmission pathways at within- and between-pride scales. A combination of social and spatial networks, coupled with consideration of subtype co-infection, was likely to be important for FIVPle transmission for the two major subtypes, but the relative contribution of each factor was strongly subtype specific. 5.Our study provides evidence that pathogen heterogeneity is important in understanding pathogen transmission, which could have consequences for how endemic pathogens are managed. Furthermore, we demonstrate that community phylogenetic ecology coupled with phylodynamic techniques can reveal insights into the differential evolutionary pressures acting on virus subtypes, which can manifest into landscape-level effects.
Data from: Split between two worlds: automated sensing reveals links between above- and belowground social networks in a free-living mammal
Many animals socialize in two or more major ecological contexts. In nature, these contexts often involve one situation in which space is more constrained (e.g. shared refuges, sleeping cliffs, nests, dens or burrows) and another situation in which animal movements are relatively free (e.g. in open spaces lacking architectural constraints). Although it is widely recognized that an individual's characteristics may shape its social life, the extent to which architecture constrains social decisions within and between habitats remains poorly understood. Here we developed a novel, automated-monitoring system to study the effects of personality, life-history stage and sex on the social network structure of a facultatively social mammal, the California ground squirrel (Otospermophilus beecheyi) in two distinct contexts: aboveground where space is relatively open and belowground where it is relatively constrained by burrow architecture. Aboveground networks reflected affiliative social interactions whereas belowground networks reflected burrow associations. Network structure in one context (belowground), along with preferential juvenile–adult associations, predicted structure in a second context (aboveground). Network positions of individuals were generally consistent across years (within contexts) and between ecological contexts (within years), suggesting that individual personalities and behavioural syndromes, respectively, contribute to the social network structure of these free-living mammals. Direct ties (strength) tended to be stronger in belowground networks whereas more indirect paths (betweenness centrality) flowed through individuals in aboveground networks. Belowground, females fostered significantly more indirect paths than did males. Our findings have important potential implications for disease and information transmission, offering new insights into the multiple factors contributing to social structures across ecological contexts.
Multilayer social networks reveal the social complexity of a cooperatively breeding bird
<p>Focal observations of Arabian babblers (<em>Argya squamiceps</em>). The folder contains adjacency matrices of nine social groups observed between 2017 and 2020. Some of the groups were observed multiple times, the complete list of group rounds is also provided. We recorded six different interaction types, hence the folder contains a total of 114 adjacency matrices.</p>
Integrierte Antwortstrategien als Reaktion auf negative Electronic Word-of-Mouth. Eine Analyse von Nestlés Beschwerdemanagement auf deutschen und US-amerikanischen Social Network Sites
<p>Die hochgeladenen sieben Dokumente beinhalten das Codebuch der Arbeit, alle Beschwerden samt Codierungen sowie alle Korrelationen zwischen Beschwerden und Antwortstrategien.</p> <p>Des weiteren finden sich alle Beschwerden und Beschwerdeantworten der vier CSR-Themen <em>Tierleid, Privatisierung von Wasser/Wasserklau, Regenwaldzerstörung </em>und <em>Kinderarbeit. </em>Bei jeder Beschwerde sind zudem die passenden qualitativ ermittelten Kategorien grafisch dargestellt.</p>
Observational Data for: Polarized information ecosystems can reorganize social networks via information cascades
<p><strong>General Information</strong></p> <p>This contains the observational data for the publication:</p> <blockquote> <p>Tokita, Guess, and Tarnita (2021). Polarized information ecosystems can reorganize social networks via information cascades. <em>Proceedings of the National Academies of Science of the United States.</em></p> </blockquote> <p>Please see the above peer-reviewed article that resulted from this data for more details.</p> <p>Please see the <em>o</em><em>bservational/ </em>directory in the Github repository <a href="http://%28https//github.com/christokita/information-cascades)">https://github.com/christokita/information-cascades</a> for the code that analyzes the Twitter data and generates derived data. Data directly pertaining to the four news sources of interest are in the <em>data/</em> directory, while data that is related to the 4,000 monitored news followers are in the <em>data_derived/</em> directory.</p> <p>The main data file is in a zipped file. We also included a read me with a description of each directory and subdirectory of the data.</p> <p><strong>Methods</strong></p> <p>For each of our four news sources of interest (CBS News, USA Today, Vox, and the Washington Examiner), we used the Twitter API to sample 3,000 random Twitter users that followed that news source's account. We then used the <em>tweetscores</em> R package to estimate the ideology of each of the 12,000 sampled users.</p> <p>From the original pool of approximately 12,000 sampled users who follow a news outlet of interest, we monitored the follower networks of 1,000 liberal followers of CBS News, 1,000 conservative followers of USA Today, 1,000 liberal followers of Vox, and 1,000 conservative followers of the Washington Examiner. When selecting the 1,000 random users to monitor for a given news outlet, we first conducted filters based on self-reported geolocation and other Twitter profile attributes to reasonably filter down to US-based users who primarily tweet in English. We then pulled the complete follower network of each of our 4,000 monitored users at the beginning and end of a 6-week period from August to September 2020, allowing us to assess who unfollowed these users over this period of time. Finally, using the initial follower networks of each monitored user, we estimated the ideology of up to 50 random followers to create a baseline for the ideological composition of each user’s follower network. </p> <p>Using the initial and final follower networks of each individual, we calculated the rate of unfollows by opposite-ideology users in their follower network. To determine whether an unfollow event was breaking a cross-ideology social tie, users and their unfollowers were classified simply as either liberal (ideology score < 0) or conservative (ideology score > 0). We excluded from analysis unfollowers for whom we could not estimate ideology, either because the account had been suspended, deleted, or made private, or because they did not follow any of the political, news, or cultural accounts needed to do the estimation. We then compared the proportion of unfollowers that were of the opposite ideology (i.e., conservative unfollowers if the focal user is liberal) against the estimated proportion of opposite-ideology followers in the focal user’s initial follower network. This allowed us to account for the fact that many user’s follower networks were not ideologically balanced and set the baseline expectation that if unfollows were random then the proportion of unfollows by opposite-ideology users should match the proportion of followers that were of the opposite ideology.</p> <p> </p> <p><strong>Abstract (for main paper)</strong></p> <p>The precise mechanisms by which the information ecosystem polarizes society remain elusive. Focusing on political sorting in networks, we develop a computational model that examines how social network structure changes when individuals participate in information cascades, evaluate their behavior, and potentially rewire their connections to others as a result. Individuals follow proattitudinal information sources but are more likely to first hear and react to news shared by their social ties and only later evaluate these reactions by direct reference to the coverage of their preferred source. Reactions to news spread through the network via a complex contagion. Following a cascade, individuals who determine that their participation was driven by a subjectively “unimportant” story adjust their social ties to avoid being misled in the future. In our model, this dynamic leads social networks to politically sort when news outlets differentially report on the same topic, even when individuals do not know others’ political identities. Observational follow network data collected on Twitter support this prediction:We find that individuals in more polarized information ecosystems lose cross-ideology social ties at a rate that is higher than predicted by chance. Importantly, our model reveals that these emergent polarized networks are less efficient at diffusing information: Individuals avoid what they believe to be “unimportant” news at the expense of missing out on subjectively “important” news far more frequently. This suggests that “echo chambers”—to the extent that they exist—may not echo so much as silence.</p>
The NGI Forward semantic social network data
<p>The <a href="https://research.ngi.eu/">NGI Forward project</a> is part of the European Union's <a href="https://www.ngi.eu">Next Generation Internet Initiative</a>. It is meant to provide European institutions with policy advice for how to shape the future, human-centric Internet. As part of it, a team of ethnographers coded a specially convened online conversation, then arranged its results into a semantic social network. This dataset encodes that conversation, as well as the results of the coding exercise, in raw data form for further exploration and replication purposes. The dataset is pseudonymized.</p> <ul> <li><a href="https://exchange.ngi.eu/">Funnel website</a> of the project.</li> <li><a href="https://journals.sagepub.com/doi/10.1177/1525822X20908236">About semantic social networks</a>.</li> <li><a href="https://edgeryders.eu/t/long-term-ssna-data-storage-documentation-manual/12786">Data export and documentation process</a> (contains links to the code used to export the data)</li> </ul>
The social network of target of rapamycin complex 1 in plants
<p>The target of rapamycin complex 1 (TORC1) is a highly conserved serine–threonine protein kinase crucial for coordinating growth according to nutrient availability in eukaryotes. It works as a central integrator of multiple nutrient inputs such as sugar, nitrogen, and phosphate and promotes growth and biomass accumulation in response to nutrient sufficiency. Studies, especially in the past decade, have identified the central role of TORC1 in regulating growth through interaction with hormones, photoreceptors, and stress-signaling machinery in plants. In this review, we comprehensively analyse the interactome and phosphoproteome of the Arabidopsis TORC1 signaling network. Our analysis highlights the role of TORC1 as a central hub kinase communicating with the transcriptional and translational apparatus, ribosomes, chaperones, protein kinases, metabolic enzymes, and autophagy and stress response machinery to orchestrate growth in response to nutrient signals. This analysis also suggests that along with the conserved downstream components shared with other eukaryotic lineages, plant TORC1 signaling underwent several evolutionary innovations and co-opted many lineage-specific components. Based on the protein–protein interaction and phosphoproteome data, we also discuss several uncharacterized and unexplored components of the TORC1 signaling network, highlighting potential links for future studies.</p>
Exploring Korean adolescent stress on social media: A semantic network analysis
<p><strong>Korean Adolescent's Stress Semantic Network Analysis Project</strong></p> <p>Semantic Network Analysis for Korean Adolescent's Stress</p> <p>Input data file</p> <ul> <li>data_news.csv : News data collected from Naver(<a href="https://www.naver.com">https://www.naver.com</a>)</li> <li>data_blog.csv : Blog data collected from Naver(<a href="https://www.naver.com">https://www.naver.com</a>) and Daum(<a href="https://www.daum.net">https://www.daum.net</a>)</li> </ul> <p>Output files</p> <ul> <li>Frequency Table of Each word in Documents (<em><strong>freq_news.csv</strong></em>, <em><strong>freq_blog.csv</strong></em>)</li> <li>TF-IDF(Term Frequency-Inverse Document Frequency) Table of Each word in Documents (<em><strong>tfidf_news.csv</strong></em>, <em><strong>tf_idf_blog.csv</strong></em>)</li> <li>Frequency Table of 30 keywords in Documents (<em><strong>freq_news_30.csv</strong></em>, <em><strong>freq_blog_30.csv</strong></em>)</li> <li>DTM(Document Term Matrix) of 30 keywords in Documents (<em><strong>DTM_news_30.csv</strong></em>, <em><strong>DTM_blog_30.csv</strong></em>)</li> <li>COM(Co-Occurrence Matrix) of 30 keywords in Documents (<em><strong>COM_news_30.csv</strong></em>, <em><strong>COM_blog_30.csv</strong></em>)</li> <li>Binary COM of 30 keywords in Documents (<em><strong>BinaryCOM_news_30.csv</strong></em>, <em><strong>BinaryCOM_blog_30.csv</strong></em>)</li> <li>Centrality Table of 30 keywords in Documents (<em><strong>centrality_news_30.csv</strong></em>, <em><strong>centrality_blog_30.csv</strong></em>)</li> </ul>
Data and code for: Generation and applications of simulated datasets to integrate social network and demographic analyses
<p class="MsoNormal"><span>Social networks are tied to population dynamics; interactions are driven by population density and demographic structure, while social relationships can be key determinants of survival and reproductive success. However, difficulties integrating models used in demography and network analysis have limited research at this interface. We introduce the R package genNetDem for simulating integrated network-demographic datasets. It can be used to create longitudinal social networks and/or capture-recapture datasets with known properties. It incorporates the ability to generate populations and their social networks, generate grouping events using these networks, simulate social network effects on individual survival, and flexibly sample these longitudinal datasets of social associations. By generating co-capture data with known statistical relationships it provides functionality for methodological research. We demonstrate its use with case studies testing how imputation and sampling design influence the success of adding network traits to conventional Cormack-Jolly-Seber (CJS) models. We show that incorporating social network effects in CJS models generates qualitatively accurate results, but with downward-biased parameter estimates when network position influences survival. Biases are greater when fewer interactions are sampled or fewer individuals are observed in each interaction. While our results indicate the potential of incorporating social effects within demographic models, they show that imputing missing network measures alone is insufficient to accurately estimate social effects on survival, pointing to the importance of incorporating network imputation approaches. genNetDem provides a flexible tool to aid these methodological advancements and help researchers test other sampling considerations in social network studies.</span></p>
Data for "Structure and drivers of social networks and their links with health in older adults"
<p>Data, supplementary material and scripts for the paper "Structure and drivers of social networks and their links with health in older adults"</p> <p>Social network is an important factor in promoting healthy aging. However, the mechanisms linking social capital to health are complex. Moreover, most of the social network analysis studies on older adults consider only participants’ relationships and not how these relationships are themselves connected. In this study, we went further than current ego-centered network studies by determining global social network metrics and the structure of relationships among older adult participants of the RECORD Cohort using the Veritas-Social questionnaire. The aim of this study is to identify key dimensions of social networks of older adults, and to evaluate how these dimensions relate to depressive symptoms, life satisfaction, and well-being. Using Principal Component Analyses (PCA), we identified four social network dimensions with psychological meanings. Dimension 1 (homophily) was positively linked with perceived accessibility to services in one’s residential neighborhood but negatively linked with the level of study. Dimension 2 (social integration) as Dimension 3 (social support) was only linked to the number of people living with ego. Dimension 4 was linked with perceived accessibility to local services. Finally, and rather surprisingly, we found that none of the four network dimensions, even the degree, was linked to the three health status metrics.</p>
Social network and fitness data from age-structured populations of forked fungus beetles
<p>We investigated the relationships between age, social behavior, and fitness at three levels of organization: the individual, the local social environment, and the population. Replicate groups of forked fungus beetles (<em>Bolitotherus cornutus</em>) were engineered to have either young- or old- biased age structures, and both social and reproductive behaviors were recorded.</p>
Data for Vasari Social Network
<p>Files to run scripts in <a href="https://github.com/ISE-FIZKarlsruhe/vasari_network">https://github.com/ISE-FIZKarlsruhe/vasari_network</a> for generating a social network weighted with Pointwise Mutual Information (PMI) from <em>The Lives of The Artists</em> (1568) by Giorgio Vasari.</p> <p>The 10 .csv files in <em>volumes/</em> contain the pages extracted from <em>The Lives of The Artists </em>edition on Project Gutenberg (<a href="https://onlinebooks.library.upenn.edu/webbin/metabook?id=livespainters">here</a>)</p>
Results for Vasari Social Network
<p>Files to run scripts in <a href="https://github.com/ISE-FIZKarlsruhe/vasari_network">https://github.com/ISE-FIZKarlsruhe/vasari_network</a> for generating a social network weighted with Pointwise Mutual Information (PMI) from <em>The Lives of The Artists</em> (1568) by Giorgio Vasari.</p> <p>The input data for running the source code is available <a href="https://doi.org/10.5281/zenodo.8395369">here</a>.</p> <p>The <em>centralities/</em> directory contains 10 files reporting the highest centraility scores in the extracted network. The network is extracted dynamically from a 10 volumes edition of the book. This means that a network is created for each volume. However, every network is the evolution of his precedent since it includes the information extracted from the previous volumes of the book. The numerical division of the results (from 0 to 9) reflects this dynamic extraction.</p>
Extracted Information for the Systematic Review of Survey Scales for Measuring Information Privacy Concerns on Social Network Sites
<p>The data set is part of a systematic literature review of survey scales for measuring information privacy concerns (IPCs) used in research on social network sites (SNSs).</p> <p>The results of this systematic literature review are available in Bartol, J., Vehovar, V., & Petrovčič, A. (2023). Systematic review of survey scales measuring information privacy concerns on social network sites. <em>Telematics and Informatics</em>, 102063. https://doi.org/10.1016/j.tele.2023.102063</p> <p>The article also includes a detailed description of the methods used in generating this data set.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.