Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
21
datasets available to search
ShareScore release 0.9.0
Dataset results
21 results for “politicians”
Query auto-completions for German politicians of the 18th Bundestag
<p><strong>bundestag.csv</strong> - UTF-8 encoded comma separated text file</p> <p>This dataset contains the members of the 18th German Bundestag in the constitution of late 2016.</p> <p><Name>: name of the politician</p> <p><Born>: birthday</p> <p><Party>: party membership of the politician</p> <p><Bundesland> state of the politician</p> <p><Gender> gender of the politician</p> <p><Age> age of the politician (as of 2017)</p> <p><Cluster 3> number of unique auto-completions assigned to topic: "location information"</p> <p><Cluster 2> number of unique auto-completions assigned to topic: "personal and emotional"</p> <p><Cluster 1> number of unique auto-completions assigned to topic: "politics and economics"</p> <p><Total> total number of unique auto-completions</p> <p> </p> <p><strong>terms.csv </strong>- UTF-8 encoded comma separated text file</p> <p>This dataset contains the unordered and pooled auto-completions for the German politicians from Bing search (http://api.bing.net/osjson.aspx), from Duck-Duck-Go (https://duckduckgo.com/ac/) and from Google search (http://clients1.google.de/complete/search). The data was crawled on (mostly) two times per day from 2017/02/03 to 2017/06/19. German language settings were used for Google and Bing, English language setting was used for Duck-Duck-Go. The API requests were sent with an IP address from Cologne, Germany. </p> <p><source>: google, bing or ddg</p> <p><queryterm>: the query term, matches the name of the politican in the file <bundestag.csv></p> <p><suggestterm>: the suggested query auto-completion</p>
btw17 query auto completion - query suggestions for German politicians and parties before the federal election 2017
<p>The dataset contains the query suggestions for 5 major German parties (terms: "afd", "csu", "dielinke", "fdp", "grüne", "spd") and ten popular politicians and party leaders (terms: "Alexander Gauland", "Alice Weidel", "Angela Merkel", "Cem Özdemir", "Christian Lindner", "Dietmar Bartsch", "Katrin Göring-Eckardt", "Martin Schulz", "Sahra Wagenknecht").</p> <p>The data was crawled on (mostly) two times per day from Tue Aug 04, 2017 to Tue Oct 31, 2017. The dataset contains 20001 suggestions from Bing search (http://api.bing.net/osjson.aspx), 11935 suggestions from Duck-Duck-Go (https://duckduckgo.com/ac/) and 33521 suggestions from Google search (http://clients1.google.de/complete/search). Note, that for some terms and dates no suggestions were returned by some of the APIs.</p> <p>German language settings were used for Google and Bing, English language setting was used for Duck-Duck-Go. The API requests were sent with an IP address from Cologne, Germany. </p> <p>The UTF-8 encoded comma separated text file contains the following columns:</p> <p><source>: google, bing or ddg</p> <p><queryterm>: the query term</p> <p><date>: the date and time of the API call formatted as ISO8601</p> <p><suggestterm>: the suggested query completion (the query term was removed from the suggestion)</p> <p><position>: the position of the query suggestion within the list returned by the API (ranges from 0 to 19)</p> <p> </p> <p> </p> <p><br> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p>
Canadian Politicians on YouTube
<div> <h4>Description of Columns</h4> </div> <p><code>Honorific title</code>: MPs who are members of the Canadian Privy Council and use the title “The Honourable” (Hon.)</p> <p><code>First name</code>: MPs’ first name</p> <p><code>Last name</code>: MPs’ last name</p> <p><code>Username</code>: MPs’ YouTube username</p> <p><code>Profile URL</code>: URL for MPs’ YouTube account</p> <p><code>Status</code>: Whether an MP’s YouTube account is <em>Active</em> or <em>Inactive</em></p> <ul> <li> <p><code>Active</code>: At least one short or full length video was posted in 2025</p> </li> <li> <p><code>Inactive</code>: The last short or full length video was posted on December 31, 2024 or earlier</p> </li> </ul> <p><code>Gender</code>: MPs’ gender, as categorized by the <a href="https://www.ourcommons.ca/Members/en/search" rel="nofollow">House of Commons</a></p> <p><code>Political Affiliation</code>: MPs’ political party affiliation, as categorized by the <a href="https://www.ourcommons.ca/Members/en/search" rel="nofollow">House of Commons</a></p> <p><code>Constituency</code>: Name of the MPs’ constituency (as of the 45th Parliament, following redistribution)</p> <p><code>Province/Territory</code>: Where the MPs’ constituency is located</p>
Brazilian Politician dataset for Record Linkage
<p>The Brazilian political dataset from the Tribunal Superior Eleitoral (TSE) is a comprehensive and valuable resource. It contains information about Brazilian politicians, including their names, political parties, electoral districts, and personal data such as race, gender, address, place, and date of birth. The TSE has published a version of this dataset every two years from 1992 until now, updating the information of the politicians.</p> <p>The structure of the TSE dataset can be utilized to create rich and meaningful datasets for record linkage projects. This means the data can link information from different versions (presented over the year) and create a more comprehensive understanding of politician data evolution. Observing the data's evolution makes it possible to identify patterns and connections that might not be apparent otherwise. This can also aid in identifying irregularities or illegal activities related to political campaigns. Overall, the TSE dataset's structure lends itself well to record linkage and populational projects, making it a valuable resource for researchers and analysts.</p> <p>We use the politician TSE dataset to build a dataset that can be used in the record linkage and privacy-preserving record linkage context. Moreover, we leverage the modification/updates in the politician's personal information (reflected in the dataset) over time to build our dataset. For example, a politician can marry and change his/her name, or the record could be inserted with a typo in the political affiliations.</p> <p>We created a dataset for the record linkage application by utilizing the variations in the original TSE dataset, which resulted in real-world linkage errors. This dataset can be useful for training classifiers and measuring the accuracy of both record linkage and privacy-preserving record linkage.</p> <p> </p>
Politicians dataset
<p>This dataset covers 77,954 politicians around the world.</p> <p>This dataset can also be found here for updates and more configuration parameters: <a href="https://www.workwithdata.com/dataset?entity=politicians">https://www.workwithdata.com/dataset?entity=politicians</a></p>
Wikidata Dump Indian_Politician_Properties
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> <br> <a href="https://tools.wmflabs.org/wdumps/dump/94">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
Wikidata Dump Indian_Politician_Properties
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> <br> <a href="https://tools.wmflabs.org/wdumps/dump/94">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
Wikidata Dump politician
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> <br> <a href="https://tools.wmflabs.org/wdumps/dump/546">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
Wikidata Dump French politicians
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> <br> <a href="https://tools.wmflabs.org/wdumps/dump/2061">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
Wikidata Dump politicians
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> politicians<br> <a href="https://tools.wmflabs.org/wdumps/dump/2065">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
Wikidata Dump politicians
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> politicians<br> <a href="https://tools.wmflabs.org/wdumps/dump/2065">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
Wikidata Dump Politician / Entrepreneur / Businessperson (en)
<p>RDF dump of wikidata produced with <a href="//wdumps.toolforge.org/">wdumper</a>.</p><p><br><a href="//wdumps.toolforge.org/dump/2607">View on wdumper</a></p><p><b>entity count<b>: 579092, <b>statement count</b>: 11735776, <b>triple count</b>: 12839535</b></b></p>
One million articles from five post socialist countries with extracted features: sentiment, basic emotions, LDA topics and presence of influential domestic politicians
<p>This is a replication data for my paper under blind review.<br> <br> This paper develops a new prediction model for media content presence on a website. It analyses a new corpus of one million articles from five countries: Poland, Russia, Belarus, Kazakhstan and Ukraine, in two languages, Polish and Russian. These articles were scraped daily from seventeen websites in 2017-2020 period. The research applies a wide range of natural language processing methods to automatically derive several properties of each article: its topic, sentiment, basic emotions, mentions of influential domestic politicians. The articles’ embeddings and their cosine similarity are used to calculate the news context, such as how an article differs from the daily issue main themes. These features are used to estimate a logistic regression assessing the likelihood that the same or slightly modified, as measured by cosine similarity, article will remain on the main web page the next day. The key, and somewhat unexpected result is that articles with negative sentiment polarity are less likely to be published for more than one day. This result holds for all countries analyzed. It means that the negative news bias documented in the literature is partly offset by their shorter life cycle.<br> <br> Data is in the Python pickle format. Should be read into Python using the pickle.load() function. Each element (row) is the data frames or list represents one news article. Each file has the same format. Loading a pickle file returns a list of four elements:<br> 1. A dummy variable equal to 1 when the article was published the next day, with the text being identical<br> 2. A dummy variable equal to 1 when the article was published the next day, but we allow for small text modifications (cosine similarity > 0.99)<br> 3. Dataframe with extracted features, described below.<br> 4. List with texts of articles in Polish or Russian<br> <br> Ad 3. The columns of the dataframe are as follows (we refer to row number i in description):<br> - pandas index (may appear once or twice in the datafame)<br> - maxcosine: maximum cosine similarity between art i and all articles published next day<br> - cosine_diff: cosine similarity between article i and the elementwise average of embeddings of all articles in the current issue. Measure how similar is the article i to the core narrative of the current issue<br> - cosine_std: std. dev. of cosine similarity measures between all pairs of articles in the current issue. Measures how focused or dispersed is the current issue news coverage<br> - thirteen LDA topic groups: politics, legislation and legal affairs (POL); economy, finance, various sectors of the economy (ECO); military, war, protests, crime, security threats (MIL); international affairs, specific issues concerning foreign countries (INT); technology (TECH); family issues, culture, sport, education (FAM); regional issues and housing (REG); health issues and the Covid-19 pandemic (HEA); media (MED); accidents (ACC); religion (REL); the Soviet Union (USSR); and articles for which no topic could be determined (MISC).<br> - rsent.c: relative sentiment that is dictionary based sentiment of articles i minus the average sentiment of the newspaper. This approach eliminates newspaper or country idiosyncratic sentiment factors. c stands for Covid, the sentiment lexicon was augmented with Covid related terms<br> - dip_*: Variable measuring if influential domestic politicians are mentioned in article i, * represent a country acronym. If N is equal to the number of occurrences of the names of influential domestic politicians in the article i, dip_* = 0 if N=0, dip_* = 1+ log(N) if N>0.<br> - three or four names of news portals from which the data was scraped.<br> - names of six basic emotions and the article i emotion scores calculated using zero-shot learning and the large version of the XLM (Conneau et al., 2019) model from the huggingface transformers library available at https://huggingface.co/vicgalle/xlm-roberta-large-xnli-anli<br> Names of the politicians used to calculate dip variables<br> Russia<br> "putin" "medvedev" "vaino" "shoigu" "bortnikov" "lavrov" "mishustin" "kirienko" "sechin"<br> Ukraine<br> "zelensky" "shmygal" "akhmetov" "avakov" "ermak" "poroshenko" "medvedchuk" "groisman"<br> Kazakhstan<br> "sagyntaev" "mamin" "tokayev" "nnazarbayev" "dnazarbayeva" "kulibayev" "masimov"<br> Belarus<br> "alukashenko" "vakulchik" "vlukashenko" "kobyakov" "makei" "myasnikovich" "rumas" "golovchenko"<br> Poland<br> "kaczynski" "duda" "morawiecki" "ziobro"<br> Data coverage<br> Country, news portal, numbr of articles<br> Russia iz.ru 43,782<br> Russia kommersant.ru 46,070<br> Russia novayagazeta.ru 29,357<br> Russia vedomosti.ru 27,797<br> Kazakhstan informburo.kz 29,375<br> Kazakhstan nur.kz 67,350<br> Kazakhstan tengrinews.kz 44,285<br> Kazakhstan zakon.kz 109,442<br> Belarus bdg.by 33,447<br> Belarus belgazeta.by 21,995<br> Belarus sb.by 83,685<br> Ukraine kp.ua 194,792<br> Ukraine segodnya.ua 45,835<br> Ukraine vesti.ua 90,559<br> Poland gazeta.pl 53,321<br> Poland rp.pl 49,587<br> Poland wpolityce.pl 76,625<br> <br> In the provided dataframes the number of observations is smaller, because the issues for which there was no next day issue, were removed.<br> <br> Data was scraped daily between 2017 or 2018 (depending on the country) and January 2021.</p>
Transkripte von drei TV-Debatten mit Schweizer PolitikerInnen vor eidgenössischen Abstimmungen. Transcripts of three TV debates with Swiss politicians before popular votes.
<p>Der Datensatz enthält Transkripte (doc, htm, PDF) von drei 'Arena'-Abstimmungssendungen mit Schweizer PolitikerInnen vor eidgenössischen Abstimmungen. Die TV-Sendungen wurden nach GAT 2 transkribiert / The dataset contains transcripts (doc, htm, PDF) of three 'Arena' TV-debates with Swiss politicians before popular votes. The debates were transcribed according to GAT 2.</p>
Spanish politicians salaries
<p>Dataset that contains the salary of the Spanish politicians as of today (November 8th 2021). </p>
Like or dislike? The impact of politicians' gender on vertical affective polarization in Flanders (Belgium) (replication data)
<p>Replication data for Devroe, R. & Wauters, B. (2023). Like or dislike? The impact of politicians’ gender on vertical affective polarization in Flanders (Belgium) </p>
Wikidata Dump Indian_Politician_Properties
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> <br> <a href="https://tools.wmflabs.org/wdumps/dump/94">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
Wikidata Dump Indian_Politician_Properties
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> <br> <a href="https://tools.wmflabs.org/wdumps/dump/94">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
Wikidata Dump Indian_Politician_Properties
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> <br> <a href="https://tools.wmflabs.org/wdumps/dump/94">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
Wikidata Dump Indian_Politicians
<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> <br> <a href="https://tools.wmflabs.org/wdumps/dump/93">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.