Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13
datasets available to search
ShareScore release 0.9.0
Dataset results
13 results for “European Parliament”
Gender codification of 412 national (general) and European Parliament elections in six European countries (2003-2021)
<p>This dataset has been produced by applying the Manifesto Gender Analysis (MGA) codebook to 412 national (general) and European Parliament elections in the six countries participating in the UNTWIST project (Denmark, Germany, Hungary, Spain, Switzerland, and the UK) from 2003 to 2021.</p> <p> The Manifesto Gender Analysis coding procedure, developed by WP4 of the UNTWIST consortium, aims to analyse gender-related content in party manifestos. It relies on existing manifestos collected by MARPOR and EM projects from 2003-2021 in six national contexts: Denmark, Germany, Hungary, Spain, Switzerland, and the United Kingdom. The process involves splitting manifestos into quasi-sentences, coding them based on a scheme inspired by previous projects and feminist typology, and completing an expert survey. This method ensures comprehensive analysis and potential scalability through computational methods. </p> <p>The coding procedure involves a series of essential steps, divided in two main activities: the classification of manifestos’ quasi-sentences, and the completion of a survey dedicated to more general concepts which can be gauged by evaluating the content of the entire documents. In the latter case, then, the unit of measure of each coder consists in the manifesto document, whereas in the former the units of measure are quasi-sentences - i.e., arguments denoting a verbal expression of a political idea or issue. Coders are instructed to split sentences containing multiple arguments into quasi-sentences and ensure that each quasi-sentence encapsulates a single political idea or issue. </p> <p>Once the manifestos are split into said units, coders classify the arguments following the MGA coding scheme. The coding scheme (MGA) consists of 5 domains and 25 coding categories, covering various aspects of gender-related issues. Each domain includes an "other" category for relevant statements that do not fit precisely into the defined categories. Apart from coding categories related to specific themes, the coding scheme then includes additional dimensions. The classification process consists of seven steps: (1) assessing whether the quasi-sentence addresses gender-related issues, (2) defining both the domain and coding category, (3) determining whether the quasi-sentence refers to a specific recipient or group based on gender and/or sexual orientation, (4) evaluating intersectionality, (5)<strong> </strong>assigning the sentiment or connotation, (6) determining if it's related to a goal, issue, or policy, and (7) characterising the policy if applicable.</p> <p>After completing the classification of the quasi-sentences in a given manifesto, coders fill in a survey for each manifesto document. The surveys provide information that cannot be directly inferred from the quasi-sentences, focusing on the gender ontology of a manifesto, the degree to which a manifesto entails a binary conception of sexes, the extent to which a manifesto promotes a patriarchal conception of the society, and how much a manifesto promotes heterosexuality as the only normal and socially acceptable sexual orientation of individuals. While the last four characteristics are gauged relying on quasi-interval measures (scales ranging from 0 to 10), the first one, gender ontology, consists in a categorical variable which distinguishes between manifestos with an essentialist ontology – gender and sex are the same and inseparable –, a constructivist ontology – biological sex is mediated through social construction of femininity and masculinity –, and other or undefined ontologies.</p>
Convex inference for community discovery in signed networks (European Parliament Voting Dataset)
<p>This repository contains the necessary tools to reproduce the experiments of the paper</p> <ul> <li>G. Santatmaría, V. Gómez (2015)<br> Convex inference for community discovery in signed networks.<br> NIPS 2015 Workshop: Networks in the Social and Information Sciences</li> </ul> <p>The method first maps the MAP problem on the Potts model as a hinge-loss minimization problem (see the paper for details). To run the code you need to install psl (included here) and if you want to additionally compare with other inference methods, such as max prod belief propagation or junction tree, you need to install the libDAI library (also included here)</p> <p>The directory europeanCongressData/ (~500 Mb) contains the votings of the EU parlament, including 300 votings events from the actual term, from May 2014 to June 2015, obtained from http://www.votewatch.eu/</p> <ul> <li>data/ : json files with the european votes</li> <li>network.net : signed network built from the votes</li> <li>political_parties.txt : "ground truth" party</li> <li>community_results/ : results for different number of communities and initial vertices</li> <li>dataComputations.py : used to build the signed network</li> <li>dataProcessing.py : used to build the signed network</li> </ul> <p>We would appreciate if you cite the paper after using the data or the code.</p> <p>DEPENDENCIES</p> <p>The code has been tested in Linux Mint 18.1 Serena and Ubuntu 14.04</p> <p>- For PSL library, you need to have<br> java 1.8<br> you may need to export JAVAHOME='/usr/lib/jvm/YOURJAVA1.8FOLDER'<br> maven 3.x</p> <p>- For libDAI you will need:<br> make doxygen graphviz libboost-dev libboost-graph-dev libboost-program-options-dev libboost-test-dev libgmp-dev cimg-dev libgmp-dev</p> <p>CODE TO RUN THE FOLLOWING EXPERIMENTS:</p> <p>Compare the performance in terms of structural balance of max prod bp and our method against an exact inference method (junction tree), with different number of communities</p> <p>INSTALL</p> <p>To install the experiments you have to follow the next steps:</p> <p>1 Build the libdai library by doing: make -B on the folder (libdai)</p> <p>2 Generate the class path of the groovy project:<br> mvn clean install<br> mvn dependency:build-classpath-Dmdep.outputFile=classpath.out</p> <p>on the psl root folder (You need to have java 1.8 and maven 3.x installed)</p> <p>3 Grant exec permissions to the run.sh script</p> <p>Options</p> <p>The main python file to run the experiments is</p> <p>evaluatebalanceon_sn.py.</p> <p>It accepts the following parameters:</p> <p>1 (Int) Nodes of the graph. In order to run the junction tree we recommend to set this paremeter to 150 or less<br> 2 (Int) The number of underlying communities<br> 3 (Float) The maximum amount of unbalance for the experiments. We recommend 0.45<br> 4 (Bool) Whether to use an heuristic to find the initial node for each community or to use directly random nodes from the ground truth communities. This heuristic looks alternatively for the nodes with highest negative degree and highest positive degree. For the case when the number of communities is equal to 2 (Ising Model), the heuristic is used by default.</p> <p>An example of execution would be:</p> <p>python evaluate_balance_on_sn.py 120 3 0.45 True True</p> <p>The results of the experiments are save in the folder results/<br> Scripts</p> <p>The main script of the hinge-loss method can be found in the folder psl/psl-example/src/main/java/edu/umd/cs/example/PottsCommunities.groovy</p> <p>Authors:</p> <p>Guillermo Santamaria & Vicenc Gomez<br> Mar 5, 2017</p> <p>For further questions, please contact vicen.gomez@upf.edu</p>
Questions in European Parliament on Brain Drain
<p>This is data collated for a discourse analysis of the treatment of questions on "brain drain" in the European Parliament. Collated by Jacob A. Hasselbalch for the ENLIGHTEN project (H2020 #649456).</p>
Directory of the European Parliament members
<p>Over the past twenty-five years, a field of research into the careers of Members of European Parliament (MEPs) has developed. Drawing on a massive amount of accessible open data, we have assembled an updated database comprising all MEPs between 1979 and 2025.</p> <p>This dataset contains (some) socio-demographic informations about MEP’s and their carreer paths in EP. Data were scraped from the EP website.</p> <p>Metadata are filled separately in a rich text format (.rtf) document.</p>
Multiple Partitioning of Multiplex Signed Networks: Application to European Parliament Votes
<p><strong>Presentation. </strong>For more than a decade, graphs have been used to model the voting behavior taking place in parliaments. However, the methods described in the literature suffer from several limitations. The two main ones are that 1) they rely on some temporal integration of the raw data, which causes some information loss; and/or 2) they identify groups of antagonistic voters, but not the context associated with their occurrence. In this article, we propose a novel method taking advantage of multiplex signed graphs to solve both these issues. It consists in first partitioning separately each layer, before grouping these partitions by similarity. We show the interest of our approach by applying it to a European Parliament dataset. Particularly, we study the voting behavior of French and Italian MEPs on "Agriculture and Rural Development" (AGRI) during the 2012-13 legislative year.</p> <p>These are the data used in the following paper:</p> <ul> <li>N. Arınık, R. Figueiredo, and V. Labatut, “Multiple partitioning of multiplex signed networks: Application to European Parliament votes,” <em>Social Networks</em>, vol. 60, pp. 83–102, 2020. DOI: <a href="http://doi.org/10.1016/j.socnet.2019.02.001">10.1016/j.socnet.2019.02.001</a> ⟨<a href="https://hal.archives-ouvertes.fr/hal-02082574">hal-02082574</a>⟩</li> </ul> <p><strong>Source code.</strong> The code source is accessible on GitHub: <a href="https://github.com/CompNet/MultiNetVotes">https://github.com/CompNet/MultiNetVotes</a></p> <p><strong>Citation. </strong>If you use these data our this source code, please cite the above paper.</p> <p><br><code>@Article{Arinik2020,</code><br><code> author = {Arınık, Nejat and Figueiredo, Rosa and Labatut, Vincent},</code><br><code> title = {Multiple Partitioning of Multiplex Signed Networks: Application to {E}uropean {P}arliament Votes},</code><br><code> journal = {Social Networks},</code><br><code> year = {2020},</code><br><code> volume = {60},</code><br><code> pages = {83-102},</code><br><code> doi = {10.1016/j.socnet.2019.02.001},</code><br><code>}</code><br><br>----------------------------------------------<br><strong>Details.</strong><br><br><strong># RAW INPUT FILES</strong><br>The 'itsyourparliament' folder contains all raw input files for further data processing. This is the same raw data that can be found in our previous Figshare repository: https://doi.org/10.6084/m9.figshare.5785833<br>The folder structure is as follows:<br>* itsyourparliament/<br>** domains: There are 28 domain files. Each file corresponds to a domain (such as Agriculture, Economy, etc.) and contains corresponding vote identifiers and their "itsyourparliament.eu" links.<br>** meps: There are 870 Members of Parliament (MEP) files. Each file contains the MEP information (such as name, country, address, etc.)<br>** votes: There are 7513 vote files. Each file contains the votes expressed by MEPs<br><br><strong># ROLLCALL NETWORKS</strong><br>This folder contains two separate zip files regarding rollcall networks:<br>- rollcall-networks: This folder contains only the rollcall networks that are used in the article.<br>- all-rollcall-networks: For those who are interested in other countries or domains, we make available all rollcall networks that we can extract from raw data.<br>Note that these rollcall networks constitute the layers of the input signed multplex network, as illustrated in Figure 1 of the article. Note also that we consider three vote types in our network extraction process: FOR, AGAINST and ABSTAIN.<br><br><strong># ROLLCALL PARTITIONS</strong><br>Note that MEPs who voted similarly are connected together by positive links, and are connected by negative links to MEPs that voted differently from them. MEPs who did not vote at all (ABSENT) are isolates (nodes without any<br>neighbor). We identify the factions of similarly voting MEPs in the graph by solving the Correlation Clustering problem (CC).<br>The rollcall partitions correspond to voting patterns, as illustrated in Figure 1 of the article.<br><br><strong># ROLLCALL CLUSTERING</strong><br>This folder contains the results of Steps 3 and 4 of our workflow (see Figure 1 in the article). The structure of this folder is as follows:<br>|__ votetypes=FAA/: 'FAA' means we consider three vote types in our analysis: FOR, AGAINST and ABSTAIN.<br>|__ F.purity-k=2-sil=SILHOUETTE_SCORE<br>|__ clu=CLUSTER_NO/<br>|__ network: It corresponds to the network created through the similarity network-based approach, as explained in Section 4.4 of the article.<br>|__ partition: It corresponds to the characteristic voting pattern, as explained in Section 4.4 of the article.<br>----------------------------------------------</p> <p>Funding: this research benefited from the support of the Agorantic FR 3621, as well as the FMJH Program PGMO and from the support to this program from EDF-THALES-ORANGE-CRITEO.</p>
MERICS China Podcast: China and the European Parliament election, with Ivana Karásková and Grzegorz Stec
<p>Ahead of the European Parliament election on June 6-9, 2024, this episode looks at the role of the European Parliament in EU-China relations and the possible impact of the election results on the European “de-risking” agenda among other topics. </p> <p><strong>Johannes Heller-John</strong> talks to <strong>Ivana Karásková</strong> and <strong>Grzegorz Stec</strong>. Ivana is a European China Policy Fellow at MERICS and the founder of MapInfluenCE and China Observers in Central and Eastern Europe (CHOICE) at the Association for International Affairs (AMO) in Prague. Grzegorz is the Head of the MERICS Brussels Office.</p> <p>Recently, Ivana co-authored two reports, one on <a href="https://www.amo.cz/en/foreign-electoral-interference-affecting-eu-democratic-processes/" target="_blank" rel="noopener noreferrer">foreign electoral interference in the EU</a> and one on the <a href="https://www.amo.cz/en/from-the-fringes-to-the-forefront-how-extreme-parties-in-the-european-parliament-can-shape-eu-china-relations/" target="_blank" rel="noopener noreferrer">rise of fringe parties in the EP and their impact on EU-China relations</a>. Grzegorz has published articles on <a href="https://www.merics.org/en/merics-briefs/how-ep-parties-see-china-ev-exports-trade-and-technology-council" target="_blank" rel="noopener">how EP parties see China</a> and on <a href="https://www.merics.org/en/comment/meps-key-lessons-eu-china-policy-during-last-mandate" target="_blank" rel="noopener">key lessons learned by Members of the EP during the last mandate</a>.</p>
European Parliament Interpreting Corpus (EPIC)
<p>EPIC v2.0 is a parallel and trilingual (English, Italian, and Spanish) corpus of European Parliament (EP) speeches and their simultaneous interpretations. The data were collected from EP sessions in Feb–Apr, and July 2004, including the speeches of 175 speakers and an unknown number of interpreters. The current version of the EPIC (v2.0) contains 692,585 tokens (source: 247,385, target: 445,200) and 83 h 36 min 14 s of audiovisual recordings (source videos: 27 h 32 min 20 s, target audio: 56 h 3 min 54 s). It is fully transcribed, annotated, and aligned at the recording–transcript level.</p>
EuroparlExtract - Directional Parallel Corpora Extracted from the European Parliament Proceedings Parallel Corpus
<p>This dataset contains directional parallel corpora extracted from the European Parliament Proceedings Corpus (Europarl) v7 created by Philipp Koehn (see http://www.statmt.org/europarl/). For the extraction, the EuroparlExtract corpus processing toolkit by Michael Ustszewski (2017) was used. EuroparlExtract is freely available under the MIT License (see https://github.com/mustaszewski/europarl-extract).</p>
EuroparlExtract - Comparable Corpora Extracted from the European Parliament Proceedings Parallel Corpus
<p>This dataset contains comparable translational corpora extracted from the European Parliament Proceedings Corpus (Europarl) v7 created by Philipp Koehn (see http://www.statmt.org/europarl/). For the extraction, the EuroparlExtract corpus processing toolkit by Michael Ustszewski (2017) was used. Europarl Extract is freely available under the MIT License (see https://github.com/mustaszewski/europarl-extract).</p>
List of European Parliament plenary speeches selected for the corpus together with speakers' names (Nov-2014 to Apr-2018); examples of collocations of "refugee(s)", "refugié(s)", "Flüchtling(e)" and "menekült(ek)"
<p>This data relates to the article "Hidden Patterns in interpreted xenophobic discourse in the European Parliament" [in print].</p> <p>The data contains a chronological list of the plenary debates from which the speeches were taken as well as the names of each speaker. It also also contains examples of verbs collocating with the term <em>refugee(s</em>), <em>refugié(s)</em>, <em>Flüchtling(e)</em> and <em>menekült(ek)</em> in the four language versions (English, French, German, Hungarian). These collocations were identified by the author of the paper.</p> <p>The speeches were downloaded from the Multimedia Center on the European Parliament's pubilc website: <a href="https://multimedia.europarl.europa.eu/en/home">https://multimedia.europarl.europa.eu/en/home</a>.</p>
parliamentr: speeches from european parliaments in a standardized, machine-readable format
<p>Data accompagnying the R package parliamentr. Here on zenodo is the dataset, and the R package holds the code used to scrape & clean the data, as well as code to download and use this dataset.</p> <p>Currently under development </p>
Multilingual test set for language identification and speech recognition from European Parliament recordings
<p>This test set for language identification and speech recognition is composed by multilingual extracts from European Parliament sessions recordings. </p> <p><strong>Dataset description</strong></p> <p>Audio files and official transcripts were downloaded from: https://www.europarl.europa.eu/plenary/en/debates-video.html</p> <p>The test set has a duration of 02h 56m 34s, composed by 15 multilingual audio files of around 12 minutes, selected from the original material to maximize the number of language changes. </p> <p>Official language labels were manually reviewed to fix start/end timestamps, and official text transcripts, where present, were added to the annotation.</p> <p>The test set covers 19 languages in total.</p> <p>The test set is presented in the following paper:</p> <p>M. Valente, F. Brugnara, G. Morrone, E. Zovato, L. Badino, "Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech", accepted to Interspeech 2024.</p> <p>For more information please refer to the README.txt in the testset .zip archive.</p> <p><strong>License and copyright</strong></p> <p>The data is released with CC0 license: https://creativecommons.org/public-domain/cc0/<br>For the raw data, see also European Parliament's legal notice: https://www.europarl.europa.eu/legal-notice/en/</p>
Supplementary Materials for article entitles 'Nominal and verbal syntax in translation and interpreting. Evidence from English speeches made in the European Parliament and their German translations and interpretations', submitted to Languages
<p>The Supplementary Materials contain the transcriptions (raw and tagged, 'sample_df.tsv'), the data frames with POS-frequencies, with ('pos_freqs_PART_split.tsv') and without ('pos_freqs.tsv') the PART-split in the German data, the data frames for the identification of interpreters ('voice_embeddings.csv'), an R-script for the statistical analysis and generation of plots ('stats.R') as well as the plots (folder 'plots').</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.