Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
40
datasets available to search
ShareScore release 0.9.0
Dataset results
40 results for “Parliament”
Aggregate Dataset on Descriptive Representation in the Austrian Parliament (2017-2024)
<p>This is the dataset aggregated at the legislative/party level by the CSIC team from the individual-level data provided by the PLUS team on descriptive representation in the Austrian lower chamber of Parliament for WP4 of the ActEU project. </p>
PPAST-GR Dataset: Greek Parliament Proceedings post WWII Analysis and Recognition
<p>The first post-WWII years in Greece were devastating. After a brutal Nazi occupation, the Greek Civil War (1946–1949) erupted. It wrecked the economy and the country’s infrastructure and altered politics and the social fabric for decades to come. For the tense and unstable first years of the conflict (1946–1947), a study of the issues discussed in the Greek parliament could facilitate our understanding of the society at the time. An obstacle is that parliament proceedings are publicly available in a machine-readable form from 1989; before that only scanned images of the original records exist. We show that text recognition followed by natural language processing can unlock this corpus for historical research. Using Transkribus, we trained a text recogniser (1.5% CER) that we applied to 3,156 images from 1946 and 1947. As low-quality recognition is inevitable, we trained a language model on the transcribed text and applied it to recognised text, discarding records with high average cross-entropy. Using information extraction techniques, we sampled speeches that were applauded and we introduce the first quantification of issues that were thus received. All our resources will be made public.</p> <p>Our model is publicly available through the Transkribus platform (http://www.transkribus.org/) under the name "Greek Parliamentary Proceedings 1946".</p>
Greek Parliament Proceedings, 1989-2020
<p>This dataset is produced on behalf of <a href="http://www.imedd.org/">iMEdD</a> by PhD Student in Machine Learning <a href="https://orcid.org/0000-0003-3395-2182">Konstantina Dritsa</a>, with the contribution of data journalist and <a href="https://lab.imedd.org/">iMEdD Lab</a> Project Manager <a href="https://devlab.imedd.org/lab-author/%ce%ba%ce%ad%ce%bb%ce%bb%cf%85-%ce%ba%ce%b9%ce%ba%ce%ae/">Kelly Kiki</a>. iMEdD (incubator for Media Education and Development) is a non-profit journalism organisation that supports and promotes transparency, credibility and independence in journalism. Lab is iMEdD’s content production division which publishes original interactive investigative and data-driven stories by experimenting with new forms and tools in journalism. </p> <p>This dataset is the next version of a<a href="https://zenodo.org/record/2587904#.X7fMx80zaUn"> previous upload</a>, which originated from the work implemented during the course of the Master thesis entitled "<a href="http://www.pyxida.aueb.gr/index.php?op=view_object&object_id=6387">Speech quality and sentiment analysis on the Hellenic Parliament proceedings</a>" at the Athens University of Economics & Business in 2018 under the supervision of the Associate Professor <a href="https://orcid.org/0000-0002-3971-4612">Panagiotis Louridas</a>.</p> <p>This dataset includes 1,280,918 speeches (rows) of Greek parliament members with a total volume of 2.4 GB, that were exported from 5,355 parliamentary sitting record files. They extend chronologically from 1989 up to late July 2020. The dataset consists of a .csv file in UTF-8 encoding and includes the following columns of data:</p> <ul> <li> <p><strong>member_name</strong>: the official name of the parliament member who talked during a sitting.</p> </li> <li> <p><strong>sitting_date</strong>: the date that the sitting took place.</p> </li> <li> <p><strong>parliamentary_period</strong>: the name and/or number of the parliamentary period that the speech took place in. A parliamentary period includes multiple parliamentary sessions.</p> </li> <li> <p><strong>parliamentary_session</strong>: the name and/or number of the parliamentary session that the speech took place in. A parliamentary session includes multiple parliamentary sittings.</p> </li> <li> <p><strong>parliamentary_sitting</strong>: the name and/or number of the parliamentary sitting that the speech took place in.</p> </li> <li> <p><strong>political_party</strong>: the political party that the speaker belonged to the moment of their speech.</p> </li> <li> <p><strong>government</strong>: the government in force when the speech took place.</p> </li> <li> <p><strong>member_region</strong>: the electoral district the speaker belonged to.</p> </li> <li> <p><strong>roles</strong>: information about the parliamentary roles and/or government position of the speaker the moment of their speech.</p> </li> <li> <p><strong>member_gender</strong>: the sex of the speaker</p> </li> <li> <p><strong>speech</strong>: the speech that the member made during the parliamentary sitting</p> </li> </ul> <p>The methodology followed for the production of this dataset is described in the iMEdD Lab's article entitled "<a href="https://devlab.imedd.org/i-dimiourgia-tou-dataset-me-ta-koinovouleftika-praktika/">The creation of a dataset with the parliament proceedings within 31 years</a>". Scripts and relevant documentation are available on <a href="https://github.com/iMEdD-Lab/Greek_Parliament_Proceedings">GitHub</a>. </p>
Number and proportion of social workers in the cantonal and national parliaments of Switzerland in 2024
<p>Dieser Datensatz trägt für den Stichtag 1.1.2024 die Anzahl und Anteile von Sozialarbeitenden in allen 26 kantonalen Parlamenten sowie im nationalen Parlament (National- und Ständerat) der Schweiz zusammen. Die Anzahl und Anteile werden mit anderen Berufsangaben (z.B. Jurist:innen, Landwirt:innen, Lehrpersonen) verglichen. Die Berufsangaben wurden den kantonalen Staatskalendern oder den offiziellen Webseiten der jeweiligen Parlamente entnommen. Die Berufsangaben wurden von den jeweiligen Parlamentsmitgliedern selbst gemacht und sind teils nicht eindeutig zuzuordnen (z.B. "Unternehmer:in", "Geschäftsleiter:in" etc.).</p> <p>Für eine zusammenfassende Darstellung der Ergebnisse siehe: Kindler, T. (2024). Democratic advocacy: Representation of social workers in the cantonal and national parliaments of Switzerland. Eastern Switzerland University of Applied Sciences. https://doi.org/10.5281/zenodo.10413049</p>
Supplementary Data Files of the Greek Parliament Proceedings Dataset
<p>The dataset includes supplementary files of the previous upload "<a href="https://zenodo.org/record/6626316">A Greek Parliament Proceedings Dataset for Computational Linguistics and Political Analysis</a>". Specifically, it includes the 5,355 sitting record files from which speeches were extracted in the form of the conversation that took place in the Greek Parliament as well as the previous versions of the tell_all_cleaned.csv before preprocessing and cleaning, namely tell_all.csv and tell_all_FILLED.csv.</p> <ul> <li>original_data: A folder of the original record files downloaded from the website of the Greek Parliament (https://www.hellenicparliament.gr/Praktika/Synedriaseis-Olomeleias). The filenames are edited to follow the naming format "recordDate_id_periodNo_sessionNo_sittingNo.ext".</li> <li>_data: A folder of the record files converted to text format with filenames translated to English.</li> <li>tell_all.csv: The initial file of all extracted speeches before preprocessing and cleaning. The file includes the following columns: member_name, sitting_date, parliamentary_period, parliamentary_session, parliamentary_sitting, political_party, government, member_region, roles, member_gender, speaker_info, speech.</li> <li>tell_all_FILLED.csv: This file is an intermediate step of preprocessing of the tell_all.csv file. In this file, missing names of chairmen of various parliamentary sittings are filled. It includes the same columns as the tell_all.csv file.</li> </ul> <p>-------------</p> <p><strong><em>Acknowledgments:</em></strong></p> <p>This work was supported by the European Union’s Horizon 2020 research and innovation program ``FASTEN'' under grant agreement No 825328 and the non profit data journalism organization iMEdD.org.</p>
A Greek Parliament Proceedings Dataset for Computational Linguistics and Political Analysis
<p>The dataset is a new version of the previous upload and includes the following files:</p> <p>1. <strong>dataset_versions/tell_all.csv: </strong>The initial dataset of 1,280,927 extracted speeches, before preprocessing and cleaning. The speeches extend chronologically from July 1989 up to July 2020 and were exported from 5,355 parliamentary sitting record files. The file has a total volume of 2.5 GB and includes the following columns:</p> <ul> <li>member_name: the name of the individual who spoke during a sitting.</li> <li>sitting_date: the date the sitting took place.</li> <li>parliamentary_period: the name and/or number of the parliamentary period that the speech took place in. A parliamentary period is defined as the time span between one general election and the next. A parliamentary period includes multiple parliamentary sessions.</li> <li>parliamentary_session: the name and/or number of the parliamentary session that the speech took place in. A session is defined as a time span of usually 10 months within a parliamentary period during which the parliament can convene and function as stipulated by the constitution. A session can fall into the following categories: regular, extraordinary or special. In the intervals between the sessions the parliament is in recess. A parliamentary session includes multiple parliamentary sittings.</li> <li>parliamentary_sitting: the name and/or number of the parliamentary sitting that the speech took place in. A sitting is defined as a meeting of parliament members.</li> <li>political_party: the political party of the speaker.</li> <li>government: the government in force when the speech took place.</li> <li>member_region: the electoral district the speaker belonged to.</li> <li>roles: information about the parliamentary roles and/or government position of the speaker.</li> <li>member_gender: the gender of the speaker</li> <li>speech: the speech that the individual gave during the parliamentary sitting.</li> </ul> <p>2. <strong>dataset_versions/tell_all_FILLED.csv: </strong>This file is an intermediate version of the dataset that includes improvements in the consistency and completeness of the dataset, with a total volume of 2.5 GB. Specifically, this file is produced by filling the missing names of chairmen of various parliamentary sittings of the "tell_all.csv". It includes the same columns as the "tell_all.csv" file.</p> <p>3.<strong> dataset_versions/tell_all_cleaned.csv: </strong>This version of the dataset is the result of further cleaning and preprocessing and is used for our word usage change study. It consists of 1,280,918 speech fragments of Greek parliament members in the order of the conversation that took place, with a total volume of 2.12 GB. It includes the same columns as the aforementioned versions. The preprocessing includes the replacement of all references to political parties with the symbol "@" followed by an abbreviation of the party name, using regular expressions that capture different grammatical cases and variations. It also includes the removal of accents, strings with length less than 2 characters, all punctuation except full stops, and the replacement of stopwords with "@sw".</p> <p>4. <strong>wiki_data</strong>: A folder of modern Greek female and male names and surnames and their available grammatical cases crawled from the entries of the Wiktionary Greek names category (https://en.wiktionary.org/wiki/Category:Greek_names). We produced the grammatical cases of the missing grammatical entries according to the rules of the Greek grammar and saved the files in the same folder by adding to their filenames the string "_populated.json".</p> <p>5. <strong>parl_members_activity_1989onwards_with_gender.csv</strong>: The Greek Parliament website provides a<br> <a href="https://www.hellenicparliament.gr/Vouleftes/Diatelesantes-Vouleftes-Apo-Ti-Metapolitefsi-Os-Simera/">list</a> of all the elected members of parliament since the fall of the military junta in Greece, in 1974. We collected and cleaned the data, added the gender and kept the elected members from 1989 onwards, matching the available parliament proceeding records. This dataset includes the full names of the members, the date range of their service, the political party they served, the electoral district they belonged to and their gender.</p> <p>6. <strong>formatted_roles_gov_members_data.csv</strong>: As government members we refer to individuals in ministerial or other government posts, regardless of whether they were elected in the parliament. This information is available in the website of the <a href="https://gslegal.gov.gr/?page_id=776&sort=time">Secretariat General for Legal and Parliamentary Affairs</a>. The government members dataset includes the full names of the official individuals, the name of the role they were given, the date range of their service at each specific role and their gender.</p> <p>7. <strong>governments_1989onwards.csv</strong>: A dataset of government information including the names of governments since 1989, their start and end dates, and a URL that points to the respective official government web page of each past government. The data is crawled from the website of the <a href="https://gslegal.gov.gr/?page_id=776&sort=time">Secretariat General for Legal and Parliamentary Affairs</a>.</p> <p>8. <strong>extra_roles_manually_collected.csv</strong>: A dataset with manually collected information from Wikipedia about additional government or parliament posts such as Chairman of the Parliament, party leaders, opposition leaders and other information.</p> <p>9. <strong>all_members_activity.csv</strong>: A dataset of all the information of the aforementioned files 3,4,5,6 merged. Each row of the file includes the full name of the individual, the start and end date of their term of office, the political party and electoral district they belonged to, their gender, the parliamentary and/or government positions that they held along with start and end dates, and the name of the government that was in power during their term of office. An individual can change political parties or become an independent member of the parliament during a parliamentary period, thus having more than one entries/rows in the file.</p> <p>10. <strong>freqs_for_semantic_shift_cleaned_data_decade1990.csv & freqs_for_semantic_shift_cleaned_data_decade2010.csv</strong>: Files of frequencies of words in the corpora of the decades 1990-1999 and 2010-2019.</p> <p>11. <strong>compass_top100.csv:</strong> Top 100 most changed words between the decades 1990-1999 and 2010-2019, as computed with the use of the Compass tool by V. D. Carlo et. al. [1].</p> <p>12. <strong>compass_fc_top100.csv</strong>: Top 100 most changed words between the decades 1990-1999 and 2010-2019, as computed with the use of the Compass tool [1] in combination with the frequency cut-offs of the Gonen et. al. approach [3]. For the frequency cut-offs, the files in bullet 8 are used.</p> <p>13. <strong> procrustes_top100.csv</strong>: Top 100 most changed words between the decades 1990-1999 and 2010-2019, as computed with the use of the Orthogonal Procrustes approach of Hamilton et. al. [2].</p> <p>14. <strong>nn_top100.csv</strong>: Top 100 most changed words between the decades 1990-1999 and 2010-2019, as computed with the use of the Gonen et. al. approach [3].</p> <p>15. <strong>second_order_top100.csv</strong>: Top 100 most changed words between the decades 1990-1999 and 2010-2019, as computed with the use of the Second-Order Similarity approach by Hamilton et. al. [4].</p> <p>16. <strong>top100_minfreq50.xls</strong>: An .xls file for convinient viewing of the top 100 most changed words per approach with minimum frequency of 50 occurrences, produced by merging the aforementioned files 11, 12, 13, 14, 15 and 16.</p> <p>17. <strong>freqs_for_semantic_shift_cleaned_data_period1997_2007.csv & freqs_for_semantic_shift_cleaned_data_period2008_2018.csv</strong>: Files of frequencies of words in the corpora of the decades before (1997_2007) and during (2008_2018) the Greek economic crisis.</p> <p>18. <strong>semantic_shifts_dichotomy_crisis_compass_1997_2007_2008_2018_atleast50.csv</strong>: A file with the top 100 most changed words between between the decades before (1997-2007) and during (2008-2018) the Greek economic crisis. The computations are implemented with the use of the Compass tool.</p> <p>19. <strong>selected_topics_shift_per_period_compass.csv</strong>: The usage change of selected topics/words of generic political interest between pairs of consecutive parliamentary periods. The computations are implemented with the use of the Compass tool.</p> <p>20. <strong>semantic_shifts_party_embeddings_per_period_merged_compass.csv</strong>: The usage change of selected political party names that have played an important role in recent political history, namely New Democracy (ND), the Panhellenic Socialist Movement (PASOK), the Coalition of the Radical Left - Progressive Alliance (SYRIZA), the Communist Party of Greece (KKE), the Coalition of the Left, of Movements and Ecology (SYN) and Golden Dawn (GD).</p> <p>-------------</p> <p><strong><em>Citations:</em></strong></p> <p>[1] Valerio Di Carlo, Federico Bianchi, and Matteo Palmonari. Training Temporal Word Em- beddings with a Compass. In <em>Proceedings of the Thirty–Third AAAI Conference on Artificial Intelligence</em>, AAAI’19, pages 6326–6334, 2019. doi: 10.1609/aaai.v33i01.33016326.</p> <p>[2] William L. Hamilton, Jure Leskovec, and Dan Jurafsky. Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change. In <em>Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</em>, ACL 2016, pages 1489– 1501, Berlin, Germany, August 2016. Association for Computational Linguistics. doi: 10. 18653/v1/P16-1141. URL https://www.aclweb.org/anthology/P16-1141.</p> <p>[3] Hila Gonen, Ganesh Jawahar, Djamé Seddah, and Yoav Goldberg. Simple, Interpretable and Stable Method for Detecting Words with Usage Change across Corpora. In <em>Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</em>, ACL 2020, pages 538– 555, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl- main.51. URL https://aclanthology.org/2020.acl-main.51.</p> <p>[4] William L. Hamilton, Jure Leskovec, and Dan Jurafsky. Cultural Shift or Linguistic Drift? Comparing Two Computational Measures of Semantic Change. In <em>Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing</em>, EMNLP 2016, pages 2116–2121, Austin, Texas, November 2016. Association for Computational Linguistics. doi: 10.18653/v1/D16-1229. URL https://www.aclweb.org/anthology/D16-1229.</p> <p>-------------</p> <p><strong><em>Acknowledgments:</em></strong></p> <p>This work was supported by the European Union’s Horizon 2020 research and innovation program ``FASTEN'' under grant agreement No 825328 and the non profit data journalism organization iMEdD.org.</p>
damascus-parliament-perchoir-pattern
Hexagon based pattern with mostly squared parts, reinterpreted as an isometric view of a cubic 3D structure with mostly diagonal parts. 2D reference pictures: [](https://www.vectorstock.com/941333) [](http://hostmaster.syrianhistory.com/en/photos/8262) [](http://hostmaster.syrianhistory.com/en/photos/7820) See also: • https://regolo54.tumblr.com/post/177195904237/fatepur-sikri • https://www.facebook.com/591752864907764 • https://www.facebook.com/184576975625357 same basic shape with additional pathes (Iran). The underlying shape is less obvious without coloring the kite polygons, in [photo 38](https://books.google.fr/books?id=o9IxDwAAQBAJ&pg=PA68) and [fig. 112a-b](https://books.google.fr/books?id=o9IxDwAAQBAJ&pg=PA247) from J. Bonner reference book. Source: Objaverse 1.0 / Sketchfab
Parliament of Victoria, Melbourne
This was a 3D model I worked on from September 6th to September 9th, 2019. This is the Parliament of Victoria, located in Melbourne, Australia. It was made for a VR project set in the city of Melbourne produced by Samurai Punk. This also happened to be my first commission as a 3D artist, which I am exceptionally happy about. Check out more here: https://christian-romero.com/3d-modeling-animation/victoria-parliament/ Source: Objaverse 1.0 / Sketchfab
Air Raid Shelter Parliament Square
The entrance to a WW2 air raid shelter in the side of the Treasury Building on Parliament Square, London. This is one of five along this road. This is the furthest east. 158 photos taken in July 2020 with a Sony a6000 and processed in Reality Capture. Source: Objaverse 1.0 / Sketchfab
EuroparlExtract - Directional Parallel Corpora Extracted from the European Parliament Proceedings Parallel Corpus
<p>This dataset contains directional parallel corpora extracted from the European Parliament Proceedings Corpus (Europarl) v7 created by Philipp Koehn (see http://www.statmt.org/europarl/). For the extraction, the EuroparlExtract corpus processing toolkit by Michael Ustszewski (2017) was used. EuroparlExtract is freely available under the MIT License (see https://github.com/mustaszewski/europarl-extract).</p>
EuroparlExtract - Comparable Corpora Extracted from the European Parliament Proceedings Parallel Corpus
<p>This dataset contains comparable translational corpora extracted from the European Parliament Proceedings Corpus (Europarl) v7 created by Philipp Koehn (see http://www.statmt.org/europarl/). For the extraction, the EuroparlExtract corpus processing toolkit by Michael Ustszewski (2017) was used. Europarl Extract is freely available under the MIT License (see https://github.com/mustaszewski/europarl-extract).</p>
Electric car subsidies in social media and parliament 2022–2023: analysis results
<p>Electric car subsidies in social media and the Finnish parliament 2022–2023: comparative analysis results obtained using three methods for qualitative classification of textual data.</p>
An annotated dataset of Central Acts enacted by the Indian Parliament for legal research
<p>This dataset consists of 858 annotated Central Acts enacted by the Indian Parliament from the year 1838 to 2020. The Central Acts are available in a portable document format (PDF) on <a href="https://www.indiacode.nic.in/">India Code</a> website. This website has been developed by the Legislative Department under the Ministry of Law and Justice in the Government of India and is a digital repository of all enforced Central and State Acts enacted by the Indian Parliament.</p> <p>We used <em>pdfminer.six</em>, a text extraction python library for PDF documents, to convert these unstructured PDFs into a structured JSON format. Furthermore, we used regular expressions to remove the noisy text and extract meta-information (e.g., initial portions of the document containing act title, enactment date, and other meta information) from these acts. The result was a dataset of 858 structured JSON files corresponding to each of the Central Acts with relevant metadata.</p> <p>The annotation schema for the Central Act dataset consists of the following fields:</p> <ul> <li><strong>Act Title:</strong> The title, usually called the “short title," is the name by which an act is known. The short title means the term by which an act or resolution may be cited and often references the date of commencement.</li> <li><strong>Act ID:</strong> It includes the number of the act and the year of its enactment.</li> <li><strong>Enactment Date:</strong> Specifies the date on which the bill becomes an act through approval.</li> <li><strong>Act Definition:</strong> Also known as the "long title," is a summarized breakdown of the act's purpose. It may be presented in only a few words, in some cases, while in others, it can run to several pages.</li> <li><strong>Chapters and Parts:</strong> Chapters or Parts are subdivisions used to arrange the information in an act or other piece of legislation. There is no set pattern as to how these are applied in an act. Sometimes there are only Chapters, or only Parts and other times unlabeled headings. In some cases, Parts are divided into Chapters. So there is no hard-and-fast rule regarding the use of Chapters and Parts in Indian legislation. Furthermore, we define the Chapter (or Part) ID, for example, <em>CHAPTER IV, PART I</em> and Chapter (or Part) Name to identify by its heading. If the Chapter (or Part) has been omitted or repealed within the act, then the ID field will contain <em>[Omitted]</em> or <em>[Repealed] </em>respectively. </li> <li><strong>Sections:</strong> An arrangement of Sections (this is not part of the law, but assists in navigating through an Act or other piece of legislation, especially lengthy documents).</li> <li><strong>Subheadings:</strong> This field is optional and is contained within Chapters or Parts. It is further nested into Sections.</li> <li><strong>Schedule, Annexure, Appendix and Forms:</strong> Many acts have schedules attached at the end, which generally add more details (such as maps or fees), give examples of forms or sets of rules to be used under the act, or list sections of other acts amended by this one. The majority of them will contain amendments and/or repeals of legislation. However, the Schedule to an act always contains supplementary information of importance in meeting the act's objectives.</li> <li><strong>Footnotes:</strong> Footnotes are notes placed at the bottom of a page. They cite references or comment on a designated part of the text above it. This field consist of a key, value pair of page number and footnote text, respectively. It can be empty if there are no footnotes in an act document.</li> </ul>
Semantically tagged Finnish parliament discussions 1991-2015
<p><strong>Semantically tagged Finnish parliament discussions 1991-2015</strong></p> <p>The original data is this (Rauh et al, 2017):</p> <p>https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/E4RSP9</p> <p>We have imported the raw text out of the original data set without speaker and party tags. The Finnish text has been first tagged with UD2 parser using the Mylly service of the Language Bank of Finland. After UD2 parse, semantic tags have been added to the text with FiST (Kettunen, 2019).</p> <p><strong>Output form</strong></p> <p># newdoc</p> <p># newpar</p> <p># sent_id = 1</p> <p># text = Arvoisa herra puhemies!</p> <p>Arvoisa arvoisa Z99 amod</p> <p>herra#herra#Noun#S2.2m S9 compound:nn</p> <p>puhemies#puhemies#Noun#G1.1/S2 root</p> <p>! PUNCT</p> <p> </p> <p>The output consists of the 1) token word form of the text, 2) lemma of the token, 3) POS, 4) semantic tag and 5) UD2 grammatical relation of the word in the sentence.</p> <p>Our semantic tagger does not resolve ambiguity. If the lexeme has multiple possible semantic tags, they are all included in the output (e.g., herra#herra#Noun#<strong>S2.2m S9</strong> compound:nn). Slash notation in the semantic tags (e.g., puhemies#puhemies#Noun#<strong>G1.1/S2</strong>) indicates that the word can belong to two or more categories (Löfbeg, 2017). If the semantic category of the word is not recognized, the word is tagged with Z99.</p> <p>The data consists of 4 036 269 sentences and ca. 65.247 million words. Lexical coverage of FiST for the data is 86.05 %, i.e. 86% of the words are known for the tagger and marked with a semantic tag.</p> <p><strong>References</strong></p> <p>Kettunen, Kimmo (2019). FiST – towards a Free Semantic Tagger of Modern Standard Finnish. IWCLUL2019, http://aclweb.org/anthology/W19-0306</p> <p>Rauh, Christian; De Wilde, Pieter; Schwalbach, Jan, 2017, "Corp_Eduskundta.Rdata", The ParlSpeech data set: Annotated full-text vectors of 3.9 million plenary speeches in the key legislative chambers of seven European states, https://doi.org/10.7910/DVN/E4RSP9/U8VZHK, Harvard Dataverse, V1.</p> <p>Lofberg, L. (2017). Creating large semantic lexical resources for the Finnish language. [Doctoral Thesis, Lancaster University]. Lancaster University. https://doi.org/10.17635/lancaster/thesis/3</p> <p>UCREL Semantic Analysis System (USAS). https://ucrel.lancs.ac.uk/usas/</p>
Trinity College Dublin and Parliament House
Source: Objaverse 1.0 / Sketchfab
List of European Parliament plenary speeches selected for the corpus together with speakers' names (Nov-2014 to Apr-2018); examples of collocations of "refugee(s)", "refugié(s)", "Flüchtling(e)" and "menekült(ek)"
<p>This data relates to the article "Hidden Patterns in interpreted xenophobic discourse in the European Parliament" [in print].</p> <p>The data contains a chronological list of the plenary debates from which the speeches were taken as well as the names of each speaker. It also also contains examples of verbs collocating with the term <em>refugee(s</em>), <em>refugié(s)</em>, <em>Flüchtling(e)</em> and <em>menekült(ek)</em> in the four language versions (English, French, German, Hungarian). These collocations were identified by the author of the paper.</p> <p>The speeches were downloaded from the Multimedia Center on the European Parliament's pubilc website: <a href="https://multimedia.europarl.europa.eu/en/home">https://multimedia.europarl.europa.eu/en/home</a>.</p>
parliamentr: speeches from european parliaments in a standardized, machine-readable format
<p>Data accompagnying the R package parliamentr. Here on zenodo is the dataset, and the R package holds the code used to scrape & clean the data, as well as code to download and use this dataset.</p> <p>Currently under development </p>
Multilingual test set for language identification and speech recognition from European Parliament recordings
<p>This test set for language identification and speech recognition is composed by multilingual extracts from European Parliament sessions recordings. </p> <p><strong>Dataset description</strong></p> <p>Audio files and official transcripts were downloaded from: https://www.europarl.europa.eu/plenary/en/debates-video.html</p> <p>The test set has a duration of 02h 56m 34s, composed by 15 multilingual audio files of around 12 minutes, selected from the original material to maximize the number of language changes. </p> <p>Official language labels were manually reviewed to fix start/end timestamps, and official text transcripts, where present, were added to the annotation.</p> <p>The test set covers 19 languages in total.</p> <p>The test set is presented in the following paper:</p> <p>M. Valente, F. Brugnara, G. Morrone, E. Zovato, L. Badino, "Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech", accepted to Interspeech 2024.</p> <p>For more information please refer to the README.txt in the testset .zip archive.</p> <p><strong>License and copyright</strong></p> <p>The data is released with CC0 license: https://creativecommons.org/public-domain/cc0/<br>For the raw data, see also European Parliament's legal notice: https://www.europarl.europa.eu/legal-notice/en/</p>
Supplementary Materials for article entitles 'Nominal and verbal syntax in translation and interpreting. Evidence from English speeches made in the European Parliament and their German translations and interpretations', submitted to Languages
<p>The Supplementary Materials contain the transcriptions (raw and tagged, 'sample_df.tsv'), the data frames with POS-frequencies, with ('pos_freqs_PART_split.tsv') and without ('pos_freqs.tsv') the PART-split in the German data, the data frames for the identification of interpreters ('voice_embeddings.csv'), an R-script for the statistical analysis and generation of plots ('stats.R') as well as the plots (folder 'plots').</p>
Mental Health and Well-being of People Who Seek Help From Their Member of Parliament
ClinicalTrials.gov study NCT04203966. IPD Sharing: NO. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.