Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
79
datasets available to search
ShareScore release 0.9.0
Dataset results
79 results for “debate”
Extracted patterns about transport from the French Great National Debate (Grand Débat National)
<p>This data set is composed by 5 geojson files, that can be used to generate maps of mainland France :</p> <ul> <li>motifs_all.geojson : pattern about transport extracted from contributions of the French Great National Debate (Grand Débat National). Original dataset : https://granddebat.fr/pages/donnees-ouvertes</li> <li>bikeway_fr.geojson and railroad_fr.geojson : cycleways and railways of mainland France, from Open Street Map. Original dataset : https://download.geofabrik.de/europe/france.html</li> <li>trainstations.geojson : train stations and halts of mainland France, from Open Street Map. Original dataset : https://download.geofabrik.de/europe/france.html</li> <li>au2010_carto.geojson : categorized urban areas of mainland France. Original dataset : https://www.insee.fr/fr/information/2115011</li> <li>communesimportantes.geojson : the main cities of mainland France</li> </ul> <p>The data set is in French.</p> <p><em>Ce jeu de données est composé de 5 fichiers geojson qui peuvent être utilisés pour générer des cartes en France métropolitaine :</em></p> <ul> <li><em>motifs_all.geojson : motifs à propos du transport extraient des contributions en ligne au Grand Débat National. Jeu de données d'origine : https://granddebat.fr/pages/donnees-ouvertes</em></li> <li><em>bikeway_fr.geojson and railroad_fr.geojson : pistes cyclables et voies ferrées en France métropolitaine, venant d'Open Street Map. Jeu de données d'origine : https://download.geofabrik.de/europe/france.html</em></li> <li><em>trainstations.geojson : gares et petites gares en France métropolitaine, from Open Street Map. Original dataset : https://download.geofabrik.de/europe/france.html</em></li> <li><em>au2010_carto.geojson : aires urbaines catégorisées en France métropolitaine, définies par l'INSEE. Jeu de données d'origine : https://www.insee.fr/fr/information/2115011</em></li> <li><em>communesimportantes.geojson : principales villes de France métropolitaine</em></li> </ul>
SRBCorp: Corpus of Parliamentary Debates in Serbia
<p>The repository contains a cleaned and pre-processed corpus of parliamentary debates from the National Assembly of Serbia. The corpus is accompanied by the metadata on elected representatives and their political parties. It covers the period of 1997-2020 (eight terms) and counts over 300 thousand speeches.</p> <p><strong>If you use the dataset, please cite</strong>: Mochtak, Michal, Josip Glaurdić, and Christophe Lesschaeve (2022): SRBCorp: Corpus of Parliamentary Debates in Serbia (v1.1.1), https://doi.org/10.5281/zenodo.6521648.</p> <p>v1.1.1 (<strong>latest version</strong>)<br> - added the concept DOI to codebooks (DOI was generated only after the repository was published)</p> <p>v1.1.0<br> - added a new variable for policy category "tag" using ML model trained on known tags of agenda points in the parliaments of Croatia and Bosnia-Herzegovina</p> <p>v1.0.0<br> - originally posted on GESIS repository (https://doi.org/10.7802/2389); migrated to ZENODO due to limitations concerning the concept DOI</p>
CROCorp: Corpus of Parliamentary Debates in Croatia
<p>The repository contains a cleaned and pre-processed corpus of parliamentary debates from the Croatian Parliament (Sabor). The corpus is accompanied by the metadata on elected representatives and their political parties. It covers the period of 2003-2020 (five complete terms) and counts over 500 thousand speeches.</p> <p><strong>If you use the dataset, please cite</strong>: Mochtak, Michal, Josip Glaurdić, and Christophe Lesschaeve (2022): CROCorp: Corpus of Parliamentary Debates in Croatia (v1.1.1), https://doi.org/10.5281/zenodo.6521372.</p> <p>v1.1.1 (<strong>latest version</strong>)<br> - added the concept DOI to codebooks (DOI was generated only after the repository was published)</p> <p>v1.1.0<br> - improved coding of dummy variable "moderator" (using less error-prone alghoritm for detecting the modertor role)<br> - fixed issue with agenda points which are conncatenated while preserving a unique web link<br> - recoded agenda points tags using better ML model (transformer architecture)</p> <p>v1.0.0<br> - originally posted on GESIS repository (migrated to ZENODO due to limitations concerning the concept DOI)</p>
BiHCorp: Corpus of Parliamentary Debates in Bosnia and Herzegovina
<p>The repository contains a cleaned and pre-processed corpus of parliamentary debates from the Parliamentary Assembly of Bosnia and Herzegovina. The corpus is accompanied by the metadata on elected representatives and their political parties. It covers the period of 1998-2018 (six complete terms) and counts over 127 thousand speeches.</p> <p><strong>If you use the dataset, please cite</strong>: Mochtak, Michal, Josip Glaurdić, Christophe Lesschaeve, and Ensar Muharemović (2022): BiHCorp: Corpus of Parliamentary Debates in Bosnia and Herzegovina (v1.1.1),<br> https://doi.org/10.5281/zenodo.6517697.</p> <p>v1.1.1 (<strong>latest version</strong>)<br> - added the concept DOI to codebooks (DOI was generated only after the repository was published)</p> <p>v1.1.0<br> - fixed a typo in one of the debates' date<br> - fixed minor inconsistencies in the tag column</p> <p>v1.0.0<br> - originally posted on GESIS repository (https://doi.org/10.7802/2387); migrated to ZENODO due to limitations concerning the concept DOI</p>
User study Data: Boosting Intellectual Humility During Search on Debated Topics
<pre><strong>User study data </strong> The following column headers correspond to the following study variables: Intervention = Intervention (CONTROL= control, DUMMYCONTROL = ATI control, PRIME = prime, QUESTIONNAIRE = remind, FULL = reinforce) DV1_AC_Clicks = Attitude confirming clicks DV2_Lowest_Rank = Lowest rank clicked DV3_Dwell_Time = Dwell time DV4_Task_Completion = Task completion time DV5_Cumulative_Clicks = Cumulative clicks IH = Intellectual Humility Ranking = Ranking Topic = Topic rationale = Rationale for behavior (free text) rationale_category = Rationale for behavior (category, one of IH_driven = driven by IH, ranking_driven = ranking, bias_driven = confirmation bias, content/form_driven = content/form, task_driven/unclear = task/unclear)) Att_change = Attitude change Knowledge = Knowledge gain (1 = no knowledge gain, 5 = substantial knowledge gain) NASA.Mental = Reflection on search task, mental demand NASA.Temporal = Reflection on search task, temporal demand NASA.Performance = Reflection on search task, performance NASA.Effort = Reflection on search task, effort NASA.Frustration = Reflection on search task, frustration</pre>
Cartolabe Grand Debat 01-03/2019
<p>The Grand Débat dataset shows data from the <a href="https://granddebat.fr/">citizen's debate initiative</a> which took place in France from January to March 2019. It shows propositions from citizens in response to government questions in regard to societal and political topics.</p>
Corpus of political tweets UK-EU-DEBATE-20-21
<p> </p> <p>The <em>UK-EU-DEBATE-20-21</em> corpus was collected within the framework of the collaborative research project OLiNDiNUM (<em><a href="https://olindinum.huma-num.fr">Observatoire LINguistique du DIscours NUMérique</a> / </em>Linguistic Observatory of Online Debate) to be part of a shared research archive of shared corpora and resources. </p> <p>The corpus was selected with a view to examining the UK-EU media debate on the COVID-19 vaccination campaign following a specific transformative moment: the signature of the Brexit withdrawal agreement by the UK and the EU at the end of January 2021.</p> <p>The data were retrieved through the Application Programming Interface of the social networking site Twitter, using the accounts of key political actors in the UK government and EU institutions over a period of 14 months (1 February 2020–31 March 2021). The composition of the corpus is illustrated in the table.</p> <p> </p> <table> <tbody> <tr> <td><em>Political Actor</em></td> <td><em>Role</em></td> <td><em>Account</em></td> <td><em>Tweets</em></td> </tr> <tr> <td>Boris Johnson</td> <td>UK Prime Minister</td> <td>@BorisJohnson</td> <td>1186</td> </tr> <tr> <td>Dominic R. Raab</td> <td>UK Foreign Secretary</td> <td>@DominicRaab</td> <td>1468</td> </tr> <tr> <td>Priti Patel</td> <td>UK Home Secretary</td> <td>@pritipatel</td> <td>941</td> </tr> <tr> <td>Ursula von der Leyen</td> <td>President of the European Commission</td> <td>@vonderleyen</td> <td>1338</td> </tr> <tr> <td>David Sassoli</td> <td>President of the European Parliament</td> <td>@EP_President</td> <td>554</td> </tr> <tr> <td>Charles Michel</td> <td>President of the Council of the European Union</td> <td>@eucopresident</td> <td>675</td> </tr> </tbody> </table> <p> </p> <p>The data are supplied in separate .csv files (tab-delimited format). Each row contains the text of the tweet (<em>data__text</em>) and the tweet identifier (<em>data__id</em>) as a header. The tweet identifier enables swift retrieval of the original tweet by searching https://twitter.com/anyuser/status/<em>data__id. </em></p> <p> </p>
Methodological Appendix for FEUTURE Online Paper No. 28 "Narratives of a Contested Relationship: Unravelling the Debates in the EU and Turkey"
<p>This is the methodological appendix for the narrative analysis conducted by the researchers from the University of Cologne (UzK) and Middle East Technical University (METU) within the scope of the ongoing research project, which is entitled “The Future of EU-Turkey Relations: Mapping Dynamics and Testing Scenarios” (FEUTURE) and funded by the European Union’s Horizon 2020 Research and Innovation Programme. The appendix is designed to provide comprehensive information on the operationalization of the qualitative research that was carried out for the FEUTURE Online Paper No. 28 “Narratives of a Contested Relationship: Unravelling the Debates in the EU and Turkey” published in February 2019.</p> <p>The following sections present details on the selected actors, data sampling and collection, codebook and variables, and overall time span of the research.</p>
Reddit Climate Change Debate Dataset
<p> </p> <p>This dataset contains pairwise interactions between Reddit users debating climate change on general-purpose subreddits. Each account is enriched with information about the stance concerning climate change (e.g., whether one denies or believes climate change exists) estimated by a deep neural model. </p> <p>All data is anonymized, and no personally identifiable information is released.</p> <h2><strong>Dataset</strong></h2> <p>Interactions are stored in four files, each encompassing 3 months of interactions in 2022. <span>Each interaction file <strong><em>climatechange-X.csv</em></strong> contains three columns identifying source, target, and weight, respectively.</span> The resulting graphs are directed.</p> <p>The <strong>climatechange-opinions.csv</strong> file contains three columns identifying node, opinion, and time window. Opinions are stored as floats in [-1,1] such that 1 implies maximum adherence with deniers, -1 implies maximum adherence with supporters, and 0 implies neutrality. Thus, a line like 42,0.99,2 should be read as "node 42 is a climate change denier in the second quarter of 2022".</p> <p>Code to reproduce the experiments in the paper is released in a jupyter notebook.</p> <p>For further information on fields and volumes, please refer to the data paper.</p> <h3><strong>Citation</strong></h3> <p>If used for research purposes, please cite the following paper describing the dataset details:</p> <p><em>TBD</em></p> <h3><strong>Acknowledgements</strong></h3> <p>This work is supported by:</p> <ul> <li>the European Union – Horizon 2020 Program under the scheme “INFRAIA-01-2018-2019 – Integrating Activities for Advanced Communities”,<br>Grant Agreement n.871042, “SoBigData++: European Integrated Infrastructure for Social Mining and Big Data Analytics” (http://www.sobigdata.eu); </li> <li>SoBigData.it which receives funding from the European Union – NextGenerationEU – National Recovery and Resilience Plan (Piano Nazionale di Ripresa e Resilienza, PNRR) – Project: “SoBigData.it – Strengthening the Italian RI for Social Mining and Big Data Analytics” – Prot. IR0000013 – Avviso n. 3264 del 28/12/2021;</li> <li>EU NextGenerationEU programme under the funding schemes PNRR-PE-AI FAIR (Future Artificial Intelligence Research). </li> </ul> <p> </p> <p> </p>
AustroParl Corpus of Parliamentary Debates
<p>The <em>AustroParl Corpus of Parliamentary Debates</em>, prepared in the <a href="http://polmine.github.io">PolMine Project</a>, comprises all protocols of plenary sessions in the Austrian <em>Nationalrat </em>between 1996 and 2019. The corpus is built based on pdf documents issued by the <em>Nationalrat</em>. The R package <a href="https://polmine.github.io/frappp_slides/slides_en.html">frappp</a> has been used to extract structural information from the orginal text and to prepare an XML version of the corpus (preliminary TEI format). The structural annotation comprises speaker, party affiliation, parliamentary group affiliation, role, legislative period, session, date, interjections, year and agenda item.</p> <p>This release offers a linguistically annotated and indexed format of the corpus. As part of the corpus preparation pipeline, the data has been linguistically annotated (using the <a href="https://www.cis.uni-muenchen.de/~schmid/tools/TreeTagger/">TreeTagger</a> and <a href="https://stanfordnlp.github.io/stanfordnlp/">StanfordNLP</a>) and imported into the <a href="http://cwb.sourceforge.net/">Corpus Workbench (CWB)</a>. The linguistic annotation comprises POS-tagging and lemmatization.</p> <p>This language resource is still very much in development and comes without any guarantees.</p>
ParisParl Corpus of Parliamentary Debates
<p>The <em>ParisParl Corpus of Parliamentary Debates</em>, prepared in the <a href="http://polmine.github.io">PolMine Project</a>, comprises all protocols of plenary sessions in the French <em>Assemblée nationale</em> between 1996 and 2019. The corpus is built based on pdf documents issued by the <em>Assemblée nationale</em>. The R package <a href="https://polmine.github.io/frappp_slides/slides_en.html">frappp</a> has been used to extract structural information from the orginal text and to prepare an XML version of the corpus (preliminary TEI format). The structural annotation comprises speaker, party affiliation, parliamentary group affiliation, role, legislative period, session, date, interjections, year and agenda item.</p> <p>This release offers a linguistically annotated and indexed format of the corpus. As part of the corpus preparation pipeline, the data has been linguistically annotated (using the <a href="https://www.cis.uni-muenchen.de/~schmid/tools/TreeTagger/">TreeTagger</a> and <a href="https://stanfordnlp.github.io/stanfordnlp/">StanfordNLP</a>) and imported into the <a href="http://cwb.sourceforge.net/">Corpus Workbench (CWB)</a>. The linguistic annotation comprises POS-tagging and lemmatization.</p> <p>This language resource is still very much in development and comes without any guarantees.</p>
RegioParl Corpus of Parliamentary Debates in Germany's Regional Parliaments (2000-2012)
<p>RegioParl is a linguistically annotated and indexed variant (CWB data format) of a corpus of parliamentary debates in Germany's regional parliaments. The corpus includes all parliamentary protocols of all "Landtage" and the "Bundesrat" between 2000 and early 2012. Some regional parliaments are represented in the corpus with data that predate the year 2000. The corpus preparation resulted from a cooperation project with the Institut für Deutsche Sprache Mannheim (IDS).</p>
Conclusiones del debate sobre evaluación online. Jornadas Conversación casUSAL
<p>Conclusiones del debate sobre evaluación <em>online</em> celebrado el 2 de julio de 2020 en las Jornadas Conversación casUSAL (<a href="https://facultadcero.org/encuentroUSAL/">https://facultadcero.org/encuentroUSAL/</a>)</p>
Conclusions of the online assessment debate held on July 2, 2020, at the University of Salamanca (Spain)
<p>Conclusions of the online assessment debate held on July 2, 2020, at the University of Salamanca (Spain)</p>
Cartolabe Debat RUA
<p>Debat RUA (french)</p>
VivesDebate: A New Annotated Multilingual Corpus of Argumentation in a Debate Tournament
<p>The application of the latest Natural Language Processing breakthroughs in computational argumentation has shown promising results which have raised the interest in this area of research. However, the available corpora with argumentative annotations are often limited to a very specific purpose or are not of adequate size to take advantage of state-of-the-art deep learning techniques (e.g., deep neural networks). In this paper, we present VivesDebate, a large, richly annotated, and versatile professional debate corpus for computational argumentation research. The corpus has been created from 29 transcripts of a debate tournament in Catalan and has been machine-translated into Spanish and English. The annotation contains argumentative propositions, argumentative relations, debate interactions, and professional evaluations of the arguments and argumentation. The presented corpus can be useful for research on a heterogeneous set of computational argumentation underlying tasks such as argument mining, argument analysis, argument evaluation, or argument generation among others. All this makes VivesDebate a valuable resource for computational argumentation research within the context of massive corpora aimed at Natural Language Processing tasks.</p>
Extracted opinions from the French National Great Debate (Grand Débat National) about transport and socio-economic description of municipalities participating to the debate
<p>This data set contains :</p> <p>- motifs_with_geom.json : extracted propositions from the answers to the online public consultation related to the French National Great Debate (Grand Débat National) which took place in 2019. The propositions consist of automatically extracted patterns. The propositions are related to transport. Extracted propositions are also georeferenced by postal codes and associated coordinates (WGS 84).</p> <p>- communes_with_all_data.geojson : socio-economic features which describes municipalities in France (mainland + overseas territories). These data are also georeferenced by postal codes and associated coordinates (WGS 84). data_geo.zip are the source files to build this file. The features are described in data_cadrage_meta.csv</p> <p>- motifsxcommunes_without_geom.csv : joined file of extracted propositions and socio-economic features of municipalities, without coordinates. These data are used to performed statistic analysis.</p> <p>- transportonto4.owl : ontology describing transport in French, used to extract propositions related to transport.</p>
LadiesDebating-KG: A Knowlege Graph for representing the "Edinburgh Ladies' Debating Society Digital Collection" (1865 - 1880)
<p>This Knowlege Graph represents the information of the "Edinburgh Ladies’ Debating Society<strong>"</strong> (years: 1865 - 1880) collection in RDF (ttl format). This collection consists of the complete runs of two Edinburgh journals, <strong>‘The Attempt’ (10 volumes, 1865-74)</strong> and its successor ‘<strong>The Ladies’ Edinburgh Magazine’ (6 volumes, 1875-80)</strong>. These publications were produced by a leading Edinburgh women’s club, known during the period as the Edinburgh Essay Society or the Ladies’ Edinburgh Essay Society, but subsequently as the Ladies’ Edinburgh Debating Society. The Society existed from 1865 to 1935. The raw dataset is provided by the NLS in this <a href="https://data.nls.uk/data/digitised-collections/edinburgh-ladies-debating-society/">link</a>. As other NLS data collections, they are originally provided using two XMLs schemas: METS for descriptive, structural, technical and administrative metadata (Title, Author, Publisher, etc); and ALTO for encoding the OCR text of a page.</p> <p>In this work, we have extracted the information from METS and ALTO XMLS using <a href="https://github.com/francesNLP/defoe">defoe</a> tool and developed a <a href="https://github.com/francesNLP/defoe/blob/master/defoe/nls/queries/write_metadata_pages_yml.py">new information extraction defoe query</a> , and created a new Knowlege Graph called LadiesDebating-KG. The LadiesDebating-KG uses the <a href="https://francesnlp.github.io/NLS-ontology/doc/index-en.html">NLS Ontology </a>to represent the information extracted. Furthermore, during the information extraction phase, we have employed several techniques to mitigate two common OCR errors: long-S and the line-break hyphenation.</p> <p>The LadiesDebating-KG contains 38,279 RDF triples. It has information from 2 series and 16 volumes: <strong>'The attempt' </strong>serie has 10 volumes and <strong>'The Ladies' </strong>serie<strong> </strong>has 6 volumes . Each serie has an Editor, mmsid, Shelf-Locator, publication year, etc. A Volume has several Pages, with text in them. The data model of the LadiesDebating-KG can be found <a href="https://francesnlp.github.io/NLS-ontology/doc/dataModel.png">here</a>.</p> <pre> </pre>
Disentangling Web Search on Debated Topics - User Study Data
<p>Data of an exploratory, open-ended user study (N = 255) to advance knowledge and uncover relations between the different facets of web search on debated topics. We explored the relations between factors inherent to the searcher and search system (user characteristics, exposure bias), search intercations (confirmation bias, position bias, search effort), and post-search epistemic states (attitude change, knowledge gain). This data set contains the following variables for each of the 255 participants: SERP ranking bias, prior knowledge, attitude strength, receptiveness to opposing views, attitude-confirming clicks, click rank deviation, number of clicks, time on SERP, hover depth, attitude change, knowledge gain.</p>
Data from the MA Thesis "Does She Talk Differently?: Exploring Implications of Gender in US Presidential and Vice Presidential Debates" and Coded Transcriptions of the 7 Analyzed Debates
<p>The raw data collected for the master's thesis "Does She Talk Differently?: Exploring Implications of Gender in US Presidential and Vice Presidential Debates", the tables and graphs created based on the data as well as the transcriptions for the seven debates analyzed for the research paper can be found in the files.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.