Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3
datasets available to search
ShareScore release 0.9.0
Dataset results
3 results for “Bundestag”
Corpus der Plenarprotokolle des Deutschen Bundestages (CPP-BT)
<h3>Überblick</h3> <p>Das <strong>Corpus der Plenarprotokolle des Deutschen Bundestages (CPP-BT)</strong> ist einer der größten, frei verfügbaren Datensätze von Plenarprotokollen des Deutschen Bundestages. Er ist eine Zusammenstellung aller Plenarprotokolle von der 1. Wahlperiode bis zur aktuellsten 21. Wahlperiode die im XML-Format auf dem <a href="https://www.bundestag.de/services/opendata">Open Data Portal</a> des Deutschen Bundestages und dem <a href="https://dip.bundestag.de/">Dokumentations- und Informationssystem für parlamentarische Materialien (DIP)</a> bis zum jeweiligen Stichtag veröffentlicht waren.</p> <p><em>Bitte beachten Sie das beiliegende Codebook!</em> Es enthält wichtige Informationen zur korrekten Nutzung des Datensatzes. Es hilft auch bei der Entscheidung, welche Variante für Sie am besten geeignet ist. In der Regel empfehle ich für quantitative Forschung die CSV-Dateien und für traditionelle Forschung die TXT-Sammlung. Parquet-Dateien sind für Big Data-Anwendungen verfügbar.</p> <p>Der CPP-BT ist der <strong>Zwillings-Korpus</strong> des<strong> </strong><a href="https://doi.org/10.5281/zenodo.4643065"><strong>Corpus der Drucksachen des Deutschen Bundestages (CDRS-BT)</strong></a><strong>. </strong> Durch die Verbindung beider Korpora können Sie Plenarprotokolle und Drucksachen — und damit alle Vorgänge des Bundestages — in einheitlichen Analysen untersuchen.</p> <p> </p> <h3>Aktualisierung</h3> <p>Dieser Datensatz wird mehrmals pro Wahlperiode aktualisiert. Benachrichtigungen über neue und aktualisierte Datensätze veröffentliche ich immer zeitnah auf Mastodon unter <a href="https://fediscience.org/@seanfobbe">@seanfobbe@fediscience.org</a></p> <p> </p> <h3>NEU in Version 2025-05-24</h3> <ul> <li>Vollständige Aktualisierung der Daten (bis einschließlich aktuellste Wahlperiode)</li> <li>Neukonzeptionierung des Datensatzes als deklarative {targets} Pipeline</li> <li><strong>Wichtige Änderung: </strong>Variable "nummer_original" zu "protokoll_nr" umbenannt</li> <li><strong>Wichtige Änderung:</strong> Variable "datum" zu "sitzung_datum" umbenannt</li> <li>Neues Feature: Alle Einzelreden des Bundestages in tabellarischem Format mit vielen neuen Metadaten verfügbar (ab 18. Wahlperiode)</li> <li>Neues Feature: Datensatz im Parquet-Format verfügbar</li> <li>Neues Feature: Zusätzlicher Bericht zur Qualitätskontrolle</li> <li>Inhaltiche Erweiterung und Verbesserung der TXT-Variante</li> <li>Viele zusätzliche Tests zur Qualitätsprüfung</li> <li>Pipeline ruft automatisch die tagesaktuell neuesten Bundestagsprotokolle ab (API Key notwendig)</li> <li>Pipeline speichert viele Checkpoints und kann jederzeit unterbrochen und fortgesetzt werden</li> <li>Delta Updates möglich</li> <li>Grundlegende Überarbeitung des Codebooks</li> </ul> <p> </p> <h3>Features</h3> <ul> <li>Insgesamt bis zu 35 Variablen in der CSV-Variante</li> <li>Plenarprotokolle von der 1. Wahlperiode bis zur neuesten Wahlperiode am Stichtag</li> <li>Aufteilung in Einzelreden u.a. mit ID, Name, Fraktion und Amt der Redner:in (ab 18. Wahlperiode)</li> <li>Aufteilung in Protokollbestandteile: Inhaltsverzeichnis, Sitzungsverlauf, Anlagen, Rednerliste (ab 18. Wahlperiode)</li> <li>Fortlaufende Aktualisierung (Datensatz kann zusätzlich via Pipeline täglich aktualisiert werden)</li> <li>Urheberrechtsfreiheit</li> <li>Offene und plattformunabhängige Formate (PDF, TXT, CSV, XML, Parquet)</li> <li>Linguistische Kennzahlen</li> <li>Umfangreiches Codebook</li> <li>Compilation Report, um den Erstellungs-Prozess zu erläutern</li> <li>Dutzende Diagramme und Tabellen für alle Zwecke (im ZIP-Archiv 'ANALYSE')</li> <li>Diagramme liegen jeweils in einem für den Druck (PDF) und das Web (PNG) optimierten Format vor</li> <li>Tabellen sind im CSV-Format bereitgestellt und sind damit sowohl für Menschen als auch für Maschinen gut lesbar</li> <li>Kryptographische Signaturen</li> <li><a href="https://doi.org/10.5281/zenodo.4542661">Veröffentlichung des Source Codes</a></li> </ul> <h3> </h3> <h3>Eckdaten</h3> <p><em>Stichtag:</em> 24. Mai 2025</p> <p><em>Inhaltlicher Umfang</em>: 4566 Plenarprotokolle / ~362 Millionen Tokens</p> <p><em>Zeitlicher Umfang:</em> 1949 bis 2025</p> <p><em>Wahlperioden:</em> 1. bis 21. Wahlperiode</p> <p><em>Formate:</em><strong> </strong>CSV, TXT, XML und Parquet</p> <h3> </h3> <h3>Source Code und Compilation Report</h3> <p>Der gesamte Erstellungs-Prozess ist vollautomatisiert und detailliert dokumentiert. Mit jeder Kompilierung des vollständigen Datensatzes wird auch ein umfangreicher Compilation Report in einem attraktiv designten PDF-Format erstellt (ähnlich dem Codebook). Zudem werden Qualitätskontrollen auf Vollständigkeit und Plausibilität durchgeführt und in einem separaten Bericht dokumentiert.</p> <p>Der Compilation Report enthält den Source Code für die Daten-Pipeline, dokumentiert relevante Rechenergebnisse, gibt sekundengenaue Zeitstempel an und ist mit einem klickbaren Inhaltsverzeichnis versehen. Wenn Sie sich für Details des Erstellungs-Prozesses interessieren, lesen Sie diesen bitte zuerst.</p> <p>Der vollständige <em>Source Code,</em> der <em>Compilation Report</em> und die <em>Robustness Checks</em> sind <em>öffentlich einsehbar und dauerhaft erreichbar</em> im wissenschaftlichen Archiv des CERN unter diesem Link hinterlegt: <a href="https://doi.org/10.5281/zenodo.4542661">https://doi.org/10.5281/zenodo.4542661</a></p> <h3> </h3> <h3>Kryptographische Signaturen</h3> <p>Die Integrität und Echtheit der einzelnen Archive des Datensatzes sind durch eine <em>Zwei-Phasen-Signatur</em> sichergestellt.</p> <p>In <em>Phase I</em> werden während der Kompilierung für jedes ZIP-Archiv, das Codebook und die Robustness Checks Hash-Werte in zwei verschiedenen Verfahren (SHA2-256 und SHA3-512) berechnet und in einer CSV-Datei dokumentiert.</p> <p>In <em>Phase II</em> werden diese CSV-Datei und der Compilation Report mit meinem persönlichen geheimen GPG-Schlüssel signiert. Dieses Verfahren stellt sicher, dass die Kompilierung von jedermann durchgeführt werden kann, insbesondere im Rahmen von Replikationen, die persönliche Gewähr für Ergebnisse aber dennoch vorhanden ist.</p> <p>Die während der Kompilierung des Datensatzes erstellte CSV-Datei mit den Hash-Prüfsummen ist mit meiner <em>persönlichen GPG-Signatur</em> versehen. Der mit dieser Version korrespondierende Public Key ist sowohl mit dem Datensatz als auch mit dem Source Code hinterlegt. Er hat folgende Kenndaten:</p> <p><em>Name:</em> Sean Fobbe (fobbe-data@posteo.de)</p> <p><em>Fingerabdruck:</em> FE6F B888 F0E5 656C 1D25 3B9A 50C4 1384 F44A 4E42</p> <h3> </h3> <h3>Kein Urheberrecht: Public Domain</h3> <p>An den Plenarprotokollen besteht gem. § 5 Abs. 2 UrhG <em>kein </em>Urheberrecht, da sie amtliche Werke sind. § 5 UrhG ist auf amtliche Datenbanken analog anzuwenden (BGH, Beschluss vom 28.09.2006 - I ZR 261/03, "Sächsischer Ausschreibungsdienst"). Alle eigenen Beiträge (z.B. durch Zusammenstellung und Anpassung der Metadaten) und damit den gesamten Datensatz stelle ich gemäß einer <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">CC0 1.0 Universal Public Domain License</a> vollständig urheberrechtsfrei.</p> <p> </p> <h3>Disclaimer</h3> <p>Dieser Datensatz ist eine private wissenschaftliche Initiative und steht in keiner Verbindung zum Deutschen Bundestag oder anderen amtlichen Stellen der Bundesrepublik Deutschland.</p> <p> </p> <h3>Alternativen</h3> <ul> <li>Blaette, A., & Leonhardt, C. (2024). GermaParl Corpus of Plenary Protocols (v2.1.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.12794676">https://doi.org/10.5281/zenodo.12794676</a></li> <li>Richter, F., Koch, P., Franke, O., Kraus, J., Kuruc, F., Thiem, A., Högerl, J., Heine, S., & Schöps, K. (2020). Open Discourse (Version V4) [dataset]. Harvard Dataverse. <a href="https://doi.org/doi:10.7910/DVN/FIKIBO">https://doi.org/doi:10.7910/DVN/FIKIBO</a></li> <li>Rauh, Christian; Schwalbach, 2020, "The ParlSpeech V2 data set: Full-text corpora of 6.3 million parliamentary speeches in the key legislative chambers of nine representative democracies", <a href="https://doi.org/10.7910/DVN/L4OAKN">https://doi.org/10.7910/DVN/L4OAKN</a>, Harvard Dataverse, V1</li> <li>Open Knowledge Foundation, "Offenes Parlament", <a href="https://offenesparlament.de/daten/">https://offenesparlament.de/daten/</a></li> </ul> <p> </p> <h3>Weitere Open Access Veröffentlichungen (Fobbe)</h3> <p>Website<em> </em>—<em> </em><a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p> <p>Open Data — <a href="https://zenodo.org/communities/sean-fobbe-data/">zenodo.org/communities/sean-fobbe-data/</a></p> <p>Source Code — <a href="https://zenodo.org/communities/sean-fobbe-code/">zenodo.org/communities/sean-fobbe-code/</a></p> <p>Volltexte regulärer Publikationen — <a href="https://zenodo.org/communities/sean-fobbe-publications/">zenodo.org/communities/sean-fobbe-publications/</a></p> <p> </p> <h3>Kontakt</h3> <p>Fehler gefunden? Anregungen? Melden Sie diese entweder im Issue Tracker auf Codeberg oder kontaktieren Sie mich über <a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p>
Query auto-completions for German politicians of the 18th Bundestag
<p><strong>bundestag.csv</strong> - UTF-8 encoded comma separated text file</p> <p>This dataset contains the members of the 18th German Bundestag in the constitution of late 2016.</p> <p><Name>: name of the politician</p> <p><Born>: birthday</p> <p><Party>: party membership of the politician</p> <p><Bundesland> state of the politician</p> <p><Gender> gender of the politician</p> <p><Age> age of the politician (as of 2017)</p> <p><Cluster 3> number of unique auto-completions assigned to topic: "location information"</p> <p><Cluster 2> number of unique auto-completions assigned to topic: "personal and emotional"</p> <p><Cluster 1> number of unique auto-completions assigned to topic: "politics and economics"</p> <p><Total> total number of unique auto-completions</p> <p> </p> <p><strong>terms.csv </strong>- UTF-8 encoded comma separated text file</p> <p>This dataset contains the unordered and pooled auto-completions for the German politicians from Bing search (http://api.bing.net/osjson.aspx), from Duck-Duck-Go (https://duckduckgo.com/ac/) and from Google search (http://clients1.google.de/complete/search). The data was crawled on (mostly) two times per day from 2017/02/03 to 2017/06/19. German language settings were used for Google and Bing, English language setting was used for Duck-Duck-Go. The API requests were sent with an IP address from Cologne, Germany. </p> <p><source>: google, bing or ddg</p> <p><queryterm>: the query term, matches the name of the politican in the file <bundestag.csv></p> <p><suggestterm>: the suggested query auto-completion</p>
The #BTW17 Twitter Dataset - Recorded Tweets of the Federal Election Campaigns of 2017 for the 19th German Bundestag
<p>The German Bundestag elections are the most important democratic elections of Germany. This dataset comprises Twitter interactions related with German politicians of the most important political parties over several months in the (pre-)phase of the German election campaigns in 2017. The Twitter accounts of 364 politicians (that is approximately half of the German parliament, the German Bundestag) were followed for almost half a year. The collected data comprise of about 10 GB of Twitter raw data generated by more than 120.000 active Twitter users generating more than 1.200.000 tweets during the pre- and hot-phase of the election campaigns for the 19th German Bundestag. <br> The dataset can be used to study how political parties, their followers and supporters make use of social media channels like Twitter in the context of political election campaigns and what kind of content is shared.</p> <p>The following files contain relevant context information:</p> <ul> <li><strong>crawled-pages.json</strong> contains the URLs of the official party faction websites of the 18th German Bundestag that were crawled to identify the Twitter screennames of German politicians of all Bundestag factions. Because the <em>Alternative für Deutschland (AfD)</em> and the <em>Freie Demokratische Partei (FDP)</em> were not part of the 18th German Bundestag (but it was likely that they will enter the 19th German Bundestag) other official websites were selected to crawl for relevant and representative politicians for these both parties (in case of the <em>AfD</em> this was the website of the directorate of the <em>AfD</em> federal party and the list of members of the European Parliament, in case of the <em>FDP</em> this was the website of the executive committee of the <em>FDP</em> federal party of Germany).</li> <li><strong>followed-accounts.json</strong> contains the (manually checked and edited) crawling result of 327 Twitter screennames of politicians that have been observed via the Twitter streaming API to collect this dataset.</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.