Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3 results for “bundestag”

Learn how ShareScore rates datasets ↗
zenodo48/100

Corpus der Plenarprotokolle des Deutschen Bundestages (CPP-BT)

<h3>&Uuml;berblick</h3> <p>Das <strong>Corpus der Plenarprotokolle des Deutschen Bundestages (CPP-BT)</strong> ist einer der gr&ouml;&szlig;ten, frei verf&uuml;gbaren Datens&auml;tze von Plenarprotokollen des Deutschen Bundestages. Er ist eine Zusammenstellung aller Plenarprotokolle von der 1. Wahlperiode bis zur aktuellsten 21. Wahlperiode die im XML-Format auf dem <a href="https://www.bundestag.de/services/opendata">Open Data Portal</a> des Deutschen Bundestages und dem <a href="https://dip.bundestag.de/">Dokumentations- und Informationssystem f&uuml;r parlamentarische Materialien (DIP)</a> bis zum jeweiligen Stichtag ver&ouml;ffentlicht waren.</p> <p><em>Bitte beachten Sie das beiliegende Codebook!</em> Es enth&auml;lt wichtige Informationen zur korrekten Nutzung des Datensatzes. Es hilft auch bei der Entscheidung, welche Variante f&uuml;r Sie am besten geeignet ist. In der Regel empfehle ich f&uuml;r quantitative Forschung die CSV-Dateien und f&uuml;r traditionelle Forschung die TXT-Sammlung. Parquet-Dateien sind f&uuml;r Big Data-Anwendungen verf&uuml;gbar.</p> <p>Der CPP-BT ist der <strong>Zwillings-Korpus</strong> des<strong> </strong><a href="https://doi.org/10.5281/zenodo.4643065"><strong>Corpus der Drucksachen des Deutschen Bundestages (CDRS-BT)</strong></a><strong>.&nbsp;</strong> Durch die Verbindung beider Korpora k&ouml;nnen Sie Plenarprotokolle und Drucksachen &mdash; und damit alle Vorg&auml;nge des Bundestages &mdash; in einheitlichen Analysen untersuchen.</p> <p>&nbsp;</p> <h3>Aktualisierung</h3> <p>Dieser Datensatz wird mehrmals pro Wahlperiode aktualisiert. Benachrichtigungen &uuml;ber neue und aktualisierte Datens&auml;tze ver&ouml;ffentliche ich immer zeitnah auf Mastodon unter <a href="https://fediscience.org/@seanfobbe">@seanfobbe@fediscience.org</a></p> <p>&nbsp;</p> <h3>NEU in Version 2025-05-24</h3> <ul> <li>Vollst&auml;ndige Aktualisierung der Daten (bis einschlie&szlig;lich aktuellste Wahlperiode)</li> <li>Neukonzeptionierung des Datensatzes als deklarative {targets} Pipeline</li> <li><strong>Wichtige &Auml;nderung: </strong>Variable "nummer_original" zu "protokoll_nr" umbenannt</li> <li><strong>Wichtige &Auml;nderung:</strong> Variable "datum" zu "sitzung_datum" umbenannt</li> <li>Neues Feature: Alle Einzelreden des Bundestages in tabellarischem Format mit vielen neuen Metadaten verf&uuml;gbar (ab 18. Wahlperiode)</li> <li>Neues Feature: Datensatz im Parquet-Format verf&uuml;gbar</li> <li>Neues Feature: Zus&auml;tzlicher Bericht zur Qualit&auml;tskontrolle</li> <li>Inhaltiche Erweiterung und Verbesserung der TXT-Variante</li> <li>Viele zus&auml;tzliche Tests zur Qualit&auml;tspr&uuml;fung</li> <li>Pipeline ruft automatisch die tagesaktuell neuesten Bundestagsprotokolle ab (API Key notwendig)</li> <li>Pipeline speichert viele Checkpoints und kann jederzeit unterbrochen und fortgesetzt werden</li> <li>Delta Updates m&ouml;glich</li> <li>Grundlegende &Uuml;berarbeitung des Codebooks</li> </ul> <p>&nbsp;</p> <h3>Features</h3> <ul> <li>Insgesamt bis zu 35 Variablen in der CSV-Variante</li> <li>Plenarprotokolle von der 1. Wahlperiode bis zur neuesten Wahlperiode am Stichtag</li> <li>Aufteilung in Einzelreden&nbsp;u.a. mit ID, Name, Fraktion und Amt der Redner:in (ab 18. Wahlperiode)</li> <li>Aufteilung in Protokollbestandteile: Inhaltsverzeichnis, Sitzungsverlauf, Anlagen, Rednerliste (ab 18. Wahlperiode)</li> <li>Fortlaufende Aktualisierung (Datensatz kann zus&auml;tzlich via Pipeline t&auml;glich aktualisiert werden)</li> <li>Urheberrechtsfreiheit</li> <li>Offene und plattformunabh&auml;ngige Formate (PDF, TXT, CSV, XML, Parquet)</li> <li>Linguistische Kennzahlen</li> <li>Umfangreiches Codebook</li> <li>Compilation Report, um den Erstellungs-Prozess zu erl&auml;utern</li> <li>Dutzende Diagramme und Tabellen f&uuml;r alle Zwecke (im ZIP-Archiv 'ANALYSE')</li> <li>Diagramme liegen jeweils in einem f&uuml;r den Druck (PDF) und das Web (PNG) optimierten Format vor</li> <li>Tabellen sind im CSV-Format bereitgestellt und sind damit sowohl f&uuml;r Menschen als auch f&uuml;r Maschinen gut lesbar</li> <li>Kryptographische Signaturen</li> <li><a href="https://doi.org/10.5281/zenodo.4542661">Ver&ouml;ffentlichung des Source Codes</a></li> </ul> <h3>&nbsp;</h3> <h3>Eckdaten</h3> <p><em>Stichtag:</em> 24. Mai 2025</p> <p><em>Inhaltlicher Umfang</em>: 4566 Plenarprotokolle / ~362 Millionen Tokens</p> <p><em>Zeitlicher Umfang:</em> 1949 bis 2025</p> <p><em>Wahlperioden:</em> 1. bis 21. Wahlperiode</p> <p><em>Formate:</em><strong> </strong>CSV, TXT, XML und Parquet</p> <h3>&nbsp;</h3> <h3>Source Code und Compilation Report</h3> <p>Der gesamte Erstellungs-Prozess ist vollautomatisiert und detailliert dokumentiert. Mit jeder Kompilierung des vollst&auml;ndigen Datensatzes wird auch ein umfangreicher Compilation Report in einem attraktiv designten PDF-Format erstellt (&auml;hnlich dem Codebook). Zudem werden Qualit&auml;tskontrollen auf Vollst&auml;ndigkeit und Plausibilit&auml;t durchgef&uuml;hrt und in einem separaten Bericht dokumentiert.</p> <p>Der Compilation Report enth&auml;lt den Source Code f&uuml;r die Daten-Pipeline, dokumentiert relevante Rechenergebnisse, gibt sekundengenaue Zeitstempel an und ist mit einem klickbaren Inhaltsverzeichnis versehen. Wenn Sie sich f&uuml;r Details des Erstellungs-Prozesses interessieren, lesen Sie diesen bitte zuerst.</p> <p>Der vollst&auml;ndige <em>Source Code,</em> der <em>Compilation Report</em> und die <em>Robustness Checks</em> sind <em>&ouml;ffentlich einsehbar und dauerhaft erreichbar</em> im wissenschaftlichen Archiv des CERN unter diesem Link hinterlegt:&nbsp;<a href="https://doi.org/10.5281/zenodo.4542661">https://doi.org/10.5281/zenodo.4542661</a></p> <h3>&nbsp;</h3> <h3>Kryptographische Signaturen</h3> <p>Die Integrit&auml;t und Echtheit der einzelnen Archive des Datensatzes sind durch eine <em>Zwei-Phasen-Signatur</em> sichergestellt.</p> <p>In <em>Phase I</em> werden w&auml;hrend der Kompilierung f&uuml;r jedes ZIP-Archiv, das Codebook und die Robustness Checks Hash-Werte in zwei verschiedenen Verfahren (SHA2-256 und SHA3-512) berechnet und in einer CSV-Datei dokumentiert.</p> <p>In <em>Phase II</em> werden diese CSV-Datei und der Compilation Report mit meinem pers&ouml;nlichen geheimen GPG-Schl&uuml;ssel signiert. Dieses Verfahren stellt sicher, dass die Kompilierung von jedermann durchgef&uuml;hrt werden kann, insbesondere im Rahmen von Replikationen, die pers&ouml;nliche Gew&auml;hr f&uuml;r Ergebnisse aber dennoch vorhanden ist.</p> <p>Die w&auml;hrend der Kompilierung des Datensatzes erstellte CSV-Datei mit den Hash-Pr&uuml;fsummen ist mit meiner <em>pers&ouml;nlichen GPG-Signatur</em> versehen. Der mit dieser Version korrespondierende Public Key ist sowohl mit dem Datensatz als auch mit dem Source Code hinterlegt. Er hat folgende Kenndaten:</p> <p><em>Name:</em> Sean Fobbe (fobbe-data@posteo.de)</p> <p><em>Fingerabdruck:</em> FE6F B888 F0E5 656C 1D25 3B9A 50C4 1384 F44A 4E42</p> <h3>&nbsp;</h3> <h3>Kein Urheberrecht: Public Domain</h3> <p>An den Plenarprotokollen besteht gem. &sect; 5 Abs. 2 UrhG <em>kein </em>Urheberrecht, da sie amtliche Werke sind. &sect; 5 UrhG ist auf amtliche Datenbanken analog anzuwenden (BGH, Beschluss vom 28.09.2006 - I ZR 261/03, "S&auml;chsischer Ausschreibungsdienst"). Alle eigenen Beitr&auml;ge (z.B. durch Zusammenstellung und Anpassung der Metadaten) und damit den gesamten Datensatz stelle ich gem&auml;&szlig; einer <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">CC0 1.0 Universal Public Domain License</a> vollst&auml;ndig urheberrechtsfrei.</p> <p>&nbsp;</p> <h3>Disclaimer</h3> <p>Dieser Datensatz ist eine private wissenschaftliche Initiative und steht in keiner Verbindung zum Deutschen Bundestag oder anderen amtlichen Stellen der Bundesrepublik Deutschland.</p> <p>&nbsp;</p> <h3>Alternativen</h3> <ul> <li>Blaette, A., &amp; Leonhardt, C. (2024). GermaParl Corpus of Plenary Protocols (v2.1.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.12794676">https://doi.org/10.5281/zenodo.12794676</a></li> <li>Richter, F., Koch, P., Franke, O., Kraus, J., Kuruc, F., Thiem, A., H&ouml;gerl, J., Heine, S., &amp; Sch&ouml;ps, K. (2020). Open Discourse (Version V4) [dataset]. Harvard Dataverse. <a href="https://doi.org/doi:10.7910/DVN/FIKIBO">https://doi.org/doi:10.7910/DVN/FIKIBO</a></li> <li>Rauh, Christian; Schwalbach, 2020, "The ParlSpeech V2 data set: Full-text corpora of 6.3 million parliamentary speeches in the key legislative chambers of nine representative democracies", <a href="https://doi.org/10.7910/DVN/L4OAKN">https://doi.org/10.7910/DVN/L4OAKN</a>, Harvard Dataverse, V1</li> <li>Open Knowledge Foundation, "Offenes Parlament", <a href="https://offenesparlament.de/daten/">https://offenesparlament.de/daten/</a></li> </ul> <p>&nbsp;</p> <h3>Weitere Open Access Ver&ouml;ffentlichungen (Fobbe)</h3> <p>Website<em> </em>&mdash;<em> </em><a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p> <p>Open Data&nbsp; &mdash;&nbsp; <a href="https://zenodo.org/communities/sean-fobbe-data/">zenodo.org/communities/sean-fobbe-data/</a></p> <p>Source Code&nbsp; &mdash;&nbsp; <a href="https://zenodo.org/communities/sean-fobbe-code/">zenodo.org/communities/sean-fobbe-code/</a></p> <p>Volltexte regul&auml;rer Publikationen&nbsp; &mdash;&nbsp; <a href="https://zenodo.org/communities/sean-fobbe-publications/">zenodo.org/communities/sean-fobbe-publications/</a></p> <p>&nbsp;</p> <h3>Kontakt</h3> <p>Fehler gefunden? Anregungen? Melden Sie diese entweder im Issue Tracker auf Codeberg oder kontaktieren Sie mich &uuml;ber <a href="https://www.seanfobbe.de">www.seanfobbe.de</a></p>

opencc-zeroFeb 2021View details →
zenodo44/100

Query auto-completions for German politicians of the 18th Bundestag

<p><strong>bundestag.csv</strong> - UTF-8 encoded comma separated text file</p> <p>This dataset contains the members of the 18th German Bundestag in the constitution of late 2016.</p> <p>&lt;Name&gt;: name of the politician</p> <p>&lt;Born&gt;: birthday</p> <p>&lt;Party&gt;: party membership of the politician</p> <p>&lt;Bundesland&gt; state of the politician</p> <p>&lt;Gender&gt; gender of the politician</p> <p>&lt;Age&gt; age of the politician (as of 2017)</p> <p>&lt;Cluster 3&gt; number of unique auto-completions assigned to topic: &quot;location information&quot;</p> <p>&lt;Cluster 2&gt; number of unique auto-completions assigned to topic: &quot;personal and emotional&quot;</p> <p>&lt;Cluster 1&gt; number of unique auto-completions assigned to topic: &quot;politics and economics&quot;</p> <p>&lt;Total&gt; total number of unique auto-completions</p> <p>&nbsp;</p> <p><strong>terms.csv </strong>- UTF-8 encoded comma separated text file</p> <p>This dataset contains the unordered and pooled auto-completions for the German politicians from Bing search (http://api.bing.net/osjson.aspx), from Duck-Duck-Go (https://duckduckgo.com/ac/) and from Google search (http://clients1.google.de/complete/search). The data was crawled on (mostly) two times per day from 2017/02/03 to 2017/06/19. German language settings were used for Google and Bing, English language setting was used for Duck-Duck-Go. The API requests were sent with an IP address from Cologne, Germany.&nbsp;</p> <p>&lt;source&gt;: google, bing or ddg</p> <p>&lt;queryterm&gt;: the query term, matches the name of the politican in the file &lt;bundestag.csv&gt;</p> <p>&lt;suggestterm&gt;: the suggested query auto-completion</p>

opencc-by-4.0Sep 2019View details →
zenodo40/100

The #BTW17 Twitter Dataset - Recorded Tweets of the Federal Election Campaigns of 2017 for the 19th German Bundestag

<p>The German Bundestag elections are the most important democratic elections of Germany. This dataset comprises Twitter interactions related with German politicians of the most important political parties over several months in the (pre-)phase of the German election campaigns in 2017. The Twitter accounts of 364&nbsp;politicians (that is approximately half of the German parliament, the German Bundestag) were followed for almost half a year. The collected data comprise of about 10 GB of Twitter raw data generated by more than 120.000 active Twitter users generating more than 1.200.000 tweets during the pre- and hot-phase of the election campaigns for the 19th German Bundestag.&nbsp;<br> The dataset can be used to study how political parties, their followers and supporters make use of social media channels like Twitter in the context of&nbsp;political election campaigns and what kind of content is shared.</p> <p>The following files contain relevant context information:</p> <ul> <li><strong>crawled-pages.json</strong>&nbsp;contains the URLs of the official party faction websites of the 18th German Bundestag that were crawled to identify the Twitter screennames of German politicians of all Bundestag factions. Because the <em>Alternative f&uuml;r Deutschland (AfD)</em> and the <em>Freie Demokratische Partei&nbsp;(FDP)</em> were not part of the 18th German Bundestag (but it was likely that they will enter the 19th German Bundestag) other official websites were selected to crawl for relevant and representative politicians for these both parties (in case of the <em>AfD</em> this was the website of the directorate of the <em>AfD</em> federal party and the list of members of the European Parliament, in case of the <em>FDP</em> this was the website of the executive committee of the <em>FDP</em> federal party of Germany).</li> <li><strong>followed-accounts.json</strong>&nbsp;contains the (manually checked and edited) crawling result of 327&nbsp;Twitter screennames of &nbsp;politicians that have been observed via the Twitter streaming API to collect this dataset.</li> </ul>

opencc-by-4.0Sep 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record