Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5
datasets available to search
ShareScore release 0.9.0
Dataset results
5 results for “computational linguistics”
A cross-linguistic computational approach on chance resemblances
<p>Material for the "A cross-linguistic computational approach on chance resemblances" presentation, with raw wordlists, lexeme analyses, and distribution statistics.</p>
Computational linguistics based text emotion analysis using enhanced beetle antenna search with deep learning during COVID-19 pandemic
Open the record for dataset details and reuse information.
A Greek Parliament Proceedings Dataset for Computational Linguistics and Political Analysis
<p>The dataset is a new version of the previous upload and includes the following files:</p> <p>1. <strong>dataset_versions/tell_all.csv: </strong>The initial dataset of 1,280,927 extracted speeches, before preprocessing and cleaning. The speeches extend chronologically from July 1989 up to July 2020 and were exported from 5,355 parliamentary sitting record files. The file has a total volume of 2.5 GB and includes the following columns:</p> <ul> <li>member_name: the name of the individual who spoke during a sitting.</li> <li>sitting_date: the date the sitting took place.</li> <li>parliamentary_period: the name and/or number of the parliamentary period that the speech took place in. A parliamentary period is defined as the time span between one general election and the next. A parliamentary period includes multiple parliamentary sessions.</li> <li>parliamentary_session: the name and/or number of the parliamentary session that the speech took place in. A session is defined as a time span of usually 10 months within a parliamentary period during which the parliament can convene and function as stipulated by the constitution. A session can fall into the following categories: regular, extraordinary or special. In the intervals between the sessions the parliament is in recess. A parliamentary session includes multiple parliamentary sittings.</li> <li>parliamentary_sitting: the name and/or number of the parliamentary sitting that the speech took place in. A sitting is defined as a meeting of parliament members.</li> <li>political_party: the political party of the speaker.</li> <li>government: the government in force when the speech took place.</li> <li>member_region: the electoral district the speaker belonged to.</li> <li>roles: information about the parliamentary roles and/or government position of the speaker.</li> <li>member_gender: the gender of the speaker</li> <li>speech: the speech that the individual gave during the parliamentary sitting.</li> </ul> <p>2. <strong>dataset_versions/tell_all_FILLED.csv: </strong>This file is an intermediate version of the dataset that includes improvements in the consistency and completeness of the dataset, with a total volume of 2.5 GB. Specifically, this file is produced by filling the missing names of chairmen of various parliamentary sittings of the "tell_all.csv". It includes the same columns as the "tell_all.csv" file.</p> <p>3.<strong> dataset_versions/tell_all_cleaned.csv: </strong>This version of the dataset is the result of further cleaning and preprocessing and is used for our word usage change study. It consists of 1,280,918 speech fragments of Greek parliament members in the order of the conversation that took place, with a total volume of 2.12 GB. It includes the same columns as the aforementioned versions. The preprocessing includes the replacement of all references to political parties with the symbol "@" followed by an abbreviation of the party name, using regular expressions that capture different grammatical cases and variations. It also includes the removal of accents, strings with length less than 2 characters, all punctuation except full stops, and the replacement of stopwords with "@sw".</p> <p>4. <strong>wiki_data</strong>: A folder of modern Greek female and male names and surnames and their available grammatical cases crawled from the entries of the Wiktionary Greek names category (https://en.wiktionary.org/wiki/Category:Greek_names). We produced the grammatical cases of the missing grammatical entries according to the rules of the Greek grammar and saved the files in the same folder by adding to their filenames the string "_populated.json".</p> <p>5. <strong>parl_members_activity_1989onwards_with_gender.csv</strong>: The Greek Parliament website provides a<br> <a href="https://www.hellenicparliament.gr/Vouleftes/Diatelesantes-Vouleftes-Apo-Ti-Metapolitefsi-Os-Simera/">list</a> of all the elected members of parliament since the fall of the military junta in Greece, in 1974. We collected and cleaned the data, added the gender and kept the elected members from 1989 onwards, matching the available parliament proceeding records. This dataset includes the full names of the members, the date range of their service, the political party they served, the electoral district they belonged to and their gender.</p> <p>6. <strong>formatted_roles_gov_members_data.csv</strong>: As government members we refer to individuals in ministerial or other government posts, regardless of whether they were elected in the parliament. This information is available in the website of the <a href="https://gslegal.gov.gr/?page_id=776&sort=time">Secretariat General for Legal and Parliamentary Affairs</a>. The government members dataset includes the full names of the official individuals, the name of the role they were given, the date range of their service at each specific role and their gender.</p> <p>7. <strong>governments_1989onwards.csv</strong>: A dataset of government information including the names of governments since 1989, their start and end dates, and a URL that points to the respective official government web page of each past government. The data is crawled from the website of the <a href="https://gslegal.gov.gr/?page_id=776&sort=time">Secretariat General for Legal and Parliamentary Affairs</a>.</p> <p>8. <strong>extra_roles_manually_collected.csv</strong>: A dataset with manually collected information from Wikipedia about additional government or parliament posts such as Chairman of the Parliament, party leaders, opposition leaders and other information.</p> <p>9. <strong>all_members_activity.csv</strong>: A dataset of all the information of the aforementioned files 3,4,5,6 merged. Each row of the file includes the full name of the individual, the start and end date of their term of office, the political party and electoral district they belonged to, their gender, the parliamentary and/or government positions that they held along with start and end dates, and the name of the government that was in power during their term of office. An individual can change political parties or become an independent member of the parliament during a parliamentary period, thus having more than one entries/rows in the file.</p> <p>10. <strong>freqs_for_semantic_shift_cleaned_data_decade1990.csv & freqs_for_semantic_shift_cleaned_data_decade2010.csv</strong>: Files of frequencies of words in the corpora of the decades 1990-1999 and 2010-2019.</p> <p>11. <strong>compass_top100.csv:</strong> Top 100 most changed words between the decades 1990-1999 and 2010-2019, as computed with the use of the Compass tool by V. D. Carlo et. al. [1].</p> <p>12. <strong>compass_fc_top100.csv</strong>: Top 100 most changed words between the decades 1990-1999 and 2010-2019, as computed with the use of the Compass tool [1] in combination with the frequency cut-offs of the Gonen et. al. approach [3]. For the frequency cut-offs, the files in bullet 8 are used.</p> <p>13. <strong> procrustes_top100.csv</strong>: Top 100 most changed words between the decades 1990-1999 and 2010-2019, as computed with the use of the Orthogonal Procrustes approach of Hamilton et. al. [2].</p> <p>14. <strong>nn_top100.csv</strong>: Top 100 most changed words between the decades 1990-1999 and 2010-2019, as computed with the use of the Gonen et. al. approach [3].</p> <p>15. <strong>second_order_top100.csv</strong>: Top 100 most changed words between the decades 1990-1999 and 2010-2019, as computed with the use of the Second-Order Similarity approach by Hamilton et. al. [4].</p> <p>16. <strong>top100_minfreq50.xls</strong>: An .xls file for convinient viewing of the top 100 most changed words per approach with minimum frequency of 50 occurrences, produced by merging the aforementioned files 11, 12, 13, 14, 15 and 16.</p> <p>17. <strong>freqs_for_semantic_shift_cleaned_data_period1997_2007.csv & freqs_for_semantic_shift_cleaned_data_period2008_2018.csv</strong>: Files of frequencies of words in the corpora of the decades before (1997_2007) and during (2008_2018) the Greek economic crisis.</p> <p>18. <strong>semantic_shifts_dichotomy_crisis_compass_1997_2007_2008_2018_atleast50.csv</strong>: A file with the top 100 most changed words between between the decades before (1997-2007) and during (2008-2018) the Greek economic crisis. The computations are implemented with the use of the Compass tool.</p> <p>19. <strong>selected_topics_shift_per_period_compass.csv</strong>: The usage change of selected topics/words of generic political interest between pairs of consecutive parliamentary periods. The computations are implemented with the use of the Compass tool.</p> <p>20. <strong>semantic_shifts_party_embeddings_per_period_merged_compass.csv</strong>: The usage change of selected political party names that have played an important role in recent political history, namely New Democracy (ND), the Panhellenic Socialist Movement (PASOK), the Coalition of the Radical Left - Progressive Alliance (SYRIZA), the Communist Party of Greece (KKE), the Coalition of the Left, of Movements and Ecology (SYN) and Golden Dawn (GD).</p> <p>-------------</p> <p><strong><em>Citations:</em></strong></p> <p>[1] Valerio Di Carlo, Federico Bianchi, and Matteo Palmonari. Training Temporal Word Em- beddings with a Compass. In <em>Proceedings of the Thirty–Third AAAI Conference on Artificial Intelligence</em>, AAAI’19, pages 6326–6334, 2019. doi: 10.1609/aaai.v33i01.33016326.</p> <p>[2] William L. Hamilton, Jure Leskovec, and Dan Jurafsky. Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change. In <em>Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</em>, ACL 2016, pages 1489– 1501, Berlin, Germany, August 2016. Association for Computational Linguistics. doi: 10. 18653/v1/P16-1141. URL https://www.aclweb.org/anthology/P16-1141.</p> <p>[3] Hila Gonen, Ganesh Jawahar, Djamé Seddah, and Yoav Goldberg. Simple, Interpretable and Stable Method for Detecting Words with Usage Change across Corpora. In <em>Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</em>, ACL 2020, pages 538– 555, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl- main.51. URL https://aclanthology.org/2020.acl-main.51.</p> <p>[4] William L. Hamilton, Jure Leskovec, and Dan Jurafsky. Cultural Shift or Linguistic Drift? Comparing Two Computational Measures of Semantic Change. In <em>Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing</em>, EMNLP 2016, pages 2116–2121, Austin, Texas, November 2016. Association for Computational Linguistics. doi: 10.18653/v1/D16-1229. URL https://www.aclweb.org/anthology/D16-1229.</p> <p>-------------</p> <p><strong><em>Acknowledgments:</em></strong></p> <p>This work was supported by the European Union’s Horizon 2020 research and innovation program ``FASTEN'' under grant agreement No 825328 and the non profit data journalism organization iMEdD.org.</p>
Data from: The appropriateness of language found in research consent form templates: a computational linguistic analysis
Background: To facilitate informed consent, consent forms should use language below the grade eight level. Research Ethics Boards (REBs) provide consent form templates to facilitate this goal. Templates with inappropriate language could promote consent forms that participants find difficult to understand. However, a linguistic analysis of templates is lacking. Methods: We reviewed the websites of 124 REBs for their templates. These included English language medical school REBs in Australia/New Zealand (n=23), Canada (n=14), South Africa (n=8), the United Kingdom (n=34), and a geographically-stratified sample from the United States (n=45). Template language was analyzed using Coh-Metrix linguistic software (v.3.0, Memphis, USA). We evaluated the proportion of REBs with five key linguistic outcomes at or below grade eight. Additionally, we compared quantitative readability to the REBs' own readability standards. To determine if the template's country of origin or the presence of a local REB readability standard influenced the linguistic variables, we used a MANOVA model. Results: Of the REBs who provided templates, 0/94 (0%, 95% CI=0-3.9%) provided templates with all linguistic variables at or below the grade eight level. Relaxing the standard to a grade 12 level did not increase this proportion. Further, only 2/22 (9.1%, 95% CI= 2.5-27.8) REBs met their own readability standard. The country of origin (DF= 20, 177.5, F=1.97, p=0.01), but not the presence of an REB-specific standard (DF=5, 84, F=0.73, p=0.60), influenced the linguistic variables. Conclusions: Inappropriate language in templates is an international problem. Templates use words that are long, abstract, and unfamiliar. This could undermine the validity of participant informed consent. REBs should set a policy of screening templates with linguistic software.
Data from: The appropriateness of language found in research consent form templates: a computational linguistic analysis
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.