Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
58
datasets available to search
ShareScore release 0.9.0
Dataset results
58 results for “keywords”
Co-occurrences of trending keywords in popular tech media (01.2016-04.2021)
<p>Sources with weights</p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5 %</li> <li>IEEE Spectrum 5 %</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p>Methodology</p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues have been selected (e.g. 'metoo', 'gdpr')</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul>
Keyword frequencies in popular tech media (01.2016-04.2021)
<p>Sources with weights</p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5 %</li> <li>IEEE Spectrum 5 %</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p>Methodology</p> <ul> <li>Frequency of appearances for all unigrams and bigrams in the texts</li> <li>Frequency: number of appearances of every term divided by the number of all terms (for every month and source) </li> <li>Several media sources: a representative index is calculated with weighted average (weights as above)</li> <li>Average monthly change in the analised term's frequency is calculated by OLS regressions</li> <li>The dependent variable of the estimation is the frequency index, while the number of months since the beginning of the analysed period (January 2016) is the independent variable</li> <li>The regression coefficient (referred to as coef) shows by how much on average the analysed expression’s frequency changed with every observed month (marginal change of the frequency), revealing which keywords had the biggest monthly growth</li> </ul> <p>Columns</p> <p>freq_months (e.g. freq_2019-04): the average frequency of the term</p> <p>coef: the regression coefficient</p> <p>coef_norm: the regression coefficient divided by the mean frequency of the keyword</p>
Co-occurrences of trending keywords in popular tech media during the COVID-19 pandemic (01.2020-06.2020)
<p>Sources: </p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe </li> <li>IEEE Spectrum </li> <li>Techforge </li> <li>Fastcompany </li> <li>The Guardian (Tech) </li> <li>Arstechnica </li> <li>Reuters </li> <li>Gizmodo </li> <li>ZDNet </li> <li>The Register </li> <li>The Verge </li> <li>TechCrunch </li> </ul> <p>Methodology</p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues and technologies have been selected (e.g. covid19)</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul>
Keyword frequencies in popular tech media during the COVID-19 pandemic (01.2020-06.2020)
<p>Sources: </p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe </li> <li>IEEE Spectrum </li> <li>Techforge </li> <li>Fastcompany </li> <li>The Guardian (Tech) </li> <li>Arstechnica </li> <li>Reuters </li> <li>Gizmodo </li> <li>ZDNet </li> <li>The Register </li> <li>The Verge </li> <li>TechCrunch </li> </ul> <p>Methodology is modified relative to the regular trend analysis due to the short period of analysis (weekly freqiencies)</p> <ul> <li>Frequency of appearances for all unigrams and bigrams in the texts</li> <li>Frequency: number of appearances of every term divided by the number of all terms (for every week) </li> <li>Several media sources: all articles are treated equally</li> <li>Average monthly change in the analised term's frequency is calculated by OLS regressions</li> <li>The dependent variable of the estimation is the frequency index, while the number of weeks since the beginning of the analysed period (January 2020) is the independent variable</li> <li>The regression coefficient (referred to as coef) shows by how much on average the analysed expression’s frequency changed with every observed week (marginal change of the frequency), revealing which keywords had the biggest weekly growth</li> </ul> <p>Columns</p> <p>freq_2020_weeks (e.g. freq_2020_ww0): the average frequency of the term</p> <p>coef: the regression coefficient</p> <p>coef_norm: the regression coefficient divided by the mean frequency of the keyword</p>
Corpus and list of keywords from Improving sustainable crop protection using population genetics concepts
<p>Corpus extracted in April 2021 from the ISI Web of Science portal (https://www.webofscience.com) with the following request: ‘Plant AND Resistan* AND Durab*’. A first corpus of 2522 articles was built considering all publication years for this extraction. This collection was then refined by categories to remove articles outwith the scope of our search (e.g. related to durable resistant materials for constructions). We also kept only articles cited at least once. The final corpus was composed of 1783 articles from 1979 to 2021:</p> <ul> <li>CORPUS_plant_resistance_durability.zip</li> </ul> <p>List of keywords used for the network presented in the article:</p> <ul> <li>keywords_list.csv</li> </ul>
Co-occurrences of trending keywords in popular tech media
<p><strong>Co-occurrences of trending keywords in the tech media (01.2016-03.2019)</strong></p> <p><strong>Sources</strong></p> <ul> <li>Gigaom 0.5%</li> <li>Euractiv 0.9%</li> <li>The Conversation 1.3%</li> <li>Politico Europe 1.3%</li> <li>IEEE Spectrum 1.8%</li> <li>Techforge 4.3%</li> <li>Fastcompany 4.5%</li> <li>The Guardian (Tech) 9.2%</li> <li>Arstechnica 10.0%</li> <li>Reuters 11%</li> <li>Gizmodo 17.5%</li> <li>ZDNet 18.3%</li> <li>The Register 19.5%</li> </ul> <p><strong>Methodology</strong></p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues have been selected (e.g. 'metoo', 'gdpr')</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul> <p><strong>Files</strong></p> <p>unigram-unigram co-occurrences: cooc11weighted.csv</p> <p>unigram-bigram co-occurrences: cooc12weighted.csv</p> <p>bigram-unigram co-occurrences: cooc21weighted.csv</p> <p>bigram-bigram co-occurrences: cooc22weighted.csv<br> </p> <p> </p>
Keyword frequency in popular tech media
<p><strong>Keywords trending in the tech media (01.2016-03.2019)</strong></p> <p><strong>Sources</strong></p> <ul> <li>Gigaom 0.5%</li> <li>Euractiv 0.9%</li> <li>The Conversation 1.3%</li> <li>Politico Europe 1.3%</li> <li>IEEE Spectrum 1.8%</li> <li>Techforge 4.3%</li> <li>Fastcompany 4.5%</li> <li>The Guardian (Tech) 9.2%</li> <li>Arstechnica 10.0%</li> <li>Reuters 11%</li> <li>Gizmodo 17.5%</li> <li>ZDNet 18.3%</li> <li>The Register 19.5%</li> </ul> <p><strong>Methodology</strong></p> <ul> <li>Frequency of appearances for all unigrams and bigrams in the texts</li> <li>Frequency: number of appearances of every term divided by the number of published articles (for every month and source)</li> <li>This measure reveals how many times an expression has been mentioned on average per article</li> <li>Several media sources: a representative index is calculated with weighted average</li> <li>Average monthly change in the analised term's frequency is calculated by OLS regressions</li> <li>The dependent variable of the estimation is the frequency index, while the number of months since the beginning of the analysed period (January 2016) is the independent variable</li> <li>The regression coefficient (referred to as coef) shows by how much on average the analysed expression’s frequency changed with every observed month (marginal change of the frequency), revealing which keywords had the biggest monthly growth</li> </ul> <p><strong>Files</strong></p> <ul> <li>unigrams: coefs_1weighted_site.csv</li> <li>bigrams: coefs_2weighted_site.csv</li> </ul> <p> </p>
Unpacking the concept of "educators' data literacy in Higher Education" - Systematic Review of the literature and Keyword Map
<p>As algorithmic decision-making and data collection become pervasive within higher education, how can educators make sense of the systems that shape life and learning in the 21st century? Through a systematic review of the literature, the paper investigates the gaps in the literature, which prevent the formulation of potential pathways and principles on which educators’ data literacy can - and should - be developed and fostered. The analysis of 137 papers through the methods of classification under relevant categories, and key words mapping, showed that there is little attention on HE teachers, and most approaches to educators’ data literacy address management and technical abilities for data processing, with less concern on critical, ethical and personal approaches to datafication in education.</p> <p>The present dataset shows the full list of articles analysed.</p> <p>The dataset, and ods file, is composed by the following sheets:</p> <ol> <li>Codebook</li> <li>List of articles extracted from SCOPUS</li> <li>List of articles extracted from WOS</li> <li>List of articles extracted from ERIC</li> <li>List of articles extracted from DOAJ</li> <li>Interrater Agreement</li> <li>PRISMA workflow</li> <li>Analysis - First Level (classification of 137 articles selected)</li> <li>Analysis - Second Level (List of articles relating faculty development)</li> <li>Supplementary tables (counting articles in relation to the categories of analysis).</li> </ol> <p>As for the Keywords' Map, a second file .csv displays the text over which basis was performed the keyword maps analysis. A .txt file shows notes relating the analysis procedures using the software VOS-Viewer <a href="http://www.vosviewer.com/">http://www.vosviewer.com/</a></p> <p> </p>
Keyword counts from US Presidential State of the Union Addresses and Presidential Budget Messages
<p>Keyword counts from US Presidential State of the Union Addresses and Presidential Budget Messages. This was done using the Python scripts provided under <a href="https://github.com/JeremySilver/KeywordCountsPresidentialMessages">https://github.com/JeremySilver/KeywordCountsPresidentialMessages</a>. The raw text data is from <a href="http://www.presidency.ucsb.edu/">The American Presidency Project</a> (<a href="http://www.ucsb.edu/">UCSB</a>), with some Presidential Budget Messages being extracted from US Federal Budget documents available through <a href="https://fraser.stlouisfed.org/">FRASER</a> (a digital library of U.S. economic, financial, and banking history) or, for the more recent documents the website of the <a href="https://www.whitehouse.gov/">White House</a>.</p> <p>The data headings are:</p> <ul> <li>pid: in most cases, this is the index for the text document as archived on <a href="http://www.presidency.ucsb.edu/">The American Presidency Project</a> website. In some cases, this was the filename of a plain-text file read directly.</li> <li>year: Year that the message was delivered.</li> <li>date: Date that the message was delivered.</li> <li>name: Name of the US President delivering the message.</li> <li>count_of_all_words: Count of all words in the document.</li> <li>count_of_keywords: Count of all keywords encountered in that document.</li> <li>Keyword specific columns - three per keyword. For example, for the 'energy' keyword, the 'energy' column gives the number of times the 'energy' keyword was counted in the message, 'energy_pct_of_keywords' gives this count as a percentage of all keywords, and 'energy_pct_of_all_words' gives this count as a percentage of all words</li> </ul> <p>Below is the list of keywords that match when the search is applied to a dictionary file containing over 99,000 US English words.</p> <ul> <li>energy: 'energy'</li> <li>tax: 'nontaxable', 'overtax', 'overtaxed', 'overtaxes', 'overtaxing', 'surtax', 'surtaxed', 'surtaxes', 'surtaxing', 'surtaxs', 'tax', 'taxable', 'taxation', 'taxations', 'taxed', 'taxes', 'taxing', 'taxpayer', 'taxpayers', 'taxs'</li> <li>defense: 'defend', 'defense'</li> <li>education: 'education'</li> <li>employment: 'employ', 'employable', 'employe', 'employed', 'employee', 'employees', 'employer', 'employers', 'employes', 'employing', 'employment', 'employments', 'employs', 'underemployed', 'unemployable', 'unemployed', 'unemployeds', 'unemployment', 'unemployments'</li> <li>research: 'research', 'researched', 'researcher', 'researchers', 'researches', 'researching', 'researchs'</li> <li>shooting: 'shooting'</li> <li>space: 'space'</li> <li>nuclear: 'nuclear'</li> <li>natural resources: 'natural resources'</li> <li>racism: 'racism', 'civil rights'</li> <li>crime: 'crime', 'crimes', 'criminal', 'criminally', 'criminals', 'decriminalization', 'decriminalizations', 'decriminalize', 'decriminalized', 'decriminalizes', 'decriminalizing'</li> <li>environment: 'environment', 'environmental', 'environmentalism', 'environmentalisms', 'environmentalist', 'environmentalists', 'environmentally', 'environments'</li> <li>religion: 'faith', 'god', 'prayer', 'religion'</li> <li>health: 'health', 'healthful', 'healthfully', 'healthfulness', 'healthfulnesss', 'healthier', 'healthiest', 'healthily', 'healthiness', 'healthinesss', 'healths', 'healthy', 'unhealthful', 'unhealthier', 'unhealthiest', 'unhealthy'</li> <li>terror: 'terror', 'terrorism', 'terrorisms', 'terrorist', 'terrorists', 'terrorize', 'terrorized', 'terrorizes', 'terrorizing', 'terrors'</li> <li>war: 'war', 'warrior', 'warriors', 'wars'</li> <li>economy: 'economic', 'economical', 'economically', 'economics', 'economicss', 'economy', 'economys', 'microeconomics', 'microeconomicss', 'socioeconomic', 'uneconomic', 'uneconomical'</li> <li>jobs: 'jobs'</li> <li>business: 'agribusiness', 'agribusinesses', 'agribusinesss', 'business', 'businesses', 'businesslike', 'businessman', 'businessmans', 'businessmen', 'businesss', 'businesswoman', 'businesswomans', 'businesswomen'</li> <li>drugs: 'drugs', 'narcotics'</li> <li>inflation: 'inflation'</li> <li>climate: 'climate'</li> <li>science: 'science', 'sciences', 'scientific', 'scientifically', 'scientist', 'scientists'</li> <li>gun: 'gun', 'gunfire', 'gunman', 'guns', 'handgun', 'rifle', 'shotgun'</li> <li>tech: 'biotechnology', 'biotechnologys', 'technical', 'technological', 'technologically', 'technologies', 'technologist', 'technologists', 'technology', 'technologys'</li> <li>military: 'military'</li> <li>security: 'security'</li> <li>housing: 'housing'</li> <li>pollution: 'pollution'</li> </ul> <p>The dictionary file used is a standard file among Linux systems, and the version used was provided with version 7.1-1 of the Ubuntu 'wamerican' package. Two extra phrases, which do not appear in the dictionary file, are added to the list: 'civil rights' (under the 'racism' keyword) and 'natural resources' (under the 'natural resources' theme).</p>
Dataset: Comparative evaluation of a keyword based search and semantic search in a data portal for biodiversity research.
<p>Supplementary material for a comparative evaluation of a keyword based search and semantic search in a data portal for biodiversity research. We conducted a relevance evaluation with 6 users over 19 search questions in two search interfaces.</p> <p>The users provided up to five search questions and relevant keywords from their research background. We setup a dataset search over a corpus of ~92,000 randomly selected metadata files from GFBio (<a href="https://www.gfbio.org">https://www.gfbio.org</a>). For each of their own search queries, the users got two result sets presented. The first one displayed results obtained from a keyword search. The second panel contained dataset results from a prototypical semantic search. Instead of results with exact mentions of the query terms, the semantic search also presented related results with synonyms and more specific terms or terms obtained from concept nodes of a higher hierarchy level.</p> <p>Each user rated the relevance of his/her own search queries on a 7-point Likert scale for both search results.<br> In addition, users also assessed the expanded keywords for each question.</p> <p>More information can be found in our publication:</p> <p>Löffler, F. and Klan, F. (2016): Does Term Expansion Matter for the Retrieval of Biodiversity Data? in Joint Proceedings of the Posters and Demos Track of the 12th International Conference on Semantic Systems - SEMANTiCS2016 and the 1st International Workshop on Semantic Change & Evolving Semantics (SuCCESS'16), co-located with the 12th International Conference on Semantic Systems (SEMANTiCS 2016),2016, <a href="http://ceur-ws.org/Vol-1695/paper2.pdf">http://ceur-ws.org/Vol-1695/paper2.pdf</a></p> <p> </p>
Artifacts for Keyword Extraction From Specification Documents for Planning Security Mechanisms
<p>This dataset contains the data used for evaluating VDocScan - a keyword extraction based security vulnerability prediction method. The repository includes an extensive list of Products and Vulnerability reports from CVE, a custom created dataset mapping vulnerability reports to product documentations, as well as, intermediate results from the study such as decision trees rendered for each vulnerability, correlation matrix of vulnerabilities etc. </p>
A mapping of keywords from published papers on alien squirrels to biological invasion research themes
<p><strong>Context</strong></p> <p>This dataset was used to produce the worldl and the graphs in the editorial to the research topic <a href="https://www.frontiersin.org/research-topics/29270/ecology-impact-and-management-of-squirrel-invasions"><em>Ecology, impact and management of squirrel invasions</em></a> (La Morgia et al. 2023).</p> <p><strong>Contents of the dataset</strong></p> <p>The dataset contains the keywords of papers since 2000 harvested with a Web of Science search (performed on 29/05/2023) using the advanced search string TS=(invasive squirrel) OR TI=(invasive squirrel) OR AB=(invasive squirrel). We screened the search results, excluding papers irrelevant to alien squirrels, for example, papers on computer science or physiology, medical or other aspects without any bearing to conservation science. To do this, we checked the abstract and keywords of the papers. Out of the 401 initial papers, after this first screening, we kept 217 in this dataset. The keywords of these papers were manually assigned to alien squirrel research topics by the authors of this dataset (using an own categorisation) and then mapped to the seven broad themes of invasive alien species research of <a href="https://doi.org/10.1007/s10530-023-03067-7">Stevenson et al. (2023)</a>: </p> <ol> <li>Ecosystems: topics which discuss a specific region, or biome, or focused on a particular species strongly associated with one ecosystem type;</li> <li>Monitoring: topics regarding all aspects of monitoring, including detection, identification, and distributional mapping;</li> <li>Management and decision-making: topics discussing the management and socio-political aspects of invasion science, such as prevention, control, and policy;</li> <li>Interactions: topics discussing the interactions with native species, or the effects of those interactions</li> <li>Assessing change: topics focused on studying and analysing temporal and ecological change;</li> <li>Traits: topics that explored the characteristics of alien squirrels;</li> <li>Invasion mechanisms: topics discussing dispersal pathways and drivers of spread.</li> </ol> <p><strong>Dataset description</strong></p> <p>Every row (N = 1275) in the comma-separated .csv represents one original keyword with reference to the paper in which that keyword appears and mapped to the research topics on invasive squirrels and the broad themes in invasion biology research. The .csv contains the following fields:</p> <ul> <li>ID: a unique ID assigned to the combination of an original keyword and the corresponding paper harvested from the WoS search</li> <li>original_keyword: the original keywords associated with the paper (WoS search)</li> <li>keyword_topic: categorization of original keywords into topics related to invasive squirrel research by La Morgia et al. (2023)</li> <li>mapped_category: mapping to one of the seven broad themes of invasive alien species research of <a href="https://doi.org/10.1007/s10530-023-03067-7">Stevenson et al. (2023)</a> as listed and described above</li> <li>authors: author(s) of the paper (WoS search)</li> <li>year: publication year of paper (WoS search)</li> <li>title: title of the paper (WoS search)</li> <li>journal: full journal name (WoS search)</li> <li>doi: full doi of the paper (WoS search)</li> </ul> <p><strong>Potential applications of the dataset</strong></p> <p>This dataset can be used to reproduce the graphs in La Morgia et al. (2023) or to perform more in-depth review or analysis of the literature on alien squirrel invasions. For more information and graph code, we refer to <a href="https://github.com/Vale-LaMo/squirrels">this GitHub repository</a>.</p>
Using Bidirected Graphs to Map Keywords
<p>This study attempts to demonstrate the significance of considering two-way relationships by proposing a keyword network formed using bidirected graphs and association rules to examine the two-way relationship of two or more keywords. A web application to visualize is accessible at <a href="http://www.coconut-libtool.com">www.coconut-libtool.com</a></p>
The US LTER Thesaurus: Contents and Keyword Use Statistics in LTER Data Packages in 2006 and 2018
This dataset contains raw data and statistical summaries that reflect use of keywords in LTER Datasets in May 2018 and 2006. Specific summaries include: Number of uses and number sites by keyword (LTERVocabKeywordSummary.csv), Summary of keyword use by data package (LTERVocabDataPackageSummary.csv), Summary of Keyword Use by LTER Site in 2018(LTERVocabSiteSummary.csv), Summary of Keyword Use by LTER Site in 2006(KeyStats2006.csv). Raw data includes XML files containing the US LTER Thesaurus in Moodle format and the ResultSet containing the information for each dataset from the Environmental Data Initiative PASTA repository.
Agriculture Keywords Dataset
<p>This dataset consists of 193 agricultural keywords in English and Luganda. 64 keywords are in English and 129 keywords are in Luganda. The list of keywords was compiled by obtaining the counts of the most used agricultural words in Luganda radio discussions and online newspaper in Uganda. The keywords are categorized into crops, diseases, fertilisers, herbicides and general agriculture-related keywords. The data consists of folders containing .wav files with unique IDs as file names. The name of each folder in dataset refers to the name of the keyword. The dataset consists of 5290 keyword utterances in Luganda. These were collected from different age groups and gender. </p>
The Use of Keywords in Archaeornithology Literature Appendix B in JSON
<p>Appendix B – Vocabulary Matching Tool output with Getty Art & Architecture Thesaurus terms. Formatted in JSON.</p>
No Man's Sky Patch Keywords for the article "Adapting the Harris Matrix for Software Stratigraphy"
<p>Full list of <em>No Man's Sky</em> patch keywords for the article "Adapting the Harris Matrix for Software Stratigraphy" published in <em>Advances in Archaeological Practice.</em></p>
Co-occurrences of trending keywords in popular tech media (01.2016-04.2019)
<p>Sources with weights</p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5 %</li> <li>IEEE Spectrum 5 %</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p>Methodology</p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues have been selected (e.g. 'metoo', 'gdpr')</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul> <p> </p>
Keyword frequencies in popular tech media (01.2016-04.2019)
<p><strong>Sources with weights</strong></p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5 %</li> <li>IEEE Spectrum 5 %</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p><strong>Methodology</strong></p> <ul> <li>Frequency of appearances for all unigrams and bigrams in the texts</li> <li>Frequency: number of appearances of every term divided by the number of published articles (for every month and source)</li> <li>This measure reveals how many times an expression has been mentioned on average per article</li> <li>Several media sources: a representative index is calculated with weighted average (weights as above)</li> <li>Average monthly change in the analised term's frequency is calculated by OLS regressions</li> <li>The dependent variable of the estimation is the frequency index, while the number of months since the beginning of the analysed period (January 2016) is the independent variable</li> <li>The regression coefficient (referred to as coef) shows by how much on average the analysed expression’s frequency changed with every observed month (marginal change of the frequency), revealing which keywords had the biggest monthly growth</li> </ul> <p><strong>Files</strong></p> <p>The dataset contains two files:</p> <p>Unigrams: coefs_1weighted_site.csv</p> <p>Bigrams: coefs_2weighted_site.csv</p> <p><strong>Columns</strong></p> <p>freq_months (e.g. freq_2019-04): the average frequency of the term</p> <p>coef: the regression coefficient</p> <p>coef_norm: the regression coefficient divided by the mean frequency of the keyword</p> <p>coef_norm_max: the regression coefficient divided by the maximum frequency of the keyword</p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p>
Keyword Spotting with African Languages
<p><strong>Keyword spotting refers to the task of learning to detect spoken keywords. It interfaces all modern voice-based virtual assistants on the market: Amazon’s Alexa, Apple’s Siri, and the Google Home device. Contrarily to speech recognition models, keyword spotting doesn’t run on the cloud, but directly on the device. </strong></p> <p><strong>The motivation of this paper is to extend the Speech commands dataset (Warden 2018) with African languages. In particular, we are going to focus on 6 Senegalese languages: Wolof, Pulaar, Serer, Mandinka, Diola, Soninke. </strong></p> <p><strong>The choice of these languages is guided, on the one hand, by their status as languages considered to be the languages of the first generation, that is to say, the first codified languages (endowed with a writing system and considered by the state of Senegal as national languages) with decree n ° 68-871 of July 24, 1968. On the other hand, they represent the languages that are most spoken in Senegal.</strong></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.