Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

58

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

58 results for “keywords”

Learn how ShareScore rates datasets ↗
zenodo44/100

Co-occurrences of trending keywords in popular tech media (01.2016-04.2021)

<p>Sources with weights</p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5&nbsp;%</li> <li>IEEE Spectrum 5&nbsp;%</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p>Methodology</p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues have been selected (e.g. &#39;metoo&#39;, &#39;gdpr&#39;)</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Keyword frequencies in popular tech media (01.2016-04.2021)

<p>Sources with weights</p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5&nbsp;%</li> <li>IEEE Spectrum 5&nbsp;%</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p>Methodology</p> <ul> <li>Frequency of appearances for all unigrams and bigrams in the texts</li> <li>Frequency: number of appearances of every term divided by the number of all terms&nbsp;(for every month and source)&nbsp;</li> <li>Several media sources: a representative index is calculated with weighted average (weights as above)</li> <li>Average monthly change in the analised term&#39;s frequency is calculated by OLS regressions</li> <li>The dependent variable of the estimation is the frequency index, while the number of months since the beginning of the analysed period (January 2016) is the independent variable</li> <li>The regression coefficient (referred to as coef) shows by how much on average the analysed expression&rsquo;s frequency changed with every observed month (marginal change of the frequency), revealing which keywords had the biggest monthly growth</li> </ul> <p>Columns</p> <p>freq_months (e.g. freq_2019-04):&nbsp;the average frequency of the term</p> <p>coef:&nbsp;the regression coefficient</p> <p>coef_norm:&nbsp;the regression coefficient divided by the mean frequency of the keyword</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Co-occurrences of trending keywords in popular tech media during the COVID-19 pandemic (01.2020-06.2020)

<p>Sources:&nbsp;</p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe&nbsp;</li> <li>IEEE Spectrum&nbsp;</li> <li>Techforge&nbsp;</li> <li>Fastcompany&nbsp;</li> <li>The Guardian (Tech)&nbsp;</li> <li>Arstechnica&nbsp;</li> <li>Reuters&nbsp;</li> <li>Gizmodo&nbsp;</li> <li>ZDNet&nbsp;</li> <li>The Register&nbsp;</li> <li>The Verge&nbsp;</li> <li>TechCrunch&nbsp;</li> </ul> <p>Methodology</p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues and technologies have been selected (e.g. covid19)</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Keyword frequencies in popular tech media during the COVID-19 pandemic (01.2020-06.2020)

<p>Sources:&nbsp;</p> <ul> <li>Euractiv</li> <li>The Conversation</li> <li>Politico Europe&nbsp;</li> <li>IEEE Spectrum&nbsp;</li> <li>Techforge&nbsp;</li> <li>Fastcompany&nbsp;</li> <li>The Guardian (Tech)&nbsp;</li> <li>Arstechnica&nbsp;</li> <li>Reuters&nbsp;</li> <li>Gizmodo&nbsp;</li> <li>ZDNet&nbsp;</li> <li>The Register&nbsp;</li> <li>The Verge&nbsp;</li> <li>TechCrunch&nbsp;</li> </ul> <p>Methodology is modified relative to the regular trend analysis due to the short period of analysis (weekly freqiencies)</p> <ul> <li>Frequency of appearances for all unigrams and bigrams in the texts</li> <li>Frequency: number of appearances of every term divided by the number of all terms&nbsp;(for every week)&nbsp;</li> <li>Several media sources: all articles are treated equally</li> <li>Average monthly change in the analised term&#39;s frequency is calculated by OLS regressions</li> <li>The dependent variable of the estimation is the frequency index, while the number of weeks since the beginning of the analysed period (January 2020) is the independent variable</li> <li>The regression coefficient (referred to as coef) shows by how much on average the analysed expression&rsquo;s frequency changed with every observed week&nbsp;(marginal change of the frequency), revealing which keywords had the biggest weekly&nbsp;growth</li> </ul> <p>Columns</p> <p>freq_2020_weeks&nbsp;(e.g. freq_2020_ww0):&nbsp;the average frequency of the term</p> <p>coef:&nbsp;the regression coefficient</p> <p>coef_norm:&nbsp;the regression coefficient divided by the mean frequency of the keyword</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Corpus and list of keywords from Improving sustainable crop protection using population genetics concepts

<p>Corpus extracted in April 2021 from the ISI Web of Science portal (https://www.webofscience.com) with the following request: &lsquo;Plant AND Resistan* AND Durab*&rsquo;. A first corpus of 2522 articles was built considering all publication years for this extraction. This collection was then refined by categories to remove articles outwith the scope of our search (e.g. related to durable resistant materials for constructions). We also kept only articles cited at least once. The final corpus was composed of 1783 articles from 1979 to 2021:</p> <ul> <li>CORPUS_plant_resistance_durability.zip</li> </ul> <p>List of keywords used for the network presented in the article:</p> <ul> <li>keywords_list.csv</li> </ul>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Co-occurrences of trending keywords in popular tech media

<p><strong>Co-occurrences of trending keywords in the tech media (01.2016-03.2019)</strong></p> <p><strong>Sources</strong></p> <ul> <li>Gigaom 0.5%</li> <li>Euractiv 0.9%</li> <li>The Conversation 1.3%</li> <li>Politico Europe 1.3%</li> <li>IEEE Spectrum 1.8%</li> <li>Techforge 4.3%</li> <li>Fastcompany 4.5%</li> <li>The Guardian (Tech) 9.2%</li> <li>Arstechnica 10.0%</li> <li>Reuters 11%</li> <li>Gizmodo 17.5%</li> <li>ZDNet 18.3%</li> <li>The Register 19.5%</li> </ul> <p><strong>Methodology</strong></p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues have been selected (e.g. &#39;metoo&#39;, &#39;gdpr&#39;)</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul> <p><strong>Files</strong></p> <p>unigram-unigram co-occurrences: cooc11weighted.csv</p> <p>unigram-bigram co-occurrences: cooc12weighted.csv</p> <p>bigram-unigram co-occurrences: cooc21weighted.csv</p> <p>bigram-bigram co-occurrences: cooc22weighted.csv<br> &nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2019View details →
zenodo44/100

Keyword frequency in popular tech media

<p><strong>Keywords trending in the tech media (01.2016-03.2019)</strong></p> <p><strong>Sources</strong></p> <ul> <li>Gigaom 0.5%</li> <li>Euractiv 0.9%</li> <li>The Conversation 1.3%</li> <li>Politico Europe 1.3%</li> <li>IEEE Spectrum 1.8%</li> <li>Techforge 4.3%</li> <li>Fastcompany 4.5%</li> <li>The Guardian (Tech) 9.2%</li> <li>Arstechnica 10.0%</li> <li>Reuters 11%</li> <li>Gizmodo 17.5%</li> <li>ZDNet 18.3%</li> <li>The Register 19.5%</li> </ul> <p><strong>Methodology</strong></p> <ul> <li>Frequency of appearances for all unigrams and bigrams in the texts</li> <li>Frequency: number of appearances of every term divided by the number of published articles (for every month and source)</li> <li>This measure reveals how many times an expression has been mentioned on average per article</li> <li>Several media sources: a representative index is calculated with weighted average</li> <li>Average monthly change in the analised term&#39;s frequency is calculated by OLS regressions</li> <li>The dependent variable of the estimation is the frequency index, while the number of months since the beginning of the analysed period (January 2016) is the independent variable</li> <li>The regression coefficient (referred to as coef) shows by how much on average the analysed expression&rsquo;s frequency changed with every observed month (marginal change of the frequency), revealing which keywords had the biggest monthly growth</li> </ul> <p><strong>Files</strong></p> <ul> <li>unigrams: coefs_1weighted_site.csv</li> <li>bigrams: coefs_2weighted_site.csv</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jan 2019View details →
zenodo44/100

Unpacking the concept of "educators' data literacy in Higher Education" - Systematic Review of the literature and Keyword Map

<p>As algorithmic decision-making and data collection become pervasive within higher education, how can educators make sense of the systems that shape life and learning in the 21st century? Through a systematic review of the literature, the paper investigates the gaps in the literature, which prevent the formulation of potential pathways and principles on which educators&rsquo; data literacy can - and should - be developed and fostered. The analysis of 137 papers through the methods of classification under relevant categories, and key words mapping, showed that there is little attention on HE teachers, and most approaches to educators&rsquo; data literacy address management and technical abilities for data processing, with less concern on critical, ethical and personal approaches to datafication in education.</p> <p>The present dataset shows the full list of articles analysed.</p> <p>The dataset, and ods file, is composed by the following sheets:</p> <ol> <li>Codebook</li> <li>List of articles extracted from SCOPUS</li> <li>List of articles extracted from WOS</li> <li>List of articles extracted from ERIC</li> <li>List of articles extracted from DOAJ</li> <li>Interrater Agreement</li> <li>PRISMA workflow</li> <li>Analysis - First Level (classification of 137 articles selected)</li> <li>Analysis - Second Level (List of articles relating faculty development)</li> <li>Supplementary tables (counting articles in relation to the categories of analysis).</li> </ol> <p>As for the Keywords&#39; Map, a second file .csv displays the&nbsp;text&nbsp; over which basis was performed the keyword maps analysis. A .txt file shows notes relating the analysis procedures using the software VOS-Viewer&nbsp;<a href="http://www.vosviewer.com/">http://www.vosviewer.com/</a></p> <p>&nbsp;</p>

opencc-by-4.0Jun 2019View details →
zenodo44/100

Keyword counts from US Presidential State of the Union Addresses and Presidential Budget Messages

<p>Keyword counts from US Presidential State of the Union Addresses and Presidential Budget Messages. This was done using the Python scripts provided under&nbsp;<a href="https://github.com/JeremySilver/KeywordCountsPresidentialMessages">https://github.com/JeremySilver/KeywordCountsPresidentialMessages</a>. The raw text data is from&nbsp;<a href="http://www.presidency.ucsb.edu/">The American Presidency Project</a>&nbsp;(<a href="http://www.ucsb.edu/">UCSB</a>), with some&nbsp;Presidential Budget Messages being extracted from US Federal Budget documents available through&nbsp;<a href="https://fraser.stlouisfed.org/">FRASER</a>&nbsp;(a digital library of U.S. economic, financial, and banking history) or, for the more recent documents the website of the&nbsp;<a href="https://www.whitehouse.gov/">White House</a>.</p> <p>The data headings are:</p> <ul> <li>pid: in most cases, this is the index for the text document as archived on&nbsp;<a href="http://www.presidency.ucsb.edu/">The American Presidency Project</a>&nbsp;website. In some cases, this was the filename of a plain-text file read directly.</li> <li>year: Year that the message was delivered.</li> <li>date: Date that the message was delivered.</li> <li>name: Name of the US President delivering the message.</li> <li>count_of_all_words: Count of all words in the document.</li> <li>count_of_keywords: Count of all keywords encountered in that document.</li> <li>Keyword specific columns - three per keyword. For example, for the &#39;energy&#39; keyword, the&nbsp;&#39;energy&#39; column gives the number of times the &#39;energy&#39; keyword was counted in the message, &#39;energy_pct_of_keywords&#39; gives this count as a percentage of all keywords, and &#39;energy_pct_of_all_words&#39; gives this count as a percentage of all words</li> </ul> <p>Below is the list of keywords that match when the search is applied to a dictionary file containing over 99,000 US English words.</p> <ul> <li>energy: &#39;energy&#39;</li> <li>tax: &#39;nontaxable&#39;, &#39;overtax&#39;, &#39;overtaxed&#39;, &#39;overtaxes&#39;, &#39;overtaxing&#39;, &#39;surtax&#39;, &#39;surtaxed&#39;, &#39;surtaxes&#39;, &#39;surtaxing&#39;, &#39;surtaxs&#39;, &#39;tax&#39;, &#39;taxable&#39;, &#39;taxation&#39;, &#39;taxations&#39;, &#39;taxed&#39;, &#39;taxes&#39;, &#39;taxing&#39;, &#39;taxpayer&#39;, &#39;taxpayers&#39;, &#39;taxs&#39;</li> <li>defense: &#39;defend&#39;, &#39;defense&#39;</li> <li>education: &#39;education&#39;</li> <li>employment: &#39;employ&#39;, &#39;employable&#39;, &#39;employe&#39;, &#39;employed&#39;, &#39;employee&#39;, &#39;employees&#39;, &#39;employer&#39;, &#39;employers&#39;, &#39;employes&#39;, &#39;employing&#39;, &#39;employment&#39;, &#39;employments&#39;, &#39;employs&#39;, &#39;underemployed&#39;, &#39;unemployable&#39;, &#39;unemployed&#39;, &#39;unemployeds&#39;, &#39;unemployment&#39;, &#39;unemployments&#39;</li> <li>research: &#39;research&#39;, &#39;researched&#39;, &#39;researcher&#39;, &#39;researchers&#39;, &#39;researches&#39;, &#39;researching&#39;, &#39;researchs&#39;</li> <li>shooting: &#39;shooting&#39;</li> <li>space: &#39;space&#39;</li> <li>nuclear: &#39;nuclear&#39;</li> <li>natural&nbsp;resources: &#39;natural&nbsp;resources&#39;</li> <li>racism: &#39;racism&#39;, &#39;civil rights&#39;</li> <li>crime: &#39;crime&#39;, &#39;crimes&#39;, &#39;criminal&#39;, &#39;criminally&#39;, &#39;criminals&#39;, &#39;decriminalization&#39;, &#39;decriminalizations&#39;, &#39;decriminalize&#39;, &#39;decriminalized&#39;, &#39;decriminalizes&#39;, &#39;decriminalizing&#39;</li> <li>environment: &#39;environment&#39;, &#39;environmental&#39;, &#39;environmentalism&#39;, &#39;environmentalisms&#39;, &#39;environmentalist&#39;, &#39;environmentalists&#39;, &#39;environmentally&#39;, &#39;environments&#39;</li> <li>religion: &#39;faith&#39;, &#39;god&#39;, &#39;prayer&#39;, &#39;religion&#39;</li> <li>health: &#39;health&#39;, &#39;healthful&#39;, &#39;healthfully&#39;, &#39;healthfulness&#39;, &#39;healthfulnesss&#39;, &#39;healthier&#39;, &#39;healthiest&#39;, &#39;healthily&#39;, &#39;healthiness&#39;, &#39;healthinesss&#39;, &#39;healths&#39;, &#39;healthy&#39;, &#39;unhealthful&#39;, &#39;unhealthier&#39;, &#39;unhealthiest&#39;, &#39;unhealthy&#39;</li> <li>terror: &#39;terror&#39;, &#39;terrorism&#39;, &#39;terrorisms&#39;, &#39;terrorist&#39;, &#39;terrorists&#39;, &#39;terrorize&#39;, &#39;terrorized&#39;, &#39;terrorizes&#39;, &#39;terrorizing&#39;, &#39;terrors&#39;</li> <li>war: &#39;war&#39;, &#39;warrior&#39;, &#39;warriors&#39;, &#39;wars&#39;</li> <li>economy: &#39;economic&#39;, &#39;economical&#39;, &#39;economically&#39;, &#39;economics&#39;, &#39;economicss&#39;, &#39;economy&#39;, &#39;economys&#39;, &#39;microeconomics&#39;, &#39;microeconomicss&#39;, &#39;socioeconomic&#39;, &#39;uneconomic&#39;, &#39;uneconomical&#39;</li> <li>jobs: &#39;jobs&#39;</li> <li>business: &#39;agribusiness&#39;, &#39;agribusinesses&#39;, &#39;agribusinesss&#39;, &#39;business&#39;, &#39;businesses&#39;, &#39;businesslike&#39;, &#39;businessman&#39;, &#39;businessmans&#39;, &#39;businessmen&#39;, &#39;businesss&#39;, &#39;businesswoman&#39;, &#39;businesswomans&#39;, &#39;businesswomen&#39;</li> <li>drugs: &#39;drugs&#39;, &#39;narcotics&#39;</li> <li>inflation: &#39;inflation&#39;</li> <li>climate: &#39;climate&#39;</li> <li>science: &#39;science&#39;, &#39;sciences&#39;, &#39;scientific&#39;, &#39;scientifically&#39;, &#39;scientist&#39;, &#39;scientists&#39;</li> <li>gun: &#39;gun&#39;, &#39;gunfire&#39;, &#39;gunman&#39;, &#39;guns&#39;, &#39;handgun&#39;, &#39;rifle&#39;, &#39;shotgun&#39;</li> <li>tech: &#39;biotechnology&#39;, &#39;biotechnologys&#39;, &#39;technical&#39;, &#39;technological&#39;, &#39;technologically&#39;, &#39;technologies&#39;, &#39;technologist&#39;, &#39;technologists&#39;, &#39;technology&#39;, &#39;technologys&#39;</li> <li>military: &#39;military&#39;</li> <li>security: &#39;security&#39;</li> <li>housing: &#39;housing&#39;</li> <li>pollution: &#39;pollution&#39;</li> </ul> <p>The dictionary file used is a standard file among Linux systems, and the version used was provided with version 7.1-1 of the Ubuntu &#39;wamerican&#39; package.&nbsp;Two extra phrases, which do not appear in the dictionary file, are added to the list: &#39;civil rights&#39; (under the &#39;racism&#39; keyword) and &#39;natural&nbsp;resources&#39; (under the &#39;natural&nbsp;resources&#39; theme).</p>

opencc-by-4.0Jun 2019View details →
zenodo44/100

Dataset: Comparative evaluation of a keyword based search and semantic search in a data portal for biodiversity research.

<p>Supplementary material for a comparative evaluation of a keyword based search and semantic search in a data portal for biodiversity research. We conducted a relevance evaluation with 6 users over 19 search questions in two search interfaces.</p> <p>The users provided up to five search questions and relevant keywords from their research background. We setup a dataset search over a corpus of ~92,000 randomly selected metadata files from GFBio (<a href="https://www.gfbio.org">https://www.gfbio.org</a>). For each of their own search queries, the users got two result sets presented. The first one displayed results obtained from a keyword search. The second panel contained dataset results from a prototypical semantic search. Instead of results with exact mentions of the query terms, the semantic search also presented related results with synonyms and more specific terms or terms obtained from concept nodes of a higher hierarchy level.</p> <p>Each user rated the relevance of his/her own search queries on a 7-point Likert scale for both search results.<br> In addition, users also assessed the expanded keywords for each question.</p> <p>More information can be found in our publication:</p> <p>L&ouml;ffler, F. and Klan, F. (2016): Does Term Expansion Matter for the Retrieval of Biodiversity Data? in Joint Proceedings of the Posters and Demos Track of the 12th International Conference on Semantic Systems - SEMANTiCS2016 and the 1st International Workshop on Semantic Change &amp; Evolving Semantics (SuCCESS&#39;16), co-located with the 12th International Conference on Semantic Systems (SEMANTiCS 2016),2016, <a href="http://ceur-ws.org/Vol-1695/paper2.pdf">http://ceur-ws.org/Vol-1695/paper2.pdf</a></p> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Artifacts for Keyword Extraction From Specification Documents for Planning Security Mechanisms

<p>This dataset contains the data used for evaluating VDocScan - a keyword extraction based security vulnerability prediction method. The repository includes an extensive list of Products and Vulnerability reports from CVE, a custom created dataset mapping vulnerability reports to product documentations, as well as, intermediate results from the study such as decision trees rendered for each vulnerability, correlation matrix of vulnerabilities etc.&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo44/100

A mapping of keywords from published papers on alien squirrels to biological invasion research themes

<p><strong>Context</strong></p> <p>This dataset was used to produce the worldl and the graphs in the editorial to the research topic <a href="https://www.frontiersin.org/research-topics/29270/ecology-impact-and-management-of-squirrel-invasions"><em>Ecology, impact&nbsp;and management of squirrel invasions</em></a>&nbsp;(La Morgia et al. 2023).</p> <p><strong>Contents of the dataset</strong></p> <p>The dataset contains the keywords of papers since 2000 harvested with a Web of Science search (performed on 29/05/2023) using the advanced search string&nbsp;TS=(invasive squirrel) OR TI=(invasive squirrel) OR AB=(invasive squirrel). We screened the search results, excluding papers irrelevant to alien squirrels, for example, papers on computer science or physiology, medical or other aspects without any bearing to conservation science. To do this, we checked the abstract and keywords of the papers.&nbsp;Out of the 401&nbsp;initial papers, after this first screening, we kept 217 in this dataset.&nbsp;The&nbsp;keywords of these papers were manually assigned to alien squirrel research topics by the authors of this dataset (using an own categorisation) and then mapped to the seven broad themes of invasive alien species research of <a href="https://doi.org/10.1007/s10530-023-03067-7">Stevenson et al. (2023)</a>:&nbsp;</p> <ol> <li>Ecosystems: topics which discuss a specific region, or biome, or focused on a particular species strongly associated with one ecosystem type;</li> <li>Monitoring: topics regarding all aspects of monitoring, including detection, identification, and distributional mapping;</li> <li>Management and decision-making: topics discussing the management and socio-political aspects of invasion&nbsp;science, such as prevention, control, and policy;</li> <li>Interactions: topics discussing the interactions with native species, or the effects of those interactions</li> <li>Assessing change: topics focused on studying and analysing temporal and ecological change;</li> <li>Traits: topics that explored the characteristics of alien squirrels;</li> <li>Invasion mechanisms: topics discussing dispersal pathways and drivers of spread.</li> </ol> <p><strong>Dataset description</strong></p> <p>Every row (N = 1275)&nbsp;in the comma-separated .csv represents one original keyword with reference to the paper in which that keyword appears and mapped to the research topics on invasive squirrels and the broad themes in invasion biology research. The .csv contains the following fields:</p> <ul> <li>ID: a unique ID assigned to the combination of an original keyword and the corresponding paper&nbsp;harvested&nbsp;from the WoS search</li> <li>original_keyword: the original keywords associated with the paper&nbsp;(WoS search)</li> <li>keyword_topic: categorization&nbsp;of original keywords into topics related to invasive squirrel research by La Morgia et al. (2023)</li> <li>mapped_category:&nbsp;mapping to one of the seven broad themes of invasive alien species research of <a href="https://doi.org/10.1007/s10530-023-03067-7">Stevenson et al. (2023)</a>&nbsp;as listed and described above</li> <li>authors: author(s) of the paper (WoS search)</li> <li>year: publication year of paper&nbsp;(WoS search)</li> <li>title: title of the paper (WoS search)</li> <li>journal: full journal name (WoS search)</li> <li>doi: full doi of the paper&nbsp;(WoS search)</li> </ul> <p><strong>Potential applications of the dataset</strong></p> <p>This dataset can be used to reproduce the graphs in La Morgia et al. (2023) or to perform more in-depth review or analysis of the literature on alien squirrel invasions. For more information and graph code, we refer to <a href="https://github.com/Vale-LaMo/squirrels">this GitHub repository</a>.</p>

opencc-zeroJun 2023View details →
zenodo44/100

Using Bidirected Graphs to Map Keywords

<p>This study attempts to demonstrate the significance of considering two-way relationships by proposing a keyword network formed using bidirected graphs and association rules to examine the two-way relationship of two or more keywords. A web application to visualize is accessible at <a href="http://www.coconut-libtool.com">www.coconut-libtool.com</a></p>

opencc-by-4.0Jun 2023View details →
edi44/100

The US LTER Thesaurus: Contents and Keyword Use Statistics in LTER Data Packages in 2006 and 2018

This dataset contains raw data and statistical summaries that reflect use of keywords in LTER Datasets in May 2018 and 2006. Specific summaries include: Number of uses and number sites by keyword (LTERVocabKeywordSummary.csv), Summary of keyword use by data package (LTERVocabDataPackageSummary.csv), Summary of Keyword Use by LTER Site in 2018(LTERVocabSiteSummary.csv), Summary of Keyword Use by LTER Site in 2006(KeyStats2006.csv). Raw data includes XML files containing the US LTER Thesaurus in Moodle format and the ResultSet containing the information for each dataset from the Environmental Data Initiative PASTA repository.

openCustomMay 2018View details →
zenodo40/100

Agriculture Keywords Dataset

<p>This dataset consists of 193 agricultural keywords in English and Luganda. 64&nbsp;keywords are in English and 129 keywords are in Luganda.&nbsp;The list of keywords was compiled by obtaining the counts of the most used agricultural words in Luganda radio discussions and online&nbsp;newspaper in Uganda. The keywords are&nbsp;categorized into crops, diseases, fertilisers, herbicides and general agriculture-related&nbsp;keywords.&nbsp;The data consists of folders containing .wav files with unique IDs as file names. The name of each folder in dataset refers to the name of the keyword. The dataset consists of 5290 keyword utterances in Luganda. These were collected from different age groups and gender.&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

The Use of Keywords in Archaeornithology Literature Appendix B in JSON

<p>Appendix B &ndash; Vocabulary Matching Tool output with Getty Art &amp; Architecture Thesaurus terms. Formatted in JSON.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

No Man's Sky Patch Keywords for the article "Adapting the Harris Matrix for Software Stratigraphy"

<p>Full list of&nbsp;<em>No Man&#39;s Sky</em>&nbsp;patch keywords for the article&nbsp;&quot;Adapting the Harris Matrix for Software Stratigraphy&quot; published in&nbsp;<em>Advances in Archaeological Practice.</em></p>

opencc-by-4.0Jan 2018View details →
zenodo40/100

Co-occurrences of trending keywords in popular tech media (01.2016-04.2019)

<p>Sources with weights</p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5&nbsp;%</li> <li>IEEE Spectrum 5&nbsp;%</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p>Methodology</p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues have been selected (e.g. &#39;metoo&#39;, &#39;gdpr&#39;)</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jun 2019View details →
zenodo40/100

Keyword frequencies in popular tech media (01.2016-04.2019)

<p><strong>Sources with weights</strong></p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5&nbsp;%</li> <li>IEEE Spectrum 5&nbsp;%</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p><strong>Methodology</strong></p> <ul> <li>Frequency of appearances for all unigrams and bigrams in the texts</li> <li>Frequency: number of appearances of every term divided by the number of published articles (for every month and source)</li> <li>This measure reveals how many times an expression has been mentioned on average per article</li> <li>Several media sources: a representative index is calculated with weighted average (weights as above)</li> <li>Average monthly change in the analised term&#39;s frequency is calculated by OLS regressions</li> <li>The dependent variable of the estimation is the frequency index, while the number of months since the beginning of the analysed period (January 2016) is the independent variable</li> <li>The regression coefficient (referred to as coef) shows by how much on average the analysed expression&rsquo;s frequency changed with every observed month (marginal change of the frequency), revealing which keywords had the biggest monthly growth</li> </ul> <p><strong>Files</strong></p> <p>The dataset contains two files:</p> <p>Unigrams: coefs_1weighted_site.csv</p> <p>Bigrams: coefs_2weighted_site.csv</p> <p><strong>Columns</strong></p> <p>freq_months (e.g. freq_2019-04):&nbsp;the average frequency of the term</p> <p>coef:&nbsp;the regression coefficient</p> <p>coef_norm:&nbsp;the regression coefficient divided by the mean frequency of the keyword</p> <p>coef_norm_max:&nbsp;the regression coefficient divided by the maximum frequency of the keyword</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2019View details →
zenodo40/100

Keyword Spotting with African Languages

<p><strong>Keyword spotting refers to the task of learning to detect spoken keywords. It interfaces all modern voice-based virtual assistants on the market: Amazon&rsquo;s Alexa, Apple&rsquo;s Siri, and the Google Home device. Contrarily to speech recognition models, keyword spotting doesn&rsquo;t run on the cloud, but directly on the device.&nbsp;</strong></p> <p><strong>The motivation of this paper is to extend the Speech commands dataset (Warden 2018) with African languages. In particular, we are going to focus on 6 Senegalese languages: Wolof, Pulaar, Serer, Mandinka, Diola, Soninke.&nbsp;</strong></p> <p><strong>The choice of these languages is guided, on the one hand, by their status as languages considered to be the languages of the first generation, that is to say, the first codified languages (endowed with a writing system and considered by the state of Senegal as national languages) with decree n &deg; 68-871 of July 24, 1968. On the other hand, they represent the languages that are most spoken in Senegal.</strong></p>

opencc-by-4.0Apr 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record