Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
58
datasets available to search
ShareScore release 0.9.0
Dataset results
58 results for “keywords”
Keywords and Terms from the LTER Network - 2006
This dataset contains a list of keywords and keyterms used in U.S. LTER documents and datasets. It was jointly created by the Information Management Committee of the U.S. LTER.
Co-occurrences of trending keywords in popular tech media (01.2016-02.2020)
<p><strong>Sources with weights</strong></p> <pre> Arstechnica: 1/8, Euractiv: 1/8, Fastcompany: 1/8, The Register: 1/8, Techcrunch: 1/8, The Guardian: 1/8, Venturebeat: 1/8, The Verge: 1/8</pre> <p><strong>Methodology</strong></p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues and technologies have been selected (e.g. 'gdpr', '5G')</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul>
Keyword frequencies in popular tech media (01.2016-02.2020)
<p><strong>Sources with weights</strong></p> <pre> Arstechnica: 1/8, Euractiv: 1/8, Fastcompany: 1/8, The Register: 1/8, Techcrunch: 1/8, The Guardian: 1/8, Venturebeat: 1/8, The Verge: 1/8 </pre> <p><strong>Methodology</strong></p> <ul> <li>Frequency of appearances for all unigrams and bigrams in the texts</li> <li>Frequency: number of appearances of every term divided by the number of published articles (for every month and source)</li> <li>This measure reveals how many times an expression has been mentioned on average per article</li> <li>Several media sources: a representative index is calculated with weighted average (weights as above)</li> <li>Average monthly change in the analised term's frequency is calculated by OLS regressions</li> <li>The dependent variable of the estimation is the frequency index, while the number of months since the beginning of the analysed period (January 2016) is the independent variable</li> <li>The regression coefficient (referred to as coef) shows by how much on average the analysed expression’s frequency changed with every observed month (marginal change of the frequency), revealing which keywords had the biggest monthly growth</li> </ul> <p><strong>Columns</strong></p> <p>freq_months (e.g. freq_2019-04): the average frequency of the term</p> <p>coef: the regression coefficient</p> <p>coef_norm: the regression coefficient divided by the mean frequency of the keyword</p> <p>coef_norm_max: the regression coefficient divided by the maximum frequency of the keyword</p>
Thesaurus for KeyWords Plus for MEJ-24 2000-2019
<p>This is a supplemental file for the STI2024 submission "Delineating the field of medical education over time: A case study on interdisciplinarity and interuniversity collaboration patterns, 2000-2019". This file can be used when generating the KeyWords Plus networks of the Web of Science data outputs for the MEJ-24 2000-2019.<br> </p>
Metadata matrix: year 2017 by keywords, from webQDA.
<p>Part of the thesis of Sonia Verdugo Castro (USAL, Spain). <br> Title of the article: "The gender gap in STEM higher education studies: visualisation of literature".<br> Authors of the publication: Sonia Verdugo-Castro, Mª Cruz Sánchez-Gómez, Alicia García-Holgado, Francisco J. García-Peñalvo</p>
Metadata matrix: year 2015 by keywords, from webQDA.
<p>Part of the thesis of Sonia Verdugo Castro (USAL, Spain). <br> Title of the article: "The gender gap in STEM higher education studies: visualisation of literature".<br> Authors of the publication: Sonia Verdugo-Castro, Mª Cruz Sánchez-Gómez, Alicia García-Holgado, Francisco J. García-Peñalvo</p>
Keyword frequencies in arXiv and SSRN working papers
<p>The dataset contains the raw results of the trend analysis performed on two working paper repositories: ArXiv and SSRN.</p> <p>ArXiv: the dataset consists of working papers acquired via ArXiv’s API (<a href="https://arxiv.org/help/api/index">https://arxiv.org/help/api/index</a>). Working papers have been collected from the Computer Science discipline (all CS categories)</p> <p>SSRN (The Social Science Research Network): Working papers have been collected from two broad categories: 1. Information Systems & eBusiness, 2. Innovation. </p> <p>For each repository, there are two separate analyses: for the period 2016.01-2019.12 and for 2020.01-2020.06 (COVID-19).</p> <p>Methodology:</p> <ul> <li>Frequency of appearances for all unigrams and bigrams in the texts</li> <li>Frequency: number of appearances of every term divided by the number of all terms (for every month and in case of COVID - every week due to the shorter time period) </li> <li>Average monthly / weekly change in the analised term's frequency is calculated by OLS regressions</li> <li>The dependent variable of the estimation is the frequency index, while the number of months since the beginning of the analysed period (January 2016) is the independent variable (in the case of COVID: weeks since January 2020)</li> <li>The regression coefficient (referred to as coef) shows by how much on average the analysed expression’s frequency changed with every observed week (marginal change of the frequency), revealing which keywords had the biggest weekly growth</li> </ul> <p> </p> <p> </p>
Identified journal descriptor (JD), semantic type (ST), and MAUI keywords for PubMed/MEDLINE articles
<p>The Journal descriptor (JD) and Semantic type (ST) of PubMed articles were identified using a tool, called Journal descriptor indexing (JDI).</p> <p>MAUI keywords are identified by the MAUI tool and they can be used as a complement to MESH terms, as MeSH keywords are not always available in all PubMed articles.</p> <p>The methods of building the datasets can be found in the two articles: "Author name disambiguation in MEDLINE based on journal descriptors and semantic types" and " Exploring author name disambiguation on PubMed-scale"</p> <p>The PubMed database used is the 2019 baseline version, the number of articles in the two datasets are as follows:</p> <p>$ wc -l pubmed-paper-jd-st.tsv <br> 29796281 pubmed-paper-jd-st.tsv</p> <p>$ wc -l pubmed-paper-maui-keywords.tsv <br> 26509210 pubmed-paper-maui-keywords.tsv<br> </p>
Identity Lexicon and Keyword Counts for "The Life of a Tie: Social Origins of Network Diversity"
<p>Prototype identity lexicon for the categories of occupation (e.g., "reporter" at Boston Globe), familial roles (e.g., proud "father"), political affiliation (life-long "democrat"), and cultural and sports interests (e.g., "hiphop", "NFL").</p> <p>Also includes a CSV file with counts of each identity keywords matched for all 572K users in our dataset. </p> <p> </p>
humanities_keywords Dataset
<p>The WE1S humanities_keywords dataset contains word-frequency and other non-consumptive-use data about 474,930 unique documents (no duplicate or close variants) mentioning the word "humanities" in English-language news sources. and other keywords related to the humanities in English-language news sources. Other keywords include "liberal arts," "the arts," "literature," "history," and "philosophy." The documents came from 850 U.S. and 437 international news sources with their associated blogs (including student newspapers) published mostly during 1989-2019. <em>(See <a href="https://we1s.ucsb.edu/research/we1s-materials/">WE1S Research Materials Overview</a> for the relation between the project's "datasets" and "collections.")</em></p>
Programming language keyword frequencies extracted from 16,000,000 public GitHub repositories (October 2016)
<p>Origin</p> <p>16,000,000 repositories on GitHub as of October 2016, classified with github/linguist and parsed with Pygments. Token.Keyword tokens were filtered and MapReduce-d. Fuzzy duplicate repositories were discarded.</p> <p>Some languages, e.g. Haskell, are parsed wrong, resulting in <strong>many</strong> keywords. Still they were not removed since we are not familiar with such languages.</p> <p>Format</p> <p>Triples [language name]\t[keyword]\t[frequency]</p> <p>Tabs and new lines in keywords are escaped as \t and \n respectively.</p>
Transparency in Keyword Faceted Search: a dataset of Google Shopping html pages
<p>This dataset contains a collection of around 2,000 HTML pages: these web pages contain the search results obtained in return to queries for different products, searched by a set of synthetic users surfing Google Shopping (US version) from different locations, in July, 2016.</p> <p>Each file in the collection has a name where there is indicated the location from where the search has been done, the userID, and the searched product: <em>no_email_LOCATION_USERID.PRODUCT.shopping_testing.#.html</em></p> <p>The locations are Philippines (PHI), United States (US), India (IN). The userIDs: 26 to 30 for users searching from Philippines, 1 to 5 from US, 11 to 15 from India.</p> <p>Products have been choice following 130 keywords (e.g., MP3 player, MP4 Watch, Personal organizer, Television, etc.).</p> <p>In the following, we describe how the search results have been collected.</p> <p>Each user has a fresh profile. The creation of a new profile corresponds to launch a new, isolated, web browser client instance and open the Google Shopping US web page.</p> <p>To mimic real users, the synthetic users can browse, scroll pages, stay on a page, and click on links.</p> <p>A fully-fledged web browser is used to get the correct desktop version of the website under investigation. This is because websites could be designed to behave according to user agents, as witnessed by the differences between the mobile and desktop versions of the same website.</p> <p>The prices are the retail ones displayed by Google Shopping in US dollars (thus, excluding shipping fees).</p> <p>Several frameworks have been proposed for interacting with web browsers and analysing results from search engines. This research adopts OpenWPM. OpenWPM is automatised with <a href="http://www.seleniumhq.org/">Selenium</a> to efficiently create and manage different users with isolated Firefox and Chrome client instances, each of them with their own associated cookies.</p> <p>The experiments run, on average, 24 hours. In each of them, the software runs on our local server, but the browser's traffic is redirected to the designated remote servers (i.e., to India), via tunneling in SOCKS proxies. This way, all commands are simultaneously distributed over all proxies. The experiments adopt the Mozilla Firefox browser (version 45.0) for the web browsing tasks and run under Ubuntu 14.04. Also, for each query, we consider the first page of results, counting 40 products. Among them, the focus of the experiments is mostly on the top 10 and top 3 results.</p> <p>Due to connection errors, one of the Philippine profiles have no associated results. Also, for Philippines, a few keywords did not lead to any results: videocassette recorders, totes, umbrellas. Similarly, for US, no results were for totes and umbrellas.</p> <p>The search results have been analyzed in order to check if there were evidence of price steering, based on users' location.</p> <p><strong>One term of usage applies:</strong></p> <p>In any research product whose findings are based on this dataset, please cite</p> <pre>@inproceedings{DBLP:conf/ircdl/CozzaHPN19, author = {Vittoria Cozza and Van Tien Hoang and Marinella Petrocchi and Rocco {De Nicola}}, title = {Transparency in Keyword Faceted Search: An Investigation on Google Shopping}, booktitle = {Digital Libraries: Supporting Open Science - 15th Italian Research Conference on Digital Libraries, {IRCDL} 2019, Pisa, Italy, January 31 - February 1, 2019, Proceedings}, pages = {29--43}, year = {2019}, crossref = {DBLP:conf/ircdl/2019}, url = {https://doi.org/10.1007/978-3-030-11226-4\_3}, doi = {10.1007/978-3-030-11226-4\_3}, timestamp = {Fri, 18 Jan 2019 23:22:50 +0100}, biburl = {https://dblp.org/rec/bib/conf/ircdl/CozzaHPN19}, bibsource = {dblp computer science bibliography, https://dblp.org} } </pre> <p> </p> <p> </p>
Keyword impressions, category and location
<p>Data used in EW-Shopp (<a href="https://www.ew-shopp.eu/">https://www.ew-shopp.eu/</a>) project.</p> <p>Also, theese data sets have been used in the tutorial titled “SEMANTIC DATA ENRICHMENT FOR DATA SCIENTISTS” held by Matteo Palmonari (University of Milan-Bicocca, IT), Dumitru Roman (SINTEF, NO), Vincenzo Cutrona (University of Milan-Bicocca, IT), Nikolay Nikolov (SINTEF, NO), Aljaž Košmerlj (Jozef Stefan Institute, SI) at the Sixteenth Extended Semantic Web Conference (ESWC 2019), June 2019, Portoroz, SI. </p> <p>Link to the tutorial is: <a href="https://ew-shopp.github.io/eswc2019-tutorial/">https://ew-shopp.github.io/eswc2019-tutorial/</a></p>
Co-occurrences of trending keywords in popular tech media (01.2016-12.2019)
<p>Sources with weights</p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5 %</li> <li>IEEE Spectrum 5 %</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p>Methodology</p> <ul> <li>Exploring the relationship between topics</li> <li>Pairs of terms which are mentioned together in media articles</li> <li>Most trending social issues and technologies have been selected (e.g. 'gdpr', '5G')</li> <li>The co-occurrence analysis is calculated for pairs consisting of emerging social issues and trending uni/bigrams</li> <li>The number of times the terms appear in articles together with a social issue is divided by the number of times the social issue is mentioned across all articles</li> <li>A single index is constructed for all word pairs by weighted average (taking into account the prevalence of the given source)</li> </ul>
Keyword frequencies in popular tech media (01.2016-12.2019)
<p>Sources with weights</p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5 %</li> <li>IEEE Spectrum 5 %</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p>Methodology</p> <ul> <li>Frequency of appearances for all unigrams and bigrams in the texts</li> <li>Frequency: number of appearances of every term divided by the number of all terms (for every month and source) </li> <li>Several media sources: a representative index is calculated with weighted average (weights as above)</li> <li>Average monthly change in the analised term's frequency is calculated by OLS regressions</li> <li>The dependent variable of the estimation is the frequency index, while the number of months since the beginning of the analysed period (January 2016) is the independent variable</li> <li>The regression coefficient (referred to as coef) shows by how much on average the analysed expression’s frequency changed with every observed month (marginal change of the frequency), revealing which keywords had the biggest monthly growth</li> </ul> <p>Columns</p> <p>freq_months (e.g. freq_2019-04): the average frequency of the term</p> <p>coef: the regression coefficient</p> <p>coef_norm: the regression coefficient divided by the mean frequency of the keyword</p> <p> </p> <p> </p>
COVID-19 Twitter data, keyword stream 2020-01-13 to 2020-06-06
<p>Twitter data was collected through the Twitter API, specifically through the filter streaming endpoint, using the Crowdbreaks platform (<a href="http://crowdbreaks.org">crowdbreaks.org</a>) The data used in this work consists of a total of 353,993,900 tweets (thereof 267,026,740 retweets) posted by 26,262,332 users in a 146 day observation period, i.e. from January 13 to June 7, 2020. These tweets have been identified by Twitter to be in English language and match one or more of the keywords "wuhan", "ncov", "coronavirus", "covid" and "sars-cov-2".</p> <p>The data is complete with respect to these keywords, except during a period between mid-March to mid-April when volume exceeded the 1% threshold imposed by Twitter and was subsampled by an (unknown) degree.</p> <p>The following fields are published:</p> <ul> <li>id: Tweet ID</li> <li>is_retweet: Whether or not tweet is a retweet</li> <li>num_retweets: Number of retweets</li> <li>user.id: Id of tweeting user</li> <li>country_code: country code as predicted by local-geocode (https://github.com/mar-muel/local-geocode)</li> </ul>
WiP: "Keywords" on PubChem content
<p>Work in progress: exploring the creation of groups of chemicals based on PubChem annotation content in the form of "keywords" (the exact term is still a subject of debate ... for the moment keyword is the placeholder). This is a file deposition corresponding to the code base on the <a href="https://git-r3lab.uni.lu/eci/pubchem/-/tree/master/annotations/keywords">ECI GitLab</a> pages to create these files.</p> <p>Part of this work was performed at <a href="https://www.biohackathon-europe.org/">BioHackathon Europe</a> 2020 #BioHackEU20</p>
number of times keyword "Bayesian" appears in NASA/ADS entries
<p>number of times keyword "Bayesian" appears in NASA/ADS entries</p>
TopicTracker keywords and MeSH terms resulting from the analysis of papers on autonomy, equity, privacy, proportionality and trust in the context of Covid-19
<p>This dataset contains normalized keywords and MeSH terms contained in articles retrieved with 5 separate querioes on Covid-19 and autonomy, equity, privacy, proportionality, trust.</p>
Keyword-Matching-for-Canadian-Mechanical-Engineering-Programs-2023-2024
<p>This data repository contributed to the survey of the prevalence of Artificial Intelligence (AI) in current Canadian engineering curricula, which will be published in the proceedings of IDETC2024:</p> <ol> <li>Web-scraping course information (course codes, names, and descriptions) from the webites of engineering programs</li> <li>List of AI keywords</li> <li>Developing a keyword-matching algorithm to look for the keywords within every web-scraped course description and to generate a list of courses whose descriptions matched with at least one keyword</li> </ol> <p>The final results and scripts for each of these projects can be found in their respective folders in this repository</p> <p>The keyword-matching script can be directly run on the commandline and the webscraping scripts can be opened via Jupyter Notebooks or Google Colab.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.