Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

58

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

58 results for “keywords”

Learn how ShareScore rates datasets ↗
zenodo32/100

Embeddings for the paper "The Zapatista Semantic Struggle: Analysing the Linguistic Innovation of the EZLN with Semantic Difference Keywords (SDKs)"

<p>Embeddings from word2vec models described in "The Zapatista Semantic Struggle: Analysing the Linguistic Innovation of the EZLN with Semantic Difference Keywords (SDKs)". Full reference TBD.</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

List of 403 Keywords Related to Visual Impairment, Ocular Diseases and Eye Conditions

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
Figshare32/100

Keyword Analysis

<p>Keyword analysis, using 20 articles on digital files.</p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

Research Data for Beyond Keywords: Intent-Driven Semantic Code Search in Software Ecosystems

<p>This contains the software repository for the intent-enhanced code search engine as well as the baseline code search engine used in the Master Thesis: &#39;Beyond Keywords: Intent-Driven Semantic Code Search in Software Ecosystems&#39;</p>

openJun 2023View details →
zenodo28/100

Japanese Sample Tweets, COVID-19 Keywords and Emotions from 2020-01-01 to 2020-06-30 (88,495,817 tweets and 47,539,139 retweets)

<p><strong>Data</strong></p> <p>Tweets_YYYY-MM.tsv.gz:<br> The first column is the tweet id, the second column is the date and time (JST) when the tweet was posted, the third column is the tweet id of the mention destination, the fourth column is the tweet id of the retweet source, the fifth column is the place id, the sixth column is the country code, the seventh column is the prefecture code if the country code is JP, and the eighth column is the COVID-19-related keyword included in the tweet. Columns with no information are empty. For example, a tweet with an empty eighth column is not a COVID-19-related tweet.<br> This data was collected using statuses/sample of the Twitter Streaming API, narrowed down by language=ja. Therefore, most of the tweets are Japanese tweets. Also, due to a failure of the data collection server, a large number of tweets on January 22 are missing :(<br> We have used 肺炎, コロナ and COVID (case insensitive) as keywords related to COVID-19.</p> <p>Emotions_YYYY-MM.tsv.gz:<br> The first column is the tweet id, the second and subsequent columns are the number of occurrences of each emotional keyword. Column names (types of emotion) are shown in the first row.<br> We used <a href="http://arakilab.media.eng.hokudai.ac.jp/~ptaszynski/repository/mlask.htm">mlask43-simple</a> (Perl implementation of <a href="http://doi.org/10.5334/jors.149">ML-Ask</a>) with dictionaries used in <a href="https://github.com/ikegami-yukino/pymlask">pymlask</a> to extract emotional keywords from the tweet.</p> <p><strong>Publication</strong></p> <p>This data set was created for my study. If you make use of this data set, please cite:<br> Mitsuo Yoshida. The State of Social Media During the COVID-19 Pandemic: Japan&#39;s Situation, Research Trends and Public Datasets. <em>Journal of Japanese Society for Artificial Intelligence (in Japanese)</em>. vol.35, no.5, pp.644-653, 2020.<br> 吉田光男. <a href="http://id.nii.ac.jp/1004/00010708/">COVID-19流行下におけるソーシャルメディア ―日本での状況と研究動向・公開データセット―</a>. <em>人工知能</em>. vol.35, no.5, pp.644-653, 2020.</p>

opencc-zeroAug 2020View details →
zenodo28/100

Keywords derived from Publications for SMS "Design Practices in Visualization Driven Data Exploration for Non-Expert Audiences".

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
dryad28/100

Data from: An automated approach to identifying search terms for systematic reviews using keyword co-occurrence networks

1. Systematic review, meta-analysis, and other forms of evidence synthesis are critical to strengthen the evidence base concerning conservation issues and to answer ecological and evolutionary questions. Synthesis lags behind the pace of scientific publishing, however, due to time and resource costs which partial automation of evidence synthesis tasks could reduce. Additionally, current methods of retrieving evidence for synthesis are susceptible to bias towards studies with which researchers are familiar. In fields that lack standardized terminology encoded in an ontology, including ecology and evolution, research teams can unintentionally exclude articles from the review by omitting synonymous phrases in their search terms. 2. To combat these problems, we developed a quick, objective, reproducible method for generating search strategies that uses text mining and keyword co-occurrence networks to identify the most important terms for a review. The method reduces bias in search strategy development because it does not rely on a predetermined set of articles and can improve search recall by identifying synonymous terms that research teams might otherwise omit. 3. When tested against the search strategies used in published environmental systematic reviews, our method performs as well as the published searches and retrieves gold-standard hits that replicated versions of the original searches do not. Because the method is quasi-automated, the amount of time required to develop a search strategy, conduct searches, and assemble results is reduced from approximately 17-34 hours to under 2 hours. 4. To facilitate use of the method for environmental evidence synthesis, we implemented the method in the R package litsearchr, which also contains a suite of functions to improve efficiency of systematic reviews by automatically deduplicating and assembling results from separate databases.

opencc-zeroJul 2019View details →
zenodo28/100

Sankey diagram - Auths - Keywords – Journals

<p>Repository of image, part of PhD Thesis in Management.</p>

opencc-by-4.0Nov 2022View details →
dryad28/100

Data from: An automated approach to identifying search terms for systematic reviews using keyword co-occurrence networks

Open the record for dataset details and reuse information.

publicJul 2019View details →
zenodo24/100

Tweets containing the keyword 'bucha' during the Russian invasion of Ukraine

<p>Comprehensive dataset of Tweets containing the keyword 'bucha' around the Ukraine Invasion in February 2022.</p><p>The user handle column has been excluded to protect deleted accounts that have not been retweeted or replied to.</p><p>Tweets have been collected via the Academic API using the Search endpoint in four languages:</p><p>&nbsp;</p><p><strong>Languages</strong></p><p><i>English</i></p><p>search term: "bucha AND lang='en'"</p><p><i>German</i></p><p>search term: "(bucha OR butscha) AND lang='de'"</p><p><i>Russian</i></p><p>search term: (Бу́ча OR bucha) AND lang:ru</p><p><i>Ukrainian</i></p><p>search term: (Бу́ча OR bucha) AND lang:uk</p><p>&nbsp;</p><p><strong>Timeframe</strong></p><p>Data starts 1. March 2022 and ends on 27. May 2023.</p><p>&nbsp;</p><p><strong>Collection dates</strong></p><p>Details on collection dates per Tweet (e.g. to compare with creation dates) as well as the IDs of Tweets for consistency checks can be found here: <a href="https://github.com/Leibniz-HBI/ukraine_twitter_data">https://github.com/Leibniz-HBI/ukraine_twitter_data</a> (<a href="https://doi.org/10.17605/OSF.IO/RTQXN">https://doi.org/10.17605/OSF.IO/RTQXN</a>)</p>

restrictedNov 2023View details →
zenodo24/100

KWX: scholarly articles with keywords from arXiv

<p>A dataset based on&nbsp;<a href="https://www.kaggle.com/datasets/Cornell-University/arxiv">arXiv Dataset</a>, where keywords of papers are added in this dataset.</p>

opencc-by-4.0May 2023View details →
zenodo24/100

Video of the metadata matrix: year navigation by keyword, from webQDA.

<p>Part of the thesis of Sonia Verdugo Castro (USAL, Spain).&nbsp;<br> Title of the article: &quot;The gender gap in STEM higher education studies: visualisation of literature&quot;.<br> Authors of the publication: Sonia Verdugo-Castro, M&ordf; Cruz S&aacute;nchez-G&oacute;mez, Alicia Garc&iacute;a-Holgado, Francisco J. Garc&iacute;a-Pe&ntilde;alvo</p>

opencc-by-4.0Oct 2021View details →
zenodo20/100

Keywords Analyse und Research von ranklike SEO

<u>Source</u>: Flickr <br><u>4DCity URL</u>: <a href="https://4dcity.org/imgupload/1677443821.6646.jpg">https://4dcity.org/imgupload/1677443821.6646.jpg</a> <br><u>Original Image URL</u>: <a href="https://live.staticflickr.com/65535/52283387967_d05ee4e218_m.jpg">https://live.staticflickr.com/65535/52283387967_d05ee4e218_m.jpg</a>

restrictedFeb 2023View details →
nasa20/100

KEYWORD SEARCH IN TEXT CUBE: FINDING TOP-K RELEVANT CELLS

KEYWORD SEARCH IN TEXT CUBE: FINDING TOP-K RELEVANT CELLS BOLIN DING*, YINTAO YU*, BO ZHAO*, CINDY XIDE LIN*, JIAWEI HAN*, AND CHENGXIANG ZHAI* Abstract. We study the problem of keyword search in a data cube with text-rich dimension(s) (so-called text cube). The text cube is built on a multidimensional text database, where each row is associated with some text data (e.g., a document) and other structural dimensions (attributes). A cell in the text cube aggregates a set of documents with matching attribute values in a subset of dimensions. A cell document is the concatenation of all documents in a cell. Given a keyword query, our goal is to find the top-k most relevant cells (ranked according to the relevance scores of cell documents w.r.t. the given query) in the text cube. We define a keyword-based query language and apply IR-style relevance model for scoring and ranking cell documents in the text cube. We propose two efficient approaches to find the top-k answers. The proposed approaches support a general class of IR-style relevance scoring formulas that satisfy certain basic and common properties. One of them uses more time for pre-processing and less time for answering online queries; and the other one is more efficient in pre-processing and consumes more time for online queries. Experimental studies on the ASRS dataset are conducted to verify the efficiency and effectiveness of the proposed approaches.

restrictednotspecifiedApr 2025View details →
nasa20/100

Efficient Keyword-Based Search for Top-K Cells in Text Cube

Previous studies on supporting free-form keyword queries over RDBMSs provide users with linked-structures (e.g.,a set of joined tuples) that are relevant to a given keyword query. Most of them focus on ranking individual tuples from one table or joins of multiple tables containing a set of keywords. In this paper, we study the problem of keyword search in a data cube with text-rich dimension(s) (so-called text cube). The text cube is built on a multidimensional text database, where each row is associated with some text data (a document) and other structural dimensions (attributes). A cell in the text cube aggregates a set of documents with matching attribute values in a subset of dimensions. We define a keyword-based query language and an IR-style relevance model for coring/ranking cells in the text cube. Given a keyword query, our goal is to find the top-k most relevant cells. We propose four approaches, inverted-index one-scan, document sorted-scan, bottom-up dynamic programming, and search-space ordering. The search-space ordering algorithm explores only a small portion of the text cube for finding the top-k answers, and enables early termination. Extensive experimental studies are conducted to verify the effectiveness and efficiency of the proposed approaches. Citation: B. Ding, B. Zhao, C. X. Lin, J. Han, C. Zhai, A. N. Srivastava, and N. C. Oza, “Efficient Keyword-Based Search for Top-K Cells in Text Cube,” IEEE Transactions on Knowledge and Data Engineering, 2011.

restrictednotspecifiedMar 2025View details →
zenodo16/100

Movies with mention to LGBTQ+ on plot-keywords [1909-2019]

<p>IMDb was the only source from which data was extracted. The sample was constructed using the tool search engines, filtering &ldquo;feature films&rdquo; (over 45 min of lentgh), excluding &ldquo;adult titles&rdquo; and excluding &ldquo;released&rdquo;. In order to obtain a comprehensive length of time, all productions from 1895 to December 2019 were included according to a search carried out in March 2020. To collect only those films that could have detailed information, the number of items was limited to those with over 50 user ratings (N = 119809) as a way of minimally controlling the popularity of the published work.</p> <p>In the resulting sample (N = 1768), the presence of descriptors in the field &ldquo;plot&rdquo; was coded, designating different terms related to the LGBTQ+ community. Those included the following terms and their variants with similar etymology: homosexual (homo/homosexuality), gay, lesbian, trans (transsexual, transgender) and queer. After adding those productions that had a descriptor term from this list in the keywords field, the total of the items corresponding to these categories was 9409 films. This decision responds to the need to reflect in the sample films in which there is representation of the group, but it is not necessarily part of the plot or is revealed through the course of it. It is also supported by the documentary tradition, according to which keywords tend to overlap with the plot or summary (La Barre &amp; de Novais Cordeiro, 2012, p. 241).&nbsp;</p> <p>The following information was extracted from each item (movie):</p> <p>&bull;&nbsp;&nbsp;&nbsp;&nbsp; Production by country and production by language in each year. In co-productions, only the first producing country was considered and the other countries discarded.</p> <p>&bull;&nbsp;&nbsp;&nbsp;&nbsp; Identity of the group (gay, lesbian, trans, etc.) as per the plot keywords. Several identities and expressions such as transsexuality and transgender have been included under the label &ldquo;trans&rdquo; as it was impossible to recognize the correct expression from the labels provided by IMDb.</p> <p>&bull;&nbsp;&nbsp;&nbsp;&nbsp; Cinema genres. Using the first two descriptors, a list of genre pairs was created, which were later grouped into 13 generic categories, as per the formal qualities of the theme: Drama (any combinations of the drama category that were not included in other categories), Comedy (combinations including comedy that were not considered in other categories), Action/Adventure, Melodrama (Drama + Comedy), Horror, Crime (including thrillers), Fantasy (including Science Fiction), Biography (in both fictional and documentary forms), Documentary (excluding biographies and fake documentaries but including News), Romantic Comedy (Romance + Comedy), Animation (excluding documentary formats) and Music/Musical feature films.</p> <p>&bull;&nbsp;&nbsp;&nbsp;&nbsp; Age ratings. Parental Advisory guidelines have changed significantly over the decades, from the first classifications in the United Kingdom, Germany, and the United States to today. In order to establish a suitable comparison, the descriptor provided by IMDb, which is usually established by the MPAA (Motion Picture Association of America), was used. When this was omitted, the descriptor used was determined as per the recommended age: Universal, Parental Guidance (PG), 12-13, 14-16, 17-18, as well as X and banned, according to the historical equivalence as provided by IMDd (2020).</p> <p>Data was pooled to consider evolution by historical periods and trends in a single group or correlational ex post facto design. Subsequently, the data was analyzed with the statistical package SPSS v.26. To visualize the main trends from the data, Tableau 2020 software was used.</p>

restrictedMay 2020View details →
zenodo12/100

Co-Occurence Keyword Network

<p>A bibliometric map&nbsp;displaying&nbsp;various nodes (word items) and the links by which they are connected conceived on VOSviewer.&nbsp;In this publication, we focused on words indicative of deficit thinking (abnormality, behavior problem, cognitive deficit, cognitive impairment, deficit, disorder, impairment, language difficulty, problem, risk, risk factor and attention problem). This map can be viewed and explored by uploading the attached json file on app.vosviewer.com. Readers can&nbsp;evaluate specific nodes by using the &quot;find&quot; feature on the left-handside menu.&nbsp;</p> <p>&nbsp;</p>

restrictedMay 2023View details →
zenodo12/100

Estudos incluídos : Keywords

<p>Palavras chave dos estudos inclu&iacute;dos no projeto de disserta&ccedil;&atilde;o:&nbsp;<strong>Modelos de avalia&ccedil;&atilde;o de compet&ecirc;ncias informacionais em contexto acad&eacute;mico: Uma revis&atilde;o sistem&aacute;tica da literatura</strong></p>

restrictedJul 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record