Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3
datasets available to search
ShareScore release 0.9.0
Dataset results
3 results for “query suggestions”
btw17 query auto completion - query suggestions for German politicians and parties before the federal election 2017
<p>The dataset contains the query suggestions for 5 major German parties (terms: "afd", "csu", "dielinke", "fdp", "grüne", "spd") and ten popular politicians and party leaders (terms: "Alexander Gauland", "Alice Weidel", "Angela Merkel", "Cem Özdemir", "Christian Lindner", "Dietmar Bartsch", "Katrin Göring-Eckardt", "Martin Schulz", "Sahra Wagenknecht").</p> <p>The data was crawled on (mostly) two times per day from Tue Aug 04, 2017 to Tue Oct 31, 2017. The dataset contains 20001 suggestions from Bing search (http://api.bing.net/osjson.aspx), 11935 suggestions from Duck-Duck-Go (https://duckduckgo.com/ac/) and 33521 suggestions from Google search (http://clients1.google.de/complete/search). Note, that for some terms and dates no suggestions were returned by some of the APIs.</p> <p>German language settings were used for Google and Bing, English language setting was used for Duck-Duck-Go. The API requests were sent with an IP address from Cologne, Germany. </p> <p>The UTF-8 encoded comma separated text file contains the following columns:</p> <p><source>: google, bing or ddg</p> <p><queryterm>: the query term</p> <p><date>: the date and time of the API call formatted as ISO8601</p> <p><suggestterm>: the suggested query completion (the query term was removed from the suggestion)</p> <p><position>: the position of the query suggestion within the list returned by the API (ranges from 0 to 19)</p> <p> </p> <p> </p> <p><br> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p>
Qbias – A Dataset on Media Bias in Search Queries and Query Suggestions
<p>We present Qbias, two novel datasets that promote the investigation of bias in online news search as described in</p> <blockquote> <p>Fabian Haak and Philipp Schaer. 2023. 𝑄𝑏𝑖𝑎𝑠 - A Dataset on Media Bias in Search Queries and Query Suggestions. In Proceedings of ACM Web Science Conference (WebSci’23). ACM, New York, NY, USA, 6 pages. <a href="https://doi.org/10.1145/3578503.3583628">https://doi.org/10.1145/3578503.3583628</a>.</p> </blockquote> <p><strong>Dataset 1: AllSides Balanced News Dataset (allsides_balanced_news_headlines-texts.csv)</strong></p> <p>The dataset contains 21,747 news articles collected from <a href="https://www.allsides.com/headline-roundups">AllSides balanced news headline</a> roundups in November 2022 as presented in our publication. The AllSides balanced news feature three expert-selected U.S. news articles from sources of different political views (left, right, center), often featuring spin bias, and slant other forms of non-neutral reporting on political news. All articles are tagged with a bias label by four expert annotators based on the expressed political partisanship, left, right, or neutral. The AllSides balanced news aims to offer multiple political perspectives on important news stories, educate users on biases, and provide multiple viewpoints. Collected data further includes headlines, dates, news texts, topic tags (e.g., "Republican party", "coronavirus", "federal jobs"), and the publishing news outlet. We also include AllSides' neutral description of the topic of the articles.<br> Overall, the dataset contains 10,273 articles tagged as left, 7,222 as right, and 4,252 as center.</p> <p>To provide easier access to the most recent and complete version of the dataset for future research, we provide a scraping tool and a regularly updated version of the dataset at <a href="https://github.com/irgroup/Qbias">https://github.com/irgroup/Qbias</a>. The repository also contains regularly updated more recent versions of the dataset with additional tags (such as the URL to the article). We chose to publish the version used for fine-tuning the models on Zenodo to enable the reproduction of the results of our study. </p> <p> </p> <p><strong>Dataset 2: Search Query Suggestions (suggestions.csv)</strong></p> <p>The second dataset we provide consists of 671,669 search query suggestions for root queries based on tags of the AllSides biased news dataset. We collected search query suggestions from Google and Bing for the 1,431 topic tags, that have been used for tagging AllSides news at least five times, approximately half of the total number of topics. The topic tags include names, a wide range of political terms, agendas, and topics (e.g., "communism", "libertarian party", "same-sex marriage"), cultural and religious terms (e.g., "Ramadan", "pope Francis"), locations and other news-relevant terms. On average, the dataset contains 469 search queries for each topic. In total, 318,185 suggestions have been retrieved from Google and 353,484 from Bing.</p> <p>The file contains a "root_term" column based on the AllSides topic tags. The "query_input" column contains the search term submitted to the search engine ("search_engine"). "query_suggestion" and "rank" represents the search query suggestions at the respective positions returned by the search engines at the given time of search "datetime". We scraped our data from a US server saved in "location".</p> <p>We retrieved ten search query suggestions provided by the Google and Bing search autocomplete systems for the input of each of these root queries, without performing a search. Furthermore, we extended the root queries by the letters a to z (e.g., "democrats" (root term) >> "democrats a" (query input) >> "democrats and recession" (query suggestion)) to simulate a user's input during information search and generate a total of up to 270 query suggestions per topic and search engine. The dataset we provide contains columns for root term, query input, and query suggestion for each suggested query. The location from which the search is performed is the location of the Google servers running Colab, in our case Iowa in the United States of America, which is added to the dataset. </p> <p><strong>AllSides Scraper</strong></p> <p>At <a href="https://github.com/irgroup/Qbias">https://github.com/irgroup/Qbias</a>, we provide a scraping tool, that allows for the automatic retrieval of all available articles at the AllSides balanced news headlines. </p> <p>We want to provide an easy means of retrieving the news and all corresponding information. For many tasks it is relevant to have the most recent documents available. Thus, we provide this Python-based scraper, that scrapes all available AllSides news articles and gathers available information. By providing the scraper we facilitate access to a recent version of the dataset for other researchers.</p> <p> </p>
query suggestion with siri - query suggestions for instagram accounts
<p>Das Datenset enthält Query Suggestions zu 20 Instragram-Accounts bzw. Personen oder Firmen dahinter. (Instagram, Cristiano Ronaldo, Ariana Grande, Selena Gomez, The Rock, Kim Kardashian West, Kylie Jenner, Beyoncé, Taylor Swift, Leo Messi, Neymar jr, Kendall Jenner, Justin Bieber, National Geographic, Barbie, Khloe Kardashian, Jennifer Lopez, Miley Cyrus, Nike, Katy Perry)</p> <p>Die Datenerhebung erfolgte vom 26.05.2019 - 23.06.2019. Enthalten sind Vorschläge der Suchmaschinen Bing, DuckDuckGo, Google und der Suche in Siri (an einem Tablet und in einer XCode Simulation), wobei die Anzahl der zurückgelieferten Vorschläge variiert.</p> <p>Das Datenset enthält folgende Spalten: Plattform, Suchbegriff, Datum, Vorschlag Position</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.