Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
26
datasets available to search
ShareScore release 0.7.1
Dataset results
26 results for “web search”
User study data: Nudges to Mitigate Confirmation Bias during Web Search for Opinion Formation, automatic vs. reflective study
<p>Data of two user studies (282 and 307 participants), investigating the risks and benefits of warning labels with and without obfuscations to mitigate confirmation bias during web search on debated topics.</p> <p> </p> <p>Study Variables (study 1 and study 2)</p> <p> </p> <p> display_con: Search result display<br> - Study 1<br> - 1: targeted warning label with obfuscation<br> - 2: random warning label with obfuscation<br> - 3: regular (no intervention)<br> - Study 2<br> - 1: targeted warning label with obfuscation<br> - 2: targeted warning label without obfuscation<br> - 3: random warning label with obfuscation<br> - 4: random warning label without obfuscation<br> - 5: regular (no intervention)<br>- CRT_cat: Cognitive reflection<br> - 1: intuitive<br> - 2: analytic<br>- topic: Assigned debated topic<br> - 1: Is drinking milk healthy for humans? <br> - 2: Is homework beneficial?<br> - 3: Should people become vegetarian?<br> - 4: Should students have to wear school uniforms?<br>- clicksup_prop: Clicks on attitude-confirming (AC) search results (proportion of all clicks)<br>- clickwarn_prop: Clicks on warning label (WL) search results (proportion of all clicks)<br>- show_clicked: Clicks on show-button (number of clicks, only in conditions with obfuscation)<br>- accuracy_bias: Accuracy bias estimation (Difference between a) observed bias (as the proportion of attitude-confirming clicks) and b) perceived bias (reported in the post-interaction questionnaire and re-coded into values from 0 to 1), positive values indicate an overestimation of bias)<br>- att_change: Attitude change (Difference between attitude reported in the pre-interaction questionnaire and the post-interaction questionnaire. Negative values indicate an attitude change in the attitude-opposing direction, while positive values indicate an attitude strengthening in the attitude-supporting direction.)<br>- knowledge_1: Self-reported prior knowledge (Reported on a seven-point Likert scale ranging from non-existent to excellent as a response to how they would describe their knowledge on the topic they were assigned to)<br>- N_clicks: Cumulative clicks (Number of all clicks on search results)<br>- NFC: Need for Cognition (Mean response to 4-item subset of the NFC questionnaire)<br>- UX_usability: Usability (Mean of responses on a seven-point Likert scale to the module "usability"from the meCUE 2.0 questionnaire)<br>- UX_usefulness: Usefulness (Mean of responses on a seven-point Likert scale to the module "usefulness"from the meCUE 2.0 questionnaire)</p>
Code and data associated with: Searching the web builds fuller picture of arachnid trade
<p>Data and code used in the paper: Searching the web builds fuller picture of arachnid trade. Throughout the methods we have indicated the stage of analysis each data component was used and the code script connected. We have numbered to code and data supplements to reflect as closely as possible the order in which data generation and summary was undertaken. The following provide additional details linked to each of the data files.</p> <p>Data S1 - Website data: lang = language of the search engine used, ad hoc websites had language described after discovery; engine = the search engine used; page = the page on which the website appeared from the search engine; searchdate = search date in YYYY-mm-dd HH:MM:SS; link = link to the webpage, redacted to protect website identity; reviewdate = date revewied for arachnids being sold and search strategy; sells = whether the website sells arachnids (1 == sells); allow = whether the site explcicilt forbids automated searching (1 == allows, NA when search method was not fully automated, e.g., single page); type = the type of the website (e.g., trade, classified ads); order = whether arachnids where organised in a particular ways; target = a refined target URL to start search; method = the search method chosen, see methods for details; refine = any refinement or filter than could constrain the scope of the website to be searched; spages = the number of pages required to cycle through to cover the entire stock (also separated by ; if multiple cycles where needed or multiple single pages could be easily collected); prelimCheck = whether the website passed initial checks for arachnid selling; notes = any details that might need special attention during searches; webID = code used for subsequent data summary.</p> <p>Data S2 - Raw keyword searches outputs: species keywords. sp = the modern species or genus that a keyword is associated with; page = the number of the page the keyword was detected on; keyw = the exact keyword that was detected; spORgen = whether the keyword was a species binomial or just genus; termsSurrounding = the words surrounding a genus keyword detection (only applies to Data S3); webID = the website ID.</p> <p>Data S3 – Raw keyword searches outputs: genus keywords. sp = the modern species or genus that a keyword is associated with; page = the number of the page the keyword was detected on; keyw = the exact keyword that was detected; spORgen = whether the keyword was a species binomial or just genus; termsSurrounding = the words surrounding a genus keyword detection (multiple detections separated by ;); webID = the website ID.</p> <p>Data S4 - Raw keyword search outputs: temporal sample. sp = the modern species or genus that a keyword is associated with; page = the number of the page the keyword was detected on; keyw = the exact keyword that was detected; spORgen = whether the keyword was a species binomial or just genus; termsSurrounding = the words surrounding a genus keyword detection (multiple detections separated by ;); webID = the website ID; timestamp.parse = the timestamp extracted from the archived web page; year = a simplified timestamp including only the year.</p> <p>Data S5 - LEMIS data used. An arachnid filtered version of <sup>74,75</sup>.</p> <p>Data S6 - CITES trade database data used <sup>76</sup>.</p> <p>Data S7 - CITES appendices data used <sup>77</sup>.</p> <p>Data S8 - IUCN Redlist data used <sup>78</sup>.</p> <p>Data S9 - Compiled final dataset, with data deriving from WSC, Scorpion files, ITIS, WAM and the data collection process. speciesId = a numeric code, one per species; clade = the clade the species belongs to; family = the family the species belongs to; genus = the genus of the species; species = the species epithet; author = the species authority name; year = the species authority year; parentheses = whether parentheses are needed with the authority; distribution = WSC original distribution descriptions; invalid = whether the species is considered valid; source = the species source, either World Spider Catalogue, Scorpion files, ITIS or WAM; accName = the species binomial being used as our accepted name; allNames = the accepted species binomial and all synonyms; allGenera = the accepted genus, and all other genera the species has belonged to at one point; onlineTradeSnap = whether the species was detected via a match to the accName in the snapshot data; onlineTradeSnap_Any = whether the species was detected via any synonym in the snapshot data; onlineTradeSnap_genus = whether the genus was detected via a match to the genus in the snapshot data; onlineTradeSnap_genusAny = whether the genus was detected via any synonym in the snapshot data; onlineTradeTemp = whether the species was detected via a match to the accName in the temporal data; onlineTradeTemp_Any = whether the species was detected via any synonym in the temporal data; onlineTradeTemp_genus = whether the genus was detected via a match to the genus in the temporal data; onlineTradeTemp_genusAny = whether the genus was detected via any synonym in the temporal data; onlineTradeEither = whether the species was detected via a match to the accName in the temporal data or snapshot data; onlineTradeEither_Any = whether the species was detected via any synonym in the temporal data or snapshot data; LEMIStrade = whether the species was detected via a match to the accName in the LEMIS data; LEMIStrade_Any = whether the species was detected via any synonym in the LEMIS data; LEMIStrade_genus = whether the genus was detected via any synonym in the LEMIS data; LEMIStrade_genusAny = whether the genus was detected via any synonym in the LEMIS data; CITEStrade = whether the species was detected via a match to the accName in the CITES trade database data; CITEStrade_Any = whether the species was detected via any synonym in the CITES trade database data; CITEStrade_genus = whether the genus was detected via any synonym in the CITES trade database data; CITEStrade_genusAny = whether the genus was detected via any synonym in the CITES trade database data; CITESapp = the CITES appendix the species is listed under using an exact match to the accName; CITESapp_Any = the CITES appendix the species is listed under using any match to any of the species’ synonyms; redlist = the IUCN Redlist category the species is listed under using an exact match to the accName; redlist_Any = the IUCN Redlist category the species is listed under using any match to any of the species’ synonyms; extactMatchTraded = the species is detected in any of the trade sources via a match to the accName; anyMatchTraded = the species is detected in any of the trade sources via a match to any species’ synonym.</p> <p>Data S10 - Forum listings of “What species are you currently keeping” from an online fora posted between 9th September 2021 and 9th October 2021, to provide an idea of online discussions. Each user with a separate list is provided in a separate tab. Morph_collector is the same as poster1, but the potential cryptic species or morphs are noted separately to make them clearer.</p> <p>Data S11 – Distribution information for spiders. Only two columns used in summaries: accName = the accepted name used throughout summaries; NAME = the country name the spider occurs in.</p> <p>Data S12 - Distribution information for scorpions. species = the accepted name used throughout summaries; NAME = the country name the scorpions occurs in.</p> <p>Code S1 - Search URL Extract.R</p> <p>Code S2 - Retrieve web data.R</p> <p>Code S3 - Temporal Classified Ads.R</p> <p>Code S4 - Keyword Generation.R</p> <p>Code S5 - Keyword Search.R</p> <p>Code S6 - LEMIS filter and summary.R</p> <p>Code S7 - Compiling results.R</p> <p>Code S8 - Summary Figures.R</p> <p>Code S9 - Temporal Figures.R</p> <p>Code S10 - New description figure.R</p> <p>Code S11 - Term exploration.R</p> <p>Code S12 - LEMIS summary and mapping.R</p>
PLAE web app enables powerful searching and multiple visualizations across one million unified single-cell ocular transcriptomes
<p>Supplementary Data for "PLAE web app enables powerful searching and multiple visualizations across one million unified single-cell ocular transcriptomes"</p> <p> </p>
Data set of the article: Using Machine Learning for Web Page Classification in Search Engine Optimization
<p>Data of investigation published in the article: "Using Machine Learning for Web Page Classification in Search Engine Optimization"</p> <p>Abstract of the article:</p> <p>This paper presents a novel approach of using machine learning algorithms based on experts’ knowledge to classify web pages into three predefined classes according to the degree of content adjustment to the search engine optimization (SEO) recommendations. In this study, classifiers were built and trained to classify an unknown sample (web page) into one of the three predefined classes and to identify important factors that affect the degree of page adjustment. The data in the training set are manually labeled by domain experts. The experimental results show that machine learning can be used for predicting the degree of adjustment of web pages to the SEO recommendations—classifier accuracy ranges from 54.59% to 69.67%, which is higher than the baseline accuracy of classification of samples in the majority class (48.83%). Practical significance of the proposed approach is in providing the core for building software agents and expert systems to automatically detect web pages, or parts of web pages, that need improvement to comply with the SEO guidelines and, therefore, potentially gain higher rankings by search engines. Also, the results of this study contribute to the field of detecting optimal values of ranking factors that search engines use to rank web pages. Experiments in this paper suggest that important factors to be taken into consideration when preparing a web page are page title, meta description, H1 tag (heading), and body text—which is aligned with the findings of previous research. Another result of this research is a new data set of manually labeled web pages that can be used in further research. </p>
Disentangling Web Search on Debated Topics - User Study Data
<p>Data of an exploratory, open-ended user study (N = 255) to advance knowledge and uncover relations between the different facets of web search on debated topics. We explored the relations between factors inherent to the searcher and search system (user characteristics, exposure bias), search intercations (confirmation bias, position bias, search effort), and post-search epistemic states (attitude change, knowledge gain). This data set contains the following variables for each of the 255 participants: SERP ranking bias, prior knowledge, attitude strength, receptiveness to opposing views, attitude-confirming clicks, click rank deviation, number of clicks, time on SERP, hover depth, attitude change, knowledge gain.</p>
Confirmation Bias in Web-Based Search: A Randomized Online Study on the Effects of Expert Information and Social Tags on Information Search and Evaluation
<p>ABSTRACT</p> <p>Background: The public typically believes psychotherapy to be more effective than pharmacotherapy for depression treatments. This is not consistent with current scientific evidence, which shows that both types of treatment are about equally effective.</p> <p>Objective: The study investigates whether this bias towards psychotherapy guides online information search and whether the bias can be reduced by explicitly providing expert information (in a blog entry) and by providing tag clouds that implicitly reveal experts’ evaluations.</p> <p>Methods: A total of 174 participants completed a fully automated Web-based study after we invited them via mailing lists. First, participants read two blog posts by experts that either challenged or supported the bias towards psychotherapy. Subsequently, participants searched for information about depression treatment in an online environment that provided more experts’ blog posts about the effectiveness of treatments based on alleged research findings. These blogs were organized in a tag cloud; both psychotherapy tags and pharmacotherapy tags were popular. We measured tag and blog post selection, efficacy ratings of the presented treatments, and participants’ treatment recommendation after information search.</p> <p>Results: Participants demonstrated a clear bias towards psychotherapy (mean 4.53, SD 1.99) compared to pharmacotherapy (mean 2.73, SD 2.41; <em>t</em><sub>173</sub>=7.67, <em>P</em><.001, <em>d</em>=0.81) when rating treatment efficacy prior to the experiment. Accordingly, participants exhibited biased information search and evaluation. This bias was significantly reduced, however, when participants were exposed to tag clouds with challenging popular tags. Participants facing popular tags challenging their bias (n=61) showed significantly less biased tag selection (<em>F</em><sub>2,168</sub>=10.61, <em>P</em><.001, partial eta squared=0.112), blog post selection (<em>F</em><sub>2,168</sub>=6.55, <em>P</em>=.002, partial eta squared=0.072), and treatment efficacy ratings (<em>F</em><sub>2,168</sub>=8.48, <em>P</em><.001, partial eta squared=0.092), compared to bias-supporting tag clouds (n=56) and balanced tag clouds (n=57). Challenging (n=93) explicit expert information as presented in blog posts, compared to supporting expert information (n=81), decreased the bias in information search with regard to blog post selection (<em>F</em><sub>1,168</sub>=4.32, <em>P</em>=.04, partial eta squared=0.025). No significant effects were found for treatment recommendation (<em>P</em>s>.33).</p> <p>Conclusions: We conclude that the psychotherapy bias is most effectively attenuated—and even eliminated—when popular tags implicitly point to blog posts that challenge the widespread view. Explicit expert information (in a blog entry) was less successful in reducing biased information search and evaluation. Since tag clouds have the potential to counter biased information processing, we recommend their insertion.</p>
Bird predation on Roseau cane scale as revealed by a web image search and querying a citizen monitoring database
<p>NA</p>
Supplementary Material for Nested Segmentation of Web Search Queries
<p>1. Query test of SGCL12 [SGCL12QueryTestSet.txt]</p> <p>2. Outputs of 16 nesting algorithms on the query test set of SGCL12 [Input flat segmention: Saha Roy et al., SIGIR 2012] [Nested Segmentation Outputs SGCL12.zip]</p> <p>3. Query set of TREC-WT [TREC-WTQueryTestSet.txt]</p> <p>4. Outputs of 16 nesting algorithms on the query test set of TREC-WT [Input flat segmention: Saha Roy et al., SIGIR 2012] [Nested Segmentation Outputs TREC-WT.zip]</p> <p>5. Code and executables for generating the nested segmentations for a set of queries [Code and execs for generating nested segmentations.zip]</p> <p>6. Code and executables for IR-based evaluation of a nested segmentation [Code and execs for IR evaluation of a nested segmentation.zip]</p> <p>7. Readme.txt</p>
A Novel Algorithm for Estimating Web Page Ranking in Search Engine Results Pages
<p><em><strong>Abstract:</strong> </em>Search engine optimization (SEO) can make a big improvement in the traffic to a web page. Because search engines keep their main rules of ranking undeclared, it’s important to develop models that can estimate the ranking of a web page in the search engine to be able to optimize web pages to rank higher in the search engine. The available research methodologies used machine learning algorithms to provide solutions for this target with the help of generated datasets by scraping the search engine results pages (SERP) and crawling web pages. Their proposed models suffered from the inability to be updated dynamically if the search engine updated its ranking algorithm, and their input data did not include the diversity of web pages and languages. This research will propose a novel original rank estimation algorithm that’s able to overcome other research challenges, with a set of comparative experiments and complexity analysis. Results will show that the proposed algorithm could achieve higher values of accuracy, precision, and recall.</p> <p><strong><em>Dataset: </em></strong></p> <p>For research purpose, the dataset will play two roles, first, it will act the role of search engine result pages (SERP), and second, it will be used to test algorithms and calculate performance measurements. Dataset is consisting of 9930 web pages, aimed to identify search results pages, focusing on the top 3 pages of SERP, with 31 extracted attributes that's related to search engine optimization (SEO). The distribution of examples between class labels was balanced, with changes due to scraping operation issues, but not significantly different, with fractions of 39.9%, 34.6%, and 25.5% for the class labels page1, page2, and page 3. Feature names are: 'Title 1 Length', 'Title 2 Length', 'Meta Description 1 Length', 'Meta Description 2 Length', 'Meta Keywords 1 Length', 'H1-1 Length', 'H1-2 Length', 'H2-1 Length', 'H2-2 Length', 'Size (bytes)', 'Word Count', 'Text Ratio', 'Inlinks', 'Unique Inlinks', 'Unique JS Inlinks', '% of Total', 'Outlinks', 'Unique Outlinks', 'Unique JS Outlinks', 'External Outlinks', 'Unique External Outlinks', 'Unique External JS Outlinks', 'Response Time', 'Status Code', 'Keyword in MetaDescription1', 'Keyword in Title1', 'Keyword in MetaKeywords1', 'Keyword in URL', 'Has LastModified', 'Keyword in Headers', and 'Keyword in Emphasized Text'.</p> <p>The process of dataset generation involved scraping the search engine, extracting URLs for selected keywords, focusing on feature extraction, cleaning and preprocessing, and generating new attributes related to keywords in web pages. It involved also removing missing values, duplicates, and data type conversions to obtain a comprehensive dataset.<br> Keyword selection involves selecting keywords from various categories and considering diversity, including high and low traffic, long-term and short-term keywords, and generic and branded keywords. Apify online tool was used for search engine scraping with default language and US country, resulting in 388 selected keywords with 30 results per keyword. Dataset included extracted SEO features from 9991 web pages using screamingFrog desktop software and Rapidminer desktop software, determining page SEO-friendliness and comparing it to SERP rankings. Dataset cleaning involved removing redundant attributes, removing paid SERP results, replacing missing values, and converting data types. Rapidminer was used for data cleaning and preprocessing, generating new attributes related to keyword usage in web pages.<br> </p>
Weathering the storm for love? The mate-searching behaviour of wild male Sydney funnel-web spiders (Atrax robustus)
Open the record for dataset details and reuse information.
Data from: Bird predation on Roseau cane scale as revealed by a web image search and querying a citizen monitoring database
Open the record for dataset details and reuse information.
Data from: How common road salts and organic additives alter freshwater food webs: in search of safer alternatives
The application of deicing road salts began in the 1940s and has increased drastically in regions where snow and ice removal is critical for transportation safety. The most commonly applied road salt is sodium chloride (NaCl). However, the increased costs of NaCl, its negative effects on human health, and the degradation of roadside habitats has driven transportation agencies to seek alternative road salts and organic additives to reduce the application rate of NaCl or increase its effectiveness. Few studies have examined the effects of NaCl in aquatic ecosystems, but none have explored the potential impacts of road salt alternatives or additives on aquatic food webs. We assessed the effects of three road salts (NaCl, MgCl2 and ClearLane™) and two road salts mixed with organic additives (GeoMelt™ and Magic Salt™) on food webs in experimental aquatic communities, with environmentally relevant concentrations, standardized by chloride concentration. We found that NaCl had few effects on aquatic communities. However, the microbial breakdown of organic additives initially reduced dissolved oxygen. Additionally, microbial activity likely transformed unusable phosphorus from the organic additives to usable phosphorus for algae, which increased algal growth. The increase in algal growth led to an increase in zooplankton abundance. Finally, MgCl2 – a common alternative to NaCl – reduced compositional differences of zooplankton, and at low concentrations increased the abundance of amphipods. Synthesis and applications. Our results indicate that alternative road salts (to NaCl), and road salt additives can alter the abundance and composition of organisms in freshwater food webs at multiple trophic levels, even at low concentrations. Consequently, road salt alternatives and additives might alter ecosystem function and ecosystem services. Therefore, transportation agencies should use caution in applying road salt alternatives and additives. A comprehensive investigation of road salt alternatives and road salt additives should be conducted before wide-scale use is implemented. Further research is also needed to determine the impacts of salt additives and alternatives on higher trophic levels, such as amphibians and fish.
1st International Workshop on Open Web Search #wows2024 at ECIR 2024: Document Processors
<p><a href="https://opensearchfoundation.org/en/events-osf/wows2024/#osf-callforcontributions">The First International Workshop on Open Web Search</a> (WOWS) hosted at [ECIR 2024](https://www.ecir2024.org/) aimed to promote and discuss ideas and approaches to open up the web search ecosystem so that small research groups and young startups can leverage the web to foster an open and diverse search market. The workshop had two calls that support collaborative and open web search engines: (1) for scientific contributions, and (2) for open-source implementations. This repository collects the outputs of all submitted document processing components on public datasets for the second call aims to gather open-source prototypes and gain practical experience with collaborative, cooperative evaluation of search engines and their components using the [TIREx Information Retrieval Evaluation Platform](https://www.tira.io/tirex) hosted on [TIRA](https://www.tira.io).</p> <p> </p> <p> </p> <h2>Citations</h2> <p>If you reuse the resources, please ensure to cite TIRA and TIREx and the corresponding datasets, the corresponding bib-entries are:</p> <p>For TIREx:</p> <pre><code>@InProceedings{froebe:2023e,<br> author = {Maik Fr{\"o}be and {Jan Heinrich} Reimer and Sean MacAvaney and Niklas Deckers and Simon Reich and Janek Bevendorff and Benno Stein and Matthias Hagen and Martin Potthast},<br> booktitle = {46th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2023)},<br> doi = {10.1145/3539618.3591888},<br> editor = {Hsin{-}Hsi Chen and Wei{-}Jou (Edward) Duh and Hen{-}Hsen Huang and Makoto P. Kato and Josiane Mothe and Barbara Poblete},<br> ids = {potthast:2023t},<br> isbn = {9781450394086},<br> month = jul,<br> numpages = 11,<br> pages = {2826--2836},<br> publisher = {ACM},<br> site = {Taipei, Taiwan},<br> title = {{The Information Retrieval Experiment Platform}},<br> url = {https://dl.acm.org/doi/10.1145/3539618.3591888},<br> year = 2023<br>}<br></code></pre> <p>for TIRA:</p> <pre><code>@InProceedings{froebe:2023b,<br> address = {Berlin Heidelberg New York},<br> author = {Maik Fr{\"o}be and Matti Wiegmann and Nikolay Kolyada and Bastian Grahm and Theresa Elstner and Frank Loebe and Matthias Hagen and Benno Stein and Martin Potthast},<br> booktitle = {Advances in Information Retrieval. 45th European Conference on {IR} Research ({ECIR} 2023)},<br> doi = {10.1007/978-3-031-28241-6_20},<br> editor = {Jaap Kamps and Lorraine Goeuriot and Fabio Crestani and Maria Maistro and Hideo Joho and Brian Davis and Cathal Gurrin and Udo Kruschwitz and Annalina Caputo},<br> ids = {potthast:2023h},<br> month = apr,<br> pages = {236--241},<br> publisher = {Springer},<br> series = {Lecture Notes in Computer Science},<br> site = {Dublin, Irland},<br> title = {{Continuous Integration for Reproducible Shared Tasks with TIRA.io}},<br> url = {https://link.springer.com/chapter/10.1007/978-3-031-28241-6_20},<br> year = 2023<br>}</code><br><br>All query processors are described in the corresponding WOWS paper, please cite the papers and underlying approaches accordingly.</pre> <p>Forthermore, please cite the datasets that you use.</p> <h3>Args.me</h3> <p>If you re-use the Args.me indices, please additionally cite:</p> <pre><code>@InProceedings{bondarenko:2021d,<br> address = {Berlin Heidelberg New York},<br> author = {Alexander Bondarenko and Lukas Gienapp and Maik Fr{\"o}be and Meriem Beloucif and Yamen Ajjour and Alexander Panchenko and Chris Biemann and Benno Stein and Henning Wachsmuth and Martin Potthast and Matthias Hagen},<br> booktitle = {Experimental IR Meets Multilinguality, Multimodality, and Interaction. 12th International Conference of the CLEF Association (CLEF 2021)},<br> editor = {{K. Sel{\c{c}}uk} Candan and Bogdan Ionescu and Lorraine Goeuriot and Henning M{\"u}ller and Alexis Joly and Maria Maistro and Florina Piroi and Guglielmo Faggioli and Nicola Ferro},<br> ids = {potthast:2021t},<br> month = sep,<br> pages = {450-467},<br> publisher = {Springer},<br> series = {Lecture Notes in Computer Science},<br> site = {Bucharest, Romania},<br> title = {{Overview of Touch{\'e} 2021: Argument Retrieval}},<br> volume = 12880,<br> year = 2021<br>}<br></code><br><code>@InProceedings{bondarenko:2022f,<br> address = {Berlin Heidelberg New York},<br> author = {Alexander Bondarenko and Maik Fr{\"o}be and Johannes Kiesel and Shahbaz Syed and Timon Gurcke and Meriem Beloucif and Alexander Panchenko and Chris Biemann and Benno Stein and Henning Wachsmuth and Martin Potthast and Matthias Hagen},<br> booktitle = {Experimental IR Meets Multilinguality, Multimodality, and Interaction. 13th International Conference of the CLEF Association (CLEF 2022)},<br> editor = {Alberto Barr{\'o}n-Cede{\~n}o and Giovanni Da San Martino and Mirko Degli Esposti and Fabrizio Sebastiani and Craig Macdonald and Gabriella Pasi and Allan Hanbury and Martin Potthast and Guglielmo Faggioli and Nicola Ferro},<br> ids = {potthast:2022j},<br> month = sep,<br> numpages = 29,<br> publisher = {Springer},<br> series = {Lecture Notes in Computer Science},<br> site = {Bologna, Italy},<br> title = {{Overview of Touch{\'e} 2022: Argument Retrieval}},<br> year = 2022<br>}</code></pre> <h3>Antique</h3> <p>If you re-use the Antique indices, please additionally cite:</p> <pre><code>@inproceedings{hashemi:2020,<br> author = {Helia Hashemi and Mohammad Aliannejadi and Hamed Zamani and W. Bruce Croft},<br> editor = {Joemon M. Jose and Emine Yilmaz and Jo{\~{a}}o Magalh{\~{a}}es and Pablo Castells and Nicola Ferro and M{\'{a}}rio J. Silva and Fl{\'{a}}vio Martins},<br> title = {{ANTIQUE:} {A} Non-factoid Question Answering Benchmark},<br> booktitle = {Advances in Information Retrieval - 42nd European Conference on {IR} Research, {ECIR} 2020, Lisbon, Portugal, April 14-17, 2020, Proceedings, Part {II}},<br> series = {Lecture Notes in Computer Science},<br> volume = {12036},<br> pages = {166--173},<br> publisher = {Springer},<br> year = {2020},<br>}</code><br><br></pre> <h3>CORD-19</h3> <p>If you re-use the CORD-19 indices, please additionally cite:</p> <pre><code>@article{voorhees:2020,<br> author = {Ellen M. Voorhees and Tasmeer Alam and Steven Bedrick and Dina Demner{-}Fushman and William R. Hersh and Kyle Lo and Kirk Roberts and Ian Soboroff and Lucy Lu Wang},<br> title = {{TREC-COVID:} constructing a pandemic information retrieval test collection},<br> journal = {{SIGIR} Forum},<br> volume = {54},<br> number = {1},<br> pages = {1:1--1:12},<br> year = {2020},<br>}<br><br>@article{wang:2020,<br> author = {Lucy Lu Wang and Kyle Lo and Yoganand Chandrasekhar and Russell Reas and Jiangjiang Yang and Darrin Eide and Kathryn Funk and Rodney Kinney and Ziyang Liu and William Merrill and Paul Mooney and Dewey A. Murdick and Devvret Rishi and Jerry Sheehan and Zhihong Shen and Brandon Stilson and Alex D. Wade and Kuansan Wang and Chris Wilhelm and Boya Xie and Douglas Raymond and Daniel S. Weld and Oren Etzioni and Sebastian Kohlmeier},<br> title = {{CORD-19:} The Covid-19 Open Research Dataset},<br> journal = {CoRR},<br> volume = {abs/2004.10706},<br> year = {2020},<br> eprinttype = {arXiv},<br> eprint = {2004.10706},<br>}<br><br></code></pre> <h3>Cranfield</h3> <p>If you re-use the Cranfield indices, please additionally cite:</p> <pre><code>@inproceedings{cleverdon:1967,<br> title={The {C}ranfield tests on index language devices},<br> author={Cleverdon, Cyril},<br> booktitle={{ASLIB} Proceedings},<br> year={1967},<br> pages = {173--192},<br> organization={MCB UP Ltd. (Reprinted in Readings in Information Retrieval, Karen Sparck-Jones and Peter Willett, editors, Morgan Kaufmann, 1997)}<br>}<br><br>@inproceedings{cleverdon:1991,<br> author = {Cyril W. Cleverdon},<br> editor = {Abraham Bookstein and Yves Chiaramella and Gerard Salton and Vijay V. Raghavan},<br> title = {The Significance of the {C}ranfield Tests on Index Languages},<br> booktitle = {Proceedings of the 14th Annual International {ACM} {SIGIR} Conference on Research and Development in Information Retrieval. Chicago, Illinois, USA, October 13-16, 1991 (Special Issue of the {SIGIR} Forum)},<br> pages = {3--12},<br> publisher = {{ACM}},<br> year = {1991},<br>}</code></pre> <h3>Medline TREC Genomics</h3> <p>If you re-use the Medline TREC Genomics indices, please additionally cite:</p> <pre><code>@inproceedings{hersh:2004,<br> author = {William R. Hersh and Ravi Teja Bhupatiraju and L. Ross and Aaron M. Cohen and Dale Kraemer and Phoebe Johnson},<br> editor = {Ellen M. Voorhees and Lori P. Buckland},<br> title = {{TREC} 2004 Genomics Track Overview},<br> booktitle = {Proceedings of the Thirteenth Text REtrieval Conference, {TREC} 2004, Gaithersburg, Maryland, USA, November 16-19, 2004},<br> series = {{NIST} Special Publication},<br> volume = {500-261},<br> publisher = {National Institute of Standards and Technology {(NIST)}},<br> year = {2004},<br>}<br><br>@inproceedings{hersh:2005,<br> author = {William R. Hersh and Aaron M. Cohen and Jianji Yang and Ravi Teja Bhupatiraju and Phoebe M. Roberts and Marti A. Hearst},<br> editor = {Ellen M. Voorhees and Lori P. Buckland},<br> title = {{TREC} 2005 Genomics Track Overview},<br> booktitle = {Proceedings of the Fourteenth Text REtrieval Conference, {TREC} 2005, Gaithersburg, Maryland, USA, November 15-18, 2005},<br> series = {{NIST} Special Publication},<br> volume = {500-266},<br> publisher = {National Institute of Standards and Technology {(NIST)}},<br> year = {2005},<br>}</code><br><br><br></pre> <h3>Medline TREC Precision Medicine</h3> <p>If you re-use the Medline TREC Precision Medicine indices, please additionally cite:</p> <pre><code>@inproceedings{roberts:2017,<br> author = {Kirk Roberts and Dina Demner{-}Fushman and Ellen M. Voorhees and William R. Hersh and Steven Bedrick and Alexander J. Lazar and Shubham Pant},<br> editor = {Ellen M. Voorhees and Angela Ellis},<br> title = {Overview of the {TREC} 2017 Precision Medicine Track},<br> booktitle = {Proceedings of The Twenty-Sixth Text REtrieval Conference, {TREC} 2017, Gaithersburg, Maryland, USA, November 15-17, 2017},<br> series = {{NIST} Special Publication},<br> volume = {500-324},<br> publisher = {National Institute of Standards and Technology {(NIST)}},<br> year = {2017},<br>}<br><br>@inproceedings{roberts:2018,<br> author = {Kirk Roberts and Dina Demner{-}Fushman and Ellen M. Voorhees and William R. Hersh and Steven Bedrick and Alexander J. Lazar},<br> editor = {Ellen M. Voorhees and Angela Ellis},<br> title = {Overview of the {TREC} 2018 Precision Medicine Track},<br> booktitle = {Proceedings of the Twenty-Seventh Text REtrieval Conference, {TREC} 2018, Gaithersburg, Maryland, USA, November 14-16, 2018},<br> series = {{NIST} Special Publication},<br> volume = {500-331},<br> publisher = {National Institute of Standards and Technology {(NIST)}},<br> year = {2018},<br>}</code><br><br></pre> <h3>MS MARCO (TREC Deep Learning 2019 and 2020</h3> <p>If you re-use the MS MARCO indices, please additionally cite:</p> <pre><code>@inproceedings{craswell:2019,<br> author = {Nick Craswell and Bhaskar Mitra and Emine Yilmaz and Daniel Campos and Ellen M. Voorhees},<br> booktitle = {28th International Text Retrieval Conference, {TREC} 2019, Gaithersburg, Maryland, USA},<br> editor = {{Ellen M.} Voorhees and Angela Ellis},<br> month = nov,<br> title = {{Overview of the {TREC} 2019 Deep Learning Track}},<br> publisher = {National Institute of Standards and Technology (NIST)},<br> series = {NIST Special Publication},<br> year = {2019}<br>}<br><br>@inproceedings{craswell:2020,<br> author = {Nick Craswell and Bhaskar Mitra and Emine Yilmaz and Daniel Campos},<br> editor = {Ellen M. Voorhees and Angela Ellis},<br> title = {{Overview of the {TREC} 2020 Deep Learning Track}},<br> booktitle = {Proceedings of the 29th Text REtrieval Conference, {TREC} 2020, Virtual Event, Gaithersburg, MD, USA, November 16-20, 2020},<br> series = {{NIST} Special Publication},<br> volume = {1266},<br> publisher = {National Institute of Standards and Technology {(NIST)}},<br> year = {2020},<br>}</code></pre> <pre> </pre> <h3>NFCorpus</h3> <p>If you re-use the next indices, please additionally cite:</p> <pre><code>@inproceedings{boteva:2016,<br> author = {Vera Boteva and Demian Gholipour Ghalandari and Artem Sokolov and Stefan Riezler},<br> editor = {Nicola Ferro and Fabio Crestani and Marie{-}Francine Moens and Josiane Mothe and Fabrizio Silvestri and Giorgio Maria Di Nunzio and Claudia Hauff and Gianmaria Silvello},<br> title = {A Full-Text Learning to Rank Dataset for Medical Information Retrieval},<br> booktitle = {Advances in Information Retrieval - 38th European Conference on {IR} Research, {ECIR} 2016, Padua, Italy, March 20-23, 2016. Proceedings},<br> series = {Lecture Notes in Computer Science},<br> volume = {9626},<br> pages = {716--722},<br> publisher = {Springer},<br> year = {2016},<br>}</code></pre> <h3>LongEval</h3> <p>Please cite the [corresponding dataset](https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-5151).:</p> <pre><code>@misc{11234/1-5151, title = {{LongEval} Click-Model Relevance Judgements (Qrels)}, author = {Galu{\v s}{\v c}{\'a}kov{\'a}, Petra and Devaud, Romain and Gonzalez-Saez, Gabriela and Mulhem, Philippe and Goeuriot, Lorraine and Piroi, Florina and Popel, Martin}, url = {http://hdl.handle.net/11234/1-5151}, note = {{LINDAT}/{CLARIAH}-{CZ} digital library at the Institute of Formal and Applied Linguistics ({{\'U}FAL}), Faculty of Mathematics and Physics, Charles University}, copyright = {Qwant {LongEval} Attribution-{NonCommercial}-{ShareAlike} License}, year = {2023} }</code><br><br></pre> <p>The index can be re-used in the [LongEval 2024](https://clef-longeval.github.io/) shared task hosted at [CLEF 2024](https://clef2024.imag.fr/). The documents (and thereby the derived PyTerrier Index are under the <a href="https://lindat.mff.cuni.cz/repository/xmlui/page/Qwant_LongEval_BY-NC-SA_License">Qwant LongEval Attribution-NonCommercial-ShareAlike License</a> and by reusing the indices you also accept and aggree to do this under the sharealike qwant license.</p>
User Evaluation and Metrics Analysis of a Prototype Web-based Federated Search Engine for Art and Cultural Heritage
<p>This dataset includes the quantitative data of the usage during the evaluation phase of a prototype web-based federated search engine for art and cultural heritage related content. The metrics which resulted in the dataset were in the form of a timeline of actions taken from a user (evaluator) in the course of a single session of interaction with the platform. A total of 20 different metrics were being monitored regarding the usage of the search engine, including submitting a query, a voice query, preforming a visual search, viewing a result, viewing a visual search result, updating an avatar, editing a user profile or changing user preferences, bookmarking and removing bookmarks of results and visual search results, using text to speech of all the various elements, opening the source view of a result and clicking a concept tag. All metrics included the timestamp of the event taking place and the value of the related event (e.g. the term of a search query).</p>
BIOFRUITNET - Boosting Innovation in Organic FRUIT production through knowledge NETworks - WP2 - Bibliographic search in Web of Science
<p>In the project BIOFRUITNET (Boosting Innovation in Organic FRUIT production through knowledge NETworks, HORIZON 2020 Grant Agreement No. 862850), nine search strings covering <strong>stone fruits, pome fruits and citrus fruits</strong>, within the selected main topics: <strong>pests, plant diseases and plant nutrition </strong>were developed, to target the most relevant scientific literature within the topics. The bibliographic searches were conducted on October 19, 2020 in Web of Science. The databases presented are the complete lists of references included from the WoS search. These lists will be processed to select the most relevant materials ready for practice in a following project work package. Project homepage: <a href="https://biofruitnet.eu/">https://biofruitnet.eu/</a></p>
FIG. 2 in Trolling for water striders: active searching for prey and the evolution of reduced webs in the spider Wendilgarda sp. (Araneae, Theridiosomatidae)
FIG. 2. Probable sequence of attachments by W. sp. near the water surface (schematic and not to scale). It was not clear from behavioural observations whether line a±b was doubled (as in drawing A) or broken and replaced (as it probably is in W. clara); nor was it certain whether point c was on the surface of the water or, more likely, just above it (see ®gure 7). It was con®rmed repeatedly, however, that the sticky line c±d (with balls in drawing C) was added to the non-sticky line rather than replacing it, as the sticky line was seen sagging brie¯y away from the straight vertical line. In two cases favourable lighting angles and background allowed con®rmation that line b±c was added to rather than replaced the line or lines laid just previously (a±b).
FIG. 7 in Trolling for water striders: active searching for prey and the evolution of reduced webs in the spider Wendilgarda sp. (Araneae, Theridiosomatidae)
FIG. 7. Microscopic views of silk on slides (stippled 5 puddles of sticky material; black masses 5 attachment discs, except for (E) where they 5 sticky material). (A) Lines of W. sp. ¯are away from central vertical line within masses of sticky material. (B) Attachment of W. sp. to water, showing attachment disc (presumed initiation of sticky silk) at bottom tip of vertical line. (C) Attachment of BCI creek W. sp. to water, showing attachment disc higher on vertical line and more sparse radial lines. (D) Attachment disc on a slack vertical line of W. sp. above section with sticky balls (e.g. d in ®gure 2), showing greater curliness of non-sticky lines. (E) Puddles of sticky material on a vertical line of W. sp., showing strong concentration of material at lower end of the line. Scale for (A), (C) and (D) at upper left.
FIG. 9 in Trolling for water striders: active searching for prey and the evolution of reduced webs in the spider Wendilgarda sp. (Araneae, Theridiosomatidae)
FIG. 9. Stages of the production of a second vertical line (A±D) and the initiation of a third line (E) by W. clara. The tight new vertical line (A, B) pulled the suspension line downward as the spider made the second descent; the tension then diminished and the angle in the suspension line became less acute (C; compare with B) as the spider extended the vertical line and then moved along the suspension line toward the previous vertical line (C). The spider apparently reeled up the suspension line as it then moved away from line 2 (D), because the white speck at the top of the new vertical line (dots in A±C) disappeared, and a white speck (presumably the accumulated reeled up silk) moved with the spider to the site where the next vertical line was laid (E). The positions of the white specks on the suspension line with respect to the vertical line in (A) ±(C) were not determined by direct observation; they are guesses based on the directions in which the specks moved and new lines were carried.
FIG. 5 in Trolling for water striders: active searching for prey and the evolution of reduced webs in the spider Wendilgarda sp. (Araneae, Theridiosomatidae)
FIG. 5. Probable mechanism used by spiders to jerk objects up out of the water (schematic). The spider ®rst reeled up the line, and thus tensed the entire line. When it then released this silk suddenly, the greater elasticity of the much longer line above the spider caused the spider to be displaced upward (arrow). The momentum of its body produced an upward jerk on the object when the line below its body became tight. This interpretation is tentative, because it was not possible to verify directly that the reeled up line was not broken (as in the drawing).
FIG. 8 in Trolling for water striders: active searching for prey and the evolution of reduced webs in the spider Wendilgarda sp. (Araneae, Theridiosomatidae)
FIG. 8. Tips of radial lines of an attachment of W. sp. to water, showing how they were progressively thinner near their tips (scale 5 0.05 mm).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.