Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
24
datasets available to search
ShareScore release 0.9.0
Dataset results
24 results for “html”
PanDDA analysis of BRD1 screened against 3D-Fragment-Consortium Fragment Library (HTML Summary)
<p>Interactive summary page for "PanDDA analysis of BRD1 screened against 3D-Fragment-Consortium Fragment Library".</p> <p><strong>Please click on "0_index.html" in the "Files" section to open the interactive summary.</strong></p> <p>All datasets are also available as combined zip files from https://zenodo.org/record/48769 .</p> <p> </p>
PanDDA analysis of JMJD2D screened against Zenobia Fragment Library (HTML Summary)
<p>Interactive summary page for "PanDDA analysis of JMJD2D screened against Zenobia Fragment Library".</p> <p><strong>Please click on "0_index.html" in the "Files" section to open the interactive summary.</strong></p> <p>All datasets are also available as combined zip files from https://zenodo.org/record/48770 .</p>
PanDDA analysis of SP100 screened against selection of Maybridge Fragment Library (HTML Summary)
<p>Interactive summary page for "PanDDA analysis of SP100 screened against selection of Maybridge Fragment Library".</p> <p><strong>Please click on "0_index.html" in the "Files" section to open the interactive summary.</strong></p> <p>All datasets are also available as combined zip files from https://zenodo.org/record/48771 .</p>
PanDDA analysis of BAZ2B screened against Zenobia Fragment Library (HTML Summary)
<p>Interactive summary page for "PanDDA analysis of BAZ2B screened against Zenobia Fragment Library".</p> <p><strong>Please click on "0_index.html" in the "Files" section to open the interactive summary.</strong></p> <p>All datasets are also available as combined zip files from https://zenodo.org/record/48768 .</p>
The BigGrams: the semi-supervised information extraction system from HTML: an improvement in the wrapper induction - dataset
<p><strong>Brief description</strong></p> <p>The zip file contains two folders. The <strong>"websites"</strong> folder includes crawled web pages from real websites, like a agatameble.pl (an e-shop website), filmweb.pl (a website about films), and ptaki.info (a website about birds). The <strong>"reference-seeds"</strong> folder contains three subfolders, i.e. agatameble.pl, filmweb.pl, and ptaki.info. Each subfolder contains reference-seeds.csv file. The file contains data, i.e. reference instances - carefully labelled ground-truth of corresponding values in each web page of given websites mentioned above.</p> <p><strong>Reference</strong></p> <p>I would appreciate it if you cite the following paper when using the dataset:</p> <p>Marcin Mirończuk The BigGrams: the semi-supervised information extraction system from HTML: an improvement in the wrapper induction, Knowledge and Information Systems, Volume 54, Issue 3, p. 711–776, 2018, (pdf Open Access – http://rdcu.be/u88F lub DOI http://dx.doi.org/10.1007/s10115-017-1097-2)</p>
Prototyping 3D Virtual Learning Environments with X3D-based Content and Visualization Tools-Figure 8. The 3D model of the faculty building in a HTML page
<p>The pipeline processing was the following: a) the 3D model from Sketchup was saved as a Collada file; b) this file has been imported in MeshLab (MESHLAB 2017) and converted to VRML97 format (wrl); c) aopt utility was used to convert wrl files to X3D and HTML5 files. The model was visualized in the OpenSim virtual world setting using an external browser (Figure 8).</p>
Prototyping 3D Virtual Learning Environments with X3D-based Content and Visualization Tools-Figure 7. The HTML source (partial) code, integrating the X3D model
<p>To integrate the model into a web page, a model conversion to X3D format and an X3DOM output under the form of an HTML5 encoded webpage (Figure 7) were needed. Instant Reality distribution provides a command line transcoding tool, named Avalon Optimizer (aopt), that was used to convert a VRML format (wrl extension) of the model to X3D. </p>
IMP HTML reports (preprint)
<p>ZIP file contains the HTML files which are referred to as the "Additional file 1" in the IMP manuscript.<br> </p>
IMP HTML reports
<p>This file contains all the HTML reports generated by IMP for the analysis of datasets reported in the article.</p>
PanDDA analysis of JMJD2D screened against Zenobia Fragment Library - HTML Summary
<p>De-methylase JMJD2D screened against the Zenobia Fragment Library by X-ray Crystallography.</p>
JavaScript and html Validity Errors
<p>This data is captured by tool ‘Artemis’, which is a test generation tool for javascript based web applications. The data contained here has benchmarks and html validity errors found out by Artemis tool, when ran against live applications. This data can be used to analyse the html coding errors committed by the programmers and see how we can prevent these errors.</p>
GRETIL - Göttingen Register of Electronic Texts in Indian Languages. HTML
<p>GRETIL - Göttingen Register of Electronic Texts in Indian Languages. HTML</p>
HTML Status Codes of Publications citing ICPSR
<p>This dataset contains the HTML status codes of ICPSR citing literature, which were tested in 2020 for a PhD thesis on research data and software (re)use indications in scholarly works.</p>
Transparency in Keyword Faceted Search: a dataset of Google Shopping html pages
<p>This dataset contains a collection of around 2,000 HTML pages: these web pages contain the search results obtained in return to queries for different products, searched by a set of synthetic users surfing Google Shopping (US version) from different locations, in July, 2016.</p> <p>Each file in the collection has a name where there is indicated the location from where the search has been done, the userID, and the searched product: <em>no_email_LOCATION_USERID.PRODUCT.shopping_testing.#.html</em></p> <p>The locations are Philippines (PHI), United States (US), India (IN). The userIDs: 26 to 30 for users searching from Philippines, 1 to 5 from US, 11 to 15 from India.</p> <p>Products have been choice following 130 keywords (e.g., MP3 player, MP4 Watch, Personal organizer, Television, etc.).</p> <p>In the following, we describe how the search results have been collected.</p> <p>Each user has a fresh profile. The creation of a new profile corresponds to launch a new, isolated, web browser client instance and open the Google Shopping US web page.</p> <p>To mimic real users, the synthetic users can browse, scroll pages, stay on a page, and click on links.</p> <p>A fully-fledged web browser is used to get the correct desktop version of the website under investigation. This is because websites could be designed to behave according to user agents, as witnessed by the differences between the mobile and desktop versions of the same website.</p> <p>The prices are the retail ones displayed by Google Shopping in US dollars (thus, excluding shipping fees).</p> <p>Several frameworks have been proposed for interacting with web browsers and analysing results from search engines. This research adopts OpenWPM. OpenWPM is automatised with <a href="http://www.seleniumhq.org/">Selenium</a> to efficiently create and manage different users with isolated Firefox and Chrome client instances, each of them with their own associated cookies.</p> <p>The experiments run, on average, 24 hours. In each of them, the software runs on our local server, but the browser's traffic is redirected to the designated remote servers (i.e., to India), via tunneling in SOCKS proxies. This way, all commands are simultaneously distributed over all proxies. The experiments adopt the Mozilla Firefox browser (version 45.0) for the web browsing tasks and run under Ubuntu 14.04. Also, for each query, we consider the first page of results, counting 40 products. Among them, the focus of the experiments is mostly on the top 10 and top 3 results.</p> <p>Due to connection errors, one of the Philippine profiles have no associated results. Also, for Philippines, a few keywords did not lead to any results: videocassette recorders, totes, umbrellas. Similarly, for US, no results were for totes and umbrellas.</p> <p>The search results have been analyzed in order to check if there were evidence of price steering, based on users' location.</p> <p><strong>One term of usage applies:</strong></p> <p>In any research product whose findings are based on this dataset, please cite</p> <pre>@inproceedings{DBLP:conf/ircdl/CozzaHPN19, author = {Vittoria Cozza and Van Tien Hoang and Marinella Petrocchi and Rocco {De Nicola}}, title = {Transparency in Keyword Faceted Search: An Investigation on Google Shopping}, booktitle = {Digital Libraries: Supporting Open Science - 15th Italian Research Conference on Digital Libraries, {IRCDL} 2019, Pisa, Italy, January 31 - February 1, 2019, Proceedings}, pages = {29--43}, year = {2019}, crossref = {DBLP:conf/ircdl/2019}, url = {https://doi.org/10.1007/978-3-030-11226-4\_3}, doi = {10.1007/978-3-030-11226-4\_3}, timestamp = {Fri, 18 Jan 2019 23:22:50 +0100}, biburl = {https://dblp.org/rec/bib/conf/ircdl/CozzaHPN19}, bibsource = {dblp computer science bibliography, https://dblp.org} } </pre> <p> </p> <p> </p>
Versiunile Doinei: aliniere în TEI-xml, html pentru Versioning Machine, capturi ecran VM
<p>Dosarul cuprinde fisiere xml și html care au rezultat în urma procesului de aliniere a versiunilor <em>Doinei</em> de Mihai Eminescu după principiul de segmentare paralelă și de marcare cu coduri TEI </p>
Additional HTML Figures: Exploding biplots with density axes in Plotly
<p>Two figures as seen in <em>Exploding biplots with density axes in Plotly, </em>in the HTML format produced by the code for this paper.</p>
IMP HTML reports
<p>ZIP file contains the HTML files which are referred to as the "Additional file 1" in the manuscript.</p>
FIGURE. Central Asian species of Salvia: A. S. korolkovii (photo by A. Gaziev, https://www.plantarium.ru/page/view/item/33503. html); B. S. aethiopis (photo by F. Celep) C. S. vvedenskyi (photo by Georgy Lazkov, https://www.plantarium.ru/page/view/item/33562. html). D. S. drobovii (photo by N. Beshko, https://www.plantarium.ru/page/view/item/33477.html); E. S. lilacinocoerulea (photo by N. Beshko, https://www.plantarium.ru/page/view/item/33505.html). F. S. karelinii (photo by D. Polevoy, https://www.plantarium.ru/page/ view/item/27351.html) G. S. insignis (photo by O. Turdiboev), H. S. verticillata subsp. amasiaca (photo by F. Celep). in Synopsis of the Central Asian Salvia species with identification key
FIGURE. Central Asian species of Salvia: A. S. korolkovii (photo by A. Gaziev, https://www.plantarium.ru/page/view/item/33503. html); B. S. aethiopis (photo by F. Celep) C. S. vvedenskyi (photo by Georgy Lazkov, https://www.plantarium.ru/page/view/item/33562. html). D. S. drobovii (photo by N. Beshko, https://www.plantarium.ru/page/view/item/33477.html); E. S. lilacinocoerulea (photo by N. Beshko, https://www.plantarium.ru/page/view/item/33505.html). F. S. karelinii (photo by D. Polevoy, https://www.plantarium.ru/page/ view/item/27351.html) G. S. insignis (photo by O. Turdiboev), H. S. verticillata subsp. amasiaca (photo by F. Celep).
Gretil quotations html files
<p>These are html-tables of possible quotations/parallel passages within the etexts contained in the gretil-collection.</p> <p>For the code and description see: https://github.com/sebastian-nehrdich/gretil-quotations</p> <p>There is a file called 0_index.html which consists of a table of all html-files and their corresponding text names(which have been scraped from the GRETIL-headers and might look not very readable or complete at times).</p> <p>The threshold for this calcaltion was set rather low, so there might be passages listed as parallel/quotations which are not really similar. </p> <p> </p> <p>After unzipping, the files will use about 4.7GB of disk space.</p>
Gretil quotations html files
<p>These are html-tables of possible quotations/parallel passages within the etexts contained in the gretil-collection.</p> <p>For the code and description see: https://github.com/sebastian-nehrdich/gretil-quotations</p> <p>There is a file called 0_index.html which consists of a table of all html-files and their corresponding text names(which have been scraped from the GRETIL-headers and might look not very readable or complete at times).</p> <p>After unzipping, the files will use about 5.4GB of disk space.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.