Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,609
datasets available to search
ShareScore release 0.9.0
Dataset results
2,609 results for “web”
Design of an Ontology-Driven Constraint Tester (ODCT) and Application to SAREF & Smart Energy Appliances: Datasets, SHACL Shapes, Demo Video of Web Application, and Detailed Performance Reports
<h2>Description</h2> <p>This repository presents the resources used for validating the compliance of <strong>smart energy appliances</strong> against the <strong>Smart Appliances REFerence (SAREF)</strong> ontology and its extension <strong>SAREF4ENER</strong>, as part of the <strong>Ontology-Driven Constraint Tester (ODCT)</strong> project. The ODCT tool is specifically designed to ensure <strong>semantic interoperability</strong> and adherence to standardized ontological frameworks, which are crucial for integrating smart devices into modern energy management systems.</p> <h2>ODCT Overview</h2> <p>The <strong>Ontology-Driven Constraint Tester (ODCT)</strong> is a robust framework created to validate datasets against ontologies defined by <strong>SAREF</strong> and <strong>SAREF4ENER</strong>, both established under ETSI SmartM2M. This tool has been applied to the <strong>Flexible Start use case</strong> from the <strong>Joint Research Centre’s (JRC) Code of Conduct for Energy Smart Appliances</strong>. The ODCT tool ensures that smart devices like energy-efficient washing machines, thermostats, and connected lighting operate in compliance with established ontologies, thereby enhancing their <strong>interoperability</strong> within energy management systems and smart grids.</p> <h2>Repository Contents</h2> <p>This repository contains essential resources used in the ODCT compliance testing process:</p> <ul> <li> <p><strong>Compliant Dataset</strong>: This dataset represents a fully compliant scenario where no errors are present in the smart energy appliances’ profiles, demonstrating the ODCT’s accuracy under ideal conditions.</p> </li> <li> <p><strong>Modified Datasets</strong>: These datasets introduce various types of errors to showcase ODCT’s ability to handle diverse compliance scenarios:</p> <ol> <li><strong>Modified Dataset 1</strong>: Introduces type mismatches and spelling errors in key attributes.</li> <li><strong>Modified Dataset 2</strong>: Contains extraneous properties and missing required properties, including details about energy consumption and efficiency class.</li> <li><strong>Modified Dataset 3</strong>: Includes both extraneous and missing properties, and additional priority levels for energy profiles.</li> </ol> </li> <li> <p><strong>SHACL Shapes</strong>: The SHACL shapes used in the compliance testing for both SAREF and SAREF4ENER ontologies are included in this repository to allow reproducibility of the validation process.</p> </li> </ul> <ul> <li> <p><strong>Error Detection Results and Performance Reports</strong>: After conducting compliance tests using ODCT we got the Results and Performance Reports, the repository includes comprehensive reports detailing the results. These reports highlight the types of errors detected and provide a performance analysis of the tool under various scenarios.</p> </li> <li> <p><strong>Demonstration Video</strong>: A video is provided to guide users through the <strong>ODCT web application</strong>, showcasing how the tool detects errors and generates detailed compliance reports based on smart energy appliance datasets.</p> </li> </ul> <h2>Background</h2> <p>The integration of smart energy appliances into modern power grids is key to improving <strong>energy management</strong> and supporting <strong>sustainability goals</strong> like the <strong>European Green Deal</strong>. However, ensuring that these devices communicate effectively and conform to <strong>standardized protocols</strong> is a challenge. The <strong>ODCT</strong> tool addresses this challenge by providing a rigorous, ontology-based validation framework that is both <strong>protocol-agnostic</strong> and <strong>technology-flexible</strong>.</p> <p>This work is grounded in the broader context of <strong>global warming</strong> and the need for <strong>energy efficiency</strong> and <strong>demand-side flexibility</strong> in energy systems. By ensuring compliance with <strong>SAREF</strong> and <strong>SAREF4ENER</strong>, ODCT supports the EU’s ambitions for <strong>carbon neutrality</strong> by 2050, contributing to a connected, efficient, and sustainable energy ecosystem.</p> <h2>Methodology</h2> <p>ODCT uses a structured methodology that involves:</p> <ol> <li><strong>Generating relevant datasets</strong> for validation.</li> <li><strong>Defining SHACL shape constraints</strong> based on ontologies.</li> <li><strong>Developing a user-friendly web application</strong> to facilitate compliance testing.</li> <li><strong>Performing compliance tests</strong> that validate datasets against SHACL shapes, ensuring interoperability and adherence to energy management standards.</li> </ol> <h2>Why It Matters</h2> <p>Researchers and developers working on smart energy appliances will benefit from ODCT by:</p> <ul> <li>Ensuring their devices meet standardized ontological requirements for <strong>interoperability</strong>.</li> <li>Reducing <strong>compliance issues</strong> in the development phase, leading to smoother integration into energy management systems.</li> <li>Supporting the <strong>sustainability efforts</strong> by enhancing device communication in <strong>smart grids</strong>.</li> </ul> <p>This repository showcases the potential of ODCT in fostering <strong>data accuracy</strong>, <strong>semantic interoperability</strong>, and <strong>compliance</strong> with essential energy standards. It offers comprehensive resources for furthering research and development in the field of smart energy appliances and energy management.</p>
Data in: Reduced predation and energy flux in soil food webs by introduced tree species
<p>The introduction of non-native tree species has become a global concern and may disruptnative communities and related ecosystem functions. Soil food webs regulate organic matter decomposition and nutrient cycling in forests with their feeding activities, butevaluating consequences of tree species introduction on soil invertebrates is challengingdue to the complex trophic structure and wide range in body size of soil invertebrates. Here, we employed an energetic food web approach, and estimated the energy flux in soil food webs using a four-node model including soil meso- and macrofauna decomposers and predators. We examined pure and mixed stands of native European beech (<em>Fagus sylvatica</em>), introduced Douglas fir (<em>Pseudotsuga menziesii</em>) and native range-expanding Norway spruce (<em>Picea abies</em>) across site conditions. Compared to native forests, introduced tree species reduced total mass of macrofauna predators by 92% at sandy sites but not that of decomposers, suggesting trophic downgrading in soil food webs by Douglas fir. The energy flux in mixed forests was intermediate between respective monocultures, suggesting that tree mixtures mitigate potential negative impacts of introduced tree species on food web functioning. Across size classes, soil macrofauna responded more sensitively to changes in environmental conditions than soil mesofauna. Despite the lower total mass, the energy flux through mesofauna outweighed that through macrofauna when consideringenergy loss to predators, highlighting the importance of mesofauna for decomposition processes in forest soil food webs. Additionally, total energy flux positively correlated with species richness, pointing to the significance of soil biodiversity for trophic functionality. Overall, the study emphasizes the critical role of tree species composition, site conditionsand soil biodiversity in driving energy flux through soil food webs and maintaining forest ecosystem functions.</p>
Minimal data set for "Cohort profile: The ENTWINE iCohort Study, a multinational longitudinal web-based study of informal care"
<p><strong>Title:</strong></p> <p>Minimal Data Set for the Reproduction of Findings in "Elayan et al., Cohort Profile: The ENTWINE iCohort Study, a Multinational Longitudinal Web-Based Study of Informal Care".</p> <p> </p> <p><strong>Study Summary:</strong></p> <p>The data sets provided herein are derived from the ENTWINE iCohort Study, a multinational web-based cohort study employing an intensive longitudinal design. The study integrates a two-wave panel survey (baseline and 6-month follow-up) with optional weekly diary assessments. The cohort comprises caregivers and care recipients from nine countries: the United Kingdom, the Netherlands, Italy, Sweden, Israel, Germany, Greece, Poland, and Ireland. The study aimed to examine the influence of personal, psychological, social, economic, and geographic factors on caregiving experiences.</p> <p>Participants were eligible if they met the following criteria: 1) residency in a participating country; 2) capability to respond to surveys in English, Swedish, German, Dutch, Italian, Greek, Hebrew, or Polish; 3) access to the internet and ability to use it; 4) at least 18 years of age; 5) self-declared cognitive and physical capacity to complete the surveys; 6) either providing care to an adult (aged ≥ 18 years) with a chronic health condition, disability, or other care need, or receiving care from an adult due to similar conditions.</p> <p>The detailed methodology and results of the study can be found in the associated manuscript. For the complete survey questionnaires, please refer to: Morrison V, Zarzycki M, Vilchinsky N, Sanderman R, Lamura G, Fisher O, et al. A Multinational Longitudinal Study Incorporating Intensive Methods to Examine Caregiver Experiences in the Context of Chronic Health Conditions: Protocol of the ENTWINE-iCohort. Int J Environ Res Public Health. 2022;19. doi: <a href="https://doi.org/10.3390/ijerph19020821">10.3390/ijerph19020821</a></p> <p> </p> <p><strong>Data files:</strong></p> <p>The repository contains the following data files:</p> <ol> <li>"cg_minimal_dataset" (available in dta, sav, rds, and xlsx formats): This is a minimal data set containing de-identified and processed data derived from the ENTWINE iCohort Caregiver Baseline Survey. The variables present in this data set are detailed in the associated codebook, "cg_minimal_dataset_codebook".</li> <li>"cr_minimal_dataset" (available in dta, sav, rds, and xlsx formats): This is a minimal data set containing de-identified and processed data derived from the ENTWINE iCohort Care Recipient Baseline Survey. The variables present in this data set are detailed in the associated codebook, "cr_minimal_dataset_codebook".</li> </ol>
Data from "Evaluating top-down, bottom-up, and environmental drivers of pelagic food web dynamics along an estuarine gradient"
Synthesized fish, benthic invertebrate, and water quality dataset used for analysis in: Rogers, T., S. Bashevkin, C. Burdi, D. Colombano, P. Dudley, B. Mahardja, L. Mitchell, S. Perry, and P. Saffarinia. 2022. Evaluating top-down, bottom-up, and environmental drivers of pelagic food web dynamics along an estuarine gradient. preprint, EcoEvoRxiv. https://doi.org/10.32942/X2MK5Z
SGS-LTER Long-Term Monitoring Project: Vegetation Cover on Small Mammal Trapping Webs on the Central Plains Experimental Range, Nunn, Colorado, USA 1999 -2006, ARS Study Number 118 (Reformatted to the ecocomDP Design Pattern)
This data package is formatted as an ecocomDP (Ecological Community Data Pattern). For more information on ecocomDP see https://github.com/EDIorg/ecocomDP. This Level 1 data package was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-sgs/140/17. The abstract below was extracted from the Level 0 data package and is included for context: This data package was produced by researchers working on the Shortgrass Steppe Long Term Ecological Research (SGS-LTER) Project, administered at Colorado State University. Long-term datasets and background information (proposals, reports, photographs, etc.) on the SGS-LTER project are contained in a comprehensive project collection within the Digital Collections of Colorado (http://digitool.library.colostate.edu/R/?func=collections&collection_id=3429). The data table and associated metadata document, which is generated in Ecological Metadata Language, may be available through other repositories serving the ecological research community and represent components of the larger SGS-LTER project collection. Additional information and referenced materials can be found: http://hdl.handle.net/10217/83458. The abundance and diversity of small mammals in shortgrass steppe is strongly influenced by the structure and composition of vegetation. Vegetation structure provides cover from predators and harsh abiotic conditions. Plant species composition affects the types of seeds and herbaceous material available to granivores and herbivores, and influences arthropod populations, which are important prey for the omnivorous species that dominate in shortgrass steppe. Both vegetation structure and plant community composition are sensitive to the availability of precipitation as well as the activity of large mammalian herbivores. In 1999, we began measuring vegetation structure and plant community composition on the three grassland and three shrubland trapping webs where we live-trap small mammals
SGS-LTER Long-Term Montioring Project: Arthropod Pitfall Trapping on Small Mammal Trapping Webs on the Central Plains Experimental Range, Nunn, Colorado, USA 1998-2006, ARS Study Number 118 (Reformatted to the ecocomDP Design Pattern)
This data package is formatted as an ecocomDP (Ecological Community Data Pattern). For more information on ecocomDP see https://github.com/EDIorg/ecocomDP. This Level 1 data package was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-sgs/134/17. The abstract below was extracted from the Level 0 data package and is included for context: This data package was produced by researchers working on the Shortgrass Steppe Long Term Ecological Research (SGS-LTER) Project, administered at Colorado State University. Long-term datasets and background information (proposals, reports, photographs, etc.) on the SGS-LTER project are contained in a comprehensive project collection within the Digital Collections of Colorado (http://digitool.library.colostate.edu/R/?func=collections&collection_id=3429). The data table and associated metadata document, which is generated in Ecological Metadata Language, may be available through other repositories serving the ecological research community and represent components of the larger SGS-LTER project collection. Additional information and referenced materials can be found: http://hdl.handle.net/10217/83450. With the exception of heteromyids, eg kangaroo rats and pocket mice, most small rodents in shortgrass steppe are omnivorous. Depending on season, arthropods (insects and arachnids) make up 40-85% of the diet of grasshopper mice and thirteen-lined ground squirrels, the most widespread rodents in northern shortgrass steppe. Small mammals are among the most important predators of ground-dwelling macroarthropods and herbivorous insects provide a direct resource link between weather and plant production. Understanding temporal variability in the abundance of arthropods is central to determining the mechanisms that drive small rodent populations. At present, there are no long-term studies of arthropods in shortgrass steppe, despite the important role that these taxa play in grassland food w
SGS-LTER Long-Term Monitoring Project: Small Mammals on Trapping Webs on the Central Plains Experimental Range, Nunn, Colorado, USA 1994 -2006, ARS Study Number 118 (Reformatted to the ecocomDP Design Pattern)
This data package is formatted as an ecocomDP (Ecological Community Data Pattern). For more information on ecocomDP see https://github.com/EDIorg/ecocomDP. This Level 1 data package was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-sgs/137/17. The abstract below was extracted from the Level 0 data package and is included for context: This data package was produced by researchers working on the Shortgrass Steppe Long Term Ecological Research (SGS-LTER) Project, administered at Colorado State University. Long-term datasets and background information (proposals, reports, photographs, etc.) on the SGS-LTER project are contained in a comprehensive project collection within the Digital Collections of Colorado (http://digitool.library.colostate.edu/R/?func=collections&collection_id=3429). The data table and associated metadata document, which is generated in Ecological Metadata Language, may be available through other repositories serving the ecological research community and represent components of the larger SGS-LTER project collection. Additional information and referenced materials can be found: http://hdl.handle.net/10217/83452. Small mammals (rabbits, rodents) are integral components of semiarid ecosystems because of their roles as consumers of plants, seeds and arthropods, as soil disturbance agents, and as food for raptors, snakes and mammalian carnivores. Because of their vagility and intermediate trophic position, populations of small mammals may track changes in vegetation and the abiotic environment that may result from shifts in land-use and other anthropogenic disturbances. However, these populations are variable over space and time, and their response to environmental changes may not be immediately apparent given their behavioral flexibility and relatively long life-spans and generation times. Patterns in the distribution and abundance of small mammals thus may simultaneously reflect and affect the
Alligator pond food-web sampling in Shark River Slough and Taylor Slough, Everglades National Park, Florida, USA, 2018–2019
These datasets were used to investigate if American Alligators engineer differences in nutrient availability and changes to community structure by their creation of “alligator ponds” compared to the surrounding phosphorus (P)-limited oligotrophic marsh in the Everglades. We used a halo sampling design of three distinct habitats extending outward from ten active alligator ponds across a hydrological gradient. We performed nutrient analysis on basal food-web resources and quantitative community analyses, and stoichiometric analyses on plants and animals. These data underly the work in Strickland et al. (2023). An apex predator engineers wetland food-web heterogeneity through nutrient enrichment and habitat modification. Journal of Animal Ecology.
Stomach contents and stable isotopes from the aquatic food web in Everglades National Park, Florida, USA, 2018-2019
Stable isotope samples were collected among habitats and between seasons at multiple sites in both Shark River Slough and Taylor Slough of Everglades National Park to represent as much of the aquatic food web as possible. Stomach contents were also recorded for many of the collected vertebrate samples. Sampling occurred between October 2018 and April 2019. These data are a follow up study to the FCE1263 data package (https://portal.edirepository.org/nis/mapbrowse?scope=knb-lter-fce&identifier=1263) following the invasion of several fishes, most notably the African Jewelfish (Rubricatochromis letourneuxi), although there are some differences in habitats sampled and overall spatial coverage. R code for analyses associated with these data are available at https://github.com/pjflood/diet_and_pond_enrichment. Data collection is complete.
Experimental Understory Food Web data in the El Verde area of the Luquillo Experimental Forest
These data include date, treatment, block number, number of coquies, number of anoles, number of insects collected on two sticky traps, number of insects counted on 4 Piper glabrescens and 4 Manilkara bidentata seedlings, percent herbivory on the aforementioned plants and number of spiders. All these measurements were taken within exclosures for closed controls, anole exclusions, coqui exclusions and total exclusions. Open controls were sampled from an area of similar dimension not enclosed in an exclosure. Support for this work was provided by grants BSR-8811902, DEB-9411973, DEB-9705814 , DEB-0080538, DEB-0218039 , DEB-0620910 , DEB-1239764, DEB-1546686, and DEB-1831952 from the National Science Foundation to the University of Puerto Rico as part of the Luquillo Long-Term Ecological Research Program. Additional support provided by the University of Puerto Rico and the International Institute of Tropical Forestry, USDA Forest Service.
PIE LTER, stable isotope chemistry (carbon, nitrogen and sulfur) for food web analysis of functional groups in the Plum Island Sound Estuary, Massachusetts.
Stable isotopes of primary producers will be compared to stable isotopes of functional groups of organisms at primarily three sites within the estuary that have different dominant sources of organic matter. The three sites are: Lower (IBYC, SO-3, mouth of Plum Island Sound, 2-3 km upstream of the mouth of the estuary), Middle (OTL, PR-10.1, upper Sound, lower Parker, 8-11 km upstream of the mouth of the estuary) and Upper (P2, PR-21.9, upper Parker, above Middle Rd Bridge (22 km upstream of the mouth of the estuary). The Lower site is dominated by marine phytoplankton, the Middle site is dominated by a mixture of salt marsh and phytoplankton and the Upper site is dominated by oligohaline phytoplankton and fresh marsh. Ten functional groups will be sampled at each site (Surface sediment, benthic diatoms, Nereis, mummichog, ribbed mussels, POM, blue mussels, pelagic copepods (Acartia), silversides and soft shell clams (Mya). Marsh, benthic algae and phytoplankton inputs or benthic vs. pelagic pathways will be evident in the isotopic signals of these functional groups. Samples will be collected between the middle and end of August to reflect a growing season using recently produced OM. Samples of 15 – 20 individuals will be pooled for analysis. Some silverside samplings will have 3 different pooled samples for determination of variance. Often times the same species are not collected at each site due to habitat differences (salinity/discharge) so additional species are collected to try to accomplish task of getting functional groups collected.
Hypernym-LIBre: A free Web-based corpus from Hypernym Detection [ Hearst Pattern extractions from Hypernym-LIBre]
<p>Hypernym-LIBre ( DOI: 10.5281/zenodo.3662204 ) is a free Web-based corpus for Hypernym detection.</p> <p>Its part-of-speech tagged and dependency annotated version is present at this: (DOI: 10.5281/zenodo.3689303)</p> <p>Here we provide the hypernym-hyponym pairs that were extracted from Hypernym-LIBre using Hearst patterns. This is to further the usage of these extractions with more techniques and methods. We also provide the counts of each pattern in a separate file.</p> <p> </p> <p>Format:</p> <p>hyponym \t hypernym</p> <p> </p> <p>Format for the counts file:</p> <p>pair \t frequency of extraction</p> <p> </p> <p>There are 2 files, one with the pairs, one with unique pair and its counts. Both total ~430MB.</p>
Sample of facial mask N95 and FFP2 pricing on retail webs and time evolution per country
<p>We've gathered - for a Data Science educational project - the pricing of several face mask for breathing protection in a given period of time.</p> <p>Countries : Spain', 'USA', 'France', 'UK', 'Germany', 'Italy', 'Netherlands', 'Australia'</p> <p> </p> <p>'asin' type: STRING "Código de identifícación único de product equivalente de AMAZON"</p> <p>'description' type: STRING 'Texto descriptivo del producto'</p> <p>'dateTime' type: TIMESTAMP 'Cadena de carateres que contiene fecha y hora GMT'</p> <p>'date' Type: DATETIME ' Formato diferente de la misma fecha / hora de captura '</p> <p>'country' type: STRING 'Pais al que pertenece a distribución del producto 'Valores posibles: '</p>
[Dataset] FP-Redemption: Measuring Browser Fingerprinting Adoption for the Sake of Web Security
<p>Full dataset for the paper "FP-Redemption: Measuring Browser Fingerprinting Adoption for the Sake of Web Security"</p> <p>5 files are provided:</p> <ul> <li>dataset.csv. The raw elements collected when browsing the web. Each entry corresponds to one attribute being accessed with one parameter combination by one script on one webpage. A single attribute with the same parameters can be accessed several times. It is represented with the key <em>nbTimes</em></li> <li>domainTags.csv: For each website, it provides its category and country tag.</li> <li>webpageTags.csv: For each webpage, it provides its type.</li> <li>fingerprinters.zip/<filenumber>.js: Our fingerprinters. Out of the 199 we requested, 7 are missing, leading in 192 js files.</li> <li>mapping.csv. 3 columns CSV file: <ul> <li>The first one lists the 199 fingerprinters detected by our algorithm.</li> <li>The second one gives the <filenumber> used to link a fingerprinter and its file in the directory.</li> <li>The third one gives the groups the fingerprinters belongs to. By default, each fingerprinter belongs to his own group. However, several fingerprinters are belonging to the same group as we evaluate there were duplicates. Thus, the number of distinct groups corresponds to the distinct fingerprinters we measured in our dataset: 169.</li> </ul> </li> </ul>
TESTAR Test results extracted while executing MyThaiStar as web system under test
<p>TESTAR test results datasets extracted with TESTAR tool using MyThaiStar web application as System Under Test (SUT). These datasets have been generated to be used as an example to be automatically generated and introduced locally in DECODER PKM, from H2020 DECODER Project.</p> <p>TESTAR tool is an open source tool (www.testar.org) for automated testing through graphical user interface (GUI) currently being developed by the Universitat Politecnica de Valencia and the Open University of the Netherlands.</p> <p>MyThaiStar (<a href="https://github.com/devonfw/my-thai-star">github.com/devonfw/my-thai-star</a>) is the reference application that Capgemini uses internally to promote best programming practices and the correct use of last technologies. It’s is developed with Devon Framework, the standard tool for development at the company.</p> <p>PKM is the Persistent Knowledge Monitor developed as main infrastructure from H2020 DECODER Project (www.decoder-project.eu) under grant agreement number 824231.</p> <p>As TESTAR explores automatically the SUT, it will apply a couple of oracles to automatically check if any error or exception is detected at the Document Object Model (DOM) level extracted from MyThaiStar SUT.</p> <p>- MyThaiStar_TestResults_dataset.rar: All the information obtained through the DOM is used graphically and semantically to create screenshots, logs and html reports that indicate how TESTAR has navigated in the exploration and indicates if any error has been detected. </p> <p>- ArtefactTestResults_MyThaiStar_2020.1_2020-06-15_12h14m24s.json: for DECODER project purposes, TESTAR test results knowledge has been summarized and referenced in an artifact JSON file to adapt to PKM input requirements.</p> <p> </p> <p> </p> <p> </p>
Webis-Web-Archive-17
<p>The Webis-Web-Archive-17 comprises a total of 10,000 web page archives from mid-2017 that were carefully sampled from the Common Crawl to involve a mixture of high-ranking and low-ranking web pages. The dataset contains the web archive files, HTML DOM, and screenshots of each web page, as well as per-page annotations of visual web archive quality. See <a href="https://webis.de/data.html?q=tags%3Awebis-web-archive-17">this overview</a> for all datasets that built upon this one. If you use this dataset in your research, please cite it using <a href="https://webis.de/publications.html?q=10.1145%2F3239574">this paper</a>.</p>
Covid-on-the-Web dataset
<p>This RDF dataset provides two main knowledge graphs produced by processing the scholarly articles of the <a href="https://www.semanticscholar.org/cord19">COVID-19 Open Research Dataset</a> (CORD-19), a resource of articles about COVID-19 and the coronavirus family of viruses.</p> <p>The <em>CORD-19 Named Entities Knowledge Graph</em> describes named entities identified and disambiguated by NCBO BioPortal annotator, Entity-fishing and DBpedia Spotlight. The <em>CORD-19 Argumentative Knowledge Graph</em> describes argumentative components and PICO elements extracted from the articles by the Argumentative Clinical Trial Analysis platform (ACTA).</p> <p>Homepage: <a href="https://github.com/Wimmics/CovidOnTheWeb">https://github.com/Wimmics/CovidOnTheWeb</a></p> <p>License: see the LICENCE file in the archive.</p>
Práctica web scraping motorflash_v1
<p>Esta práctica se ha realizado en el contexto de la asignatura de 'tipología y ciclo de vida de los datos' del Master en Ciencia de Datos de la Universitat Oberta de Catalunya. En ella, se aplican técnicas de web scraping utilizando el lenguaje de programación Python y la librería scrapy. Los datos se han extraído de la página web de anuncios de coches de segunda mano '<a href="https://www.motorflash.com/">https://www.motorflash.com/</a>', de la que se obtienen datos generales y características del vehículo anunciado. </p> <p>Para conocer más a fondo el proceso de extracción puede visitar el repositorio del proyecto <a href="https://github.com/CarlosRea/MotorflashScraper">https://github.com/CarlosRea/MotorflashScraper</a> </p>
Japanese Trend Queries and Its' Web Trends
<p><strong>Abstract</strong> (our paper)</p> <p>Many researchers work on studies for discovering trend keywords and queries on the web, i.e., search frequency and social media. Moreover, studies on trend query classifications are being conducted. However, the behavior of trend queries for various web resources is unclear. In this study, we investigate how trend queries appear in different resources on the web. We clarify the following. (1) Most trend queries are not registered with online dictionary services. (2) The trend converges in approximately two days. (3) Social media websites (such as Twitter) are responsive to trend queries.</p> <p><strong>Data</strong></p> <p><strong>Publication</strong></p> <p>This data set was created for our study. If you make use of this data set, please cite:<br /> Mitsuo Yoshida, Yuki Arase. Trend Query Analysis on Heterogeneous Web Resources. <em>IPSJ Transactions on Databases (in Japanese)</em>. vol.9, no.1, pp.20-30, 2016.</p>
WikiMuTe: A web-sourced dataset of semantic descriptions for music audio
<p>This upload contains the supplementary material for our <a href="https://arxiv.org/abs/2312.09207" target="_blank" rel="noopener">paper</a> presented at the <a href="https://mmm2024.org/" target="_blank" rel="noopener">MMM2024 conference</a>.</p> <h2>Dataset</h2> <p>The dataset contains rich text descriptions for music audio files collected from Wikipedia articles.</p> <p>The audio files are freely accessible and available for download through the URLs provided in the dataset.</p> <h3>Example</h3> <p>A few hand-picked, simplified examples of the dataset. </p> <table> <tbody> <tr> <td> <p><strong>file</strong></p> </td> <td> <p><strong>aspects</strong></p> </td> <td> <p><strong>sentences</strong></p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/7/7a/Bongo_sound.wav" target="_blank" rel="noopener"><strong>🔈 Bongo sound.wav</strong></a></p> </td> <td> <p>['bongoes', 'percussion instrument', 'cumbia', 'drums']</p> </td> <td> <p>['a loop of bongoes playing a cumbia beat at 99 bpm']</p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/4/46/Example_of_double_tracking_in_a_pop-rock_song_%283_guitar_tracks%29.ogg" target="_blank" rel="noopener"><strong>🔈 Example of double tracking in a pop-rock song (3 guitar tracks).ogg</strong></a></p> </td> <td> <p>['bass', 'rock', 'guitar music', 'guitar', 'pop', 'drums']</p> </td> <td> <p>['a pop-rock song']</p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/6/62/OriginalDixielandJassBand-JazzMeBlues.ogg" target="_blank" rel="noopener"><strong>🔈 OriginalDixielandJassBand-JazzMeBlues.ogg</strong></a></p> </td> <td> <p>['jazz standard', 'instrumental', 'jazz music', 'jazz']</p> </td> <td> <p>['Considered to be a jazz standard', 'is an jazz composition']</p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/5/58/Colin_Ross_-_Etherea.ogg" target="_blank" rel="noopener"><strong>🔈 Colin Ross - Etherea.ogg</strong></a></p> </td> <td> <p>['chirping birds', 'ambient percussion', 'new-age', 'flute', 'recorder', 'single instrument', 'woodwind']</p> </td> <td> <p>['features a single instrument with delayed echo, as well as ambient percussion and chirping birds', 'a new-age composition for recorder']</p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/8/8b/Belau_rekid_%28instrumental%29.oga" target="_blank" rel="noopener"><strong>🔈 Belau rekid (instrumental).oga</strong></a></p> </td> <td> <p>['instrumental', 'brass band']</p> </td> <td> <p>['an instrumental brass band performance']</p> </td> </tr> <tr> <td> <p><strong>...</strong></p> </td> <td> <p>...</p> </td> <td> <p>...</p> </td> </tr> </tbody> </table> <h3>Dataset structure</h3> <p>We provide three variants of the dataset in the <code>data</code> folder.</p> <p>All are described in the paper.</p> <ol> <li><code>all.csv</code> contains all the data we collected, without any filtering.</li> <li><code>filtered_sf.csv</code> contains the data obtained using the <em>self-filtering</em> method.</li> <li><code>filtered_mc.csv</code> contains the data obtained using the <em>MusicCaps</em> dataset method.</li> </ol> <h3>File structure</h3> <p>Each CSV file contains the following columns:</p> <ul> <li><code>file</code>: the name of the audio file</li> <li><code>pageid</code>: the ID of the Wikipedia article where the text was collected from</li> <li><code>aspects</code>: the short-form (tag) description texts collected from the Wikipedia articles</li> <li><code>sentences</code>: the long-form (caption) description texts collected from the Wikipedia articles</li> <li><code>audio_url</code>: the URL of the audio file</li> <li><code>url</code>: the URL of the Wikipedia article where the text was collected from</li> </ul> <h3>Citation</h3> <div> <p>If you use this dataset in your research, please cite the following paper:</p> <div> <pre><code>@inproceedings{wikimute,</code><br><code> title = {WikiMuTe: {A} Web-Sourced Dataset of Semantic Descriptions for Music Audio},</code><br><code> author = {Weck, Benno and Kirchhoff, Holger and Grosche, Peter and Serra, Xavier},</code><br><code> booktitle = "MultiMedia Modeling",</code><br><code> year = "2024",</code><br><code> publisher = "Springer Nature Switzerland",</code><br><code> address = "Cham",</code><br><code> pages = "42--56",</code><br><code> doi = {10.1007/978-3-031-56435-2_4},</code><br><code> url = {https://doi.org/10.1007/978-3-031-56435-2_4},</code><br><code>}</code></pre> </div> </div> <h3>License</h3> <p>The data is available under the <a href="https://creativecommons.org/licenses/by-sa/3.0/" target="_blank" rel="noopener">Creative Commons Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0) license</a>.</p> <p>Each entry in the dataset contains a URL linking to the article, where the text data was collected from.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.