Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
190
datasets available to search
ShareScore release 0.9.0
Dataset results
190 results for “knowledge base”
Knowledge base for NBS for water treatment and stormwater management
<p>Five tables containing:</p> <ol> <li>nbs_catalog.csv: A catalogue of nature-based solutions for wastewater treatment and stormwater management. For each solution there is information on its performance, types of water, cobenefits, barriers and cost.</li> <li>sci_publications.csv: A list of scientific publications focused on one or several technologies of the above catalogue.</li> <li>sci_publications_treatment_details: For solutions for water treatment, a second table containing data about treatment performance extracted from previous scientific publications.</li> <li>description_nbs_catalog.csv: Descriptors for the catalogue.</li> <li>description_sci_publications_treatment_details.csv: Descriptors for the treatment performance data.</li> </ol> <p>The most updated version of each table can be queried from https://snappapi-v2.icradev.cat/</p>
Data from: A FAIR and modular image-based workflow for knowledge discovery in the emerging field of imageomics
<p>Data and results from the Imageomics Workflow. These include data files from the Fish-AIR repository (https://fishair.org/) for purposes of reproducibility and outputs from the application-specific imageomics workflow contained in the Minnow_Segmented_Traits repository (https://github.com/hdr-bgnn/Minnow_Segmented_Traits).</p> <p>Fish-AIR:<br> This is the dataset downloaded from Fish-AIR, filtering for Cyprinidae and the Great Lakes Invasive Network (GLIN) from the Illinois Natural History Survey (INHS) dataset. These files contain information about fish images, fish image quality, and path for downloading the images. The data download ARK ID is dtspz368c00q. (2023-04-05). The following files are unaltered from the Fish-AIR download. We use the following files:</p> <p>extendedImageMetadata.csv: A CSV file containing information about each image file. It has the following columns: ARKID, fileNameAsDelivered, format, createDate, metadataDate, size, width, height, license, publisher, ownerInstitutionCode. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>imageQualityMetadata.csv: A CSV file containing information about the quality of each image. It has the following columns: ARKID, license, publisher, ownerInstitutionCode, createDate, metadataDate, specimenQuantity, containsScaleBar, containsLabel, accessionNumberValidity, containsBarcode, containsColorBar, nonSpecimenObjects, partsOverlapping, specimenAngle, specimenView, specimenCurved, partsMissing, allPartsVisible, partsFolded, brightness, <br> uniformBackground, onFocus, colorIssue, quality, resourceCreationTechnique. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>multimedia.csv: A CSV file containing information about image downloads. It has the following columns: ARKID, parentARKID, accessURI, createDate, modifyDate, fileNameAsDelivered, format, scientificName, genus, family, batchARKID, batchName, license, source, ownerInstitutionCode. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>meta.xml: A XML file with the metadata about the column indices and URIs for each file contained in the original downloaded zip file. This file is used in the fish-air.R script to extract the indices for column headers.</p> <p>The outputs from the Minnow_Segmented_Traits workflow are:</p> <p>sampling.df.seg.csv: Table with tallies of the sampling of image data per species during the data cleaning and data analysis. This is used in Table S1 in Balk et al. </p> <p>presence.absence.matrix.csv: The Presence-Absence matrix from segmentation, not cleaned. This is the result of the combined outputs from the presence.json files created by the rule “create_morphological_analysis”. The cleaned version of this matrix is shown as Table S3 in Balk et al.</p> <p>heatmap.avg.blob.png and heatmap.sd.blob.png: Heatmaps of average area of biggest blob per trait (heatmap.avg.blob.png) and standard deviation of area of biggest blob per trait (heatmap.sd.blob.png). These images are also in Figure S3 of Balk et al.</p> <p>minnow.filtered.from.iqm.csv: Filtered fish image data set after filtering (see methods in Balk et al. for filter categories).</p> <p>burress.minnow.sp.filtered.from.iqm.csv: Fish image data set after filtering and selecting species from Burress et al. 2017.</p>
A map selection of wigeon stopover sites (core areas) based on wetland expert knowledge
<p>Stopover areas (core areas only) along the migration route of wigeons tracked with GPS transmitters were selected when they exhibited forests on more than 50% of their total surface or had less than 50% cover by water and/or wetland on the ESA’s global land cover map. We created a sample of 5,630 regions of interest (3,403 for training and 2,227 for validation), delineated with polygons assigned to land classes listed in the Table 1. We used archives of Google Earth, ESRI, and BING satellites for the photointerpretation of the land classes as described in Table 1. The classification was performed with a Sentinel-2 MultiSpectral Instrument, Level-2A image collection in Google Earth Engine (GEE) through the R-package Rgee to create a batch process applying the GEE Random forest classifier to each selected core home range. The cloudless (maximum 3%) images were selected within the period from 01/06/2021 to 30/09/2021. The optimal number of trees was estimated at 100 for an out of bag error of 14%. The overall accuracy on the validation sample was 82 %. </p>
Database of permacultural adoption responses in Mexicali, BC, Mexico. based on Circular Economy, Knowledge Management, and Sustainability policies
<p>Database documenting the perspectives of citizens in Mexicali, Baja California, Mexico, regarding the adoption of permaculture practices. The study is analyzed through the lenses of Knowledge Management, Circular Economy, and Sustainability Policies. The data was collected during the summer of 2024. </p>
MUHAI Benchmark : Task 2 (Credibility of knowledge-based generated gossip stories)
<p><strong>Meaning and Understanding in Human-Centric AI (MUHAI) Benchmark<br> Task 2 (Credibility of knowledge-based generated gossip stories)</strong></p> <p> </p> <p>This dataset aims at investigating whether the use of Knowledge Graphs has an impact on the credibililty of automatically-generated stories.</p> <p><br> The submission includes the following data:</p> <ol> <li>Generated stories (.txt)</li> <li>Story generation template </li> <li>A tsv file with entities and triples (to be used for generating stories)</li> <li>Evaluation description : Questions and metrics submitted to the users</li> </ol> <p>The "gossip stories" are generated with the T5 languge model fine-tuned on the WebNLG challenge. The model takes the triples (file 3) as input and generates one sentence each. A link prediction algorithm based on Jaccard's similarity learns the likelihood of two entities to be related (3). Then, the narrative continues with automatically generated celebrity background descriptions.</p> <p>The credibility of the story is evaluated using a questionnaire based on Gaziano et. al. The questionnaire was filled in by the test subjects after reading each generated article. One for a KG-generated text where links were predicted using the link prediction and one for text that was generated using triples of random entities (celebrities). </p> <p>Full code available at : https://github.com/kmitd/muhai-credibility-KR</p>
MineDojo Internet Knowledge Base (YouTube)
<p><strong>Project website:</strong> <a href="https://minedojo.org">minedojo.org</a></p> <p><strong>Paper:</strong> <a href="https://arxiv.org/abs/2206.08853">arxiv.org/abs/2206.08853</a></p> <p><strong>GitHub:</strong> <a href="https://github.com/MineDojo/MineDojo">github.com/MineDojo/MineDojo</a></p> <p>Minecraft is among the most streamed games on YouTube. Human players have demonstrated a stunning range of creative activities and sophisticated missions that take hours to complete. We collect 730K+ narrated Minecraft videos, which add up to <strong>33 years of duration and 2.2B words</strong> in English transcripts. The time-aligned transcripts enable the agent to ground free-form natural language in video pixels and learn the semantics of diverse activities without laborious human labeling.</p> <p>There are two files in our YouTube knowledge base.</p> <ul> <li><strong>youtube_tutorial.json</strong> (tutorial videos): <p>Minecraft tutorial videos include step-by-step demonstrations and sometimes detailed verbal explanations. They also serve as a rich source of creative missions that humans find interesting. We harvest thousands of tasks from these videos in our benchmarking suite. </p> </li> <li><strong>youtube_full.json</strong> (general gameplay videos): <p>Unlike tutorials, general gameplay videos do not necessarily provide guidance on particular tasks. Instead, they capture the “in-the-wild” human experiences that are much larger in quantity, diverse in contents, and rich in learning signals.</p> </li> </ul> <p>Data Structure</p> <pre><code class="language-python">list[ { "id": str, # video id "title": str, # video title "link": str, # video link "view_count": int # number of times the video has been viewed "like_count": int # number of users who have indicated that they liked the video "duration": float # video duration in seconds "fps": float, # video FPS } ]</code></pre> <p>Check out our paper!</p> <pre><code class="language-markdown">@article{fan2022minedojo, title = {MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge}, author = {Linxi Fan and Guanzhi Wang and Yunfan Jiang and Ajay Mandlekar and Yuncong Yang and Haoyi Zhu and Andrew Tang and De-An Huang and Yuke Zhu and Anima Anandkumar}, year = {2022}, journal = {arXiv preprint arXiv: Arxiv-2206.08853} }</code></pre> <p> </p>
Context-based Entity Recommendation on Real-Life Knowledge Work in Context (RLKWiC dataset)
<h2><a href="../records/11059573">RLKWiC</a> Add-on: Benchmarking Dataset for Entity Recommendation</h2> <p>This benchmark, built on top of the Real-Life Knowledge Work in Context (<a href="../records/11059573">RLKWiC</a>) dataset, is designed to evaluate context-based entity recommendation by simulating a scenario where participants receive entities extracted from their activities across their defined contexts. </p> <p>In total, 1850 entity recommendations were generated across 56 contexts. After deduplication, these entities were presented to participants for explicit relevance assessment on a 3-point scale:</p> <ul> <li>0 [Irrelevant]: Signifying a lack of relevance between the recommended entity and the context.</li> <li>1 [Relevant] Denoting a connection between the entity and the context, although it may not fully represent it.</li> <li>2 [Representative]: The entity closely aligns with the context, indicating a high level of relevance where the context can be inferred to be about this entity.</li> </ul> <p>Participants could also suggest additional relevant entities. The resulting dataset comprises 1067 entities with explicit relevance scores, offering a resource for benchmarking entity recommendation in real-life knowledge work.</p> <h3><strong>Paper: </strong><a href="https://dl.acm.org/doi/10.1145/3640457.3688068" target="_blank" rel="noopener">Context-based Entity Recommendation for Knowledge Workers: Establishing a Benchmark on Real-life Data</a></h3>
FaaS Characteristics and Constraints Knowledge Base
<p>YAML-formatted and timestamped description of Function-as-a-Service (FaaS) service characteristics and constraints such as maximum execution time and pricing. This dataset allows for adaptive software and workflow generation in dynamically evolving Serverless Computing environments. We envision the inclusion of the dataset into code generators, code transformers, workflow schedulers and compatibility modes of open source FaaS runtimes.</p> <p>Furthermore, due to evidences of evolving values being given by hyperlinks, the dataset will serve as single source of truth about the technological development in the FaaS space.</p>
Iterative Bleaching Extends Multiplexity (IBEX) Knowledge-Base
<p>The Iterative Bleaching Extends Multiplexity (IBEX) imaging method is an iterative immunolabeling and chemical bleaching method that enables highly multiplexed imaging of diverse tissues. Development of the <a href="https://doi.org/10.1038/s41596-021-00644-9">IBEX method</a> and <a href="https://github.com/niaid/imaris_extensions">related software</a> was led by Dr. Andrea Radtke and Dr. Ziv Yaniv. <a href="https://doi.org/10.1073/pnas.2018488117">IBEX</a> and related methods, <a href="https://doi.org/10.1073/pnas.1708981114">Ce3D</a>, <a href="https://doi.org/10.1111/imr.13052">Ce3D-IBEX</a>, <a href="https://doi.org/10.1073/pnas.2018488117">Opal-plex</a>, were originally developed in the laboratory of <a href="https://www.niaid.nih.gov/research/ronald-n-germain-md-phd">Dr. Ronald N. Germain</a>, US National Institutes of Health.</p><p>The IBEX Imaging Community is an international group of scientists committed to sharing knowledge related to multiplexed imaging in a transparent and collaborative manner. This open, global repository is a central resource for reagents, protocols, panels, publications, software, and datasets. In addition to IBEX, we support standard, single cycle multiplexed imaging (Multiplexed 2D imaging), volume imaging of cleared tissues with clearing enhanced 3D (Ce3D), highly multiplexed 3D imaging (Ce3D-IBEX), and extension of the IBEX dye inactivation protocol to the Leica Cell DIVE (Cell DIVE-IBEX). This dataset contains the current state of knowledge with respect to the IBEX microscopy imaging protocol.</p><p>How to use the Knowledge-Base:</p><ol><li>Save a copy to your computer.</li><li>To find a reagent: Open the reagent_resources.csv file found in the data directory. Use a spreadsheet application to filter the columns based on target name, target species, vendor, etc.</li><li>To view a complete list of fluorescent probes tested by the IBEX imaging community: Open the fluorescent_probes.csv file. This file reports the spectral properties and inactivation conditions of each fluorescent probe.</li><li>To import publications cited in the Knowledge-Base, import the publications.bib file found in the data directory to your reference manager.</li><li>To view a local copy of the website: Open the index.md file found in the docs directory using a markdown editor such as the free <a href="https://code.visualstudio.com/">Visual Studio Code</a>.</li><li>To view supporting information for a reagent (images, publications, notes): Open a specific target-conjugate-orcid combination under the docs-supporting_material directory structure using a markdown editor. This can also be visualized from the <a href="https://ibeximagingcommunity.github.io/ibex_imaging_knowledge_base/reagent_resources.html">Reagent Resources page</a> and filtered using a catalog number or other unique identifier in your web browser.</li></ol><p></p><p>Join the <a href="https://ibeximagingcommunity.github.io/ibex_imaging_knowledge_base/">online IBEX Imaging community</a> and contribute your knowledge. For more details on how to contribute, see <a href="https://ibeximagingcommunity.github.io/ibex_imaging_knowledge_base/contrib.html">these instructions</a>.</p><p>This research was supported by:</p><ul><li>The Intramural Research Program of the NIH, National Institute of Allergy and Infectious Diseases and National Cancer Institute, under grants 1ZIAAI001290-02, 1ZIAAI000545-33, 1ZIAAI000758-24, 1ZIAAI000974-16, 1ZIAAI001034-14.</li><li> The Wellcome Trust, under grant 224586/Z/21/Z.</li><li> The National Institute of Allergy and Infectious Diseases, NIH, under grant 1ZIAAI001343-01.</li></ul><p></p>
Qualitative dataset based on ancestral knowledge about coffee crops
<p> </p> <p>The qualitative dataset is about coffee pests based on the ancestral knowledge of coffee farmers in the Department of Cauca, Colombia. The dataset has been obtained from a survey applied to coffee growers with 432 records and 41 variables collected weekly from September 2020 to August 2021. The qualitative dataset includes climatic conditions, productive activities, external conditions, and coffee bio-aggressors. This dataset allows researchers to find patterns for coffee crop protection by means of ancestral knowledge not detected by real-time agricultural sensors. As far as we are concerned, there are no datasets like the one presented in this paper with similar characteristics of qualitative value that express the empirical knowledge of coffee farmers used to detect triggers of causal behaviors of pests and diseases in coffee crops.</p>
Fig. 16 in Contributions to the knowledge of Formicidae (Hymenoptera, Aculeata): a new diagnosis of the family, the first global male-based key to subfamilies, and a treatment of early branching lineages
Fig. 16. Male representatives of three subfamilies. A–B. Formica wheeleri, Formicinae (U.S.A., CASENT0173024, A. Nobile). C–D. Rhytidoponera, Ectatomminae, ectaheteromorph clade (Australia, CASENT0004610, A. Nobile). E–F. Pogonomyrmex rastratus (Argentina, CASENT0172673, A. Nobile). Scale bars: A, C = 0.5 mm, B, D, F = 1.0 mm, E = 0.2 mm.
Fig. 9 in Contributions to the knowledge of Formicidae (Hymenoptera, Aculeata): a new diagnosis of the family, the first global male-based key to subfamilies, and a treatment of early branching lineages
Fig. 9. Apomyrma CD01, male, photomicrographs. A. Forewing. B. Hindwing. C. Abdominal sternum IX, ventral view. D. Genital capsule, dorsal view. E. Genital valves, slightly splayed and without cupula, ventral view. F. Genital capsule, lateral view. G. Volsella and paramere, mesal view. H. Penisvalva in situ, mesal view. Scale bars: A–B = 0.5 mm, C–H = 0.1 mm. Abbreviations: see Material and Methods.
Fig. 10 in Contributions to the knowledge of Formicidae (Hymenoptera, Aculeata): a new diagnosis of the family, the first global male-based key to subfamilies, and a treatment of early branching lineages
Fig. 10. Representative males of Leptanillinae, lateral view A. Protanilla "TH01" (Thailand, CASENT0119776, A. Nobile), arrow indicates loss of abdominal segment II petiolation. B. Protanilla "TH03" (Thailand, CASENT0119791, E. Prado). C. Leptanilla swani (Australia, CASENT0172318, A. Nobile). D. Protanilla sp. (Indonesia, CASENT0178838, A. Nobile), arrow indicates basolateral basimeral process. E. Scyphodon sp. (Indonesia, MCZ155112w, A. Nobile). F. Noonilla sp., used with permission from Petersen (1968). Scale bars: A, C, E–F = 0.2 mm, D = 0.5 mm, B = 1.0 mm.
Fig. 12 in Contributions to the knowledge of Formicidae (Hymenoptera, Aculeata): a new diagnosis of the family, the first global male-based key to subfamilies, and a treatment of early branching lineages
Fig. 12. Martialis heureka Rabeling & Verhaagh, 2008, male, wing photomicrographs and genitalia illustrations, genital membranes not shown. A. Forewing. B. Hindwing. C. Abdominal sternum IX, ventral view. D. Genital capsule, dorsal view. E. Genital capsule, ventral view. F. Genital capsule, lateral view. G. Volsella and paramere, mesal view. H. Penisvalva in situ, mesal view. Scale bars: A–B = 0.5 mm, C–H = 0.1 mm. Abbreviations: see Material and Methods.
Fig. 15 in Contributions to the knowledge of Formicidae (Hymenoptera, Aculeata): a new diagnosis of the family, the first global male-based key to subfamilies, and a treatment of early branching lineages
Fig. 15. Male representatives of three subfamilies. A, D. Frontal view. B–C, E. Lateral view. — A–B. Pseudomyrmex holmgreni, Pseudomyrmecinae (Paraguay, CASENT0173758, A. Nobile). C. Aneuretus simoni, Aneuretinae, used with permission from Wilson et al. (1956). D–E. Technomyrmex difficilis, Dolichoderinae (Madagascar, CASENT0049968, A. Nobile). Scale bars: A, D = 0.2 mm, B = 1.0 mm, E = 0.5 mm, no scale available for C.
Fig. 14 in Contributions to the knowledge of Formicidae (Hymenoptera, Aculeata): a new diagnosis of the family, the first global male-based key to subfamilies, and a treatment of early branching lineages
Fig. 14. Male representatives of three subfamilies. A, C, E. Frontal view. B, D, F. Lateral view. — A–B. Proceratium creek, Proceratiinae (U.S.A., CASENT010441, A. Nobile). C–D. Acanthostichus, Dorylinae (French Guiana, CASENT0056970, A. Nobile). E–F. Myrmecia chasei, Myrmeciinae (Australia, CASENT0903663, W. Ericson). Scale bars: A, C = 0.2 mm, B, E = 0.5 mm, D = 1.0 mm, E = 2.0 mm.
Fig. 4. Male morphology. A–C. Abdominal sternum IX in Contributions to the knowledge of Formicidae (Hymenoptera, Aculeata): a new diagnosis of the family, the first global male-based key to subfamilies, and a treatment of early branching lineages
Fig. 4. Male morphology. A–C. Abdominal sternum IX in ventral view. B. Oblique. D–E. Petiole in lateral view. F. Forewing in dorsal view. G–H. Propodeum in lateral view. — A. Aenictogiton indet. (Zambia, CASENT0106126, M. Branstetter). B. Cerapachys "parasyscia" lineage (Kenya, B. Boudinot). C. Emeryopone buttelreepeni (Thailand, CASENT0278779, B. Boudinot). D. Paraponera clavata (?Panama, B. Boudinot). E. Cerapachys lividus (Madagascar, CASENT0138502, D. Raharinjanahary). F. Phaulomyrma indet. (Thailand, UCRENT150358, A. Nobile). G. Leptanillinae indet. (Thailand, CASENT0156249, B. Boudinot); arrow indicates dorsal margin of petiolar foramen. H. Adelomyrmex dentivagans; arrow indicates propodeal lobe. All scale bars = 0.2 mm, except D–E = 1.0 mm.
Fig. 11 in Contributions to the knowledge of Formicidae (Hymenoptera, Aculeata): a new diagnosis of the family, the first global male-based key to subfamilies, and a treatment of early branching lineages
Fig. 11. Martialis heureka Rabeling & Verhaagh, 2008, male, photomicrographs. A. Head, frontal view. B. Head, anteroventral oblique view. C. Body, lateral view. D. Body, dorsal view. Scale bars: A–B = 0.2 mm, C–D = 0.5 mm.
Fig. 8 in Contributions to the knowledge of Formicidae (Hymenoptera, Aculeata): a new diagnosis of the family, the first global male-based key to subfamilies, and a treatment of early branching lineages
Fig. 8. Apomyrma CD01, male, photomicrographs (CASENT0086073, E. Prado). A. Head, frontal view. B. Body, lateral view. C. Body, dorsal view. Scale bars: A = 0.2 mm, B–C = 0.5 mm.
Fig. 2 in Contributions to the knowledge of Formicidae (Hymenoptera, Aculeata): a new diagnosis of the family, the first global male-based key to subfamilies, and a treatment of early branching lineages
Fig. 2. Pterothoracic venter morphology of representative hymenopterans. A. Tenthredinidae, female. B. Polistes (Vespidae), worker. C. Paraponera clavata gyne (Paraponerinae, Formicidae). Scale bars = 1.0 mm. Abbreviations: see Material and Methods.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.