Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6,157
datasets available to search
ShareScore release 0.7.1
Dataset results
6,157 results for “knowledge”
MUHAI Benchmark : Task 1 (Short story generation with Knowledge Graphs)
<p><strong>Meaning and Understanding in Human-Centric AI (MUHAI) Benchmark<br> Task 1 (Short story generation with Knowledge Graphs and Language Models)</strong> </p> <p>The dataset can be used to test understandability of text generated through the combination of knowledge graphs and language models without using knowledge graph embeddings.<br> <br> The task here is to generate 5-sentence stories from a set of <em>subject-predicate-object</em> triples that are extracted from a knowledge graph. Two steps need to be performed:</p> <p>1. Language model fine-tuning (SVO triple extraction + model fine-tuning)<br> 2. Story generation (knowledge enrichment + text generation) <br> <br> The submission includes the following data:</p> <ol> <li>Original ROC stories corpus (100 stories)</li> <li>ROC stories encoded with relevant triples (extracted through SpaCy, 2 versions, with and without coreference resolution)</li> <li>Stories generated by the pre-trained model (GPT2-simple)</li> <li>Stories generated by the fine-tuned model (DICE + ConceptNet + DBpedia )</li> <li>Stories generated by the fine-tuned model (DICE + ConceptNet + DBpedia + WordNet )</li> <li>Stories generated by the GPT-2-keyword-generation (an open-source software that uses GPT-2 to generate text pertaining to the specified keywords)</li> <li>Model results</li> <li>Evaluation metrics description</li> <li>User-evaluation questionnaire </li> </ol> <p>Code : https://github.com/kmitd/muhai-dice_story</p>
Knowledge base for NBS for water treatment and stormwater management
<p>Five tables containing:</p> <ol> <li>nbs_catalog.csv: A catalogue of nature-based solutions for wastewater treatment and stormwater management. For each solution there is information on its performance, types of water, cobenefits, barriers and cost.</li> <li>sci_publications.csv: A list of scientific publications focused on one or several technologies of the above catalogue.</li> <li>sci_publications_treatment_details: For solutions for water treatment, a second table containing data about treatment performance extracted from previous scientific publications.</li> <li>description_nbs_catalog.csv: Descriptors for the catalogue.</li> <li>description_sci_publications_treatment_details.csv: Descriptors for the treatment performance data.</li> </ol> <p>The most updated version of each table can be queried from https://snappapi-v2.icradev.cat/</p>
The OREGANO knowledge graph for computational drug repurposing
<p>The files here are data files from the OREGANO project, which consists of building a holistic knowledge graph on drugs, including natural compounds. Here is the list of files:</p><p> </p><p>- OREGANO_V2.tsv : The triplet file used for link prediction. 3 columns : Subjet ; Predicate ; Object</p><p>- oreganov2.1_metadata_complet.ttl : The OREGANO knowledge graph in turtle format with the names and cross-references of the various integrated entities.</p><p> </p><p>The following files contain the cross-references of OREGANO entities according to their type. They are all organised as follows: the external sources are the titles of the columns and each line begins with the identifier of the entity in OREGANO :</p><p>- TARGET.tsv: Cross-reference table of the 22,096 targets.<br>- PHENOTYPES.tsv: Cross-reference table of the 11,605 phenotypes.<br>- DISEASES.tsv: Cross-reference table of the 18,333 diseases.<br>- PATHWAYS.tsv: Cross-reference table of the 2,129 pathways.<br>- GENES.tsv: Cross-reference table of the 35,794 genes.<br>- COMPOUND.tsv: Cross-reference table of the 90,868 compounds.<br>- INDICATIONS.tsv: Cross-reference table of the 2,714 indications.<br>- SIDE_EFFECT.tsv: Cross-reference table of the 6,060 side-effects.<br>- ACTIVITY.tsv: Names of the 78 activities.<br>- EFFECT.tsv: Names of the 171 effects.</p><p>The OREGANO knowledge graph is composed of 11 types of nodes and 19 types of links. The current version of the graph contains 88,937 nodes and 824,231 links.</p><p>A SPARQL endpoint has been provided to enable users to retrieve and explore the knowledge graph at <a href="http://91.121.148.199:8889/bigdata/#query">OREGANO SPARQL endpoint</a> .</p><p> </p><p>The integration files and the knowledge graph are available on the GitHub of the OREGANO project in the Integration folder: <a href="https://gitub.u-bordeaux.fr/erias/oregano">Gitub repository</a> .</p>
Event-QA: A Dataset for Event-Centric Question Answering over Knowledge Graphs
<p>Event-QA dataset contains 1000 semantic queries and the corresponding verbalisations for EventKG - a recently proposed event-centric knowledge graph containing over 970 thousand events.</p>
CETAF-DiSSCo/COVID19-TAF biodiversity-related knowledge hub working group: indexed biotic interactions and review summary
<p>This data publication originated as part of developing a biodiversity-related knowledge hub on COVID-19 via COVID19-TAF - Communities Taking Action (https://cetaf.org/covid19-taf-communities-taking-action), a community-rooted initiative raised jointly by the Consortium of European Taxonomic Facilitaties (CETAF, https://cetaf.org) and Distributed Systems of Scientific Collections (DiSSCo, https://www.dissco.eu/).</p> <p>This archive contains the biodiversity datasets of interest identified in period 14 April-6 October 2020 through COVID19-TAF activities and subsequently indexed by Global Biotic Interactions (GloBI, https://globalbioticinteractions.org). GloBI provides open access to finding species interaction data (e.g., predator-prey, pollinator-plant, virus-host, parasite-host) by combining existing open datasets using open source software.</p> <p>These identified datasets (see references and reviews below) add to a growing collection of open species interaction datasets already indexed by GloBI. So, this data publication only includes a small subset of indexed datasets and include only datasets that were added as a direct consequence of COVID19-TAF activities of the biodiversity-related knowledge hub working group.</p> <p>If you have questions or comments about this publication, please open an issue at https://github.com/ParasiteTracker/tpt-reporting or contact the authors by email.</p> <p>Funding:<br> The creation of this archive was made possible in part by reporting software developed as part of the National Science Foundation award "Collaborative Research: Digitization TCN: Digitizing collections to trace parasite-host associations and predict the spread of vector-borne disease," Award numbers DBI:1901932 and DBI:1901926 . Also, this material is based upon work supported by the National Science Foundation under Grant No. DGE-1545433 .</p> <p>References:<br> Jorrit H. Poelen, James D. Simons and Chris J. Mungall. (2014). Global Biotic Interactions: An open infrastructure to share and analyze species-interaction datasets. Ecological Informatics. https://doi.org/10.1016/j.ecoinf.2014.08.005.</p> <p>GloBI Data Review Report</p> <p>Datasets under review:<br> - Geiselman, Cullen K. & Sarah Younger. 2020. Bat Eco-Interactions Database. www.batbase.org accessed via https://github.com/globalbioticinteractions/batbase/archive/9c65cfeee1a054f9db8cd8bf6892017fd1b3c840.zip on 2020-10-04T22:53:45.576Z<br> - Geiselman, Cullen K. and Tuli I. Defex. 2015. Bat Eco-Interactions Database. www.batplant.org accessed via https://github.com/globalbioticinteractions/batplant/archive/a2e1b57052244d5251d17e96ea61f58bea88975e.zip on 2020-10-04T22:54:28.727Z<br> - Daniel Becker, Gregory F Albery, Anna R Sjodin, Timothee Poisot, Tad Dallas, Evan A. Eskew, Maxwell J. Farrell, Sarah Guth, Barbara A Han, Nancy B Simmons, Colin J Carlson. 2020. Predicting wildlife hosts of betacoronaviruses for SARS-CoV-2 sampling prioritization. bioRxiv 2020.05.22.111344; doi: https://doi.org/10.1101/2020.05.22.111344 accessed via https://github.com/globalbioticinteractions/becker2020/archive/47c6ad28e1c5058f3c13ca69a59fdf21229e8d7f.zip on 2020-10-04T22:54:46.723Z<br> - Chen L, Liu B, Yang J, Jin Q, 2014. DBatVir: the database of bat-associated viruses. Database (Oxford). 2014:bau021. doi:10.1093/database/bau021 accessed via https://github.com/globalbioticinteractions/dbatvir/archive/a906d76e362484d3ca1edbe9683f672838ab70b0.zip on 2020-10-04T22:56:13.913Z<br> - Chen L, Liu B, Wu Z, Jin Q, Yang J, 2017. DRodVir: A resource for exploring the virome diversity in rodents. J Genet Genomics. 44(5):259-264. accessed via https://github.com/globalbioticinteractions/drodvir/archive/0346c0e8d4d66c6400e9965bd6a6aeed24cd7586.zip on 2020-10-04T23:06:04.368Z<br> - Agosti, Donat. 2020. Transcription of Linné, C. von, 1758. Systema naturae per regna tria naturae secundum classes, ordines, genera, species, cum characteribus, differentiis, synonymis, locis. Available at: http://dx.doi.org/10.5962/bhl.title.542 . accessed via https://github.com/globalbioticinteractions/linnaeus1758/archive/a818060080fa04a88dac6df1ae5b897304ae8877.zip on 2020-10-05T00:46:04.852Z<br> - Mollentze, Nardus, & Streicker, Daniel G. (2019). Viral zoonotic risk is homogenous among taxonomic orders of mammalian and avian reservoir hosts (Version 1.0.0) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.3516613 accessed via https://github.com/globalbioticinteractions/mollentze2019/archive/ad12dc74d03c3d992618f16c37cafb7f7ffd9d01.zip on 2020-10-05T00:50:55.878Z<br> - Eneida L. Hatcher, Sergey A. Zhdanov, Yiming Bao, Olga Blinkova, Eric P. Nawrocki, Yuri Ostapchuck, Alejandro A. Schäffer, J. Rodney Brister, Virus Variation Resource – improved response to emergent viral outbreaks, Nucleic Acids Research, Volume 45, Issue D1, January 2017, Pages D482–D490, https://doi.org/10.1093/nar/gkw1065 . accessed via https://github.com/globalbioticinteractions/ncbi-virus/archive/531a8d743d7adcf1153a19087e5d3c5b76750e3e.zip on 2020-10-05T00:53:53.646Z<br> - Olival, K. J., Hosseini, P. R., Zambrana-Torrelio, C., Ross, N., Bogich, T. L., & Daszak, P. (2017). Host and viral traits predict zoonotic spillover from mammals. Nature, 546(7660), 646–650. doi:10.1038/nature22975 accessed via https://github.com/globalbioticinteractions/olival2017/archive/f61070a5339d0e6c6e76d7eb4e2102decb52317d.zip on 2020-10-05T00:56:43.356Z<br> - Pensoft Darwin Core Archives with associateTaxa columns accessed via https://github.com/globalbioticinteractions/pensoft-dwca/archive/ee8831a2a391203f4fa8c05a0ddd927202b234bf.zip on 2020-10-05T00:56:51.868Z<br> - Pensoft Darwin Core Archives available via Integrated Publication Toolkit accessed via https://github.com/globalbioticinteractions/pensoft-ipt/archive/4ad4b47978324681289e36f8c2b247b1bcc97b1a.zip on 2020-10-05T00:58:01.912Z<br> - De Rojas M, Doña J, Dimov I (2020) A comprehensive survey of Rhinonyssid mites (Mesostigmata: Rhinonyssidae) in Northwest Russia: New mite-host associations and prevalence data. Biodiversity Data Journal 8: e49535. https://doi.org/10.3897/BDJ.8.e49535 accessed via https://github.com/globalbioticinteractions/pensoft-table/archive/3488e0397ca4e083d5eca6949951e426a75713e3.zip on 2020-10-05T00:58:03.647Z<br> - Marcus Guidoti, Tatiana Ruschel, Donat Agosti. 2020. Corona virus related biotic associations manually extracted from literature. Plazi. accessed via https://github.com/globalbioticinteractions/plazi-covid19/archive/326578b0d9f974760dcd2e962d86636a6487a6c0.zip on 2020-10-05T00:58:08.025Z<br> - Shaw, LP, Wang, AD, Dylus, D, et al. The phylogenetic range of bacterial and viral pathogens of vertebrates. Mol Ecol. 2020; 29: 3361– 3379. https://doi.org/10.1111/mec.15463 accessed via https://github.com/globalbioticinteractions/shaw2020/archive/bb9ab857b7fdbb4e931752d01b43d37b3ada77cf.zip on 2020-10-05T01:05:23.554Z<br> - OpenBiodiv. 2020. Annotated biotic interaction tables from Pensoft publications. accessed via https://github.com/pensoft/pensoft-interaction-tables/archive/bb7d1dc9f2eba220a61502e06e6114053fd30788.zip on 2020-10-05T03:03:23.372Z<br> - Quentin J. Groom. 2020. Bat interation data manually extracted from literature. accessed via https://github.com/qgroom/batinterations/archive/70108945f9014aa0ac1db920191867f7e151c793.zip on 2020-10-05T03:04:11.533Z</p> <p>Generated on:<br> 2020-10-06</p> <p>by:<br> GloBI's Elton 0.10.2<br> (see https://github.com/globalbioticinteractions/elton).</p> <p> </p> <p>Note that all files ending with .tsv are files formatted<br> as UTF8 encoded tab-separated values files.</p> <p>https://www.iana.org/assignments/media-types/text/tab-separated-values</p> <p><br> Included in this review archive are:</p> <p>README:<br> This file.</p> <p>review_summary.tsv:<br> Summary across all reviewed collections of total number of distinct review comments.</p> <p>review_summary_by_collection.tsv:<br> Summary by reviewed collection of total number of distinct review comments.</p> <p>indexed_interactions_by_collection.tsv:<br> Summary of number of indexed interaction records by institutionCode and collectionCode.</p> <p>review_comments.tsv.gz:<br> All review comments by collection.</p> <p>indexed_interactions_full.tsv.gz:<br> All indexed interactions for all reviewed collections.</p> <p>indexed_interactions_simple.tsv.gz:<br> All indexed interactions for all reviewed collections selecting only sourceInstitutionCode, sourceCollectionCode, sourceCatalogNumber, sourceTaxonName, interactionTypeName and targetTaxonName.</p> <p>datasets_under_review.tsv:<br> Details on the datasets under review.</p> <p>elton.jar:<br> Program used to update datasets and generate the review reports and associated indexed interactions.</p> <p><br> datasets.zip:<br> source datasets collected by elton in process of executing the generate_report.sh script.</p> <p>generate_report.sh:<br> program used to generate the report</p> <p>generate_report.log:<br> log file generated as part of running the generate_report.sh script</p>
The OREGANO knowledge graph for computational drug repurposing
<p>The files here are data files from the OREGANO project, which consists of building a holistic knowledge graph on drugs, including natural compounds. Here is the list of files:</p><p> </p><p>- OREGANO_V2.tsv : The triplet file used for link prediction. 3 columns : Subjet ; Predicate ; Object</p><p>- oreganov2.1_metadata_complet.ttl : The OREGANO knowledge graph in turtle format with the names and cross-references of the various integrated entities.</p><p> </p><p>The following files contain the cross-references of OREGANO entities according to their type. They are all organised as follows: the external sources are the titles of the columns and each line begins with the identifier of the entity in OREGANO :</p><p>- TARGET.tsv: Cross-reference table of the 22,096 targets.<br>- PHENOTYPES.tsv: Cross-reference table of the 11,605 phenotypes.<br>- DISEASES.tsv: Cross-reference table of the 18,333 diseases.<br>- PATHWAYS.tsv: Cross-reference table of the 2,129 pathways.<br>- GENES.tsv: Cross-reference table of the 35,794 genes.<br>- COMPOUND.tsv: Cross-reference table of the 90,868 compounds.<br>- INDICATIONS.tsv: Cross-reference table of the 2,714 indications.<br>- SIDE_EFFECT.tsv: Cross-reference table of the 6,060 side-effects.<br>- ACTIVITY.tsv: Names of the 78 activities.<br>- EFFECT.tsv: Names of the 171 effects.</p><p>The OREGANO knowledge graph is composed of 11 types of nodes and 19 types of links. The current version of the graph contains 88,937 nodes and 824,231 links.</p><p>A SPARQL endpoint has been provided to enable users to retrieve and explore the knowledge graph at <a href="http://91.121.148.199:8889/bigdata/#query">OREGANO SPARQL endpoint</a> .</p><p> </p><p>The integration files and the knowledge graph are available on the GitHub of the OREGANO project in the Integration folder: <a href="https://gitub.u-bordeaux.fr/erias/oregano">Gitub repository</a> .</p><p> </p>
Data from: A FAIR and modular image-based workflow for knowledge discovery in the emerging field of imageomics
<p>Data and results from the Imageomics Workflow. These include data files from the Fish-AIR repository (https://fishair.org/) for purposes of reproducibility and outputs from the application-specific imageomics workflow contained in the Minnow_Segmented_Traits repository (https://github.com/hdr-bgnn/Minnow_Segmented_Traits).</p> <p>Fish-AIR:<br> This is the dataset downloaded from Fish-AIR, filtering for Cyprinidae and the Great Lakes Invasive Network (GLIN) from the Illinois Natural History Survey (INHS) dataset. These files contain information about fish images, fish image quality, and path for downloading the images. The data download ARK ID is dtspz368c00q. (2023-04-05). The following files are unaltered from the Fish-AIR download. We use the following files:</p> <p>extendedImageMetadata.csv: A CSV file containing information about each image file. It has the following columns: ARKID, fileNameAsDelivered, format, createDate, metadataDate, size, width, height, license, publisher, ownerInstitutionCode. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>imageQualityMetadata.csv: A CSV file containing information about the quality of each image. It has the following columns: ARKID, license, publisher, ownerInstitutionCode, createDate, metadataDate, specimenQuantity, containsScaleBar, containsLabel, accessionNumberValidity, containsBarcode, containsColorBar, nonSpecimenObjects, partsOverlapping, specimenAngle, specimenView, specimenCurved, partsMissing, allPartsVisible, partsFolded, brightness, <br> uniformBackground, onFocus, colorIssue, quality, resourceCreationTechnique. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>multimedia.csv: A CSV file containing information about image downloads. It has the following columns: ARKID, parentARKID, accessURI, createDate, modifyDate, fileNameAsDelivered, format, scientificName, genus, family, batchARKID, batchName, license, source, ownerInstitutionCode. Column definitions are defined https://fishair.org/vocabulary.html and the persistent column identifiers are in the meta.xml file.</p> <p>meta.xml: A XML file with the metadata about the column indices and URIs for each file contained in the original downloaded zip file. This file is used in the fish-air.R script to extract the indices for column headers.</p> <p>The outputs from the Minnow_Segmented_Traits workflow are:</p> <p>sampling.df.seg.csv: Table with tallies of the sampling of image data per species during the data cleaning and data analysis. This is used in Table S1 in Balk et al. </p> <p>presence.absence.matrix.csv: The Presence-Absence matrix from segmentation, not cleaned. This is the result of the combined outputs from the presence.json files created by the rule “create_morphological_analysis”. The cleaned version of this matrix is shown as Table S3 in Balk et al.</p> <p>heatmap.avg.blob.png and heatmap.sd.blob.png: Heatmaps of average area of biggest blob per trait (heatmap.avg.blob.png) and standard deviation of area of biggest blob per trait (heatmap.sd.blob.png). These images are also in Figure S3 of Balk et al.</p> <p>minnow.filtered.from.iqm.csv: Filtered fish image data set after filtering (see methods in Balk et al. for filter categories).</p> <p>burress.minnow.sp.filtered.from.iqm.csv: Fish image data set after filtering and selecting species from Burress et al. 2017.</p>
A map selection of wigeon stopover sites (core areas) based on wetland expert knowledge
<p>Stopover areas (core areas only) along the migration route of wigeons tracked with GPS transmitters were selected when they exhibited forests on more than 50% of their total surface or had less than 50% cover by water and/or wetland on the ESA’s global land cover map. We created a sample of 5,630 regions of interest (3,403 for training and 2,227 for validation), delineated with polygons assigned to land classes listed in the Table 1. We used archives of Google Earth, ESRI, and BING satellites for the photointerpretation of the land classes as described in Table 1. The classification was performed with a Sentinel-2 MultiSpectral Instrument, Level-2A image collection in Google Earth Engine (GEE) through the R-package Rgee to create a batch process applying the GEE Random forest classifier to each selected core home range. The cloudless (maximum 3%) images were selected within the period from 01/06/2021 to 30/09/2021. The optimal number of trees was estimated at 100 for an out of bag error of 14%. The overall accuracy on the validation sample was 82 %. </p>
Database of permacultural adoption responses in Mexicali, BC, Mexico. based on Circular Economy, Knowledge Management, and Sustainability policies
<p>Database documenting the perspectives of citizens in Mexicali, Baja California, Mexico, regarding the adoption of permaculture practices. The study is analyzed through the lenses of Knowledge Management, Circular Economy, and Sustainability Policies. The data was collected during the summer of 2024. </p>
UB2030 | The Future of Research Libraries as Knowledge Hubs | Interviews
<p>UB2030 is a podcast about the innovation of the university library through technological changes and shifts in research and education. In this podcast, David Oldenhof and Maurice Vanderfeesten discuss different subjects. They will be accompanied by guests who bring in an outside perspective.</p> <p>This release contains 11 episodes.</p> <p>Follow for more at <a href="https://ubvu.github.io/ub2030/">https://ubvu.github.io/ub2030/</a></p> <p>📖 <a href="https://doi.org/10.5281/zenodo.14615659"><strong>CLICK TO READ THE REPORT</strong></a></p> <p>🎧<strong>Listen on</strong> <a href="https://soundcloud.com/vu-library-live/sets/ub2030-the-future-of-research-libraries"><strong>SoundCloud</strong></a>, <a href="https://open.spotify.com/show/7dgTKn69lE3cnvs7CKv59v"><strong>Spotify</strong></a>, or your favorite <a href="https://antennapod.org/">(open)</a> podcast app.</p> <ul> <li>Authors: Maurice Vanderfeesten, David Oldenhof</li> <li>Client: Joeri Both</li> <li>Organization: <a href="https://vu.nl/nl/over-de-vu/diensten/universiteitsbibliotheek">University Library, Vrije Universiteit Amsterdam</a></li> <li>Date: 2023-05-19</li> </ul> <p><a href="https://doi.org/10.5281/zenodo.14615659" target="_blank" rel="noopener">DOI:10.5281/zenodo.14615659</a> (Rapport)</p> <p><a href="https://doi.org/10.5281/zenodo.10666049" target="_blank" rel="noopener">DOI:10.5281/zenodo.10666049</a> (Data)</p> <p><a href="https://ubvu.github.io/ub2030/">Project Page</a> | <a href="https://soundcloud.com/vu-library-live/sets/ub2030-the-future-of-research-libraries">Listen on SoundCloud</a> | <a href="https://feeds.soundcloud.com/users/soundcloud:users:527805591/sounds.rss">Podcast RSS</a> | <a href="https://forms.office.com/e/KX08BEenpu">Listener Feedback</a></p> <h2>Reason (Why Now)</h2> <p>The world is rapidly changing technologically, around AI, blockchain (NFTs), and linked data. As a library, you want to remain relevant for state-of-the-art research and education. We need to not only implement existing projects but also explore the horizon of opportunities and threats that await us. The client for this project is Joeri Both.</p> <h2>Project Goal (Why and How)</h2> <ul> <li>This project provides input for the next multi-year plan, creating a roadmap with a horizon up to 2030.</li> <li>With this project, we aim to increase the knowledge level of the UB by identifying key innovative/disruptive developments, to become a full-fledged partner for providers of (new) technological applications.</li> <li>From there, we translate innovative developments/trends into UB practice in broad terms, offering suggestions for workable/realistic pilot projects that contribute to the UB ambitions for researchers.</li> <li>The approach is to deliver an innovation sub-report each month, including a podcast episode, with the aim of making innovation an actively discussed topic within the UB.</li> <li>This gives the management team insight into the wide range of possibilities and developments, allowing them to make strategic choices for starting pilots and better embedding the innovation process in the organization.</li> </ul> <h2>Project Scope (What is and isn't included)</h2> <p>Through interviews, we gather information and ideas from each department and external experts. This information is linked to the ambition themes in the multi-year plan, focusing mainly on technological developments and their impact on our work processes and product/service offerings. The collected information is available in the form of an audio recording/podcast and interview report. For the Research Support department, these ideas are further developed into pilot proposals on three implementation levels: short, medium, and long term. The interviews are scheduled per department, divided into the themes of the ambitions in the multi-year plan.</p> <h2>Episodes</h2> <ul> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-01-introductie/"><strong>Episode 01 -- Introduction</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-02-rdm/"><strong>Episode 02 -- Future of Research Data Management and Research Software Management</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-03-research_intelligence/"><strong>Episode 03 -- Future of Research Intelligence</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-04-open_science/"><strong>Episode 04 -- Future of Open Science</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-05-research_support/"><strong>Episode 05 -- Future of Research Support</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-06-digital_services_and_infrastructures/"><strong>Episode 06 -- Future of Digital Services and Infrastructures</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-07-education_support/"><strong>Episode 07 -- Future of Education Support</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-08-special_collections/"><strong>Episode 08 -- Future of Special Collections</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-09-information_services/"><strong>Episode 09 -- Future of Information Services (not public)</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-10-library_desk_services/"><strong>Episode 10 -- Future of Library Desk Services</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-11-aquisition_and_metadata/"><strong>Episode 11 -- Future of Acquisition and Metadata (canceled)</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-12-society/"><strong>Episode 12 -- Future of the Library in Society</strong></a></li> <li><a href="https://github.com/ubvu/ub2030/blob/main/ub2030-bonus-01-desci/"><strong>Episode 13 -- BONUS DeSci: Future of Open Science Ecosystems</strong></a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/ubvu/ub2030/compare/v1.1...v1.6">https://github.com/ubvu/ub2030/compare/v1.1...v1.6</a></p>
Figure 2: Direct and indirect paths of knowledge transfer to New Zealand to manage sand drifting in the nineteenth and twentieth centuries.
<p>Figure 2 of article: Managing Coastal Sand Drift in the Anthropocene: A Case Study of the Manawatū-Whanganui Dune Field, New Zealand, 1800s–2020s</p> <p>DOI zenodo: 10.5281/zenodo.5075980</p>
Local Ecological Knowledge and folk medicine in historical Esthonia, Livonia, Courland and Galicia, 1805-1905
<p>Background: Historical ethnobotanical data can provide valuable information about past human-nature relationships as well as serve as a basis for diachronic analysis. This thesis aims to document medicinal plant uses in the 19th century mentioned in German-language sources in the historical regions of Esthonia, Livonia, Courland and Galicia to analyse the gathered data in regard to plant families and medicinal use categories and finally to qualitatively compare the results with various studies from the study area and surrounding regions with recently acquired data as well as historical data.</p> <p>Methods: Data was mainly obtained by systematic manual search in various relevant historical German-language works focused on the medicinal use of plants. Data about plant and non-plant constituents, their usage, the mode of administration, used plant parts and their German and local names was extracted and collected into a database in the form of Use Reports.</p>
MUHAI Benchmark : Task 2 (Credibility of knowledge-based generated gossip stories)
<p><strong>Meaning and Understanding in Human-Centric AI (MUHAI) Benchmark<br> Task 2 (Credibility of knowledge-based generated gossip stories)</strong></p> <p> </p> <p>This dataset aims at investigating whether the use of Knowledge Graphs has an impact on the credibililty of automatically-generated stories.</p> <p><br> The submission includes the following data:</p> <ol> <li>Generated stories (.txt)</li> <li>Story generation template </li> <li>A tsv file with entities and triples (to be used for generating stories)</li> <li>Evaluation description : Questions and metrics submitted to the users</li> </ol> <p>The "gossip stories" are generated with the T5 languge model fine-tuned on the WebNLG challenge. The model takes the triples (file 3) as input and generates one sentence each. A link prediction algorithm based on Jaccard's similarity learns the likelihood of two entities to be related (3). Then, the narrative continues with automatically generated celebrity background descriptions.</p> <p>The credibility of the story is evaluated using a questionnaire based on Gaziano et. al. The questionnaire was filled in by the test subjects after reading each generated article. One for a KG-generated text where links were predicted using the link prediction and one for text that was generated using triples of random entities (celebrities). </p> <p>Full code available at : https://github.com/kmitd/muhai-credibility-KR</p>
Knowledge gaps on trade-offs of soil carbon sequestration related to soil management strategies
<p>The database contains 87 unique literature items (29 reviews, 42 meta-analyses, 16 original papers) describing the effect of a soil management strategy (tillage management, cropping systems, water management, cover crops, crop residues, livestock manure, slurry, compost, biochar, liming) on the trade-offs between soil carbon sequestration or SOC change and N2O emission, CH4 emission and nitrogen leaching. Since some literature items describe effects of several SMS categories, the database_summary tab comprises a total of 112 unique inputs. For each input it is indicated in the Database_summary tab if it was used as input for the "Soil management effect assessment" in Maenhout et al. (2024) [Maenhout, P., Di Bene, C., Cayuela, M. L., Diaz-Pines, E., Govednik, A., Keuper, F., Mavsar, S., Mihelic, R., O'Toole, A., Schwarzmann, A., Suhadolc, M., Syp, A., & Valkama, E. (2024). Trade-offs and synergies of soil carbon sequestration: Addressing knowledge gaps related to soil management strategies. European Journal of Soil Science, 75(3), e13515. https://doi.org/10.1111/ejss.13515] and/or to define knowledge gaps ("Knowledge gap in tab"-column). Knowledge gaps and research recommendations are gouped per soil management strategy in different tabs in this database. Per soil management strategy, knowledge gaps are clustered per theme in groups. These themes include: the specific soil management strategy, pedoclimatic conditions, establishment of experiments, other soil management strategies, meta-analysis, modelling and other</p>
Data from a three-phase Delphi study used to investigate Knowledge Infrastructure for Research Data in Norway, KIRDN_Data; PhD project
<p>A modified three-phase Delphi study was used to explore the knowledge infrastructure for research data in Norway. The study includes different stakeholders involved in research data sharing. A Delphi study is characterised by the use of an expert panel to elicit opinions on a shared reality from different perspectives. Data collection is performed in several rounds with the intention of reaching consensus or solving an issue. </p> <p>A group of 24 experts took part in the study. The group consisted of policy-makers, representatives of national service providers, and researchers and research support staff from four Norwegian universities. The participants were invited based on their involvement in the development of policies, infrastructure or data-related research support. The research support staff were recruited to include representatives from different research support services at the universities, including libraries, research offices and IT departments. While the researchers were selected from based on their receival of EU funding with requirements of data management plans.</p> <p>Data were collected in three phases. The first phase, the ‘exploration phase’, was conducted using open interviews lasting approximately one hour in January/February 2018. The purpose of this phase was to obtain an initial overview of the panel members opinions’ on issues regarding research data management.</p> <p>In the second phase, the ‘evaluation phase’, conducted in August/September 2018, participants answered a survey containing nine questions on topics such as data stewardship, DMPs, ethical aspects of data sharing and core functions in a research data infrastructure. The survey was designed to further explore issues and tensions uncovered in the first interviews. Several of the questions were formulated as statements that the participants were asked to agree or disagree upon. </p> <p>The third, ‘concluding phase’ was conducted using interviews in March/April 2019. These interviews lasted approximately 30 minutes and were based on results from the questionnaire as well as the first interview. Participants were asked whether they had thoughts on the preliminary findings of the study. </p> <p>Based on requests from some of the participants, the questions were sent to all participants prior to the data collection, in all three phases. The participants were also sent the transcripts from the interviews and were asked for permission to share the complete material or parts of the data material to which they contributed. </p>
An analysis of meiofauna knowledge generated by Latin American researchers
<p>Bibliographic databases used to analyse the document production of benthic meiofauna in Latin American countries. To be opened on R, bibliometrix package.</p> <p> </p>
Dataset for Creating a Scholarly Knowledge Graph from Survey Article Tables
<ul> <li><strong>Selected papers.csv</strong><br> This file lists all selected survey papers used to create the knowledge graph. The file contains paper titles, table references (of the tables that are extracted), sources and the reference to the survey paper (either a DOI, or a full textual reference)<br> </li> <li><strong>ORKG comparisons.csv</strong><br> All comparisons imported in the ORKG are listed in this file. Per survey paper, multiple tables could be extracted, and therefore multiple comparisons are created. The files lists the internal IDs and the URL to the comparisons. <br> </li> <li><strong>Ingested papers.csv</strong><br> A full list of all individual papers extracted from the survey articles. The file contains paper titles and their respective URL in the ORKG. </li> </ul>
RDF Reification Benchmark (REF) using the Biomedical Knowledge Repository (BKR)
<p>This resource can be used for benchmarking different RDF modelling solutions for statement-level metadata, namely: </p> <p>- RDF Reification,</p> <p>- Singleton Property,</p> <p>- RDF* (RDF-star). </p> <p> </p> <p>More details about this resource can be found in the following publication:</p> <p>Fabrizio Orlandi, Damien Graux, Declan O'Sullivan, "Benchmarking RDF Metadata Representations: Reification, Singleton Property and RDF*", <em>15th IEEE International Conference on Semantic Computing (ICSC)</em>, 2021.</p> <p>Pre-print available at: http://fabriziorlandi.net/pdf/2021/ICSC2021_REF-Benchmark.pdf</p> <p> </p> <p>The dataset contains 3 different versions of the Biomedical Knowledge Repository (BKR) knowledge graph, as described in:</p> <p>Vinh Nguyen, Olivier Bodenreider, Amit Sheth. "Don't Like RDF Reification? Making Statements About Statements Using Singleton Property" WWW 2014, doi: 10.1145/2566486.2567973.</p> <p>and,</p> <p>Satya S. Sahoo, Olivier Bodenreider, Pascal Hitzler, Amit Sheth and Krishnaprasad Thirunarayan. "Provenance Context Entity (PaCE): Scalable Provenance Tracking for Scientific RDF Data" in Sci Stat Database Manag. 2010; 6187: 461–470. doi: 10.1007/978-3-642-13818-8_32</p> <p> </p> <p>The 3 knowledge graphs dumps are packaged as Gzipped RDF files in Turtle (and Turtle*) syntax. </p> <p>BKR-R-fullKGdump.ttl.gz for the Reification method,</p> <p>BKR-S-fullKGdump.ttl.gz for the Singleton method,</p> <p>BKR-star-fullKGdump.ttls.gz for the RDF* (RDF-star) method.</p> <p> </p> <p>The RDF REiFication Benchmark (REF) includes also a set of SPARQL (and SPARQL*) queries that can be used to compare the performance of different triplestores.</p> <p>Details about the SPARQL queries, and the queries themselves, are included in the "REF-Benchmark.tar.gz" archive. The queries are named after the dataset they are designed for (BKR-R or BKR-S or BKR-star), plus they include a letter identifying the query set, and a query number. </p> <p>E.g. the query in the file "BKR-R_F-Q3.rq" is for the BKR-R (standard reification) dataset, it is part of the query set "F" and it is the number 3 of that set "F". Hence, the same query, but translated for the RDF* dataset in SPARQL* syntax, is contained in "BKR-star_F-Q3.rq".</p> <p>Sets "A" and "B" are derived from the queries introduced by V. Nguyen et al. in: "Don't Like RDF Reification? Making Statements About Statements Using Singleton Property" WWW 2014, doi: 10.1145/2566486.2567973. Set "F" has been designed more with RDF* in mind as part of this benchmark (see [Orlandi et al., ICSC 2021]) </p> <p> </p> <p> </p> <p> </p>
Node2Vec model - Czech Wikidata (knowledge graph / concepts / l80 / rw40)
<p>Node2Vec embedding model trained on Czech wikidata (from October 2020) concepts using gensim implementation of Word2Vec with the following parameters for random walks:</p> <ul> <li>length of walk = 80</li> <li>number of random walks = 40</li> </ul>
Node2Vec model - Czech Wikidata (knowledge graph / concepts / l40 / rw10)
<p>Node2Vec embedding model trained on Czech wikidata (from October 2020) concepts using gensim implementation of Word2Vec with the following parameters for random walks:</p> <ul> <li>length of walk = 40</li> <li>number of random walks = 10</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.