Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
190
datasets available to search
ShareScore release 0.9.0
Dataset results
190 results for “Knowledge Base”
FIGURE 8 in Contribution to the knowledge of Chinese Gryllacrididae (Orthoptera: Ensifera: Stenopelmatoidea) XXVI: New additions based on the specimens from Southern China
FIGURE 8. Larnaca (Larnaca) nigrimargis sp. nov. Male: A. head in frontal view; B–C. head and pronotum: B. dorsal view, C. lateral view; D. wings in lateral view; E. second and third abdominal tergites in lateral view; F–I. apex of abdomen: F. lateral view, G. apical and ventral view, H. apical view, I. ventral view.
FIGURE 7 in Contribution to the knowledge of Chinese Gryllacrididae (Orthoptera: Ensifera: Stenopelmatoidea) XXVI: New additions based on the specimens from Southern China
FIGURE 7. Glolarnaca nigrimacula Yang, Lu & Bian, 2021. Female: A. head in frontal view; B–C. head and pronotum: B. dorsal view, C. lateral view; D. second and third abdominal tergites in lateral view; E. subgenital plate in ventral view; F. seventh abdominal sternite in ventral view; G. apex of abdomen in lateral view.
FIGURE 5 in Contribution to the knowledge of Chinese Gryllacrididae (Orthoptera: Ensifera: Stenopelmatoidea) XXVI: New additions based on the specimens from Southern China
FIGURE 5. Dracogryllacris zhoui (Pang, Zhang & Bian, 2023). Female: A. head in frontal view; B–C. head and pronotum: B. dorsal view, C. lateral view; D. second and third abdominal tergites in lateral view; E. hind tibiae in lateral view; F. apex of abdomen in lateral view; G. apices of ovipositor in lateral view; H. seventh abdominal sternite and subgenital plate in ventral view.
FIGURE 6 in Contribution to the knowledge of Chinese Gryllacrididae (Orthoptera: Ensifera: Stenopelmatoidea) XXVI: New additions based on the specimens from Southern China
FIGURE 6. Glolarnaca flata sp. nov. Female: A. head in frontal view; B–C. head and pronotum: B. dorsal view, C. lateral view; D. second and third abdominal tergites in lateral view; E. subgenital plate in ventral view; F. seventh abdominal sternite in ventral view; G. apex of abdomen in lateral view.
FIGURE 2 in Contribution to the knowledge of Chinese Gryllacrididae (Orthoptera: Ensifera: Stenopelmatoidea) XXVI: New additions based on the specimens from Southern China
FIGURE 2. Apterolarnaca nigrifrontis Bian & Shi, 2016. Female: A. head in frontal view; B–C. head and pronotum: B. dorsal view, C. lateral view; D. second and third abdominal tergites in lateral view; E. hind femur in lateral view; F. ovipositor in lateral view; G. seventh abdominal sternite and subgenital plate in ventral view.
FIGURE 3 in Contribution to the knowledge of Chinese Gryllacrididae (Orthoptera: Ensifera: Stenopelmatoidea) XXVI: New additions based on the specimens from Southern China
FIGURE 3. Apterolarnaca quadrimaculata Bian & Shi, 2016. Male: A. head in frontal view; B–C. head and pronotum: B. dorsal view, C. lateral view; D. second and third abdominal tergites in lateral view; E. hind leg in lateral view; F–I. apex of abdomen: F. lateral view, G. dorsal view, H. apical and ventral view, I. ventral view.
FIGURE 4 in Contribution to the knowledge of Chinese Gryllacrididae (Orthoptera: Ensifera: Stenopelmatoidea) XXVI: New additions based on the specimens from Southern China
FIGURE 4. Apterolarnaca quadrimaculata Bian & Shi, 2016. Female: A. head in frontal view; B–C. head and pronotum: B. dorsal view, C. lateral view; D. second and third abdominal tergites in lateral view; E. hind leg in lateral view; F. seventh abdominal sternite and subgenital plate in ventral view; G. apex of abdomen in lateral view.
FIGURE 1 in Contribution to the knowledge of Chinese Gryllacrididae (Orthoptera: Ensifera: Stenopelmatoidea) XXVI: New additions based on the specimens from Southern China
FIGURE 1. Apterolarnaca nigrifrontis Bian & Shi, 2016. Male: A. head in frontal view; B–C. head and pronotum: B. dorsal view, C. lateral view; D. right hind leg in lateral view; E. second and third abdominal tergites in lateral view; F–I. apex of abdomen: F. lateral view, G. dorsal and slightly apical view, H. apical view, I. ventral view.
Knowledge base for rice plant disease diagnosis
<p>This dataset describes symptoms, disease, and their relationships in rice plant disease.</p>
FreeCiv games played by Knowledge-based Reinforcement Learning
<p>The dataset contains 600 fully played games of FreeCiv game. </p>
Descriptive Statistics for Estimates on Knowledge of Gender-Based Violence in Early Childhood
<p><strong>Table 2:</strong></p> <p><strong>2 of 12</strong></p> <p><strong>Descriptive Statistics for Estimates on Knowledge of Gender-Based Violence in Early Childhood</strong></p>
Artifacts for paper "CITYWALK: Enhancing LLM-Based C++ Unit Test Generation via Project-Dependency Awareness and Language-Specific Knowledge" submitted to TOSEM
<p>The project includes the data and code used in the submitted TOSEM paper titled "CITYWALK: Enhancing LLM-Based C++ Unit Test Generation via Project-Dependency Awareness and Language-Specific Knowledge"</p>
FIGURE 5 in Enabling comparisons of characters using an Xper2 based knowledge-base of fern morphology
FIGURE 5. Phylogenetic character written in the required syntax. The first "Echantillonnage" (i.e., sampling) section indicates the taxa included in the analysis. The second "Representation-3ia" section gives the list of phylogenetic characters labeled with a number between square brackets (e.g., [1]). Hierarchical characters are given in parenthetical syntax. Each state is composed of two numbers: the first is the descriptor label in the Xper² knowledge base (e.g., 59 stands for the "cauline symmetry" descriptor), the second is the descriptor state (e.g., 59:0 stands for the "0: radial" state of the 59th descriptor).
FIGURE 4. Interactive identification key. A in Enabling comparisons of characters using an Xper2 based knowledge-base of fern morphology
FIGURE 4. Interactive identification key. A: List of descriptors sorted according to their discriminant power. B: Definition of the selected descriptor. C: List of states for the selected descriptor. D: Illustration of the different descriptor states. E: List of the discarded taxa during the process of identification. F: List of the remaining taxa, still candidate to be identified. G: Identified taxon displayed at the end of the identification process.
FIGURE 1 in Enabling comparisons of characters using an Xper2 based knowledge-base of fern morphology
FIGURE 1. Descriptions provided as interactive windows thanks to Xper² interface. A: List of taxa. B: Taxon complementary data (e.g., author's name). C: Descriptive model (i.e., list of characters with parent-child relationships). D: Definition and reference of the character, ontology ID for major concepts and related terms. E: Summary of the parent-child relationships for the character. F: List of Character states with the value(s) selected for each taxon (including unknown data). G: Character state complementary data (e.g., illustration).
A unified DTI prediction framework based on knowledge graph and recommendation system
<p>## A unified DTI prediction framework based on knowledge graph and recommendation system</p> <p> </p> <p># Code and data description</p> <p>## Scripts</p> <p>- `kge_nfm.py`: the complement of the KGE_NFM & NFM methods.</p> <p>- `kge_rf.py`: the complement of the KGE_RF & RF methods.</p> <p>- `deepdit.py`: the complement of the MPNN_CNN & DeepDTI methods.</p> <p>- the complement of DTINet and DTiGEMS is tested based on their source packages (more in Prerequisites)</p> <p><br> </p> <p>## `data/` directory</p> <p>#### `yamanishi_08/` directory</p> <p>- `data_folds/`: 10 folds training set and test set in the three scenarios</p> <p>- `warm_start_1_1/`</p> <p>- `warm_start_1_10/`</p> <p>- `drug_coldstart/`</p> <p>- `protein_coldstart/`</p> <p>- `kg_data/`: supporting knowledge graph data</p> <p>- `dt_all_08.csv`: whole DTI dataset</p> <p>- `791drug_struc.csv`: drugbank id and smiles of drugs</p> <p>- `989proseq.csv`: kegg id and sequences of proteins</p> <p>- `morganfp.txt`: list of drug morgan fingerprints</p> <p>- `pro_ctd.txt`: list of protein descriptors</p> <p> </p> <p>#### `BioKG/` directory</p> <p>- `data_folds/`: 10 folds training set and test set in the three scenarios</p> <p>- `warm_start_1_10/`</p> <p>- `drug_coldstart/`</p> <p>- `protein_coldstart/`</p> <p>- `kg.csv`: supporting knowledge graph data</p> <p>- `dti.csv`: whole DTI dataset</p> <p>- `comp_struc.csv`: drugbank id and smiles of drugs</p> <p>- `pro_seq.csv`: sequences of proteins</p> <p>- `fp_df.csv`: list of drug morgan fingerprints</p> <p>- `prodes_df.csv`: list of protein descriptors</p> <p> </p> <p>#### `hetionet/` directory</p> <p>- `data_folds/`: 10 folds training set and test set in the three scenarios</p> <p>- `warm_start_1_10/`</p> <p>- `drug_coldstart/`</p> <p>- `protein_coldstart/`</p> <p>- `kg.csv`: supporting knowledge graph data</p> <p>- `dti.csv`: whole DTI dataset</p> <p>- `map_drugs_df`: drugbank id and smiles of drugs</p> <p>- `pro_seq.csv`: sequences of proteins</p> <p>- `fp_df.csv`: list of drug morgan fingerprints</p> <p>- `prodes_df.csv`: list of protein descriptors</p> <p> </p> <p>#### `luo's_dataset/` directory</p> <p>- `data_folds/`: 10 folds training set and test set in the three scenarios</p> <p>- `warm_start_1_1/`</p> <p>- `warm_start_1_10/`</p> <p>- `drug_coldstart/`</p> <p>- `protein_coldstart/`</p> <p>- `mapping/`: related mappings and similarity matrix (https://github.com/luoyunan/DTINet)</p> <p>- `protein.txt`: list of protein names</p> <p>- `disease.txt`: list of disease names</p> <p>- `se.txt`: list of side effect names</p> <p>- `drug_dict_map`: a complete ID mapping between drug names and DrugBank ID</p> <p>- `protein_dict_map`: a complete ID mapping between protein names and UniProt ID</p> <p>- `Similarity_Matrix_Drugs.txt` : Drug similarity scores based on chemical structures of drugs</p> <p>- `Similarity_Matrix_Proteins.txt` : Protein similarity scores based on primary sequences of proteins</p> <p>- `feature/`: related features used in methods</p> <p>- `drug_smiles.csv`: drugbank id and smiles</p> <p>- `seq.txt`: list of protein sequences</p> <p>- `morganfp.txt`: list of drug morgan fingerprints</p> <p>- `pro_ctd.txt`: list of protein descriptors</p> <p> </p> <p>#### `eg_model/` directory</p> <p>We provided a pre-trained kge model for example.</p> <p>- `dismult_400_warm_1_10.pkl`</p> <p><br> </p> <p># Prerequisites</p> <p>#### Operating system: Linux</p> <p>#### Programing language: python</p> <p>#### KGE_NFM & NFM dependencies</p> <p>```</p> <p>- python 3.6</p> <p>- pandas '1.1.5'</p> <p>- numpy '1.18.4'</p> <p>- scikit-learn '0.24.1'</p> <p>- tensorflow '1.15.0'</p> <p>- ampligraph '1.3.2'</p> <p>- deepctr '0.8.4'</p> <p>```</p> <p>#### baseline dependencies</p> <p>- RF & KGE_RF (included in KGE_NFM&NFM dependencies)</p> <p>- MPNN_CNN & DeepDTI:</p> <p>- source: https://github.com/kexinhuang12345/DeepPurpose</p> <p>```</p> <p>- deeppurpose '0.0.9'</p> <p>- torch '1.6.0+cu101'</p> <p>```</p> <p>- DTINet:</p> <p>- source: https://github.com/luoyunan/DTINet</p> <p>- note: in this work, we run the DTINet in a python environment, which need Linux system and python2. Importantly, this method requires the [Inductive Matrix Completion](http://bigdata.ices.utexas.edu/software/inductive-matrix-completion/) (IMC) library. More detailed information about the installation of this method could be found in the source code of the DTINet.</p> <p>- DTiGEMS:</p> <p>- source: https://github.com/MahaThafar/DTiGEMSplus</p> <p>- TriModel:</p> <p>- source: http://drugtargets.insight-centre.org/</p> <p><br> <br> </p> <p># Example (kge_nfm.py)</p> <p> </p> <p>#### A brief presentation of the results:</p> <p>- return average loss when training kge model</p> <p>```</p> <p>Average Loss: 0.475181: 2%|###3 | 1/50 [01:10<57:31, 70.44s/epoch]</p> <p>```</p> <p>- return performance(mrr) on training set of DTI for early stopping (kge_model in `eg_model/`)</p> <p>```</p> <p>In [35]: roc = roc_auc(test_label,test_score)</p> <p>...: pr = pr_auc(test_label,test_score)</p> <p>...: print(roc)</p> <p>...: print(pr)</p> <p>0.8731770833333332</p> <p>0.44079654835037246</p> <p>```</p> <p> </p> <p>- nfm training process (`patience=10`)</p> <p> </p> <p>```</p> <p>In [45]: roc_nfm,pr_nfm,pred_y = train_nfm(feature_columns,train_model_input,train_label,test_model_input,test_label,patience)</p> <p>Train on 44851 samples</p> <p>Epoch 1/2000</p> <p>44851/44851 - 2s - loss: 0.5332 - precision: 0.0976</p> <p>Epoch 2/2000</p> <p>44851/44851 - 1s - loss: 0.4143 - precision: 0.0000e+00</p> <p>Epoch 3/2000</p> <p>44851/44851 - 1s - loss: 0.3456 - precision: 0.0000e+00</p> <p>Epoch 4/2000</p> <p>44851/44851 - 1s - loss: 0.3443 - precision: 0.0000e+00</p> <p>Epoch 5/2000</p> <p>44851/44851 - 1s - loss: 0.3470 - precision: 0.0000e+00</p> <p>Epoch 6/2000</p> <p>44851/44851 - 1s - loss: 0.3382 - precision: 0.0000e+00</p> <p>......</p> <p>Epoch 279/2000</p> <p>44851/44851 - 1s - loss: 0.0758 - precision: 0.9248</p> <p>Epoch 280/2000</p> <p>44851/44851 - 1s - loss: 0.0753 - precision: 0.9327</p> <p>Epoch 281/2000</p> <p>44851/44851 - 1s - loss: 0.0796 - precision: 0.9155</p> <p>Epoch 282/2000</p> <p>44851/44851 - 1s - loss: 0.0764 - precision: 0.9276</p> <p>Epoch 283/2000</p> <p>44851/44851 - 1s - loss: 0.0739 - precision: 0.9127</p> <p>```</p> <p> </p> <p>- reutrn results as type of roc_auc & pr_auc</p> <p>```</p> <p>0.9812476679104477</p> <p>0.8803416284646345</p> <p>```</p>
Dataset and code for 'AI-based Knowledge Extraction from the Bioprinting Literature for identifying technology trends'
<p>Zip file containing the dataset and code for the paper 'AI-based Knowledge Extraction from the Bioprinting Literature for identifying technology trends'. The dataset is composed of:</p> <ul> <li>A train_data.csv file, containing all annotated keywords used for classifier training.</li> <li>A filt_ls.pkl file, containing the sentences used to train the embeddings model.</li> <li>A train.py file, to train the composite keyword annotation model.</li> </ul> <p>The authors acknowledge the supported by the European Union’s Horizon 2020 research and innovation program under the project GIOTTO: “Giotto: Active ageing and osteoporosis: The next challenge for smart nanobiomaterials and 3D technologies,” grant agreement no. 814410.</p>
Relevant Datasets and Software Used for Paper "KGML-xDTD: A Knowledge Graph-based Machine Learning Framework for Drug Treatment Prediction and Mechanism Description"
<p>This repository contains relevant datasets and software used in a paper "KGML-xDTD: A Knowledge Graph-based Machine Learning Framework for Drug Treatment Prediction and Mechanism Description". They are used to run the code of <em>KGML-xDTD </em>stored on <a href="https://github.com/chunyuma/KGML-xDTD">Github</a> and support the results of this paper.</p> <p><strong>About the datasets</strong></p> <p>1. <em>bkg_rtxkg2c_v2.7.3.tar.gz</em></p> <p>This tar.gz file contains three sub-folders: tsv_files, scripts, and relevant_dbs. The "tsv_files" sub-folder has the input files that the neo4j software uses. The "scripts" sub-folder contains a shell script with a relevant python script to construct the biomedical knowledge graph. The "relevant_dbs" sub-folder stores two auxiliary databases that <em>KGML-xDTD</em> needs to use. </p> <p>2. <em>indication_paths.yaml</em></p> <p>This file contains the <a href="https://sulab.github.io/DrugMechDB">DrugMechDB</a> MOA paths that we used to evaluate the predicted MOA paths by <em>KGML-xDTD. </em>It is downloaded from the official <a href="https://github.com/SuLab/DrugMechDB">GitHub repository</a> of DrugMechDB.</p> <p>3. <em>training_data.tar.gz</em></p> <p>This tar.gz file contains the processed training data of four data sources (e.g., <a href="https://mychem.info">MyChem</a>, <a href="https://lhncbc.nlm.nih.gov/ii/tools/SemRep_SemMedDB_SKR/SemMedDB_download.html">SemMedDB</a>, <a href="https://bioportal.bioontology.org/ontologies/NDFRT">NDF-RT</a>, <a href="https://unmtid-shinyapps.net/shiny/repodb/">RepoDB</a>) mentioned in the paper. These processed drug-disease pairs have been matched to the identifiers of biological entities used in our biomedical knowledge graph and respectively split into true positive (tp) sets and true negative (tn) sets. We also provide the names of these drug identifiers and disease identifiers under a sub-folder "translated _to_name".</p> <p><strong>About the software</strong></p> <p><em>neo4j-community-3.5.26.tar.gz</em></p> <p>This tar.gz is the Neo4j community version 3.5.26 downloaded from <a href="https://neo4j.com/download-center/#community">Neo4j Download Center</a>. Although the newer versions are available, due to their big changes in the Neo4j setting that are not compatible with our scripts on Github, we provide the version that we used in our research. If you would like to use the newer version, modifications to our script will be required to import the biomedical knowledge graph into your local Neo4j database with the new setting.</p>
Mobile Feature-oriented Knowledge Base Generation Using Knowledge Graphs
<p>Sample study data set and derived metadata and evaluation files for the research titled 'Mobile Feature-oriented Knowledge Base Generation Using Knowledge Graphs' (MApp-KG resource)</p>
A federated learning framework based on transfer learning and knowledge distillation for targeted advertising-Click-Through Rate Prediction Dataset
<p>https://www.kaggle.com/c/avazu-ctr-prediction</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.