Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
49
datasets available to search
ShareScore release 0.9.0
Dataset results
49 results for “Integrated database”
Integrating QSAR models predicting acute contact toxicity and mode of action profiling in honey bees (A. mellifera): Data curation using open source databases, performance testing and validation
<p>This excel file (DOI: <a href="https://doi.org/10.5281/zenodo.3755675">https://doi.org/10.5281/zenodo.3755675</a>) provides the collection of raw data used for developing the first integrative Quantitative Structure-Activity Relationship (QSAR) model using EFSA's OpenFoodTox, US-EPA ECOTOX and Pesticide Properties DataBase i) to predict acute contact toxicity (LD<sub>50</sub>) and ii) to profile the Mode of Action (MoA) of pesticides active substances in honey bees (<em>Apis mellifera</em>)<em>. </em>Chemical identifiers (e.g. SMILES, CAS n., InChI) and acute contact toxicity data (LD<sub>50</sub>) on honey bees were used to develop and validate i) a two-category QSAR model (toxic/non-toxic; n=411) (sensitivity =0.93), specificity =0.85), balanced accuracy =0.90), Matthews correlation coefficient MCC=0.78), and ii) a regression-based model (n=113) (R2=0.74; MAE=0.52). Similarly, current study proposes the first MoA profiling for 113 pesticides active substances and the first harmonised MoA classification scheme for acute contact toxicity in honey bees, including LD<sub>50s</sub> data points from three different databases such as EFSA's OpenFoodTox, US-EPA ECOTOX and Pesticide Properties DataBase. Such classification allows to further define MoAs and the target site of Plant Protection Products (PPPs) active substances, thus enabling regulators and scientists to refine chemical grouping and toxicity extrapolations for single chemicals and component-based mixture risk assessment of multiple chemicals.</p> <p>The full data collection and analysis of QSAR models, toxicity data (LD<sub>50</sub>) and Mode of Action (Moa) data are described in Carnesecchi et al., 2020 (DOI: doi.org/10.1016/j.scitotenv.2020.139243).</p> <p>This work was supported by the European Food Safety Authority (EFSA) [contract number: OC/EFSA/SCER/2018/01 and NP/EFSA/AFSCO/2016/02 (Edoardo Carnesecchi)].</p>
ToxicoDB: an integrated database to mine and visualize large-scale toxicogenomic datasets (TGGATEs human dataset)
<p>This data was generated by Igarashi Y, Nakatsu N, Yamashita T, Ono A, Ohno Y, Urushidani T, Yamada H. Open TG-GATEs: a large-scale toxicogenomics database. Nucleic Acids Res [Internet]. 2015 Jan;43(Database issue):D921–7. Available from: http://dx.doi.org/10.1093/nar/gku955 PMCID: PMC4384023. The data have been curated and analyzed using our open-source R package, <em>ToxicoGx</em> (<a href="http://bioconductor.org/packages/devel/bioc/html/ToxicoGx.html">http://bioconductor.org/packages/devel/bioc/html/ToxicoGx.html</a>), and are available publicly in the <em>ToxicoDB </em>web application (<a href="http://www.toxicodb.ca">www.toxicodb.ca</a>).</p>
Data from: Towards an integrated database on Canadian ocean resources: benefits, current states, and research gaps
Oceanic ecosystem services support a range of human benefits, and Canada has extensive research networks producing growing data sets. We present a first effort to compile, link, and harmonize available information to provide new perspectives on the status of Canadian ocean ecosystems and corresponding research. The metadata database currently includes 1094 individual assessments and data sets from government (n = 716), nongovernment (n = 320), and academic sources (n = 58), comprising research on marine species, natural drivers and resources, human activities, ecosystem services, and governance, with data sets spanning 1979–2012 on average. Overall, research shows a strong prevalence towards single-species fishery studies, with an underrepresentation of economic and social aspects, and of the Arctic region in general. Nevertheless, the number of studies that are multispecies or ecosystem-based have increased since the 1960s. We present and discuss two illustrative case studies — marine protected area establishment in Canada and herring resource use by the Heiltsuk First Nation — highlighting the potential use of multidisciplinary data sets drawn from metadata records. Identifying knowledge gaps is key to achieving the comprehensive, accessible and interdisciplinary data sets and subsequent analyses necessary for new sustainability policies that meet both ecological and socioeconomic needs.
Integrated Protein-Ligand Interaction Database
<p><strong>IPLID</strong> integrates protein-ligand interaction data from multiple well-known resources, including BindingDB, ChEMBL, DrugBank, GPCRDB, PubChem, LINCS-HMS KinomeScan, and four published kinome assay results. Our database can facilitate projects in <em>machine learning or deep learning-based drug development </em>and other applications by providing integrated data sets appropriate for many research interests. Our database can be utilized for small-scale (e.g. kinases or GPCRs only) and large-scale (e.g. proteome-wide), qualitative or quantitative projects. With its ease of use and straightforward data format, IPLID offers a great educational resource for computer science and data science trainees who lack familiarity with chemistry and biology.</p> <p> </p> <p>Data statistics</p> <p>Target (data type) Activities | Unique chemicals | Unique proteins | File name</p> <p>All (binary) 96318 | 18107 | 3107 | integrated_binary_activity.tsv</p> <p>All (numerical) 2798365 | 683009 | 5876 | integrated_continuous_activity.tsv</p> <p>CYP450 (binary) 67552 | 17273 | 47 | integrated_cyp450_binary.tsv</p> <p>CRT (binary) 4152 | 1219 | 412 | integrated_cancer_related_targets_binary.tsv</p> <p>CDT (binary) 519 | 349 | 88 | integrated_cardio_targets_binary.tsv</p> <p>DRT (binary) 4433 | 1325 | 852 | integrated_disease_related_targets_binary.tsv</p> <p>FDA (binary) 6217 | 1521 | 592 | integrated_fda_approved_targets_binary.tsv </p> <p>GPCR (binary) 1958 | 545 | 129 | integrated_gpcr_binary.tsv</p> <p>NR (binary) 1335 | 657 | 264 | integrated_nr_binary.tsv</p> <p>PDT (binary) 1469 | 674 | 404 | integrated_potential_drug_targets_binary.tsv</p> <p>TF (binary) 1966 | 998 | 304 | integrated_tf_binary.tsv</p> <p> </p> <p>*Abbreviations: CYP450 (Cytochrome P450), CRT (Cancer-Related Target), CDT (Cardiovascular Disease candidate Target), DRT (Disease-Related Target), FDA (FDA-approved target), GPCR (G-Protein Coupled Receptor), NR (Nuclear Receptor), PDT (Potential Drug Target), TF (Transcription Factor)</p> <p>*These protein classifications are from UniProt database and the Human Protein Atlas (<a href="https://www.proteinatlas.org/">https://www.proteinatlas.org/</a>)</p> <p><a href="https://github.com/XieResearchGroup/DrugTargetInteraction/blob/master/iplid/IPLID_data_stat.png">IPLID data statistics</a></p>
Integrated Protein-Ligand Interaction Database
<p>Computational prediction of genome-wide protein-ligand interactions plays a key role in drug discovery, toxicology, and in many other applications. Despite recent advances in <em>deep learning</em>, the large quantity of high-quality data required for training and evaluating models has impeded its applications in computational drug development. To obviate this problem, we have developed an <em>integrated Protein-Ligand Interaction Database </em>(<strong>IPLID</strong>). <strong>IPLID</strong> integrates protein-ligand interaction data from multiple well-known resources, including BindingDB, ChEMBL, DrugBank, GPCRDB, PubChem, LINCS-HMS KinomeScan, and four published kinome assay results. <strong>IPLID</strong> is enabled with search functionalities specifically designed for machine learning, particularly deep learning projects. Users can retrieve numerically or binary labeled (e.g. pki, pkd, or binary) protein-ligand interaction data for different classes of proteins (e.g. GPCRs, kinases, FDA-approved targets, protein products of cancer-related genes, etc.). To facilitate the development of benchmarks for training and testing of machine learning algorithms, it also provides chemical-chemical structure similarity scores calculated by a well-established method, Tanimoto coefficient (Jaccard similarity) of two Extended Connectivity Fingerprint (ECFP4) molecular representations. Protein sequence similarities by BLAST score comparison and position-specific scoring matrices against UniRef50 sequence database are also available for more complicated protein-ligand interaction modeling projects. We believe our database can facilitate projects in <em>machine learning or deep learning-based drug development</em> and other applications by providing integrated data sets appropriate for many research interests. Our database can be utilized for small-scale (e.g. kinases or GPCRs only) and large-scale (e.g. proteome-wide), qualitative or quantitative projects. With its ease of use and straightforward data format, <strong>IPLID</strong> offers a great educational resource for computer science and data science trainees who lack familiarity with chemistry and biology.</p> <p> </p> <ul> <li>Activities are in <em>tab-delimited</em> text file formats (.tsv).</li> <li>Binary activities are under '<em>binary_activity</em>' directory, and numerical activities are under '<em>numerical_activity</em>' directory.</li> <li>File names are in "(<strong>targets</strong>)_(<strong>activity_type</strong>).tsv"</li> <li>Long target names are abbreviated; abbreviations listed below.</li> <li>Ligand-ligand similarity scores are under '<em>ligand_info</em>' directory.</li> <li>Protein-protein similarity scores and position-specific scoring matrices are under '<em>protein_info</em>' directory.</li> <li>Primary ligand-id and protein-id are <em>InChIKey</em> and <em>UniProt ID</em>, respectively.</li> </ul> <p>*<strong>Abbreviations</strong>: CYP450 (Cytochrome P450), CRT (Cancer-Related Target), CDT (Cardiovascular Disease candidate Target), DRT (Disease-Related Target), FDA (FDA-approved target), GPCR (G-Protein Coupled Receptor), NR (Nuclear Receptor), PDT (Potential Drug Target), TF (Transcription Factor)</p> <p>*These protein classifications are from UniProt database and the Human Protein Atlas (<a href="https://www.proteinatlas.org/">https://www.proteinatlas.org/</a>)</p>
APPENDIX. List of sequenced specimens of Triphosa, with identification, Sampling sites collecting data, Accession numbers, and process ID in BOLD database. Data taken from BOLD and generated by Axel Hausmann (1); Bernd Müller (2); Dirk Stadie (3); Iva Mihoci 4); Marco Infusino, Stefano Scalercio (5); Norbert Poell (6); Wanke et al. (7). in An integrative taxonomic revision of the genus Triphosa Stephens, 1829 (Geometridae: Larentiinae) in the Middle East and Central Asia, with description of two new species
APPENDIX. List of sequenced specimens of Triphosa, with identification, Sampling sites collecting data, Accession numbers, and process ID in BOLD database. Data taken from BOLD and generated by Axel Hausmann (1); Bernd Müller (2); Dirk Stadie (3); Iva Mihoci 4); Marco Infusino, Stefano Scalercio (5); Norbert Poell (6); Wanke et al. (7).
TCMID: Traditional Chinese Medicine integrative database for herb molecular mechanism analysis
<p><strong>ABSTRACT: </strong>Traditional Chinese Medicines Integrated Database and the description about Chinese herbs, including English and Latin names, properties, meridians, medicinal parts, herbal effect and indication. Traditional Chinese Medicine (TCM) is a system of healthcare and healing that has been practiced for thousands of years in China. It is based on a holistic approach that views the human body and its various systems as interconnected. TCM encompasses a wide range of practices, including herbal medicine, acupuncture, massage (tui na), exercise (qigong), and dietary therapy.</p> <p><strong>Instruction: </strong></p> <p>Data was cleaned and duplicates were removed.</p> <p><strong>Inspiration: </strong>The dataset was uploaded to UBRITE for "DGR_DEPOT" summer 2023 team project</p> <p><strong>Acknowledgements: </strong>Ruichao Xue 1, Zhao Fang, Meixia Zhang, Zhenghui Yi, Chengping Wen, Tieliu Shi</p> <p>TCMID: Traditional Chinese Medicine integrative database for herb molecular mechanism analysis. Nucleic Acids Res. 2013 Jan;41(Database issue):D1089-95. doi: 10.1093/nar/gks1100. Epub 2012 Nov 29. PMID: 23203875; PMCID: PMC3531123.</p> <p><strong>U-BRITE LAST UPDATED June 19, 2023</strong></p>
Multi-omics Database for Integrative Microbiome Analysis in a Cohort of Korean Patients With Ankylosing Spondylitis
ClinicalTrials.gov study NCT06076083. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.
Belize terrestrial mammal database from: An integrated approach to measure hunting intensity and assess its impacts on mammal populations
Open the record for dataset details and reuse information.
Data from: Towards an integrated database on Canadian ocean resources: benefits, current states, and research gaps
Open the record for dataset details and reuse information.
ToxicoDB: an integrated database to mine and visualize large-scale toxicogenomic datasets (TGGATEs rat dataset)
<p>This data was generated by Igarashi Y, Nakatsu N, Yamashita T, Ono A, Ohno Y, Urushidani T, Yamada H. Open TG-GATEs: a large-scale toxicogenomics database. Nucleic Acids Res [Internet]. 2015 Jan;43(Database issue):D921–7. Available from: http://dx.doi.org/10.1093/nar/gku955 PMCID: PMC4384023. The data have been curated and analyzed using our open-source R package, <em>ToxicoGx</em> (<a href="http://bioconductor.org/packages/devel/bioc/html/ToxicoGx.html">http://bioconductor.org/packages/devel/bioc/html/ToxicoGx.html</a>) and are available publicly in the <em>ToxicoDB </em>web application (<a href="http://www.toxicodb.ca">www.toxicodb.ca</a>).</p>
ToxicoDB: an integrated database to mine and visualize large-scale toxicogenomic datasets (drug matrix dataset)
<p>This data was generated by Ganter B, Snyder RD, Halbert DN, Lee MD. Toxicogenomics in drug discovery and development: mechanistic analysis of compound/class-dependent effects using the DrugMatrix database. Pharmacogenomics [Internet]. 2006 Oct;7(7):1025–1044. Available from: http://dx.doi.org/10.2217/14622416.7.7.1025 PMID: 17054413. The data have been curated and analyzed using our open-source R package, <em>ToxicoGx</em> (<a href="https://cran.r-project.org/web/packages/ToxicoGx/">cran.r-project.org/web/packages/ToxicoGx</a>) and are available publicly in the <em>ToxicoDB </em>web application (<a href="http://www.toxicodb.ca">www.toxicodb.ca</a>).</p> <p> </p> <p> </p>
Text-fig. 2. "Site screen" scheme of complete results of the IPR-vegetation analysis derived from the database. in The Integrated Plant Record Vegetation Analysis: Internet Platform And Online Application
Text-fig. 2. "Site screen" scheme of complete results of the IPR-vegetation analysis derived from the database.
Demo for InCliniGene, graph database for vector integration sites
<p>Demo video to configure, install, and query the graph database.</p>
YM500: An integrative small RNA sequencing (smRNA-seq) database for microRNA research
GEO Series GSE39841. Homo sapiens. 34 samples. Type: Non-coding RNA profiling by high throughput sequencing.
CilioGenics: an integrated method and database for predicting novel ciliary genes
<p>These include the files and codes used for generating the analysis in the paper CilioGenics: an integrated method and database for predicting novel ciliary genes.</p>
Figure 4 from: Schuh R (2012) Integrating specimen databases and revisionary systematics. ZooKeys 209: 255-267. https://doi.org/10.3897/zookeys.209.3288
Figure 4 - Report of specimens examined, including unique specimen identifiers.
Figure 2 from: Schuh R (2012) Integrating specimen databases and revisionary systematics. ZooKeys 209: 255-267. https://doi.org/10.3897/zookeys.209.3288
Figure 2 - Diagram of specimen data connections and work flows.
Figure 3 from: Schuh R (2012) Integrating specimen databases and revisionary systematics. ZooKeys 209: 255-267. https://doi.org/10.3897/zookeys.209.3288
Figure 3 - Map of species distributions in western North America created using the Simple Mapper.
Figure 1 from: Schuh R (2012) Integrating specimen databases and revisionary systematics. ZooKeys 209: 255-267. https://doi.org/10.3897/zookeys.209.3288
Figure 1 - Linear barcode label (left), matrix code label (right).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.