Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,151
datasets available to search
ShareScore release 0.9.0
Dataset results
4,151 results for “discovery”
Cancer related protein visualized by using Discovery Studio Visualizer
<p>Cancer related Protein visualized by using Discovery Studio Visualizer</p>
Big Data to Knowledge (BD2K) Training Coordinating Center (TCC) Educational Resource Discovery Index (ERuDIte) as Linked Data
<p>This is a release of the Big Data to Knowledge (BD2K) Training Coordinating Center (TCC) Educational Resource Discovery Index (ERuDIte) as Linked Data.<br> <br> ERuDIte contains over 11,000 training resources on data science including courses (MOOCs), video tutorials, conference talks, and other materials. The metadata of these resources is described uniformly using schema.org. In addition, we use machine learning techniques to tag each resource with concepts from the Data Science Education Ontology (DSEO), which we developed to further describe the contents of the training resources. Resource relevance and tags are curated by experts to ensure high quality. Finally, we map the references to people and organizations in the learning resource metadata to entities in DBpedia, DBLP, and ORCID, thus embedding our collection in the web of linked data. Our collection is continually growing. We hope that ERuDIte will provide a framework to foster open linked educational resources on the web.<br> <br> Distributed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (https://creativecommons.org/licenses/by-nc-sa/4.0/)</p>
answered questionnaire to Bachelor Thesis "Wie sinnvoll ist die Ergänzung des Resource Discovery Systems an der Bibliothek des Max-Planck-Instituts für evolutionäre Anthropologie durch einen zusätzlichen, externen Index?"
<p>The dataset contains the answers that were given in the online questionnaire that was conducted as part of the Bachelor Thesis "Wie sinnvoll ist die Ergänzung des Resource Discovery Systems an der Bibliothek des Max-Planck-Instituts für evolutionäre Anthropologie durch einen zusätzlichen, externen Index?"</p> <p>The questionnaire and further information can be found in the Bachelor Thesis, which is linked uner Related Works.</p>
Discovery and characterization of pyridine and furan substituted ligands of choline acetyltransferase
<p><span>This repository contains datasets for the manuscript "Discovery and characterization of pyridine and furan substituted ligands of choline acetyltransferase"</span></p> <ul> <li><span>Data set of 1.4 million compounds used for virtual screening are freely available at </span><span><a href="https://vitasmlab.biz/downloads"><span>https://vitasmlab.biz/downloads</span></a></span><span>. Vina-MPI used for the virtual screening protocol is freely available at </span><span><a href="https://github.com/mokarrom/mpi-vina"><span>https://github.com/mokarrom/mpi-vina</span></a></span><span>. </span></li> <li><span>The docking pose and docking score for the screened library with Vina-MPI is available in PDBQT format with their docking scores in the folder “vitas_virtual_screening_800K”.</span></li> <li><span>Selected 5958 compounds from the virtual screening are given as PDBQT with docking scores in folder “top_5K_hits”.</span></li> <li><span>Re-docked docking score (Top_5K_re_docking.sdf) and MMGBSA (Top_250_MMGBSA.sdf) calculation are also available with the structures in SDF format.</span></li> </ul>
[NGC5084] SAUNAS II: Discovery of Cross-shaped X-ray Emission and a Rotating Circumnuclear Disk in the Supermassive S0 Galaxy NGC 5084
<p>The contained FITS files represent the processed Chandra/ACIS X-ray surface brightness maps of the NGC5084 galaxy, observed with Chandra/ACIS and analyzed with the SAUNAS pipeline as described in Borlaff et al. 2024b (https://ui.adsabs.harvard.edu/abs/2024arXiv240810449B/abstract). All the images have the photometric calibrations (in units of photons cm-2 s-1 pixel-1) and have been astrometrically aligned. </p> <div>Each file contains four FITS extensions as detailed below: </div> <div>----</div> <div>EXTENSION NAME TYPE SIZE DETAILS </div> <div>----</div> <div>0 INFO no-data 0 BLANK EXTENSION. <br>1 SB_FLUX float64 512x512 X-RAY SURFACE BRIGHTNESS MAP. [photons cm-2 s-1 pixel-1] <br>2 STD_SB_FLUX float64 512x512 X-RAY SURFACE BRIGHTNESS NOISE MAP [photons cm-2 s-1 pixel-1]<br>3 SNR float64 512x512 SIGNAL-TO-NOISE RATIO [ - ]</div> <div>-----------</div> <div> </div> <div>Use the SNR extension (extension #3) to determine if your source of interest in the SB_FLUX map (extension #1) is statistically significant over the background limit.</div> <div> </div> <div> </div> <div> </div> <div> </div>
First discovery and confirmation of PN candidates found from AI and deep learning techniques applied to VPHAS+ survey data
<div> <div> <div> <div> <div> <div> <p>Appendices: VPHAS+ detected PNG images labelled by red boxes, SHS Hα-Rband quotient images, VPHAS+ Hα-Rband quotient images, SAAO spectra with spectral lines labelled.</p> </div> </div> </div> <p>Table: Parameters of the observed PN candidates and any associated nebulosity or outflows.</p> <p>Reduced spectra.</p> </div> </div> </div>
Development of a spectral library for the discovery of altered genomic events in Mycobacterium avium associated with virulence using mass spectrometry-based proteogenomic analysis
<p><em>Mycobacterium avium</em> is one of the prominent disease-causing bacteria in humans. It causes lymphadenitis, chronic and extrapulmonary, and disseminated infections in adults, children, and immunocompromised patients. <em>M. avium</em> has ~4,500 predicted protein-coding regions on an average, which can be helpful in discovering several variants at the proteome level. Many of them are potentially associated with virulence, thus identifying such proteins can be a helpful feature in the development of panel-based theranostics. In line with such a long-term goal, we carried out an in-depth proteomic analysis of <em>M. avium</em> with both data-dependent and data-independent acquisition methods. Further, a set of proteogenomic investigations were carried out using the protein database for <em>Mycobacterium tuberculosis,</em> and a genome six-frame translated database and a variant protein database of <em>M. avium</em>. A search of mass spectrometry data analysis against <em>M. avium</em> protein database resulted in the identification of 2,954 proteins. Further, proteogenomic analyses aided in the identification of 1,301 novel peptide sequences and correction of translation start sites for 15 proteins. At the end, we created a spectral library of <em>M. avium</em> proteins including novel genome search-specific peptides and variant peptides detected in this study. We validated the spectral library by a data-independent acquisition of the <em>M. avium</em> proteome. Thus, we present a <em>M. avium </em>spectral library of 29,033 peptide precursors supported by 0.4 million fragment ions for further use by the biomedical community.</p>
Mask or Enhance: Data Curation Aiding the Discovery of Piezoresponse Force Microscopy Contributors
<p>This repository contains the data used in the corresponding study:</p> <p>Mask or Enhance: Data Curation Aiding the Discovery of Piezoresponse Force Microscopy Contributors</p> <p><strong>Abstract</strong></p> <p>Piezoresponse force microscopy (PFM) is routinely used to probe the nanoscale electromechanical response of ferroelectric and piezoelectric materials. However, many challenges remain in the interpretation of the recovered signal. Specifically, many non-ferroelectric contributions affect the measured response, ranging from electrostatics, to charge injection and trapping, and topographic cross-talk. Recently, machine learning (ML) has been utilized to identify multiple contributors within complex data systems, such as PFM response. A substantial advancement in ML approaches for PFM techniques is offered by dimensional stacking, enabling encoding of physical and/or chemical correlations within the materials’ response across different data dimensions spanning varying ranges. However, dimensional stacking requires appropriate scaling for each dimension (before ML analysis) to minimize undesired information loss. Here, the impact of clustering globally and locally scaled parameters in polarization switching experiments via resonant PFM (RPFM) are discussed. Specifically, dimensional stacking of scaled parameters can mask or enhance ferroelectric and non-ferroelectric behaviors, and aid identification of various physical phenomena contributing to the measured RPFM response. This study highlights the importance of data curation for ML, and its role in identifying signal contributors to scanning probe microscopy (SPM)-based techniques with multidimensional data, such as resonant and/or spectroscopic SPM.</p>
MISATO - Machine learning dataset for structure-based drug discovery
<p>Developments in Artificial Intelligence (AI) have had an enormous impact on scientific research in recent years. Yet, relatively few robust methods have been reported in the field of structure-based drug discovery. To train AI models to abstract from structural data, highly curated and precise biomolecule-ligand interaction datasets are urgently needed. We present MISATO, a curated dataset of almost 20000 experimental structures of protein-ligand complexes, associated molecular dynamics traces, and electronic properties. Semi-empirical quantum mechanics was used to systematically refine protonation states of proteins and small molecule ligands. Molecular dynamics traces for protein-ligand complexes were obtained in explicit water. The dataset is made readily available to the scientific community via simple python data-loaders. AI baseline models are provided for dynamical and electronic properties. This highly curated dataset is expected to enable the next-generation of AI models for structure-based drug discovery. Our vision is to make MISATO the first step of a vibrant community project for the development of powerful AI-based drug discovery tools.</p>
MBC and ECBL Libraries: outstanding tools for drug discovery
<p><strong>UPDATE</strong>. New in this revision: python scripts to process DBs and calculate the percentage of molecules which pass the Veber and Ghose filters. Two new DBs were also added and considered for the analysis.</p> <p>Data and scripts to reproduce all the graphics reported in the Manuscript entitled: "MBC and ECBL Libraries: outstanding tools for drug discovery".</p> <p><strong>List of analyzed DBs:</strong></p> <ol> <li>MBC2016 (Total entries: 1,096 cmpds; 7.39% excluded from properties analysis - QikProp failure).</li> <li>MBC2022 (Total entries: 2,577 cmpds; 3.14% excluded from properties analysis - QikProp failure).</li> <li>ECBL (Total entries: 101,021 cmpds; 0.20% excluded from properties analysis - QikProp failure).</li> <li>ChEMBL v.31 (Total entries 1,908,325 cmpds; 2.97% excluded from properties analysis - QikProp failure).</li> <li>DrugBank v.5.0 (Total entries 10,981 cmpds; 4.13% excluded from properties analysis - QikProp failure).</li> <li>ZINC20 (Total entries 10,723,360 cmpds; 0.61% excluded from properties analysis - QikProp failure).</li> <li>NuBBE (Total entries 2,223 cmpds) - <strong>NEW</strong></li> <li>Approved drugs (Total entries: 3,140 cmpds) - <strong>NEW</strong></li> </ol> <p><strong>Files:</strong></p> <p><em>QikProp_properties.docx</em>: doc file containing the full list of QikProp properties calculated for each analyzed DB.</p> <p><em>DATA_comparison.xlsx</em>: excel file containing data used to reproduce plots in <strong>Figure 4</strong> of the MS.</p> <ul> <li><em>Murcko_scaffold_percentages</em>: distribution (%) of the first 50 most populated Murcko scaffolds for MBC2016, MBC2022 and ECBL.</li> <li><em>Murcko_scaffolds_comparison</em>: distribution (count) of the first 94 common Murcko scaffolds for MBC2016, MBC2022 and ECBL.</li> </ul> <p>QikProp properties for all the analyzed DBs (8 files; CSV format).</p> <p>SMILES codes for all the analyzed DBs (8 files; SMI format). </p> <p><em>joinplots.py</em>: python script to generate the 2D plots in <strong>Figure 2</strong> of the MS.</p> <p><em>fingerprint_similarity.py</em>: python script to run and generate the Tanimoto similarity plots in <strong>Figure 3</strong> of the MS.</p> <p><em>calc_kde.py</em>: python script to run kernel density analysis reported in <strong>Figure 5 </strong>of the MS.</p> <p><em>Veber_filter.py: python script to generate </em>data presented in <strong>Table 1 </strong>of the MS. <strong>(NEW)</strong></p> <p><em>Ghose filter.py: python script to generate </em>data presented in <strong>Table 1 </strong>of the MS. <strong>(NEW)</strong></p>
GE Discovery TOF MI PET NEMA IQ projector benchmark listmode data
<p>## LIST0000.BLF</p> <p>listmode file from GE Discvoery MI PET/CT containing all acquired emission events (HDF5)<br> of a single bed position NEMA IQ phantom acq.</p> <p>## corrections.h5</p> <p>file containing all quantitative corrections estimate using GE's duetto tool box (HDF5)</p> <p>- correction_lists/sens -> sensivity value for acquired events<br> - correction_lists/atten -> attenuation value for acquired events<br> - correction_lists/contam -> additive contaminations (randoms + scatter) for all acquired events<br> - all_xtals/atten -> attenuation values for all possible crystal combinations<br> - all_xtals/sens -> sensitivity values for all possible crystal combinations<br> - all_xtals/xtal_ids -> all possible crystal combinations</p> <p> </p>
Data for: Discovery of low-metallicity stars in the central parsec of the Milky Way
<p>Spectroscopic data from Stostad et al. (2015) and Do et al. (2015). This file contains K-band spectra from the Gemini NIFS instrument of stars in the central parsec of the Galactic center. The files are in FITS format. Additional descriptions are in in Stostad et al. (2015). </p> <p>If using this data, please cite Do et al. (2015) and Stostad et al. (2015)</p> <p>https://ui.adsabs.harvard.edu/abs/2015ApJ...808..106S/abstract<br>https://ui.adsabs.harvard.edu/abs/2015ApJ...809..143D/abstract</p> <p> </p> <p>Contributors to the creation of the spectra from this dataset include:</p> <p>Tuan Do</p> <p>Morten Stostad</p>
Figure 2. Combined species discovery curve for 726 in Canopy assemblages and species richness of planthoppers (Hemiptera: Fulgoroidea) in the Ecuadorian Amazon
Figure 2. Combined species discovery curve for 726 planthopper canopy fogging samples from Onkone Gare (three collecting years) including select estimators of diversity. Total observed morphospecies was 573, with 26% represented as singletons. The averaged value of the diversity estimators is 740. Curves for species observed and diversity estimators failed to reach an asymptote.
Databases for exploratory mode of RRE-Finder: A Genome-Mining Tool for Class-Independent RiPP Discovery
<p>RREFinder is a bioinformatic tool for the detection of RiPP Recognition Elements (RREs). See "RRE-Finder: A Genome-Mining Tool for Class-Independent RiPP Discovery".</p> <p>This database contains the required databases to run exploratory mode of the tool.</p>
Data for Bovine breed-specific augmented reference graphs facilitate accurate sequence read mapping and unbiased variant discovery
<p><strong>Description of the datasets</strong></p> <p>Data are organized as folders and compressed with tar.gz.</p> <p>There are two compressed data folder: <strong>data </strong>which used for cattle genome graphs experiment and <strong>data_human</strong> which we used for human genome graphs experiment. </p> <p><strong>Cattle genome graphs experiments</strong></p> <p>First you need to unzip the file using command <em>tar -xvzf data.tar.gz</em>. After unzipping, the data folder is organized as follows:</p> <ul> <li>Utilities: contain bovine ARS-UCD 1.2 fasta reference with the accompanying index.</li> <li>Bin: contain the softwares used in the paper (vg, liftover, vcf2diploid)</li> <li>Part1: data for analysis in variant prioritization section, further subdivided into: <ul> <li>vcf_sim: variant files from four animal in each breed used to simulate reads</li> <li>reads_sim: simulated short reads used for read mapping</li> <li>vcf_freq: variants augmented to graphs filtered based on allele frequency</li> </ul> </li> <li>Part2: data used for analysis in the section of graph mapping with breeds-filtered variants, further subdivided into: <ul> <li>vcf_breed: variant files used to graphs construction.</li> </ul> </li> <li>Part3: data used for analysis in the section of consensus genome, further subdivided into: <ul> <li>read_sims: simulated reads as in the part1, but the coordinates are liftovered to the new consensus genomes.</li> <li>reference: contain the original reference and consensus references.</li> <li>vcf_consensus: contain major allele variants to construct consensus genomes.</li> </ul> </li> <li>Part4: data analysis in the section of whole genome graph construction and variant genotyping. <ul> <li>vcf_construct: variants from chromosome 1-29 from 82 Brown Swiss used to construct BSW whole genome graph.</li> <li>BSW_graph: whole genome Brown Swiss graph with the three accompanying indexes (xg,gcsa, and gbwt).</li> </ul> </li> </ul> <p><strong>Human genome graphs experiments</strong></p> <p>First you need to unzip the <em>data_human</em> file using command <em>tar -xvzf data</em><em>_hum.tar.gz</em>. After unzipping, the data folder is organized as follows:</p> <ul> <li>reference: the g1k_v37 reference used as a graph backbone</li> <li>vcf_sim: variant files from four individuals in each population used to simulate reads</li> <li>reads_sim: simulated short reads used for read mapping</li> <li>vcf_freq: variants augmented to graphs filtered based on allele frequency</li> </ul>
Fig. 3 in Rare sponges from marine caves: discovery of Neophrissospongia nana nov. sp. (Demospongiae, Corallistidae) from Sardinia with an annotated checklist of Mediterranean lithistids
Fig. 3. Neophrissospongia nana nov. sp., holotype MSNG 54599, spicular complement. A. Dichotriaene. B. Cladome of a tubercled dichotriaene (top view). C. Streptaster/amphiaster microscleres with tubercles. D. Cladome of a smooth dichotriaene (top view). E. Dicranoclone desma. F. Styles/sub-tylostyles.
Fig. 2 in Rare sponges from marine caves: discovery of Neophrissospongia nana nov. sp. (Demospongiae, Corallistidae) from Sardinia with an annotated checklist of Mediterranean lithistids
Fig. 2. Neophrissospongia nana nov. sp., holotype MSNG 54599, photomicrographs of the skeletal architecture and spicular complement. A-C. Subectosomal skeletal network of dicranoclone desmas. D. Dichotriaene (top) and simple triaene (bottom, arrows) radially arranged in the ectosome with tubercled cladomes. E. Articulation among desmas with scattered styles/sub-tylostyles in the sub-ectosome. F. Detail showing the arrangement of silica in dicranoclone desmas (cross section) and styles/sub-tylostyles (arrows). G-I. Streptaster/amphiaster microscleres with spines/tubercles.
Fig. 1 in Rare sponges from marine caves: discovery of Neophrissospongia nana nov. sp. (Demospongiae, Corallistidae) from Sardinia with an annotated checklist of Mediterranean lithistids
Fig. 1. Habitus of the cave-dwelling Mediterranean Neophrissospongia nana nov. sp. from north-western Sardinia. A. In situ plate-like growth form in the type locality Grotta delle Terrazze (photo by R. Barbieri), B. Holotype MSNG 54599 (top view) (photo by G. Delitala).
Fig. 6. A–D in Unexpected discovery of six new species of Aphyosemion (Cyprinodontiformes, Aplocheilidae) in the Wonga-Wongué Presidential Reserve in Gabon
Fig. 6. A–D. Aphyosemion aurantiacum Chirio, Legros & Agnèse sp. nov. A. Adult, ♁, from locality 12, not preserved. Photo O. Buisson. B. Adult, ♀, from locality 12, not preserved. Photo O. Buisson. C. Holotype, adult, ♁, from Wézé Spring (MRAC 2016-019-P-64). D. Paratype, adult, ♀, from Wézé Spring (MRAC 2016-019-P-65-73). E–H. A. rubrogaster Chirio, Legros & Agnèse sp. nov. E. Adult, ♁, from locality 16, not preserved. Photo O. Buisson. F. Adult, ♀, from locality 16, not preserved. Photo O. Buisson. G. Holotype, adult, ♁, from the Niengé River (MRAC 2016-019-P-93). H. Paratype, adult, ♀, from the Niengé River (MRAC 2016-019-P-94-108).
Fig. 5. A–D in Unexpected discovery of six new species of Aphyosemion (Cyprinodontiformes, Aplocheilidae) in the Wonga-Wongué Presidential Reserve in Gabon
Fig. 5. A–D. Aphyosemion barakoniense Chirio, Legros & Agnèse sp. nov. A. Adult, ♁, from locality 7, not preserved. B. Adult, ♀ from locality 7, not preserved. C. Holotype, adult, ♁, from the lower Barakonié River (MRAC 2016-019-P-19). D. Paratype, adult, ♀, from the lower Barakonié River (MRAC 2016-019-P-20-36). E–F. A. pusillum Chirio, Legros & Agnèse sp. nov. E. Adult, ♁, from locality 10, not preserved. F. Holotype, adult, ♁, from the Okoyo River (MRAC 2016-019-P-57).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.