Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
53
datasets available to search
ShareScore release 0.9.0
Dataset results
53 results for “Structure Discovery”
Structural variant discovery and genotyping in next-generation sequencing data
<p>Code, logs, data, and summaries for detection and genotyping of genomic structural variants in the D.melanogaster Sussex LHM hemiclones (and one in-house reference line individual), using Genomestrip/2.0</p> <p>The unfiltered CNV pipleline results are lhm_gs.cnvs.raw.vcf.gz</p> <p>Filtered CNV results (including removal of bad samples) are filtered.goodS.lhm_gs.cnvs.raw.vcf.gz</p> <p>The file uploaded to NCBI dbVAR (which comprises of the filtered CNVs and indels >50bp from the HaplotypeCaller method) is lhm_sx16.dbVAR.vcf.gz</p> <p>The NCBI dbVAR accession number is nstd134. Code, logs and summary data are in the zipped archives, named accordingly. The archive reference_data.zip contains additional input files required for Genomestrip, including a shell script for making some of them. The file gstrip_lhm_RG_bams.list is also an input for Genomestrip, indicating bam file names and paths.</p> <p>The pre-print manuscript for this data is available on biorxiv: "Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample" http://biorxiv.org/content/early/2016/10/17/081554 doi: http://dx.doi.org/10.1101/081554</p> <p> </p>
MISATO - Machine learning dataset for structure-based drug discovery
<p>Developments in Artificial Intelligence (AI) have had an enormous impact on scientific research in recent years. Yet, relatively few robust methods have been reported in the field of structure-based drug discovery. To train AI models to abstract from structural data, highly curated and precise biomolecule-ligand interaction datasets are urgently needed. We present MISATO, a curated dataset of almost 20000 experimental structures of protein-ligand complexes, associated molecular dynamics traces, and electronic properties. Semi-empirical quantum mechanics was used to systematically refine protonation states of proteins and small molecule ligands. Molecular dynamics traces for protein-ligand complexes were obtained in explicit water. The dataset is made readily available to the scientific community via simple python data-loaders. AI baseline models are provided for dynamical and electronic properties. This highly curated dataset is expected to enable the next-generation of AI models for structure-based drug discovery. Our vision is to make MISATO the first step of a vibrant community project for the development of powerful AI-based drug discovery tools.</p>
Data for the "Discovery of Dehydroamino Acid Residues in the Capsid and Matrix Structural Proteins of HIV-1"
<p>Bottom-up mass spectrometry-based proteomic analysis (trypsin) was performed on four biological replicates of HIV-1 virions. These virions were isolated from HEK293T cells transfected with a HIV-1 proviral plasmid derived from the pNL4-3 molecular clone, rendered biosafe due to inactivating point mutations in both the env and vpr reading frames. There are 8 total spectra, 4 are from unlabeled aliquots of sample, and 4 are from aliquots of sample treated with glutathione to label dehydroamino acids (Spectra can be accessed on MassIVE (MSV000088220). All data was analyzed using MetaMorpheus version 0.0.319 (https://github.com/smith-chem-wisc/MetaMorpheus). Provided here are the results of this analysis.</p>
Interactive Causal Structure Discovery with Hyytiälä measurements (experiment code and data)
<p>This archive contains code and data required to reproduce the results presented in the following two papers.</p> <p>Interactive Causal Structure Discovery in Earth System Sciences<br> published in Proceedings of The KDD'21 Workshop on Causal Discovery, 2021.</p> <p>Technical note: incorporating expert domain knowledge into causal structure discovery workflows<br> published in Biogeosciences, 2022</p> <p>The archive contains a README markdown document detailing the contents and how to run the experiments.</p> <p> </p>
Structure-Based Discovery of Mouse Trace Amine-Associated Receptor 5 Antagonists
<p>Topology, parameter and coordinates files of the Molecular dynamics (MD) simulations of mTAAR5 3D models in complex with the three novel antagonists (Compound 9, 10 and 11, Table S4). We used ACEMD3 (v3.5.1) as a molecular engine, CHARMM36 as force field. The three replicas of 200 ns (xtc files) were concatened in a single trajectory for each system. Water molecules, ions, and membrane atoms (POPC: phosphatidylcholine) atoms were removed from the original trajectories and topology before the upload.</p>
Microgeographic population structuring in a genus of California trapdoor spiders and discovery of an enigmatic new species (Euctenizidae: Promyrmekiaphila korematsui sp. nov.)
<p>The recognition and delineation of cryptic species remains a perplexing problem in systematics, evolution, and species delimitation. Once recognized as such, cryptic species complexes provide fertile ground for studying genetic divergence within the context of phenotypic and ecological divergence (or lack thereof). Herein we document the discovery of a new cryptic species of trapdoor spider, <em>Promyrmekiaphila korematsui </em>sp. nov. Using subgenomic data obtained via target enrichment, we document the phylogeography of the California endemic genus <em>Promyrmekiaphila </em>and<em> </em>its constituent species, which also includes <em>P. clathrata </em>and <em>P. winnemem</em>. Based on these data we show a pattern of strong geographic structuring among populations but cannot entirely discount recent gene flow among populations that are parapatric, particularly for deeply diverged lineages within <em>P. clathrata</em>.<em> </em>The genetic data, in addition to revealing a new undescribed species, also allude to a pattern of potential phenotypic differentiation where species likely come into contact. Alternatively, phenotypic cohesion among genetically divergent <em>P. clathrata </em>lineages suggests that some level of gene flow is ongoing or occurred in the recent past. Despite considerable field collection efforts over many years, additional sampling in potential zones of contact for both species and lineages is needed to completely resolve the dynamics of divergence in <em>Promyrmekiaphila</em> at the population-species interface.</p>
Comparison of Cosine, Modified Cosine, and Neutral Loss Based Spectral Alignment For Discovery of Structurally Related Molecules
<p>Spectral libraries and analysis results of the evaluation between the cosine similarity, modified cosine similarity, and neutral loss matching for the discovery of structurally related molecules.</p> <p>Spectral libraries used as input:<br> - MassIVE-KB peptide spectral library (version 2018/06/15): LIBRARY_CREATION_AUGMENT_LIBRARY_TEST-82c0124b-download_filtered_mgf_library-main.mgf<br> - GNPS community spectral libraries (downloaded on 2022/05/12): ALL_GNPS_NO_PROPOGATED.mgf<br> - GNPS bile acids spectral library (downloaded on 2022/05/12): BILELIB19.mgf</p> <p>Analysis output results:<br> - massivekb_peptide_mods.csv: 955,228 peptide MS/MS spectrum pairs from MassIVE-KB<br> - gnps_libraries.csv: 10 million small molecule MS/MS spectrum pairs from the GNPS community spectral libraries<br> - gnps_libraries_metadata.csv: structural (InChI, SMILES) and class information (computed using Classyfire) for 58,165 small molecule spectra from the GNPS community spectral libraries<br> - gnps_bilelib.csv: 340,637 bile acids MS/MS spectrum pairs from the GNPS bile acids spectral library</p> <p>For more information, see: https://github.com/bittremieux/cosine_neutral_loss/</p>
Data for paper "Automated Structure Discovery for Scanning Tunneling Microscopy"
<p>Contents of the dataset:</p> <ul> <li>band.h5 -- keys are molecule indices, each molecule has the following keys:<br> <ul> <li>eigs: KS eigenvalues for each state</li> <li>coefs: KS eigenvectors for each basis set</li> <li>xyz: atomic positions</li> <li>Z: atomic species</li> <li>qs: mulliken point charges</li> </ul> </li> <li>rotations_210611.pickle -- keys train/val/test <ul> <li>Each set is a dict containing id -- rotation pairs</li> <li>rotations are 3x3 numpy arrays</li> </ul> </li> <li>disks.pt -- a pretrained model for Atomic Disks predictions</li> </ul>
Processed CODEX Datasets from - Discovery and Generalization of Tissue Structures from Spatial Omics Data
<p>This entry provides access to processed CODEX data files of four studies analyzed in the article "Discovery and Generalization of Tissue Structures from Spatial Omics Data". Details of datasets can be found in the STAR Methods section of the article.</p> <p>For each dataset, a zip file containing multiple comma-separated values (CSV) files is included.</p> <p>Each region is assigned an unique identifier (e.g., DKD_kidney_001), and its related data files are:</p> <ul> <li>`{region_id}.cell_data.csv`, a table containing three columns: "CELL_ID", "X", and "Y". This table provides centroid locations for all cells segmented in this region.</li> <li>`{region_id}.expression.csv`, a table containing multiple columns: "CELL_ID", "DAPI", "CD45", etc. This table provides detailed protein biomarker expression quantified for all cells in this region.</li> <li>`{region_id}.scgp_annotations.csv`, a table containing two columns: "CELL_ID" and "SCGP". This table provides SCGP/SCGP-Extension annotations for all cells in this region.</li> </ul> <p>Code base for SCGP is also included in this entry. Please refer to <a href="https://gitlab.com/enable-medicine-public/scgp">https://gitlab.com/enable-medicine-public/scgp</a> for the latest codes, questions, and/or issues. Raw CODEX data and images will be accessible through links posted at the code base. Raw data will also be available from lead contact (A.E.T.) upon request.</p>
Workshop Material - 3D-e-Chem Structural Cheminformatics Workflows for Computer-Aided Drug Discovery
<p>The workshop at the KNIME user meeting (Berlin 9th of March 2018) is set up to stimulate participants with varying degrees of experience in cheminformatics to learn and apply the different structural cheminformatics tools and workflows developed within the context of the 3D-e-Chem project. You will learn how to construct and apply integrated cheminformatics workflows using the 3D-e-Chem KNIME nodes for the exploitation of G protein-coupled receptor and kinase data (two important pharmaceutical target classes) to obtain useful information for drug discovery.</p> <p>Information on the 3D-e-Chem KNIME nodes and workflows can be found online:</p> <p>3D-e-Chem GitHub website: <a href="http://3d-e-chem.github.io/">http://3d-e-chem.github.io/</a></p>
TAD-fusion score: discovery and ranking the contribution of deletions to genome structure
<p>Datasets and code of the manuscript:</p> <p>Huynh L. & Hormozdiari F., TAD-fusion score: discovery and ranking the contribution of deletions to genome structure.</p> <p> </p>
Molecular structure discovery for untargeted metabolomics using biotransformation rules and global molecular networking
<p>Comparative analysis of SIRIUS to evaluate our method, Biotransformation-based Annotation Method (BAM). This dataset includes all scripts, data, and results relevant to this analysis. BAM can be found on GitHub (https://github.com/HassounLab/BAM). </p> <p> </p> <p> </p>
Microgeographic population structuring in a genus of California trapdoor spiders and discovery of an enigmatic new species (Euctenizidae: Promyrmekiaphila korematsui sp. nov.)
Open the record for dataset details and reuse information.
AFM data of ice clusters on Cu(111) and Au(111) in paper "Structure discovery in Atomic Force Microscopy imaging of ice"
<p>Frequency shift CO-tip atomic force microscopy data of small ice clusters on Cu(111) and Au(111) surfaces as they appear in the paper "Structure discovery in Atomic Force Microscopy imaging of ice".</p><p>The data are saved in a compressed .tar.gz archive. The unpacked archive contains each experiment as a Numpy .npz file. Each file contains the measurement data as a 3D array in the key 'data' and the physical extent of the scan region in the x and y directions in Ånströms in the keys 'lengthX' and 'lengthY'.</p>
Hartree potentials and geometries of relaxed on-surface ice clusters in "Structure discovery in Atomic Force Microscopy imaging of ice"
<p>Hartree potentials and geometries of on-surface DFT-relaxed ice clusters used in the paper "Structure discovery in Atomic Force Microscopy imaging of ice".</p><p>The data are saved in a compressed .tar.gz archive. The unpacked archive contains the data for each ice cluster in the .xsf format. The density functional theory (DFT) calculations were done using the Vienna Ab-initio Simulation Package with the optB86b-vdW density functional.</p>
Unleashing the power of data through organization: Structure and connections for meaning, learning, and discovery
<p><span>Knowledge organization is needed everywhere. Its importance is marked by its pervasiveness. This paper will show many areas, tasks, and functions where proper use of Knowledge Organization, construed as broadly as the term implies, provides support for learning and understanding, for sensemaking and meaning making, for inference, and for discovery by people and computer programs and thereby will make the world a better place. The paper focuses not on metadata but rather on structuring and representing the actual data or knowledge itself and argues for more communication between the largely separated KO, Ontology, Data Modeling, and Semantic Web communities to address the many problems that need better solutions. In particular, the paper discusses the application of knowledge organization in Knowledge bases for question answering and cognitive systems; Knowledge bases for information extraction from text or multimedia; Linked data; Big data and data analytics; Electronic health records as one example; Influence diagrams (causal maps), dynamic system models, process diagrams, concept maps, and other node-link diagrams; Information systems in organizations; Knowledge organization for understanding and learning; and Knowledge transfer between domains. The paper argues for moving beyond triples to a more powerful representation using entities and multi-way relationships but not attributes.</span></p>
Discovery of thermostable fluorescently responsive glucose biosensors by structure-assisted function extrapolation
<p>Accurate assignment of protein function from sequence remains a fascinating and difficult challenge. The periplasmic binding protein (PBP) superfamily present an interesting case of function prediction, because they are both ubiquitous in prokaryotes, and they tend to diversify through gene duplication "explosions" that can lead to large numbers of paralogs in a genome. An engineered version of the moderately thermostable glucose-binding PBP from <i>Escherichia coli</i> has been used successfully as a reagentless fluorescent biosensor both <i>in vitro</i> and <i>in vivo</i>. To develop more robust sensors that meet the challenges of real-world applications, we report the discovery of thermostable homologs that retain a glucose-mediated conformationally coupled fluorescence response. Accurately identifying a glucose-binding PBP homolog among closely related paralogs is challenging. We demonstrate that a structure-based method that filters sequences by residues that bind glucose in an archetype structure is highly effective. Using fully sequenced bacterial genomes we found that this filter reduced high paralog numbers to single hits in a genome, consistent with accurate separation of glucose binding from other functions. We expressed engineered proteins for eight homologs, chosen to represent different degrees of sequence identity and tested their glucose-mediated fluorescence responses. We accurately predicted the presence of glucose binding down to 31% sequence identity. We also have successfully identified suitable candidates for next-generation robust, fluorescent glucose sensors.</p>
FIGURE. Coprolites preserved in an early Permian fern mesophyll. A, Gross morphology of a fragmentary fern frond, specimen PB23532. B, Basal part of a penultimate pinna showing sphenopteroid vegetative pinnules. C, Polished surface showing two sporangia with typical annulus structures (white arrowheads). D, SEM image showing an in situ trilete spore. E, The fertile pinnule which contains numerous coprolites along a transverse wound area. F, Enlargement showing coprolites filled with brown to black contents. G, SEM image of the same part in E. H, SEM image showing locally preserved epidermal cells and nearby coprolites. in Discovery of coprolites in an Early Permian fern mesophyll
FIGURE. Coprolites preserved in an early Permian fern mesophyll. A, Gross morphology of a fragmentary fern frond, specimen PB23532. B, Basal part of a penultimate pinna showing sphenopteroid vegetative pinnules. C, Polished surface showing two sporangia with typical annulus structures (white arrowheads). D, SEM image showing an in situ trilete spore. E, The fertile pinnule which contains numerous coprolites along a transverse wound area. F, Enlargement showing coprolites filled with brown to black contents. G, SEM image of the same part in E. H, SEM image showing locally preserved epidermal cells and nearby coprolites.
Structure-guided discovery of potent antifungals that prevent Ras signaling by inhibiting protein farnesyltransferase
<p>Infections by fungal pathogens are difficult to treat due to a paucity of antifungals and emerging resistances. Next-generation antifungals therefore are needed urgently. We have developed compounds that prevent farnesylation of <em>Cryptoccoccus</em> <em>neoformans</em> Ras protein by inhibiting protein farnesyltransferase with 3–4 nanomolar affinities. Farnesylation directs Ras to the cell membrane and is required for infectivity of this lethal pathogenic fungus. Our high-affinity compounds inhibit fungal growth with 3–6 micromolar minimum inhibitory concentrations, 4- to 8-fold better than Fluconazole, an antifungal commonly used in the clinic. Compounds bound with distinct inhibition mechanisms at two alternative, partially overlapping binding sites, accessed via different inhibitor conformations. We showed that antifungal potency depends critically on the selected inhibition mechanism, because this determines the efficacy of an inhibitor at low <em>in</em> <em>vivo</em> levels of enzyme and farnesyl substrate. We elucidated how chemical modifications of the antifungals encode the desired inhibitor conformation and concomitant inhibitory mechanism.</p>
Few juveniles or males were collected. Only four males from groups 7, 8, 9, and 11, all in clade D, were included in the dataset. The male in Fig. 13E–H conforms to the general morphological description of males in Lobocriconema with an undifferentiated labial region, the absence of a stylet, a degenerate pharyngeal region, a FIGURE 7. SEM images of specimens representing clades D (A–H) and B (I). NID numbers are associated with unique specimens, all are females except image C. A) Lobocriconema sp., face view with conspicuous labial disc surrounded by irregular labial structure, Nine-Mile Prairie, Nebraska, NID 4533. B) Lobocriconema sp., face view lacking submedian lobes and displaying subcuticular labial structure, Big Thicket National Preserve, Texas, NID 4560. C) Lobocriconema sp., juvenile, head with visible submedian lobes, body scales with fine terminal projections, Spring Creek Prairie, Nebraska, NID 4514. D) Lobocriconema sp., face view lacking submedian lobes and displaying subcuticular labial structure, Nine-Mile Prairie, Nebraska, NID 4527 E) Lobocriconema sp., cephalic profile with protruding stylet, Nine-Mile Prairie, Nebraska, NID 4529. F) Lobocriconema sp., head profile lacking submedian lobes, Tunica Hills, Louisiana, NID 4574. G) Lobocriconema sp., tail with closed vulva, Nine-Mile Prairie, Nebraska, NID 4533. H) Lobocriconema sp., tail with closed vulva, Nine-Mile Prairie, Nebraska, NID 4526. I) Lobocriconema sp., face view lacking submedian lobes, Great Smoky Mountains National Park, Purchase Knob, NID 4570. in Species discovery and diversity in Lobocriconema (Criconematidae: Nematoda) and related plant-parasitic nematodes from North American ecoregions
Few juveniles or males were collected. Only four males from groups 7, 8, 9, and 11, all in clade D, were included in the dataset. The male in Fig. 13E–H conforms to the general morphological description of males in Lobocriconema with an undifferentiated labial region, the absence of a stylet, a degenerate pharyngeal region, a FIGURE 7. SEM images of specimens representing clades D (A–H) and B (I). NID numbers are associated with unique specimens, all are females except image C. A) Lobocriconema sp., face view with conspicuous labial disc surrounded by irregular labial structure, Nine-Mile Prairie, Nebraska, NID 4533. B) Lobocriconema sp., face view lacking submedian lobes and displaying subcuticular labial structure, Big Thicket National Preserve, Texas, NID 4560. C) Lobocriconema sp., juvenile, head with visible submedian lobes, body scales with fine terminal projections, Spring Creek Prairie, Nebraska, NID 4514. D) Lobocriconema sp., face view lacking submedian lobes and displaying subcuticular labial structure, Nine-Mile Prairie, Nebraska, NID 4527 E) Lobocriconema sp., cephalic profile with protruding stylet, Nine-Mile Prairie, Nebraska, NID 4529. F) Lobocriconema sp., head profile lacking submedian lobes, Tunica Hills, Louisiana, NID 4574. G) Lobocriconema sp., tail with closed vulva, Nine-Mile Prairie, Nebraska, NID 4533. H) Lobocriconema sp., tail with closed vulva, Nine-Mile Prairie, Nebraska, NID 4526. I) Lobocriconema sp., face view lacking submedian lobes, Great Smoky Mountains National Park, Purchase Knob, NID 4570.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.