Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
356
datasets available to search
ShareScore release 0.9.0
Dataset results
356 results for “In silico”
A voltage-dependent fluorescent indicator for optogenetic applications, archaerhodopsin-3: Structure and optical properties from in silico modeling
<p>seq.fasta and P96787_1UAZ.fasta are the input files for I-TASSER suite.</p> <p>The command of running I-TASSER suite:<br> path-to-I-TASSER-dir/I-TASSERmod/runI-TASSER.pl -pkgdir path-to-I-TASSER-dir -libdir path-to-I-TASSER-lib-dir -seqname P96787 -datadir path-to-workdir -outdir path-to-outdir -java_home /usr/ -restraint2 P96787_1UAZ.fasta</p> <p>For detailed data please refer to I-TASSER documentation.</p> <p>The P96787_1UAZ_IT.zip is an archieve with I-TASSER output for the best achaerhodopsin-3 prediciton.</p> <p>SCRIPTS_GENERAL.zip contains all the scripts for the structure postprocessing: addition of hydrogen atoms,<br> chromophore insertion and equilibration, internal waters addition.</p> <p>P96787_wi.pdb is a final prepared structure.</p> <p>Input files for spectra calculations are rhodopsin_gaussian.inp and rhodopsin_orca.inp.</p> <p> </p>
In silico modeling of arch-3: Medeller and RosettaCM model building.
<p>Here are all the files nessesary for prediction of archaerhodopsin-3 structure using Medeller or RosettaCM algorithms. </p> <p>For modeling using Medeller please read README_MEDELLER file for all the information. For RosettaCM:</p> <p>Please, before using these scripts adjust them for you cluster. You will need the recent version of Rosetta package installed.</p> <p>Let's take for example stucture P96787 (arch-3) as a query, 1UAZ_full.pdb as a template.</p> <p>Also, you'll need grishex.script, thread.script, hybridize.script, relaxate.script, stage1_membrane.wts, stage2_membrane.wts, stage3_rlx_membrane.wts, uuu.grishin, rosetta_cm.options, rosetta_cm.xml, relax.options, cluster.options that are located in the "GENERAL" folder here. Let's assume they are located in the workfolder.</p> <p>Create an alignment file. For pairwise alignement use AlignMe/MP-T. Copy the alignment information in the file P96787_1UAZ.aln in the workfolder.</p> <p>Put in the workfolder attached P96787_3.frags P96787_9.frags attached here -- fragment files for Rosetta.<br> Put in the workfolder attached P96787.octopus.</p> <p>mv 1UAZ_full.pdb 1UAZ.pdb</p> <p>Open the grish.script and check: "target" and "templateA" which describe each line of the alignment file change on the names of the lines in your alignment file!<br> For example for AlignMe it will be P96787 1UAZ.<br> ./grishex.script P96787 1UAZ</p> <p>./thread.script P96787 1UAZ</p> <p><br> Run the hybridize script: ./hybridize.script P96787 1UAZ 500<br> It will take several days.</p> <p>It will create file hybridized_P96787.out -- a binary silent file of 500 structures with the rebuilt loops and inserted fragments, if<br> there were gaps. It will take several days also.</p> <p>Run the ./relax.script P96787 50<br> <br> It will create file relaxed_P96787.out -- a binary silent file of 25000 structures -- 50 for each hybridized.</p> <p>Time to evaluate the results.</p> <p>Clustering.</p> <p>Run ./do.script P96787<br> open LISTIK<br> delete the _0001 in the end of each line.<br> Run ./resulting1.script</p> <p>In the folder PDBSS open cluster_summary.txt.<br> The first structure in the list is the best one. It is located in the folder PDBSS.</p> <p><br> </p>
Supplementary information for: "A voltage-dependent fluorescent indicator for optogenetic applications, archaerhodopsin-3: Structure and optical properties from in silico modeling".
<p>This is supplementary data for F1000Research article: A voltage-dependent fluorescent indicator for optogenetic applications, archaerhodopsin-3: Structure and optical properties from in silico modeling.</p> <p>Here are files for modeling archaerhodopsin-3 with I-TASSER, Medeller and RosettaCM algorithms, structure postprocessing and spectra calculations.</p>
In silico database of ~500 000 virtual patients with diverse cardiovascular disease
<p>The presented dataset contains a large virtual cohort of more than 50'000 heart failure patients characterized by realistic traces of volumes, pressures, flows, and regional mechanics. We used the well-established CircAdapt model <a href="https://www.circadapt.org/">(https://www.circadapt.org/</a> , <a href="http://framework.circadapt.org/">http://framework.circadapt.org/</a>) of the human heart and circulation to simulate a large cohort of virtual patients, covering a wide range of HF-related disease heterogeneity and severity. <br>In the construction of virtual patients, generating a variety of parameter sets for the initial population is essential and can be accomplished using various sampling techniques. In our investigation, we utilized the Sobol-low discrepancy sequence to ensure uniformity across the high-dimensional parameter space. For a more detailed explanation please refer to the document <em>"MARCIUS_deliverable_D1.5_VirtualDatabase.pdf"</em></p> <p><strong>Acknowledgments:</strong> This work was supported by the European Union's Horizon 2020 Research and Innovation program under the Marie Skłodowska-Curie grant agreement No. 86074, <strong>"MARCIUS - MARie Curie Intelligent UltraSound"(<a href="https://www.marcius-project.com/">https://www.marcius-project.com/</a>)</strong>. MARCIUS rationale was to develop a comprehensive in silico simulation platform comprising both the generation of virtual patients (<em>presented dataset</em>) and their associated realistic image data. Such approach would make it possible to learn the most relevant patterns within a wide representative set of patients to lead the training of ML-based image processing algorithms in order to analyze real-world clinical data.</p> <p> </p>
Direct synthesis, characterization, in vitro and in silico studies of simple chalcones as potential antimicrobial and antileishmanial agents
<p>Chalcone represents a vital biosynthetic scaffold owing to its numerous therapeutic effects. The present study was intended to synthesize seventeen chalcone derivatives <strong>(3a-q)</strong> by direct coupling of substituted acetophenones and benzaldehyde. The target chalcones were characterized by spectroscopic analyses followed by their <em>in vitro</em> antimicrobial, and antileishmanial investigations with reference to standard drugs. The majority of the chalcones displayed good to excellent biological activities. Chalcone <strong>3q</strong> (1000 μg/mL) exhibited the most potent antibacterial effect with its zone of inhibition values of 30, 33, and 34 mm versus <em>Staphylococcus aureus,</em> <em>Escherichia coli, and Pseudomonas aeruginosa </em>respectively. The results also confirmed chalcone <strong>3q</strong> to be the most potent versus <em>Leishmania major </em>with the lowest IC<sub>50</sub> value of 0.59±0.12 μg/mL.<em> </em>Chalcone <strong>3i</strong> (500 μg/mL) was noticed to be the most potent antifungal agent with its zone of inhibition being 29 mm against <em>Candida albicans</em>. Computational studies of chalcones <strong>3i </strong>and<strong> 3q</strong> supported the preliminary <em>in vivo</em> results. The existence of the amino moiety and bromine atom on ring-A and methoxy moieties on ring-B caused better biological effects of the chalcones. In brief, the investigations reveal that chalcones (<strong>3i </strong>and<strong> 3q) </strong>can be employed as building blocks<strong> </strong>to discover novel antimicrobial agents.</p>
Dataset for "In silico-labeled ghost cytometry"
<p>Data and code for the analyses in Ugawa <em>et al.</em> "In silico<em>-</em>labeled ghost cytometry" eLife.</p>
In-vitro Major Arterial Cardiovascular Simulator: Benchmark Data Set for in-silico Model Validation
<p><strong>Background</strong><br> <br> The data described here supplements the paper "In-vitro Major Arterial Cardiovascular Simulator to generate Benchmark Data Sets for in-silico Model Validation" (to be submitted). It was created at Technische Hochschule Mittelhessen (THM) in Germany and uploaded to Zenodo. Please cite the paper M. Wisotzki, A. Mair, P. Schlett, B. Lindner, M. Oberhardt, S. Bernhard, In Vitro Major Arterial Cardiovascular Simulator to Generate Benchmark Data Sets for In Silico Model Validation (2022), Data 7(11), DOI: 10.3390/data7110145 and the Zenodo doi when using this dataset.</p> <p><strong>General description / Dataset Structure</strong></p> <p>Each mat-File describes a different stenosis degree at the popliteal artery of the in-vitro simulator MACSim (details can be found in the paper). There are 17 pressure signals for different positions, one flow sensor close to the stenosis location and one monitor signal of the proportional valve use to control the input curve. Total duration of each signal is 60s with a sampling rate of 1000 Hz. Each mat-file contains a header structure with metadata and struct array for signals of each sensor. Signals in each mat-File are aligned with respect to a common time axis, but this is not guaranteed between different measurements/files. The file format can either be loaded directly in Matlab or in Python with scipy's loadmat function.</p> <p>The different stenosis degrees for each degree are:<br> ScenarioI: 100 % Area fraction (no stenosis)<br> ScenarioII: 37,5 % Area fraction<br> ScenarioIII: 23,4 % Area fraction<br> ScenarioIV: 6,56 % Area fraction</p> <p><strong>Data fields for each file</strong></p> <table> <caption>headerStruct</caption> <thead> <tr> <th scope="col">field</th> <th scope="col">description</th> </tr> </thead> <tbody> <tr> <td>rate</td> <td>sampling rate in Hz</td> </tr> <tr> <td>description</td> <td>name of the scenario according to the paper, corresponds to filename</td> </tr> <tr> <td>configuration</td> <td>parameters of the trapezoidal input curve (offset and amplitude in mmHg, ascend times and descend times and smoothing window in a fraction the time period (1.2s))</td> </tr> </tbody> </table> <p> </p> <table> <caption>signalStruct</caption> <thead> <tr> <th scope="col">field</th> <th scope="col">description</th> </tr> </thead> <tbody> <tr> <td>nodeId</td> <td>corresponds to numbered nodes at which the sensor is placed, the corresponding location can be found in the paper (node numbering, not sensor numbers) or in the software SISCA (https://gitlab.com/agbernhard.lse.thm/sisca) in the example database.</td> </tr> <tr> <td>type</td> <td>'p' ... pressure or 'q' ... flow</td> </tr> <tr> <td>data</td> <td>double array, time series of each sensor, unit mmHg for type 'p' and ml/s for type 'q' </td> </tr> <tr> <td>anatomicalPosition</td> <td> <p>name of the corresponding anatomical position</p> </td> </tr> </tbody> </table>
Raw data and metadata of SiO2 NP physicochemical characterisation, in vitro investigations and in silico predictions on protein corona formation
<p>Raw data and metadata of SiO<sub>2</sub> NP physicochemical characterisation, <em>in vitro</em> investigations and<em> in silico</em> predictions on protein corona formation. Data repository for Hasenkopf I. et al., 2022. Please note that “pristine” NPs in metadata are referred to as “bare” NPs in the paper.</p> <p>1. NP phys-chem raw data (xlsx-formatted) + corresponding metadata (xlsx-formatted)</p> <p>2. NP-protein binding raw data (xlsx-formatted) + corresponding metadata (xlsx-/pdf-formatted, ImageLab-formatted analyses, pdf-formatted ImageLab reports)</p> <p>3. huRBL mediator release raw data (xlsx-formatted) + corresponding metadata (pdf-formatted)</p> <p>4. <em>in silico</em> modelling raw data (xlsx-formatted) + corresponding metadata (config-/map-formatted)</p> <p> </p>
LibINVENT: Reaction-based Generative Scaffold Decoration for in Silico Library Design
<p>Training datasets used for LibINVENT publication <a href="https://doi.org/10.1021/acs.jcim.1c00469">https://doi.org/10.1021/acs.jcim.1c00469</a>.</p> <p>The datasets used for training of the prior:</p> <ul> <li><strong>purged_chembl_sliced.smi.gz:</strong> The CHEMBl 27 compounds, filtered according to the rules described in the manuscript and sliced according to the reaction rules.</li> <li><strong>chembl_train.smi.gz:</strong> The purged, sliced dataset used for model training. The DRD2 compounds are removed as described in the manuscript.</li> </ul>
IN SILICO SCREENING OF TWO-PHOTON ABSORPTION PROPERTIES OF A LARGE SET OF BIS-BF2-DYES
<p>This data set is a part of Supporting Info for our work entitled “IN SILICO SCREENING OF TWO-PHOTON ABSORPTION PROPERTIES OF A LARGE SET OF BIS-BF2-DYES” which is published in ChemPhotoChem peer-reviewed scientific journal (<a href="https://doi.org/10.1002/cptc.202200137">doi.org/10.1002/cptc.202200137</a>). In this data set, we provide the XYZ coordinates for the 280 compounds studied in this work.</p>
Use of A Molecular Switch Probe to Activate or Inhibit GIRK1 Heteromers In Silico Reveals a Novel Gating Mechanism
<p>GIRK channel structure models (PDB structure files) used for Molecular Dynamics simulations.</p>
Additional datasets and code accompanying the article "Evaluation of in silico algorithms for use with ACMG/AMP clinical variant interpretation guidelines"
<p> </p> <p> </p> <p>Datasets and code accompanying the manuscript titled “Evaluation of in silico algorithms for use with ACMG/AMP clinical variant interpretation guidelines” by Rajarshi Ghosh, Ninak Oak and Sharon E. Plon.</p>
Machine learning-based pulse wave analysis for classification of circle of Willis topology: an in silico study with 30,618 virtual subjects (database: Complete CoW)
<p>This repository contains the dataset for the complete CoW described in the article with the same name. MATLAB and Python codes for post-processing the dataset and the code for training and testing all machine learning models using the open-source library TensorFlow 2.12, the Keras application programming interface, and the Scikit-learn Python package can be found in here (<a href="https://zenodo.org/records/12519322" target="_blank" rel="noopener">https://zenodo.org/records/12519322</a>).</p>
Figure 1. Kaempferol-3-O in Analysis of the toxicological and pharmacokinetic profile of Kaempferol-3-O-β-D-(6"-E-p-coumaryl) glucopyranoside - Tiliroside: in silico, in vitro and ex vivo assay
Figure 1. Kaempferol-3-O-β-D-(6"-E-p-coumaryl) glucopyranoside – tiliroside.
Dataset to reproduce the figures of article "Information dynamics of in silico EEG Brain Waves"
<p>Data and code to reproduce results figures.</p>
In silico mock communities for evaluation of taxonomic profilers across prokaryotes and viruses
<p><em>In silico </em>mock communities generated with CAMISIM for benchmarking the performance of taxonomic profilers across prokaryotic (50 communities), eukaryotic (30 communities), and viral communities (10 communities) of the human microbiome. Metagenomes were generated using CAMISIM (Fritz et al., 2019), which simulates 2.1 Gb of Illumina 2 ×150 bp paired end reads with the default HiSeq 2500 error profile and a mean insert size of 200 bp. To assess profiling performance for a range of sequencing depths, the 50 <em>in silico</em> metagenomes were also rarefied with seqtk (-s100) to sequencing depths of 20, 5, 2, 1, 0.5, 0.25 and 0.1 million read pairs. Counts are provided for rarefied metagenomes.</p> <p><strong>Prokaryotic communities<br></strong>For prokaryotic benchmarking, 10 body site-representative prokaryotic metagenomes were simulated for each of the following five body sites: adult gut, infant gut, oral, skin, and vagina. Genome accession ids for prokaryotic species found in each human body site were identified from published literature (Bäckhed et al., 2015; Proctor et al., 2019; Saheb Kashaf et al., 2021). </p> <p>Adult Gut: pro_gut_adult.zip<br>Infant Gut: pro_gut_infant.zip<br>Oral: pro_oral.zip<br>Skin: pro_skin_1.zip, pro_skin_2.zip, pro_skin_3.zip<br>Vaginal: pro_vaginal.zip</p> <p>Downsized counts: </p> <p><strong>Eukaryotic communities<br></strong>30 eukaryotic <em>in silico</em> metagenomes comprising up to 200 randomly sampled genomes from a set of 113 eukaryotic species (See Supplementary Table 2 from the paper) corresponding to the eukaryotic species within both CHAMP and MetaPhlAn 4 (Blanco-Míguez et al., 2023) databases.</p> <p>Eukaryotic data is deposited here: <a href="https://doi.org/10.5281/zenodo.12090449" target="_blank" rel="noopener">doi: 10.5281/zenodo.12090449</a></p> <p><strong>Viral communities</strong><br>10 viral communities were simulated with 95% of the reads from bacteria and 5% of the reads originating from phages. Each community consisted of 200 randomly selected bacterial genomes from GTDB with species-level annotation and 200 viral genomes from the Gut Phage Database (GPD, Camarillo-Guerrero et al., 2021). </p> <p>Counts: phage_communities_counts.zip<br>FastQ, forward reads: camisimu_[1-10].fq.1.gz<br>FastQ, reverse reads: camisimu_[1-10].fq.2.gz</p> <p><strong>References</strong></p> <p>Bäckhed, F., Roswall, J., Peng, Y., Feng, Q., Jia, H., Kovatcheva-Datchary, P., et al. (2015). Dynamics and Stabilization of the Human Gut Microbiome during the First Year of Life. <em>Cell Host Microbe</em> 17, 690–703. doi: 10.1016/J.CHOM.2015.04.004</p> <p>Blanco-Míguez, A., Beghini, F., Cumbo, F., McIver, L. J., Thompson, K. N., Zolfo, M., et al. (2023). Extending and improving metagenomic taxonomic profiling with uncharacterized species using MetaPhlAn 4. <em>Nature Biotechnology 2023 41:11</em> 41, 1633–1644. doi: 10.1038/s41587-023-01688-w</p> <p>Camarillo-Guerrero, L. F., Almeida, A., Rangel-Pineros, G., Finn, R. D., and Lawley, T. D. (2021). Massive expansion of human gut bacteriophage diversity. <em>Cell</em> 184, 1098. doi: 10.1016/J.CELL.2021.01.029</p> <p>Fritz, A., Hofmann, P., Majda, S., Dahms, E., Dröge, J., Fiedler, J., et al. (2019). CAMISIM: Simulating metagenomes and microbial communities. <em>Microbiome</em> 7, 1–12. doi: 10.1186/S40168-019-0633-6/FIGURES/5</p> <p>Proctor, L. (2019). Priorities for the next 10 years of human microbiome research. <em>Nature 2021 569:7758</em> 569, 623–625. doi: 10.1038/d41586-019-01654-0</p> <p>Saheb Kashaf, S., Proctor, D. M., Deming, C., Saary, P., Hölzer, M., Mullikin, J., et al. (2021). Integrating cultivation and metagenomics for a multi-kingdom view of skin microbiome diversity and functions. <em>Nature Microbiology 2021 7:1</em> 7, 169–179. doi: 10.1038/s41564-021-01011-w</p>
An in-silico analysis of information sharing systems for adaptable resources management: a case study of oyster farmers
<p>Model and data outcomes --> We developed an agent-based models involving oyster farmers sharing information to adapt to an ill-understood virus. Various scenarios of heterogeneity and information sharing (through social networks and centralized information system) are simulated.</p>
Uncovering the structure and function of Pseudomonas aeruginosa periplasmic proteins by an in silico approach
<p><strong>Caprari_et_al_models_and_docking: </strong>Predicted three-dimensional structures and docking simulations results described in the manuscript "Uncovering the structure and function of Pseudomonas aeruginosa periplasmic proteins by an in silico approach" by Caprari et al. More informations are given in the "README.txt" file.</p>
In silico prediction of high-resolution Hi-C interaction matrices (part III)
<p>The uploaded files are source datasets for the HiC-Reg approach. HiC-Reg is a regression based method that predict contact counts from one-dimensional regulatory signals such as epigenetic marks and regulatory protein binding. See more details here (<a href="https://github.com/Roy-lab/HiC-Reg">https://github.com/Roy-lab/HiC-Reg</a>). There are a total of six files in this dataset: Data.tgz, Gm12878.tgz, Hmec.tgz, K562.tgz, Huvec.tgz and Nhek.tgz. The Data.tgz include predictions and other downstream analysis such as feature importance analysis, significant interaction calling, and data files for select figures. The Gm12878.tgz, K562.tgz, Huvec.tgz, Hmec.tgz and Nhek.tgz contain trained models, predictions, feature files for two chromosomes for in each cell line.</p> <p>This is part III of the dataset which contains Nhek.tgz and Data.tgz.</p>
In silico prediction of high-resolution Hi-C interaction matrices (part I)
<p>The uploaded files are source datasets for the HiC-Reg approach. HiC-Reg is a regression based method that predict contact counts from one-dimensional regulatory signals such as epigenetic marks and regulatory protein binding. See more details here (<a href="https://github.com/Roy-lab/HiC-Reg">https://github.com/Roy-lab/HiC-Reg</a>). There are a total of six files in this dataset: Data.tgz, Gm12878.tgz, Hmec.tgz, K562.tgz, Huvec.tgz and Nhek.tgz. The Data.tgz include predictions and other downstream analysis such as feature importance analysis, significant interaction calling, and data files for select figures. The Gm12878.tgz, K562.tgz, Huvec.tgz, Hmec.tgz and Nhek.tgz contain trained models, predictions, feature files for two chromosomes for in each cell line.</p> <p>This is part I of the dataset which contains Gm12878.tgz and Hmec.tgz.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.