Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,773

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,773 results for “predictive modeling”

Learn how ShareScore rates datasets ↗
zenodo44/100

Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic Prediction - Datasets

<p>Datasets to NeurIPS 2021 accepted paper &quot;Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic Prediction&quot;.</p> <p>Datasets are pytorch files containing a dictionary with training, validation and test sets. Train, validation and test sets are custom dataset classes which inherit from the standard torch dataset class. Corresponding code an be found at https://github.com/HSG-AIML/NeurIPS_2021-Weight_Space_Learning.</p> <p>Datasets 41, 42, 43 and 44 are our dataset format wrapped around the zoos from Unterthiner et al, 2020 (https://github.com/google-research/google-research/tree/master/dnn_predict_accuracy)<br> <br> Abstract:<br> Self-Supervised Learning (SSL) has been shown to learn useful and information-preserving representations. Neural Networks (NNs) are widely applied, yet their weight space is still not fully understood. Therefore, we propose to use SSL to learn neural representations of the weights of populations of NNs. To that end, we introduce domain specific data augmentations and an adapted attention architecture. Our empirical evaluation demonstrates that self-supervised representation learning in this domain is able to recover diverse NN model characteristics. Further, we show that the proposed learned representations outperform prior work for predicting hyper-parameters, test accuracy, and generalization gap as well as transfer to out-of-distribution settings.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Output data of the models used in "Comparison of six approaches to predicting droplet activation of surface active aerosol. Part 1: moderately surface active organics" by Vepsäläinen et al. (2022)

<p>Output data of the different models used in &quot;Comparison of six approaches to predicting droplet activation of surface active aerosol. Part 1: moderately surface active organics&quot; by Veps&auml;l&auml;inen et al. (2022).</p> <p>Output data is included for 50 nm particles containing malonic acid (mna), succinic acid (sca) and glutaric acid (glutarica), mixed with ammonium sulphate (AS) in different organic mass fractions.&nbsp;</p> <p>A plotter that allows the user to plot the K&ouml;hler curves, surface tensions and organic<br> partitioning factors during droplet growth from the model output data provided is included.&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Snow equi-temperature metamorphism described by a phase-field model applicable on micro-tomographic images: prediction of microstructural and transport properties

<p>This dataset provides data described and used in the article submitted to Journal of Advances in Modeling Earth Systems &quot;Snow equi-temperature metamorphism described by a phase-field model applicable on micro-tomographic images: prediction of microstructural and transport properties&quot;.</p> <p>It contains .csv files with different properties computed on outputs of the model Snow3D simulating equi-temperature metamorphism. This micro-scale model was used here with experimental micro-tomographic snow images as input and returns series of 3-D images of snow showing features of equi-temperature metamorphism at different time steps as output.</p> <p>In this dataset, you will find two types of files:</p> <p>- the microstructural properties (density, specific surface area, covariance lengths, mean curvature) computed on&nbsp; the simulated images at different time steps.</p> <p>- the transport properties (effective conductivity, normalizes effective vapor diffusion coefficient, permeability) of the simulated images at different time steps.</p> <p>Finally, metadata_simulations.csv gather the information relative to the simulations.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

COLONOMICS - predictive models for normal colon gene expression and DNA methylation for TWAS and MWAS

<p>We provide&nbsp;significant SNP prediction models derived from the COLONOMICS data (<a href="https://www.colonomics.org">https://www.colonomics.org</a>). Genotypes were obtained by Affymetrix 6.0 array, imputed to TopMed panel. Gene expression was obtained from Affymetrix U219 array, DNA methylation was obtained with Illuminan 450K array and miRNA expression was obtained by NGS.&nbsp;We provide SNP prediction models for 1,758 genes, 30,530 CpG probes and 38 miRNAs obtained from colon normal biopsy samples. These features can be predicted from SNPs located within &plusmn;1Mb, which we assumed they act through cis mechanisms. We include&nbsp;the model&rsquo;s summary statistics and corresponding SNP weights in SQLite objects. Models were trained using the elastic net procedure employed in the PredictDB pipeline (<a href="https://predictdb.org/">https://predictdb.org</a>), according to which only models with a predictive performance p-value &lt; 0.05 and R<sup>2</sup> &gt; 0.1 are considered significant. We adjusted the models by basic covariates, i.e., sex, age, tissue type and colon anatomic location where biopsies were collected (left and right colon). Genome coordinates refer to GRCh37/hg19.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Data, scripts, and figures of the article: Processing weights of chickens determined by Dual-Energy X-Ray Absorptiometry. 2. Developing prediction models

<p>This data set contains the data, JMP scripts, and figures of the article titled &quot;Processing weights of chickens determined by Dual-Energy X-Ray Absorptiometry. 2. Developing prediction models&quot; to be published in the journal Animal - Open Space.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Input features of E. coli proteome for predicting and modeling protein-protein interactions with AF2Complex

<p>Input features to be used with AF2Complex for predicting protein-protein interactions among ~4400 E. coli proteins. A pickled feature file was generated by the feature data pipeline of AF2Complex for each E. coli protein. To reduce storage size, we limited up to 10,000 MSA sequences and up to 10 structural templates from the Protein Data Bank. The cutoff date for sequence libraries and the Protein Data Bank releases used for feature generation is no later than 11-30-2021.</p> <ul> <li>ecoli_af2c_fea.txt -- A list of all E coli protein with pre-generated input features</li> <li>af2c_fea_ecoli_220331_msa10ktem10.tar&nbsp;-- Input features named after the UniProt ID of each proteins. Note that after untar the tarball, you may use the gzipped feature pickle files directly with AF2Complex w/o gunzip.</li> </ul>

opencc-by-4.0Aug 2022View details →
zenodo44/100

VEL-Ar trajectory prediction model linear, co-seismic, and post-seismic grids for interpolation

<p>VEL-Ar trajectory prediction model linear, co-seismic, and post-seismic interpolation grids in ASCII format. The generation of these grids is described in http://doi.org/10.1007/s00190-015-0871-8</p>

opencc-by-4.0Dec 2015View details →
zenodo44/100

ΔvapHm-VOC: Standard Molar Vaporization Enthalpy Database for Machine Learning Prediction Models

<p>We present the full database of the article "Data-Driven, Explainable Machine Learning Model for Predicting Volatile Organic Compounds&rsquo; Standard Vaporization Enthalpy".</p> <p>This is the database used for data driven, explainable supervised ML model to predict &Delta;<sub>vap</sub><em>H</em><sub>m</sub>&deg; of VOCs. The model was built on an established experimental database of 2410 unique molecules and 223 VOCs categorized by chemical groups. Using supervised ML regression algorithms, the Random Forest successfully predicted VOCs&rsquo; &Delta;<sub>vap</sub><em>H</em><sub>m</sub>&deg; with a mean absolute error of 3.02 kJ mol<sup>-1</sup> and a 94% test score. The model was successfully validated through the prediction of &Delta;<sub>vap</sub><em>H</em><sub>m</sub>&deg; for a known database of VOCs and through molecular group hold-out tests.</p> <div> <div> <div> <div> <p>The model's database was built with a variety of molecules from diverse chemical families with known experimental &Delta;<sub>vap</sub><em>H</em><sub>m</sub>&deg; values. Entries were collected from <a href="https://doi.org/10.1063/1.3309507" target="_blank" rel="noopener">Acree and Chickos&rsquo; 2010 compilation</a>, curated by <a href="https://doi.org/10.1016/j.fluid.2013.09.021" target="_blank" rel="noopener">Gharagheizi (2013)</a>, with experimental vaporization enthalpy at the standard temperature of 298.15 K. This database was selected as it is an open-access repository, generally presenting experimental values with low uncertainties and corrected for the real-to-ideal behavior of the gas phase. We introduced a routine to convert and present each chemical entry into a SMILES string, along with chemical family categorization. For VOCs, we built a specific database of compounds documented in a VOC regulatory environmental guideline (<a href="https://www.s-t-a.org/Files%20Public%20Area/Documents/The%20Categorisation%20of%20Volatile%20Organic%20Compounds%20HMIP%20(1996).pdf">Marlowe <em>et al</em>., 1995</a>), and we used our web-scrapping routine to gather experimental &Delta;<sub>vap</sub><em>H</em><sub>m</sub>&deg; values. The external dataset for validation studies was also collected from <a href="https://doi.org/10.1016/j.fluid.2013.09.021" target="_blank" rel="noopener">Gharagheizi (2013)</a>.</p> <p>Along with &Delta;<sub>vap</sub><em>H</em><sub>m</sub>&deg; experimental values, each molecule is represented by its CAS number, SMILES string and InChlKey. We generated 106 chemical descriptors for every molecule in the database, using <a href="http://http//www.rdkit.org/">RDKit</a> software version 2022.09.4, running on top of Python 3.9. Descriptors were calculated from the &ldquo;MolFromSmiles&rdquo; function in &ldquo;RDKIT.Chem&rdquo; as descriptors with non-numerical values were removed. The descriptors encode significant chemical information and are used to present physicochemical characteristics of compounds, building a relationship between structure and &Delta;<sub>vap</sub><em>H</em><sub>m</sub>&deg;.</p> </div> </div> </div> </div> <p>Through chemical feature importance analysis, the explainable model revealed that VOC polarizability, connectivity indexes and electrotopological state are key for the model&rsquo;s prediction accuracy. We thus present a replicable and explainable model, which can be further expanded towards the prediction of other thermodynamic properties of VOCs.</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Data archive and code for "Predicting September Arctic Sea Ice: A Multi-Model Seasonal Skill Comparison"

<p>This upload contains data and code related to the paper "Predicting September Arctic Sea Ice: A Multi-Model Seasonal Skill Comparison" by M. Bushuk, S. Ali, D. Bailey, Q. Bao, L. Batte, U. S. Bhatt, E. Blanchard-Wrigglesworth, E. Blockley, G. Cawley, J. Chi, F. Counillon, P. Goulet Coulombe, R. Cullather, F. X. Diebold, A. Dirkson, E. Exarchou, M. Gobel, W. Gregory, V. Guemas, L. Hamilton, B. He, S. Horvath, M. Ionita, J. E. Kay, E. Kim, N. Kimura, D. Kondrashov, Z. M. Labe, W. Lee, Y. J. Lee, C. Li, X. Li, Y. Lin, Y. Liu, W. Maslowski, F. Massonnet, W. N. Meier, W. J. Merryfield, H. Myint, J. C. Acosta Navarro, A. Petty, F. Qiao, D. Schroder, A. Schweiger, Q. Shu, M. Sigmond, M. Steele, J. Stroeve, N. Sun, S. Tietsche, M. Tsamados, K. Wang, J. Wang, W. Wang, Y. Wang, Y. Wang, J. Williams, Q. Yang, X. Yuan, J. Zhang, and Y. Zhang, published in the Bulletin of the American Meteorological Society, DOI: https://doi.org/10.1175/BAMS-D-23-0163.1.</p> <p>See README.txt for a description of the datasets and code.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Predictive models for off-target binding profiles generation

<p>Models for predicting off-target binding, built with Conformal Prediction, and the <a href="http://cpsign-docs.genettasoft.com">CPSign software</a>. The dataset is part of an upcoming publication (Manuscript in preparation), which will provide more details.</p> <p>The dataset is a GZipped Tar archive, with the models as Java Archive (JAR) files. For every JAR-file, there is also a corresponding audit log, with the extension &quot;.audit.json&quot;, produced by the workflow software (<a href="http://scipipe.org">SciPipe</a>) used to train the models. This audit file contains all the shell commands used in the workflow that produced the models.</p>

opencc-by-4.0Jun 2018View details →
zenodo44/100

Replication package for "An Empirical Assessment of Best-Answer Prediction Models in Technical Q&A Sites" (EMSE 2018)

<p>Replication package for the paper:</p> <blockquote> <p>F. Calefato, F. Lanubile, and N. Novielli (2018) &ldquo;<a href="http://collab.di.uniba.it/fabio/wp-content/uploads/sites/5/2018/07/EMSE-D-17-00159_R3.compressed.pdf">An Empirical Assessment of Best-Answer Prediction Models in Technical Q&amp;A Sites</a>.&rdquo;&nbsp;Empirical Software Engineering Journal, DOI:&nbsp;<a href="http://dx.doi.org/10.1007/s10664-018-9642-5">10.1007/s10664-018-9642-5</a>.</p> </blockquote>

openother-openFeb 2019View details →
zenodo44/100

Consensus models to predict oral rat acute toxicity and validation on a dataset coming from the industrial context

<p>We report predictive models of acute oral systemic toxicity representing a follow-up of our previous work in the framework of the NICEATM project. It includes the update of original models through the addition of new data and an external validation of the models using a dataset relevant for the chemical industry context. A regression model for LD50 and classification model for toxicity classes according to the Global Harmonized System categories were prepared. ISIDA descriptors were used to encode molecular structures. Machine learning algorithms included Support Vector Machine (SVM), Random Forest (RF) and Na&iuml;ve Bayesian. Selected individual models were combined in consensus.</p> <p>The different datasets were compared using the Generative Topographic Mapping approach. It appeared that the NICEATM datasets were lacking some relevant chemotypes for chemical industry. The new models trained on enlarged data sets have applicability domain (AD) sufficiently large to accommodate industrial compounds. The fraction of compounds inside the models&rsquo; AD increased from 58 % (NICEATM model) to 94 % (new model). Yet, the increase of training sets only slightly improved of the models&rsquo; prediction performance: RMSE values decreased from 0.56 to 0.47 and balanced accuracies increased from 0.69 to 0.71 for NICEATM and new models, respectively.</p>

opencc-by-4.0Jul 2019View details →
zenodo44/100

Data Sets and Prediction Models Created using MLP, RNN and LSTM

<p>Binary Data Sets</p> <p>-&nbsp;2018tbi219_shuffled3.csv and&nbsp;2018tbi219_shuffled5.csv</p> <p>- these are stratified data sets that were able to produce prediction models with high prediction rates.</p> <p>Models.zip</p> <p>- this zip file contains the prediction models created using Keras&nbsp;Deep Learning Algorithms: MLP, RNN and LSTM</p> <p>Model Creation - Python Code Snippets</p> <p>- Code snippets for creating the prediction models</p>

opencc-by-4.0Sep 2019View details →
zenodo44/100

QSPR models for bioconcentration factor (BCF): Are they able to predict data of industrial interest?

<p>This dataset is described and studied in the article&nbsp;</p> <p>&quot;QSPR models for bioconcentration factor (BCF): Are they able to predict data of industrial interest?&quot;</p> <p>published in <em>SAR and QSAR Environmental Research</em> (Taylor&amp;Francis).</p> <p>Files description:</p> <p>SI_BCFtrainset.xlsx: a collection of 1129 chemical structures and CAS identifiers with their logBCF values extracted from various literature sources.</p> <p>SI_BCFtestset.xlsx: a collection of 204 chemical structures for which the logBCF is considered of lower reliability and used as an external test set.</p> <p>SI_FullDataset_rawdata.csv: the raw data composed of 15372 entries with the following columns:&nbsp;CASRN, Tissue, Duration [d], Test organism, Exposure type, Steady state, RESPONSE, RESPONSE UNIT, Media&nbsp;type, TakenFrom, TITLE, AUTHOR, YEAR, SOURCE, SMILES</p> <p>SI_ExcludedOutliers34.csv: 34 chemical structures that have been identified as suspicious during analysis.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2018View details →
zenodo44/100

DeepPredSpeech: computational models of predictive speech coding based on deep learning

<p>This dataset contains all data, source code, pre-trained computational predictive&nbsp;models and experimental&nbsp;results related to:&nbsp;&nbsp;</p> <p>Hueber&nbsp;T., Tatulli E., Girin L., Schwatz, J-L&nbsp;&quot;How predictive can be predictions in the neurocognitive processing of auditory and audiovisual speech? A deep learning study.&quot; (<a href="https://doi.org/10.1101/471581">biorXiv preprint&nbsp;https://doi.org/10.1101/471581</a>).&nbsp;</p> <ul> <li>Raw data are extracted from the publicly available database NTCD-TIMIT (10.5281/zenodo.260228).&nbsp; <ul> <li>Audio recordings are available in the audio_clean/ directory</li> <li>Post-processed lip image sequences are available in the lips_roi/ directory (67x67 pixels, 8bits, obtained by lossless inverse DCT-2D transform from the DCT feature available in the original repository of NTCD-TIMIT)</li> <li>Phonetic segmentation (extracted from NTCD-TIMIT original zenodo repository) is available in the HTK MLF file volunteer_labelfiles.mlf</li> </ul> </li> <li>Audio features (MFCC-spectrogram and log-spectrogram) are available in the mfcc_16k/ and fft_16k/ directories.&nbsp;</li> <li>Models (audio-only, video-only and audiovisual, based on deep feed-forward neural networks and/or convolutional neural network, in&nbsp;.h5 format, trained with Keras 2.0 toolkit) and data normalization parameters (in .dat scikit-learn&nbsp;format)&nbsp;are available in models_mfcc/ and models_logspectro/ directories</li> <li>Predicted and target (ground truth) MFCC-spectro (resp.&nbsp;log-spectro) for the test databases (1909 sentences), and for the different values of <span class="math-tex">\(\tau_p\)</span>&nbsp;or&nbsp;<span class="math-tex">\(\tau_f\)</span> are available in pred_testdb_mfccspectro/ (resp.&nbsp;pred_testdb_logspectro/) directory</li> </ul> <p>Source code for extracting audio features, training and evaluating the models is available on GitHub&nbsp;https://github.com/thueber/DeepPredSpeech/</p> <p>All directories have been zipped before upload.</p> <p>Feel free to contact me for more details.</p> <p>Thomas Hueber, Ph. D., CNRS research fellow, GIPSA-lab, Grenoble, France,&nbsp;thomas.hueber@gipsa-lab.fr&nbsp;</p>

opencc-by-4.0Nov 2018View details →
zenodo44/100

Data from "Predicting Global Ground Geoelectric Field With Coupled Geospace and Three‐Dimensional Geomagnetic Induction Models"

<p>Data presented in http://dx.doi.org/10.1029/2018SW001859 excluding the first and last hours which were determined to contain partially unphysical results and probably should not be used.</p> <p>Each file contains the external ground magnetic field components calculated on 5x5 degree geographic grid and the results of induction modeling using 1d and 3d ground conductivity models: total ground magnetic field components, horizontal ground electric field components. Times are given in UTC.</p> <p>To reproduce Figure 8 in above reference use:</p> <p>&nbsp;</p> <p>import numpy<br> import matplotlib.pyplot<br> data = numpy.load(&#39;2006-12-14T22:58:00.npz&#39;) # or 2006-12-14T22_58_00.npz<br> matplotlib.pyplot.colorbar(<br> &nbsp;&nbsp; &nbsp;matplotlib.pyplot.imshow(<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;data[&#39;B_3D_north&#39;],<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;cmap = matplotlib.pyplot.get_cmap(&#39;bwr&#39;),<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;vmin = -800, vmax = 800,<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;extent = (-180, 180, -90, 90)<br> &nbsp;&nbsp; &nbsp;),<br> &nbsp;&nbsp; &nbsp;format = &#39;%.0f&#39;, fraction = 0.02, pad = 0.03<br> )<br> matplotlib.pyplot.show()</p> <p>&nbsp;</p> <p>and substitute B_3D_east, E_3D_north, etc. for the different panels.</p>

opencc-by-4.0Jul 2018View details →
zenodo44/100

RDF version of the data from Choi, JS. et al. Towards a generalized toxicity prediction model for oxide nanomaterials using integrated data from different sources (2018)

<p>This is an RDFied version of the dataset published in&nbsp;Choi, JS., Ha, M.K., Trinh, T.X. et al. Towards a generalized toxicity prediction model for oxide nanomaterials using integrated data from different sources. Sci Rep 8, 6110 (2018)</p> <p>The original dataset publication DOI:&nbsp;<a href="https://doi.org/10.1038/s41598-018-24483-z">https://doi.org/10.1038/s41598-018-24483-z</a></p> <p>The Original publication authors:&nbsp;Jang-Sik Choi, My Kieu Ha, Tung Xuan Trinh, Tae Hyun Yoon &amp; Hyung-Gi Byun</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

RDF version of the data from Anastasios G. et al. Computational enrichment of physicochemical data for the development of a zeta-potential read-across predictive model with Isalos Analytics Platform. NanoImpact (2021).

<p>This is an RDFied version of the dataset published by&nbsp;Anastasios G. et al. Computational enrichment of physicochemical data for the development of a zeta-potential read-across predictive model with Isalos Analytics Platform. NanoImpact (2021).</p> <p>The original dataset publication DOI:&nbsp;<a href="https://doi.org/10.1016/j.impact.2021.100308">https://doi.org/10.1016/j.impact.2021.100308</a></p> <p>The Original publication authors:&nbsp;Anastasios G. Papadiamantis, Antreas Afantitis, Andreas Tsoumanis, Eugenia Valsami-Jones, Iseult Lynch, Georgia Melagraki</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Data for Predictive Modelling of Laminated Composite Plates

<p>Two different problems, i.e. a low-dimensional (LD) and a high-dimensional (HD) problems are considered. The LD problem has 2 variables for a 4-ply symmetric square composite laminate. Similarly, the HD problem consists of 16 variables for a 32-ply symmetric square composite laminate. The value of <em>h </em>for LD and HD problems is taken as 0.005 and 0.04 respectively.</p> <p>For each problem, three different types of sampling technique, i.e. random sampling (RS), Latin hypercube sampling (LHS) [1] and Hammersley sampling (HS) [2] are adopted. The RS, LHS and HS primarily differ in the uniformity of sample points over the design space such that RS has the least and HS has the maximum uniform distributions of sample points. Based on the recommendations of Jin et al. [3], and Zhao and Xue [4], 72 and 612 sample points are considered in each training dataset of LD and HD problems respectively.</p> <p>Based on the FE formulation, several high-fidelity datasets for the LD and HD problems are generated, as presented in the Supplementary Material file &ldquo;Predictive modelling of laminated composite plates.xlsx&rdquo; in nine sheets that are organized as detailed out in Table 1.</p> <p>References:</p> <p>1.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; McKay, M. D.; Beckman, R. J.; Conover, W. J. A comparison of three methods for selecting values of input variables in the analysis of output from a computer code. <em>Technometrics</em>, <strong>2000</strong>, <em>42</em>, 55-61.</p> <p>2.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Hammersley, J. M. Monte Carlo methods for solving multivariable problems. <em>Annals of the New York Academy of Sciences</em>, <strong>1960</strong>, <em>86</em>, 844-874.</p> <p>3.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Jin, R.; Chen, W.; Simpson, T. W. Comparative studies of metamodelling techniques under multiple modelling criteria. <em>Structural and Multidisciplinary Optimization</em>, <strong>2001</strong>, <em>23</em>, 1-13.</p> <p>4.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Zhao, D.; Xue, D. A comparative study of metamodeling methods considering sample quality merits. <em>Structural and Multidisciplinary Optimization</em>, <strong>2010</strong>, <em>42</em>, 923-938.</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Dataset related to article "New in silico models to predict in vitro micronucleus induction as marker of genotoxicity"

<p>The .txt file contains the dataset of the in silico&nbsp;model for genotoxicity as induction of micronuclei.</p> <p>The .doc file contains the descriptors of the models and the structural alerts.</p>

opencc-by-4.0Jul 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record