Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
486
datasets available to search
ShareScore release 0.7.1
Dataset results
486 results for “curation”
High-quality large curated dataset of protein sequences (1.83 million) and their corresponding Position Specific Scoring Matrices
<p>As part of his master thesis at the Rostlab, which is located at the Technical University of Munich (TUM), Mr. Issar Arab developed the first language model that encodes evolutionary information of proteins explicitly. The pre-training involved the creation of a novel high-quality dataset of protein sequences (around 1.83 million proteins, or ~0.8 Billion amino acids) with their corresponding Position Specific Scoring Matrices (PSSMs). Those matrices reflect the relative frequency of each amino acid at each position in a protein and is derived from evolutionarily related proteins.</p> <p>Mr. Arab makes this work publicly available to help other researchers speed up their work to leverage AI to learn the representation of protein evolutionary information more explicitly. The set of sequences was derived by extracting all PSSMs from the <a href="https://predictprotein.org/">PredictProtein</a> (PP) cache, which were also part o the UniProt Reference Cluster with 50% sequence identity (uniref50 2019_12). The overlap between PP and uniref50 was further filtered to only include high-quality samples, e.g. only multiple sequence alignments with a certain number of aligned sequences were considered. The processing led to a training set of 1.83 Million sequences, a validation set of 879 instances, and a test set of 879 entries. The training data of proteins is reduced to 40% sequence identity, with respect to the validation/test sets, and contains sequences ranging between 18 and 9858 residues in length.</p> <p>Refer to the Jupyter notebook for a detailed description of the files' structure and a Python code snippet to correctly manipulate this data.</p> <p>To access the full original work, please visit the following link: <a href="https://mediatum.ub.tum.de/node?id=1579236">Manuscript</a> <br><br><strong>Note:</strong> The dataset was recently used to fine tune a protein sequence language model (<a href="https://github.com/issararab/PEvoLM">PEvoLM</a>). The work was presented at the CIBCB'23 conference. If you use PEvoLM or this dataset in your work, please cite the following publication:</p> <p>- Issar Arab, <strong>PEvoLM: Protein Sequence Evolutionary Information Language Model</strong>, <em>IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), Eindhoven, Netherlands</em>, (2023), pp. 1-8, doi:<a href="https://ieeexplore.ieee.org/document/10264890">10.1109/CIBCB56990.2023.10264890</a></p>
Hacker News Curated Comments Dataset
<p>A curated dataset from fh-bigquery:hackernews.stories</p> <p>Only HN stories with more than 10 comments are included, and only comments from users with more than 10 comments are included.</p>
modelforge curated dataset: ANI-2x
<h1><strong>Modelforge Curated ANI-2x Dataset:</strong><br><strong>- 1000 configuration test set</strong><strong><br></strong><strong>- </strong><strong>Version: nc_1000_v1.1</strong></h1> <p>This provides a curated hdf5 file for a subset of the ANI-2x dataset designed to be compatible with <a href="https://github.com/choderalab/modelforge">modelforge</a>, an infrastructure to implement and train NNPs. This dataset contains 101 unique records for 1000 total configurations, with a maximum of 10 per system. Note, configurations are parititioned into records based on the array of atomic species appearing in sequence in the source data file.</p> <p>When applicable, the units of properties are provided in the datafile, encoded as strings compatible with the <a href="https://github.com/openforcefield/openff-units/">openff-units</a> package. For more information about the structure of the data file, please see the following:</p> <p>This dataset is compatible with modelforge hdf5 schema 2.</p> <h3><strong>Properties Included: </strong></h3> <ul> <li>atomic_numbers </li> <li>positions <ul> <li>"per_atom"</li> <li> "nanometer"</li> </ul> </li> <li>forces <ul> <li>"per_atom"</li> <li> "kilojoule_per_mole / nanometer"</li> </ul> </li> <li>energies <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> </ul> <h3><strong>Source Dataset:</strong></h3> <div> <p>The ANI-2x data set includes properties for small organic molecules that contain H, C, N, O, S, F, and Cl. This dataset contains 9651712 conformers. This data was generated with the wB97X/631Gd level of theory used in the original ANI-2x paper, calculated using Gaussian 09.</p> </div> <h3>Citations:</h3> <p><em>ANI-2x publication:</em></p> <ul> <li> <div> <p>Devereux, C, Zubatyuk, R., Smith, J. et al. "Extending the applicability of the ANI deep learning molecular potential to sulfur and halogens." Journal of Chemical Theory and Computation 16.7 (2020): 4192-4202. <a href="https://doi.org/10.1021/acs.jctc.0c00121">https://doi.org/10.1021/acs.jctc.0c00121</a></p> </div> </li> </ul> <p><em>Source dataset, released with CC Attribution 4.0 International license:</em></p> <ul> <li>Huddleston, K., Zubatyuk, R., Smith, J., Roitberg, A., Isayev, O., Pickering, I., Devereux, C., & Barros, K. (2023). ANI-2x Release [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.10108942" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.10108942</a></li> </ul>
modelforge curated dataset: tmQM
<h1><strong>Modelforge Curated tmQM Dataset: <br>- D</strong><strong>ataset restricted to [Pd, Zn, Fe, or Cu] and [C, H, P, S, O, N, F, Cl, or Br]<br></strong><strong>- </strong><strong>Version: PdZnFeCu_CHPSONFClBr_nc_1000_v1.1</strong></h1> <div> <div> <p>This dataset contains a subset of the tmQM dataset, with 1000 unique systems with 1000 total configurations (1 configuration per system). This dataset is limited to systems that contain transition metals Pd, Zn, Fe, or Cu, and also only contain elements C, H, P, S, O, N, F, Cl, or Br.<br><br></p> </div> </div> <div>The full tmQM dataset contains the geometries and properties of 108,541 mononuclear complexes extracted from the Cambridge Structural Database, including Werner, bioinorganic, and organometallic complexes based on a large variety of organic ligands and 30 transition metals (the 3d, 4d, and 5d from groups 3 to 12). All complexes are closed-shell, with a formal charge in the range {+1, 0, −1}e. </div> <div> </div> <div>This provides a curated hdf5 file for a subset of the tmQM dataset (release 13Aug2024) designed to be compatible with <a href="https://github.com/choderalab/modelforge">modelforge</a>, an infrastructure to implement and train NNPs. This dataset contains 1000 unique systems with 1000 total configurations (1 configuration per system), designed to be used for testing. </div> <p>The original tmQM repository (<a href="https://github.com/uiocompcat/tmQM">https://github.com/uiocompcat/tmQM</a>) was forked and a release made that corresponds to the data committed on 13 August 2024 (<a href="https://github.com/chrisiacovella/tmQM/releases/tag/2024Aug13">https://github.com/chrisiacovella/tmQM/releases/tag/2024Aug13</a>).</p> <p>When applicable, the units of properties are provided in the datafile, encoded as strings compatible with the <a href="https://github.com/openforcefield/openff-units/">openff-units</a> package. </p> <p>This dataset is compatible with modelforge hdf5 schema 2.</p> <h2>Original Citation:</h2> <p>David Balcells and Bastian Bjerkem Skjelstad,<br>tmQM Dataset—Quantum Geometries and Properties of 86k Transition Metal Complexes<br>Journal of Chemical Information and Modeling 2020 60 (12), 6135-6146<br>DOI: 10.1021/acs.jcim.0c01041</p> <h2>Properties Included: </h2> <ul> <li>atomic_numbers </li> <li>positions <ul> <li>"per_atom"</li> <li>"nanometer"</li> </ul> </li> <li>partial_charges <ul> <li>"per_atom"</li> <li>"elementary_charge"</li> </ul> </li> <li>total_charge <ul> <li>"per_system"</li> <li>"elementary_charge"</li> </ul> </li> <li>spin_multiplicities <ul> <li>"per_system"</li> <li>"dimensionless"</li> </ul> </li> <li>electronic_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>dispersion_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>total_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>dipole_moment_magnitude <ul> <li>"per_system"</li> <li>"elementary_charge * nanometer"</li> </ul> </li> <li>dipole_moment_computed <ul> <li>"per_system"</li> <li>"elementary_charge * nanometer"</li> </ul> </li> <li>dipole_moment_computed_scaled <ul> <li>"per_system"</li> <li>"elementary_charge * nanometer"</li> </ul> </li> <li>energy_of_lumo <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>energy_of_homo <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>homo_lumo_gap <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>stoichiometry <ul> <li>"meta_data"</li> </ul> </li> <li>metal_n_ligands <ul> <li> "meta_data"</li> </ul> </li> <li>metal_center_charge <ul> <li>"meta_data"</li> </ul> </li> </ul>
modelforge curated dataset: QM9
<h1><strong>Modelforge Curated QM9 Dataset:</strong><br><strong>- Full dataset</strong><br><strong>- Version: full_dataset_v1.1:</strong></h1> <p>This provides a curated hdf5 file for a subset of the QM9 dataset to be used for testing purposes, designed to be compatible with <a href="https://github.com/choderalab/modelforge">modelforge</a>, an infrastructure to implement and train NNPs. This test dataset contains 133885 systems with a single configuration for each unique molecule.</p> <p>When applicable, the units of properties are provided in the datafile, encoded as strings compatible with the <a href="https://github.com/openforcefield/openff-units/">openff-units</a> package. For more information about the structure of the data file, please see the following:</p> <ul> <li><a href="https://github.com/choderalab/modelforge/wiki/Dataset-and-curation#curation-module">https://github.com/choderalab/modelforge/wiki/Dataset-and-curation#curation-module</a></li> </ul> <h3><strong>Properties Included: </strong></h3> <ul> <li>atomic_numbers </li> <li>positions <ul> <li> "per_atom"</li> <li>"nanometer"</li> </ul> </li> <li>partial_charges <ul> <li>"per_atom"</li> <li>"elementary_charge"</li> </ul> </li> <li>polarizability <ul> <li>"per_system"</li> <li>"nanometer ** 3"</li> </ul> </li> <li>dipole_moment_per_system <ul> <li>"per_system"</li> <li>"elementary_charge * nanometer"</li> </ul> </li> <li>dipole_moment_scalar_per_system <ul> <li>"per_system"</li> <li>"elementary_charge * nanometer"</li> </ul> </li> <li>energy_of_homo <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>lumo-homo_gap <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>zero_point_vibrational_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>internal_energy_at_298.15K <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>internal_energy_at_0K <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>enthalpy_at_298.15K <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>free_energy_at_298.15K <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>heat_capacity_at_298.15K <ul> <li>"per_system"</li> <li>"kilojoule_per_mole / kelvin"</li> </ul> </li> <li>rotational_constants <ul> <li>"per_system"</li> <li>"gigahertz"</li> </ul> </li> <li>harmonic_vibrational_frequencies <ul> <li>"per_system"</li> <li>"1 / centimeter"</li> </ul> </li> <li>electronic_spatial_extent <ul> <li>"per_system"</li> <li>"nanometer ** 2"</li> </ul> </li> <li>smiles_gdb-17 <ul> <li> "meta_data"</li> <li>"meta_data"</li> </ul> </li> <li>inchi_corina <ul> <li>"meta_data"</li> </ul> </li> <li>inchi_b3lyp <ul> <li>"meta_data"</li> </ul> </li> <li>idx <ul> <li>"meta_data"</li> </ul> </li> <li>tag <ul> <li>"meta_data"<strong><br></strong></li> </ul> </li> </ul> <h2>Original Source:</h2> <p>The QM9 dataset includes 133,885 organic molecules with up to nine total heavy atoms (C,O,N,or F; excluding H) original published by Ramakrishnan, et al. Properties in the QM9 dataset were calculated at the B3LYP/6-31G(2df,p) level of quantum chemistry.</p> <h3>Citations:</h3> <p><em>Original publication:</em></p> <ul> <li>Ramakrishnan, R., Dral, P., Rupp, M. et al."Quantum chemistry structures and properties of 134 kilo molecules." Sci Data 1, 140022 (2014). <a href="https://doi.org/10.1038/sdata.2014.22">https://doi.org/10.1038/sdata.2014.22</a></li> </ul> <p><em>Source dataset, released with CCO 1.0 Universal license:</em></p> <ul> <li>Ramakrishnan, Raghunathan; Dral, Pavlo; Rupp, Matthias; Anatole von Lilienfeld, O. (2014). Quantum chemistry structures and properties of 134 kilo molecules. figshare. Collection. <a href="https://doi.org/10.6084/m9.figshare.c.978904.v5">https://doi.org/10.6084/m9.figshare.c.978904.v5</a></li> </ul>
modelforge curated dataset: OpenFF PhAlkEthOH Dataset
<h1><strong>Modelforge Curated OpenFF PhAlkEthOH Dataset:<br>- 1000 configuration test set, final energy minimized configuration only<br>- Version: nc_1000_minimal_v1.1</strong></h1> <p>This provides a curated hdf5 file for a subset of the OpenFF PhAlkEthOH dataset designed to be compatible with <a href="https://github.com/choderalab/modelforge">modelforge</a>, an infrastructure to implement and train NNPs. The This dataset contains 1000 unique records for 1000 total configurations. This contains a single configuration corresponding to the final configuration in the optimization trajectory (i.e., the energy minimized configuration).</p> <p>This excludes any configurations where the magnitude of any forces on the atoms are greater than 1 hartree/bohr.</p> <p>When applicable, the units of properties are provided in the datafile, encoded as strings compatible with the <a href="https://github.com/openforcefield/openff-units/">openff-units</a> package. For more information about the structure of the data file, please see the following:</p> <h2><strong>Source Dataset:</strong></h2> <p>PhAlkEthOH: <strong>Ph</strong>enyls, <strong>Alk</strong>anes, <strong>Eth</strong>ers, and alcohols (<strong>OH</strong>) </p> <p>The PhAlkEthOH dataset contains a collection of optimization trajectories of linear and cyclic molecules containing phyl rings, small alkanes, ethers, and alcohols containing only elements carbon, oxygen and hydrogen. For each unique molecule, configurations correspond to snapshots from the optimization trajectory. All QM datapoints retrieved from <a href="https://qcarchive.molssi.org/">The MolSSI qcarchive</a> and were generated using B3LYP-D3BJ/DZVP level of theory, the default theory used for force field development by the Open Force Field Initiative.</p> <h2><strong>Properties Included:</strong></h2> <ul> <li>atomic_numbers </li> <li>positions <ul> <li>"per_atom"</li> <li>"nanometer"</li> </ul> </li> <li>dispersion_correction_gradient <ul> <li>"per_atom"</li> <li>"kilojoule_per_mole / nanometer"</li> </ul> </li> <li>dispersion_correction_force <ul> <li>"per_atom"</li> <li>"kilojoule_per_mole / nanometer"</li> </ul> </li> <li>dft_total_gradient <ul> <li>"per_atom"</li> <li>"kilojoule_per_mole / nanometer"</li> </ul> </li> <li>dft_total_force <ul> <li>"per_atom"</li> <li>"kilojoule_per_mole / nanometer"</li> </ul> </li> <li>total_charge <ul> <li>"per_system"</li> <li>"elementary_charge"</li> </ul> </li> <li>dispersion_correction_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>dft_total_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>scf_dipole <ul> <li>"per_system"</li> <li>"elementary_charge * nanometer"</li> </ul> </li> <li>source <ul> <li>"meta_data"</li> </ul> </li> <li>molecular_formula <ul> <li>"meta_data"</li> </ul> </li> <li>canonical_isomeric_explicit_hydrogen_mapped_smiles <ul> <li>"meta_data"</li> </ul> </li> </ul> <h3>Citations:</h3> <p><em>Related manuscripts:</em></p> <ul> <li>Bannan CC, Mobley D. ChemPer: An Open Source Tool for Automatically Generating SMIRKS Patterns. ChemRxiv. 2019; doi:<a href="https://dx.doi.org/10.26434/chemrxiv.8304578.v1">10.26434/chemrxiv.8304578.v1</a></li> <li>Wang Y, Fass J, Kaminow B, Herr JE, Rufa D, Zhang I, Pulido I, Henry M, Macdonald HE, Takaba K, Chodera JD. End-to-end differentiable construction of molecular mechanics force fields. Chemical Science. 2022;13(41):12016-33. doi:<a href="https://doi.org/10.1039%2Fd2sc02739a" target="_blank" rel="noopener noreferrer">10.1039/d2sc02739a</a></li> </ul> <p><em>Source dataset:</em></p> <ul> <li>Gokey, T,., "OpenFF Sandbox CHO PhAlkEthOH v1.0", 2020, <a href="https://github.com/openforcefield/qca-dataset-submission/tree/master/submissions/2020-09-18-OpenFF-Sandbox-CHO-PhAlkEthOH">https://github.com/openforcefield/qca-dataset-submission/tree/master/submissions/2020-09-18-OpenFF-Sandbox-CHO-PhAlkEthOH</a></li> </ul>
modelforge curated dataset: ANI-1x
<h1><strong>Modelforge Curated ANI-1x Dataset:<br></strong><strong>- Full dataset<br>- V</strong><strong>ersion: full_dataset_v1.1:</strong></h1> <p>This provides a curated hdf5 file for the ANI-1x dataset designed to be compatible with <a href="https://github.com/choderalab/modelforge">modelforge</a>, an infrastructure to implement and train NNPs. This dataset contains <span>3114 unique records for 4956005 </span>unique configurations. </p> <p>When applicable, the units of properties are provided in the datafile, encoded as strings compatible with the <a href="https://github.com/openforcefield/openff-units/">openff-units</a> package. </p> <p>This datafile is compatible with modelforge HDF5 schema 2. </p> <p>For more information about the structure of the data file, please see the following:</p> <ul> <li><a href="https://github.com/choderalab/modelforge/wiki/Dataset-and-curation#curation-module">https://github.com/choderalab/modelforge/wiki/Dataset-and-curation#curation-module</a></li> </ul> <h2><strong>Properties Included: </strong></h2> <ul> <li>atomic_numbers </li> <li>positions <ul> <li>"per_atom"</li> <li>"nanometer"</li> </ul> </li> <li>wb97x_dz_forces <ul> <li>"per_atom"</li> <li>"kilojoule_per_mole / nanometer"</li> </ul> </li> <li>wb97x_tz_forces <ul> <li>"per_atom"</li> <li>"kilojoule_per_mole / nanometer"</li> </ul> </li> <li>wb97x_dz_cm5_charges <ul> <li>"per_atom"</li> <li>"elementary_charge"</li> </ul> </li> <li>wb97x_dz_hirshfeld_charges <ul> <li>"per_atom"</li> <li>"elementary_charge"</li> </ul> </li> <li>wb97x_tz_mbis_charges <ul> <li>"per_atom"</li> <li>"elementary_charge"</li> </ul> </li> <li>wb97x_dz_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>wb97x_tz_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>ccsd(t)_cbs_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>hf_dz_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>hf_tz_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>hf_qz_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>npno_ccsd(t)_dz_corr_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>npno_ccsd(t)_tz_corr_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>tpno_ccsd(t)_dz_corr_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>mp2_dz_corr_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>mp2_tz_corr_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>mp2_qz_corr_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>wb97x_dz_dipole <ul> <li>"per_system"</li> <li>"elementary_charge * nanometer"</li> </ul> </li> <li>wb97x_tz_dipole <ul> <li>"per_system"</li> <li>"elementary_charge * nanometer"</li> </ul> </li> <li>wb97x_dz_quadrupole <ul> <li>"per_system"</li> <li>"elementary_charge * nanometer ** 2"<strong><br></strong></li> </ul> </li> </ul> <h2>Source Dataset:</h2> <div> <p>The ANI-1x data set includes properties for small organic molecules that contain H, C, N, and O. This dataset contains nearly 5 million conformers. This data was generated with the wB97X/631Gd level of theory . A subset of the the conformers (~500K) with accurate coupled cluster methods (ANI-1xcc).</p> </div> <h3>Citations:</h3> <p><em>ANI-1x publications:</em></p> <ul> <li> <p>ANI-1x dataset</p> <p>Smith, J. S.; Nebgen, B.; Lubbers, N.; Isayev, O.; Roitberg, A. E. Less Is More: Sampling Chemical Space with Active Learning. J. Chem. Phys. 2018, 148 (24), 241733. <br><a href="https://doi.org/10.1063/1.5023802" rel="nofollow">https://doi.org/10.1063/1.5023802</a></p> </li> <li> <p>ANI-1ccx dataset</p> <p>Smith, J. S.; Nebgen, B. T.; Zubatyuk, R.; Lubbers, N.; Devereux, C.; Barros, K.; Tretiak, S.; Isayev, O.; Roitberg, A. E. Approaching Coupled Cluster Accuracy with a General-Purpose Neural Network Potential through Transfer Learning. Nat. Commun. 2019, 10 (1), 2903. <br><a href="https://doi.org/10.1038/s41467-019-10827-4" rel="nofollow">https://doi.org/10.1038/s41467-019-10827-4</a></p> </li> <li> <p>wB97x/def2-TZVPP data</p> <p>Zubatyuk, R.; Smith, J. S.; Leszczynski, J.; Isayev, O. Accurate and Transferable Multitask Prediction of Chemical Properties with an Atoms-in-Molecules Neural Network. Sci. Adv. 2019, 5 (8), eaav6490. <br><a href="https://doi.org/10.1126/sciadv.aav6490" rel="nofollow">https://doi.org/10.1126/sciadv.aav6490</a></p> </li> <li> <div> </div> </li> </ul> <p><em>Source dataset, released with CCO 1.0 Universal License:</em></p> <ul> <li>Smith, Justin S; Zubatyuk, Roman; Nebgen, Benjamin; Lubbers, Nicholas; Barros, Kipton; Roitberg, Adrian; et al. (2020). The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules. figshare. Collection. <a href="https://doi.org/10.6084/m9.figshare.c.4712477.v1">https://doi.org/10.6084/m9.figshare.c.4712477.v1</a></li> </ul> <p><em>Github repository:</em></p> <ul> <li><a href="https://github.com/aiqm/ANI1x_datasets"><em>https://github.com/aiqm/ANI1x_datasets</em></a></li> </ul>
modelforge curated dataset: SPICE 1
<h1><strong>Modelforge Curated SPICE 1 Dataset:</strong><br><strong>-full dataset, limited to [H, C, N, O, F, Cl, S]</strong><br><strong>-Version: <span>full_dataset_HCNOFClS_v1.1</span></strong></h1> <p>This provides a curated hdf5 file for the SPICE 1 dataset (release v1.1.4) designed to be compatible with <a href="https://github.com/choderalab/modelforge">modelforge</a>, an infrastructure to implement and train NNPs. <span>This dataset contains 16565 unique records for 976408 </span><span>total configurations.</span> The dataset is limited to the elements that are compatible with the ANI2x NNP: [H, C, N, O, F, Cl, S].</p> <p>When applicable, the units of properties are provided in the datafile, encoded as strings compatible with the <a href="https://github.com/openforcefield/openff-units/">openff-units</a> package. </p> <p>This is compatible with modelforge HDF5 schema 2. </p> <p>For more information about the structure of the data file, please see the following:</p> <ul> <li><a href="https://github.com/choderalab/modelforge/wiki/Dataset-and-curation#curation-module">https://github.com/choderalab/modelforge/wiki/Dataset-and-curation#curation-module</a></li> </ul> <h2><strong>Properties Included:</strong></h2> <ul> <li>atomic_numbers </li> <li>positions <ul> <li>"per_atom"</li> <li>"nanometer"</li> </ul> </li> <li>dft_total_force <ul> <li>"per_atom"</li> <li>"kilojoule_per_mole / nanometer"</li> </ul> </li> <li>mbis_charges <ul> <li>"per_atom"</li> <li>"elementary_charge"</li> </ul> </li> <li>mbis_dipoles <ul> <li>"per_atom"</li> <li>"elementary_charge * nanometer"</li> </ul> </li> <li>mbis_quadrupoles <ul> <li>"per_atom"</li> <li>"elementary_charge * nanometer ** 2"</li> </ul> </li> <li>mbis_octupoles <ul> <li>"per_atom"</li> <li>"elementary_charge * nanometer ** 3"</li> </ul> </li> <li>mayer_indices <ul> <li>"per_atom"</li> <li>"dimensionless"</li> </ul> </li> <li>wiberg_lowdin_indices <ul> <li>"per_atom"</li> <li> "dimensionless"</li> </ul> </li> <li>total_charge <ul> <li>"per_system"</li> <li>"elementary_charge"</li> </ul> </li> <li>dft_total_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>formation_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>scf_dipole <ul> <li>"per_system"</li> <li>"elementary_charge * nanometer"</li> </ul> </li> <li>scf_quadrupole <ul> <li>"per_system"</li> <li>"elementary_charge * nanometer ** 2"</li> </ul> </li> <li>smiles <ul> <li>"meta_data"<strong><br></strong></li> </ul> </li> </ul> <h2><strong>Source Dataset:</strong></h2> <p>Small-molecule/Protein Interaction Chemical Energies (SPICE).</p> <p>The SPICE dataset contains 1.1 million conformations for a diverse set of small molecules, dimers, dipeptides, and solvated amino acids. It includes 15 elements, charged and uncharged molecules, and a wide range of covalent and non-covalent interactions. It provides both forces and energies calculated at the ωB97M-D3(BJ)/def2-TZVPPD level of theory, using Psi4 1.4.1 along with other useful quantities such as multipole moments and bond orders.</p> <p> </p> <h3>Citations:</h3> <p><em>Original publication:</em></p> <ul> <li>Eastman, P., Behara, P.K., Dotson, D.L. et al. SPICE, A Dataset of Drug-like Molecules and Peptides for Training Machine Learning Potentials. Sci Data 10, 11 (2023). <a href="https://doi.org/10.1038/s41597-022-01882-6">https://doi.org/10.1038/s41597-022-01882-6</a></li> </ul> <p><em>Source dataset, released with CCO 1.0 Universal license:</em></p> <ul> <li> <div> <div> <div> <div> <div>Eastman, P., Behara, P. K., Dotson, D., Galvelis, R., Herr, J., Horton, J., Mao, Y., Chodera, J., Pritchard, B., Wang, Y., De Fabritiis, G., & Markland, T. (2022). SPICE 1.1.4 (1.1.4) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.8222043" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.8222043</a></div> </div> </div> </div> </div> </li> </ul>
modelforge curated dataset: SPICE 1 OpenFF
<h1><strong>Modelforge Curated SPICE 1 OpenFF Dataset:</strong><br><strong>- full dataset, limited to [H, C, N, O, F, Cl, S]</strong><br><strong>- Version: full_dataset_HCNOFClS_v2.1<br></strong></h1> <p>This provides a curated hdf5 file for the SPICE 1 OpenFF dataset (Open Force Field initiative default level of theory) designed to be compatible with <a href="https://github.com/choderalab/modelforge">modelforge</a>, an infrastructure to implement and train NNPs. This dataset contains 100 unique records for 1000 total configurations, with a maximum of 10 configurations per record.<br>The dataset is limited to the elements that are compatible with ANI2x NNP: [H, C, N, O, F, Cl, S]. This excludes any configurations where the magnitude of any forces <br>on the atoms are greater than 1 hartree/bohr.</p> <p>When applicable, the units of properties are provided in the datafile, encoded as strings compatible with the <a href="https://github.com/openforcefield/openff-units/">openff-units</a> package. </p> <p>This is compatible with modelforge HDF5 schema 2.</p> <p>For more information about the structure of the data file, please see the following:</p> <ul> <li><a href="https://github.com/choderalab/modelforge/wiki/Dataset-and-curation#curation-module">https://github.com/choderalab/modelforge/wiki/Dataset-and-curation#curation-module</a></li> </ul> <h2><strong>Properties Included:</strong></h2> <ul> <li>atomic_numbers </li> <li>positions <ul> <li>"per_atom"</li> <li>"nanometer"</li> </ul> </li> <li>dft_force <ul> <li>"per_atom"</li> <li>"kilojoule_per_mole / nanometer"</li> </ul> </li> <li>mbis_charges <ul> <li>"per_atom"</li> <li>"elementary_charge"</li> </ul> </li> <li>dispersion_correction_force <ul> <li>"per_atom"</li> <li>"kilojoule_per_mole / nanometer"</li> </ul> </li> <li>dft_total_force <ul> <li>"per_atom"</li> <li>"kilojoule_per_mole / nanometer"</li> </ul> </li> <li>total_charge <ul> <li>"per_system"</li> <li>"elementary_charge"</li> </ul> </li> <li>dft_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>scf_dipole <ul> <li>"per_system"</li> <li>"elementary_charge * nanometer"</li> </ul> </li> <li>dispersion_correction_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>dft_total_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>source <ul> <li>"meta_data"</li> </ul> </li> <li>molecular_formula <ul> <li>"meta_data"</li> </ul> </li> <li>canonical_isomeric_explicit_hydrogen_mapped_smiles <ul> <li> "meta_data"<strong><br></strong></li> </ul> </li> </ul> <h2><strong>Source Dataset:</strong></h2> <p>Small-molecule/Protein Interaction Chemical Energies (SPICE).</p> <p>The SPICE dataset contains 1.1 million conformations for a diverse set of small molecules, dimers, dipeptides, and solvated amino acids. It includes 15 elements, charged and uncharged molecules, and a wide range of covalent and non-covalent interactions. </p> <p>It provides both forces and energies calculated using B3LYP-D3BJ/DZVP level of theory, using Psi4 1.4.1. This is the default theory used for force field development by the <a href="https://openforcefield.org/">Open Force Field Initiative</a>. </p> <div> <p>This includes the following collections from the <a href="https://qcarchive.molssi.org">MolSSI qcarchive</a> (these are also included in the standard SPICE 1 dataset):</p> <ul> <li>"SPICE Solvated Amino Acids Single Points Dataset v1.1",</li> <li>"SPICE Dipeptides Single Points Dataset v1.2",</li> <li>"SPICE DES Monomers Single Points Dataset v1.1",</li> <li>"SPICE DES370K Single Points Dataset v1.0",</li> <li>"SPICE PubChem Set 1 Single Points Dataset v1.2",</li> <li>"SPICE PubChem Set 2 Single Points Dataset v1.2",</li> <li>"SPICE PubChem Set 3 Single Points Dataset v1.2",</li> <li>"SPICE PubChem Set 4 Single Points Dataset v1.2",</li> <li>"SPICE PubChem Set 5 Single Points Dataset v1.2",</li> <li>"SPICE PubChem Set 6 Single Points Dataset v1.2",</li> </ul> </div> <p>This does not include the following collections (which are part of the standard SPICE 1 dataset):</p> <div> <ul> <li>"SPICE Ion Pairs Single Points Dataset v1.1",</li> <li>"SPICE DES370K Single Points Dataset Supplement v1.0",</li> </ul> </div> <h3><strong>Citations:</strong></h3> <p><em>Original SPICE 1 publication:</em></p> <ul> <li>Eastman, P., Behara, P.K., Dotson, D.L. et al. SPICE, A Dataset of Drug-like Molecules and Peptides for Training Machine Learning Potentials. Sci Data 10, 11 (2023). <a href="https://doi.org/10.1038/s41597-022-01882-6">https://doi.org/10.1038/s41597-022-01882-6</a></li> </ul>
modelforge curated dataset: SPICE 2
<h1><strong>Modelforge Curated SPICE 2 Dataset:<br>- Full dataset limited to elements [H, C, N, O, F]<br>- Version: full_dataset_HCNOF_v1.1</strong></h1> <p>This provides a curated hdf5 file for a subset of the SPICE 2 dataset (release v2.0.1) designed to be compatible with <a href="https://github.com/choderalab/modelforge">modelforge</a>, an infrastructure to implement and train NNPs. This dataset contains 57037 unique records for 928073 total configurations. The dataset is limited to the elements ['H', 'C', 'N', 'O', 'F'].</p> <p>When applicable, the units of properties are provided in the datafile, encoded as strings compatible with the <a href="https://github.com/openforcefield/openff-units/">openff-units</a> package. </p> <p>This is compatible with modelforge hdf5 schema 2.</p> <p>For more information about the structure of the data file, please see the following:</p> <ul> <li><a href="https://github.com/choderalab/modelforge/wiki/Dataset-and-curation#curation-module">https://github.com/choderalab/modelforge/wiki/Dataset-and-curation#curation-module</a></li> </ul> <h2><strong>Source Dataset:</strong></h2> <p>Small-molecule/Protein Interaction Chemical Energies (SPICE).</p> <p>The SPICE 2 dataset contains roughly 2 million conformations for a diverse set of small molecules, dimers, dipeptides, and solvated amino acids. It includes 17 elements, charged and uncharged molecules, and a wide range of covalent and non-covalent interactions. SPICE 2 is an update to spice 1, roughly double the total dataset size and including 2 additional elements. It provides both forces and energies calculated at the ωB97M-D3(BJ)/def2-TZVPPD level of theory, using Psi4 1.4.1 along with other useful quantities such as multipole moments and bond orders.</p> <div> <pre> </pre> </div> <h2><strong>Properties Included:</strong></h2> <ul> <li>atomic_numbers </li> <li>positions <ul> <li>"per_atom"</li> <li>"nanometer"</li> </ul> </li> <li>dft_total_force <ul> <li>"per_atom"</li> <li>"kilojoule_per_mole / nanometer"</li> </ul> </li> <li>mbis_charges <ul> <li>"per_atom"</li> <li>"elementary_charge"</li> </ul> </li> <li>mbis_dipoles <ul> <li>"per_atom"</li> <li>"elementary_charge * nanometer"</li> </ul> </li> <li>mbis_quadrupoles <ul> <li>"per_atom"</li> <li>"elementary_charge * nanometer ** 2"</li> </ul> </li> <li>mbis_octupoles <ul> <li>"per_atom"</li> <li>"elementary_charge * nanometer ** 3"</li> </ul> </li> <li>mayer_indices <ul> <li>"per_atom"</li> <li>"dimensionless"</li> </ul> </li> <li>wiberg_lowdin_indices <ul> <li>"per_atom"</li> <li>"dimensionless"</li> </ul> </li> <li>total_charge <ul> <li>"per_system"</li> <li>"elementary_charge"</li> </ul> </li> <li>dft_total_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>formation_energy <ul> <li>"per_system"</li> <li>"kilojoule_per_mole"</li> </ul> </li> <li>scf_dipole <ul> <li>"per_system"</li> <li>"elementary_charge * nanometer"</li> </ul> </li> <li>scf_quadrupole <ul> <li>"per_system"<br>"elementary_charge * nanometer ** 2"</li> </ul> </li> <li>smiles <ul> <li>"meta_data"</li> </ul> </li> </ul> <p> </p> <h2><strong>Citations:</strong></h2> <div> <p><em>Spice 2 publication:</em></p> <ul> <li> <p>Eastman, P., Pritchard, B. P., Chodera, J. D., & Markland, T. E. Nutmeg and SPICE: models and data for biomolecular machine learning. Journal of chemical theory and computation, 20(19), 8583-8593 (2024). <a href="https://doi.org/10.1021/acs.jctc.4c00794">https://doi.org/10.1021/acs.jctc.4c00794</a></p> </li> </ul> </div> <p><em>Original SPICE 1 publication:</em></p> <ul> <li>Eastman, P., Behara, P.K., Dotson, D.L. et al. SPICE, A Dataset of Drug-like Molecules and Peptides for Training Machine Learning Potentials. Sci Data 10, 11 (2023). <a href="https://doi.org/10.1038/s41597-022-01882-6">https://doi.org/10.1038/s41597-022-01882-6</a></li> </ul> <p><em>Source dataset, released with CCO 1.0 Universal license:</em></p> <ul> <li> <div> <div> <div> <div> <div>Eastman, P., Behara, P. K., Dotson, D., Galvelis, R., Herr, J., Horton, J., Mao, Y., Chodera, J., Pritchard, B., Wang, Y., De Fabritiis, G., & Markland, T. (2024). SPICE 2.0.1 (2.0.1) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.10975225" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.10975225</a></div> </div> </div> </div> </div> </li> </ul>
PlayMyData: a curated dataset of multi-platform videogames
<h2><strong>About</strong></h2><p>This repository contains the source code implementation used to replicate the experimental results obtained in the submitted to the 21st International Conference on Mining Software Repositories (MSR204).</p><p><i>"PlayMyData: a curated dataset of multi-platform videogames"</i></p><p>authored by:</p><p>Andrea D'Angelo(1), Claudio Di Sipio, Cristiano Politowsky (2) and Riccardo Rubei</p><p>(1) Università degli Studi dell'Aquila, Italy</p><p>(2) University of Montreal, Canada</p><h2><strong>Introduction </strong></h2><p>PlayMyData is a multi-purpose, comprehensive videogame dataset of videogames released from 1993 up to November 2023. It contains metadata like titles, platforms, a summary of the story, and release data. It also integrates data from HowLongToBeat on completion times.</p><h2><strong>Data description</strong></h2><p>The dataset is structured as follows: </p><ul><li><strong>all_games_PlayStation.csv</strong>: It contains IGDB metadata collected for the PlayStation platforms.</li><li><strong>screenshots.zip</strong>: It contains the collected screenshots from IGDB, grouped by genre</li><li><strong>all_games_Xbox.csv</strong>: It contains IGDB metadata collected for the Xbox platforms.</li><li><strong>all_games_PC.csv</strong>: It contains IGDB metadata collected for the PC.</li><li><strong>all_games_Nintendo.csv</strong>: It contains IGDB metadata collected for the Nintendo platforms.</li><li><strong>platforms.csv</strong>: It contains all gaming platforms available on IGDB and the corresponding ID.</li><li><strong>genres.csv</strong>: It contains the list of genres available on IGDB and the corresponding ID.</li><li><strong>all_videos.csv</strong>: It contains the gameplay URLs and the corresponding name as it appears on IGDB, e.g. "Gameplay Video" </li><li><strong>video_ids.csv</strong>: It contains the mapping with games and the list of gameplay videos. Some entries are missing.</li></ul><h2><strong>How to collect PlayMyData</strong></h2><p>To collect PlayMyData, please refer to the supporting GitHub repo available at: https://github.com/riccardoRubei/MSR2024-Data-Showcase</p><p> </p>
Curated dataset for analysis for the paper "Decision Support Systems Adoption in Pesticide Management"
<p>Dataset created from farmer responses to a survey on the decision support systems adoption for intergrated pest management in the framework of the EU funded project IPM Decisions.</p>
BTyperDB: a community-curated, global atlas of Bacillus cereus sensu lato genomes for epidemiological surveillance
<p>The ability to cause foodborne illness, anthrax, and other infections has been attributed to numerous lineages within <em>Bacillus cereus sensu lato</em> (<em>s.l.</em>). However, existing pathogen surveillance databases facilitate dangerous pathogen misidentifications when applied to <em>B. cereus s.l.</em>, potentially hindering outbreak or bioterrorism attack response efforts. To address this, we developed BTyperDB (<a href="http://www.btyper.app/">www.btyper.app</a>), an atlas of <em>B. cereus s.l.</em> genomes with standardized, community-curated metadata. BTyperDB aggregates all publicly available <em>B. cereus s.l.</em> genomes (including >2,600 previously unassembled genomes) with novel genomes donated by laboratories around the world, nearly doubling the number of publicly available <em>B. cereus s.l.</em> genomes. To showcase its utility for pathogen surveillance, we use BTyperDB to identify emerging anthrax toxin- and capsule-harboring lineages. Overall, our study provides insight into the epidemiology of an under-studied group of emerging pathogens and highlights the benefits of inclusive, community-driven metadata FAIRification efforts.</p>
AAA-100: A Curated Dataset of 3D Watertight Abdominal Aortic Aneurysm Models
<p>An abdominal aortic aneurysm (AAA) is a local dilatation of the abdominal aorta exceeding 30 mm that might rupture, with fatal outcomes in 70-80% of cases. Personalized 3D models of AAAs, including surrounding vasculature such as iliac and renal arteries play an important role in tailored clinical decision-making for AAA patients. Models could be used for, e.g., AAA growth modeling, stentgraft sizing and positioning for endovascular aorta repair (EVAR) procedures, or 3D printing for surgical practice. Extracting high-quality 3D arterial models from imaging modalities such as computed tomography angiography (CTA) is a time-consuming and challenging problem. For downstream applications such as computational fluid dynamics (CFD) or shape analysis, models should have sub-voxel accuracy, be watertight, and adhere to topological constraints. We present the AAA-100 dataset, containing 100 detailed 3D AAA models with consistent anatomical boundaries acquired semi-automatically from pre-operative CTA scans. These models span a wide range of possible AAA pathology. Moreover, all models are carefully curated to be anatomically and topologically correct. </p> <p>A detailed description of the data set and file structure is provided in description.pdf.</p> <p>We kindly ask you to cite the following works when using the AAA-100 dataset in your research</p> <blockquote> <p>Alblas, D., Suk, J., Brune, C., Yeung, K. K., & Wolterink, J. M. (2025). SIRE: Scale-invariant, rotation-equivariant estimation of artery orientations using graph neural networks. <em>Medical Image Analysis</em>, 103467.</p> <p>Rygiel, P., Alblas, D., Brune, C., Yeung, K. K., & Wolterink, J. M. (2024). Global Control for Local SO (3)-Equivariant Scale-Invariant Vessel Segmentation. <em>arXiv preprint arXiv:2403.15314</em>.</p> </blockquote>
modelforge curated dataset: tmQM
<h1>Curated tmQM Dataset:</h1> <h2>Full dataset, version "full_dataset_v1":</h2> <p>This provides a curated hdf5 file for the tmQM dataset (release 13Aug2024) designed to be compatible with <a href="https://github.com/choderalab/modelforge">modelforge</a>, an infrastructure to implement and train NNPs. This datafile includes 108541 unique molecules. Note, only a single configuration per unique molecule is provided. </p> <p>Change from full_dataset_v0: fixed minor labeling bug and scaling issue in the scaled version of the computed dipole moment.</p> <p>When applicable, the units of properties are provided in the datafile, encoded as strings compatible with the <a href="https://github.com/openforcefield/openff-units/">openff-units</a> package. For more information about the structure of the data file, please see the following:</p> <ul> <li><a href="https://github.com/choderalab/modelforge/wiki/Dataset-and-curation#curation-module">https://github.com/choderalab/modelforge/wiki/Dataset-and-curation#curation-module</a></li> </ul> <p>This curated dataset was generated using the modelforge software at commit <add commit>:</p> <ul> <li>Link to the source code at this commit: <add commit></li> <li>Link to the script file used to generate the dataset: <add commit></li> </ul>
Curated new database list from the 2024 NAR Database issue - Table 1
<p>This is a curated list of 90 new databases recently published in the <a href="https://doi.org/10.1093/nar/gkad1173">2024 Nucleic Acids Research (NAR) Database issue</a>. These databases were specifically categorized as the "new databases" (not previously published) in Table 1 of the <a href="https://doi.org/10.1093/nar/gkad1173">2024 compilation paper</a>. In this dataset, we curated a set of data columns to facilitate knowledge modeling and database integration efforts:</p> <ul> <li>"has API": does the database provide API access, Yes or No.</li> <li>"Primary Entity Type": the primary biomedical entity types captured in the database.</li> <li>"Downloadable": does the database provide direct file downloads, Yes or No.</li> <li>"license specified?": does the database specify a data license explicitly, Yes or No.</li> <li>"knowledge level": the level of knowledge expressed in the database, example values are specified in <a href="https://biolink.github.io/biolink-model/knowledge_level/">the BioLink model</a>.</li> <li>"agent type": the high-level category of agent who originally generated a statement of knowledge or other type of information, example values are specified in <a href="https://biolink.github.io/biolink-model/agent_type/">the BioLink model</a>.</li> <li>"notes": optional extra notes</li> </ul>
Highly curated hERG dataset of 8879 unique molecular compounds with corresponding potency values
<p>This dataset was built during a research project, in the field of Computer-Aided Drug Discovery (CADD), funded by the Natural Sciences and Engineering Research Council of Canada (NSERC) Discovery grant. The aim of the project was to build descriptor-based machine learning models for hERG cardiotoxicity liability predictions. The dataset includes a total of 8879 unique molecular compounds gathered from ChEMBL and PubChem publicly available bioactivity databases, as well as from literature mining. The list is split into 2 sets, 8380 for training and 499 for testing. All molecular compounds are represented in their SMILE format with their corresponding PIC50 potency values.</p> <p>To access the full original work, please visit the following link: <a href="https://arxiv.org/abs/2112.13467">Manuscript</a></p> <p><strong>Note:</strong> Upon usage of this data, kindly cite the original manuscript describing the curation process:<br><br>Arab, Issar, and Khaled Barakat. "ToxTree: descriptor-based machine learning models for both hERG and Nav1. 5 cardiotoxicity liability predictions." <em>arXiv preprint arXiv:2112.13467</em> (2021).</p> <p> </p> <p><strong>Refer to our latest manually curated and a much larger dataset here: <a href="../record/8359714">link</a> </strong></p>
Highly curated Nav1.5 dataset of 1723 unique molecular compounds with corresponding potency values
<p>This dataset was built during a research project, in the field of Computer-Aided Drug Discovery (CADD), funded by the Natural Sciences and Engineering Research Council of Canada (NSERC) Discovery grant. The aim of the project was to build descriptor-based machine learning models for Nav1.5 cardiotoxicity liability predictions. The dataset includes a total of 1723 unique molecular compounds gathered from ChEMBL and PubChem publicly available bioactivity databases. The list is split into 2 sets, 1550 for training and 173 for testing. All molecular compounds are represented in their SMILE format with their corresponding PIC50 potency values.</p> <p>To access the full original work, please visit the following link: <a href="https://arxiv.org/abs/2112.13467">Manuscript</a> <br><strong>Refer to a much larger and latest dataset here: <a href="../record/8359714">link</a> </strong></p> <p><br><strong>Note:</strong> Upon usage of this data, kindly cite both the dataset and the original manuscript describing the curation process as written below.<br><br><em>Arab I, Barakat K. ToxTree: Descriptor-based machine learning models for both hERG and Nav1. 5 cardiotoxicity liability predictions. arXiv preprint arXiv:2112.13467. 2021</em></p> <p><em>Arab, Issar, & Barakat, Khaled. (2021). Highly curated Nav1.5 dataset of 1723 unique molecular compounds with corresponding potency values [Data set]. Zenodo. https://doi.org/10.5281/zenodo.5807731</em></p>
MetaPro: a web-based metabolomics application for MS data batch inspection and library curation
<p>MetaPro is a metabolomics web analysis platform built on the Aird data format with high performance and high compression. This platform includes a series of necessary functions for metabolomics analysis such as quality control, retention time(RT) alignment, target analysis, untarget analysis, manual integration, batch inspection, MS2 library establishment, and report export, providing efficient data analysis, management and visualization capabilities</p>
miRNAture v.1.1: updated dataset with curated metazoan miRBase v.22.1 and Rfam-14.4 miRNA families.
<p>This folder contains the updated data to annotate miRNAs using miRNAture v.1.1. As usual, it includes CMs, HMMs and required pre-calculated data to validate mature miRNAs (corrected hairpins, mature files and miRBase genomes). It contains a 1034 and 1124 metazoan miRNA families curated from miRBase v.22.1 and Rfam 14.4, respectively. To use this dataset, refer the path of this uncompressed folder with the -dataF/-datadir flag in miRNature v.1.1.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.