Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

21,281

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

21,281 results for “Molecular”

Learn how ShareScore rates datasets ↗
edi60/100

Ramped Pyrolysis Oxidation (RPO) coupled radiocarbon (14C-DOC) and stable carbon (13C-DOC), high-resolution molecular composition (FT-ICR MS), and biodegradable dissolved organic carbon (BDOC) of groundwater, river water, and lagoon water in northeast Alaska, 2017

Supra-permafrost groundwater (SPGW), river water, and lagoon water were sampled near Kaktovik, AK to assess the reactivity and origin of dissolved organic matter (DOM) across interconnected hydrologic systems during late summer. Water samples were collected on August 17th 2017 from SPGW along the beach of Jago Lagoon (Jago GW), surface water from the Jago River’s main channel above tidal influence (Jago R), and from the water column of Kaktovik Lagoon at 2–3 m depth (KA LW). Measurements were made from grab samples for river and lagoon water, and from a composite sample for SPGW gathered from 10 individual shoreline locations. Data include dissolved organic carbon concentration (DOC, mg C L-1), Ramped Pyrolysis Oxidation (RPO) derived fraction compositions of 13C-DOC (δ13C ‰), 14C-DOC (in fraction modern), and method/instrumental error in the 14C and 13C results, and Fourier Transform Ion Cyclotron Resonance Mass Spectrometry (FT-ICR MS) molecular composition and summarized compound classes. Biodegradable DOC (BDOC) bottle experiments were performed using all three sample types, where DOC concentration was subsequently measured at 2, 7, 14, and 28 days. FT-ICR MS composition was subsequently measured at the 28-day timepoint to track changes in molecular formulae and compound class relative abundance following biodegradation. Data from RPO serial thermal oxidation include temperature and normalized CO2 profiles for each background sample. Thermal-oxidation profiles of CO2 were transformed into non-parametric activation energy (E) distributions using an inverse model. Model output includes C mass of oxidized CO2 (µg C), Tmax (K), Emax (kJ mol-1), Emean (kJ mol-1), Estd (kJ mol-1), and p(0,E)max of user-defined sample fractions. FT-ICR MS results include a summary table of the relative abundance of compound classes (e.g., unsaturated phenolic, polyphenolic, aliphatic, condensed aromatics, peptide-like) and elemental groupings (e.g., CHO-type, CHON-type, CHOS-type, CHON

openCC0Jan 2026View details →
edi60/100

Fluxes of Molecular Hydrogen (H2) at Harvard Forest EMS Tower 2010-2012

Molecular hydrogen (H2) is an atmospheric trace gas with a large microbe-mediated soil sink, yet cycling of this compound throughout ecosystems is poorly understood. Measurements of the sources and sinks of H2 in various ecosystems are sparse, resulting in large uncertainties in the global H2 budget. Constraining the H2 cycle is critical to understanding its role in atmospheric chemistry and climate. We measured H2 fluxes at high frequency in a temperate mixed deciduous forest for 15 months using a tower-based flux-gradient approach to determine both the soil-atmosphere and the net ecosystem flux of H2. The data presented here along with other data available at Harvard Forest can be used for efforts to model the H2 soil sink.

openCC0Dec 2023View details →
edi56/100

Molecular composition of dissolved organic matter in NTL-LTER lakes detected by Fourier-transform ion cyclotron resonance mass spectrometry

The composition of dissolved organic matter (DOM) varies widely in the environment due to distinct sources of the material and subsequent processing. DOM composition drives its reactivity in terms of many processes including photochemical reactions, microbial metabolism, and carbon cycling within water bodies. This study uses ultra-high resolution mass spectrometry via a Fourier-transform ion cyclotron resonance mass spectrometer (FT-ICR MS) to evaluate DOM composition at the molecular level to determine differences in DOM composition among the NTL-LTER lakes. Whole water samples were collected from the surface of each lake near the shore on August 18th and 19th in 2016 in. Ultraviolet-visible spectra were recorded as light absorbance can also give information about DOM composition. Additionally, concentrations of anions, cations, and pH were measured waters because these can all alter DOM reactivity in the environment. Both water chemistry and DOM composition vary widely among the lakes with the bogs displaying the most terrestrial-like signature in DOM and the oligotrophic lakes show more microbial-like or environmentally processed DOM.

openCC (other)Dec 2022View details →
edi56/100

Molecular composition of dissolved organic matter from Lake Mendota from June – November 2017, analyzed by Fourier-transform ion cyclotron resonance mass spectrometry

Dissolved organic matter (DOM) is a complex mixture of organic compounds found in all natural waters. Its composition affects its reactivity towards numerous processes. Its composition is a function of both its source (e.g., allochthonous or autochthonous) as well as the extent of environmental processing it has undergone (e.g., chemical or biological degradation). Fourier-transform ion cyclotron resonance mass spectrometry (FT-ICR MS) allows for the characterization of dissolved organic matter at the molecular level. The water sample was collected near the NTL-LTER research buoy on Lake Mendota. Formula assignments were made to raw mass to charge ratios detected in the mass spectrum using a custom processing script and resulting in a list of chemical formulas making up the DOM sample.

openCC (other)Dec 2022View details →
zenodo52/100

Supplementary Material to article "Molecular Diversity of Mycobacterium avium subsp. paratuberculosis in Four Dairy Goat Herds from Thuringia (Germany)"

<p>These data (supplementary material) belong to the publication "Molecular Diversity of <i>Mycobacterium avium</i> subsp. <i>paratuberculosis</i> in Four Dairy Goat Herds from Thuringia (Germany)". The study determined the diversity of <i>Mycobacterium avium</i> subsp. <i>paratuberculosis</i> (MAP) isolated from four goat herds affected by paratuberculosis in Thuringia (Germany), as well as the detailed distribution of MAP genotypes among the animals and their environment in one herd (herd 1). A combination of three methods was used to genotype isolates from fecal samples of infected goats, from various intestinal and other tissues of clinically affected goats, and from environmental samples. The six MAP-C genotypes identified could be assigned to five different phylogenetic subgroups. The results suggest individual infection strains within each herd. In herd 1, one predominant strain was found, and two strains occurred sporadically. The identified genotypes were not goat specific.</p>

opencc-by-4.0Nov 2023View details →
zenodo52/100

Dataset of "Neutron imaging and molecular simulation of systems from methane and p‑xylene"

<p>The dataset contains parameterizations, and input files for molecular dynamics simulations used in the study of methane dissolution in p-xylene. For selected conditions, full simulation data, i.e., trajectories and energetics are provided. All used simulation results data are provided in the table, along with the measured experimental data.</p>

opencc-by-4.0Dec 2024View details →
zenodo52/100

Dataset of "Molecular dynamics of evaporative cooling of water clusters"

<p>The cooling of water clusters through evaporation into a vacuum is studied using classical molecular dynamics with the SPC water model, and the results are compared with semimacroscopic theory. A model based on the Hertz&ndash;Knudsen equation underestimates the cooling rates. A modified approach, which accounts for the Kelvin equation, provides better results. While the rotational temperature of the clusters is in equilibrium with their internal temperature, the translational temperature of the clusters &ldquo;as individual particles&rdquo; remains unchanged.</p>

opencc-by-4.0Oct 2024View details →
zenodo52/100

Dataset of "Molecular Dynamics Simulations Unveil the Aggregation Patterns and Salting out of Polyarginines at Zwitterionic POPC Bilayers in Solutions of Various Ionic Strengths"

<p>Molecular dynamics simulations are performed for a series of model cell-penetrating peptides (in particular nona-arginines) in aqueous solutions, in contact with model phosphocholine (POPC) membranes in conditions of different ionic strengths. The unusual aggregation properties of peptides at model lipid bilayers are analyzed and different sizes and lifetimes of aggregates are presented.<br>This dataset contains molecular dynamics simulation data with trajectories, input files, and topology files for all studied systems. They contain low peptide concentration in water, low NaCl concentration, high NaCl concentration, low CaCl2 concentration, and high CaCl2 concentration.<br>In addition to low peptide concentration, high peptide concentration in water, low NaCl concentration, high NaCl concentration, low CaCl2 concentration, and high CaCl2 concentration are also studied.</p>

opencc-by-4.0May 2024View details →
zenodo52/100

Liquid Chromatography - Tandem Mass Spectrometry (LC-MS/MS) and Gas Chromatography - Mass Spectrometry (GC-MS) Reference Libraries from Global Natural Products Social Molecular Networking (GNPS) and National Institute of Standards and Technology (NIST) WebBook Processed for Spectral Library Matching

<div>In order to obtain a high-quality LC-MS/MS reference database for spectral library matching, we selected 22 high-quality GNPS tandem mass spectrometry databases generated under the positive ion mode. Further preprocessing similar to Huber et al involving mass-to-charge (m/z) and intensity filtering yields the database found in the file LCMS_GNPS_reference_library.csv which contains 14,705 electrospray ionization (ESI) mass spectra, each of which corresponds to a unique compound. The NIST WebBook database was used to construct GC-MS database contained in the file GCMS_NIST_WebBook.csv. This database contains 23,721 electron ionization (EI) mass spectra, each of which corresponds to a unique non-hyphenated Chemical Abstract Service (CAS) Registry Number.</div> <div>&nbsp;</div> <div>Both LC-MS/MS and GC-MS databases are organized into three columns: one for the identifier, one for the m/z values, and one for the intensity values. For example, if spectrum A has 20 ion fragments, then there will be 20 rows corresponding to spectrum A in the corresponding database with the identifier A repeated 20 times with the corresponding m/z and intensity values.</div>

opencc-by-4.0Jul 2024View details →
zenodo48/100

Computational Supporting Information for How Chemical Environment Activates Anthralin and Molecular Oxygen for Direct Reaction

<p>The updated version of the dataset contains all original computational results, including validation of the level of theory, molecular structures, and analysis spreadsheets that are in support of our experimental observations of spontaneous reactivity of anthralin/dithranol molecule with molecular oxygen without any catalyst or co-substrate.<br> The paper was published in Journal of Organic Chemistry, 2020, 85(2), 1315&ndash;1321 (DOI: 10.1021/acs.joc.9b03133).</p> <p>In the meantime, the science was also also presented at the 8th ELSI Symposium, Tokyo Institute of Technology, Tokyo (Japan); February 3-7, 2020 in the context of molecular catalysis and their role in the chemical evolution of the building blocks of life.</p> <p>This version also has an important update that is being exclusively published here on Zenodo. The selected level of theory (MN15 functional with triple-zeta quality basis set supplemented with BOTH diffuse and polarization basis functions) is further confirmed to be one of the most reasonable one among 98 commonly used functionals.</p>

opencc-by-4.0Apr 2020View details →
zenodo48/100

Molecular datasets from "SMILES-Based Deep Generative Scaffold Decorator for De-Novo Drug Design"

<p>Herein find the molecular datasets from &quot;<a href="https://chemrxiv.org/articles/SMILES-Based_Deep_Generative_Scaffold_Decorator_for_De-Novo_Drug_Design/11638383">SMILES-Based Deep Generative Scaffold Decorator for De-Novo Drug Design</a>&quot;. These were generated with&nbsp;SMILES-based scaffold decorator generative models&nbsp;trained with two training sets (DRD2 and ChEMBL). These generative models require a partially-built molecule (scaffold) as input and output several possible completions for each scaffold. Each dataset corresponds to a model trained with the&nbsp; ChEMBL or DRD2&nbsp;sets, wither multi-step (ms) or single-step (ss) and the provenance of the scaffolds (validation set, or non-dataset).</p> <p>The molecules generated are annotated with a set of descriptors. The DRD2 datasets have the predicted probability of each molecule to be active&nbsp;on DRD2 (p)&nbsp;obtained from a Random Forest model. The ChEMBL model&#39;s descriptors are related to the synthesizability of the molecules (see manuscript). Also, the datasets decorated from validation set scaffolds are annotated whether they are part of the validation set (in_validation).</p>

opencc-by-4.0Jan 2020View details →
zenodo48/100

Bulk and single-cell gene expression profiling of SARS-CoV-2 infected human cell lines identifies molecular targets for therapeutic intervention

<p>Single cell RNA seq datasets used for analysis in the&nbsp;Bulk and single-cell gene expression profiling of SARS-CoV-2 infected human cell lines identifies molecular targets for therapeutic intervention</p>

opencc-by-4.0Sep 2020View details →
zenodo48/100

Data for: Temperature-controlled Molecular Bonding Hysteresis: Interphase Dynamics of a Nanoparticle-modified Polymer Network

<p>The data is supplementary to the publication "Temperature-controlled Molecular Bonding Hysteresis: Interphase Dynamics of a Nanoparticle-modified Polymer Network", DOI: <a title="DOI URL" href="https://doi.org/10.1021/acs.jpclett.4c00406">10.1021/acs.jpclett.4c00406</a></p> <p>Key words: Thermal volume expansion, Interphase dynamics, Temperature-modulated optical refractometry, Nanoparticles, Optical Remanence, Hysteresis, Refractive index</p> <p>The data sets contain measured and processed data on the interphase dynamics of a nanoparticle modified epoxy resin collected via Temperature-modulated optical refractometry (TMOR).</p> <p>Material details:</p> <ul> <li>Cycloaliphatic epoxy resin + Anhydride curing agent + 1-methylimidazole</li> <li>Core-shell rubber nanoparticles, 100 nm, dispersed in a cycloaliphatic epoxy carrier resin</li> </ul> <p>Funding received from:</p> <ul> <li>German Research Foundation (DFG), project number: 521902629.</li> </ul>

opencc-by-4.0Mar 2024View details →
zenodo48/100

Data of publication Ultra-narrow Optical Linewidths in Rare-Earth Molecular Crystals

<p>Data corresponding to main text Figures, Extended Data figures, and Supplementary Figures in publication &#39;Ultra-narrow Optical Linewidths in Rare-Earth Molecular Crystals, by D. Serrano, S. Kumar Kuppusamy, B. Heinrich, O. Fuhr, M. Ruben and P. Goldner.</p>

opencc-by-4.0Nov 2021View details →
zenodo48/100

Inactive to active transition of human Thymidine Kinase 1 revealed by Molecular Dynamics simulations

<p>The trajectories and input files for the manuscript <em>Inactive to active transition of human Thymidine</em></p> <p><em>Kinase 1 revealed by Molecular Dynamics simulations</em> (<a href="https://doi.org/10.1021/acs.jcim.1c01157">https://doi.org/10.1021/acs.jcim.1c01157</a>)&nbsp;</p> <p>ABSTRACT</p> <p>Despite its importance for the nucleoside (and nucleoside prodrug) metabolism, the structure<br> of the active conformation of human Thymidine Kinase 1 (hTK1) remains elusive. We perform<br> microsecond molecular dynamics simulations of the inactive enzyme form bound to a<br> bisubstrate inhibitor that was shown experimentally to activate another TK1-like kinase,<br> Thermotoga maritima TK (TmTK). Our results are in excellent agreement with the<br> experimental findings for the TmTK closed-to-open state transition. We show that the inhibitor<br> induces an increase of the enzyme radius of gyration due to the expansion on one of the dimer<br> interfaces; the structural changes observed, including the active site pocket volume increase,<br> decrease in monomer-monomer buried surface area and of the number of hydrogen bonds (as<br> compared to the inactive enzyme control simulation), show that the catalytically competent<br> (open) conformation of hTK1 can be assumed in the presence of an activating ligand.</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Morphometric data from: Incongruent molecular and morphological variation in the crab spider Synema globosum (Araneae: Thomisidae) in Europe

<p>Here we provide the complete set of files used by <a href="https://doi.org/10.3897/zookeys.1078.64116">Urfer et al. (2021</a>, see References section below for the complete citation of the publication) for the morphometric and the molecular analysis. In particular, we provide the following documents:</p> <p><br> PART 1: MORPHOMETRIC ANALYSIS</p> <p>- 1_Synema_data_multiple_imputation_mice.R: R-script used for replacing NAs.</p> <p>- 1_Synema_data_NA_imputed.csv: Dataset with raw values (in millimeters) of all 28 specimens used for the morphometric analysis. Each specimen was measured 4 times. NAs replaced using the R-script &quot;Synema_multiple_imputation_mice.R&quot; above. This is the datafile used for all morphometric analyses.</p> <p>- 1_Synema_data_with_NA.csv: Dataset with raw values (in millimeters) of all 28 specimens. Each specimen was measured 4 times. NAs not replaced.<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;<br> - 1_Synema_Reliability.R: R-script for calculating reliability.<br> &nbsp;&nbsp; &nbsp;<br> - 1_Synema_Reliability_supplementary_figure.pdf: Results of reliability analysis presented in a bar plot.</p> <p>- 1_Synema_Reliability_supplementary_table.txt: Results of reliability analysis presented in a table.<br> &nbsp;&nbsp; &nbsp;<br> - 1_Synema_Shape_PCA_and_PCA_Ratio_Spectrum.R: R-script for calculating the shape PCA and the PCA Ratio Spectrum of the first shape PC. You may get the necessary MRA source script from http://doi.org/10.5281/zenodo.4250142<br> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<br> - Synema_globosum_AR9379_PV.jpg, Synema_globosum_AR9379_PV.jpg, Synema_globosum_AR9379_PV.jpg, etc.: Photographs taken with a LEICA M205 C stere-omicroscope.</p> <p>&nbsp;&nbsp;&nbsp; 1. Numbers after AR_ refer to the inventory number of the specimens in the Natural History Musuem Bern (NMBE). The specimen number was also used in the data file.<br> &nbsp;&nbsp;&nbsp; 2. The photo named &quot;Synema_globosum_AR9163_with_measurements&quot; shows the position of the measurements. Otherwise, the measurements are not indicated in the raw photos.</p> <p><br> Example image&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Character name&nbsp;&nbsp; &nbsp;Definition<br> Synema_globosum_AR9163_with_measurements&nbsp;&nbsp; &nbsp;cym.l&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Cymbium lenght&nbsp;&nbsp; &nbsp;Distance of the anterior margin to the tip of the cymbium<br> Synema_globosum_AR9163_with_measurements&nbsp;&nbsp; &nbsp;cym.b&nbsp;&nbsp; &nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp; Cymbium breadth&nbsp;&nbsp; &nbsp;widest breadth of the cymbium<br> Synema_globosum_AR9163_with_measurements&nbsp;&nbsp; &nbsp;bul.b&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Bulb breadth&nbsp;&nbsp; &nbsp;widest breadth of the genital bulbus<br> Synema_globosum_AR9163_with_measurements&nbsp;&nbsp; &nbsp;tib.b&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Tibia breadth&nbsp;&nbsp; &nbsp;breadth of the tibia base at the patella joint</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

The predator problem and PCR primers in molecular dietary analysis: swamped or silenced; depth or breadth? - Dataset

<p>Raw sequencing data and other metadata files are associated with Cuff et al. (2022), available at&nbsp;https://doi.org/10.5281/zenodo.4708418</p> <p>The associated code, files and description pertain&nbsp;to the non-metric multi-dimensional scaling plot presented in this review (Figure 4). The code and data required for the boxplot (Figure 3) are given at the Zenodo link above (for Cuff et al. 2022).</p> <p>Data were collected and processed according to Cuff&nbsp;et al., (2022) up to the point of aggregating the two primer pair datasets. Binary matrices for prey detections were combined for the two primer pairs, but each sample represented separately for each primer pair (i.e., not aggregated by sample). Instances where taxa were only identified to genus (or lower, e.g., family) level by only one of the primer pairs resulted in aggregation for the other primer pair at that taxonomic level, except for species within those groups that were reliably identified to species level by both primers. Samples for which only one primer pair generated prey data were removed. The non-metric multidimensional scaling spider plot was created using &lsquo;metaMDS&rsquo; with a Jaccard distance matrix and 999 tries in the &lsquo;vegan&rsquo; package (Oksanen et al., 2016). Outliers that obscured the overall patterns were removed, the final plot having a stress of 0.061. Point colours were assigned using the &lsquo;set1&rsquo; palette of the &lsquo;RColorBrewer&rsquo; package (Neuwirth, 2014) and the final plot created using &lsquo;ggplot2&rsquo; (Wickham, 2016).</p>

opencc-by-4.0Apr 2022View details →
zenodo48/100

WS22 database: combining Wigner Sampling and geometry interpolation towards configurationally diverse molecular datasets

<p>The WS22&nbsp;database provides a collection of molecular datasets that explores a broad configurational space of flexible organic molecules with varying sizes and complexity.&nbsp;It includes several chemical properties calculated with a quantum chemical (QM) method. Complementary to the structured datasets, this repository also provides the&nbsp;molecular geometries for the equilibrium structures together with the corresponding output of the QM frequency calculations.&nbsp;Details about the methodology, content, and structure of the WS22 datasets are provided in the README file included in this repository.</p>

opencc-by-4.0Aug 2022View details →
zenodo48/100

Atomistic trajectories from ab-initio molecular dynamics simulations of wetted TiO2 nanoparticle

<p>This repository&nbsp;contains atomistic trajectories from ab-initio molecular dynamics simulations of water and TiO2 nanoparticle described in the paper:</p> <p>E. G. Brandt, L. Agosta and A.P.Lyubartsev, &quot; Reactive wetting properties&nbsp; of TiO2 nanoparticles predicted by ab initio molecular dynamics simulations&quot;, Nanoscale, 8, 13385-13398 (2016) DOI: 10.1039/c6nr02791a</p> <p>The trajectories are saved in the .xtc format, and initial structures with specification of atom types are given in the .pdb format.</p> <p>The name of each file contains brief information about the simulated system:</p> <p>TiO2 : composition of the nanoparticle<br> n24 &nbsp;: number of TiO2 units in the nanoparticle<br> anatase/brookite/rutile : type of crystall structure<br> - a number 0 - 30 : number of water molecules in the simulation<br> 2fs - the time step</p> <p>For more details, see the referred paper</p>

opencc-by-4.0Jan 2018View details →
zenodo48/100

Dataset for "Large Language Models as molecular design engines"

<ol> <li><strong>claude-gpt-paper.zip :</strong><br><br>This dataset contains data and results associated with the paper "Large Language Models as molecular design<br>engines" The paper investigates the use of large language models, specifically Claude 3 Opus, for generating and analyzing chemical structures based on various prompts from A-H (as mentioned in the manuscript), and guided design related to electron-withdrawing groups (EWG), electron-donating groups (EDG).</li> </ol> <p>The dataset includes:</p> <ol> <li>PM7 MOPAC energy calculations for generated molecules, along with their SMILES representations and molecule IDs.</li> <li>PM7-calculated charges for the generated molecules.</li> <li>Output files from the Claude 3 Opus language model for each prompt category along.</li> <li>Original dataset (subset of ZINC database) used to build common keys and the initial design space.</li> <li>JSON file containing common keys for featurizing unknown SMILES.</li> <li>PCA object to convert molecule embeddings to 3-dimensional embeddings.</li> </ol> <p>The data is organized into the following folders:</p> <ul> <li><code>pm7_charge_results</code>: Contains HOMO-LUMO energy differences for plotting.</li> <li><code>pm7_charge_calculation</code>: Contains PM7 MOPAC energy calculations and charges.</li> <li><code>out</code>: Contains output files from the Claude 3 Opus language model.</li> <li><code>fact-dropbox</code>: Contains the original dataset, common keys, and PCA object file.</li> </ul> <p>The data can be used to reproduce the results presented in the paper and serve as a foundation for further research in this area.</p> <p>For a detailed description of the folder structure and contents, please refer to the File_descriptions.md file included in the dataset.<br><br><br>2. llm-visulizer-dashapp.zip<br><br>This is the code for the visualizer app for viewing the molecules generated by the LLM. The README.md file has details about running the app.</p> <p>3. claude-gpt-paper-codes.zip&nbsp;</p> <p>This contains the notebook GPT_modification_just_plots.ipynb for plotting, and other codes. The README.md file has details about running the main notebook for getting the plots.</p>

opencc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record