Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

351

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

351 results for “jetting”

Learn how ShareScore rates datasets ↗
zenodo44/100

OmniFold Weights | CMS 2011A Open Data | Jet Primary Dataset | pT 375-700 GeV

<p>Unfolding weights corresponding to a selection of jets from the <a href="https://doi.org/10.5281/zenodo.3340205">Jet Primary Dataset of the CMS 2011A Open Data in MOD HDF5 format</a>&nbsp;and associated simulated datasets. The unfolding is performed in a high-dimensional manner&nbsp;using the <a href="https://arxiv.org/abs/1911.09107">OmniFold</a> method, which can unfold all observables simultaneously.&nbsp;<a href="https://arxiv.org/abs/1810.05165">Particle Flow Networks</a> are used in Step 1 and Step 2 of the OmniFold method&nbsp;to process the full phase space information. The datasets and neural networks&nbsp;were accessed/built via the <a href="https://energyflow.network/">EnergyFlow Python package</a>. An upcoming version of the package will contain an example/demo demonstrating how to use these weights.</p> <p>The phase space selections for the data, sim, and gen datasets (using the terminology of the OmniFold paper) are:</p> <ul> <li>data:&nbsp;<span class="math-tex">\(p_T^{\rm jet}\in [375, 700]\)</span>&nbsp;GeV,&nbsp;<span class="math-tex">\(|\eta^{\rm jet}|&lt;2.4\)</span>, jet quality&nbsp;<span class="math-tex">\(\ge\)</span>&nbsp;2</li> <li>sim:&nbsp;<span class="math-tex">\(p_{T,\text{corr}}^{\rm jet} \in [375, 700]\)</span>&nbsp;GeV,&nbsp;<span class="math-tex">\(|\eta^{\rm jet}| &lt; 2.4\)</span>, gen jet matched (&#39;gen_jet_pts != -1&#39; in EnergyFlow), jets from the <a href="https://doi.org/10.5281/zenodo.3341500">170</a> and <a href="https://doi.org/10.5281/zenodo.3341772">1800</a> MC datasets are excluded</li> <li>gen: Matched to sim jet</li> </ul> <p>The omnifold_weights.npz file contains two arrays, &#39;wssim&#39; corresopnding to the Step 1 weights&nbsp;<span class="math-tex">\(\omega_n\)</span>, and &#39;wsgen&#39; corresponding to the Step 2 weights&nbsp;<span class="math-tex">\(\nu_n\)</span>,&nbsp;for iteration&nbsp;<span class="math-tex">\(n\)</span>. The shape of each of these arrays is (6, 16489054), with the first axis being the iteration axis and the second axis being the event axis.&nbsp;There are 5 iterations, but 6 sets of weights in each array, with the 0th entry being the starting weights.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Magnetosheath Jets MMS1 (5/2015 - 6/2020) - Classified

<p><strong>README</strong></p> <p>This dataset contains the time and the class of magnetosheath jets observed by MMS 1 during 05/2015 &ndash; 06/2020.&nbsp;</p> <p>More information about the different classes can be found in the articles:</p> <ol> <li>https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2019JA027754</li> <li>https://www.frontiersin.org/articles/10.3389/fspas.2020.00024/full</li> </ol> <p>While an extension of the classification process has been published in:</p> <ul> <li>https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2021JA029269</li> </ul> <p><strong>INFO</strong>&nbsp;</p> <p>For the attached text file. The first column is the time of maximum dynamic pressure in UTC. The second column is the class of the jet:</p> <p><em>Main Categories:</em><br>11 =&nbsp;Quasi-parallel jet<br>22 =&nbsp;Quasi-perpendicular jet &nbsp;<br>3 =&nbsp;Boundary jet<br>5 =&nbsp;Encapsulated jet&nbsp;</p> <p><em>Secondary Categories:</em><br>1 =&nbsp;Possibly Quasi-parallel jet<br>2 =&nbsp;Possibly Quasi-parallel jet<br>4 =&nbsp;&nbsp;Possibly Boundary jet<br>6 =&nbsp;Possibly Encapsulated jet<br>7 =&nbsp;Close to Magnetopause or Bow Shock jet<br>0 =&nbsp;Unclassified jet<br>8 =&nbsp;Data gap jet</p> <p>Burst availability is given in the third column as:</p> <p>2 =&nbsp;Full burst availability<br>1 = Partial burst availability<br>0 = No burst availability</p> <p>For more information, please contact the author (savvasraptis@gmail.com)&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Liquid-Jet Photoemission Spectroscopy as a Structural Tool: Site-Specific Acid-Base Chemistry of Vitamin C

<p>Dataset pertaining to the article "Liquid-Jet Photoemission Spectroscopy as a Structural Tool: Site-Specific Acid-Base Chemistry of Vitamin C", submitted to Physical Chemistry Chemical Physics. Here, we demonstrate how liquid-jet photoemission spectroscopy can be systematically used for chemical analysis, probing tautomeric forms and deprotonation sites in aqueous vitamin C. We also present a fast and reliable computational protocol to model the spectra.</p> <p>&nbsp;</p> <p>Files with extension .h5 are hdf5-files structured according to the NeXus standard v2022.07, see<br>https://www.nexusformat.org/<br>https://fairmat-experimental.github.io/nexus-fairmat-proposal/50433d9039b3f33299bab338998acb5335cd8951/mpes-structure.html<br>NeXus data files can be opened with any software capable of opening hdf5-structured files. The following viewers are adapted to the specifics of the NeXus data format:<br>* nexpy (distributed with python)<br>* https://h5web.panosc.eu/h5wasm (web-based NeXus viewer maintained by the European Photon and Neutron Open Science Cloud-consortium)</p> <p><br>The following files are provided:</p> <p><strong>Photoemission data pertaining to vitamin C PES measurements:</strong></p> <p>vitamin-C.h5</p> <p>&nbsp;</p> <p><strong>Computational data for Figures 3, 5, and 6: binding energy values (plain text) from which the spectra were modeled, optimized molecular geometries (XYZ coordinates, distances in angstroms) used in the calculations, and a sample input for the binding energy calculation. All data for each figure are packed in a zip file.</strong></p> <p>COMPUTATIONAL-DATA-Fig3.zip<br>COMPUTATIONAL-DATA-Fig5.zip<br>COMPUTATIONAL-DATA-Fig6.zip<br><br><br></p> <p><strong>Numeric representations of the traces shown in the article's figures (space-separated or comma-separated ascii-files):</strong></p> <p><strong>Figure 3:</strong><br>EXPT.dat<br>TAUTOMER-A.dat<br>TAUTOMER-B.dat<br>C2.dat<br>C3.dat<br>C5.dat<br>C6.dat<br>C2+C3.dat<br>C2+C5.dat<br>C2+C6.dat<br>C3+C5.dat<br>C3+C6.dat<br>C5+C6.dat</p> <p><strong>Figure 5:</strong><br>EXPT.dat<br>TAUTOMER-A.dat<br>C3.dat<br>C2+C3.dat</p> <p><strong>Figure 6:</strong><br>EXPT.dat<br>C3.dat<br>C3-UFF-ensemble.dat<br>C3-Bondi-single.dat<br>C3-Bondi-ensemble.dat<br>TAUTOMER-A.dat<br>TAUTOMER-A-ensemble.dat<br>C2+C3.dat<br>C2+C3-ensemble.dat</p> <p><strong>Figure S1:</strong><br>EXPT.dat<br>EXPT-850eV.dat</p> <p>&nbsp;</p> <p>Version history:<br>4: Data from version 2 reuploaded<br>3: Experimental NeXus-data added<br>2: Computational data and figure traces added<br>1: Initial upload</p> <p>Contact person for questions regarding this data set: Lukas Tomanik, tomanikl@vscht.cz . If you use these data for your scientific work we are curious to learn about it.</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

IML workshop challenge on jet mass regression

<p>This dataset is associated with the LPCC IML (Lhc Physics Center at Cern Inter-experimental Machine Learning) working group.&nbsp; It was produced for the second IML annual workshop (April 2018).</p> <p>This dataset is part of a machine learning &quot;challenge&quot; on jet mass regression at future circular collider (FCC) conditions.&nbsp; Further details can be found on the challenge page, here:</p> <p><a href="https://gitlab.cern.ch/IML-WG/IML_challenge_2018/wikis/home">https://gitlab.cern.ch/IML-WG/IML_challenge_2018/wikis/home</a></p>

opencc-zeroMar 2018View details →
zenodo44/100

Z + up to 9 jets parton level events at 14 TeV in HDF5

<p>Z + up to 9 jets at 14 TeV parton level events in HDF5 file format.</p> <p>Merging scale is at 20 GeV.</p> <p>This data was used in FERMILAB-PUB-19-192-T</p>

opencc-by-4.0May 2019View details →
zenodo44/100

Wplus + up to 9 jets parton level events at 14 TeV in HDF5

<p>Wplus + up to 9 jets partonic events at 14TeV</p> <p>Merging scale is at 20 GeV.</p> <p>This data was used in FERMILAB-PUB-19-192-T</p>

opencc-by-4.0May 2019View details →
zenodo44/100

Data files for: Gigantic jet discharges evolve stepwise through the middle atmosphere

<p>Original video files and some other data belonging to the article &quot;Gigantic jet discharges evolve stepwise through the middle atmosphere&quot; published in Nature Communications on September 25th, 2019 (https://doi.org/10.1038/s41467-019-12261-y)</p> <p>The high-speed video .cine files can be read by (free) CineViewer and PCC software of Vision Research Inc. which can convert to avi files. For any questions, contact the first author.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2019View details →
zenodo44/100

Profiling results for electron-positron plus jets production with Pepper

<h2>Profiling results for electron positron plus jets production with Pepper v1.1.1</h2> <h3>Contents</h3> <p>This includes data/output for the following GPU accelerated unweighted event generation runs:</p> <ul> <li>A run over a few seconds to produce pepper-internal timing results (`*timers.csv` files)</li> <li>The same run, but with `nvprof` (`*.nvprof` files)</li> <li>A run with a single event batch, with `ncu --import-source on --set full` (`*.ncu-rep` files)</li> </ul> <p>The simulated process is electron-positron pair production with n jets, with n = 0...4, at 13 TeV. The configuration files (`pepper.ini`) are included. Also the `pepper_cache` directories are included.</p> <p>All results are for the "main" pepper variant (which uses Kokkos to utilise the GPU) and for the "native" pepper variant (which uses CUDA directly).</p> <h3>Software stack and CPU/GPU hardware details</h3> <p>The software/hardware used for the profiling are as follows:</p> <ul> <li>GPU Driver: NVIDIA-SMI: 550.100, Driver Version: 550.100, CUDA Version: 12.4</li> <li>GPU: Tesla V100S-PCIE-32GB</li> <li>Cuda compilation tools: release 11.6, V11.6.124</li> <li>Kokkos 4.3.01</li> <li>gcc (GCC) 11.4.1 20231218 (Red Hat 11.4.1-3)</li> <li>Intel(R) Xeon(R) Silver 4214R CPU @ 2.40GHz</li> </ul> <p>Note that the event generation includes write-out of event files. The LHEH5 files are written to local SSD storage. The event files are not included in this dataset.</p> <h3>Additional run parameters</h3> <table> <tbody> <tr> <td>e^+ e^-</td> <td>number of batches</td> <td>batch size</td> <td>rate (main variant)</td> <td>rate (native variant)</td> </tr> <tr> <td>+0j</td> <td>20</td> <td>1,179,648</td> <td>5.7e9</td> <td>7.4e9</td> </tr> <tr> <td>+1j</td> <td>20</td> <td>1,179,648 * 2</td> <td>1.4e10</td> <td>3.3e10</td> </tr> <tr> <td>+2j</td> <td>80</td> <td>1,179,648 / 2</td> <td>1.3e10</td> <td>5.4e10</td> </tr> <tr> <td>+3j</td> <td>20</td> <td>1,179,648</td> <td>7.7e9</td> <td>1.6e10</td> </tr> <tr> <td>+4j</td> <td>20</td> <td>1,179,648</td> <td>2.2e9</td> <td>3.1e9</td> </tr> <tr> <td>+5j</td> <td>20</td> <td>1,179,648</td> <td>3.4e8</td> <td>4.2e8</td> </tr> </tbody> </table> <p>The number of batches has been chosen such that the most important sub processes are likely to be sampled and such that the overall runtime is at least a few seconds. The batch size has been chosen by scanning over factors of two and picking the best-performing batch size with the native variant. The event rate is the number given in the Pepper output. It does not include the closing time of the generated HDF5 file, which can be significant for the two lowest multiplicities. This can be checked by inspecting the corresponding relative and absolute timing results in the `*timing.csv` files included in the dataset.</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Profiling results for top pair plus jets production with Pepper

<h2>Profiling results for top pair plus jets production with Pepper v1.1.1</h2> <h3>Contents</h3> <p>This includes data/output for the following GPU accelerated unweighted event generation runs:</p> <ul> <li>A run over a few seconds to produce pepper-internal timing results (`*timers.csv` files)</li> <li>The same run, but with `nvprof` (`*.nvprof` files)</li> <li>A run with a single event batch, with `ncu --import-source on --set full` (`*.ncu-rep` files)</li> </ul> <p>The simulated process is top pair production with n jets, with n = 0...4, at 13 TeV. The configuration files (`pepper.ini`) are included. Also the `pepper_cache` directories are included.</p> <p>All results are for the "main" pepper variant (which uses Kokkos to utilise the GPU) and for the "native" pepper variant (which uses CUDA directly).</p> <h3>Software stack and CPU/GPU hardware details</h3> <p>The software/hardware used for the profiling are as follows:</p> <ul> <li>GPU Driver: NVIDIA-SMI: 550.100, Driver Version: 550.100, CUDA Version: 12.4</li> <li>GPU: Tesla V100S-PCIE-32GB</li> <li>Cuda compilation tools: release 11.6, V11.6.124</li> <li>Kokkos 4.3.01</li> <li>gcc (GCC) 11.4.1 20231218 (Red Hat 11.4.1-3)</li> <li>Intel(R) Xeon(R) Silver 4214R CPU @ 2.40GHz</li> </ul> <p>Note that the event generation includes write-out of event files. The LHEH5 files are written to local SSD storage. The event files are not included in this dataset.</p> <h3>Additional run parameters</h3> <table> <tbody> <tr> <td>TTBar</td> <td>number of batches</td> <td>batch size</td> <td>rate (main variant)</td> <td>rate (native variant)</td> </tr> <tr> <td>+0j</td> <td>20</td> <td>1,179,648</td> <td>1.4e10</td> <td>2.0e10</td> </tr> <tr> <td>+1j</td> <td>20</td> <td>1,179,648</td> <td>1.5e10</td> <td>2.4e10</td> </tr> <tr> <td>+2j</td> <td>20</td> <td>1,179,648 * 2</td> <td>1.3e10</td> <td>2.7e10</td> </tr> <tr> <td>+3j</td> <td>20</td> <td>1,179,648 / 2</td> <td>6.4e9</td> <td>1.4e10</td> </tr> <tr> <td>+4j</td> <td>20</td> <td>1,179,648</td> <td>2.2e9</td> <td>2.2e9</td> </tr> </tbody> </table> <p>The number of batches has been chosen such that the most important sub processes are samples and such that the overall runtime is at least a few seconds. The batch size has been chosen by scanning over factors of two and picking the best-performing batch size with the native variant. The event rate is the number given in the Pepper output. It does not include the closing time of the generated HDF5 file, which can be significant for the two lowest multiplicities. This can be checked by inspecting the corresponding relative and absolute timing results in the `*timing.csv` files included in the dataset.</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Z+2Jet Unfolding Data Set

<p>Inspired by the recent application of ML-based unfolding to ATLAS data, we will unfold events from pp to Z + 2 jets production from the reco-level to the pre-detector or gen-level. We generate the events with Madgraph 5, shower and hadronization are simulated with Pythia 8.311, and detector effects are included via Delphes 3.5.0 using the default CMS card. Jets are clustered at gen-level and reco-level using an anti-kT algorithm with R=0.4 implemented in FastJet~3.3.4.</p> <p>We apply a set of cuts resembling the ATLAS analysis. For details on the cuts please see the associated paper.&nbsp;<br>All events must pass all cuts on gen and reco level.&nbsp;</p> <p>The training set consists of 1.5M events, the test set of 400k events.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Single-pulse hard x-ray holograms of an exploding water jet

<p>This h5 file contains the holograms of an exploding micro-fluidic jet recorded at the MID setup at EuXFEL. The holograms were recorded with single-pulse illumination of 17.8 keV hard x-rays.</p> <p>The file contains only one&nbsp; group (/frames/pixels) with 4499 frames recorded with the Andor Zyla 5.5 camera used in this experiment.</p> <p>The first 352 frames and frames 4233 to 4499 can be used as empty beam / reference frames. The frames in between are data frames. The microfluidic jet is put in the field of view and is pumped with an IR laser with pumping offset from 5 to -35 ns with respect to the arrival of the XFEL pulse.&nbsp;</p> <p>The data set has been published in:</p> <p>J. Hagemann, M. Vassholz, H. Hoeppe, M. Osterhoff, J. M. Rossell&oacute;, R. Mettin, F. Seiboth, A. Schropp, J. M&ouml;ller, J. Hallmann, C. Kim, M. Scholz, U. Boesenberg, R. Schaffer, A. Zozulya, W. Lu, R. Shayduk, A. Madsen, C. G. Schroer, and T. Salditt, &ldquo;Single-pulse phase-contrast imaging at free-electron lasers in the hard X-ray regime,&rdquo; Journal of Synchrotron Radiation 28(1), 52&ndash;63 (2021).<br> &nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

Durability in materials jetting: A long-term study

<p>This file contains data for the paper &quot;Durability in materials jetting: A long-term study&quot;<br> For more information, please contact Ali Payami Golhin (payami.ag@gmail.com)</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Dataset for Appearance evaluation of digital materials in material jetting

<p>Dataset for Appearance evaluation of digital materials in material jetting: Including dataset for gloss, haze, scattering, specular BRDF, reflectance, and transmittance</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Dataset for the optical properties of tilted surfaces in material jetting

<p>Dataset for the optical properties of tilted surfaces in material jetting:</p> <p>Including dataset for gloss, haze, scattering, specular BRDF, reflectance, transmittance, and statistical analysis</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Top Jet W-Momentum Reconstruction Dataset

<p><strong>Overview</strong></p> <p>A set of Monte Carlo simulated events, for the evaluation of top quarks' (and their&nbsp;child particles') momentum reconstruction, produced using the HEPData4ML package [1]. Specifically, the entries in this dataset correspond with top quark jets, and the momentum of the jets' constituent particles.&nbsp;This is a newer version of the "Top Quark Momentum Reconstruction Dataset" [2], but with sufficiently large changes to warrant this separate posting.</p> <p>The dataset&nbsp;is saved in HDF5 format, as sets of arrays with keys (as detailed below). There are ~1.5M events, approximately broken down into the following sets:</p> <ul> <li>Training: 700k events (files with "_train" suffix)</li> <li>Validation: 200k events (files with "_valid" suffix)</li> <li>Testing (small): 100k events (files with "_test" suffix)</li> <li>Testing (large): 500k events (files with "_test_large" suffix)</li> </ul> <p>The two separate types of testing files -- small and large -- are independent from one another, the former for conveniently running quicker testing and the latter for testing with a larger sample.</p> <p>There are four version of the dataset present, with the versions indicated by the filenames. The different versions correspond with whether or not fast detector simulation was performed (versus truth-level jets), and whether or not the W-boson mass was modified: One version of the dataset uses the nominal value of &nbsp;<span>\(m_W = 80.385 \text{ GeV}\)</span>&nbsp;as used by Pythia8 [3], whereas another uses a variable mW taking on 101 values evenly-spaced as&nbsp;<span>\(m_W \in \{ 64.308,96.462 \} \text{ GeV}\)</span>. The dataset naming scheme is as follows:</p> <ul> <li>train.h5 : jets clustered from truth-level, nominal mW</li> <li>train_mW.h5: jets clustered from truth-level, variable mW</li> <li>train_delphes.h5: jets clustered from Delphes outputs, nominal mW</li> <li>train_delphes_mW.h5: jets clustered from Delphes outputs, variable mW</li> </ul> <p><strong>Description</strong></p> <ul> <li>13 TeV&nbsp;center-of-mass energy, fully hadronic top quark decays, simulated with Pythia8. (<span>\(t \rightarrow W \, b, \; W\rightarrow q \, q'\)</span>) <ul> <li>Events are generated with leading top quark pT in [550,650] GeV. (set via Pythia8's <span>\(\hat{p}_{T,\text{ min}}\)</span>&nbsp;and&nbsp;<span>\(\hat{p}_{T,\text{ max}}\)</span> variables)</li> <li>No inital- or final-state radiation (ISR/FSR), nor multi-parton interactions (MPI)</li> <li>Where applicable,&nbsp;detector simulation is done using DELPHES [4], with the ATLAS detector card.</li> </ul> </li> <li>Clustering of particles/objects is done via FastJet [5], using the anti-k<sub>T</sub> algorithm, with&nbsp;<span>\(R=0.8\)</span>&nbsp;. <ul> <li>For the truth-level data, inputs to jet clustering are truth-level, final-state particles (i.e. clustering "truth jets").</li> <li>For the data with detector simulation, the inputs are calorimeter towers from DELPHES. <ul> <li>`Tower` objects from DELPHES (<em>not</em>&nbsp;E-flow objects, no tracking information)</li> </ul> </li> <li>Each entry in the dataset corresponds with a single top quark jet, extracted from a <span>\(t\bar{t}\)</span>&nbsp;event. <ul> <li>All jets are matched to a parton-level top quark within&nbsp;<span>\(\Delta R &lt; 0.8\)</span>&nbsp;. We choose the jet&nbsp;<em>nearest</em>&nbsp;the parton-level top quark.</li> <li>Jets are required to have <span>\(|\eta| &lt; 2\)</span>, and&nbsp;<span>\(p_{T} &gt; 15 \text{ GeV}\)</span>.</li> <li>The 200 leading (highest-p<sub>T</sub>) jet constituent four-momenta are stored in Cartesian coordinates&nbsp;(E,px,py,pz), sorted by decreasing&nbsp;pT, with zero-padding.</li> <li>The jet four-momentum is stored in Cartesian coordinates (E, p<sub>x</sub>, p<sub>y</sub>, p<sub>z</sub>), as well&nbsp;&nbsp;as in cylindrical coordinates <span>\((p_T,\eta,\phi,m)\)</span>.</li> <li>The truth (parton-level) four-momenta of the top quark, the bottom quark the W-boson, and the quarks to which the W-boson decays, are stored in Cartesian coordinates. <ul> <li>In addition, the momenta of the 120 leading stable daughter particles of the W-boson are stored in Cartesian coordinates.</li> </ul> </li> </ul> </li> </ul> </li> </ul> <p>Description of data fields &amp; metadata<br>Below is a brief description of the various fields in the dataset. The dataset also contains metadata fields, <a href="https://docs.h5py.org/en/stable/high/attr.html">stored using HDF5's "attributes"</a>. This is used for fields that are common across many events, and stores information such as generator-level configurations (in principle, all the information is stored as to be able to recreate the dataset with the HEPData4ML tool).</p> <p>Note that fields whose keys have the prefix "jh_" correspond with output from the Johns Hopkins top tagger [6], as implemented in FastJet.</p> <p>Also note that for the keys corresponding with four-momenta in Cartesian coordinates, there are rotated versions of these fields -- the data has been rotated so that the W-boson is at&nbsp;<span>\((\theta=0, \phi=0)\)</span>, and the b-quark is in the&nbsp;<span>\((\theta=0, \phi &lt; 0)\)</span>&nbsp;plane. This rotation is potentially useful for visualizations of the events.</p> <ul> <li> <p>Nobj: The number of constituents in the jet.</p> </li> <li> <p>Pmu: The four-momenta of the jet constituents, in (E, p<sub>x</sub>, p<sub>y</sub>, p<sub>z</sub>). Sorted by decreasing p<sub>T</sub> and zero-padded to a length of 200.</p> <ul> <li> <p>Pmu_rot: Rotated version.</p> </li> </ul> </li> <li> <p>contained_daughter_sum_Pmu: Four-momentum sum of the stable daughter particles of the W-boson that fall within&nbsp;<span>\(\Delta R &lt; 0.8\)</span>&nbsp;of the jet centroid.</p> <ul> <li> <p>contained_daughter_sum_Pmu_rot: Rotated version.</p> </li> </ul> </li> <li> <p>cross_section: Cross-section for the corresponding process, reported by Pythia8.</p> </li> <li> <p>cross_section_uncertainty:&nbsp;Cross-section uncertainty for the corresponding process, reported by Pythia8.</p> </li> <li> <p>energy_ratio smeared: Ratio of the true energy of W-boson daughter particles contributing to this calorimeter tower, divided by the total smeared energy in this calorimeter tower.</p> <ul> <li> <p>Only relevant for the DELPHES datasets.</p> </li> </ul> </li> <li> <p>energy_ratio_truth: Ratio of the true energy of W-boson daughter particles contributing to this calorimeter tower, divided by the total true energy of particles contributing to this calorimeter tower.</p> <ul> <li> <p>The above definition is relevant only for the DELPHES datasets. For the truth-level datasets, this field is repurposed to store a value (0 or 1) indicating whether or not the given particle (whose momentum is in the `Pmu` field) is a W-boson daughter.</p> </li> </ul> </li> <li> <p>event_idx: Redundant -- used for event indexing during the event generation process.</p> </li> <li> <p>is_signal: Redundant -- indicates whether an event is signal or background, but this is a fully signal dataset. Potentially useful if combining with other datasets produced with HEPData4ML.</p> </li> <li> <p>jet_Pmu: Four-momentum of the jet, in&nbsp;(E, p<sub>x</sub>, p<sub>y</sub>, p<sub>z</sub>).</p> <ul> <li> <p>jet_Pmu_rot: Rotated version.</p> </li> </ul> </li> <li> <p>jet_Pmu_cyl: Four-momentum of the jet, in <span>\((pT_,\eta,\phi,m)\)</span>.</p> </li> <li> <p>jet_bqq_contained_dR06: Boolean flag indicating whether or not the truth-level b and the two quarks from W decay are contained within&nbsp;<span>\(\Delta R &lt; 0.6\)</span>&nbsp;of the jet centroid.</p> </li> <li> <p>jet_bqq_contained_dR08:&nbsp;Boolean flag indicating whether or not the truth-level b and the two quarks from W decay are contained within&nbsp;<span>\(\Delta R &lt; 0.8\)</span>&nbsp;of the jet centroid.</p> </li> <li> <p>jet_bqq_dr_max: Maximum of&nbsp;<span>\(\big\lbrace \Delta R \left( \text{jet},b \right), \; \Delta R \left( \text{jet},q \right), \; \Delta R \left( \text{jet},q' \right) \big\rbrace\)</span>.</p> </li> <li> <p>jet_qq_contained_dR06:&nbsp;Boolean flag indicating whether or not the two quarks from W decay are contained within&nbsp;<span>\(\Delta R &lt; 0.6\)</span>&nbsp;of the jet centroid.</p> </li> <li> <p>jet_qq_contained_dR08:&nbsp;Boolean flag indicating whether or not the two quarks from W decay are contained within&nbsp;<span>\(\Delta R &lt; 0.8\)</span>&nbsp;of the jet centroid.</p> </li> <li> <p>jet_qq_dr_max:&nbsp;Maximum of&nbsp;<span>\(\big\lbrace \Delta R \left( \text{jet},q \right), \; \Delta R \left( \text{jet},q' \right) \big\rbrace\)</span>.</p> </li> <li> <p>jet_top_daughters_contained_dR08: Boolean flag indicating whether the final-state daughters of the top quark are within&nbsp;<span>\(\Delta R &lt; 0.8\)</span>&nbsp;of the jet centroid. Specifically, the algorithm for this flag checks that the jet contains the stable daughters of both the b quark and the W boson. For the b and W each, daughter particles are allowed to be uncontained as long as (for each particle) the <span>\(p_T\)</span>&nbsp;of the sum of uncontained daughters is below&nbsp;<span>\(2.5 \text{ GeV}\)</span>.</p> </li> <li> <p>jh_W_Nobj: Number of constituents in the W-boson candidate identified by the JH tagger.</p> </li> <li> <p>jh_W_Pmu: Four-momentum of the JH tagger W-boson candidate, in&nbsp;(E, p<sub>x</sub>, p<sub>y</sub>, p<sub>z</sub>).</p> <ul> <li> <p>jh_W_Pmu_rot: Rotated version.</p> </li> </ul> </li> <li> <p>jh_W_constituent_Pmu: Four-momentum of the constituents of the JH tagger W-boson candidate, in&nbsp;(E, p<sub>x</sub>, p<sub>y</sub>, p<sub>z</sub>).</p> <ul> <li> <p>jh_W_constituent_Pmu_rot: Rotated version.</p> </li> </ul> </li> <li> <p>jh_m: Mass of the JH W-boson candidate.</p> </li> <li> <p>jh_m_resolution: Ratio of JH W-boson candidate mass, versus the true W-boson mass.</p> </li> <li> <p>jh_pt:&nbsp;<span>\(p_T\)</span>&nbsp;of the JH W-boson candidate.</p> </li> <li> <p>jh_pt_resolution:&nbsp;Ratio of JH W-boson candidate <span>\(p_T\)</span>, versus the true W-boson mass.</p> </li> <li> <p>jh_tag: Whether or not a jet was tagged by the JH tagger.</p> </li> <li> <p>mc_weight: Monte Carlo weight for this event, reported by Pythia8.</p> </li> <li> <p>process_code: Process code reported by Pythia8.</p> </li> <li> <p>rotation_matrix: Rotation matrix for rotating the events' 3-momenta as to produce the rotated copies stored in the dataset.</p> </li> <li> <p>truth_Nobj: Number of truth-level particles (saved in truth_Pmu).</p> </li> <li> <p>truth_Pdg: PDG codes of the truth-level particles.</p> </li> <li> <p>truth_Pmu: Truth-level particles: The top quark, bottom quark, W boson, q, q', and 120 leading, stable W-boson daughter particles, in (E, p<sub>x</sub>, p<sub>y</sub>, p<sub>z</sub>). A few of these are also stored in separate keys:</p> </li> <li> <ul> <li>truth_Pmu_0: Top quark. <ul> <li>truth_Pmu_0_rot: Rotated version.</li> </ul> </li> <li> <p>truth_Pmu_1: Bottom quark.</p> <ul> <li> <p>truth_Pmu_1_rot: Rotated version.</p> </li> </ul> </li> <li> <p>truth_Pmu_2: W-boson.</p> <ul> <li> <p>truth_Pmu_2_rot: Rotated version.</p> </li> </ul> </li> <li> <p>truth_Pmu_3: q from W decay.</p> <ul> <li> <p>truth_Pmu_3_rot: Rotated version.</p> </li> </ul> </li> <li> <p>truth_Pmu_4: q' from W decay.</p> <ul> <li> <p>truth_Pmu_4_rot: Rotated version.</p> </li> </ul> </li> <li> <p>truth_Pmu_0_rot: Rotated version of `truth_Pmu`.</p> </li> </ul> </li> </ul> <p>The following fields correspond with metadata -- they provide the index of the corresponding metadata entry for each event:</p> <ul> <li> <p>command_line_arguments: The command-line arguments passed to HEPData4ML's `run.py` script.</p> </li> <li> <p>config_file: The contents of the Python configuration file used for HEPData4ML. This, together with the command-line arguments, defines how the tool was run, what processes, jet clustering and post-processing was done, etc.</p> </li> </ul> <ul> <li> <p>git_hash: Git hash for HEPData4ML.</p> </li> <li> <p>timestamp: Timestamp for when the dataset was created (local).</p> </li> <li> <p>timestamp_string_utc: Timestamp for when the dataset was created (in UTC).</p> </li> <li> <p>pythia_config: Configuration file passed to Pythia8, by HEPData4ML. Defines the process.</p> </li> <li> <p>pythia_random_seed: Random seed passed to Pythia8, for initializing its random number generator.</p> </li> <li> <p>unique_id: A unique string for identifying the generated set of events (one run of the HEPData4ML tool).</p> </li> <li> <p>unique_id_short: Similar to `unique_id`, but shortened.</p> </li> </ul> <p><strong>Citations</strong></p> <p><strong>[1]</strong>: J. T. Offermann, X. Liu, and T. Hoffman,&nbsp;<a href="https://github.com/janTOffermann/HEPData4ML">HEPData4ML</a> (2023).</p> <p><strong>[2]</strong>: J. T. Offermann, A. Bogatskiy, and T. Hoffman, <a href="https://doi.org/10.5281/zenodo.7338117">Top Quark Momentum Reconstruction Dataset</a>&nbsp;(2022).</p> <p><strong>[3]</strong>:&nbsp;C. Bierlich and others,&nbsp;<a href="https://arxiv.org/abs/2203.11601"><em>A Comprehensive Guide to the Physics and Usage of PYTHIA 8.3</em></a>, (2022).</p> <p><strong>[4]</strong>:&nbsp;J. de Favereau, C. Delaere, P. Demin, A. Giammanco, V. Lema&icirc;tre, A. Mertens, and M. Selvaggi,&nbsp;<a href="https://arxiv.org/abs/1307.6346"><em>DELPHES 3, A Modular Framework for Fast Simulation of a Generic Collider Experiment</em></a>, JHEP&nbsp;<strong>02</strong>, 057 (2014).</p> <p><strong>[5]</strong>:&nbsp;M. Cacciari, G. P. Salam, and G. Soyez,&nbsp;<a href="https://arxiv.org/abs/1111.6097"><em>FastJet User Manual</em></a>, Eur. Phys. J. C&nbsp;<strong>72</strong>, 1896 (2012).</p> <p><strong>[6]</strong>:&nbsp;D. E. Kaplan, K. Rehermann, M. D. Schwartz, and B. Tweedie,&nbsp;<a href="https://arxiv.org/abs/0806.0848"><em>Top Tagging: A Method for Identifying Boosted Hadronically Decaying Top Quarks</em></a>, Phys. Rev. Lett.&nbsp;<strong>101</strong>, 142001 (2008).</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Herwig7.1 Quark and Gluon Jets

<p>Two datasets of quark and gluon jets generated with Herwig 7.1.4, one with all kinematically realizable quark jets and one that excludes charm and bottom quark jets (at the level of the hard process), analogous to <a href="https://zenodo.org/record/2658763">this dataset of Pythia jets</a>. Note that the two datasets in this record should not be combined. Generation parameters are listed below:</p> <ul> <li>Herwig 7.1.4,&nbsp;<span class="math-tex">\(\sqrt{s}=14\,\text{TeV} \)</span></li> <li>Quarks from&nbsp;<span class="math-tex">\(gq\to Z(\to\nu\bar\nu)q\)</span>, gluons from&nbsp;<span class="math-tex">\(q\bar q\to Z(\to\nu\bar\nu)g\)</span></li> <li>FastJet 3.3.0, anti-kT&nbsp;jets with R=0.4</li> <li><span class="math-tex">\(p_T^\text{jet}\in[500,550]\,\text{GeV},\,|y^\text{jet} |&lt;1.7\)</span></li> </ul> <p>There are 20 files in each dataset, each in compressed NumPy format. Files including charm and bottom jets have &#39;withbc&#39; in their filename. There are two arrays in each file</p> <ul> <li>X: (100000,M,4), exactly 50k quark and 50k gluon jets, randomly sorted, where M is the max multiplicity of the jets in that file (other jets have been padded with zero-particles), and the features of each particle are its pt, rapidity, azimuthal angle, and pdgid.</li> <li>y: (100000,), an array of labels for the jets where gluon is 0 and quark is 1.</li> </ul> <p>If you use this dataset, please cite this Zenodo record and, optionally, the <a href="https://zenodo.org/record/2658763">Pythia dataset</a> which inspired it. The datasets can be downloaded and read into python automatically using the&nbsp;<a href="https://energyflow.network/docs/datasets/#quark-and-gluon-jets">EnergyFlow Python package</a>.</p> <p>Changes:</p> <ul> <li>v1 - Renamed files from v0 to include &#39;withbc&#39; (events were also shuffled around), added files without b and c quark jets.</li> </ul>

opencc-by-4.0May 2019View details →
zenodo40/100

HLS4ML LHC Jet dataset (30 particles)

<p>Dataset of high-pT jets from simulations of LHC proton-proton collisions</p> <p>Prepared for FastML/HLS4ML studies:&nbsp;https://fastmachinelearning.org</p> <p>Includes: High level features (see&nbsp;https://arxiv.org/abs/1804.06913)</p> <p>Images: jet images with up to 30 particles/jet (see&nbsp;https://arxiv.org/abs/1908.05318)</p> <p>List: list of jet features with&nbsp;up to 30 particles/jet (see&nbsp;https://arxiv.org/abs/1908.05318)</p>

opencc-by-4.0Jan 2020View details →
zenodo40/100

HLS4ML LHC Jet dataset (50 particles)

<p>Dataset of high-pT jets from simulations of LHC proton-proton collisions</p> <p>Prepared for FastML/HLS4ML studies:&nbsp;https://fastmachinelearning.org</p> <p>Includes: High level features (see&nbsp;https://arxiv.org/abs/1804.06913)</p> <p>Images: jet images with up to 50 particles/jet (see&nbsp;https://arxiv.org/abs/1908.05318)</p> <p>List: list of jet features with&nbsp;up to 50 particles/jet (see&nbsp;https://arxiv.org/abs/1908.05318)</p>

opencc-by-4.0Jan 2020View details →
zenodo40/100

HLS4ML LHC Jet dataset (150 particles)

<p>Dataset of high-pT jets from simulations of LHC proton-proton collisions</p> <p>Prepared for FastML/HLS4ML studies:&nbsp;https://fastmachinelearning.org</p> <p>Includes: High level features (see&nbsp;https://arxiv.org/abs/1804.06913)</p> <p>Images: jet images with up to 150 particles/jet (see&nbsp;https://arxiv.org/abs/1908.05318)</p> <p>List: list of jet features with&nbsp;up to 150 particles/jet (see&nbsp;https://arxiv.org/abs/1908.05318)</p>

opencc-by-4.0Jan 2020View details →
zenodo40/100

HLS4ML LHC Jet dataset (100 particles)

<p>Dataset of high-pT jets from simulations of LHC proton-proton collisions</p> <p>Prepared for FastML/HLS4ML studies:&nbsp;https://fastmachinelearning.org</p> <p>Includes: High level features (see&nbsp;https://arxiv.org/abs/1804.06913)</p> <p>Images: jet images with up to 100 particles/jet (see&nbsp;https://arxiv.org/abs/1908.05318)</p> <p>List: list of jet features with&nbsp;up to 100 particles/jet (see&nbsp;https://arxiv.org/abs/1908.05318)</p>

opencc-by-4.0Jan 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record