Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
15
datasets available to search
ShareScore release 0.9.0
Dataset results
15 results for “Lepton”
Unsupervised New Physics detection at 40 MHz: A -> 4 leptons Signal Benchmark Dataset
<p>Unsupervised New Physics detection at 40 MHz data challenge</p> <p>Signal Benchmark Dataset consisting of A -> 4 leptons decays produced in collision events (simulation of LHC 13 TeV proton-proton collisions) pre-filtered by a requirement of a muon or electron with 23 GeV transverse momentum. Data format description available on the data challenge web page: https://mpp-hep.github.io/ADC2021/</p>
Fuτure - dataset for studies, development, and training of algorithms for reconstructing and identifying hadronically decaying tau leptons
<h1> Data description</h1> <h2>MC Simulation</h2> <p><br>The <strong>Fuτure</strong> dataset is intended for studies, development, and training of algorithms for reconstructing and identifying hadronically decaying tau leptons. The dataset is generated with Pythia 8, with the full detector simulation being performed by Geant4 with the CLIC-like detector setup CLICdet (CLIC_o3_v14) setup. Events are reconstructed using the Marlin reconstruction framework and interfaced with Key4HEP. Particle candidates in the reconstructed events are reconstructed using the PandoraPF algorithm.</p> <p>In this version of the dataset no γγ -> hadrons background is included.</p> <h2>Samples</h2> <p><br>This dataset contains e+e- samples with Z->ττ, ZH,H->ττ and Z->qq events, with approximately 2 million events simulated in each category.</p> <p>The following processes e+e- were simulated with Pythia 8 at sqrt(s) = 380 GeV:</p> <ul> <li>p8_ee_qq_ecm380 [Z -> qq events]</li> <li>p8_ee_ZH_Htautau [ZH -> Ztautau]</li> <li>p8_ee_Z_Ztautau_ecm380 [ZH -> Ztautau]</li> </ul> <p>The .root files from the MC simulation chain are eventually processed by the software found in <a href="https://github.com/HEP-KBFI/ml-tau-en-reg">Github</a> in order to create flat ntuples as the final product.</p> <h2><br>Features</h2> <p><br>The basis of the ntuples are the particle flow (PF) candidates from PandoraPF. Each PF candidate has four momenta, charge and particle label (electron / muon / photon / charged hadron / neutral hadron). The PF candidates in a given event are clustered into jets using generalized kt algorithm for ee collisions, with parameters p=-1 and R=0.4. The minimum pT is set to be 0 GeV for both generator level jets and reconstructed jets. The dataset contains the four momenta of the jets, with the PF candidates in the jets with the above listed properties.</p> <p>Additionally, a set of variables describing the tau lifetime are calculated using the software in <a href="https://github.com/HEP-KBFI/ml-tau-en-reg">Github</a>. As tau lifetime is very short, these variables are sensitive to true tau decays. In the calculation of these lifetime variables, we use a linear approximation.</p> <p>In summary, the features found in the flat ntuples are:</p> <p> </p> <table> <tbody> <tr> <td><strong>Name</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>reco_cand_p4s</td> <td>4-momenta per particle in the reco jet.</td> </tr> <tr> <td>reco_cand_charge</td> <td>Charge per particle in the jet.</td> </tr> <tr> <td>reco_cand_pdg</td> <td>PDGid per particle in the jet.</td> </tr> <tr> <td>reco_jet_p4s</td> <td>RecoJet 4-momenta.</td> </tr> <tr> <td>reco_cand_dz</td> <td>Longitudinal impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>reco_cand_dz_err</td> <td>Uncertainty of the longitudinal impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>reco_cand_dxy</td> <td>Transverse impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>reco_cand_dxy_err</td> <td>Uncertainty of the transverse impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>gen_jet_p4s</td> <td>GenJet 4-momenta. Matched with RecoJet within a cone of radius dR < 0.3.</td> </tr> <tr> <td>gen_jet_tau_decaymode</td> <td>Decay mode of the associated genTau. Jets that have associated leptonically decaying taus are removed, so there are no DM=16 jets. If no GenTau can be matched to GenJet within dR < 0.4, a fill value is used.</td> </tr> <tr> <td>gen_jet_tau_p4s</td> <td>Visible 4-momenta of the genTau. If no GenTau can be matched to GenJet within dR<0.4, a fill value is used.</td> </tr> </tbody> </table> <p>The ground truth is based on stable particles at the generator level, before detector simulation. These particles are clustered into generator-level jets and are matched to generator-level τ leptons as well as reconstructed jets. In order for a generator-level jet to be matched to generator-level τ lepton, the τ lepton needs to be inside a cone of dR = 0.4. The same applies for the reconstructed jet, with the requirement on dR being set to dR = 0.3. For each reconstructed jet, we define three target values related to τ lepton reconstruction:</p> <ul> <li> a binary flag <strong>isTau</strong> if it was matched to a generator-level hadronically decaying τ lepton. <strong>gen_jet_tau_decaymode</strong> of value -1 indicates no match to generator-level hadronically decaying τ.</li> <li> the categorical decay mode of the τ <strong>gen_jet_tau_decaymode</strong> in terms of the number of generator level charged and neutral hadrons. Possible <strong>gen_jet_tau_decaymode</strong> are {0, 1, . . . , 15}.</li> <li> if matched, the visible (neglecting neutrinos), reconstructable pT of the τ lepton. This is inferred from the <strong>gen_jet_tau_p4s</strong></li> </ul> <h2>Contents:</h2> <ul> <li>qq_test.parquet</li> <li>qq_train.parquet</li> <li>zh_test.parquet</li> <li>zh_train.parquet</li> <li>z_test.parquet</li> <li> z_train.parquet</li> <li>data_intro.ipynb</li> </ul> <h2>Dataset characteristics</h2> <p> </p> <table> <tbody> <tr> <td><strong>File</strong></td> <td><strong># Jets</strong></td> <td><strong>Size</strong></td> </tr> <tr> <td>z_test.parquet</td> <td> <pre>870 843</pre> </td> <td>171 MB</td> </tr> <tr> <td>z_train.parquet</td> <td> <pre>3 483 369</pre> </td> <td>681 MB</td> </tr> <tr> <td>zh_test.parquet</td> <td> <pre>1 068 606</pre> </td> <td>213 MB</td> </tr> <tr> <td>zh_train.parquet</td> <td> <pre>4 274 423</pre> </td> <td>851 MB</td> </tr> <tr> <td>qq_test.parquet</td> <td> <pre>6 366 715</pre> </td> <td>1.4 GB</td> </tr> <tr> <td>qq_train.parquet</td> <td> <pre>25 466 858</pre> </td> <td>5.6 GB</td> </tr> </tbody> </table> <p>The dataset consists of 6 files of 8.9 GB in total.</p> <h2>How can you use these data?</h2> <p>The .parquet files can be directly loaded with the Awkward Array Python library.<br>An example how one might use the dataset and the features is given in <strong>data_intro.ipynb</strong></p>
Ancillary files for "Reinterpreting the ATLAS bounds on heavy neutral leptons in a realistic neutrino oscillation model [arXiv: 2107.12980]"
<p><em>(Description copied from Appendix A "Ancillary files" of the companion paper)</em></p> <p>In order to simplify the interpretation of experimental results within realistic HNL models, we are including a number of data files along with the present publication. They can be used to generate the relevant signal samples, or to implement the extrapolation method presented in section 3.2.</p> <p><strong>Card files for the Monte-Carlo event generation</strong></p> <p>The /attachments/card_files folder contains the MadGraph card files (ending in .dat) and scripts (ending in .txt) for generating the signal samples used in this analysis, as well as for computing the total HNL width. Due to the OSSF veto, only processes with no opposite-charge same-flavor lepton pairs have been included. Additional relevant processes can easily be added by modifying the <em>generate</em> and <em>add process</em> lines in the *.txt files. All samples (except the ones used to compute the HNL width, which are generated at parton level) are generated at leading order, include up to two hard jets, and are showered and hadronized using Pythia 8. This is essential for obtaining a realistic W spectrum. The shower parameters could probably benefit from further tuning, and further improvements in the W spectrum accuracy are expected at NLO (using a suitable model). To allow computing the signal efficiencies, all cuts have been disabled in the run card (with the exception of the maximum <span class="math-tex">\(|\eta_{\mathrm{jet}}|\)</span> which needs to be set to 5 for correct matching).</p> <p><strong>Signal cross sections</strong></p> <p>The cross sections for the various processes considered in this analysis, as well as the total HNL width (both computed using MadGraph as described in section 3.2), are provided as JSON files in the /attachments/cross_sections folder.</p> <p>The file total_hnl_width.json contains the total HNL width <span class="math-tex">\(\hat{\Gamma}_{\alpha}(M_N)\)</span> (expressed in GeV), computed for the 5 mass points used in this analysis, and under the assumption of unit mixing with a single flavor <span class="math-tex">\(\alpha\)</span>, for each flavor. The total HNL width can then be computed for any combinations of mixing angles using eq. (3.2). The file is organized as two nested dictionaries, with the first key denoting the HNL mass <span class="math-tex">\(M_N\)</span>, and the second one the flavor <span class="math-tex">\(\alpha\)</span> for which the total width <span class="math-tex">\(\hat{\Gamma}_{\alpha}(M_N)\)</span> has been computed for a unit mixing angle <span class="math-tex">\(|\Theta_{\alpha}|^2 = 1\)</span> (with <em>Wtot_e</em> for <span class="math-tex">\(\alpha=e\)</span>, <em>Wtot_mu</em> for <span class="math-tex">\(\mu\)</span> and <em>Wtot_tau</em> for <span class="math-tex">\(\tau\)</span>).</p> <p>The file cross_sections.json contains the reference cross sections <span class="math-tex">\(\sigma_P^{\mathrm{ref}}\)</span> (in pb) for all the processes <em>P</em> considered in this analysis, expressed for <span class="math-tex">\(|\Theta|_{\mathrm{ref}}^2 = 1\)</span> and <span class="math-tex">\(\Gamma_{\mathrm{ref}} = 10^{-5}\,\mathrm{GeV}\)</span>. The file is organized as two nested dictionaries, with the first key denoting the HNL mass <span class="math-tex">\(M_N \)</span> and the second the process <em>P</em>. The correspondence between the key and the physical process can be found in table 7.</p> <p><strong>Signal efficiencies</strong></p> <p>The efficiencies resulting from the event selection described in section 3.1, as well as their parametrization according to eq. (3.6) (as discussed in section 3.3) can respectively be found in the files efficiencies.json and fitted_efficiencies.json in the /attachments/efficiencies folder.</p> <p>The file efficiencies.json is organized as follows. The data is located in a triply nested dictionary under the data key: the first level corresponds to the HNL mass hypothesis <span class="math-tex">\(M_N\)</span>, the second to the process key (cf. table 7) and the third to the <span class="math-tex">\(M(l_{\mathrm{sublead}},l')\)</span> bin for which the efficiency is computed. The values of the bottom-most dictionary are lists containing the efficiencies for a number of HNL lifetimes, as listed in meters in levels/lifetime.</p> <p>Finally, the file fitted_efficiencies.json is also organized as a triply nested dictionary, with the first level corresponding to the HNL mass <span class="math-tex">\(M_N\)</span>, the second to the process key, and where the third level denotes the fit parameter from eq. (3.6). tau0 is for <span class="math-tex">\(\tau_0\)</span>, epsilon0_total for <span class="math-tex">\(\epsilon_0\)</span> (the unbinned prompt efficiency), and epsilon0_binned is a list containing the prompt efficiencies <span class="math-tex">\(\epsilon_{0,b}\)</span> for the five <span class="math-tex">\(M(l_{\mathrm{sublead}},l')\)</span> bins <em>b</em> (in the same order as in efficiencies.json). The layout described here (or a similar one) can be used by experiments to report their signal efficiencies in a way that allows theorists to compute the expected signal for arbitrary choices of mixing angles.</p>
ttbar data with up to 4 jets and 2 leptons
<p>The ttbar data as described and required for the training and evaluation of the generative models in <a href="https://arxiv.org/abs/1901.00875">https://arxiv.org/abs/1901.00875</a></p> <p>The data contains the MET, METphi and the 4-vectors of the final state objects in the format (E, p_T, eta, phi) in the following order:</p> <p>MET, METphi, jet_1, jet_2, jet_3, jet_4, lepton_1, lepton_2 in each row of the .csv file. </p>
Simulated pp collisions at 13 TeV with 2 leptons + 1 b jet final state and selected benchmark Beyond the Standard Model signals
<p>This data-set is comprised of simulated events of pp collisions at 13 TeV with 2 leptons + 1 bottom jet sinal state, with HT > 500 GeV. It includes the following samples</p> <ul> <li>Standard-Model background (bkg), generated at leading order includes the sub-samples Z+Jets, ttbar, WW, WZ, and ZZ. <ul> <li>The processes were generated in kinematic regions to ensure good statistics across the whole phase space. The sampling was carried out using event generation filters at parton level as follows <ul> <li>ttbar: pT <100 GeV; pT in [100, 250] GeV; pT > 250 GeV</li> <li>The scalar sum of the pT of outgoing particles for Z+Jet: ST < 250 Gev; ST in [250, 500] GeV; ST > 500 GeV</li> <li>W/Z pT for dibosons: pT < 250 GeV; pT in [250, 500] GeV; pT > 500 GeV</li> </ul> </li> </ul> </li> <li>Vector-like T-quarks with masses 1.0, 1.2, 1.4 TeV (hq1000, hq1200, hq14000) pair produced either through the Standard-Model gluon (wohg) or through a BSM 3TeV heavy gluon (hg3000)</li> <li>tZ production through a Flavour Changing Neutral Current (fcnc) vertex</li> </ul> <p>The samples are provided with both a full set of features, or with a sanitised set of features. The sanitised features remove some accumulation at zeros from non-reconstructed objects (i.e. missing values). All samples were generated using MadGraph5 2.6.5 and the detector was simulated using Delphes 3 with the default CMS card. For the Standard-Model background, both Pythia 8.2 (with CMS CUETP8M1 underlying event tune and NNPDF 2.3 parton distribution functions) (pythia) and Herwig 7 (herwig) hadronisations are provided to compare the background simulation. For the BSM signals only Pythia is provided.</p> <p>For the details of the generation and on the differences between the two feature sets please refer to <a href="https://link.springer.com/article/10.1140%2Fepjc%2Fs10052-020-08807-w">Finding new physics without learning about it: anomaly detection as a tool for searches at colliders</a> for more details. Each file provides a train:validation:split with the ratios 1:1:1 to ensure equal statistical description of the events at each step of the machine learning workflow.</p>
Semi-leptonic ttbar full-event unfolding R&D dataset
<p>This dataset was generated for the purpose of developing unfolding methods that leverage generative machine learning models. It consists of two pieces: one piece contains events with the Standard Model (SM) production of a top-quark pair in the semi-leptonic decay mode, and the other contains events with top-quark pair production modified by a non-zero EFT operator. The SM dataset contains 15,015,000 events, and the EFT dataset contains 30,000,000. Both datasets store the following event configurations:</p> <ul> <li>Parton level: configurations of all partons that result from the matrix element calculation done using MadGraph</li> <li>Particle level: configurations of all “truth” jets and leptons that result from the parton shower and hadronization modelled using Pythia</li> <li>Detector level: configurations of all reconstruction level jets and leptons, measured by a detector simulated with Delphes and the default CMS detector card.</li> </ul> <p>Each of these configurations is stored in a dedicated group as described below. Throughout, the units of energy and transverse momentum are GeV. For more details on the generation of this dataset, see Ref. [1].</p> <p><strong>Parton level data:</strong></p> <ul> <li>No phase space requirements are placed on the events at parton level. </li> <li>The kinematics of the top, anti-top, W+, W-, and all decay products are contained in groups entitled top, antitop, Wp, and Wm respectively. Each of these contains the kinematics of the parton itself in a group called “particle”, as well as the kinematics of two daughter particles, in groups called “d1” and “d2”. In the case of the tops, these daughters are the W’s and b quarks. In the case of the W’s, these are two light quarks, or a lepton and a neutrino. The “pid” vector contains the PDGID for a given particle, used to identify its type. </li> <li>One detail is that the W’s “particle” description is not always the same as the description of the same W stored as the daughter of the tops. This results from when the W radiates some parton before decaying. </li> </ul> <p><strong>Particle level data:</strong></p> <ul> <li>At particle level all leptons and jets are required to have $p_T > 25$ GeV and absolute pseudo rapidity $|\eta| < 2.5$. </li> <li>Events at particle level are required to have at least one electron or muon and at least 4 jets, of which at least two are b-tagged. Event which pass or fail this criteria are marked by the vector contained in the group “mask”.</li> <li>Electrons and muons are stored in separate groups. Each group contains a vector “mask” which is true only if there is a true particle-level electron or muon in the event, and false if this entry is zero padding. </li> <li>Jets are clustered from stable particle level objects using the anti-kt algorithm with a radius parameter of 0.5. Jet information is stored in the group “jets”, and true jets in the event are again denoted by a true value in the vector “mask”, and zero-padding is marked by a false value. Jets additionally contain a vector “btag” which is 1 if the jet is b-tagged with the default Delphes prescription, and 0 if not.</li> <li>Information on the missing transverse momentum (MET) is contained in the group “met”. The “met” vector gives the magnitude, and the “phi” vector gives the direction in phi of the missing transverse momentum.</li> <li>In addition to the information on the jets, leptons, and MET, the particle level data also contain the configurations for the hadronic top, leptonic top, and ttbar system. These configurations are determined assuming the pseudo-top jet parton assignment algorithm, which is a common method used by LHC experiments when analyzing semileptonic ttbar events.</li> </ul> <p><strong>Detector level data:</strong></p> <ul> <li>Requirements for leptons and jets are the same as for the particle level data.</li> <li>The event selection is the same as the particle level data. Events which pass the selection are again denoted by a true value in the vector “mask”.</li> <li>The data for the leptons, jets, and MET are stored analogously to particle level</li> <li>The configurations of the top quarks and ttbar system are not pre-computed at detector level, since ideally a generative unfolding method would not assume a given jet-carton assignment algorithm when it is being trained. However if the user wishes to pursue such an application, the relevant configurations can be obtained by running the pseudo-top algorithm [2].</li> </ul> <p><strong>Citations:</strong></p> <p>[1] - <a href="https://arxiv.org/abs/2404.14332">https://arxiv.org/abs/2404.14332</a></p> <p>[2] - <a href="https://twiki.cern.ch/twiki/bin/view/LHCPhysics/ParticleLevelTopDefinitions">https://twiki.cern.ch/twiki/bin/view/LHCPhysics/ParticleLevelTopDefinitions</a></p>
Higgs to 4 Leptons with CMS Open data from the Large Hadron Collider
<p>Dataset for MIT 8.S50 online class details to process it can be found here: <br> https://github.com/mit-physics-data/psets/tree/main/pset2 </p>
New Physics Mining at the Large Hadron Collider: A -> 4 leptons
<p><span class="math-tex">\(A \to 4\ell\)</span> signal events reconstructed by inclusive single-muon selection.</p> <p>Events are represented as an array of physics-motivated high-level features.</p> <p>Details are given in https://arxiv.org/abs/1811.10276</p>
Supplementary material for "Two-loop QED corrections to the scattering of four massive leptons"
<p>We provide QED corrections to the scattering amplitudes of four massive leptons for Bhabha and Møller kinematics through two loops. If you use the results distributed with this repository in your research work, please cite <a href="https://arxiv.org/abs/2311.06385">2311.06385.</a> See the README.pdf file inside the archive for instructions.</p>
Higgs to tau dataset decaying to to tau leptons
<p>Dataset used in MIT lecture details can be found here: </p> <p>https://github.com/mit-physics-data/lectures/tree/main/lecture13</p>
Data from: The recumbirostran <em>Hapsidopareion lepton</em> from the Early Permian of Oklahoma reassessed through HRμCT and the effects of fossoriality on the neurocranium of Pan-Amniota
Open the record for dataset details and reuse information.
Database of tree-level completions of lepton-number-violating effective operators
<p>The complete database of unfiltered completions of the lepton-number-violating effective operators accompanying our paper "Exploding operators for Majorana neutrino masses and beyond." Please cite this paper if you use this database.</p> <p>The pickled file unfiltered.p has the same structure as democratic.p in our <a href="https://github.com/johngarg/neutrinomass">example code</a>. The raw completions are also provided in a text-based format. The path to the raw_completions directory initialises the ModelDatabase object provided in the database submodule of our example code. Please see the example notebooks in our example code for instructions on how to use and analyses these data.</p>
Semi-leptonic TTbar Production Dataset for Unfolding Studies
Open the record for dataset details and reuse information.
ttH(bb) dataset in the semi-leptonic decay channel
<p>Higgs boson dataset in the \(t\bar{t}H(b\bar{b}) \) semil-leptonic channel, used for studies of deep learning and quantum machine learning classification studies [1, 2].</p> <p>The simulation of the \(t\bar{t}H(b\bar{b}) \) semi-leptonic channel produces a data set that consists of the following features:</p> <ol> <li>Jet features: \((p_\mathrm{T}, \eta, \phi, E, \mathrm{b-tag}, p_\mathrm{x}, p_\mathrm{y}, p_\mathrm{z})\)</li> <li>Leptonic features: \((p_\mathrm{T}, \eta, \phi, E, p_\mathrm{x}, p_\mathrm{y}, p_\mathrm{z})\)</li> <li> Missing energy features: \((\phi, p_\mathrm{T}, p_\mathrm{x}, p_\mathrm{y})\)</li> </ol> <p>Before processing the data with (quantum) machine learning algorithms, the features are filtered using the following physically motivated criteria to constrain the problem in a suitable phase space. These criteria take into account the geometric acceptance of the detector and the goal of background suppression. </p> <p>The following preprocessing steps are applied in the related studies using this dataset:</p> <ul> <li>For electrons: \(p_\mathrm{T} > 30 \) GeV and \(|\eta|<2.1\)</li> <li>For muons: \(p_\mathrm{T} > 26\) GeV and \(|\eta|<2.1\)</li> <li>For jets: \(p_\mathrm{T} > 30\) GeV and \(|\eta|<2.4\)</li> <li>Isolation of the leptons with respect to jets is higher than the benchmark value of 0.1.</li> <li>Require at least 4 jets per event, at least 2 b-tagged jets, and exactly one lepton. </li> <li>The first seven most energetic jets are kept per collision event, allowing for one extra jet beyond the leading order expectation of 6 jets, to account for final state radiation.</li> </ul> <p>These criteria constrain the problem in a suitable phase space, taking into account the geometric acceptance of the CMS detector and the goal of background suppression. For more details, please see the corresponding papers.</p> <p>[1] V. Belis et al., <em>Higgs analysis with quantum classifiers, </em><a href="https://www.epj-conferences.org/articles/epjconf/abs/2021/05/epjconf_chep2021_03070/epjconf_chep2021_03070.html">EPJ Web Conf. 251, 03070 (2021)</a><em>, </em>arXiv: <a href="https://arxiv.org/abs/2104.07692">2104.07692</a>.</p> <p>[2] V. Belis et al., <em>Guided Quantum Compression for Higgs identification, </em>arXiv: <a href="https://arxiv.org/abs/2402.09524">2402.09524</a>.</p>
Leptonic SFG emission
<p>Data and code supporting publication entitled "Leptonic gamma-ray emission from star-forming galaxies".</p> <h3><strong>Data files description</strong></h3> <p><strong>sfgfit.tar.gz </strong>: data model and fit routines</p> <p>Python package "sfgfit" implementing the gamma-ray ray emission model from star-forming regions. Contains the README file with installation instructions and "pyproject.toml" file with additional package dependencies.</p> <p><strong>worktree.tar.gz </strong>: work tree with scripts, settings and data samples<strong><br></strong></p> <p>A copy of the work tree folder structure, containing the scripts, configuration files and final data products; ready to use once unpacked (provided "sfgfit" and FermiPy are installed). Underlying folder structure is as follows:</p> <p><em>worktree/</em><br><em>├── data : </em>Fermi/LAT and IACT data sets and analysis scripts for each if of the SFGs considered. Fermi/LAT data sets contain only the final data products (SEDs) as generated by FermiPy; the used intermediate data sets can be re-created using the "analysis.py" and "config.yaml" files in each corresponding subfolder from the full Fermi/LAT event lists, available at <a href="https://fermi.gsfc.nasa.gov/ssc/data/access/">https://fermi.gsfc.nasa.gov/ssc/data/access/</a>. To run the SED extraction, it is sufficient to execute the "analysis.py" script in the corresponding subfolder.<br>Example of the underlying structure for M82 galaxy:<br>│ ├── m82<br>│ │ └── data<br>│ │ ├── fermipy<br>│ │ │ ├── lc<br>│ │ │ │ └── 100mev-psf-classes-tuned-1TeV-2024<br>│ │ │ │ └── dt=560d<br>│ │ │ └── sed<br>│ │ │ └── 60mev-psf-classes-tuned-1TeV-2024-bs0.2<br>│ │ │ ├── analysis.py<br>│ │ │ ├── config.yaml<br>│ │ │ └── out<br>│ │ └── iact_sed.ecsv<br><br><em>└── scripts</em><br><em> ├── make-analysis </em>: scripts to re-generate analysis folders (with processing scripts and settings) based on universal templates<br><em> └── mcmc</em> : Monte Carlo fitting code using "sfgfit" package above alone with its configuration files and output results. Configuration files are for each source and spectral model are stored under the "<em>config/</em>" subfolder.</p> <h3>Demo (test code execution)</h3> <ul> <li>Fermi/LAT data analysis: LAT data reduction resource-consuming and takes almost a day on a regular desktop computer. Intermediate files generated by FermiPy are relatively large (around 3Gb per source) and thus are not included here. Provided FermiPy is installed and all-sky event list as well as spacecraft files are downloaded from <a href="https://fermi.gsfc.nasa.gov/ssc/data/access/">https://fermi.gsfc.nasa.gov/ssc/data/access/</a>, these can be regenerated as a part of the data analysis run as<br><em>> cd worktree/data/m82/data/fermipy/sed/60mev-psf-classes-tuned-1TeV-2024-bs0.2/<br></em><em># update the event and spacecraft files paths in the "config.yaml" file<br></em><em>> python3 analysis.py</em><br>Expected output: "<em>4fgl_j0955.7+6940_sed.fits</em>" in the output folder identical to the one provided here</li> <li>star-formation emission model fitting: model optimization procedure is MCMC-based and its full run takes up to 30 hr on a 30-core machine. The data sets required for it are provided in the corresponding "<em>worktree/data/*/fermipy/sed/60mev-psf-classes-tuned-1TeV-2024-bs0.2/out/</em>" folders. In order to run optimization for a shorter time (without convergence, only to ensure it works) for a single object (e.g. M82) one may<br><em>> cd worktree/scripts/mcmc</em><br><em># update the data path templates in "config/m82/sfr-steady-pwl.yaml"</em><br><em>> fit-sed-mcmc.py --config=config/m82/sfr-steady-pwl.yaml</em><br>Expected output: generated "out/m82_sfr2_steady_pwl_model_mcmc.h5" file with MCMC samples, similar to the one provided.</li> </ul> <p>For example, to re-do the MCMC fit for M82 it is sufficient run "<em>fit-sed-mcmc.py --config=config/m82/sfr-steady-pwl.yaml</em>" () The input files for the code are the in the "<em>worktree/data/*/fermipy/sed/60mev-psf-classes-tuned-1TeV-2024-bs0.2/out/</em>" folders.</p> <h3>System requirements</h3> <p>This software was tested on Red Hat Enterprise Linux Server release 7.7, running conda 24.4.0, Python versions 3.8.3 and 3.9.19. Installation time varies depending on system and internet connection speed and can be from few seconds for "sfgfit" to more than 10 minutes for "FermiPy" required to process the data.</p> <h3>Non-standard hardware requirements</h3> <p>None</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.