Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,140
datasets available to search
ShareScore release 0.9.0
Dataset results
1,140 results for “TOPS”
Top Quark Tagging Reference Dataset
<p>A set of MC simulated training/testing events for the evaluation of top quark tagging architectures.</p> <p>In total 1.2M training events, 400k validation events and 400k test events. Use “train” for training, “val” for validation during the training and “test” for final testing and reporting results.</p> <p><strong>Description</strong></p> <ul> <li> <p>14 TeV, hadronic tops for signal, qcd diets background, Delphes ATLAS detector card with Pythia8</p> </li> <li> <p>No MPI/pile-up included</p> </li> <li> <p>Clustering of particle-flow entries (produced by Delphes E-flow) into anti-kT 0.8 jets in the pT range [550,650] GeV</p> </li> <li> <p>All top jets are matched to a parton-level top within ∆R = 0.8, and to all top decay partons within 0.8</p> </li> <li> <p>Jets are required to have |eta| < 2</p> </li> <li> <p>The leading 200 jet constituent four-momenta are stored, with zero-padding for jets with fewer than 200</p> </li> <li> <p>Constituents are sorted by pT, with the highest pT one first</p> </li> <li> <p>The truth top four-momentum is stored as truth_px etc.</p> </li> <li> <p>A flag (1 for top, 0 for QCD) is kept for each jet. It is called is_signal_new</p> </li> <li> <p>The variable "ttv" (= test/train/validation) is kept for each jet. It indicates to which dataset the jet belongs. It is redundant as the different sets are already distributed as different files.</p> </li> </ul>
caseysaenger/ForamMgCa_PSM: files and scripts for revised version of manuscript "Calibration and validation of environmental controls on planktic foraminifera Mg/Ca using global core-top data".
<p>files and scripts for revised version of manuscript "Calibration and validation of environmental controls on planktic foraminifera Mg/Ca using global core-top data". Saenger, C. and M. N. Evans. Resubmitted to Paleoceanography and Paleoclimatology, May 3, 2019.</p>
Profiling results for top pair plus jets production with Pepper
<h2>Profiling results for top pair plus jets production with Pepper v1.1.1</h2> <h3>Contents</h3> <p>This includes data/output for the following GPU accelerated unweighted event generation runs:</p> <ul> <li>A run over a few seconds to produce pepper-internal timing results (`*timers.csv` files)</li> <li>The same run, but with `nvprof` (`*.nvprof` files)</li> <li>A run with a single event batch, with `ncu --import-source on --set full` (`*.ncu-rep` files)</li> </ul> <p>The simulated process is top pair production with n jets, with n = 0...4, at 13 TeV. The configuration files (`pepper.ini`) are included. Also the `pepper_cache` directories are included.</p> <p>All results are for the "main" pepper variant (which uses Kokkos to utilise the GPU) and for the "native" pepper variant (which uses CUDA directly).</p> <h3>Software stack and CPU/GPU hardware details</h3> <p>The software/hardware used for the profiling are as follows:</p> <ul> <li>GPU Driver: NVIDIA-SMI: 550.100, Driver Version: 550.100, CUDA Version: 12.4</li> <li>GPU: Tesla V100S-PCIE-32GB</li> <li>Cuda compilation tools: release 11.6, V11.6.124</li> <li>Kokkos 4.3.01</li> <li>gcc (GCC) 11.4.1 20231218 (Red Hat 11.4.1-3)</li> <li>Intel(R) Xeon(R) Silver 4214R CPU @ 2.40GHz</li> </ul> <p>Note that the event generation includes write-out of event files. The LHEH5 files are written to local SSD storage. The event files are not included in this dataset.</p> <h3>Additional run parameters</h3> <table> <tbody> <tr> <td>TTBar</td> <td>number of batches</td> <td>batch size</td> <td>rate (main variant)</td> <td>rate (native variant)</td> </tr> <tr> <td>+0j</td> <td>20</td> <td>1,179,648</td> <td>1.4e10</td> <td>2.0e10</td> </tr> <tr> <td>+1j</td> <td>20</td> <td>1,179,648</td> <td>1.5e10</td> <td>2.4e10</td> </tr> <tr> <td>+2j</td> <td>20</td> <td>1,179,648 * 2</td> <td>1.3e10</td> <td>2.7e10</td> </tr> <tr> <td>+3j</td> <td>20</td> <td>1,179,648 / 2</td> <td>6.4e9</td> <td>1.4e10</td> </tr> <tr> <td>+4j</td> <td>20</td> <td>1,179,648</td> <td>2.2e9</td> <td>2.2e9</td> </tr> </tbody> </table> <p>The number of batches has been chosen such that the most important sub processes are samples and such that the overall runtime is at least a few seconds. The batch size has been chosen by scanning over factors of two and picking the best-performing batch size with the native variant. The event rate is the number given in the Pepper output. It does not include the closing time of the generated HDF5 file, which can be significant for the two lowest multiplicities. This can be checked by inspecting the corresponding relative and absolute timing results in the `*timing.csv` files included in the dataset.</p>
Top 1000 boardgames in https://boardgamegeek.com/
<p>Top 1000 https://boardgamegeek.com/ board games obtained on 2024-11-09.</p> <p>The variables collected are: title, year of publication, minimum players, maximum players, minimum duration, maximum duration, minimum age, geek rating, general rating and description.</p>
Canopy top height and indicative high carbon stock maps for Indonesia, Malaysia, and Philippines
<p>Canopy top height and indicative high carbon stock maps for Indonesia, Malaysia, and Philippines. The provided land cover maps follow the high carbon stock approach (HCSA) stratifying vegetation based on the estimated carbon density (aboveground biomass). A deep convolutional neural network was trained to estimate canopy top height from Sentinel-2 optical satellite images using reference data derived from GEDI lidar waveforms. Carbon density and high carbon stock classes were derived from these dense canopy height maps using calibration data from an airborne lidar campaign in Sabah, Borneo. The resulting maps have a ground sampling distance (GSD) of 10 m and are based on images between 1st of September 2020 and 1st of March 2021.</p> <p>The style files (color_style_HCS.qml, color_style_canopy_top_height.qml) contain the color coding and can be loaded for visualization (e.g. in QGIS).</p> <p>The indicative HCS maps contain 9 land cover categories noted as "Label: name [colorcode]":</p> <p> 0: Open land (OL) [#440154]<br> 1: Scrub (S) [#404387]<br> 2: Young regenerating forest (YRF) [#29788e]<br> 3: Low density forest (LDF) [#22a884]<br> 4: Medium density forest (MDF) [#7ad251]<br> 5: High density forest (HDF) [#fde725]<br> 10: Oil palm [#fcffa4]<br> 11: Coconut [#a4feff]<br> 50: Urban [#fa0000]<br> 255: No data</p> <p><strong>Citation: </strong>Use of these data require citation of this dataset and the original research articles. These citations are as follows:</p> <p>Lang, N., Schindler, K., & Wegner, J. D. (2021). High carbon stock mapping at large scale with optical satellite imagery and spaceborne LIDAR. arXiv preprint arXiv:2107.07431.</p> <p>Rodríguez, A. C., D'Aronco, S., Schindler, K., & Wegner, J. D. (2021). Mapping oil palm density at country scale: An active learning approach. <em>Remote Sensing of Environment</em>, <em>261</em>, 112479.</p> <p>Lang, N., Rodríguez, A. C., Schindler, K., & Wegner, J. D. (2021). Canopy top height and indicative high carbon stock maps for Indonesia, Malaysia, and Philippines (Version 1.0) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.5012448</p> <p> </p>
TOPS Graphic for the 2022 PACE Applications Workshop
<p>This graphic was made to be displayed on a TV screen during the NASA 2022 PACE (Plankton, Aerosol, Cloud, ocean Ecosystem) Applications Workshop held on September 14-15, 2022. </p>
Filtered canopy top height estimates from GEDI LIDAR waveforms for 2019 and 2020
<p>Canopy top height (RH98) is estimated from GEDI L1B waveforms globally between 51.6° N & S from L1B Version 1 data for April-July 2019 and 2020. We refer to the original research article below for further information. The footprint data were filtered with respect to predictive uncertainty and MODIS non-vegetated probability.</p><p>The unfiltered data organized in hdf5 files corresponding to the orbit files of the GEDI L1B Version 1 data is available here:</p><p>April-July 2019: <a href="https://doi.org/10.5281/zenodo.5704852">https://doi.org/10.5281/zenodo.5704852</a></p><p>April-July 2020: <a href="https://doi.org/10.5281/zenodo.7737869">https://doi.org/10.5281/zenodo.7737869</a></p><p><strong>GEDI mission website</strong>: <a href="https://gedi.umd.edu/">https://gedi.umd.edu/</a>.</p><p><strong>Citation:</strong></p><p>Use of these data require citation of this dataset:</p><p>Lang, Nico, Kalischek, Nikolai, Armston, John, Schindler, Konrad, Dubayah, Ralph, & Wegner, Jan Dirk. (2021). Filtered canopy top height estimates from GEDI LIDAR waveforms for 2019 and 2020 (1.0) [Dataset]. Zenodo. <a href="https://doi.org/10.5281/zenodo.7737946">https://doi.org/10.5281/zenodo.7737946</a></p><p>Original research article:</p><p>Lang, N., Kalischek, N., Armston, J., Schindler, K., Dubayah, R., & Wegner, J. D. (2022). Global canopy height regression and uncertainty estimation from GEDI LIDAR waveforms with deep ensembles. <i>Remote Sensing of Environment</i>, <i>268</i>, 112760.</p><p>This filtered dataset (2019 and 2020) was used to develop the global canopy height model fusing Sentinel-2 and GEDI that is presented in:</p><p>Lang, N., Jetz, W., Schindler, K., & Wegner, J. D. (2023). A high-resolution canopy height model of the Earth. Nature Ecology & Evolution, 1-12, <a href="https://doi.org/10.1038/s41559-023-02206-6">https://doi.org/10.1038/s41559-023-02206-6</a></p>
Keras video classification example with a subset of UCF101 - Action Recognition Data Set (top 10 videos)
<p>Classify video clips with natural scenes of actions performed by people visible in the videos.</p> <p>See the UCF101 Dataset web page: <a href="https://www.crcv.ucf.edu/data/UCF101.php#Results_on_UCF101">https://www.crcv.ucf.edu/data/UCF101.php#Results_on_UCF101</a></p> <p>This example datasets consists of the 10 most numerous video from the UCF101 dataset. For the top 5 version, see: <a href="https://doi.org/10.5281/zenodo.7924745">https://doi.org/10.5281/zenodo.7924745</a> .</p> <p>Based on this code: <a href="https://keras.io/examples/vision/video_classification/">https://keras.io/examples/vision/video_classification/</a> (needs to be updated, if has not yet been already; see the issue: <a href="https://github.com/keras-team/keras-io/issues/1342">https://github.com/keras-team/keras-io/issues/1342</a>).</p> <p>Testing if data can be downloaded from figshare with `wget`, see: <a href="https://github.com/mojaveazure/angsd-wrapper/issues/10">https://github.com/mojaveazure/angsd-wrapper/issues/10</a></p> <p>For generating the subset, see this notebook: <a href="https://colab.research.google.com/github/sayakpaul/Action-Recognition-in-TensorFlow/blob/main/Data_Preparation_UCF101.ipynb">https://colab.research.google.com/github/sayakpaul/Action-Recognition-in-TensorFlow/blob/main/Data_Preparation_UCF101.ipynb</a> -- however, it also needs to be adjusted (if has not yet been already - then I will post a link to the notebook here or elsewhere, e.g., in the corrected notebook with Keras example).</p> <p>I would like to thank Sayak Paul for contacting me about his example at Keras documentation being out of date. </p> <p>Cite this dataset as:</p> <p>Soomro, K., Zamir, A. R., & Shah, M. (2012). UCF101: A dataset of 101 human actions classes from videos in the wild. <em>arXiv preprint arXiv:1212.0402</em>. <a href="https://doi.org/10.48550/arXiv.1212.0402">https://doi.org/10.48550/arXiv.1212.0402</a></p> <p>To download the dataset via the command line, please use:</p> <pre><code class="language-bash">wget -q https://zenodo.org/record/7882861/files/ucf101_top10.tar.gz -O ucf101_top10.tar.gz tar xf ucf101_top10.tar.gz</code></pre> <p> </p>
Keras video classification example with a subset of UCF101 - Action Recognition Data Set (top 5 videos)
<p>Classify video clips with natural scenes of actions performed by people visible in the videos.</p> <p>See the UCF101 Dataset web page: <a href="https://www.crcv.ucf.edu/data/UCF101.php#Results_on_UCF101">https://www.crcv.ucf.edu/data/UCF101.php#Results_on_UCF101</a></p> <p>This example datasets consists of the 5 most numerous video from the UCF101 dataset. For the top 10 version see: <a href="https://doi.org/10.5281/zenodo.7882861">https://doi.org/10.5281/zenodo.7882861</a> .</p> <p>Based on this code: <a href="https://keras.io/examples/vision/video_classification/">https://keras.io/examples/vision/video_classification/</a> (needs to be updated, if has not yet been already; see the issue: <a href="https://github.com/keras-team/keras-io/issues/1342">https://github.com/keras-team/keras-io/issues/1342</a>).</p> <p>Testing if data can be downloaded from figshare with `wget`, see: <a href="https://github.com/mojaveazure/angsd-wrapper/issues/10">https://github.com/mojaveazure/angsd-wrapper/issues/10</a></p> <p>For generating the subset, see this notebook: <a href="https://colab.research.google.com/github/sayakpaul/Action-Recognition-in-TensorFlow/blob/main/Data_Preparation_UCF101.ipynb">https://colab.research.google.com/github/sayakpaul/Action-Recognition-in-TensorFlow/blob/main/Data_Preparation_UCF101.ipynb</a> -- however, it also needs to be adjusted (if has not yet been already - then I will post a link to the notebook here or elsewhere, e.g., in the corrected notebook with Keras example).</p> <p>I would like to thank Sayak Paul for contacting me about his example at Keras documentation being out of date. </p> <p>Cite this dataset as:</p> <p>Soomro, K., Zamir, A. R., & Shah, M. (2012). UCF101: A dataset of 101 human actions classes from videos in the wild. <em>arXiv preprint arXiv:1212.0402</em>. <a href="https://doi.org/10.48550/arXiv.1212.0402">https://doi.org/10.48550/arXiv.1212.0402</a></p> <p>To download the dataset via the command line, please use:</p> <pre><code class="language-bash">wget -q https://zenodo.org/record/7924745/files/ucf101_top5.tar.gz -O ucf101_top5.tar.gz tar xf ucf101_top5.tar.gz</code></pre>
Dataset of Linkability Networks of Ethereum Accounts Involved in NFT trading of Top 15 NFT Collections
<p>The data is organized in 32 files:</p> <ul> <li> <p>collection_metadata.txt stores basic information about each collection that was analyzed. The graph data of the collection is stored in a file named as the nickname of the collection (slug column in collection_metadata.txt). The most important columns/properties are: rank, slug – nickname, creation date, address, and volume data, ...</p> </li> <li> <p>15 files reporting NFT ownership transfer, the names corresponding to the nicknames of the collections in collection_metadata.txt.</p> </li> <li> <p>15 files with graph data, the names correspond to the nicknames of the collections in collection_metadata.txt. Each file gives the address linkage graph of one of the top 15 collections according to monetary volume on Opensea marketplace, computed as presented in methodology, exclusively on Ethereum blockchain.</p> </li> </ul>
Sagittal otolith of Lota lota (total length=28 cm) from bottom (distal, concave, anti-sulcus side) and top (proximal, convex, with sulcus acusticus) view.
<p>This image shows the left sagittal otolith of <em>Lota lota </em>(total length=28 cm) from bottom (distal, concave, anti-sulcus side) and top (proximal, convex, with sulcus acusticus) view.</p>
Sagittal otolith of Perca fluviatilis (total length=17 cm) from bottom (distal, concave, anti-sulcus side) and top (proximal, convex, with sulcus acusticus) view.
<p>This image shows the left sagittal otolith of <em>Perca fluviatilis </em>(total length=17 cm) from bottom (distal, concave, anti-sulcus side) and top (proximal, convex, with sulcus acusticus) view.</p>
All three types of otoliths of Cyprinus carpio (total length=30 cm); left/right asterisci and lapilli from top and bottom view, and one sagittus.
<p>The image shows all three types of otoliths of <em>Cyprinus carpio</em> (total length=30 cm): left/right <em>asterisci</em> and <em>lapilli </em>from top and bottom view, and one sagittus. </p>
GitHub Top 25 Software Project Analysis
<p>Companion dataset for the paper "For a More Transparent Governance of Open Source" published in the Communications of the ACM, 66, 8, 28-30, 2023.</p> <p>Data collected on May, 13th, 2022.</p> <p><strong>Note:</strong> cell annotations are only visible in the Excel version of the dataset.</p>
Top Jet W-Momentum Reconstruction Dataset
<p><strong>Overview</strong></p> <p>A set of Monte Carlo simulated events, for the evaluation of top quarks' (and their child particles') momentum reconstruction, produced using the HEPData4ML package [1]. Specifically, the entries in this dataset correspond with top quark jets, and the momentum of the jets' constituent particles. This is a newer version of the "Top Quark Momentum Reconstruction Dataset" [2], but with sufficiently large changes to warrant this separate posting.</p> <p>The dataset is saved in HDF5 format, as sets of arrays with keys (as detailed below). There are ~1.5M events, approximately broken down into the following sets:</p> <ul> <li>Training: 700k events (files with "_train" suffix)</li> <li>Validation: 200k events (files with "_valid" suffix)</li> <li>Testing (small): 100k events (files with "_test" suffix)</li> <li>Testing (large): 500k events (files with "_test_large" suffix)</li> </ul> <p>The two separate types of testing files -- small and large -- are independent from one another, the former for conveniently running quicker testing and the latter for testing with a larger sample.</p> <p>There are four version of the dataset present, with the versions indicated by the filenames. The different versions correspond with whether or not fast detector simulation was performed (versus truth-level jets), and whether or not the W-boson mass was modified: One version of the dataset uses the nominal value of <span>\(m_W = 80.385 \text{ GeV}\)</span> as used by Pythia8 [3], whereas another uses a variable mW taking on 101 values evenly-spaced as <span>\(m_W \in \{ 64.308,96.462 \} \text{ GeV}\)</span>. The dataset naming scheme is as follows:</p> <ul> <li>train.h5 : jets clustered from truth-level, nominal mW</li> <li>train_mW.h5: jets clustered from truth-level, variable mW</li> <li>train_delphes.h5: jets clustered from Delphes outputs, nominal mW</li> <li>train_delphes_mW.h5: jets clustered from Delphes outputs, variable mW</li> </ul> <p><strong>Description</strong></p> <ul> <li>13 TeV center-of-mass energy, fully hadronic top quark decays, simulated with Pythia8. (<span>\(t \rightarrow W \, b, \; W\rightarrow q \, q'\)</span>) <ul> <li>Events are generated with leading top quark pT in [550,650] GeV. (set via Pythia8's <span>\(\hat{p}_{T,\text{ min}}\)</span> and <span>\(\hat{p}_{T,\text{ max}}\)</span> variables)</li> <li>No inital- or final-state radiation (ISR/FSR), nor multi-parton interactions (MPI)</li> <li>Where applicable, detector simulation is done using DELPHES [4], with the ATLAS detector card.</li> </ul> </li> <li>Clustering of particles/objects is done via FastJet [5], using the anti-k<sub>T</sub> algorithm, with <span>\(R=0.8\)</span> . <ul> <li>For the truth-level data, inputs to jet clustering are truth-level, final-state particles (i.e. clustering "truth jets").</li> <li>For the data with detector simulation, the inputs are calorimeter towers from DELPHES. <ul> <li>`Tower` objects from DELPHES (<em>not</em> E-flow objects, no tracking information)</li> </ul> </li> <li>Each entry in the dataset corresponds with a single top quark jet, extracted from a <span>\(t\bar{t}\)</span> event. <ul> <li>All jets are matched to a parton-level top quark within <span>\(\Delta R < 0.8\)</span> . We choose the jet <em>nearest</em> the parton-level top quark.</li> <li>Jets are required to have <span>\(|\eta| < 2\)</span>, and <span>\(p_{T} > 15 \text{ GeV}\)</span>.</li> <li>The 200 leading (highest-p<sub>T</sub>) jet constituent four-momenta are stored in Cartesian coordinates (E,px,py,pz), sorted by decreasing pT, with zero-padding.</li> <li>The jet four-momentum is stored in Cartesian coordinates (E, p<sub>x</sub>, p<sub>y</sub>, p<sub>z</sub>), as well as in cylindrical coordinates <span>\((p_T,\eta,\phi,m)\)</span>.</li> <li>The truth (parton-level) four-momenta of the top quark, the bottom quark the W-boson, and the quarks to which the W-boson decays, are stored in Cartesian coordinates. <ul> <li>In addition, the momenta of the 120 leading stable daughter particles of the W-boson are stored in Cartesian coordinates.</li> </ul> </li> </ul> </li> </ul> </li> </ul> <p>Description of data fields & metadata<br>Below is a brief description of the various fields in the dataset. The dataset also contains metadata fields, <a href="https://docs.h5py.org/en/stable/high/attr.html">stored using HDF5's "attributes"</a>. This is used for fields that are common across many events, and stores information such as generator-level configurations (in principle, all the information is stored as to be able to recreate the dataset with the HEPData4ML tool).</p> <p>Note that fields whose keys have the prefix "jh_" correspond with output from the Johns Hopkins top tagger [6], as implemented in FastJet.</p> <p>Also note that for the keys corresponding with four-momenta in Cartesian coordinates, there are rotated versions of these fields -- the data has been rotated so that the W-boson is at <span>\((\theta=0, \phi=0)\)</span>, and the b-quark is in the <span>\((\theta=0, \phi < 0)\)</span> plane. This rotation is potentially useful for visualizations of the events.</p> <ul> <li> <p>Nobj: The number of constituents in the jet.</p> </li> <li> <p>Pmu: The four-momenta of the jet constituents, in (E, p<sub>x</sub>, p<sub>y</sub>, p<sub>z</sub>). Sorted by decreasing p<sub>T</sub> and zero-padded to a length of 200.</p> <ul> <li> <p>Pmu_rot: Rotated version.</p> </li> </ul> </li> <li> <p>contained_daughter_sum_Pmu: Four-momentum sum of the stable daughter particles of the W-boson that fall within <span>\(\Delta R < 0.8\)</span> of the jet centroid.</p> <ul> <li> <p>contained_daughter_sum_Pmu_rot: Rotated version.</p> </li> </ul> </li> <li> <p>cross_section: Cross-section for the corresponding process, reported by Pythia8.</p> </li> <li> <p>cross_section_uncertainty: Cross-section uncertainty for the corresponding process, reported by Pythia8.</p> </li> <li> <p>energy_ratio smeared: Ratio of the true energy of W-boson daughter particles contributing to this calorimeter tower, divided by the total smeared energy in this calorimeter tower.</p> <ul> <li> <p>Only relevant for the DELPHES datasets.</p> </li> </ul> </li> <li> <p>energy_ratio_truth: Ratio of the true energy of W-boson daughter particles contributing to this calorimeter tower, divided by the total true energy of particles contributing to this calorimeter tower.</p> <ul> <li> <p>The above definition is relevant only for the DELPHES datasets. For the truth-level datasets, this field is repurposed to store a value (0 or 1) indicating whether or not the given particle (whose momentum is in the `Pmu` field) is a W-boson daughter.</p> </li> </ul> </li> <li> <p>event_idx: Redundant -- used for event indexing during the event generation process.</p> </li> <li> <p>is_signal: Redundant -- indicates whether an event is signal or background, but this is a fully signal dataset. Potentially useful if combining with other datasets produced with HEPData4ML.</p> </li> <li> <p>jet_Pmu: Four-momentum of the jet, in (E, p<sub>x</sub>, p<sub>y</sub>, p<sub>z</sub>).</p> <ul> <li> <p>jet_Pmu_rot: Rotated version.</p> </li> </ul> </li> <li> <p>jet_Pmu_cyl: Four-momentum of the jet, in <span>\((pT_,\eta,\phi,m)\)</span>.</p> </li> <li> <p>jet_bqq_contained_dR06: Boolean flag indicating whether or not the truth-level b and the two quarks from W decay are contained within <span>\(\Delta R < 0.6\)</span> of the jet centroid.</p> </li> <li> <p>jet_bqq_contained_dR08: Boolean flag indicating whether or not the truth-level b and the two quarks from W decay are contained within <span>\(\Delta R < 0.8\)</span> of the jet centroid.</p> </li> <li> <p>jet_bqq_dr_max: Maximum of <span>\(\big\lbrace \Delta R \left( \text{jet},b \right), \; \Delta R \left( \text{jet},q \right), \; \Delta R \left( \text{jet},q' \right) \big\rbrace\)</span>.</p> </li> <li> <p>jet_qq_contained_dR06: Boolean flag indicating whether or not the two quarks from W decay are contained within <span>\(\Delta R < 0.6\)</span> of the jet centroid.</p> </li> <li> <p>jet_qq_contained_dR08: Boolean flag indicating whether or not the two quarks from W decay are contained within <span>\(\Delta R < 0.8\)</span> of the jet centroid.</p> </li> <li> <p>jet_qq_dr_max: Maximum of <span>\(\big\lbrace \Delta R \left( \text{jet},q \right), \; \Delta R \left( \text{jet},q' \right) \big\rbrace\)</span>.</p> </li> <li> <p>jet_top_daughters_contained_dR08: Boolean flag indicating whether the final-state daughters of the top quark are within <span>\(\Delta R < 0.8\)</span> of the jet centroid. Specifically, the algorithm for this flag checks that the jet contains the stable daughters of both the b quark and the W boson. For the b and W each, daughter particles are allowed to be uncontained as long as (for each particle) the <span>\(p_T\)</span> of the sum of uncontained daughters is below <span>\(2.5 \text{ GeV}\)</span>.</p> </li> <li> <p>jh_W_Nobj: Number of constituents in the W-boson candidate identified by the JH tagger.</p> </li> <li> <p>jh_W_Pmu: Four-momentum of the JH tagger W-boson candidate, in (E, p<sub>x</sub>, p<sub>y</sub>, p<sub>z</sub>).</p> <ul> <li> <p>jh_W_Pmu_rot: Rotated version.</p> </li> </ul> </li> <li> <p>jh_W_constituent_Pmu: Four-momentum of the constituents of the JH tagger W-boson candidate, in (E, p<sub>x</sub>, p<sub>y</sub>, p<sub>z</sub>).</p> <ul> <li> <p>jh_W_constituent_Pmu_rot: Rotated version.</p> </li> </ul> </li> <li> <p>jh_m: Mass of the JH W-boson candidate.</p> </li> <li> <p>jh_m_resolution: Ratio of JH W-boson candidate mass, versus the true W-boson mass.</p> </li> <li> <p>jh_pt: <span>\(p_T\)</span> of the JH W-boson candidate.</p> </li> <li> <p>jh_pt_resolution: Ratio of JH W-boson candidate <span>\(p_T\)</span>, versus the true W-boson mass.</p> </li> <li> <p>jh_tag: Whether or not a jet was tagged by the JH tagger.</p> </li> <li> <p>mc_weight: Monte Carlo weight for this event, reported by Pythia8.</p> </li> <li> <p>process_code: Process code reported by Pythia8.</p> </li> <li> <p>rotation_matrix: Rotation matrix for rotating the events' 3-momenta as to produce the rotated copies stored in the dataset.</p> </li> <li> <p>truth_Nobj: Number of truth-level particles (saved in truth_Pmu).</p> </li> <li> <p>truth_Pdg: PDG codes of the truth-level particles.</p> </li> <li> <p>truth_Pmu: Truth-level particles: The top quark, bottom quark, W boson, q, q', and 120 leading, stable W-boson daughter particles, in (E, p<sub>x</sub>, p<sub>y</sub>, p<sub>z</sub>). A few of these are also stored in separate keys:</p> </li> <li> <ul> <li>truth_Pmu_0: Top quark. <ul> <li>truth_Pmu_0_rot: Rotated version.</li> </ul> </li> <li> <p>truth_Pmu_1: Bottom quark.</p> <ul> <li> <p>truth_Pmu_1_rot: Rotated version.</p> </li> </ul> </li> <li> <p>truth_Pmu_2: W-boson.</p> <ul> <li> <p>truth_Pmu_2_rot: Rotated version.</p> </li> </ul> </li> <li> <p>truth_Pmu_3: q from W decay.</p> <ul> <li> <p>truth_Pmu_3_rot: Rotated version.</p> </li> </ul> </li> <li> <p>truth_Pmu_4: q' from W decay.</p> <ul> <li> <p>truth_Pmu_4_rot: Rotated version.</p> </li> </ul> </li> <li> <p>truth_Pmu_0_rot: Rotated version of `truth_Pmu`.</p> </li> </ul> </li> </ul> <p>The following fields correspond with metadata -- they provide the index of the corresponding metadata entry for each event:</p> <ul> <li> <p>command_line_arguments: The command-line arguments passed to HEPData4ML's `run.py` script.</p> </li> <li> <p>config_file: The contents of the Python configuration file used for HEPData4ML. This, together with the command-line arguments, defines how the tool was run, what processes, jet clustering and post-processing was done, etc.</p> </li> </ul> <ul> <li> <p>git_hash: Git hash for HEPData4ML.</p> </li> <li> <p>timestamp: Timestamp for when the dataset was created (local).</p> </li> <li> <p>timestamp_string_utc: Timestamp for when the dataset was created (in UTC).</p> </li> <li> <p>pythia_config: Configuration file passed to Pythia8, by HEPData4ML. Defines the process.</p> </li> <li> <p>pythia_random_seed: Random seed passed to Pythia8, for initializing its random number generator.</p> </li> <li> <p>unique_id: A unique string for identifying the generated set of events (one run of the HEPData4ML tool).</p> </li> <li> <p>unique_id_short: Similar to `unique_id`, but shortened.</p> </li> </ul> <p><strong>Citations</strong></p> <p><strong>[1]</strong>: J. T. Offermann, X. Liu, and T. Hoffman, <a href="https://github.com/janTOffermann/HEPData4ML">HEPData4ML</a> (2023).</p> <p><strong>[2]</strong>: J. T. Offermann, A. Bogatskiy, and T. Hoffman, <a href="https://doi.org/10.5281/zenodo.7338117">Top Quark Momentum Reconstruction Dataset</a> (2022).</p> <p><strong>[3]</strong>: C. Bierlich and others, <a href="https://arxiv.org/abs/2203.11601"><em>A Comprehensive Guide to the Physics and Usage of PYTHIA 8.3</em></a>, (2022).</p> <p><strong>[4]</strong>: J. de Favereau, C. Delaere, P. Demin, A. Giammanco, V. Lemaître, A. Mertens, and M. Selvaggi, <a href="https://arxiv.org/abs/1307.6346"><em>DELPHES 3, A Modular Framework for Fast Simulation of a Generic Collider Experiment</em></a>, JHEP <strong>02</strong>, 057 (2014).</p> <p><strong>[5]</strong>: M. Cacciari, G. P. Salam, and G. Soyez, <a href="https://arxiv.org/abs/1111.6097"><em>FastJet User Manual</em></a>, Eur. Phys. J. C <strong>72</strong>, 1896 (2012).</p> <p><strong>[6]</strong>: D. E. Kaplan, K. Rehermann, M. D. Schwartz, and B. Tweedie, <a href="https://arxiv.org/abs/0806.0848"><em>Top Tagging: A Method for Identifying Boosted Hadronically Decaying Top Quarks</em></a>, Phys. Rev. Lett. <strong>101</strong>, 142001 (2008).</p>
Plant abundance data from the 2016 Chimney Tops 2 Fire in Gatlinburg, TN
We surveyed understory plant communities during the second growing season (2018) after the 2016 fire at sites in WUI near Gatlinburg (hereafter “exurban”) and in GSMNP (hereafter "natural"). We chose sites based on dominant forest vegetation type (available for GSMNP) and elevation to minimize the potential confounding effects of these variables. We used stratified random sampling in ESRI ArcMap to select 18 sites, nine in natural locations and nine in exurban locations; within the location type we randomly selected sites to represent fire severity categories, three no burn, three low/medium burn, and three high burn. In the field, we randomly selected two 1x1 m permanent plots to survey and identify each individual plant in the understory to at least genus level, and to count the number of individuals for each taxon. We generated taxa lists and count of individuals per taxa by aggregating plot level records by month. Individuals that could not be identified to at least genus level due to immature characteristics or herbivore damage were assigned observational taxonomic unit numbers (OTUs).
Bottom-up meets top-down: Leaf litter inputs influence predator-prey interactions in wetlands, 2011.
While the common conceptual role of resource subsidies is one of bottom-up nutrient and energy supply, inputs can also alter the structural complexity of environments. This can further impact resource flow by providing refuge for prey and decreasing predation rates. However, the direct influence of different organic subsidies on predator–prey dynamics is rarely examined. In forested wetlands, leaf litter inputs are a dominant energy and nutrient resource and they can also increase benthic surface cover and decrease water clarity, which may provide refugia for prey and subsequently reduce predation rates. In outdoor mesocosms, we investigated how inputs of leaf litter that alter benthic surface cover and water clarity influence the mortality and growth of gray treefrog tadpoles (Hyla versicolor) in the presence of free-swimming adult newts (Notophthalmus viridiscens), which are visual predators. To manipulate surface cover, we added either oak (Quercus spp.) or red pine (Pinus resinosa) litter and crossed these treatments with three levels of red maple (Acer rubrum) litter leachate to manipulate water clarity. In contrast to our predictions, benthic surface cover had no effect on tadpole survival while darkening the water caused lower survival. In addition, individual tadpole mass was lowest in the high maple leachate treatments, suggesting an interaction between bottom-up effects of leaf litter and topdown effects of predation risk that altered mortality and growth of tadpoles. Our results indicate that realistic changes in forest tree composition, which cause concomitant changes in litter inputs to wetlands, can substantially alter community interactions.
New Physics Mining at the Large Hadron Collider: top pair production
<p><span class="math-tex">\(t \bar t\)</span> background events reconstructed by inclusive single-muon selection.</p> <p>Events are represented as an array of physics-motivated high-level features.</p> <p>Details are given in https://arxiv.org/abs/1811.10276</p>
Top: puffs of cornstarch reveal dense and varied tiny cryptic webs in the Gaoligongshan. Shown here upper left to lower right are a symphytognathid Patu jidanweishi sp. n., a mysmenid Gaoligonga changya gen. n., sp. n., and an unidentified linyphiid. Bottom: this misty mountain landscape at QiQi is typical of the Gaoligongshan in The symphytognathoid spiders of the Gaoligongshan, Yunnan, China (Araneae: Araneoidea): Systematics and diversity of micro-orbweavers
Top: puffs of cornstarch reveal dense and varied tiny cryptic webs in the Gaoligongshan. Shown here upper left to lower right are a symphytognathid Patu jidanweishi sp. n., a mysmenid Gaoligonga changya gen. n., sp. n., and an unidentified linyphiid. Bottom: this misty mountain landscape at QiQi is typical of the Gaoligongshan
TOP-100 DOCKING POSES OF FDA APPROVED DRUGS AND DRUGS IN CLINICAL INVESTIGATION AT SARS-CoV2 MAIN PROTEASE
<p>7922 compounds were downloaded from NPC database (https://tripod.nih.gov/npc/). In order<br> to eliminate the non-specific binders, some criteria including molecular weight, between 100 to<br> 1000 g/mol; number of rotatable bonds, <100; number of atoms, between 10 and 100; number<br> of aliphatic and aromatic rings, <10; number of hydrogen-bond acceptor and donors, <10 were<br> set and as a result the total number of compounds was decreased to 6654. These ligands were<br> prepared using LigPrep module of Maestro at neutral pH (LigPrep, Schrodinger v.2017). In<br> molecular docking, we used following protein structure: SARS-CoV2 Main Protease, (PDB, 6LU7). The protein<br> was prepared using Protein Preparation module of Maestro. PROPKA was used for<br> determination of protonation states of amino acid residues. Restrained minimization was<br> performed with OPLS3 force field for the protein using 0.3 Å heavy atom convergence.<br> Docking was performed with Glide/SP using default settings. Top-100 docking poses were provided.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.