Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,140

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,140 results for “TOPS”

Learn how ShareScore rates datasets ↗
zenodo48/100

Top quark pair production at the LHC, with all hadronic resolved decays for solving event combinatorics

<p><strong>R&amp;D Datasets for solving event combinatorics in all hadronic top quark pair events at the LHC.</strong></p> <p>Used in the development of Topographs: Topological Reconstruction of Particle Physics Processes using Graph Neural Networks</p> <p>&nbsp;</p> <p>The datasets contain 5.8M ttbar events in the all hadronic decay channel, with jets matched to the truth partons in the top quark decays.</p> <p>&nbsp;</p> <p><strong>Event generation</strong></p> <ul> <li>Centre of Mass energy: 13 TeV</li> <li>MC Generator: MadGraph5_aMC@NLO v3.1.0, with MadSpin modelling the decays of the top quarks and W bosons.</li> <li>Parton Shower: Pythia v .243</li> <li>Detector response: Delphes v3.4.2 using ATLAS-like geometry</li> <li>Jets reconstructed with anti-kt algorithm, R=0.4, using FastJet</li> <li>b-Tagging corresponds to inclusive 70% b-jet efficiency</li> </ul> <p><strong>Event selection and truth matching</strong></p> <ul> <li>All events are required to have at least six reconstructed jets and exactly zero leptons (electrons or muons)</li> <li>Partons are matched to jets using <span class="math-tex">\(\Delta R\)</span> matching, with <span class="math-tex">\(\Delta R &lt; 0.4\)</span></li> <li>Events with partons matched to multiple jets or jets to multiple partons are discarded</li> <li>Up to 16 jets are stored per event</li> </ul> <p>In the training dataset 1,340,000 events have all partons from the ttbar decays matched to jets.</p> <p>In the validation dataset, 71,000 events have all partons from the ttbar decays matched to jets.</p> <p>In the testing dataset 76,000 events have all partons from the ttbar decays matched to jets.</p> <p><strong>Dataset format</strong></p> <p>The dataset is in h5 format and the key &#39;delphes&#39; has the following numpy arrays:</p> <pre><code>jets (16), jets_indices (16), matchability, nbjets, njets, partons (10) </code></pre> <p>Jets structured numpy array per event:</p> <ul> <li> <pre><code>(pt, eta, phi, energy, is_tagged)</code></pre> </li> </ul> <p>Jets_indices:</p> <ul> <li>Integer corresponding to the parton the jet is matched to</li> <li>From 0 to 5: b1 W1j1 W1j2 b2 W2j1 W2j2 (1= from top, 2=from antitop)</li> <li>-1 indicates not matched to a parton</li> <li>Properties of matched partons can be obtained from the partons array</li> </ul> <p>matchability:</p> <ul> <li>Which partons are matched to jets in event</li> <li>Binary representation with bits corresponding to each parton (length 6) 0b111111</li> <li>From left to right: b1 W1j1 W1j2 b2 W2j1 W2j2</li> <li>0b111000 (56) is one top fully matched, 0b000111 (7) is the other top fully matched, 0b111111 (63) is both tops fully matched</li> </ul> <p>njets, nbjets:</p> <ul> <li>How many jets/bjets in event</li> </ul> <p>partons:</p> <ul> <li>List of truth particles from ttbar decay: tops, Ws, quarks, ordered by top quark and its decays followed by anti-top and its decays</li> <li> <pre><code>PDGID, pt, eta, phi, mass</code></pre> </li> </ul> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo48/100

Regional scale surface of the top of the Variscan basement in some sector of Italy - Supplementary material

<p>The dataset represent the Supplementary material of thew manuscript entitled &quot;Map of the top of the Variscan basement in some sectors of Italy&quot; now under revision.</p> <p>The Supplementary material consist of 9&nbsp;files:</p> <ul> <li>input data: <ul> <li>dataset_CROP.csv</li> <li>deep_wells.csv</li> <li>domains.geojson</li> <li>thrusts_2.geojson</li> <li>INA_data_point.csv</li> </ul> </li> <li>output data: <ul> <li>INA_depth_1km.csv</li> <li>ONA_OA_ISA_AF_depth_5km.csv</li> <li>INA_contour.geojson</li> <li>ONA_OA_ISA_AF_contour.geojson</li> </ul> </li> </ul>

opencc-by-4.0Apr 2023View details →
edi48/100

Data from "Evaluating top-down, bottom-up, and environmental drivers of pelagic food web dynamics along an estuarine gradient"

Synthesized fish, benthic invertebrate, and water quality dataset used for analysis in: Rogers, T., S. Bashevkin, C. Burdi, D. Colombano, P. Dudley, B. Mahardja, L. Mitchell, S. Perry, and P. Saffarinia. 2022. Evaluating top-down, bottom-up, and environmental drivers of pelagic food web dynamics along an estuarine gradient. preprint, EcoEvoRxiv. https://doi.org/10.32942/X2MK5Z

openCC (other)Jan 2023View details →
edi48/100

MCR LTER: Coral Reef: Priority effects in coral-macroalgae interactions can drive alternate community paths in the absence of top-down control, data for Adam 2022 Ecology

These data were generated in support of the manuscript: Adam TC, Holbrook SJ, Burkepile DE, Speare KE, Brooks AJ, Ladd MC, Shantz AA, Thurber RLV, and Schmitt RJ, Ecology The outcomes of species interactions can vary greatly in time and space with the outcomes of some interactions determined by priority effects. On coral reefs, benthic algae rapidly colonize the disturbed substrate. In the absence of top-down control from herbivorous fishes, these algae can inhibit the recruitment of reef-building corals, leading to a persistent phase shift to a macroalgae-dominated state. Yet, corals may also inhibit colonization by macroalgae, and thus the effects of herbivores on algal communities may be strongest following disturbances that reduce coral cover. Here, we report results from experiments conducted on the fore reef of Moorea, French Polynesia, where we: 1) tested the ability of macroalgae to invade coral-dominated and coral-depauperate communities under different levels of herbivory, 2) explored the ability of juvenile corals (Pocillopora spp.) to suppress macroalgae, and 3) quantified the direct and indirect effects of fish herbivores and corallivores on juvenile corals. We found that macroalgae proliferated when herbivory was low but only in recently disturbed communities where coral cover was also low. When coral cover was < 10%, macroalgae increased 20-fold within one year under reduced herbivory conditions relative to high herbivory controls. Yet, when coral cover was high (50%), macroalgae were suppressed irrespective of the level of herbivory despite ample space for algal colonization. Once established in communities with low herbivory and low coral cover, macroalgae suppressed recruitment of coral larvae, reducing the capacity for coral replenishment. However, when we experimentally established small juvenile corals (2 cm diameter) following a disturbance, juvenile corals inhibited macroalgae from invading local neighborhoods, even in the absence of herbivore

openCC (other)Jan 2025View details →
zenodo44/100

The top performer: towards optimized parameters for Reduced graphene oxide uniformity by Spin coating

<p>This dataset contains the raw data used for the publication:</p> <p>-------------------------------------------------------------------------------------------------------------------------------------------------------<br> &quot;The top performer: towards optimized parameters for Reduced graphene oxide uniformity by Spin coating&quot;<br> by C. Reiner-Rozman, R. Hasler, J. Andersson, T. Rodrigues, A. Bozdogan and P. Aspermair<br> --------------------------------------------------------------------------------------------------------------------------------------------------------</p> <p><br> It consists of the SEM images (in .tif format) and the determined surface coverages (in .dat format) as well as the measured electrical data (in .dat format) of the prepared graphene field-effect transistor chips. Headers/information in the data files are in English. When using this data in any form please refer to the above-mentioned publication.</p> <p>The data is structured according to the figures of the paper. Each folder contains the data relevant to validate the results presented in the respective figure of the publication. The files are labeled according to the following description:</p> <p>&quot;measurement-type&quot;_&quot;chip-number&quot;_&quot;GO-concentration&quot;_&quot;spin-coating speed&quot;</p> <p>&quot;measurement-type&quot;:&nbsp;&nbsp; &nbsp;SEM, IDVG, baseline<br> &quot;chip-number&quot;: an increasing number of fabricated device (only used when needed)<br> &quot;GO-concentration&quot;:&nbsp;&nbsp; &nbsp;143/214/285 &micro;g/mL of graphene oxide (GO) in solution<br> &quot;spin-coating speed&quot;:&nbsp;&nbsp; &nbsp;in rpm</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

Películas en top 30 de varios géneros de filmaffinity

<p>El dataset contiene diversa informaci&oacute;n obtenida de la p&aacute;gina de filmaffinity (propietaria en todo caso de la informaci&oacute;n) de las pel&iacute;culas que formas el top 30 de varios g&eacute;neros cinematogr&aacute;ficos.</p> <p>El objetivo final de este dataset es dar respuesta a la pr&aacute;ctica 1 de la asignatura Tipolog&iacute;a y ciclo de vida de los datos, impartida dentro del M&aacute;ster en Ciencia de Datos de la UOC.</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Datasets to article "Selection history alters attentional filter settings persistently and beyond top-down control"

<p>Single-Subject Behavioral and ERP mean amplitude data for Experiments 1 to 3.</p>

opencc-by-4.0Feb 2017View details →
zenodo44/100

Top selling games in Steam platform

<p>The dataset is extracted from the list of best-selling games available in Spanish on the Steam platform. It contains information related to the content of the game, to its development and marketing (price and discounts), as well as different metrics derived from user opinions.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Top 120+ popular movies 2023 from RottenTomatoes

<p>[ENG] The file contains information about popular movies on the webpage Rotten Tomatoes. We suggest using R or Python to work with the dataset. The dataset has not been cleaned, so spaces, NaN values and unmatched variables types may be present.</p><p>[CAT] El fitxer conté informació sobre pel·lícules populars a la pàgina web Rotten Tomatoes. Suggerim utilitzar R o Python per a treballar amb el conjunt de dades. El conjunt de dades no s'ha netejat, de manera que els espais, els valors NaN i els tipus de variables que no coincideixn poden estar presents.</p><p>[ESP] El fichero contiene información sobre películas populares en la página web Rotten Tomatoes. Sugerimos utilizar R o Python para trabajar con el conjunto de datos. El conjunto de datos no se ha limpiado, de forma que los espacios, los valores NaN y los tipos de variables que no coincidan pueden estar presentes.</p><p>&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

User Feedback Dataset from the Top 15 Downloaded Mobile Applications

<p>This dataset comprises user feedback data collected from 15 globally acclaimed mobile applications, spanning diverse categories. The included applications are among the most downloaded worldwide, providing a rich and varied source for analysis. <i><strong>The dataset is particularly suitable for Natural Language Processing (NLP) applications</strong></i>, such as text classification and topic modeling.</p><p><strong>List of Included Applications:</strong></p><ul><li>TikTok</li><li>Instagram</li><li>Facebook</li><li>WhatsApp</li><li>Telegram</li><li>Zoom</li><li>Snapchat</li><li>Facebook Messenger</li><li>Capcut</li><li>Spotify</li><li>YouTube</li><li>HBO Max</li><li>Cash App</li><li>Subway Surfers</li><li>Roblox</li><li>Data Columns and Descriptions:</li></ul><p><strong>Data Columns and Descriptions:</strong></p><ul><li>review_id: Unique identifiers for each user feedback/application review.</li><li>content: User-generated feedback/review in text format.</li><li>score: Rating or star given by the user.</li><li>TU_count: Number of likes/thumbs up (TU) received for the review.</li><li>app_id: Unique identifier for each application.</li><li>app_name: Name of the application.</li><li>RC_ver: Version of the app when the review was created (RC).</li></ul><p><strong>Terms of Use:</strong></p><p>This dataset is open access for scientific research and non-commercial purposes. Users are required to acknowledge the authors' work and, in the case of scientific publication, cite the most appropriate reference:</p><p>M. H. Asnawi, A. A. Pravitasari, T. Herawan, and T. Hendrawati, "The Combination of Contextualized Topic Model and MPNet for User Feedback Topic Modeling," in IEEE Access, vol. 11, pp. 130272-130286, 2023, doi: <a href="https://doi.org/10.1109/ACCESS.2023.3332644">10.1109/ACCESS.2023.3332644</a>.</p><blockquote><p>Researchers and analysts are encouraged to explore this dataset for insights into user sentiments, preferences, and trends across these top mobile applications. If you have any questions or need further information, feel free to contact the dataset authors.</p></blockquote>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Rohdaten zu den Ergebnissen der ZKI Top Trends-Umfrage des ZKI-Arbeitskreises Strategie und Organisation für das Jahr 2022

<p>Der Arbeitskreis Strategie und Organisation des ZKI-Vereins f&uuml;hrt eine j&auml;hrliche Umfrage zu den wichtigsten Themen und Trends von IT-Einrichtungen aus Hochschulen und Forschungseinrichtungen durch. Die Umfrageergebnisse sollen dabei helfen, wichtige Entwicklungen, Themen und Best Practices im Blick zu behalten und bei den umfangreichen Themenfeldern der Digitalisierung und der rasanten Erneuerung von Technologien Schritt zu halten bzw. auch Inspiration f&uuml;r die weitere Ausgestaltung an der eigenen Einrichtung zu gewinnen.</p> <p>Die Kernumfrage adressiert die wichtigsten Themen und Ver&auml;nderungen im Umfragejahr in standardisierter Form. Dar&uuml;ber hinaus werden in jedem Jahr individuelle Schwerpunkte abgefragt, die viele Einrichtungen besch&auml;ftigen. Im Jahr 2022 waren die Schwerpunktfragen &uuml;ber die Kernumfrage hinaus:</p> <ul> <li>Hat die <strong>Nutzung von externen Cloud-Angeboten</strong> w&auml;hrend der Pandemie eher zugenommen oder eher abgenommen?</li> <li>In welchem <strong>Umfang</strong> setzen Sie <strong>Cloud-Technologien</strong> ein?</li> <li>Welche <strong>Ma&szlig;nahmen</strong> haben Sie im Bereich <strong>&quot;Digitale Souver&auml;nit&auml;t&quot;</strong> getroffen?</li> <li>Welche <strong>spezifischen Aspekte</strong> sehen Sie f&uuml;r Hochschulen und Forschungseinrichtungen im Bereich <strong>Digitaler Souver&auml;nit&auml;t</strong>?</li> <li>Welche <strong>Auswirkungen</strong> sehen Sie durch Corona f&uuml;r die <strong>Arbeitsplatzgestaltung</strong>?</li> <li>Welche <strong>Tools und Mechanismen</strong> haben Ihnen dabei geholfen, <strong>Zusammenarbeit</strong> und Team-Geist <strong>trotz weniger Pr&auml;senz</strong> zu erhalten?</li> <li>Welchen <strong>Prozentsatz</strong> an <strong>Home-Office</strong> sehen Sie zuk&uuml;nftig <strong>f&uuml;r die IT-Einrichtung</strong> Ihrer Hochschule?</li> <li>Wie sehen Sie die <strong>Rolle der IT</strong> an Ihrer Hochschule <strong>nach Corona</strong>?</li> </ul> <p>Weiterhin wird nach den Modellen zur IT-Governance gefragt.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Proton collision producing top pair, decaying hadronically via bottom quarks and W bosons

<p>This dataset contains the matrix element calculations for&nbsp;10,000 events of `p p &gt; t t~ , (t &gt; b W+) , (t~ &gt; b~ W-)`, as produced by MadGraph, without showering or hadronisation,&nbsp;and applying no cuts.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Criptomonedes_top_capitalització_mercat

<p>El dataset cont&eacute; la informaci&oacute; hist&ograve;rica durant l&#39;&uacute;ltim any, de les 10 criptomonedes amb m&eacute;s capitalitzaci&oacute; durant el&nbsp;dia 11/04/2022.&nbsp;</p> <p>En aquest dataset, s&#39;observen les columnes seg&uuml;ents:</p> <p>- Criptomoneda: Nom de la criptomoneda</p> <p>- Date: Dia del registre</p> <p>- Price: Preu del tancament de la criptomoneda.</p> <p>- Open: Preu d&#39;obertura de la criptomoneda.</p> <p>-- High: Preu m&agrave;xim de la criptomoneda.</p> <p>- Low: Preu m&iacute;nim de la criptomoneda.</p> <p>- Vol. : Volum de cotitizaci&oacute; de la criptomoneda.</p> <p>- Change%: Canvi diari del valor de la criptomoneda en %.</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

ECHAM6 verification without top 10 layers synchronization and with less physics schemes

<p>Verification of ECHAM6 nudging module when ECHAM6 has synchronized to its own outputs.</p> <p>Some physics schemes are switched off as the following namelist was used.</p> <p>&nbsp;</p> <p>&amp;physctl<br> &nbsp; LCOVER &nbsp; &nbsp; &nbsp; = .false.<br> &nbsp; lphys=.true.<br> &nbsp; lconv=.false.<br> &nbsp; lgwdrag=.false.<br> &nbsp; lrad=.true.<br> &nbsp; lsurf=.false.<br> &nbsp; lvdiff=.false.<br> &nbsp; lcond=.false.<br> /</p> <p>&nbsp;</p> <p>Files and description:</p> <ul> <li>ndg_197102.nc <ul> <li>ECHAM6 grb output converting to netcdf format as a reference data for ECHAM6 nudging module</li> </ul> </li> <li>diff_t_197102_gp_rmse.nc <ul> <li>Spatial root mean square errors on all levels and at all outputs (6 hourly).</li> <li>RMSE between ECHAM6 output (echam6_nudging_grb_T63_197102.01_echam) and reference data (ndg_197102.nc)</li> </ul> </li> <li>namelist.echam <ul> <li>namelist used to run ECHAM6 nudging case, in which nudging is expected to only apply to layer 11 to 47.</li> </ul> </li> <li>Four ECHAM6 outputs when nudging module is switched on <ul> <li>echam6_nudging_grb_T63_197102.01_echam</li> <li>echam6_nudging_grb_T63_197102.01_echam.codes</li> <li>echam6_nudging_grb_T63_197102.01_nudg</li> <li>echam6_nudging_grb_T63_197102.01_nudg.codes</li> </ul> </li> <li>Two log files generated by ECHAM6 executable <ul> <li>r_nudging_err</li> <li>r_nudging_log</li> </ul> </li> </ul>

opencc-by-4.0Apr 2022View details →
zenodo44/100

ECHAM6 verification without top 10 layers synchronization

<p>Verification of ECHAM6 nudging module when ECHAM6 has synchronized to its own outputs.</p> <p>Files and description:</p> <ul> <li>ndg_197102.nc <ul> <li>ECHAM6 grb output converting to netcdf format as a reference data for ECHAM6 nudging module</li> </ul> </li> <li>diff_t_197102_gp_rmse.nc <ul> <li>Spatial root mean square errors on all levels and at all outputs (6 hourly).</li> <li>RMSE between ECHAM6 output (echam6_nudging_grb_T63_197102.01_echam) and reference data (ndg_197102.nc)</li> </ul> </li> <li>namelist.echam <ul> <li>namelist used to run ECHAM6 nudging case, in which nudging is expected to only apply to layer 11 to 47.</li> </ul> </li> <li>Four ECHAM6 outputs when nudging module is switched on <ul> <li>echam6_nudging_grb_T63_197102.01_echam</li> <li>echam6_nudging_grb_T63_197102.01_echam.codes</li> <li>echam6_nudging_grb_T63_197102.01_nudg</li> <li>echam6_nudging_grb_T63_197102.01_nudg.codes</li> </ul> </li> <li>Two log files generated by ECHAM6 executable <ul> <li>r_nudging_err</li> <li>r_nudging_log</li> </ul> </li> </ul>

opencc-by-4.0Apr 2022View details →
zenodo44/100

JavaScript Libraries From Top 1 Million Sites

<p>Scraped data from top 1 million domains as reported by Majestic 1 Million on June 5th, 2022. The homepage of each domain is scraped and all encountered javascript script source URLs are extracted.</p> <p>You can find the source code at <a href="https://github.com/get-set-fetch/scraper/tree/main/datasets">github.com/get-set-fetch/scraper</a> and detailed documentation at <a href="https://getsetfetch.org">getsetfetch.org</a>.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

The top 100 most influential articles in olfactory disorder: a bibliometric analysis

<p><strong>Introduction</strong><strong>:</strong>&nbsp;Studies on olfactory disorder have varied considerably in content and field over the past four decades. Influential publications of olfactory disorder have not been analyzed quantitatively. This study aimed to analyze the top 100 most influential articles in olfactory disorder through bibliometrics.</p> <p><strong>Methods:</strong>&nbsp;The researchers searched the Web of Science Core Collection for publications from 1980 to August 31, 2021. The top 100 highly cited articles were screened in accordance with the inclusion and exclusion criteria. The journal impact factor and the SCImago Journal Rank indicator were searched for the journal in which the top-cited articles were published. The researchers analyzed the data via Microsoft Excel and SPSS 24.0 and used VOSviewer software for data visualization.</p> <p><strong>Results:</strong>&nbsp;The total citations of the top 100 articles have increased exponentially, especially in the past 2 years. The journal LARYNGOSCOPE contributed the most influential papers (n=6). The United States published the most top 100 articles (n=52), followed by Germany and the United Kingdom. The University of Pennsylvania published the most influential studies, with a total of 3,852 citations. Richard L. Doty and Thomas Hummel contributed the most influential literature, and the top 100 studies also cited their research frequently. The most important articles in the field of OD mainly focused on COVID-19, Parkinson&#39;s disease and olfactory tests.</p> <p><strong>Conclusion:</strong>&nbsp;Neurodegenerative diseases and COVID-19-related olfactory disorders have been considered main topics over the past 40 years. This study identified the most influential articles for olfactory disorder researchers, providing guidance for their research.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

High resolution spectra of the spinning-top Be star Achernar

<p>Achernar, the closest and brightest classical Be star, presents rotational flattening, gravity darkening, occasional emission lines due to a gaseous disk, and an extended polar wind. It is also a member of a close binary system with an early A-type dwarf companion.&nbsp;We aim to determine the orbital parameters of the Achernar system and to estimate the physical properties of the components.&nbsp;We monitored the relative position of Achernar B using a broad range of high angular resolution instruments of the VLT/VLTI over a period of 13 years (2006-2019). These astrometric observations are complemented with a series of more than 700 optical spectra for the period from 2003 to 2016. The present dataset contains the high resolution spectra of Achernar that were included in our study. They were&nbsp;collected using the BESO, BeSS, CHIRON, CORALIE, FEROS, HARPS, PUCHEROS, and UVES instruments. The spectra&nbsp;are provided in the form of standard FITS files, with the continuum flux normalized to unity.</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

How do Google News' top 100 sources visually represent the data centres' energy footprint?

<p><strong>By querying &quot;data centres&#39; energy footprint&quot; on Google News in incognito mode, the candidate has selected and mapped the top 100 results according to the ranking on May 15, 2022.&nbsp;</strong></p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Top quark pair production at the LHC

<p><strong>R&amp;D Datasets containing all-hadronic, semi-leptonic and di-leptonic top quark pair events at the LHC.</strong></p> <p>Used in the development of PIPPIN: Particles Into Particles with Permutation Invariant Network</p> <p>&nbsp;</p> <p>The datasets contain a total of 40M ttbar events in the all-hadronic, semi-leptonic and di-leptonic decay channels, with jets matched to the truth partons in the top quark decays.</p> <p>&nbsp;</p> <p><strong>Event generation</strong></p> <ul> <li>Centre of Mass energy: 13 TeV</li> <li>MC Generator: PYTHIA v8.307</li> <li>Parton Shower and Hadronisation: PYTHIA v8.307</li> <li>Detector response: Delphes v3.4.2 using ATLAS-like geometry</li> <li>Jets reconstructed with anti-kt algorithm, R = 0.4, using FastJet</li> </ul> <p><strong>Event selection and truth matching</strong></p> <ul> <li>All events are required to have between 2 and 16 reconstructed jets</li> <li>Jets are required to fall within |&eta;| &lt; 2.5 and to have a minimum pT &gt; 25 GeV</li> <li>Leptons are required to fall within |&eta;|&lt;2.5 and to have a minimum pT &gt; 15 GeV</li> <li>Partons are matched to jets using&nbsp;<span><span><span><span><span>&Delta;</span><span>R</span></span></span></span></span>&nbsp;matching, with&nbsp;<span><span><span><span><span>&Delta;</span><span>R </span><span>&lt; </span><span>0.4</span></span></span></span></span></li> <li>Events with partons matched to multiple jets or jets to multiple partons are discarded</li> </ul> <p>The training dataset contains 37M events.<br>The validation dataset contains 0.8M events.<br>The testing dataset contains 2.4M events.</p> <p>&nbsp;</p> <p><strong>Dataset format</strong></p> <p>The dataset is in HDF5 format and the key 'delphes' contains the following numpy arrays:</p> <p>Truth information (parton-level):</p> <ul> <li><code>truth_leptons</code>, <code>truth_neutrinos</code>, <code>truth_quarks</code>: Truth level information of the final state partons <ul> <li><em>keys:</em> <code>PDGID</code>, <code>pt</code>, <code>eta</code>, <code>phi</code>, <code>mass</code></li> </ul> </li> <li><code>truth_particles</code>: Truth level information of the final state partons and the intermediate particles <ul> <li><em>keys:</em> <code>PDGID</code>, <code>pt</code>, <code>eta</code>, <code>phi</code>, <code>mass</code></li> </ul> </li> </ul> <p>Reconstructed information (detector-level):</p> <ul> <li><code>leptons</code>: The zero padded reconstructed leptons (0 to 2), ordered by decay channel <ul> <li><em>keys:</em> <code>pt</code>, <code>eta</code>, <code>phi</code>, <code>energy</code>, <code>charge</code>, <code>type</code></li> </ul> </li> <li><code>MET</code>: The missing transverse energy <ul> <li><em>keys:</em> <code>MET</code>, <code>phi</code></li> </ul> </li> <li><code>jets</code>: The zero padded reconstructed jets (2 to 16), ordered by pT <ul> <li><em>keys:</em> <code>pt</code>, <code>eta</code>, <code>phi</code>, <code>energy</code>, <code>is_tagged</code>, <code>is_tau</code></li> </ul> </li> </ul> <p>Miscellaneous information:</p> <ul> <li><code>decay_channel</code>: The decay channel of the event <ul> <li>0b00 for all-hadronic, 0b01 for semi-leptonic (from Top), 0b10 for semi-leptonic (from Anti-Top), 0b11 for di-leptonic</li> </ul> </li> <li><code>matchability</code>: Which partons are matched to a reconstructed object <ul> <li>Binary representation with bits corresponding to each parton (length 6) 0b111111</li> <li>From left to right:&nbsp;b1, q1<sub>W1</sub>, q2<sub>W1</sub>, b2, q1<sub>W2</sub>,&nbsp;q2<sub>W2 </sub>(b1/W1 = from Top, b2/W2 = from Anti-Top)</li> <li>0b111000 means Top fully matched, 0b000111 means Anti-Top fully matched, 0b111111 means both Tops fully matched, etc.</li> </ul> </li> <li><code>jet_indices</code>: Integer corresponding to the parton a jet is matched to <ul> <li>From 0 to 5:&nbsp;b1, q1<sub>W1</sub>, q2<sub>W1</sub>, b2, q1<sub>W2</sub>,&nbsp;q2<sub>W2 </sub>(b1/W1 = from Top, b2/W2 = from Anti-Top)</li> <li>-1 indicates not matched to a parton</li> </ul> </li> <li><code>nleptons</code>,&nbsp;<code>njets</code>,&nbsp;<code>nbjets</code>: How many leptons, jets, b-jets in the event</li> </ul>

opencc-by-4.0Jun 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record