Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

82

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

82 results for “arxiv”

Learn how ShareScore rates datasets ↗
zenodo52/100

Dataset for: Mudrik, N., & Charles, A. S. (2022). Multi-Lingual DALL-E Storytime. arXiv preprint arXiv:2212.11985.

<p>This dataset represents the comprehensive collection of data generated during the study presented in the paper available at https://arxiv.org/abs/2212.11985.</p> <p>If your research incorporates this data and results in a publication - Please cite both the dataset and the paper.</p>

opencc-by-4.0Jun 2023View details →
zenodo48/100

arXiv:1704.05309: Supporting dataset

<p>This is a supporting dataset that accompanies arXiv:1704.05309. It consists of databases and parameter files used to generate the numerical results presented in the paper.</p> <p>This deposit contains the following files:</p> <p><strong>Databases</strong></p> <p>These are SQLite databases produced by the Sussex LSSEFT tool (git revision 977e5b03) containing the one-loop SPT results and counterterms used to obtain EFT predictions for the real-space and redshift-space power spectra.</p> <ul> <li><strong>Planck2015_CAMB@z=0_k10-full.sqlite</strong><br> This contains results for a Planck2015 cosmology. The initial power spectrum is generated using linear theory by CAMB at z = 50, using the parameter file documented below. The tree-level component is generated using a CAMB z = 0 power spectrum, also computed using linear theory, for which we also supply a CAMB parameter file.</li> <li><strong>MDR1_CAMB@z=0_k10-full.sqlite</strong><br> This contains results for a cosmology matching the MultiDark MDR1 simulation. The initial power spectrum is generated using linear theory by CAMB at z = 50 and the tree-level component is generated using a CAMB z = 0 power spectrum as above. The corresponding parameter files are supplied; see below.</li> <li><strong>MDR1_CAMB@z=0_k10-EdS.sqlite</strong><br> This contains results matching the <strong>MDR1_CAMB@z=0_k10-full.sqlite </strong>database, but with growth functions computed using the standard SPT Einstein-de Sitter approximation.</li> </ul> <p><strong>CAMB parameter files</strong></p> <p>Power spectra were produced using the November 2015 release of CAMB.</p> <ul> <li><strong>Planck2015_linear_init_params.ini</strong>, <strong>Planck2015_linear_final_params.ini</strong><br> Parameter files to generate initial (z = 50) and final (z = 0) linear power spectra in the Planck2015 cosmology.</li> <li><strong>Planck2015_nonlinear_params.ini</strong><br> Parameter file to generate a <em>nonlinear </em>z = 0 power spectrum using the HALOFIT prescription for the Planck2015 cosmology. This is used to renormalize the EFT power spectrum.</li> <li><strong>MDR1_linear_init_params.ini</strong>, <strong>MDR1_linear_final_params.ini</strong><br> Parameter files to generate initial (z = 50) and final (z = 0) linear power spectra in the MDR1 cosmology.</li> </ul> <p><strong>gevolution parameter files</strong></p> <p>N-body simulations were performed using gevolution 1.1.</p> <ul> <li><strong>gevolutionsettings.ini</strong><br> This is a settings file for gevolution that reproduces our 1024^3-particle, (2000 Mpc/h)-side simulation volume, used to compute the real-space power spectrum and the l = 0, 2, 4 multipoles of the redshift-space power spectrum.</li> </ul>

opencc-by-4.0Apr 2017View details →
zenodo48/100

Citation data of arXiv eprints and the associated quantitatively-and-temporally normalised impact metrics

<p><strong>Data collection</strong></p> <p>This dataset contains information on the eprints posted on arXiv from its launch in 1991 until the end of 2019 (1,589,006 unique eprints), plus the data on their citations and the associated impact metrics. Here, eprints include preprints, conference proceedings, book chapters, data sets and commentary, i.e. every electronic material that has been posted on arXiv.&nbsp;</p> <p>The content and metadata of the arXiv eprints were retrieved from the arXiv API (https://arxiv.org/help/api/) as of 21st January 2020, where the metadata included data of the eprint&rsquo;s title, author, abstract, subject category and the arXiv ID (the arXiv&rsquo;s original eprint identifier). In addition, the associated citation data were derived from the Semantic Scholar API (https://api.semanticscholar.org/) from 24th January 2020 to 7th February 2020, containing the citation information in and out of the arXiv eprints and their published versions (if applicable). Here, whether an eprint has been published in a journal or other means is assumed to be inferrable, albeit indirectly, from the status of the digital object identifier (DOI) assignment. It is also assumed that if an arXiv eprint received&nbsp;<em>c</em><sub>pre</sub>&nbsp;and&nbsp;<em>c</em><sub>pub</sub>&nbsp;citations until the data retrieval date (7th February 2020) before and after it is assigned a DOI, respectively, then the citation count of this eprint is recorded in the Semantic Scholar dataset as&nbsp;<em>c</em><sub>pre</sub>&nbsp;+&nbsp;<em>c</em><sub>pub</sub>. Both the arXiv API and the Semantic Scholar datasets contained the arXiv ID as metadata, which served as a key variable to merge the two datasets.</p> <p>The classification of research disciplines is based on that described in the arXiv.org website (https://arxiv.org/help/stats/2020_by_area/). There, the arXiv subject categories are aggregated into several disciplines, of which we restrict our attention to the following six disciplines: Astrophysics (&lsquo;astro-ph&rsquo;), Computer Science (&lsquo;comp-sci&rsquo;), Condensed Matter Physics (&lsquo;cond-mat&rsquo;), High Energy Physics (&lsquo;hep&rsquo;), Mathematics (&lsquo;math&rsquo;) and Other Physics (&lsquo;oth-phys&rsquo;), which collectively accounted for 98% of all the eprints. Those eprints&nbsp;tagged to multiple arXiv disciplines were counted independently for each discipline. Due to this overlapping feature, the current dataset contains a cumulative total of 2,011,216 eprints.&nbsp;</p> <p>Some general statistics and visualisations per research discipline are provided in the original article (Okamura, 2022), where the validity and limitations associated with the dataset are also discussed.</p> <p>&nbsp;</p> <p><strong>Description of columns (variables)</strong></p> <ul> <li><strong>arxiv_id</strong> :&nbsp;arXiv ID</li> <li><strong>category</strong> :&nbsp;Research discipline</li> <li><strong>pre_year</strong> :&nbsp;Year of posting v1 on arXiv</li> <li><strong>pub_year</strong> :&nbsp;Year of DOI acquisition</li> <li><strong>c_tot</strong> :&nbsp;No. of citations acquired during 1991&ndash;2019</li> <li><strong>c_pre</strong> :&nbsp;No. of citations acquired before and including the year of DOI acquisition</li> <li><strong>c_pub</strong> :&nbsp;No. of citations acquired after the year of DOI acquisition</li> <li><strong>c_<em>yyyy</em></strong>&nbsp;(<em>yyyy</em>&nbsp;= 1991, &hellip;, 2019) :&nbsp;No. of citations acquired in the year&nbsp;<em>yyyy</em>&nbsp;(with &lsquo;<em>yyyy</em>&rsquo; running from 1991 to 2019)</li> <li><strong>gamma</strong> :&nbsp;The quantitatively-and-temporally normalised citation index</li> <li><strong>gamma_star</strong> :&nbsp;The quantitatively-and-temporally standardised citation index</li> </ul> <p><em>Note:</em> The definition of the quantitatively-and-temporally normalised citation index (&gamma;; &lsquo;gamma&rsquo;) and that of the standardised citation index (&gamma;*; &lsquo;gamma_star&rsquo;) are provided in the original article (Okamura, 2022). Both indices can be used to compare the citational impact of papers/eprints published in different research disciplines at different times.&nbsp;</p> <p>&nbsp;</p> <p><strong>Data files</strong></p> <p>A comma-separated values file (&lsquo;<strong>arXiv_impact.csv</strong>&rsquo;) and a Stata file (&lsquo;<strong>arXiv_impact.dta</strong>&rsquo;) are provided, both containing the same information.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Supplementary Data: A global fit of the MSSM with GAMBIT (arXiv:1705.07917)

<p><strong>Supplementary Data</strong></p> <p><em>A global fit of the MSSM with GAMBIT </em><br> <em>arXiv:1705.07917 </em></p> <p>The files in this record contain data for the MSSM7 model considered in the GAMBIT “Round 1” weak-scale SUSY paper.</p> <p>The files consist of</p> <ul> <li>A number of YAML files corresponding to different sets of sampling parameters and/or priors</li> <li>MSSM7.yaml, a YAML file used for postprocessing</li> <li>StandardModel_SLHA2_scan.yaml, a universal YAML fragment included from other YAML files</li> <li>StandardModel_SLHA2_postprocessing.yaml, a YAML fragment included from MSSM7.yaml</li> <li>A final hdf5 file, containing the combined results of all sampling runs</li> <li>An example pip file, for producing plots from the hdf5 file using pippi</li> <li>gambit_preamble.py, a collection of python functions used for in-line data processing in the pip file</li> <li>SLHA1 and SLHA2 files for the best-fit point in each subregion of the fit.  These can found inside the tarball best_fits_SLHA.tar.gz.</li> </ul> <p>The different YAML files corresponding to different samplers and/or priors follow the naming scheme MSSM7_[scanner]_[prior]_[slice]_[special].yaml , where</p> <ul> <li>scanner = Diver, MN</li> <li>prior = log, flat</li> <li>slice = nM2, pM2, Afunnel, hZfunnel, sqcoann, slcoann (positive or negative M2, A/H funnel, h/Z funnel, squark co-annihilation, slepton co-annihilation)</li> <li>special = jDE, [blank] (used pure jDE, or used the default lambdajDE)</li> </ul> <p>A few caveats to keep in mind:</p> <ol> <li> <p>The final hdf5 results file included here was generated in the following way:</p> <ul> <li>carry out initial runs using YAML files following the naming scheme above</li> <li>combine the resulting hdf5 output files into a single file, using<br> gambit/Printers/scripts/combine_hdf5.py</li> <li>postprocess the samples to remove all points more than 5 sigma from the current best fit, using MSSM7_strip.yaml</li> <li>postprocess the samples to include a new likelihood term for LHC Run II searches, and to recompute the FlavBit likelihoods (these were buggy in a pre-release version of GAMBIT), using MSSM7.yaml .</li> </ul> </li> <li> <p>It is not necessary to repeat the steps listed in point 1 when running new scans; the LHC Run II likelihoods can be included in the original YAML file, so that no postprocessing step is required.</p> </li> <li> <p>The YAML files that we give here are updated compared to the ones that we used when generating the hdf5 file, in order to match the set of available options in the release version of GAMBIT 1.0.0. The included physics and numerics are however identical.</p> </li> <li> <p>The YAML files are designed to work with the tagged release of GAMBIT 1.0.0, and the pip file is tested with pippi 2.0, commit 2ab061a8. They may or may not work with later versions of either software (but you can of course always obtain the version that they do work with via the git history).</p> </li> <li> <p>The pip file is an example only. Users wishing to reproduce the more advanced plots in any of the GAMBIT papers should contact us for tips or scripts, or experiment for themselves. Many of these scripts are in multiple parts and require undocumented manual interventions and steps in order to implement various plot-specific customisations, so please don’t expect the same level of polish as for files provided here or in the GAMBIT repo.</p> </li> </ol>

opencc-by-4.0Jun 2017View details →
zenodo44/100

Supplementary Data: Global fits of GUT-scale SUSY models with GAMBIT (arXiv:1705.07935)

<p>Supplementary Data</p> <p><em>Global fits of GUT-scale SUSY models with GAMBIT</em><br> <em>arXiv:1705.07935</em></p> <p>The files in this record contain data for the CMSSM, NUHM1 and NUHM2 models considered in the GAMBIT "Round 1" GUT-scale SUSY paper.</p> <p>For each model, there are</p> <ul> <li>A number of YAML files, each corresponding to a different set of sampling parameters and/or priors</li> <li>A set of YAML files used for postprocessing: CMSSM_intermediate.yaml, CMSSM.yaml, NUHM1.yaml and NUHM2.yaml</li> <li>A final hdf5 file, containing the combined results of all sampling runs</li> <li>An example pip file, for producing plots from the hdf5 file using pippi</li> <li>SLHA1 and SLHA2 files for the best-fit point in each subregion of the fit. These can be found inside the tarball best_fits_SLHA.tar.gz.</li> </ul> <p>The record also contains</p> <ul> <li>StandardModel_SLHA2_scan.yaml and StandardModel_SLHA2_postprocessing.yaml, two universal YAML fragments included from other yaml files</li> <li>gambit_preamble.py, a collection of python functions used for in-line data processing in the pip files</li> </ul> <p>The different YAML files corresponding to different samplers and/or priors follow the naming scheme [model]_[scanner]_[prior]_[slice]_[special].yaml, where</p> <ul> <li>model = CMSSM, NUHM1, NUHM2</li> <li>scanner = Diver, MN</li> <li>prior = log, flat</li> <li>slice = pmu, nmu (positive or negative mu)</li> <li>special = sqcoann, slcoann, [blank] (squark co-annihilation, slepton co-annihilation, or bulk)</li> </ul> <p>A few caveats to keep in mind:</p> <ol> <li> <p>For each model, the final hdf5 results file included here was generated in the following way:</p> <ul> <li>carry out initial runs using YAML files following the naming scheme above</li> <li>combine the resulting hdf5 output files into a single file, using gambit/Printers/scripts/combine_hdf5.py</li> <li>postprocess the samples to remove all points more than 5 sigma from the current best fit, using [model]_strip.yaml</li> <li>postprocess the samples to include a new likelihood term for LHC Run II searches, and to recompute the FlavBit likelihoods (these were buggy in a pre-release version of GAMBIT). For the CMSSM, this happened in two steps, due to persistent flavour bugs, using CMSSM_intermediate.yaml and CMSSM.yaml. For the NUHM1 and NUHM2, this was done in a single step each, using NUHM1.yaml and NUHM2.yaml.</li> </ul> </li> <li> <p>It is not necessary to repeat the steps listed in point 1 when running new scans; the LHC Run II likelihoods can be included in the original YAML file, so that no postprocessing step is required.</p> </li> <li> <p>The YAML files that we give here are updated compared to the ones that we used when generating the hdf5 file, in order to match the set of available options in the release version of GAMBIT 1.0.0. The included physics and numerics are however identical.</p> </li> <li> <p>The YAML files are designed to work with the tagged release of GAMBIT 1.0.0, and the pip files are tested with pippi 2.0, commit 2ab061a8. They may or may not work with later versions of either software (but you can of course always obtain the version that they do work with via the git history).</p> </li> <li> <p>The pip file for each model is an example only. Users wishing to reproduce the more advanced plots in any of the GAMBIT papers should contact us for tips or scripts, or experiment for themselves. Many of these scripts are in multiple parts and require undocumented manual interventions and steps in order to implement various plot-specific customisations, so please don't expect the same level of polish as for files provided here or in the GAMBIT repo.</p> </li> </ol>

opencc-by-4.0May 2017View details →
zenodo44/100

Supplementary Data: Status of the scalar singlet dark matter model (arXiv:1705.07931)

<p>Supplementary Data</p> <p><em>Status of the scalar singlet dark matter model</em><br> <em>arXiv:1705.07931</em></p> <p>The files in this record contain data for the scalar singlet dark matter model considered in the GAMBIT "Round 1" scalar singlet paper.</p> <p>The files consist of</p> <ul> <li>Three YAML files, each corresponding to a different parameter range</li> <li>StandardModel_SLHA2_SingletDM_scan_15.yaml, a universal YAML fragment included from the other three YAML files</li> <li>Three hdf5 files. SingletDM.hdf5 contains the combined results of all sampling runs, and is the basis for the profile likelihood plots in the paper. SingletDM_TW_full.hdf5 and SingletDM_TW_lowmass.hdf5 contain the results from T-Walk scans over the full and low-mass parameter ranges, respectively. These are the bases for the marginalised posterior plots in the paper.</li> <li>An example pip file corresponding to each hdf5 file, for producing plots using pippi</li> <li>A tarball best_fits_yaml.tar.gz containing YAML files of the best-fit point in each subregion of the fit.</li> </ul> <p>The YAML files corresponding to different parameter ranges follow the naming scheme SingletDM_[slice].yaml, where slice may be full, lowmass or neck. Each of these YAML files contains entries in the Scanners node for running Diver, MultiNest, TWalk and GreAT.</p> <p>A few caveats to keep in mind:</p> <ol> <li> <p>The YAML files that we give here are updated compared to the ones that we used when generating the hdf5 file, in order to match the set of available options in the release version of GAMBIT 1.0.0. The included physics and numerics are however identical.</p> </li> <li> <p>The YAML files are designed to work with the tagged release of GAMBIT 1.0.0, and the pip file is tested with pippi 2.0, commit 2ab061a8. They may or may not work with later versions of either software (but you can of course always obtain the version that they do work with via the git history).</p> </li> <li> <p>The pip file is an example only. Users wishing to reproduce the more advanced plots in any of the GAMBIT papers should contact us for tips or scripts, or experiment for themselves. Many of these scripts are in multiple parts and require undocumented manual interventions and steps in order to implement various plot-specific customisations, so please don't expect the same level of polish as for files provided here or in the GAMBIT repo.</p> </li> </ol>

opencc-by-4.0May 2017View details →
zenodo44/100

Numerical data for "Black Hole Metamorphosis and Stabilization by Memory Burden [arXiv:2006.00011]"

<p>This is the numerical data that belongs to the paper</p> <p>G. Dvali, L. Eisemann, M. Michel, S. Zell, <em>Black Hole Metamorphosis and Stabilization by Memory Burden</em>, <a href="https://doi.org/10.1103/PhysRevD.102.103523">Phys. Rev. D <strong>102</strong> (2020) 103523</a>, <a href="https://arxiv.org/abs/2006.00011">arXiv:2006.00011</a>.</p> <p>The numerical data is generated using the computer program <em>TimeEvolver</em>, which was presented in</p> <p>M. Michel, S. Zell, <em>TimeEvolver: A Program for Time Evolution With Improved Error Bound, </em><a href="https://doi.org/10.1016/j.cpc.2022.108374">Comput. Phys. Commun. <strong>277</strong> (2022) 108374</a>, <a href="https://arxiv.org/abs/2205.15346">arXiv:2205.15346</a>.</p> <p>In the following, equation numbers refer to the latest arXiv-version of <a href="https://arxiv.org/abs/2006.00011">arXiv:2006.00011</a>, where also all relevant definitions can be found. In this paper, the procedure for generating the data, which we summarize in the following, is also described in more detail.</p> <p>First, among the five parameters N<sub>c</sub>; <sup><span class="math-tex">\(\epsilon\)</span></sup><sub>m</sub>; C<sub>0</sub>; ∆N<sub>c</sub>; K all but one are fixed (according to eq. (36)). Then for different values of the remaining unfixed parameter - subsequently called X - the following 2-step process is performed.</p> <ol> <li>Time evolution is computed (with the initial state shown in eq. (35)) for many different values of the parameter C<sub>m</sub> in the interval [0;1] (sampling step 10<sup>-3</sup>). The results for X are stored in five folders called &quot;X&quot;, the subfolders of which contain data for different values of X. For example, the subfolder &quot;01&quot; of the folder &quot;C0&quot; consists of data for C<sub>0</sub>=0.01. Please note that the folders &quot;Cgap&quot;, &quot;N0&quot; and &quot;Q&quot; correspond to X=<span class="math-tex">\(\epsilon\)</span><sub>m</sub>, X=N<sub>c</sub> and X=K, respectively. Subsequently, &quot;rewriting values&quot; of&nbsp;C<sub>m</sub> are selected as those for which the amplitude of n<sub>0</sub> is sufficiently large (1.2 times than in the case C<sub>m</sub>=0). This is done in Mathematica-notebooks &quot;_New.nb&quot;. Finally, rewriting values for different values of X are collected using Mathematica-notebooks with names that start on &quot;_Meta&quot;. These notebooks generate the 5 plots shown in figures 4(a), 5(a), 6(a), 7(a) and 8(a).</li> <li>Next finer scans are performed around some of the rewriting values determined in step 1 (new sampling step 5 10<sup>-5</sup>).&nbsp; The results are stored in five folders called &quot;XRates&quot;, where again subfolders correspond to different values of X. For example, the subfolder &quot;01&quot; of the folder &quot;C0Rates&quot; consists of finer scans in&nbsp;C<sub>m</sub> around rewriting value of&nbsp;C<sub>m</sub> for C<sub>0</sub>=0.01. Next, Mathematica-notebooks with the names&nbsp;&quot;_New.nb&quot; or&nbsp;&quot;_NewRates.nb&quot; are used to select around each rewriting value the C<sub>m</sub> that leads to the largest rate (see definition in <a href="https://arxiv.org/abs/2006.00011">arXiv:2006.00011</a>). Finally, these maximal rates for different values of X are collected using Mathematica-notebooks with names that start on &quot;_Meta&quot; and end on &quot;Rates.nb&quot;.&nbsp;These notebooks generate the 5 plots shown in figures 4(b), 5(b), 6(b), 7(b) and 8(b).</li> </ol>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Publication dates for ArXiv publication versions

<p>Lookup tables in plain JSON, mapping ArXiv publication version identifiers to their respective publications dates.</p> <p>The JSON files are archived in&nbsp;<em>arxiv-publication-dates-by-identifier-prefix.tar.gz</em>.<br>The archive contains files named after the date prefix of the ArXiv publication version identifiers they contain.<br>E.g., the file&nbsp;<em>1908.json</em> will contain the data for identifiers <em>1908.12345v1</em>, <em>1908.12345v2</em>, <em>1908.23456v1</em>, etc.<br>Publication dates are given in the format <em>YYYY-MM-DD</em>.</p> <h2>Reproducibility</h2> <p>The <a href="https://snakemake.readthedocs.io/">Snakemake</a> workflow that has produced this dataset has been archived&nbsp;and is available in <em>arxiv-publication-dates-workflow.tar.gz</em>.</p> <h3>Changes in version 1.2</h3> <p>Version 1.2 includes a JSON file that contains a JSON array with all file names included in the dataset: <em>file_names.json</em>.</p> <h3>Changes in version 1.1</h3> <p>For version 1.1, the dataset was extended manually to include a single missing date for <a href="https://arxiv.org/abs/0906.3421v3">arXiv:0906.3421v3</a>: <em>2010-02-02</em>.&nbsp;As of 2024-05-13, the date for the respective version had not been provided in the&nbsp;<em>arXivRaw</em> OAI-PMH data (<a href="http://export.arxiv.org/oai2?verb=GetRecord&amp;identifier=oai:arXiv.org:0906.3421&amp;metadataPrefix=arXivRaw">http://export.arxiv.org/oai2?verb=GetRecord&amp;identifier=oai:arXiv.org:0906.3421&amp;metadataPrefix=arXivRaw</a>).</p> <h3>Running the workflow</h3> <p>To reproduce the dataset on a Linux machine,&nbsp;you need a version of the <a href="https://conda-forge.org/"><em>conda</em></a> package manager installed on your system.</p> <p>Run the following:<br><br></p> <pre><code># Extract the archived workflow tar -xf my-workflow.tar.gz # Create conda environment from lock file conda env create -n arxiv-metadata --file conda-environment.lock.yaml # Activate the environment conda activate arxiv-metadata # Optionally, dry-run the workflow snakemake -n # Produce the output files snakemake --keep-storage-local-copies --software-deployment-method conda -c &lt;NUMBER OF CORES TO USE&gt;</code><br><br>Then, append the file <em>0906.json</em> (included in the <em>tar.gz</em> output) with value <em>2010-02-02</em> for a new key <em>0906.3421v3</em>.</pre> <h2>Workflow</h2> <p>To adapt/change the workflow, clone it from <a href="https://github.com/sdruskat/arxiv-publication-metadata">https://github.com/sdruskat/arxiv-publication-metadata</a>.<br>The workflow version used to produce this dataset is available at <a href="https://doi.org/10.5281/zenodo.11507183">https://doi.org/10.5281/zenodo.11507183</a>.&nbsp;</p>

opencc-zeroApr 2024View details →
zenodo44/100

Mapping between zbMATH Open identifiers, DOIs, ORCIDs and arXiv identifiers

<p>The second version of the mapping between zbMATH Open identifiers for&nbsp;<a href="https://www.wikidata.org/w/index.php?title=Property:P1556&amp;oldid=1755821772">authors</a> and <a href="https://www.wikidata.org/w/index.php?title=Property:P894&amp;oldid=1766254659">documents</a> and <a href="https://www.wikidata.org/wiki/Property:P356">DOIs</a> and <a href="https://www.wikidata.org/wiki/Property:P496">ORCIDs</a> in CSV format.</p> <ul> <li>The file authors.csv contains the mapping between zbMATH Open author id and ORCIDs for 38 159 authors.</li> <li>The file documents.csv contains the mapping between zbMATH Open document id and DOI for 2 813 563 documents.</li> </ul> <p>Beginning from this version, we also provide the mapping between&nbsp;<a href="https://www.wikidata.org/w/index.php?title=Property:P894&amp;oldid=1766254659">documents</a> and <a href="https://www.wikidata.org/w/index.php?title=Property:P818&amp;oldid=2151126552">arXiv</a> in CSV format</p> <ul> <li>The file arxiv.csv contains the mapping between zbMATH Open document id and arXiv identifiers for 528 640 documents.</li> </ul> <p>See https://zbmath.org/about/ (section Full Text Links) for a live version of this dataset. That version is more current but less reproducible. Moreover, the dataset here is restricted to documents with a permanent zbMATH Open identifier in the form <code>Zbl d+.d+</code>.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-sa-4.0Jul 2024View details →
zenodo44/100

Supplementary Data: Impact of vacuum stability, perturbativity and XENON1T on global fits of Z2 and Z3 scalar singlet dark matter (arXiv:1806.11281)

<p>&nbsp;</p> <p><strong>Supplementary Data</strong></p> <p>&nbsp;</p> <p><em>Impact of vacuum stability, perturbativity and XENON1T on global fits of Z<sub>2</sub> and Z<sub>3</sub> scalar singlet dark matter</em> <a href="https://arxiv.org/abs/1806.xxxxx"><em>arXiv:</em></a><em><a href="https://arxiv.org/abs/1806.11281">1806.11281</a></em></p> <p>The files in this record contain data for the scalar singlet dark matter models considered in the <a href="http://gambit.hepforge.org">GAMBIT</a> &quot;Scalar singlet Mark II&quot; paper.</p> <p>The files consist of</p> <ul> <li>30 regular YAML files</li> <li><code>StandardModel_SLHA2_scan.yaml</code>, a universal YAML fragment included from the other YAML files</li> <li>14 hdf5 files. 8 of these correspond to the complete set of combined samples for each fit. These 8 fits are generated from all binary permutations of three run properties: Z2 or Z3 model, with or without absolute vacuum stability demanded, and with constraints from the 2017 or 2018 XENON1T data. These 8 hdf5 files are used to generate the profile likelihood plots in the paper. The other 6 hdf5 files are the results of T-Walk runs, and are used to generate the posterior pdfs in the paper.</li> <li>Some example pip files for producing plots from the hdf5 files using <a href="github.com/patscott/pippi">pippi</a></li> <li>A tarball <code>best_fits_yaml.tar.gz</code> containing YAML files of the best-fit point in each of the 8 fits.</li> </ul> <p>The files follow the naming scheme <code>SingletDM_[model]_[slice]_[vacuum]_[xenon]_[prior]_[scanner].yaml</code>.</p> <ul> <li>model: <code>Z2</code> or <code>Z3</code></li> <li>slice: <code>full</code>, <code>lowmass</code>, <code>neck</code> or absent (for hdf5 files)</li> <li>vacuum: <code>ms</code> (metastable) or <code>vs</code> (absolute vacuum stability)</li> <li>prior: <code>logmu3</code>, <code>flatmu3</code> or absent (for Z<sub>2</sub> scans)</li> <li>scanner: <code>TWalk</code> or absent (implies Diver scans in the case of YAML files, and indicates merged samples potentially from both Diver and T-Walk in the case of hdf5 files)</li> </ul> <p>A few caveats to keep in mind:</p> <ol> <li> <p>The YAML files are designed to work with GAMBIT 1.2.0, commit e4d3f739, and the pip files are tested with pippi 2.1, commit c094b8c8. They may or may not work with later versions of either software (but you can of course always obtain the version that they do work with via the git history).</p> </li> <li> <p>The pip files are examples only. Users wishing to reproduce the more advanced plots in any of the GAMBIT papers should contact us for tips or scripts, or experiment for themselves. Many of these scripts are in multiple parts and require undocumented manual interventions and steps in order to implement various plot-specific customisations, so please don&#39;t expect the same level of polish as for files provided here or in the GAMBIT repo.&nbsp;</p> </li> </ol> <p>&nbsp;</p>

opencc-by-4.0Jun 2018View details →
zenodo44/100

Ancillary files for "Reinterpreting the ATLAS bounds on heavy neutral leptons in a realistic neutrino oscillation model [arXiv: 2107.12980]"

<p><em>(Description copied from Appendix A &quot;Ancillary files&quot; of the companion paper)</em></p> <p>In order to simplify the interpretation of experimental results within realistic HNL models, we are including a number of data files along with the present publication. They can be used to generate the relevant signal samples, or to implement the extrapolation method presented in section 3.2.</p> <p><strong>Card files for the Monte-Carlo event generation</strong></p> <p>The /attachments/card_files folder contains the MadGraph card files (ending in .dat) and scripts (ending in .txt) for generating the signal samples used in this analysis, as well as for computing the total HNL width. Due to the OSSF veto, only processes with no opposite-charge same-flavor lepton pairs have been included. Additional relevant processes can easily be added by modifying the <em>generate</em> and <em>add process</em> lines in the *.txt files. All samples (except the ones used to compute the HNL width, which are generated at parton level) are generated at leading order, include up to two hard jets, and are showered and hadronized using Pythia 8. This is essential for obtaining a realistic W spectrum. The shower parameters could probably benefit from further tuning, and further improvements in the W spectrum accuracy are expected at NLO (using a suitable model). To allow computing the signal efficiencies, all cuts have been disabled in the run card (with the exception of the maximum <span class="math-tex">\(|\eta_{\mathrm{jet}}|\)</span> which needs to be set to 5 for correct matching).</p> <p><strong>Signal cross sections</strong></p> <p>The cross sections for the various processes considered in this analysis, as well as the total HNL width (both computed using MadGraph as described in section 3.2), are provided as JSON files in the /attachments/cross_sections folder.</p> <p>The file total_hnl_width.json contains the total HNL width&nbsp;<span class="math-tex">\(\hat{\Gamma}_{\alpha}(M_N)\)</span> (expressed in GeV), computed for the 5 mass points used in this analysis, and under the assumption of unit mixing with a single flavor <span class="math-tex">\(\alpha\)</span>, for each flavor. The total HNL width can then be computed for any combinations of mixing angles using eq. (3.2). The file is organized as two nested dictionaries, with the first key denoting the HNL mass <span class="math-tex">\(M_N\)</span>, and the second one the flavor&nbsp;<span class="math-tex">\(\alpha\)</span> for which the total width&nbsp;<span class="math-tex">\(\hat{\Gamma}_{\alpha}(M_N)\)</span> has been computed for a unit mixing angle&nbsp;<span class="math-tex">\(|\Theta_{\alpha}|^2 = 1\)</span> (with <em>Wtot_e</em> for <span class="math-tex">\(\alpha=e\)</span>, <em>Wtot_mu</em> for <span class="math-tex">\(\mu\)</span> and <em>Wtot_tau</em> for <span class="math-tex">\(\tau\)</span>).</p> <p>The file cross_sections.json contains the reference cross sections&nbsp;<span class="math-tex">\(\sigma_P^{\mathrm{ref}}\)</span> (in pb) for all the processes <em>P</em> considered in this analysis, expressed for&nbsp;<span class="math-tex">\(|\Theta|_{\mathrm{ref}}^2 = 1\)</span> and <span class="math-tex">\(\Gamma_{\mathrm{ref}} = 10^{-5}\,\mathrm{GeV}\)</span>. The file is organized as two nested dictionaries, with the first key denoting the HNL mass <span class="math-tex">\(M_N \)</span> and the second the process <em>P</em>. The correspondence between the key and the physical process can be found in table 7.</p> <p><strong>Signal efficiencies</strong></p> <p>The efficiencies resulting from the event selection described in section 3.1, as well as their parametrization according to eq. (3.6) (as discussed in section 3.3) can respectively be found in the files efficiencies.json and fitted_efficiencies.json in the /attachments/efficiencies folder.</p> <p>The file efficiencies.json is organized as follows. The data is located in a triply nested dictionary under the data key: the first level corresponds to the HNL mass hypothesis <span class="math-tex">\(M_N\)</span>, the second to the process key (cf. table 7) and the third to the&nbsp;<span class="math-tex">\(M(l_{\mathrm{sublead}},l')\)</span> bin for which the efficiency is computed. The values of the bottom-most dictionary are lists containing the efficiencies for a number of HNL lifetimes, as listed in meters in levels/lifetime.</p> <p>Finally, the file fitted_efficiencies.json is also organized as a triply nested dictionary, with the first level corresponding to the HNL mass <span class="math-tex">\(M_N\)</span>, the second to the process key, and where the third level denotes the fit parameter from eq. (3.6). tau0 is for <span class="math-tex">\(\tau_0\)</span>, epsilon0_total for&nbsp;<span class="math-tex">\(\epsilon_0\)</span> (the unbinned prompt efficiency), and epsilon0_binned is a list containing the prompt efficiencies&nbsp;<span class="math-tex">\(\epsilon_{0,b}\)</span> for the five&nbsp;<span class="math-tex">\(M(l_{\mathrm{sublead}},l')\)</span> bins <em>b</em> (in the same order as in efficiencies.json). The layout described here (or a similar one) can be used by experiments to report their signal efficiencies in a way that allows theorists to compute the expected signal for arbitrary choices of mixing angles.</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Data_for_arXiv_1809_00822

<p>This data set contains the raw and meta data for all figures shown in arXiv&nbsp;1809.00822 (https://arxiv.org/abs/1809.00822) &quot;Observation of a broadband Lamb shift in an engineered quantum system&quot;.</p>

opencc-by-4.0Dec 2018View details →
zenodo44/100

ArXiV-Entity/Relation annotated dataset

<p>This dataset is a collection of abstracts from the CS section of ArXiV, each annotated with <a href="https://github.com/dwadden/dygiepp">DyGIE++</a> (SciERC model)</p> <p>The dataset can be used to train triple extractors or to cluster triples (in the Computer Science and AI domains).</p> <p>Supersedes the ArXiV-AIKG dataset as these triples are unconstrained (so they don&#39;t forcibly appear in AIKG)</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

arXiv abstracts and titles from 1,469 single-authored papers (100 unique authors) in computer science

<p>This dataset is meant to be used for experiments of Authorship Analysis. The dataset&nbsp;consists of abstracts of single-author papers from arXiv crawled using the arXiv&#39;s API by querying a list of computer-science-related keywords (&quot;deep learning&quot;, &quot;machine learning&quot;, &quot;information retrieval&quot;, &quot;computer science&quot;, &quot;data mining&quot;, &quot;support vector&quot;, &quot;logistic regression&quot;, &quot;artificial intelligence&quot;, &quot;supervised learning&quot;&#39;).&nbsp;The corpus somehow follows a power-law distribution, with few prolific authors and many authors accounting for very few papers each: we retained authors with at&nbsp;least 10 papers, resulting in a total of 1,469 documents from 100 authors. The most prolific authors (Peter D. Turney and Subhash Kak) have 34 abstracts to their names, the 10 most prolific authors have written 22 or more articles, while 50% of the authors have no more than 12 abstracts to their names. In order to divide the corpus into a training set and a test set we perform a stratified split, with the production of each author being split into a training set (70%) and a test set (30%). We use these documents as examples of &quot;scientific communication&quot;, characterised by a precise and compact style, with an abundance of technical terminology.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Cartolabe arXiv 2019

<p>Cartolabe arXiv 2019 dataset is built with data extracted from the <a href="https://arxiv.org/">arXiv science repository.</a> It gives access to ~1.6m articles in the fields of physics, mathematics, computer science, quantitative biology, quantitative finance, statistics, electrical engineering and systems science, and economics. Only authors with at least 4 articles are kept. This map was generated with the data extracted from arXiv on 29/11/2019.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Data and scripts for collective intelligence research (arXiv:2204.13424)

<p>This is the data and scripts for the study <strong>From Prediction Markets to Interpretable Collective Intelligence</strong> by Alexey V. Osipov and Nikolay N. Osipov (<a href="http://doi.org/10.48550/arXiv.2204.13424">arXiv:2204.13424</a> [cs.GT])</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Pre- and post-publication citations to published arXiv preprints

<p>This dataset contains citations to published preprints, both before they are published and after they are published. Details of the data are provided in the <code>README.md</code>.</p>

opencc-by-4.0Dec 2019View details →
zenodo40/100

Raw data used for arXiv 2004.02485

<p>This dataset contains the raw measurement data used for the study of spectral, temporal, thermal and magnetic field dependence of TLS in frequanty tunable resonators outlined in arxiv:2004.02485.</p>

opencc-by-4.0Aug 2020View details →
zenodo40/100

Supplementary Data: Strengthening the bound on the mass of the lightest neutrino with terrestrial and cosmological experiments (arXiv:2009.03287)

<p><strong>Supplementary Data</strong></p> <p><em>Strengthening the bound on the mass of the lightest neutrino with terrestrial and cosmological experiments (arXiv:2009.03287)</em></p> <p>The files in this record contain data from the scans of the models considered in the <a href="http://gambit.hepforge.org">GAMBIT</a> paper on neutrino masses.</p> <p>The files consist of</p> <ul> <li>21 <code>.yaml</code> files corresponding to different models, sampling parameters and/or priors</li> <li>11 final <code>.hdf5</code> files, containing the results of running GAMBIT with each yaml file</li> <li>10 <code>.margestats</code> files containing 1D credible regions for parameters and observables, obtained by running <a href="https://github.com/cmbant/getdist">getdist</a> on the hdf5 files</li> <li>An example file 3-NHB_Neff2_PC500_pp.pip file for plotting the results of a single hdf5 file with <a href="github.com/patscott/pippi">pippi</a></li> <li>A tarball including all files in this record except the hdf5 files.</li> </ul> <p>The different yaml, hdf5 and margestats files corresponding to different models, priors or setttings follow the naming scheme <code>[scan index]-[hierarchy][m_nu0 prior]_[Neff prior]_[scanner]_[extra]_[step].[extension]</code>, where</p> <ul> <li>scan index = <code>1</code>-<code>11</code></li> <li>hierarchy = <code>NH</code> (normal hierarchy), <code>IH</code> (inverted hierarchy)</li> <li>m_nu0 pior = <code>A</code> (linear-log prior on m_nu0), <code>B</code> (linear prior on m_nu0)</li> <li>Neff prior = <code>0</code> (Delta N_eff = 0), <code>1</code> (Delta N_eff &gt; 0), <code>2</code> (Delta N_eff free)</li> <li>scanner = <code>PC500</code> (Polychord with 500 live points), <code>DIV10k</code> (Diver with NP=1e4)</li> <li>extra = blank (standard likelihood combination), <code>Lyalpha</code> (likelihood also includes eBOSS DR14 Lyman-alpha BAO scale measurements)</li> <li>step = blank (main scan), <code>pp</code> (postprocessing of outputs of main scan).</li> <li>extension = <code>yaml</code>, <code>margestats</code>, <code>hdf5/hdf5.tar.gz</code></li> </ul> <p>A few caveats to keep in mind:</p> <ol> <li> <p>The YAML files are designed to work with the tagged release of GAMBIT 1.5.0, and the pip file is tested with pippi 2.1. They may or may not work with later versions of either software (but you can of course always obtain the version that they do work with via the git history).</p> </li> <li> <p>The pip file is an example only. Users wishing to reproduce the more advanced plots in any of the GAMBIT papers should contact us for tips or scripts, or experiment for themselves. Many of these scripts are in multiple parts and require undocumented manual interventions and steps in order to implement various plot-specific customisations, so please do not expect the same level of polish as for files provided here or in the GAMBIT repo.</p> </li> </ol>

opencc-by-4.0Sep 2020View details →
zenodo40/100

Likelihoods for the CTA sensitivity to a dark matter signal from the Galactic centre (A. Acharyya et al., [arXiv:2007.16129])

<p>We present likelihoods to estimate upper limits on DM pair-annihilation in the Galactic centre, based on&nbsp; the<br> Cherenkov Telescope Array (CTA) consortium publication &quot;Sensitivity of the Cherenkov Telescope Array to a<br> dark matter signal from the Galactic centre&quot; [arXiv:2007.16129]. As explained in more detail in Sec. 5.2 of that<br> article, these likelihoods are suitable for models featuring cuspy dark matter profiles and generic gamma-ray<br> spectra produced from annihilating dark matter.</p> <p>The four files contain likelihoods that have been derived with respect to the full CTA South baseline array layout<br> and the initial construction phase of CTA South. For each of these two cases, as indicated by the filename, we<br> provide tables with and without the inclusion of systematic uncertainties (where the former refers to the benchmark<br> treatment of systematic uncertainties as described in the main publication). Further, all likelihoods are based on the<br> benchmark analysis settings with respect to masking known bright gamma-ray sources and the adopted<br> interstellar emission models.</p>

opencc-by-4.0Sep 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record