Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
23
datasets available to search
ShareScore release 0.9.0
Dataset results
23 results for “arXiv.org”
Datasets with typos for testing LEA (https://arxiv.org/abs/2307.02912)
<p>Datasets for testing the generalization capacity of LEA in the presence of typos. Link to the paper: https://arxiv.org/abs/2307.02912</p> <p>Below is a list of raw public datasets and different versions of test splits to which automatically synthetically generated typos have been added by deleting and replacing characters.</p> <ul> <li>Abt-Buy</li> <li>Amazon-Google</li> <li>WDC-Computers (small, medium, large and xlarge)</li> <li>WDC-All (xlarge)</li> <li>RTE</li> <li>MRPC</li> </ul> <p> </p> <p>Reference:</p> <p>Almagro, M., Almazán, E., Ortego, D., & Jiménez, D. (2023, August). LEA: Improving Sentence Similarity Robustness to Typos Using Lexical Attention Bias. In <em>Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining</em> (pp. 36-46).</p>
Data for figures in arxiv.org/abs/2112.01864
<p>Data for plotting figures in arxiv.org/abs/2112.01864<br> Files ending in .gz have been compressed with gzip.</p> <p>velocity.txt: figure 1<br> convergence.txt: figure 2<br> modified_ae.txt: figures 3-4<br> modified_ae_hires.txt: figure 4<br> cpcp.txt: figures 5-6,14<br> cpcp_hires.txt: figure 6<br> solar_wind.txt: solar wind from GUMICS for empirical models in figures 7-19<br> lobe.txt: figures 7-9<br> magnetopause.txt: figures 10-11<br> plasma_sheet.txt: figures 12-13<br> Kp.txt: figures 15-17<br> auroral_electrojet.txt: figures 18-19</p>
ArXiV Archive: A tidy and complete archive of metadata for papers on arxiv.org, 1993-2019
<p>This is a full archive of metadata about papers on arxiv.org from 1993-2018, including abstracts. Data is tidy and packed in TSV files, in two different collections of the total dataset: per year (all categories) and per primary category (all years). This archive also includes Jupyter notebooks for unpacking and analyzing it in python. See the README.md file and https://github.com/staeiou/arxiv_archive for more information.</p> <p>This release has the exact same data as in v1.0.0, but the notebook 4-analysis-examples.ipynb is updated and fixes an analysis bug.</p>
All Computer Science Papers @ arXiv.org -- A High-Quality Gold Standard for Citation-based Tasks
<p>We propose a newly-created gold standard <strong>data set for citation-based tasks</strong>. This gold standard is based on <strong>all computer science papers in arXiv.org</strong>.</p> <p><strong>Abstract</strong>. Analyzing and recommending citations with their specific citation contexts have recently received much attention due to the growing number of available publications. Although data sets such as CiteSeerX have been created for evaluating approaches for such tasks, those data sets exhibit striking defects. This is understandable if one considers that both information extraction and entity linking as well as entity resolution need to be performed. In this paper, we propose a new evaluation data set for citation-dependent tasks based on arXiv.org publications. Our data set is characterized by the fact that it exhibits almost zero noise in the extracted content and that all citations are linked to their correct publications. Besides the pure content, available on a sentence-basis, cited publications are annotated directly in the text via global identifiers. As far as possible, referenced publications are further linked to DBLP. Our data set consists of over 15M sentences and is freely available for research purposes. It can be used for training and testing citation-based tasks, such as recommending citations, determining the functions or importance of citations, and summarizing documents based on their citations.</p> <p> </p> <p>More information can be found in our <strong>publication "<a href="http://www.lrec-conf.org/proceedings/lrec2018/pdf/283.pdf">A High-Quality Gold Standard for Citation-based Tasks</a>" (LREC'18)</strong>.</p> <p>You can cite the data set as follows:</p> <pre><code>@inproceedings{DBLP:conf/lrec/0001TJ18, author = {Michael F{\"{a}}rber and Alexander Thiemann and Adam Jatowt}, title = "{A High-Quality Gold Standard for Citation-based Tasks}", booktitle = "{Proceedings of the Eleventh International Conference on Language Resources and Evaluation}", series = "{LREC'18}", location = "{Miyazaki, Japan}", year = {2018}, url = {http://www.lrec-conf.org/proceedings/lrec2018/summaries/283.html} } </code></pre> <p> </p>
Dataset used in "Uncertainty-Aware Learning for Improvements in Image Quality of the Canada-France-Hawaii Telescope" (https://arxiv.org/abs/2107.00048)
<p>'x_train.p', 'y_train.p': pickle files for training split containing 50,757 samples</p> <p>'x_val.p', 'y_val.p': pickle file for validation split containing 5,640 samples</p> <p>'x_test.p', 'y_test.p': pickle file for test split containing 6,267 samples</p> <p>'feature_names.p': pickle file containing names of all 119 features</p>
datafiles for https://arxiv.org/abs/2010.00430
<p>Data files for the figures</p>
Preprints from arXiv.org in cs and q-bio
<p>arXiv Quantitative Biology (q-bio)</p> <p>This dataset contains all preprints with the label “q-bio” from 2003 (when the section was introduced) to 2014. Downloaded on 10 June, 2016.</p> <p>arXiv CS (cs)</p> <p>This dataset contains all preprints with the label “cs” from 2003 to 2014. Downloaded on 10 June, 2016.</p>
Observational Dataset for "Constraining Global Coronal Models with Multiple Independent Observables", Badman et al. (2022). Arxiv : https://arxiv.org/abs/2201.11818
<p>Observational Dataset for "Constraining Global Coronal Models with Multiple Independent Observables", Badman et al. (2022). Arxiv : https://arxiv.org/abs/2201.11818</p> <p>-----------------------------<br> -----------------------------</p> <p>Contact : Samuel T. Badman (he/him) samuel_badman@berkeley.edu, Space Sciences Lab, UC Berkeley.</p> <p>-----------------------------<br> -----------------------------</p> <p>License : Creative Commons Attribution 4.0 International</p> <p>-----------------------------<br> -----------------------------</p> <p>Research Goal of Dataset : Data supports the above titled work in defining a framework for evaluating the magnetic structure of global coronal models via the evaluation of three single valued metrics. This repository contains observational data products used as input for the studies described in this work with the aim to allow external coronal modelers to reproduce and evaluate their own work against the same dataset we used.</p> <p>-----------------------------<br> -----------------------------</p> <p>Structure of files : This repository contains three subfolders each containing observational data relating to the three metrics defined in Badman et. al. (2022). These are :</p> <p>-----------------------------</p> <p>1) ``Metric1_EUVCarringtonMaps''</p> <p>Content :</p> <p><carr_maps.####.final.h5> : Carrington maps of extreme ultraviolet (EUV) emission as observed by the SDO/AIA. These files contain slices of different wavelengths together, saved in hdf5 format. Maps for Carrington rotations (#### = 2210,2215,2216,2221) span the time intervals of interest in the associated work. The 193 angstrom wavelength slice from these maps were used as input into the EZSEG algorithm (see manuscript text) to generate ``observations'' of coronal hole boundaries which can then be compared via binary classification to modeled open field boundaries.</p> <p><read_plot_example_metric1.py> : A python script which demonstrates reading in the hdf5 files and viewing the names of the different slices, then plots the 193 slice. The slice name of primary interest is 193A ('map_0193'), but slices at 171,211 angstrom, and a magnetogram are included.</p> <p>-----------------------------</p> <p>2) ``Metric2_StreamerBelt''</p> <p><read_plot_example_metric2.py> : A python script which demonstrates reading and plotting an example white light carrington map from this data set, as well as overplotting the downstream data extraction of the streamer maximum brightness (SMB) line.</p> <p><br> 2a) ``Metric2_StreamerBelt/WL_CarringtonMaps''</p> <p>Content :</p> <p><WL_CRMAP_YYYYMMDDTHHmmSS_LC2_5p0Rs.fits> : Carrington maps of white light intensity extracted at 5.0Rs altitude using coronagraph images taken by SOHO/LASCO, using the method described in the manuscript and Poirier et al. (2021). Maps at a daily cadence over each 60 day time interval studied in the manuscript are included here, incorporating the new data available as the sun rotated. Here saved as fits files.</p> <p><WL_CRMAP_YYYYMMDDTHHMMSS_LC2_5p0Rs.mat> : Carrington maps as above but saved in .mat format (MATLAB).</p> <p>2b) ``Metric2_StreamerBelt/SMB_Line_Extractions''</p> <p><C2_YYYMMDDHHmmSS_5.0Rs_SMB.ascii> : Downstream processed versions of the relevant White light carrington map from which the line of maximum brightness (SMB line) has been extracted, as well as the streamer belt "thickness" at each longitude. This is tabulated as a 3d coordinate gridded evenly in longitude, and each SMB grid point as a northwards and southwards thickness, tabulated in degrees. These data are described in the header of each file and the extraction process is described in detail in the manuscript.</p> <p>-----------------------------</p> <p>3)Metric3_InSituTimeSeries</p> <p>Content :</p> <p><E##_XYZ_polarity.txt> : In situ polarity timeseries for 60 day intervals at 1 hour cadences during PSP encounters ## = [01,02,03], measured by spacecraft XYZ = [PSP,STA,OMN], Parker Solar Probe, STEREO A and OMNI (Earth-L1 dataset). Data values are +/- 1 indicating if magnetic vector is directed sunward or antisunward for each hour. This value is determined as described in the main text by finding the peak of a histogram of 1D B_R values over that hour interval and taking its sign.</p> <p><read_plot_example_metric3.py> : A python script which demonstrates reading in the in situ timeseries for encounter 1 and plotting them.</p> <p><gen_ss_footpoints_psp.py> : A python script which demonstrates an open source method to produce source surface footpoints for a given spacecraft (here PSP) which can be used to sub-sample a HCS map provided by a modeler to generate a modeled time series which can be used to produce scores for metric 3 described in the associated manuscript.</p> <p>-----------------------------<br> -----------------------------</p> <p>Python scripts included in this dataset use python packages</p> <p>astropy - https://github.com/astropy/astropy<br> h5py - https://github.com/h5py/h5py<br> astrospice - https://github.com/dstansby/astrospice<br> matplotlib - https://github.com/matplotlib/matplotlib<br> sunpy - https://github.com/sunpy/sunpy</p> <p> </p> <p> </p>
Raw web API output related to Figure 3 from https://arxiv.org/abs/2309.09893
<p>Raw web API output related to Figure 3 from https://arxiv.org/abs/2309.09893.</p> <p> </p> <p>Can also be found here https://github.com/CQCL/One_Bit_Addition_Raw_Data</p>
Segmentation datasets used in https://arxiv.org/abs/2112.12955
<p>The segmentation datasets (both training and test sets) used in https://arxiv.org/abs/2112.12955 . The training set is in the folder SKIN\SKIN_ECUTr.rar , it already contains data augmented images (augmented data have size 352*352, instead the original images have the original size), we have used this dataset for training all the networks (also ELoss101). </p> <p>For the polyp dataset, only the training set without data augmentation is available in this page. Notice that we use training images of size 352*352 before to apply data augmentation and to feed neural networks (even with Eloss101).</p> <p>The test sets for the polyp segmentation problem are available here: https://zenodo.org/record/5579392#.YdzC1f7MKUk</p> <p>The training sets, coupled with data augmentation, for the polyp segmentation problem are available here: https://zenodo.org/record/5579392#.YdzC1f7MKUk</p> <p> </p>
Dataset used for "End-to-end simulation of particle physics events with Flow Matching and generator Oversampling" , https://arxiv.org/abs/2402.13684
Open the record for dataset details and reuse information.
Subset of BioID dataset (https://www.bioid.com/facedb/) used for image denoising benchmark as used in DivNoising paper (https://arxiv.org/abs/2006.06072)
<p>The original BioID dataset comes from https://www.bioid.com/facedb/. </p> <p>A subset of original BioID dataset was used for image denoising benchmark (corrupted with zero mean Gaussian noise of std 15) as in DivNoising paper (https://arxiv.org/abs/2006.06072)</p>
Flywing (noise 10) dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)
<p>Flywing n10 dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)</p>
Flywing (noise 0) dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)
<p>Flywing n0 dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)</p>
DSB (noise 20) dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)
<p>DSB n20 dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)</p>
DSB (noise 10) dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)
<p>DSB n10 dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)</p>
DSB (noise 0) dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)
<p>DSB n0 dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)</p>
Mouse (noise 0) dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)
<p>Mouse n0 dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)</p>
Flywing (noise 20) dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)
<p>Flywing n20 dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)</p>
Mouse (noise 10) dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)
<p>Mouse n10 dataset for microscopy image denoising and segmentation benchmark as used in DenoiSeg paper (https://arxiv.org/abs/2005.02987)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.