Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
369
datasets available to search
ShareScore release 0.9.0
Dataset results
369 results for “Datasets Benchmarking”
STAVER: A Standardized Benchmark Dataset-Based Algorithm for Effective Variation Reduction in Large-Scale DIA-MS Data
<p>This project focuses on developing and applying STAVER, an innovative DIA algorithm designed to eliminate non-biological noise and variability from the large-scale DIA-MS study dataset analyses. STAVER is a flexible framework that utilizes prior knowledge regarding peptide separation coordinates (RT) and fragment ion intensities from the standard benchmark datasets, which effectively mitigates non-biological noise potential during library searches, enhancing spectrum identifications and protein quantification accuracy. Furthermore, the robustness and broad applicability of STAVER were validated in multiple large-scale DIA datasets from different platforms and laboratories, demonstrating significantly improved precision and reproducibility of protein quantification. It facilitates the comparative and integrative analysis of DIA datasets across different platforms and laboratories, enhancing the consistency and reliability of findings in clinical research. The project aims to promote the adoption of hybrid library search and improve the sensitivity and quality of DIA proteomics data through the open-source STAVER software package.</p>
Dataset for Gidden et.al. 2023 Updated AR6 Mitigation Benchmarks using National Emissions Inventories
<p>Scenario variables related to LULUCF emissions and removals in IPCC-assessed scenarios from AR6 calculated using OSCAR. Both direct fluxes, corresponding to model-reporting conventions, and indirect fluxes, which constitute an alignment factor to national inventories, are provided. See the original publication for more details (<a href="https://www.nature.com/articles/s41586-023-06724-y">https://www.nature.com/articles/s41586-023-06724-y</a>).</p><h4>Change log from version 1</h4><ul><li><i>AR6 Reanalysis|OSCARv3.2|Emissions|CO2|AFOLU</i> and <i>AR6 Reanalysis|OSCARv3.2|Carbon Removal|Land</i> are removed, users should explicitly calculate if needed but otherwise use direct and indirect fluxes.</li><li><i>AR6 Reanalysis|OSCARv3.2|Carbon Removal</i> is now calculated only using the direct flux component (i.e., <i>AR6 Reanalysis|OSCARv3.2|Carbon Removal|Land|Direct</i>)</li></ul>
AlphaPept Benchmark Dataset
<p>This is accompanying data for the benchmark of the AlphaPept software https://www.biorxiv.org/content/10.1101/2021.07.23.453379v1</p><p>Files starting with PXD028735: These are the output files and logs from <strong>Figure 7: Benchmarking AlphaPept on Thermo and Bruker mixed species datasets. </strong>Raw data was taken from PXD028735.</p><p>SLURM: Code and logs for running on SLURM cluster for the 200 HeLa proteome test, see <strong>Supplementary Table 1</strong></p><p>AWS_test: Timings when running on large AWS instances, see <strong>Supplementary Table 1</strong></p><p> </p>
LoDoInd: A Benchmark Low-dose Industrial CT Dataset - 1 of 3
<h2>Summary</h2> <p>This dataset accompanies the paper "LoDoInd: Introducing A Benchmark Low-dose Industrial CT Dataset and Enhancing Denoising with 2.5D Deep Learning Techniques". We are releasing the dataset with five different dose levels, including a reference set. All datasets are pre-registered, making them immediately suitable for deep learning applications in industrial CT.</p> <h2>Description</h2> <p>The uploaded content includes reconstructed images for noise levels 1 and 2. Each level comprises 4000 slices, with each slice being 1250x1250 pixels. Due to the 50 GB space limitation per submission on Zenodo, the rest of the dataset is available through separate links listed below:</p> <ul> <li>Noise1 and Noise2 (this one) <a href="../records/10356955" target="_blank" rel="noopener">https://zenodo.org/records/10356955</a></li> <li>Noise3 and Noise4 <a href="../records/10391277" target="_blank" rel="noopener">https://zenodo.org/records/10391277</a></li> <li>Noise5 and Reference <a href="../records/10391412" target="_blank" rel="noopener">https://zenodo.org/records/10391412</a></li> </ul> <p>The scanning parameters for all noise levels are summarized in the table below:</p> <table> <tbody> <tr> <td> </td> <td>Averaged Projs</td> <td>Exposure Time/ms</td> <td>Scan Time/min</td> <td>Voltage/kV</td> <td>Current/uA</td> </tr> <tr> <td>Reference</td> <td>6</td> <td>333</td> <td>59.3</td> <td>140</td> <td>180</td> </tr> <tr> <td>Noise Level 1</td> <td>1</td> <td>333</td> <td>9.9</td> <td>140</td> <td>180</td> </tr> <tr> <td>Noise Level 2</td> <td>1</td> <td>333</td> <td>9.9</td> <td>140</td> <td>90</td> </tr> <tr> <td>Noise Level 3</td> <td>1</td> <td>333</td> <td>9.9</td> <td>140</td> <td>45</td> </tr> <tr> <td>Noise Level 4</td> <td>1</td> <td>333</td> <td>9.9</td> <td>140</td> <td>23</td> </tr> <tr> <td>Noise Level 5</td> <td>1</td> <td>333</td> <td>9.9</td> <td>140</td> <td>12</td> </tr> </tbody> </table> <h2>Additional link</h2> <p>The code for supervised learning-based denoising is available <a href="https://github.com/jiayangshi/LoDoInd">code</a> .</p> <h2>Acknowledgment</h2> <p>This research was co-financed by the European Union H2020-MSCA-ITN-2020 under grant agreement no. 956172 (xCTing). </p>
Benchmark dataset for the Iterative Rotations and Assignments (IRA) algorithm
<p>In the publication of the Iterative Rotations and Assignments (<a href="https://doi.org/10.1021/acs.jcim.1c00567">IRA</a>) algorithm, a dataset of atomic structures was used for the benchamrking of the algorithm. Two other algorithms were included in the benchmark, <a href="https://doi.org/10.1021/acs.jcim.6b00546">ArbAlign</a> and <a href="https://doi.org/10.1021/acs.jctc.7b00543">fastoverlap</a>. The atomic structures were obtained from several other publications, for details please refer to <a href="https://doi.org/10.1021/acs.jcim.1c00567">IRA</a> publication. </p> <p>The data shared here contains the atomic structures, copies of algorithms, and all scripts used in the benchmark of the reference paper.<br>The source code of IRA algorithm is accessible on <a href="https://github.com/mammasmias/IterativeRotationsAssignments/tree/master">github</a>.</p> <p>Each algorithm and dataset contained in this archive may be subject to its own license, please refer to the README files inside.</p>
PCB-Vision: A Multiscene RGB-Hyperspectral Benchmark Dataset of Printed Circuit Boards
<p><strong>PCB-Vision Dataset</strong></p> <p>Description:</p> <p>The PCB-Vision dataset is a multiscene RGB-Hyperspectral benchmark dataset comprising 53 Printed Circuit Boards (PCBs). The RGB images are collected using a Teledyne Dalsa C4020 camera on a conveyor belt, while hyperspectral images (HSI) are acquired with a Specim FX10 spectrometer. The HSI data contains 224 bands in the VNIR range [400 - 1000]nm.</p> <p><strong>Data Format</strong></p> <ul> <li>RGB Images: .png files</li> <li>PCB Masks: .jpg files</li> <li>HSI Data: Each hyperspectral data cube is accompanied by a data file and a .hdr file.</li> </ul> <p><strong>Folder Organization</strong></p> <ul> <li>PCBVision <ul> <li>HSI/ <ul> <li>53 subfolders (one for each PCB)</li> <li>'General_masks' folder for 'General' segmentation ground truth</li> <li>'Monoseg_masks' folder for 'Monoseg' segmentation ground truth</li> <li>'PCB_Masks' folder for masks of the 53 PCBs in the hyperspectral cube</li> </ul> </li> <li>RGB/ <ul> <li>53 .jpg images</li> <li>'General' folder for RGB images 'General' segmentation ground truth</li> <li>'Monoseg_masks' folder for RGB images 'Monoseg' segmentation ground truth</li> </ul> </li> </ul> </li> </ul> <p><strong>Data Classes in Masks</strong></p> <ul> <li>Masks (both 'General' and 'Monoseg') contain 1 to 4 segmentation classes: <ul> <li>0: "Others"</li> <li>1: "IC"</li> <li>2: "Capacitors"</li> <li>3: "Connectors"</li> </ul> </li> </ul> <p><strong>Code Repository</strong></p> <p>To facilitate reading and working with the data, Python codes are available on the GitHub repository:</p> <p>https://github.com/hifexplo/PCBVision</p> <p><strong>Citation</strong></p> <p>If you use this dataset, please cite the following article:</p> <p><strong>Word</strong>:</p> <p>Arbash, Elias, Fuchs, Margret, Rasti, Behnood, Lorenz, Sandra, Ghamisi, Pedram, & Gloaguen, Richard. (2024). PCB-Vision: A Multiscene RGB-Hyperspectral Benchmark Dataset of Printed Circuit Boards (Version 1) [Data set]. Rodare. <a href="http://doi.org/10.14278/rodare.2704">http://doi.org/10.14278/rodare.2704</a></p> <p><strong>Latex:</strong></p> <p>@article{arbash2024pcb, title={PCB-Vision: A Multiscene RGB-Hyperspectral Benchmark Dataset of Printed Circuit Boards}, author={Arbash, Elias and Fuchs, Margret and Rasti, Behnood and Lorenz, Sandra and Ghamisi, Pedram and Gloaguen, Richard}, journal={arXiv preprint arXiv:2401.06528}, year={2024} }</p> <p><strong>Contact</strong></p> <p>For further information or inquiries, please visit our website:</p> <p>https://www.iexplo.space/</p> <p>Contact Email: e.arbash@hzdr.de</p>
UnCoVar: Benchmarking dataset for SARS-CoV-2 sequence processing software pipelines, Sanger sequences
Open the record for dataset details and reuse information.
SpokeN-100: A Cross-Lingual Benchmarking Dataset for The Classification of Spoken Numbers in Different Languages
<div> <div>SpokeN-100 is a novel, entirely artificially generated benchmarking dataset tailored for speech recognition, representing a core challenge in the field of tiny deep learning. SpokeN-100 consists of spoken numbers from 0 to 99 spoken by 32 different speakers in four different languages, namely English, Mandarin, German and French, resulting in 12,800 audio samples.</div> </div>
HexaLCSeg: Hexagon-based Historical Land Cover Benchmark Dataset
<p>This dataset is a research outcome of a European Research Council, Proof of Concept Grant funded (Grant Number 101100837, A GeoAI-based Land Use Land Cover Segmentation Process to Analyse and Predict Rural Depopulation, Agricultural Land Abandonment, and Deforestation in Bulgaria and Turkey, 1940-2040, <a href="https://cordis.europa.eu/project/id/101100837">GeoAI_LULC_Seg</a>) project.</p> <p>We introduce a new benchmark dataset derived from very high-resolution historical Hexagon (KH-9) reconnaissance satellite images for use in deep learning-based image segmentation tasks. Our dataset comprises high-resolution monochromatic Hexagon images from the 1970s and 1980s covering Turkish and Bulgarian territories, encompassing a large geographic area.</p> <p>Land cover (LC) classes used in this study:<br>Our dataset is inspired by the European Space Agency (ESA) WorldCover project and includes eight LC classes and related RGB codes were set for each class but we adjusted the 0-pixel value as no data and replaced the 0 values with 1 in the ESA RGB code palette. Additionally, a new sub-class for the trees, named Permanent Cropland is defined and its RGB code was set to 1-207-117. This class is important to differentiate permanent fruit trees from other trees, specifically crucial for past agricultural mapping purposes.</p> <p>The HexaLCSeg dataset comprises eight panchromatic images accompanied by corresponding 3-channel RGB Ground Truth Masks, all with 8-bit radiometric resolution and a spatial resolution of 1 meter. The dataset is organized into a total of 10,000 patches, each sized at 256x256 pixels. We split our dataset into 70% training (7000 patches), 15% validation (1500 patches), and 15% testing (1500 patches).</p> <p>Methodology:<br>In our study, we employed the geographic object-based image analysis (GEOBIA) approach to generate accurate land cover (LC) maps, which serve as the ground truth masks for our dataset.</p> <p>For deep learning-based image segmentation, we employed a total of 9 CNN models, implementing U-Net++ and DeepLabv3+ segmentation architectures with different hyperparameters, paired with SE-ResNeXt50 backbone that pre-trained with weight values from the 2012 ILSVRC ImageNet dataset.</p> <p>Models, metric results and weights:</p> <table> <tbody> <tr> <th>Model No</th> <th>Architecture</th> <th>Loss Function</th> <th>Augmentation</th> <th>Loss</th> <th>Accuracy</th> <th>IoU</th> <th>F-1 Score</th> <th>Precision</th> <th>Recall</th> </tr> </tbody> <tbody> <tr> <td>Model 1</td> <td>U-Net++</td> <td>Focal Loss</td> <td>No Augmentation</td> <td>0.1252</td> <td>0.9734</td> <td>0.8052</td> <td>0.8804</td> <td>0.8805</td> <td>0.8803</td> </tr> <tr> <td>Model 2</td> <td>U-Net++</td> <td>Focal Loss</td> <td>Horizontal Flip</td> <td>0.1253</td> <td>0.9728</td> <td>0.8008</td> <td>0.8776</td> <td>0.8778</td> <td>0.8774</td> </tr> <tr> <td>Model 3</td> <td>DeepLabv3+</td> <td>Focal Loss</td> <td>No Augmentation</td> <td>0.1255</td> <td>0.9720</td> <td>0.7959</td> <td>0.8739</td> <td>0.8744</td> <td>0.8734</td> </tr> <tr> <td>Model 4</td> <td>U-Net++</td> <td>Focal Loss</td> <td>Random BC</td> <td>0.1256</td> <td>0.9717</td> <td>0.7938</td> <td>0.8725</td> <td>0.8727</td> <td>0.8723</td> </tr> <tr> <td>Model 5</td> <td>DeepLabv3+</td> <td>Dice Loss</td> <td>Horizontal Flip</td> <td>0.1292</td> <td>0.9714</td> <td>0.7928</td> <td>0.8714</td> <td>0.8717</td> <td>0.8711</td> </tr> <tr> <td>Model 6</td> <td>DeepLabv3+</td> <td>Dice Loss</td> <td>No Augmentation</td> <td>0.1307</td> <td>0.9711</td> <td>0.7906</td> <td>0.8699</td> <td>0.8702</td> <td>0.8697</td> </tr> <tr> <td>Model 7</td> <td>DeepLabv3+</td> <td>Focal Loss</td> <td>Horizontal Flip</td> <td>0.1257</td> <td>0.9711</td> <td>0.7897</td> <td>0.8698</td> <td>0.8704</td> <td>0.8692</td> </tr> <tr> <td>Model 8</td> <td>DeepLabv3+</td> <td>Focal Loss</td> <td>Random BC</td> <td>0.1259</td> <td>0.9704</td> <td>0.7871</td> <td>0.8667</td> <td>0.8673</td> <td>0.8662</td> </tr> <tr> <td>Model 9</td> <td>DeepLabv3+</td> <td>Dice Loss</td> <td>Random BC</td> <td>0.1401</td> <td>0.9691</td> <td>0.7793</td> <td>0.8608</td> <td>0.8612</td> <td>0.8604</td> </tr> </tbody> </table> <p> </p> <p>System-specific notes and configuration:</p> <p>The code was implemented in Python (3.10) Programming Language.</p> <p>- torch == 2.1.2<br>- segmentation-models-pytorch == 0.3.3<br>- Albumentations == 1.4.0</p> <p>Apart from main data science libraries, RS-specific libraries such as GDAL, rasterio, and tifffile are also required.</p> <p>Citation:<br>Please kindly cite our paper if this code and the dataset used in the study are useful for your research.</p> <p>Elif Sertel et al., “HexaLCSeg: A Historical Benchmark Dataset from Hexagon Satellite Images for Land Cover Segmentation [Software and Data Sets],” <em>IEEE Geoscience and Remote Sensing Magazine</em> 12, no. 3 (September 2024): 197–206, <a href="https://doi.org/10.1109/MGRS.2024.3394248">https://doi.org/10.1109/MGRS.2024.3394248</a>.</p>
Performance Measurement Dataset of the HPC Benchmarks FASTEST, Kripke, and RELeARN for Cost-Effective Modeling Analysis with Extra-P
<p>Performance Measurement Dataset of the HPC Benchmarks FASTEST, Kripke, RELeARN for Scalability Studies with Extra-P. This data was used to analyze cost-effective modeling approaches presented in the IPDPS 2020 paper "Learning Cost-Effective Sampling Strategies for Empirical Performance Modeling".</p>
FEater dataset: A molecular fragment dataset to benchmark the robustness of 3D flexible object recognition
<p>This dataset is associated with the work: Benchmarking the robustness of the correct identification of flexible 3D objects using common machine learning models</p> <pre><code># Original FEater-Single and FEater_Dual dataset. FEater_Single ├── TestSet_coord.h5 ├── TrainingSet_coord.h5 └── ValidationSet_coord.h5 FEater_Dual ├── TestSet_coord.h5 ├── TrainingSet_coord.h5 └── ValidationSet_coord.h5 # Non-redundant baseline dataset FEater_Baseline ├── TestSet_Dual.h5 ├── TestSet_Single.h5 ├── TrainingSet_Dual.h5 └── TrainingSet_Single.h5 # FEater-Single and FEater_Dual in different sample size FEater_Mini200 ├── Mini200_Dual.h5 └── Mini200_Single.h5 FEater_Mini400 ├── Mini400_Dual.h5 └── Mini400_Single.h5 FEater_Mini800 ├── Mini800_Dual.h5 └── Mini800_Single.h5</code></pre> <p>For further details of the usage, please visit the original GitHub repository: <a title="FEater_repo" href="https://github.com/miemiemmmm/FEater" target="_blank" rel="noopener">https://github.com/miemiemmmm/FEater</a></p>
SFC-A68: A dataset for benchmarking space function and space access element classification methods in entire floors of multi-unit apartment buildings
<p> </p> <p><strong>SFC-A68</strong>: "A dataset for benchmarking space function and space access element classification methods in entire floors of multi-unit apartment buildings"</p> <p><strong>Authors:</strong> "Amir Ziaee, Georg Suter, Laura Keiblinger"</p> <p><strong>Copyright:</strong> "Design Computing Group TU Wien, 2024"</p> <p><strong>Credits:</strong> "Design Computing Group TU Wien"</p> <p><strong>License:</strong> "GNU GENERAL PUBLIC LICENSE Version 3"</p> <p><strong>Version:</strong> "1.0.2"</p> <p><strong>Maintainer:</strong> "Amir Ziaee"</p> <p><strong>Email:</strong> "amir.ziaee@tuwien.ac.at"</p> <p><strong>Acknowledgments:</strong> "The authors gratefully acknowledge support by Grant Austrian Science Fund (FWF): I 5171-N, Laura Keiblinger, and participants in course '259.428-2021S Architectural Morphology' at TU Wien for data collection. "</p> <p><strong>Description:</strong> "The analysis of building models for usable area, building safety, and energy efficiency requires accurate classification data of spaces and related space elements such as doors. To reduce input model preparation effort and errors, automated classification of spaces and related space elements is desirable. We introduce the SFC-A68 dataset, which models entire floors of 275 multi-unit apartment buildings in 13 countries. The SFC-A68 consists of three data representation formats: tabular, graph, and multi-view image datasets that are derived from the same source data. It covere 22 space function and six space access element classes."</p> <p> </p> <p> </p> <p> </p>
Tiny Robotics Dataset and Benchmark for Continual Object Detection
<p>Dataset for <strong>TiROD</strong>: Tiny Robotics Dataset and Benchmark for Continual Object Detection<br><br>Official Website -> <a href="https://pastifra.github.io/TiROD/">https://pastifra.github.io/TiROD/</a></p> <p>Code -> <a href="https://github.com/pastifra/TiROD_code">https://github.com/pastifra/TiROD_code</a></p> <p>Video -> <a href="https://www.youtube.com/watch?v=e76m3ol1i4I">https://www.youtube.com/watch?v=e76m3ol1i4I</a></p> <p>Paper -> <a href="https://arxiv.org/abs/2409.16215">https://arxiv.org/abs/2409.16215</a></p>
Benchmark Multi-Omics Datasets for Methods Comparison
<p><strong>Pathway Multi-Omics Simulated Data</strong></p> <p>These are synthetic variations of the TCGA COADREAD data set (original data available at <a href="http://linkedomics.org/data_download/TCGA-COADREAD/">http://linkedomics.org/data_download/TCGA-COADREAD/</a>). This data set is used as a comprehensive benchmark data set to compare multi-omics tools in the manuscript "pathwayMultiomics: An R package for efficient integrative analysis of multi-omics datasets with matched or un-matched samples".</p> <p>There are 100 sets (stored as 100 sub-folders, the first 50 in "pt1" and the second 50 in "pt2") of random modifications to centred and scaled copy number, gene expression, and proteomics data saved as compressed data files for the R programming language. These data sets are stored in subfolders labelled "sim001", "sim002", ..., "sim100". Each folder contains the following contents: 1) "indicatorMatricesXXX_ls.RDS" is a list of simple triplet matrices showing which genes (in which pathways) and which samples received the synthetic treatment (where XXX is the simulation run label: 001, 002, ...), (2) "CNV_partitionA_deltaB.RDS" is the synthetically modified copy number variation data (where A represents the proportion of genes in each gene set to receive the synthetic treatment [partition 1 is 20%, 2 is 40%, 3 is 60% and 4 is 80%] and B is the signal strength in units of standard deviations), (3) "RNAseq_partitionA_deltaB.RDS" is the synthetically modified gene expression data (same parameter legend as CNV), and (4) "Prot_partitionA_deltaB.RDS" is the synthetically modified protein expression data (same parameter legend as CNV).</p> <p> </p> <p><strong>Supplemental Files</strong></p> <p>The file "cluster_pathway_collection_20201117.gmt" is the collection of gene sets used for the simulation study in Gene Matrix Transpose format. Scripts to create and analyze these data sets available at: <a href="https://github.com/TransBioInfoLab/pathwayMultiomics_manuscript_supplement">https://github.com/TransBioInfoLab/pathwayMultiomics_manuscript_supplement</a></p> <p> </p>
Simulated and Real datasets used for benchmarking MARGARET
<p>A collection of real and simulated datasets used for benchmarking MARGARET against other state-of-the-art TI methods</p>
Dataset for benchmarking Multiple Object Tracking and Segmentation (MOTS) in an apple orchard field.
<p>A dataset of temporally consistent apple images and labels taken using UAVs and a wearable sensor in an orchard, consisting of 86000 manually annotated apple instances and 1700 frames annotated in the MOTS (Multi-object Tracking and Segmentation) style.</p> <p>Sequence 0-5 are used for training. Sequence 6-8 are used for testing/validation. Sequence 10-12 are the testing datasets that have "ignore regions" overlays.</p> <p>The code used in the paper can be found on <a href="https://git.wur.nl/said-lab/rt-obj-tracking/">our GitLab.</a></p>
Benchmark Dataset (2D/3D) of an Industrial Rotary Kiln Combustion Chamber with Refuse-Derived Fuel Particles from a Light-Field-Camera
<p>Benchmark dataset for the detection of fuel particles (refuse-derived fuels - RDF) in 2D and 3D image data in a rotary kiln combustion chamber.</p> <p>Organization:<br> 01_Images: 50 Images (2D)<br> 02_Labels: Labeled ground truth image with rotary kiln, burner flame, burning particle in air, non-burning particle in air and particle on wall.<br> 03_Particle List: Lists of the coordinates of the center of gravity of particles in image coordinates.<br> 04_Point Cloud: 3D point cloud for the 50 images.<br> 05_Matlab: Code and visualization examples.<br> 06_All_Data: Images and 3D point cloud for 2010 images of a sequence containing the images with ground truth (see TXT file).</p>
MIRA benchmarking Frankencell datasets
<p>MIRA benchmarking Frankencell datasets, scaffolds, and configuration file for regeneration.</p>
CrossLoc Benchmark Datasets
<p><span>To study the data-scarcity mitigation for learning-based visual localization methods via sim-to-real transfer, we curate and now present the </span><em>CrossLoc benchmark datasets</em><span>—a multimodal aerial sim-to-real data available for flights above nature and urban terrains. Unlike the previous computer vision datasets focusing on localization in a single domain (mostly real RGB images), the provided benchmark datasets include various multimodal synthetic cues paired to all real photos. Complementary to the paired real and synthetic data, we offer rich synthetic data that efficiently fills the flight envelope volume in the vicinity of the real data. </span></p> <p> </p> <p><span>The synthetic data rendering was achieved using the proposed data generation workflow TOPO-DataGen. The provided CrossLoc datasets were used as an initial benchmark to showcase the use of synthetic data to assist visual localization in the real world with limited real data.</span></p> <p><span>Please refer to our main paper at https://arxiv.org/abs/2112.09081 and our code at https://github.com/TOPO-EPFL/CrossLoc for details.</span></p>
A Social and News Media Benchmark Dataset for Topic Modeling
<p>A novel approach to topic modeling using PSO-based clustering is compared to traditional techniques using the 20 Newsgroups dataset and a collection of posts from the Reddit health forum r/Cancer.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.