Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

419

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

419 results for “large dataset”

Learn how ShareScore rates datasets ↗
zenodo44/100

Dataset - Survey results - Applying Model-based Requirements Engineering in Three Large European Collaborative Projects

<p>This dataset and its associated report contain the results of an online survey on using a&nbsp;model-based requirements engineering approach in three European projects.&nbsp;</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Dataset for "Large scale patterns and drivers of the diving behavior of gill-breathing large pelagic predators"

<p>This dataset includes all supporting data and scritps to generate figure panels in the paper "Large scale patterns and drivers of the diving behavior of gill-breathing large pelagic predators" (A. Nuno, J. Guiet, B. Baranek and D. Bianchi)</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Rigid and hinged very large floating structure (VLFS) dataset - Kelvin Hydrodynamics Laboratory

<p>This dataset corresponds to the measurements performed at the Kelvin Hydrodynamics Laboratory at the University of Strathclyde, in August 2022, to assess the motion performance and internal loading of a rigid and hinged very large floating structure (VLFS) under regular waves. The VLFS was constructed with three pontoons and two hinges. The hinges were replaced with aluminium steel bars to built the rigid VLFS. The dimensions of each pontoon of the VLFS were 580 mm x 580 mm x 52 mm. Each pontoon was built with 2 mm layer of carbon fibre and a 50 mm layer of PVC foam.</p><p>The VLFS was tested in regular waves at two incidences: 0 degrees and 30 degrees. For 0 degree incidence, the wave frequencies tested ranged from 0.4 to1.6 Hz in intervals of 0.1 Hz. Four wave heights were tested, h=5, 10, 20 and 40 mm. For 30 degree incidence, the same range of frequencies were tested, but only one wave height, h=5 mm. Preliminary results for some of the data at 0 degrees incidence can be found in&nbsp;the conference paper: https://doi.org/10.36688/ewtec-2023-389.&nbsp; Further analysis of this dataset and additional results are in preparation for a journal manuscript.</p><p>The following files are included as part of the dataset:</p><ol><li>Motion files (Matlab files).</li><li>Strain gauge and wave height files (Matlab files).</li><li>Data description file - Description of files.</li><li>Test matrix - Test cases summarised with nomenclature used in files.</li><li>Matlab script to sort out position of motion spheres as depicted in Figure 1.</li><li>Video of the hinged VLFS subject to a train of regular waves at f=0.8 Hz, i.e. when the wavelength is of similar length to the length of the platform, i.e. f=0.8 Hz.</li></ol><ul><li>The motion files contain the time series information recorded for each of the motion detection spheres. Because the motion raw data is not labelled sequentially, it is necessary to run the Matlab file included in the data repository to sort out the information of the spheres.</li><li>The strain gauge files contain the raw strain gauge data (8 channels) and the wave gauge data with the file number describing the corresponding test in the test matrix.</li></ul><p>The VLFS was equipped with 36 motion detection spheres and 8 strain gauges. The diagram and notation of each sphere is depicted in Figure 1. Figure 1 is available in the Data description document.</p><p>&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

A Large Scale Side-Scan Sonar Dataset of Seafloor Sediments for Self-Supervised Pretraining

<p>This dataset serves as an extension to the dataset part of "A convolutional vision transformer for semantic segmentation of side-scan sonar data" published in Ocean Engineering, Volume 86, part 2, 15 October 2023,<strong> </strong>DOI: <a href="https://www.sciencedirect.com/science/article/pii/S0029801823020310">10.1016/j.oceaneng.2023.115647</a> for self-supervised pretraining.</p><p>This dataset consists of patches of side-scan sonar waterfalls collected along the coast of Catalunya during an extensive survey. The waterfalls were partitioned in batches of 384 lines to generate images of size 384 × 384 with a 192 pixel-overlap along-track and across-track. This resulted in a total of 434,164 images capturing various seafloor types including rocky bottoms, sand ripples, detrital funds, posidonia, cymocea, mud, corals, artificial reefs etc.</p><p>Additional tools for using the data for self-supervised pretraining can be found under <a href="https://github.com/DeeperSense/deepersense-seafloorscan">https://github.com/DeeperSense/deepersense-seafloorscan</a></p><p>&nbsp;</p><p><strong>Acknowledgements</strong></p><p>The data in this repository were collected by Tecnoambiente SL as part of the project DeeperSense that received funding from the European Commission. Program H2020-ICT-2020-2 ICT-47-2020. Project Number: 101016958.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Metadata of a Large Sonar and Stereo Camera Dataset Suitable for Sonar-to-RGB Image Translation

<h1>Metadata of a Large Sonar and Stereo Camera Dataset Suitable for Sonar-to-RGB Image Translation</h1> <h2>Introduction</h2> <p>This is a set of metadata describing a large dataset of synchronized sonar and stereo camera recordings, that were captured between August 2021 and September 2023 during the project <a href="https://robotik.dfki-bremen.de/en/research/projects/deepersense/">DeeperSense</a> (https://robotik.dfki-bremen.de/en/research/projects/deepersense/), as training data for Sonar-to-RGB image translation. <a href="../records/7728089">Parts</a> <a href="../records/10220989">of</a> the sensor data have been published (https://zenodo.org/records/7728089, https://zenodo.org/records/10220989). Due to the size of the sensor data corpus, it is currently impractical to make the entire corpus accessible online. Instead, this metadatabase serves as a relatively compact representation, allowing interested researchers to inspect the data, and select relevant portions for their particular use case, which will be made available on demand. This is an effort to comply with the <a href="https://www.go-fair.org/fair-principles/">FAIR</a> principle A2 (https://www.go-fair.org/fair-principles/) that metadata shall be accessible, even when the base data is not immediately.</p> <h3>Locations and sensors</h3> <p>The sensor data was captured at four different locations, including one laboratory (Maritime Exploration Hall at DFKI RIC Bremen) and three field locations (Chalk Lake Hemmoor, Tank Wash Basin Neu-Ulm, Lake Starnberg). At all locations, a ZED camera and a Blueprint Oculus M1200d sonar were used. Additionally, a SeaVision camera was used at the Maritime Exploration Hall at DFKI RIC Bremen and at the Chalk Lake Hemmoor. The <code>examples/</code> directory holds a typical output image for each sensor at each available location.</p> <h3>Data volume per session</h3> <p>Six data collection sessions were conducted. The table below presents an overview of the amount of data captured in each session:</p> <table> <tbody> <tr> <th>Session dates</th> <th>Location</th> <th>Number of datasets</th> <th>Total duration of datasets [h]</th> <th>Total logfile size [GB]</th> <th>Number of images</th> <th>Total image size [GB]</th> </tr> <tr> <td>2021-08-09 - 2021-08-12</td> <td>Maritime Exploration Hall at DFKI RIC Bremen</td> <td>52</td> <td>10.8</td> <td>28.8</td> <td>389&rsquo;047</td> <td>88.1</td> </tr> <tr> <td>2022-02-07 - 2022-02-08</td> <td>Maritime Exploration Hall at DFKI RIC Bremen</td> <td>35</td> <td>4.4</td> <td>54.1</td> <td>629&rsquo;626</td> <td>62.3</td> </tr> <tr> <td>2022-04-26 - 2022-04-28</td> <td>Chalk Lake Hemmoor</td> <td>52</td> <td>8.1</td> <td>133.6</td> <td>1&rsquo;114&rsquo;281</td> <td>97.8</td> </tr> <tr> <td>2022-06-28 - 2022-06-29</td> <td>Tank Wash Basin Neu-Ulm</td> <td>42</td> <td>6.7</td> <td>144.2</td> <td>824&rsquo;969</td> <td>26.9</td> </tr> <tr> <td>2023-04-26 - 2023-04-27</td> <td>Maritime Exploration Hall at DFKI RIC Bremen</td> <td>55</td> <td>7.4</td> <td>141.9</td> <td>739&rsquo;613</td> <td>9.6</td> </tr> <tr> <td>2023-09-01 - 2023-09-02</td> <td>Lake Starnberg</td> <td>19</td> <td>2.9</td> <td>40.1</td> <td>217&rsquo;385</td> <td>2.3</td> </tr> <tr> <th>&nbsp;</th> <th>&nbsp;</th> <th>255</th> <th>40.3</th> <th>542.7</th> <th>3&rsquo;914&rsquo;921</th> <th>287.0</th> </tr> </tbody> </table> <h2>Data and metadata structure</h2> <h3>Sensor data corpus</h3> <p>The sensor data corpus comprises two processing stages:</p> <ul> <li>raw data streams stored in ROS bagfiles (aka <strong>logfiles</strong>),</li> <li>camera and sonar images (aka <strong>datafiles</strong>) extracted from the logfiles.</li> </ul> <p>The files are stored in a file tree hierarchy which groups them by session, dataset, and modality:</p> <pre><code>${session_key}/ ${dataset_key}/ ${logfile_name} ${modality_key}/ ${datafile_name}</code></pre> <p>A typical logfile path has this form:</p> <pre><code>2023-09_starnberg_lake/ 2023-09-02-15-06_hydraulic_drill/ stereo_camera-zed-2023-09-02-15-06-07.bag</code></pre> <p>A typical datafile path has this form:</p> <pre><code>2023-09_starnberg_lake/ 2023-09-02-15-06_hydraulic_drill/ zed_right/ 1693660038_368077993.jpg</code></pre> <p>All directory and file names, and their particles, are designed to serve as identifiers in the metadatabase. Their formatting, as well as the definitions of all terms, are documented in the file <code>entities.json</code>.</p> <h3>Metadatabase</h3> <p>The metadatabase is provided in two equivalent forms:</p> <ul> <li>as a standalone <a href="https://www.sqlite.org/index.html">SQLite</a> (https://www.sqlite.org/index.html) database file <code>metadata.sqlite</code> for users familiar with SQLite,</li> <li>as a collection of CSV files in the <code>csv/</code> directory for users who prefer other tools.</li> </ul> <p>The database file has been generated from the CSV files, so each database table holds the same information as the corresponding CSV file. In addition, the metadatabase contains a series of convenience views that facilitate access to certain aggregate information.</p> <p>An entity relationship diagram of the metadatabase tables is stored in the file <code>entity_relationship_diagram.png</code>. Each entity, its attributes, and relations are documented in detail in the file <code>entities.json</code></p> <p>Some general design remarks:</p> <ul> <li>For convenience, timestamps are always given in both a human-readable form (ISO 8601 formatted datetime strings with explicit local time zone), and as seconds since the UNIX epoch.</li> <li>In practice, each logfile always contains a single stream, and each stream is stored always in a single logfile. Per database schema however, the entities <code>stream</code> and <code>logfile</code> are modeled separately, with a &ldquo;many-streams-to-one-logfile&rdquo; relationship. This design was chosen to be compatible with, and open for, data collections where a single logfile contains multiple streams.</li> <li>A <code>modality</code> is not an attribute of a <code>sensor</code> alone, but of a <code>datafile</code>: Because a <code>sensor</code> is an attribute of a <code>stream</code>, and a single stream may be the source of multiple modalities (e.g.&nbsp;RGB vs.&nbsp;grayscale images from the same camera, or cartesian vs.&nbsp;polar projection of the same sonar output). Conversely, the same modality may originate from different sensors.</li> </ul> <p>As a usage example, the data volume per session which is tabulated at the top of this document, can be extracted from the metadatabase with the following SQL query:</p> <div> <pre><code><span><span>SELECT</span></span> <span> PRINTF(</span> <span> <span>'%s - %s'</span>,</span> <span> <span>SUBSTR</span>(session_start, <span>1</span>, <span>10</span>),</span> <span> <span>SUBSTR</span>(session_end, <span>1</span>, <span>10</span>)) <span>AS</span> <span>'Session dates'</span>,</span> <span> location_name_english <span>AS</span> Location,</span> <span> number_of_datasets <span>AS</span> <span>'Number of datasets'</span>,</span> <span> total_duration_of_datasets_h <span>AS</span> <span>'Total duration of datasets [h]'</span>,</span> <span> total_logfile_size_gb <span>AS</span> <span>'Total logfile size [GB]'</span>,</span> <span> number_of_images <span>AS</span> <span>'Number of images'</span>,</span> <span> total_image_size_gb <span>AS</span> <span>'Total image size [GB]'</span></span> <span><span>FROM</span></span> <span> location</span> <span> <span>JOIN</span> <span>session</span> <span>USING</span> (location_id)</span> <span> <span>JOIN</span> (</span> <span> <span>SELECT</span></span> <span> session_id,</span> <span> <span>COUNT</span>(dataset_id) <span>AS</span> number_of_datasets,</span> <span> <span>ROUND</span>(</span> <span> <span>SUM</span>(dataset_duration) <span>/</span> <span>3600</span>,</span> <span> <span>1</span>) <span>AS</span> total_duration_of_datasets_h,</span> <span> <span>ROUND</span>(</span> <span> <span>SUM</span>(total_logfile_size) <span>/</span> <span>10e9</span>,</span> <span> <span>1</span>) <span>AS</span> total_logfile_size_gb</span> <span> <span>FROM</span></span> <span> location</span> <span> <span>JOIN</span> <span>session</span> <span>USING</span> (location_id)</span> <span> <span>JOIN</span> dataset <span>USING</span> (session_id)</span> <span> <span>JOIN</span> view__dataset_total_logfile_size <span>USING</span> (dataset_id)</span> <span> <span>GROUP</span> <span>BY</span></span> <span> session_id</span> <span> ) <span>USING</span> (session_id)</span> <span> <span>JOIN</span> (</span> <span> <span>SELECT</span></span> <span> session_id,</span> <span> <span>COUNT</span>(datafile_id) <span>AS</span> number_of_images,</span> <span> <span>ROUND</span>(<span>SUM</span>(datafile_size) <span>/</span> <span>10e9</span>, <span>1</span>) <span>AS</span> total_image_size_gb</span> <span> <span>FROM</span></span> <span> <span>session</span></span> <span> <span>JOIN</span> dataset <span>USING</span> (session_id)</span> <span> <span>JOIN</span> stream <span>USING</span> (dataset_id)</span> <span> <span>JOIN</span> <span>datafile</span> <span>USING</span> (stream_id)</span> <span> <span>GROUP</span> <span>BY</span></span> <span> session_id</span> <span> ) <span>USING</span> (session_id)</span> <span><span>ORDER</span> <span>BY</span> session_id;</span></code></pre> </div>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Large-scale attributed graph & hypergraph datasets: TWeibo, Amazon2M, Amazon, MAG-PM

<p>Here we provide additional large-scale datasets used in our work "A Versatile Framework for Attributed Network Clustering via K-Nearest Neighbor Augmentation", along with the index files for constructing KNN graphs using ScaNN and Faiss.</p> <p>Usage:</p> <p>cd ANCKA/</p> <p>unzip ~/Download_path/ANCKA_data.zip -d data/</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Dataset - Papyrus 05.4 - A large scale curated dataset aimed at bioactivity predictions

<div> <p>This repository contains the version 05.5 of the Papyrus dataset, an aggregated dataset of small molecule bioactivities, as described in the article "Papyrus - A large scale curated dataset aimed at bioactivity predictions" <a href="https://doi.org/10.26434/chemrxiv-2021-1rxhk">https://doi.org/10.26434/chemrxiv-2021-1rxhk</a>.</p> <p>&nbsp;</p> <p>With the ongoing rapid growth of publicly available ligand-protein bioactivity data, there is a trove of valuable data that can be used to train a plethora of machine learning algorithms. However, not all data is equal in terms of size and quality and a significant portion of researchers&rsquo; time is needed to adapt the data to their needs. On top of that, finding the right data for a research question can often be a challenge on its own. To meet these challenges we have constructed the Papyrus dataset. Papyrus is comprised of around 60 million datapoints. This dataset contains multiple large publicly available datasets such as ChEMBL and ExCAPE-DB combined with several smaller datasets containing high-quality data. The aggregated data has been standardised and normalised in a manner that is suitable for machine learning. We show how data can be filtered in a variety of ways and also perform some example quantitative structure-activity relationship analyses and proteochemometric modelling. Our ambition is that this pruned data collection constitutes a benchmark set that can be used for constructing predictive models, while also providing a solid baseline for related research.</p> </div>

opencc-by-sa-4.0Apr 2022View details →
zenodo44/100

Accompanying data - Papyrus - A large scale curated dataset aimed at bioactivity predictions

<p><strong>Addition of supporting files:<br>- </strong>LICENSE.txt<strong><br>- </strong>data_types.json<strong><br>- </strong>data_size.json</p> <p>&nbsp;</p> <p><strong>Fixed version of Papyrus++ 05.5:<br>- In the previous 05.5 version&nbsp;</strong>data was incorrectly&nbsp;duplicated based on assay type. This resulted in unintended data augmentation.<br><strong>- In this&nbsp;fixed 05.5 version</strong>&nbsp;the duplicates have been eliminated, now reporting the correct amount of data per assay type.</p> <p>&nbsp;</p> <p>This repository contains the version 05.5 of the Papyrus dataset, an aggregated dataset of small molecule bioactivities, as described in the article "Papyrus - A large scale curated dataset aimed at bioactivity predictions" <a href="http://doi.org/10.1186/s13321-022-00672-x">http://doi.org/10.1186/s13321-022-00672-x</a>.</p> <p>&nbsp;</p> <p>With the ongoing rapid growth of publicly available ligand-protein bioactivity data, there is a trove of valuable data that can be used to train a plethora of machine learning algorithms. However, not all data is equal in terms of size and quality and a significant portion of researchers&rsquo; time is needed to adapt the data to their needs. On top of that, finding the right data for a research question can often be a challenge on its own. To meet these challenges we have constructed the Papyrus dataset. Papyrus is comprised of around 60 million datapoints. This dataset contains multiple large publicly available datasets such as ChEMBL and ExCAPE-DB combined with several smaller datasets containing high-quality data. The aggregated data has been standardised and normalised in a manner that is suitable for machine learning. We show how data can be filtered in a variety of ways and also perform some example quantitative structure-activity relationship analyses and proteochemometric modelling. Our ambition is that this pruned data collection constitutes a benchmark set that can be used for constructing predictive models, while also providing a solid baseline for related research.</p>

opencc-by-sa-4.0Aug 2022View details →
zenodo44/100

FoodSky: A Food-oriented Large Language Model, and FoodEarth: A Foundamental Food Corpus and Instruction Dataset

<p>Food is the cornerstone of both survival and social life. With the increasing complexity of global dietary needs and preferences, there is a growing demand for food intelligence to enable tasks like recipe recommendation and diet-disease correlation discovery. To address this, we introduce the Food-oriented Large Language Model (LLM) FoodSky, which offers fine-grained perception and reasoning of food data. We constructed a food corpus, FoodEarth, from various authoritative sources to enhance FoodSky's knowledge. We also developed the Topic-based Selective State Space Model and Hierarchical Topic Retrieval Augmented Generation algorithms to improve FoodSky's ability to capture fine-grained food semantics and generate context-aware food-relevant text. Extensive experiments show that FoodSky outperforms general-purpose LLMs on the Chinese National Chef Exam and Dietetic Exam, achieving accuracies of 67.2% and 66.4%, respectively. FoodSky not only enhances culinary creativity and promotes healthier eating patterns but also establishes a new standard for domain-specific LLMs tackling real-world food-related issues.</p>

opencc-zeroSep 2024View details →
zenodo44/100

Large-scale 3D building and tree datasets constructed from airborne LiDAR point clouds in Glasgow, UK

<p>This is the updated version of building 3D model data. The revision includes appending attributes to the lod1 and lod2 shapefile and creating cityjson file for each 3D building model. All 3D building models are available in mesh (.obj), multipath shapefile, and cityjson (.json) now.</p> <p><strong>IMPORTANT NOTE: We suggest using the building footprint, lod1, and lod2 data of this version (Version v4).</strong></p> <p>Urban Big Data Centre of the University of Glasgow generates 3D city models via the airborne LiDAR point clouds acquired between 2020-2021 on behalf of Glasgow City Council. It is a large-scale 3D city model containing 3D information on terrain, trees, and buildings in Glasgow City. This dataset comprises terrain, tree canopy, and building products derived from high-density airborne LiDAR point clouds.&nbsp;</p> <p>The terrain products include Digital Terrain Model (DTM), Digital Surface Model (DSM), and normalized Digital Surface Model (nDSM) in 0.5 m spatial resolution. The DTM and DSM rasters were provided by the vendor and nDSM rasters were obtained by subtracting DTM from DSM. Terrain products are provided in 5 km by 5 km GeoTIF format raster.</p> <p>The tree canopy products are composed of canopy height models (CHM) and tree top locations. Classified tree point clouds were applied with pit-free algorithm to generate CHM in 0.5 m grid raster in GeoTIF format [1]-[2]. Treetop locations were identified by using Local Maximum Filter based on CHM and are recorded as points in Shapefile format. The tree canopy products are provided in 5 km by 5 km tiles.</p> <p>Building 3D model products include footprint polygons with building height attributes and 3D mesh of building models in LoD1 and LoD2 levels. A series of processes such as converting building point clouds to building height models (BHM), converting BHM to polygons, and polygon regularization were conducted to obtain the building footprint polygons. Building height attributes were calculated from BHM for each footprint. The building footprint data are provided in Shapefile format. LoD1 models were generated based on the footprint and average height of the building. LoD2 models were constructed based on footprint and building point cloud with City3D tool[3]. LoD1 and LoD2 models are provided in OBJ and shapefile format. Building 3D model products are provided in 5 km by 5 km tiles. The RMSE of Euclidean distances between each point in the point cloud to the reconstructed model was calculated to evaluate the LoD2 model construction. A table of RMSE and a note for a few problematic models are provided.</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

WikiMed and PubMedDS: Two large-scale datasets for medical concept extraction and normalization research

<p>Two large-scale, automatically-created datasets of medical concept mentions, linked to the <a href="https://uts.nlm.nih.gov/uts/umls/home">Unified Medical Language System (UMLS)</a>.</p> <p><strong>WikiMed</strong></p> <p>Derived from Wikipedia data. Mappings of Wikipedia page identifiers to UMLS Concept Unique Identifiers (CUIs) was extracted by crosswalking Wikipedia, Wikidata, Freebase, and the NCBI Taxonomy to reach existing mappings to UMLS CUIs. This created a 1:1 mapping of approximately 60,500 Wikipedia pages to UMLS CUIs. Links to these pages were then extracted as mentions of the corresponding UMLS CUIs.</p> <p>WikiMed contains:</p> <ul> <li>393,618 Wikipedia page texts</li> <li>1,067,083 mentions of medical concepts</li> <li>57,739 unique UMLS CUIs</li> </ul> <p>Manual evaluation of 100 random samples of WikiMed found 91% accuracy in the automatic annotations at the level of UMLS CUIs, and 95% accuracy in terms of semantic type.</p> <p><strong>PubMedDS</strong></p> <p>Derived from biomedical literature abstracts from <a href="https://pubmed.ncbi.nlm.nih.gov/">PubMed</a>. Mentions were automatically identified using distant supervision based on Medical Subject Heading (MeSH) headers assigned to the papers in PubMed, and recognition of medical concept mentions using the high-performance <a href="https://allenai.github.io/scispacy/">scispaCy</a> model. MeSH header codes are included as well as their mappings to UMLS CUIs.</p> <p>PubMedDS contains:</p> <ul> <li>13,197,430 abstract texts</li> <li>57,943,354 medical concept mentions</li> <li>44,881 unique UMLS CUIs</li> </ul> <p>Comparison with existing manually-annotated datasets (NCBI Disease Corpus, BioCDR, and MedMentions) found 75-90% precision in automatic annotations. Please note this dataset is&nbsp;<em>not&nbsp;</em>a comprehensive annotation of medical concept mentions in these abstracts (only mentions located through distant supervision from MeSH headers were included), but is intended as data for <em>concept n</em><em>ormalization</em>&nbsp;research.</p> <p>Due to its size, PubMedDS is distributed as 30 individual files of approximately 1.5 million mentions each.</p> <p><strong>Data format</strong></p> <p>Both datasets use JSON format with one document per line. Each document has the following structure:</p> <pre><code class="language-json">{ "_id": "A unique identifier of each document", "text": "Contains text over which mentions are ", "title": "Title of Wikipedia/PubMed Article", "split": "[Not in PubMedDS] Dataset split: &lt;train/test/valid&gt;", "mentions": [ { "mention": "Surface form of the mention", "start_offset": "Character offset indicating start of the mention", "end_offset": "Character offset indicating end of the mention", "link_id": "UMLS CUI. In case of multiple CUIs, they are concatenated using '|', i.e., CUI1|CUI2|..." }, {} ] }</code></pre> <p><strong>Version history</strong></p> <table align="left"> <thead> <tr> <th scope="col">Version</th> <th scope="col">Notes</th> </tr> </thead> <tbody> <tr> <td>1.0.0</td> <td>Initial release</td> </tr> <tr> <td>1.0.1</td> <td>Corrected duplication error in WikiMed.zip file</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

BirdVox-296h: a large-scale dataset for detection and classification of flight calls

<p>BirdVox 296 hours dataset (BirdVox-296h)<br> ====================================</p> <p>Version 2.1, May 2022.</p> <p><br> Created By<br> ----------</p> <p>Andrew Farnsworth (1), Benjamin Mark Van Doren (1), Steve Kelling (1), Vincent Lostanlen (2), Justin Salamon (3), Aurora Cramer (4), Juan Pablo Bello (4)</p> <p>(1): Cornell Lab of Ornithology (CLO)<br> (2): Laboratoire des Sciences du Num&eacute;rique de Nantes (LS2N), CNRS<br> (3): Adobe Research<br> (4): New York University</p> <p>https://wp.nyu.edu/birdvox<br> <br> &nbsp;</p> <p>Description<br> ---------------</p> <p>The BirdVox-296h dataset contains 148 audio recordings, each two hours in duration. These recordings come from ROBIN autonomous recording units, placed near Ithaca, NY, USA during the fall 2015. They were captured by nine different sensors, originally numbered 1, 2, 3, 4, 5, 6, 7, 8, and 10.<br> <br> Ornithologist Andrew Farnsworth used the Raven software to pinpoint and label every avian flight call in time and frequency. He found 26138 sound events, of which 21546 are flight calls from Passeriformes. Of those, 13385 are identifiable in terms of family, and 8669 are identifiable in terms of both family and species. The annotation process took over 600 hours.</p> <p>The dataset can be used, among other things, for the research, development and testing of machine listening models for bird migration monitoring.</p> <p>&nbsp;</p> <p>Data Files<br> ------------</p> <p>The BirdVox-296h_wav folder contains 148 recordings as WAV files, sampled at 24 kHz, with a single channel (mono). Each recording lasts exactly two hours and is named according to the following format:</p> <p>YYYY-MM-DD_hh-mm-ss_unitUU.wav</p> <p>Where Y means Year, M means Month, D means Day, h means hour, m means minute, and s means second. This date format corresponds to the start time of the recording file, expressed in Coordinated Universal Time (UTC).</p> <p>The field UU contains two digits corresponding to the identifier of the autonomous recording unit (i.e., bioacoustic sensor). UU is either equal to 01, 02, 03, 04, 05, 06, 07, 08, or 10. Note that 09 is absent from the list because sensor 09 failed during the acquisition campaign.</p> <p>&nbsp;</p> <p>Metadata Files<br> -------------------</p> <p>The BirdVox-296h_csv-annotations folder contains CSV files, one for each audio file. The columns of each CSV file are:</p> <p>ID,Time (s),Frequency (Hz),Taxonomy Code,Fine Label,Medium Label,Coarse Label</p> <p><br> &quot;Taxonomy Code&quot; is compliant with the BirdVoxClassify software: github.com/BirdVox/BirdVoxClassify</p> <p>&quot;Fine Label&quot;, &quot;Medium Label&quot;, and &quot;Coarse Label&quot; most often correspond to species, family and order respectively.</p> <p>&nbsp;</p> <p>The BirdVox-296h_gps-coordinates.csv file contains the approximate GPS coordinates of the sensors (latitudes and longitudes rounded to 2 decimal points) of all nine sensors.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>Conditions of Use<br> -----------------</p> <p>Dataset created by Andrew Farnsworth, Steve Kelling, Vincent Lostanlen, Justin Salamon, Aurora Cramer, and Juan Pablo Bello.</p> <p>The BirdVox-full-night dataset is offered free of charge under the terms of the Creative &nbsp;Commons Attribution 4.0 International (CC BY 4.0) license:<br> https://creativecommons.org/licenses/by/4.0/</p> <p>The dataset and its contents are made available on an &quot;as is&quot; basis and without &nbsp;warranties of any kind, including without limitation satisfactory quality and &nbsp;conformity, merchantability, fitness for a particular purpose, accuracy or &nbsp;completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, Cornell Lab of Ornithology is not liable for, and expressly excludes all liability for, loss or damage however and whenever caused to anyone by any use of the BirdVox-full-night dataset or any part of it.</p> <p>&nbsp;</p> <p>Feedback<br> -------------</p> <p>Please help us improve BirdVox-296h by sending your feedback to:<br> vincent.lostanlen@ls2n.fr and af27@cornell.edu</p> <p>In case of a problem, please include as many details as possible.</p> <p>&nbsp;</p> <p>Acknowledgements<br> --------------------------</p> <p>Jessie Barry, Ian Davies, Tom Fredericks, Jeff Gerbracht, Sara Keen, Holger Klinck, Anne Klingensmith, Ray Mack, Peter Marchetto, Ed Moore, Matt Robbins, Ken Rosenberg, and Chris Tessaglia-Hymes.</p> <p>We acknowledge that the land on which the data was collected is the unceded territory of the Cayuga nation, which is part of the Haudenosaunee (Iroquois) confederacy.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

TBGA: A Large-Scale Gene-Disease Association Dataset for Biomedical Relation Extraction

<p>This repository contains the TBGA dataset. TBGA is a large-scale, semi-automatically annotated&nbsp;dataset&nbsp;for Gene-Disease Association (GDA) extraction. The dataset consists of three text files, corresponding to train, validation, and test sets, plus an additional JSON file containing the mapping between relation names and IDs. Each record in train, validation, or test files&nbsp;corresponds to a single GDA extracted from a sentence. Records are represented as JSON objects with the following structure:</p> <ul> <li><strong>text:</strong>&nbsp;sentence from which the GDA was extracted.</li> <li><strong>relation:</strong>&nbsp;relation name associated with the given GDA.</li> <li><strong>h:&nbsp;</strong>JSON object representing the gene entity, composed of: <ul> <li><strong>id:&nbsp;</strong>NCBI Entrez ID associated with the gene entity.</li> <li><strong>name:</strong>&nbsp;NCBI official gene symbol associated with&nbsp;the gene entity.</li> <li><strong>pos:&nbsp;</strong>list consisting of starting position and length of the gene mention within text.</li> </ul> </li> <li><strong>t:</strong>&nbsp;JSON object representing the disease entity, composed of: <ul> <li><strong>id:&nbsp;</strong>UMLS CUI associated with the disease entity.</li> <li><strong>name:</strong>&nbsp;UMLS preferred term associated with the disease entity.</li> <li><strong>pos:</strong>&nbsp;list consisting of starting position and length of the disease mention within text.</li> </ul> </li> </ul> <p>TBGA contains over 200,000 instances and 100,000 bags.<br> The zip file consists of one folder, named TBGA,&nbsp;containing the files corresponding to the dataset.</p> <p>If you use or extend our work, please cite the following:&nbsp;https://bmcbioinformatics.biomedcentral.com/articles/10.1186/s12859-022-04646-6#citeas<br> TBGA paper can be found at:&nbsp;<a href="https://rdcu.be/cKkY2">https://rdcu.be/cKkY2</a><br> TBGA code is available at:&nbsp;https://github.com/GDAMining/gda-extraction</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Smart Analyser of Variability Requirements of Unknown Spaces (SAVRUS) Dataset of a study with 5 real-world large numerical variability models.

<p>The publications and research associated to cite is in:</p><p><a href="https://doi.org/10.1016/j.knosys.2023.110558">https://doi.org/10.1016/j.knosys.2023.110558</a></p><p>In that research we detail the Smart Analyser of Variability Requirements of Unknown Spaces (SAVRUS) approach, and provide a web-tool prototype in <a href="https://hadas.caosd.lcc.uma.es/savrus">https://hadas.caosd.lcc.uma.es/savrus</a></p><p>In the study, we model 5 different real-world software product lines to then analysed them with SAVRUS:</p><p>Detailed real-world variability models ordered by their search space size, of which GEC QA is incompletely measured NVM Description #Booleans #Numericals Space QA #Measurements&nbsp;</p><p>Dune1</p><p>&nbsp;</p><p>Multi-grid solver</p><p>&nbsp;</p><p>11</p><p>&nbsp;</p><p>3</p><p>&nbsp;</p><p>2,304</p><p>&nbsp;</p><p>Complex..</p><p>&nbsp;</p><p>2,304</p><p>&nbsp;</p><p>HSMGP1</p><p>&nbsp;</p><p>Stencil-grid solver</p><p>&nbsp;</p><p>14</p><p>&nbsp;</p><p>3</p><p>&nbsp;</p><p>3,456</p><p>&nbsp;</p><p>..equation..</p><p>&nbsp;</p><p>3,456</p><p>&nbsp;</p><p>HiPAcc1</p><p>&nbsp;</p><p>Image processing framework</p><p>&nbsp;</p><p>33</p><p>&nbsp;</p><p>2</p><p>&nbsp;</p><p>13,485</p><p>&nbsp;</p><p>..solving..</p><p>&nbsp;</p><p>13,485</p><p>&nbsp;</p><p>Trimesh2</p><p>&nbsp;</p><p>Triangle mesh library</p><p>&nbsp;</p><p>13</p><p>&nbsp;</p><p>4</p><p>&nbsp;</p><p>239,360</p><p>&nbsp;</p><p>..time</p><p>&nbsp;</p><p>239,360</p><p>&nbsp;</p><p>GEC</p><p>&nbsp;</p><p>Generic edge computing</p><p>&nbsp;</p><p>552</p><p>&nbsp;</p><p>2</p><p>&nbsp;</p><p>~5.3*108</p><p>&nbsp;</p><p>Energy Consumption</p><p>&nbsp;</p><p>132500</p><p>&nbsp;</p><p>The dataset zip file contains:</p><ul><li>5 numerical variability models in Clafer format (.txt) for each software product line.</li><li>5 CSV files with the respective quality attribute measurements</li><li>An .xlsx file containing SAVRUS scalability results divided in different tabs.</li></ul><p>References:</p><p>[1] N. Siegmund, A. Grebhahn, S. Apel, C. Kastner, Performance-influence models for highly configurable systems, in: Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering, ESEC/FSE 2015, Association for Computing Machinery, New York, NY, USA, 2015, p.284–294. doi:10.1145/2786805.2786845.</p><p>[2] M. Bauer, A comparison of six constraint solvers for variability analysis, Tech. rep., University of Passau (2019).</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Datasets for "Large subglacial source of mercury from the southwestern margin of the Greenland Ice Sheet"

<p>Geochemical measurements and hydrochemical datasets linked to the publication &quot;Large subglacial source of mercury from the southwestern margin of the Greenland Ice Sheet&quot; in Nature Geoscience. Presented are (1) data for mercury concentrations in glacial meltwater outflows from the Greenland Ice Sheet taken in 2012, 2015 and 2018, (2) data for mercury concentrations in fjord waters from&nbsp;Nuup Kangerlua,&nbsp;Ameralik Fjord and&nbsp;S&oslash;ndre Str&oslash;mfjord, and (3) all associated hydrochemical data presented in the manuscript.&nbsp;For additional details (analytical techniques, precision, accuracy and limits of detection)&nbsp;please refer to the methodology in the publication.</p> <p>This third version has additional riverine data added the the 2012 dataset.&nbsp;</p>

opencc-by-4.0May 2021View details →
zenodo44/100

MOBDrone: a large-scale drone-view dataset for man overboard detection

<p><strong>Dataset</strong></p> <p>The <em>Man OverBoard Drone (MOBDrone)</em> dataset is a large-scale collection of aerial footage images. It contains 126,170 frames extracted from 66 video clips gathered from one UAV flying at an altitude of 10 to 60 meters above the mean sea level. Images are manually annotated with more than 180K bounding boxes localizing objects belonging to 5 categories --- <em>person, boat, lifebuoy, surfboard, wood</em>. More than 113K of these bounding boxes belong to the person category and localize people in the water simulating the need to be rescued.</p> <p>In this repository, we provide:</p> <ul> <li> <p>66 Full HD video clips (total size: 5.5 GB)&nbsp;</p> </li> <li> <p>126,170 images extracted from the videos at a rate of 30 FPS (total size: 243 GB)</p> </li> <li> <p>3 annotation files for the extracted images that follow the MS COCO data format (for more info see <a href="https://cocodataset.org/#format-data">https://cocodataset.org/#format-data</a>):</p> <ul> <li> <p><em>annotations_5_custom_classes.json</em>: this file contains annotations concerning all five categories; please note that class ids do not correspond with the ones provided by the MS COCO standard since we account for two new classes not previously considered in the MS COCO dataset --- <em>lifebuoy </em>and <em>wood</em></p> </li> <li> <p><em>annotations_3_coco_classes.json:</em> this file contains annotations concerning the three classes also accounted by the MS COCO dataset --- <em>person, boat, surfboard</em>. Class ids correspond with the ones provided by the MS COCO standard.</p> </li> <li> <p><em>annotations_person_coco_classes.json</em>: this file contains annotations concerning only the &#39;<em>person</em>&#39; class. Class id corresponds to the one provided by the MS COCO standard.</p> </li> </ul> </li> </ul> <p>The MOBDrone dataset is intended as a test data benchmark. However, for researchers interested in using our data also for training purposes, we provide training and test splits:</p> <ul> <li><em>Test set: </em>All the images whose filename starts with &quot;DJI_0804&quot; (total: 37,604 images)</li> <li><em>Training set:</em> All the images whose filename starts with &quot;DJI_0915&quot; (total: 88,568 images)</li> </ul> <p>More details about data generation and the evaluation protocol can be found at our MOBDrone paper: <a href="https://arxiv.org/abs/2203.07973">https://arxiv.org/abs/2203.07973</a><br> The code to reproduce our results is available at this GitHub Repository: <a href="https://github.com/ciampluca/MOBDrone_eval">https://github.com/ciampluca/MOBDrone_eval</a><br> See also&nbsp;<strong>&nbsp;</strong><a href="http://aimh.isti.cnr.it/dataset/MOBDrone">http://aimh.isti.cnr.it/dataset/MOBDrone</a></p> <p><strong>Citing the MOBDrone</strong></p> <p>The MOBDrone is released under a Creative Commons Attribution license, so please cite the MOBDrone if it is used in your work in any form.<br> Published academic papers should use the academic paper citation for our MOBDrone paper, where we evaluated several pre-trained state-of-the-art object detectors focusing on the detection of the overboard people</p> <blockquote> <pre>@inproceedings{MOBDrone2021, title={MOBDrone: a Drone Video Dataset for Man OverBoard Rescue}, author={Donato Cafarelli and Luca Ciampi and Lucia Vadicamo and Claudio Gennaro and Andrea Berton and Marco Paterni and Chiara Benvenuti and Mirko Passera and Fabrizio Falchi}, booktitle={ICIAP2021: 21th International Conference on Image Analysis and Processing}, year={2021} } </pre> </blockquote> <p>and this&nbsp;Zenodo Dataset</p> <blockquote> <pre>@dataset{donato_cafarelli_2022_5996890, author={Donato Cafarelli and Luca Ciampi and Lucia Vadicamo and Claudio Gennaro and Andrea Berton and Marco Paterni and Chiara Benvenuti and Mirko Passera and Fabrizio Falchi}, title = {{MOBDrone: a large-scale drone-view dataset for man overboard detection}}, month = feb, year = 2022, publisher = {Zenodo}, version = {1.0.0}, doi = {10.5281/zenodo.5996890}, url = {https://doi.org/10.5281/zenodo.5996890} }</pre> </blockquote> <p>Personal works, such as machine learning projects/blog posts, should provide a URL to the MOBDrone Zenodo page (<a href="https://doi.org/10.5281/zenodo.5996890">https://doi.org/10.5281/zenodo.5996890</a>), though a reference to our MOBDrone paper would also be appreciated.</p> <p>&nbsp;</p> <p><strong>Contact Information</strong></p> <p>If you would like further information about the MOBDrone or if you experience any issues downloading files, please contact us at <a href="mailto:mobdrone@isti.cnr.it">mobdrone[at]isti.cnr.it</a></p> <p>&nbsp;</p> <p><strong>Acknowledgements</strong></p> <p>This work was partially supported by NAUSICAA - &quot;NAUtical Safety by means of Integrated Computer-Assistance Appliances 4.0&quot; project funded by the Tuscany region (CUP D44E20003410009). The data collection was carried out with the collaboration of the Fly&amp;Sense Service of the CNR of Pisa - for the flight operations of remotely piloted aerial systems - and of the Institute of Clinical Physiology (IFC) of the CNR - for the water immersion operations.&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

A Large-scale Dataset of (Open Source) License Text Variants

<p>We introduce a large-scale dataset of the complete texts of free/open source software (FOSS) license variants. To assemble it we have collected from the Software Heritage archive&mdash;the largest publicly available archive of FOSS source code with accompanying development history&mdash;all versions of files whose names are commonly used to convey licensing terms to software users and developers.<br> The dataset consists of 6.5 million unique license files that can be used to conduct empirical studies on open source licensing, training of automated license classifiers, natural language processing (NLP) analyses of legal texts, as well as historical and phylogenetic studies on FOSS licensing.<br> Additional metadata about shipped license files are also provided, making the dataset ready to use in various contexts; they include: file length measures, detected MIME type, detected SPDX license (using ScanCode), example origin (e.g., GitHub repository), oldest public commit in which the license appeared.<br> The dataset is released as open data as an archive file containing all deduplicated license blobs, plus several portable CSV files for metadata, referencing blobs via cryptographic checksums.</p> <p>For more details see the included&nbsp;README file and companion paper:</p> <ul> <li>Stefano Zacchiroli.&nbsp;<a href="https://doi.org/10.1145/3524842.3528491"><em>A Large-scale Dataset of (Open Source) License Text Variants</em></a>. In proceedings of the&nbsp;<a href="https://conf.researchr.org/home/msr-2022">2022 Mining Software Repositories Conference (MSR 2022)</a>. 23-24 May 2022 Pittsburgh, Pennsylvania, United States. ACM 2022.</li> </ul> <p>If you use this dataset for research purposes, please acknowledge its use by citing the above paper.</p> <ul> </ul>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Towards a systematic approach to manual annotation of code smells - C# Dataset of Long Method and Large Class code smells

<p>This dataset includes open-source projects written in C# programing language, annotated for the presence of Long Method and God Class code smells. Each instance was manually annotated by at least two annotators.&nbsp;We explain our motivation and methodology for creating this dataset in our <a href="https://www.techrxiv.org/articles/preprint/Towards_a_systematic_approach_to_manual_annotation_of_code_smells/14159183/1">preprint</a>:</p> <p>Luburić, N., Prokić, S., Grujić, K.G., Slivka, J., Kovačević, A., Sladić, G. and Vidaković, D., 2021. Towards a systematic approach to manual annotation of code smells.&nbsp;</p> <p>The dataset contains two excel datasheets:</p> <ul> <li><em>DataSet_Large Class.xlsx</em> &ndash; C# classes annotated for the Large Class code smell severity.</li> <li><em>DataSet_Long Method.xlsx</em> &ndash; C# methods annotated for the Long method code smell severity.</li> </ul> <p>&nbsp;The columns in the datasheet represent:</p> <ul> <li><em>Code Snippet ID</em> &ndash; the full name of the code snippet.&nbsp; <ul> <li>For classes, this is the package/namespace name followed by the class name. The full name of inner classes also contains the names of any outer classes (e.g., <em>namespace.subnamespace.outerclass.innerclass</em>).</li> <li>For methods, this is the full name of the class and the methods&rsquo;s signature (e.g., <em>namespace.class.method(param1Type, param2Type)</em> ).</li> </ul> </li> <li><em>Link </em>&ndash; The GitHub link to the code snippet, including the commit and the start and end LOC.</li> <li><em>Code Smell </em>&ndash; code smell for which the code snippet is examined (Large Class or Long Method).</li> <li><em>Project Link </em>&ndash; the link to the version of the code repository that was annotated.</li> <li><em>Metrics </em>&ndash; a list of metrics for the code snippet, calculated by our <a href="https://github.com/Clean-CaDET/platform#readme">platform</a>. Our dataset provides 25 class-level metrics for Large Class detection and 18 method-level metrics for Long Method detection The list of metrics and their definitions is available <a href="https://github.com/Clean-CaDET/platform/blob/c4acff95ec00ff6c25fa62dde4818c1f40e39d39/CodeModel/CaDETModel/CodeItems/CaDETMetrics.cs">here</a>.</li> <li><em>Final annotation </em>&ndash; a single severity score calculated by a majority vote.&nbsp;</li> <li><em>Annotators </em>&ndash; each annotator&#39;s (1, 2, or 3) assigned severity score.</li> </ul> <p>To help guide their reasoning for evaluating the presence and the severity of a code smell, three annotators independently annotated whether the considered heuristics apply to an evaluated code snippet. We provide these results in two separate excel datasheets:</p> <ul> <li><em>LargeClass_Heuristics.xlsx </em>- C# classes annotated for the presence of heuristics relevant for the Large Class code smell.</li> <li><em>LongMethod_Heuristics.xlsx </em>- C# classes annotated for the presence of heuristics relevant for the Large Class code smell.</li> </ul> <p>The columns of these two datasheets are:</p> <ul> <li><em>Code Snippet ID </em>- the full name of the code snippet (matching the IDs from <em>DataSet_Large Class.xlsx </em>and <em>DataSet_Long Method.xlsx</em>)</li> <li><em>Annotators</em> &ndash; heuristics labelled by each of the annotators (1, 2, or 3).</li> <li><em>Heuristics </em>&ndash; whether the heuristic is applicable to the examined code snippet or not (Section 1.2.4 lists heuristics relevant for the Large Class detection, and Section 1.2.5 lists the heuristics relevant for the Long Method detection).</li> </ul>

opencc-by-4.0May 2022View details →
zenodo44/100

HISTORIAN: a large-scale HISTORIcal film dataset with cinematographic ANnotation

<p>Developing automated tools for sustainable film preservation of extensive historical film collections assumes an understanding of fundamental cinematographic settings. In order to be able to investigate new approaches to detect and classify cinematographic settings, this paper proposes a novel large-scale historical film dataset with cinematographic annotations (HISTORIAN), i.e., shot boundaries, shot types, camera movements. The dataset consists of 98 digitized original analog film reels related to the Second World War and 10593 film shots manually annotated by human film experts. Moreover, annotations for overscan areas such as sprocket holes are included. A baseline film analysis pipeline is introduced and evaluated. To the best of our knowledge, HISTORIAN is the first dataset that covers the challenges and characteristics of historical film documentaries and provides novel possibilities for exploring automatic film analysis tools.</p> <p>This repository presents a tiny set including a few examples for demonstration.</p> <p>A link to the Github repository (including helper scripts and readme) can be found <a href="https://github.com/dahe-cvl/historian_dataset">here</a></p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Dataset of 30 energy customers with flexibility data, and distributed generation, considering residential, small commerce, large commerce, and industrial customers

<p>The dataset has 30 customers: ten residential, ten small commerce, five large commerce, and five industrial customers. The combination of several energy customer types allows the creation of a dataset with different types of consumption profiles, generation, and flexibility, and, therefore, different values of participation in demand response events.</p> <p>The residential profiles of the considered customers use the data available in the Working Group on Intelligent Data Mining and Analysis (IDMA): https://site.ieee.org/pes-iss/data-sets/</p> <p>The values represent a week period using 15 minutes reading periods. All the values are expressed in kWh and the matrixes were created as [customer x time_period].</p> <p>&nbsp;</p> <p>We would be grateful if you could acknowledge the use of this dataset in your publications. Please use the Zenodo publication to cite this work.</p>

opencc-by-4.0Jun 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record