Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,037
datasets available to search
ShareScore release 0.9.0
Dataset results
1,037 results for “large-scale”
The Reddit Politosphere: A Large-Scale Text and Network Resource of Online Political Discourse
<p>The Reddit Politosphere is a large-scale resource of online political discourse covering more than 600 political discussion groups over a period of 12 years. Based on the <a href="https://doi.org/10.5281/zenodo.3608135">Pushshift Reddit Dataset</a>, it is to the best of our knowledge the largest and ideologically most comprehensive dataset of its type now available. One key feature of the Reddit Politosphere is that it consists of both text and network data. We also release annotated metadata for subreddits and users.</p> <p>Documentation and scripts for easy data access are provided in an associated <a href="https://github.com/valentinhofmann/politosphere">repository</a> on GitHub.</p>
Large-scale neural recordings with single neuron resolution using Neuropixels probes in human cortex
<p><span>Recent advances in multi-electrode array technology have made it possible to monitor large neuronal ensembles at cellular resolution in animal models. In humans, however, c</span>urrent approaches restrict recordings to few neurons per penetrating electrode or combine the signals of thousands of neurons in local field potential (LFP) recordings. Here, we describe a new probe variant and set of techniques which enable simultaneous recording from over 200 well-isolated cortical single units in human participants during intraoperative neurosurgical procedures using silicon Neuropixels probes. We characterized a diversity of extracellular waveforms with eight separable single unit classes, with differing firing rates, locations along the length of the electrode array, waveform spatial spread, and modulation by LFP events such as inter-ictal discharges and burst suppression. While some challenges remain in creating a turn-key recording system, high-density silicon arrays provide a path for studying human-specific cognitive processes and their dysfunction at unprecedented spatiotemporal resolution. </p>
The Piraeus AIS Dataset for Large-scale Maritime Data Analytics
<p><strong>AIS data collected by the University of Piraeus' AIS receiver</strong></p> <p> </p> <p><strong>Abstract</strong></p> <p>The advent of Big Data and streaming technologies has resulted in a swarm of voluminous, heterogeneous information, especially in the domains of Internet of Things (IoT) and transportation. Focusing on the maritime field, we present a dataset that contains vessel position information transmitted by vessels of different types and collected via the Automatic Identification System (AIS). The AIS dataset comes along with spatially and temporally correlated data about the vessels and the area of interest, including weather information. It covers a time span of over 2.5 years, from May 9<sup>th</sup>, 2017 to December 26<sup>th</sup>, 2019 and provides anonymised vessel positions within the wider area of the port of Piraeus (Greece), one of the busiest ports in Europe and worldwide. The dataset consists of over 244 million AIS records, an average of more than 10,000 records per hour, which makes it an ideal input for large-scale mobility data processing and analytics purposes.</p> <p> </p> <p><strong>Dataset related to the following publication</strong></p> <blockquote> <p>Andreas Tritsarolis, Yannis Kontoulis, Yannis Theodoridis, The Piraeus AIS dataset for large-scale maritime data analytics, Data in Brief, Volume 40, 2022, 107782, ISSN 2352-3409, <a href="https://doi.org/10.1016/j.dib.2021.107782">https://doi.org/10.1016/j.dib.2021.107782</a>.</p> </blockquote> <p> </p> <p><strong>Files Description</strong></p> <ul> </ul> <ul> <li><strong>ais_static</strong>: CSV flat files containing vessels' static information and their corresponding types</li> </ul> <ul> <li><strong>geodata</strong>: ESRI Shapefiles containing several geographic-related data (e.g. harbours, islands, etc.)</li> </ul> <ul> <li><strong>noaa_weather</strong>: ESRI Shapefiles containing weather forecast from GRIB files (as provided by NOAA)</li> </ul> <ul> <li><strong>unipi_ais_dynamic</strong>: CSV flat files containing AIS kinematic information </li> </ul> <ul> <li><strong>unipi_ais_dynamic_synopses</strong>: CSV flat files containing metadata (i.e. synopses) regarding vessels' AIS positions</li> </ul> <p> </p> <p><strong>Privacy Statement</strong></p> <p><strong>For privacy-related queries, please contact the authors</strong></p>
Supplementary GIS data - Potential and implications of automated pre-processing of LiDAR-based digital elevation models for large-scale archaeological landscape analysis
<p>A supplementary dataset related to the paper discussing preparation of a digital elevation model derived from DMR 5G (LiDAR-based DEM of the Czech Republic) cleaned of modern artificial features. It includes data used as a clipping mask and data produced during the testing phase.</p> <p>Contents:</p> <ul> <li>..\clipping_buffers.gdb\ - Clipping buffers based on ZABAGED dataset used for masking the original data stored as ESRI geodatabase.</li> <li>..\drainages\ - Drainages with Strahler order higher than four (potential watercourses) for the original and filtered DEMs. <ul> <li>drainages_filtered - Drainges identified in the filtered DEM stored as GeoTIFF.</li> <li>drainages_original - Drainges identified in the original DEM stored as GeoTIFF. </li> </ul> </li> <li>..\LSC\ - Locations with significant land surface curvature for the original and filtered DEMs. <ul> <li>LSC_filtered - Significant LSC identified in the filtered DEM stored as GeoTIFF. </li> <li>LSC_original - Significant LSC identified in the original DEM stored as GeoTIFF. </li> </ul> </li> <li>..\visibility\ - Viewsheds computed over the original and filtered DEMs. <ul> <li>Libice\ - Sample viewsheds computed for the early medieval hillfort of Libice. <ul> <li>Libice_visibility_filtered - Viewshed based on the filtered DEM stored as GeoTIFF. </li> <li>Libice_visibility_original - Viewshed based on the original DEM stored as GeoTIFF. </li> <li>observer_points - Observer points used for calculating the viewsheds.</li> </ul> </li> <li>regular_grid\ - Cumulative viewsheds calculated for regularly spaced points in a 10 x 10 km grid with a visibility radius of 5 km and an observer height of 2 m; a total of 574 viewsheds. <ul> <li>visibility_filtered - Cumulative viewshed for the filtered DEM stored as GeoTIFF.</li> <li>visibility_original - Cumulative viewshed for the original DEM stored as GeoTIFF. </li> <li>visibility_test_buffers - Buffers used for the viewshed calculations stored as ESRI shapefile.</li> <li>visibility_test_observers - Observer points used for the viewshed calculations stored as ESRI shapefile.</li> </ul> </li> </ul> </li> </ul> <p> </p> <p>Preprint version of the related paper:</p> <p>Novák, David and Pružinec, Filip, Potential and Implications of Automated Pre-Processing of Lidar-Based Digital Elevation Models for Large-Scale Archaeological Landscape Analysis. Available at SSRN: <a href="https://ssrn.com/abstract=4063514">https://ssrn.com/abstract=4063514</a></p>
Source data to publication "Benchmarking of Analysis Strategies for Data-Independent Acquisition Proteomics Using a Large-Scale Dataset Comprising Inter-Patient Heterogeneity"
<p>Source data to publication "Benchmarking of Analysis Strategies for Data-Independent Acquisition Proteomics Using a Large-Scale Dataset Comprising Inter-Patient Heterogeneity".</p> <p>Data and further information at GitHub repository https://github.com/kreutz-lab/dia-benchmarking (DOI: 10.5281/zenodo.6371925)</p>
Computational Analysis of Two-dimensional High-throughput Data from Large-scale RNAi Screens and Single-cell Transcriptomics
<p>This publication provides a singularity definition file to reproduce the computational environment along with the scripts to reproduce every figure or table in the revised manuscript using ZetaSuite Perl module and R package.</p> <p>First, generate a new folder and then download all the files into the folder.</p> <p>Then, uncompressed the files DataSets_part1.tar.gz,DataSets_part2.tar.gz,DataSets_part3.tar.gz,DataSets_part4.tar.gz, and scripts.tar.gz. within the folder.</p> <p>Next, move all the files in DataSets_part1 folder, DataSets_part2 folder,DataSets_part3 folder and DataSets_part4 folder to a new folder called DataSets.</p> <p>Finally, run the following scripts to generate the figures and tables in our manuscript.</p> <p>Regeneration of Figure2 and S2: singularity exec ZetaSuite.sif sh Figure2andS2.sh </p> <p>Regeneration of Figure3 and S3: singularity exec ZetaSuite.sif sh Figure3andS3.sh </p> <p>Regeneration of Figure4 and S4: singularity exec ZetaSuite.sif sh Figure4andS4.sh </p> <p>Regeneration of Figure5 and S5: singularity exec ZetaSuite.sif sh Figure5andS5.sh </p> <p>Regeneration of Figure6 and S6: singularity exec ZetaSuite.sif sh Figure6andS6.sh </p> <p>Regeneration of Figure7 and S7: singularity exec ZetaSuite.sif sh Figure7andS7.sh </p> <p> </p>
Supplementary Material of : Large-Scale 3D Image Segmentation Using Scattering Networks
<p>The reader will find here the supplementary material associated with the manuscript "Large-Scale 3D Image Segmentation Using<br> Scattering Networks" submitted to IEEE Transaction of Pattern Analysis and Machine Intelligence (TPAMI), 2022.</p>
Community access to rectal artesunate for malaria (CARAMAL): a large-scale observational implementation study in the Democratic Republic of the Congo, Nigeria and Uganda
<p>Datasets underlying the publication "Community access to rectal artesunate for malaria (CARAMAL): a large-scale observational implementation study in the Democratic Republic of the Congo, Nigeria and Uganda":</p> <p><strong>Figure 6: </strong>Number of children enrolled in the Patient Surveillance System (grey bars), and percentage of these children being administered rectal artesunate (RAS), by country.</p> <p><strong>Figure 8:</strong> Overall case fatality ratio (CFR) in patients with danger signs and a positive malaria test at enrolment across the entire study period, by enrolment location and country. Data for Uganda excludes enrolments at PHCs (N=34).</p>
EmoWOZ: A Large-Scale Corpus and Labelling Scheme for Emotion Recognition in Task-Oriented Dialogue Systems
<p>This is the dataset created for the paper, "EmoWOZ: A Large-Scale Corpus and Labelling Scheme for Emotion Recognition in Task-Oriented Dialogue Systems" (https://arxiv.org/abs/2109.04919).</p> <p>EmoWOZ is based on MultiWOZ, a multi-domain task-oriented dialogue dataset (https://github.com/budzianowski/multiwoz). It contains more than 11K task-oriented dialogues with more than 83K emotion annotations of user utterances. In addition to Wizard-of-Oz dialogues from MultiWOZ, we collect human-machine dialogues within the same set of domains to sufficiently cover the space of various emotions that can happen during the lifetime of a data-driven dialogue system. There are 7 emotion labels, which are adapted from the OCC emotion models.</p> <p>For data format and label definition, please refer to README.md. </p>
Data supporting "Large-scale citizen science programs can support ecological and climate change assessments"
<p>Text file of phenology observations pulled from the USA National Phenology Network's database (www.usanpn.org) and used in this analysis. </p>
Dataset of Jupyter Notebooks from the paper "A Large-Scale Comparison of Python Code in Jupyter Notebooks and Scripts"
<pre>This archive contains the dataset of properly-licensed Jupyter notebooks from the MSR'22 paper "A Large-Scale Comparison of Python Code in Jupyter Notebooks and Scripts". The dataset contains 847,881 notebooks stored in the PostgreSQL dump file. You can find the details about the database in the README file. To transform the notebooks into this convenient format and to calcuate the structural metrics, we used our library called Matroskin, which can be found here: <a href="https://github.com/JetBrains-Research/Matroskin">https://github.com/JetBrains-Research/Matroskin</a>. </pre>
EvaNIL: silver standard dataset for large-scale NIL entity linking evaluation
<p>The EvaNIL dataset can be used to train or evaluate approaches developed for NIL entity linking. It was built from several Biomedical and Life Sciences corpora:</p> <ul> <li>PubMed DS</li> <li>CRAFT corpus</li> <li>MedMentions</li> </ul> <p>These corpora contain entities associated with knowledge base concepts. To build the EvaNIL dataset, we assumed that those knowledge base concepts did not exist in the respective knowledge bases, so each entity is associated instead with the direct ancestors of those original concepts.</p> <p>The EvaNIL dataset is divided into 6 partitions including annotations from several knowledge bases:</p> <ul> <li>"medic" (CTD-MEDIC)</li> <li>"ctd_anatomy" (CTD-Anatomy)</li> <li>"ctd_chemicals" (CTD-Chemicals)</li> <li>"chebi" (ChEBI)</li> <li>"go_bp" (GO-Biological Process)</li> <li>"hp" (HPO)</li> </ul> <p> </p>
GTAV-NightRain: Photometric Realistic Large-scale Dataset for Night-time Rain Streak Removal
<p>Existing synthetic deraining datasets ignored the photometry property of rain streaks and superimposed 2D rain layer onto clean images without any interactions between rain and environment, making rain streaks unrealistic. The appearance of rain can change drastically with space due to environmental illumination, especially in night scenes where lights are not parallel and uniform. Considering photometry and three dimensional space, rain streaks rendered in GTA V look much more realistic. With proper modifications, it can turn into a good platform for collecting paired data on deraining.</p> <p>GTAV-NightRain dataset is a large-scale photometric realistic dataset for night-time rain streak removal. This dataset contains 12860 rainy images together with 1286 rain-free ground truths collected in GTA V. Two rain shape and different rain density are included.</p> <p>We also provide a small version of the dataset on <a href="https://drive.google.com/drive/folders/1Tsoh_9iCfYc2rtMCHnJACnKW5W_SOheb?usp=sharing">Google Drive</a>. The dataset will be continuously maintained and information of updates will be in informed at <a href="https://github.com/zkawfanx/GTAV-NightRain">Github</a>. Please check it for more information and stay tuned if you are interested in our dataset.</p>
JetClass: A Large-Scale Dataset for Deep Learning in Jet Physics
<p>JetClass is a new large-scale dataset to facilitate deep learning research in jet physics. It consists of 100M jets for training, 5M for validation and 20M for testing. The dataset contains 10 classes of jets, simulated with MadGraph + Pythia + Delphes. <br> <br> A detailed description of the JetClass dataset is presented in the paper <a href="https://arxiv.org/abs/2202.03772">Particle Transformer for Jet Tagging</a>. An interface to use the dataset is provided in <a href="https://github.com/jet-universe/particle_transformer">https://github.com/jet-universe/particle_transformer</a>.</p>
Modelling assumptions and input dataset for the case study of the paper "Societal Effects of Large-Scale Energy Storage in the Current and Future Day-Ahead Market: A Belgian Case Study"
<p>This data package includes the modelling assumptions and input data to replicate the results of the case study included in the paper "Societal Effects of Large-Scale Energy Storage in the Current and Future Day-Ahead Market: A Belgian Case Study". This paper is part of the 18th International Conference on the European Energy Market (EEM22).</p> <p>The case study models the Belgian day-ahead electricity market, in which the existing storage is considered, in addition to large-scale battery energy storage systems of different sizes for varying renewable energy shares. A detailed description of the case study is provided in the readme file. </p> <p>This supplementary data package includes the following files: </p> <p> --Belgium Model Input Data.xlsx: Dataset used as input in the case study of the mentioned paper<br> --Modelling Assumptions.pdf: Modelling assumptions considered in the case study<br> --readme.txt (this file): Includes a detailed description of the data package</p> <p> </p> <p>The data included in this dataset was collected from public open sources [1]-[2]. Please notice that this dataset does not replace the original open access information. For accessing the data, please visit the following websites:</p> <p>[1] “ENTSO-E Transparency Platform.” [Online]. Available: https://transparency.entsoe.eu/dashboard/show. [Accessed: 06-Jul-2022].<br> [2] “Grid data.” [Online]. Available: https://www.elia.be/en/grid-data. [Accessed: 06-Jul-2022].</p> <p><br> </p> <p> </p>
Data accompanying "A new brittle rheology and numerical framework for large-scale sea-ice models"
<p>Data accompanying "A new brittle rheology and numerical framework for<br> large-scale sea-ice models" by E. Olason et al, accepted for publication in<br> Journal of Advances in Modelling Earth Systems (2022).</p> <p>Files:<br> * CS2SMOS.tar.bz2: Contains Cryosat2/SMOS data, post-precessed and used to<br> produce figures comparing modelled thickness to observations.<br> * deformation_maps_demo.ipynb: An example jupyter notebook to read pairs.npz<br> * OlasonEtAl_BBM.tar.bz2: Thickness fields from the MEB run used to produce<br> figure 1 (netCDF).<br> * OlasonEtAl_MEB.tar.bz2: Thickness fields from the BBM run used to produce<br> figure 8 (netCDF).<br> * OlasonEtAl_mEVP.tar.bz2: Thickness fields from the mEVP run used to produce<br> figure 8 (netCDF).<br> * pairs.npz: Displacement pairs derived from the model's Lagrangian mesh used<br> to produce figures 3, 4, and 5 (numpy data file).<br> * Winter2006_7_BBM.nc.bz2: Thickness, concentration, and velocity fields from<br> the BBM run for the winter 2006-7 widely used in the paper (netCDF).<br> * Winter2006_7_mEVP.nc.bz2: Thickness, concentration, and velocity fields from<br> the mEVP run for the winter 2006-7 used for comparison in the paper<br> (netCDF).</p> <p> </p>
The use of GRDC gauging stations for calibrating large-scale hydrological models
<p>The Global Runoff Data Centre provides time series of observed discharges that are very valuable for calibrating and validating the results of hydrological models. We address a common issue in large-scale hydrology which, though investigated several times, has not been satisfactorily solved. Grid-based hydrological models need to fit the reported station location to the river network depending on the resolution, to compare simulated discharge with observed discharge. We introduce an Intersection over Union ratio approach to selected station locations on a coarser grid scale, reducing the errors in assigning stations to the wrong basin. We update the 10-year-old database of watershed boundaries with additional stations based on a high-resolution (3 arcseconds) river network, and we provide source codes and high- and low-resolution watershed boundaries.</p> <p>Same as release on Github: https://github.com/iiasa/CWATM_grdc_calibration_stations/releases/tag/V1.0</p>
A Large-scale Synthetic Pathological Dataset for Deep Learning-enabled Segmentation of Breast Cancer
<p>Dataset access for the paper: A Large-scale Synthetic Pathological Dataset for Deep Learning-enabled Segmentation of Breast Cancer</p>
Data for manuscript "rMATS-turbo: an efficient and flexible computational tool for alternative splicing analysis of large-scale RNA-seq data"
<p>Output files generated by rMATS-turbo for the two example datasets described in the manuscript titled "rMATS-turbo: an efficient and flexible computational tool for alternative splicing analysis of large-scale RNA-seq data".</p> <table> <tbody> <tr> <td>File</td> <td>Description</td> <td>Cell lines</td> <td>BioProject</td> </tr> <tr> <td>PC3E-GS689.tar.gz</td> <td>Compressed folder containing all 36 rMATS-turbo output files for Example 1 described in the manuscript</td> <td>PC3E and GS689 cell lines</td> <td>PRJNA438990</td> </tr> <tr> <td>CCLE.tar.gz</td> <td>Compressed folder containing all 36 rMATS-turbo output files for Example 2 described in the manuscript</td> <td>1,019 CCLE human cancer cell lines</td> <td>PRJNA523380</td> </tr> </tbody> </table> <p>A detailed description of the output files is available in the manuscript and the rMATS-turbo software GitHub repository (https://github.com/Xinglab/rmats-turbo).</p>
The Virtual Macaque Brain: A macaque connectome for large-scale network simulations in TheVirtualBrain
<p>A whole-cortex macaque structural connectome constructed from a combination of axonal tract-tracing and diffusion-weighted imaging data. Created for modeling brain dynamics using TheVirtualBrain platform. Website: thevirtualbrain.org</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.