Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

558

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

558 results for “training data”

Learn how ShareScore rates datasets ↗
zenodo48/100

Silva SSU taxonomic training data formatted for DADA2 (Silva version 138)

<p>These DADA2-formatted training fasta files were derived from the Silva Project&#39;s version 138 release. See https://www.arb-silva.de/documentation/release-138/ for database and citation information. The Silva 138 database is licensed under Creative Commons Attribution 4.0 (CC-BY 4.0); see file &quot;SILVA_LICENSE.txt&quot;. The fasta files were generated and checked for consistency with version 132 using the R code in the R-markdown document &quot;silva-v138.Rmd&quot;.</p> <p>Version 2 removes the dependence on preprocessed files from mothur, which results in a greater number of bacterial and archeal sequences. It also includes&nbsp;a new version of the assignTaxonomy training set&nbsp;that goes through the species level for use with longer amplicons obtained from long-read amplicon sequencing.</p> <p>If you use these files, please cite one or both of the Silva references below (or at the above link) and the DADA2 paper (reference below). I also recommend&nbsp;citing or linking to the Zenodo record for this specific version in your Methods or published source code to record&nbsp;the specific taxonomic database files used in your analysis.</p> <p><strong>NOTE:</strong><strong> </strong>These Version 2 files are intended for use in classifying prokaryotic 16S sequencing data and are not appropriate for classifying eukaryotic ASVs. The new method implemented within DADA2 for constructing these files only includes 100 eukaryotic sequences for use as an outgroup.</p> <p><strong>NOTE:</strong>&nbsp;These Version 2 files have a known problem in 10/883 families and 114/3838 genera. See https://github.com/mikemc/dada2-reference-databases/blob/main/silva-138/v2/bad-taxa.csv for a list of affected taxa and https://github.com/benjjneb/dada2/issues/1293 for more information.</p>

opencc-by-4.0Mar 2020View details →
zenodo48/100

IPBES Data Management Tutorials - Session 1.2: Introduction to IPBES tutorials and training

<p>The&nbsp;<em>IPBES data management tutorials</em>&nbsp;are short videos to help experts implement the IPBES data management Policy. They cover topics ranging from data management policy, reports, active research data, tools, and examples.</p> <p>The&nbsp;<em>Introduction to the IPBES data management policy</em>&nbsp;chapter&nbsp;provides an overview on data management within the IPBES platform, and the series of the tutorials prepared by the task force on knowledge and data that will assist experts with the implementation of the IPBES data management policy.</p> <p>This session,&nbsp;<em>Introduction to the IPBES tutorials and training</em>, provides<strong>&nbsp;</strong>a short introduction detailing the objectives of these tutorials and the responsibilities of the IPBES secretariat regarding data management.&nbsp;</p>

opencc-by-4.0Oct 2020View details →
zenodo48/100

ERA5 based training, validation and evaluation data for retrievals combining 22-58 GHz with 175-340 GHz microwave radiometer measurements during MOSAiC

<p>This data set is used for the training, validation and evaluation of retrievals of temperature and specific humidity profiles, as well as integrated water vapour from simlulated or measured microwave brightness temperatures (TBs), which are described in <strong>[1]</strong>.</p> <p>The data set consists of yearly files (2001-2018, 6-hourly resolution) that include data from the European Centre for Medium-Range Weather Forecasts's ERA5 reanalysis <strong>[2]</strong> and simulated TBs in the microwave spectrum. TB simulations were performed with PAMTRA <strong>[3,4]</strong> on the native ERA5 model level resolution at frequencies of a low frequency Humidity and Temperature Profiler (HATPRO, 22-58 GHz) and of a Low Humidity Profiler (LHUMPRO-243-340, aka MiRAC-P, 175-340 GHz). Afterwards, the ERA5 model level data has been interpolated to a new height grid (dimension 'z'), of which the lowest 43 indices equal the height grid of the retrieval that is developed with this data set. The upper 11 indices are included for additional TB simulations needed for the information content estimation performed and are not used for the retrievals to avoid the tropopause.</p> <p>The trained retrieval is applied to observations from the HATPRO and MiRAC-P that were installed onboard the research vessel Polarstern during the Multidisciplinary drifting Observatory for the Study of Arctic Climate (MOSAiC) expedition.</p> <p><strong>[1]:</strong> Walbr&ouml;l, A., Griesche, H. J., Mech, M., Crewell, S., and Ebell, K.: Combining low- and high-frequency microwave radiometer measurements from the MOSAiC expedition for enhanced water vapour products, Atmospheric Measurement Techniques, 17, 6223-6245, https://doi.org/10.5194/amt-17-6223-2024, 2024.</p> <p><strong>[2]:</strong> Hersbach, H., Bell, B., Berrisford, P., Hirahara, S., Hor&aacute;nyi, A., Mu&ntilde;oz-Sabater, J., Nicolas, J., Peubey, C., Radu, R., Schepers, D., Simmons, A., Soci, C., Abdalla, S., Abellan, X., Balsamo, G., Bechtold, P., Biavati, G., Bidlot, J., Bonavita, M., Chiara, G., Dahlgren, P., Dee, D., Diamantakis, M., Dragani, R., Flemming, J., Forbes, R., Fuentes, M., Geer, A., Haimberger, L., Healy, S., Hogan, R. J., H&oacute;lm, E., Janiskov&aacute;, M., Keeley, S., Laloyaux, P., Lopez, P., Lupu, C., Radnoti, G., Rosnay, P., Rozum, I., Vamborg, F., Villaume, S., and Th&eacute;paut, J.: The ERA5 global reanalysis, Quarterly Journal of the Royal Meteorological Society, 146, 1999&ndash;2049, https://doi.org/10.1002/qj.3803, 2020.</p> <p><strong>[3]:</strong> Mech, M., Maahn, M., Kneifel, S., Ori, D., Orlandi, E., Kollias, P., Schemann, V., and Crewell, S.: PAMTRA 1.0: the Passive and Active Microwave radiative TRAnsfer tool for simulating radiometer and radar measurements of the cloudy atmosphere, Geoscientific Model Development, 13, 4229&ndash;4251, https://doi.org/10.5194/gmd-13-4229-2020, 2020.</p> <p><strong>[4]:</strong> Mech, M., Maahn, M., Ori, D., Kneifel, S., and Orlandi, E.: PAMTRA Package &ndash; Passive and Active Microwave TRANsfer, available at: https://github.com/igmk/pamtra (last access: 6 September 2020), 2019c.</p>

opencc-by-4.0Apr 2024View details →
zenodo48/100

Experimental data for the motor learning study performed: "Towards functional robotic training: Motor learning of dynamic tasks is enhanced by haptic rendering but hampered by robotic assistance"

<p>The dataset contains the kinematic data and the questionnaire responses for a robot-assisted motor learning study performed in the Motor Learning and Neurorehabilitation Laboratory at the University of Bern. The details of the study are&nbsp;described in [doi: ]. The kinematic data for each participant is stored as a data frame inside a &ldquo;pickle&rdquo; (serialized python object) file. The questionnaire responses and population metrics&nbsp;are stored as&nbsp;&ldquo;CSV&rdquo; files. The variables inside the files are explained in &ldquo;DataframeVariableDescription.rtf&rdquo;. For questions, please contact oezhan.oezen@artorg.unibe.ch or L.MarchalCrespo@tudelft.nl.</p>

opencc-by-4.0Jul 2021View details →
zenodo48/100

Multi-stakeholder research data management training as a tool to improve the quality, integrity, reliability and reproducibility of research: Quantitative data of the post-course surveys

<p>Data contains&nbsp;doctoral students&#39; and postdoc researchers&#39; (n=168) self-ratings of their RDM competencies before and after the 3 ECTS credits &quot;Basics of Research Data Management&quot; (BRDM) trainings held 2019-2021 in the University of Turku and &Aring;bo Akademi University, Finland. Moreover, data contains respondents&#39; self-reported further learning needs.</p>

opencc-by-4.0May 2022View details →
zenodo48/100

Training data for bathymetry estimation via EO satellite - Hel Peninsula

<p>This dataset contains satellite image from Sentiel-2A (bands B2, B3, B4, B8) and reference sonar based bathymetry measurements.</p> <p>The reference data was acquired from Polish Maritime Administration (htttp://www.um.gdy.pl) and is publicly available. File reference_data_34.csv contains in-situ measrements aquired at northen shore of Hel Peninsula. Points coordinates are expressed in UTM34N coordinate system.</p> <p>The satellite and the reference datasets were preprocessed by the authors for adjust them for Machine Learning algorithms used in the research.</p> <p>&nbsp;</p>

opencc-by-2.0Jun 2022View details →
zenodo48/100

Training data for 'Upload data to ENA' (Galaxy Training Material)

<p>The data here is a subset of the data published in 10.5281/zenodo.3732359 to be used in GTN &#39;Upload data to ENA&#39; tutorial.</p> <p>Human traces have been removed following <a href="https://training.galaxyproject.org/training-material/topics/sequence-analysis/tutorials/human-reads-removal/tutorial.html">https://training.galaxyproject.org/training-material/topics/sequence-analysis/tutorials/human-reads-removal/tutorial.html</a></p> <p>We produced consensus sequences (*.fasta) for the Illumina PE data following SARS-CoV-2-PE-Illumina-WGS-variant-calling (https://workflowhub.eu/workflows/113?version=4), SARS-CoV-2-variation-reporting (https://workflowhub.eu/workflows/109?version=5) and COVID-19-consensus-construction (https://workflowhub.eu/workflows/138?version=4) workflows.</p>

opencc-by-4.0Aug 2021View details →
zenodo48/100

Training data for: CoastSat image classification

<p><strong>CoastSat image classification training data </strong></p> <p>CoastSat is an open-source global shoreline mapping toolbox, available at https://github.com/kvos/CoastSat, which enables users to extract time-series of shoreline change from 30+ years of publicly available satellite imagery (Landsat 5, 7, 8 and Sentinel-2).</p> <p>The automated shoreline extraction relies on a classifier&nbsp;(Multilayer Perceptron from scikit-learn) which labels each pixels on the images with one of four classes: sand, water, white-water and other land features.</p> <p>The data used to train the classifier is stored here, the README.md file provides information on the data organisation and content of each file.</p>

opencc-by-4.0Jul 2019View details →
zenodo48/100

Training data for neural network-based determination of nematic elastic constants

<p>Neural network training data packets (<strong><em>intensities_{i}.csv, K1K3_{i}.csv</em></strong>), each consisting of 1000 training data pairs, used in a machine learning-based method for determination of&nbsp;Frank elastic constants of nematic liquid crystals, experimental measurements of time-dependent light intensities&nbsp;(<strong><em>experimental_time</em></strong>_<strong><em>{i}.csv, experimental_intensity_{i}.csv</em></strong>), diode spectrum data (<strong><em>diode_lbd</em></strong><strong><em>.csv, diode_w.csv</em></strong>).</p> <p>These data sets are associated with the paper <a href="https://www.nature.com/articles/s41598-023-33134-x"><strong><em>[Zaplotnik et al. SciRep, 2023]</em></strong></a></p> <p>This is supplementary material for a Jupyter Notebook uploaded on&nbsp;<a href="https://zenodo.org/record/7368828">Zenodo</a>.</p>

opencc-by-4.0Nov 2022View details →
zenodo48/100

Merging Bioactivity Predictions from Cell Morphology and Chemical Fingerprint Models Using Similarity to Training Data

<p>The applicability domain of machine learning models trained on structural fingerprints for the prediction of biological endpoints is often limited by the lack of diversity of chemical space of the training data. In this work, we developed &ldquo;similarity-based merger models&rdquo; which combined the output of individual models trained on cell morphology (based on Cell Painting) and chemical structure (based on chemical fingerprints) and the structural and morphological similarities of the test compounds to training compounds. We applied these similarity-based merger models using logistic equations to weigh individual features and predicted assay hit calls of 177 assays from ChEMBL, PubChem and the Broad Institute, where the required Cell Painting annotations were available. We found that the similarity-based merger models outperformed other models with an additional 20% assays (79 out of 177 assays) with an AUC&gt;0.70 compared with 65 out of 177 assays using structural models and 50 out of 177 assays using Cell Painting models. Our results demonstrate that similarity-based merger models combining structure and cell morphology models can more accurately predict a wide range of biological assay outcomes and expand the applicability domain by better extrapolating to new structural and morphology spaces.</p>

opencc-by-4.0Jan 2023View details →
zenodo48/100

Open Soil Spectral Library (training data and calibration models)

<p><strong>Open Soil Spectral Library</strong> contains training MIR (91,631) and VisNIR (65,063) spectral scans + soil calibration data (&gt;60,000 unique locations) and calibration models. Key data set:</p> <ul> <li>ossl_all_L1_v1.2.qs: soil laboratory, site and spectra information;</li> </ul> <p>Important note: The data set spatially over-represents USA and European Union, with little training data in Asia, South America and Australia, hence calibration models reflect primarily soils of USA and Europe.</p> <p>To use the models and data please install <a href="https://hub.docker.com/r/opengeohub/r-geo">R and required packages</a>. Read more about the <strong><a href="https://github.com/traversc/qs">QS data format</a></strong> and how to convert it to CSV or similar. Modeling steps are explained in detail in: <a href="https://github.com/soilspectroscopy/ossl-models">https://github.com/soilspectroscopy/ossl-models</a>. To visualize database please use: <a href="https://explorer.soilspectroscopy.org/">https://explorer.soilspectroscopy.org/</a></p> <p>Complete OSSL documentation can be found at: <a href="https://soilspectroscopy.github.io/ossl-manual/">https://soilspectroscopy.github.io/ossl-manual/</a></p> <p><a href="https://soilspectroscopy.org/"><strong>Soil Spectroscopy for the Global Good</strong></a> is a Coordinated Innovation Network funded by USDA NIFA Food and Agriculture Cyberinformatics Tools Program (<a href="https://nifa.usda.gov/press-release/nifa-invests-over-7-million-big-data-artificial-intelligence-and-other">Award #2020-67021-32467</a>).</p> <p>Input datasets are property of the <a href="https://www.nrcs.usda.gov/wps/portal/nrcs/main/soils/research">USDA NRCS National Soil Survey Center &ndash; Kellogg Soil Survey Laboratory</a>, <a href="https://www.worldagroforestry.org/">ICRAF-World Agroforestry</a>, <a href="https://www.isric.org/">ISRIC-World Soil Information</a>, the <a href="http://africasoils.net/services/data/soil-databases/">Africa Soil Information Service</a> funded by the Bill and Melinda Gates Foundation, the <a href="https://esdac.jrc.ec.europa.eu/">European Soil Data Centre</a>, the <a href="https://www.neonscience.org/">National Ecological Observatory Network</a>, and <a href="https://sae.ethz.ch/">ETH Zurich</a>.&nbsp;</p> <p>For more advanced uses of the soil spectral libraries <strong>we advise to contact the original data producers</strong> especially to get help with using, extending and improving the original SSL data.</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Data for: Adaptive P300-Based Brain-Computer Interface for Attention Training

<p>The dataset contains EEG and behavioral data of 47 participants who completed 9 runs (i.e. copy-spelled 9 words) in a P300 speller task, as well as a random dot motion (RDM) task and questionnaires in a single experimental session. Details of the experimental protocol can be found here:</p> <p>Noble SC,&nbsp;Woods E,&nbsp;Ward T,&nbsp;Ringwood JV. &ldquo;Adaptive P300-Based Brain-Computer Interface for Attention Training: Protocol for a Randomized Controlled Trial.&rdquo; <em>JMIR Res Protoc</em> 2023, 12:e46135, doi:&nbsp;<a href="https://doi.org/10.2196/46135">10.2196/46135</a></p> <p>A journal article describing the results of the study can be found here:<br><br>Noble SC, Woods E, Ward T, Ringwood JV. &ldquo;Accelerating P300-Based Neurofeedback Training for Attention Enhancement Using Iterative Learning Control: A Randomised Controlled Trial.&rdquo; <em>J Neural Eng</em> 2024, 21(2), doi: <a href="https://doi.org/10.1088/1741-2552/ad2c9e" target="_blank" rel="noopener">10.1088/1741-2552/ad2c9e</a></p> <p>Please cite the results paper when using the data.</p> <p>Each participant folder contains:</p> <ul> <li>[xxx]-raw.[xxx] &ndash; unprocessed EEG signals (<strong>in</strong> <strong>mV</strong>) from 32 electrodes for all 9 P300 speller runs in Openvibe (.ov) and Matlab (.mat) file formats, see details of the runs below</li> <li>[xxx]-processed.[xxx] &ndash; contains 3 xDAWN components extracted by the xDAWN spatial filter according to the weights in &ldquo;spatial-filter.cfg&rdquo;</li> <li>classifier.cfg - LDA classifier weights</li> <li>spatial-filter.cfg - xDAWN spatial filter weights</li> <li>log.txt - contains the group assignment, start and end time of the experiment, and performance in the P300 speller and RDM tasks</li> </ul> <p>The&nbsp;file &ldquo;Subject Information.csv&rdquo; contains the age and gender of all participants.</p> <p>The file &ldquo;Questionnaire scores.csv&rdquo; contains the responses to the questionnaire described in the experimental protocol and the NASA Task Load Index (TLX) for all participants.</p> <p>The .ov and .mat files contain data from the following runs:</p> <table> <tbody> <tr> <th>Filename</th> <th>Word to be copy-spelled</th> <th>Number of flashes per row and column</th> <th>Feedback given to participant</th> </tr> </tbody> <tbody> <tr> <td>calibration-signal1</td> <td>THE</td> <td>12</td> <td>no</td> </tr> <tr> <td>calibration-signal2</td> <td>QUICK</td> <td>12</td> <td>no</td> </tr> <tr> <td>calibration-signals</td> <td>Concatenation of calibration-signal1 and calibration-signal2</td> </tr> <tr> <td>eval</td> <td>DOG</td> <td>12</td> <td>yes</td> </tr> <tr> <td>training-run-1</td> <td>BEAUTIFUL</td> <td>10</td> <td>yes</td> </tr> <tr> <td>training-run-2 to training-run-5</td> <td>BEAUTIFUL</td> <td>varying</td> <td>yes</td> </tr> <tr> <td>post-training-run</td> <td>DANCE</td> <td>12</td> <td>yes</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>This research is supported by the Irish Research Council under project ID GOIPG/2020/692 and Science Foundation Ireland under grant number 12/RC/2289_P2.</p>

opencc-by-4.0Jul 2023View details →
zenodo48/100

PANDEM-2 European COVID-19 training data set

<p>The PANDEM-2 COVID-19 European training dataset is a large collection of time series either of real or realistic synthetic (generated) data and indicators associated with the European pandemic response to the COVID-19 pandemic. It is intended to be used for training in pandemic management.</p> <p>&nbsp;</p> <p>This dataset is the result of a data gathering requirement process for pandemic management involving feedback and inputs from several public health and first responder professionals as well as researchers and military personnel directly involved in the European COVID-19 pandemic response. This work is part of the PANDEM-2 project funded by the <em>Horizon 2020 Secure Societies</em> program. To collect this data, an open source software was developed named PANDEM-Source allowing reproducibility and customisation of this dataset.&nbsp;</p> <p>&nbsp;</p> <p>The dataset includes indicators for cases, deaths, hospitalisation, testing and laboratory data including pathogen genomic information, vaccination, non-pharmaceutical interventions, participatory surveillance, social media, flights resources (human and material such as beds or vaccines), and contact tracing activities. When no open available data was found, realistic synthetic data and indicators were generated with the goal of producing a data set to be used for pandemic management&nbsp; training.&nbsp;</p> <p>The project received funding from the European Union&rsquo;s Horizon 2020 Research and Innovation programme under the Grant Agreement No. 883285. The material presented and views expressed here are the responsibility of the author(s) only. The EU Commission takes no responsibility for any use made of the information set out.<br> References</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

Pirate Illustrations for Research Data Management Training

<p>This collection of icons and comics was created to illustrate a workshop on Data Management Plans (<a href="https://doi.org/10.5281/zenodo.5575920">https://doi.org/10.5281/zenodo.5575920</a>). It is provided here to allow further reuse, for example to illustrate presentations.</p> <p>The theme of this collection is revolving around pirates, their accessories, and maritime items in general.</p> <p>Created by Jeanne Wilbrandt.</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Benchmark and training data for replicating financial and insurance examples

<p>This dataset contains training, validation and out-of-sample test data for two&nbsp;European calls and two examples of portfolio of&nbsp;variable annuity guarantees.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

2d U-net models trained to segment human placental maternal/fetal blood volumes and blood vessels from syncrotron micro-CT data along with a sample data volume.

<p>This dataset contains a 512 x 512 x 512 pixel volume taken from an imaging dataset of human placental tissue collected at Diamond Light Source Manchester Imaging Branchline, I13-2 on visits MG23941 and MG22562 using in-line high-resolution synchrotron-sourced phase contrast micro-computed X-ray tomography. This data is saved in HDF5 format with a uint8 datatype. Alongside this are two 2d binary U-net models that have been trained to segment this data. One model segments the data into regions of maternal/fetal blood volume, the other segments the blood vessels. Both models were trained using the fastai python package, which utilises the pytorch library. These models were used to segment the data in our paper &quot;A massively multi-scale approach to characterising tissue architecture by synchrotron micro-CT applied to the human placenta&quot; which can be found at <a href="https://www.biorxiv.org/content/10.1101/2020.12.07.411462v1">https://www.biorxiv.org/content/10.1101/2020.12.07.411462v1</a>. The code used for training the U-net models and for predicting the segmentation of the data volume can be found at <a href="https://github.com/DiamondLightSource/placental-segmentation-2dunet">https://github.com/DiamondLightSource/placental-segmentation-2dunet</a>&nbsp;and is published at&nbsp;<a href="https://doi.org/10.5281/zenodo.4252562">https://doi.org/10.5281/zenodo.4252562</a>&nbsp;</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Training material for de novo transcriptome reconstruction from RNA-seq data

<p>The data provided here are part of a Galaxy tutorial that analyzes RNA-seq data from a study published by Wu et al., 2014 (DOI:10.1101/gr.164830.113). The goal of this study was to investigate "the dynamics of occupancy and the role in gene regulation of the transcription factor Tal1, a critical regulator of hematopoiesis, at multiple stages of hematopoietic differentiation." To this end, RNA-seq libraries were constructed from multiple mouse cell types including G1E - a GATA-null immortalized cell line derived from targeted disruption of GATA-1 in mouse embryonic stem cells - and megakaryocytes. This RNA-seq data was used to determine differential gene expression between G1E and megakaryocytes and later correlated with Tal1 occupancy. This dataset (GEO Accession: GSE51338) consists of biological replicate, paired-end, polyA selected RNA-seq libraries. Because of the long processing time for the large original files, we have downsampled the original raw data files to include only reads that align to chromosome 19 and a subset of interesting genomic loci identified by Wu et al.</p>

opencc-by-4.0Jan 2017View details →
zenodo44/100

Eleven years of training data for south foehn for three regions of Western Austria

<p>This south foehn training data is suited for machine learning purposes.&nbsp;</p> <p>It was created by applying objective foehn classification (OFC, Vergeiner 2004) on hourly data of various stations in Western Austria. Three regions (Vorarlberg, Tiroler Unterland, Tiroler Oberland) and two intensities are available, where</p> <ul> <li>0.0 means no foehn on that day,</li> <li>0.5 means localised foehn on that day (up to half the stations in the region responded to OFC),</li> <li>1.0 means widespread foehn on that day (more than half the stations in the region responded to OFC),</li> </ul> <p>provided for each region individually.</p> <p>A paper, where the process of creation is described, is in preperation and will be linked as soon as it is reviewed.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2024View details →
zenodo44/100

Data-driven physics-based modeling of pedestrian dynamics - dataset: Pedestrian trajectories at Eindhoven train station

<p>Pedestrian trajectories measured at train station Eindhoven Centraal (the Netherlands) on platform 2 with acces to tracks 3 and 4.</p> <p>The dataset is partitioned in files containing 10 consecutive days each, recording 4 data fields:</p> <ul> <li><strong>time_ms:</strong> Passed time since start of the measurements. Unit: milliseconds.</li> <li><strong>object_identifier:</strong> unique id identifying an object.</li> <li><strong>x_position_mm:&nbsp;</strong>coordinates of the object along the x-axis at the given time. Unit: millimeters.</li> <li><strong>y_position_mm:</strong> coordinates of the object along the y-axis at the given time. Unit: millimeters.</li> </ul> <p>Each object resembles a pedestrian on the train platform recorded with 10 frames per second. We deliberately removed exact date and time information for privacy reasons (see additional note). The data set consists of 60 consecutive days starting at an unkown time between 00:00 AM and 01:00 AM of a random date between April 1st and May 1st 2022. An overhead image of the platform is included showing train track 3 in the bottom and train track 4 in the top of the image.</p> <p>The data set is supplemented to the paper <a title="Data-driven physics-based modeling of pedestrian dynamics" href="https://doi.org/10.48550/arXiv.2407.20794" target="_blank" rel="noopener">Data-driven physics-based modeling of pedestrian dynamics</a> and can be processed by the associated <a title="Software: Data-driven physics-based modeling of pedestrian dynamics" href="https://github.com/c-pouw/physics-based-pedestrian-modeling" target="_blank" rel="noopener">Python implementation</a> to create pedestrian models.&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Sentence/Table Pair Data from Wikipedia for Pre-training with Distant-Supervision

<p>This is the dataset used for pre-training in &quot;<em>ReasonBERT: Pre-trained to Reason with Distant Supervision</em>&quot;, EMNLP&#39;21.</p> <p>There are two files:</p> <p>sentence_pairs_for_pretrain_no_tokenization.tar.gz -&gt; Contain only sentences as evidence, Text-only</p> <p>table_pairs_for_pretrain_no_tokenization.tar.gz -&gt; At least one piece of evidence is a table, Hybrid</p> <p>The data is chunked into multiple tar files for easy loading. We use <a href="https://github.com/webdataset/webdataset">WebDataset</a>, a PyTorch Dataset (IterableDataset) implementation providing efficient&nbsp;sequential/streaming data access.</p> <p>For pre-training code, or if you have any questions, please check our GitHub repo&nbsp;https://github.com/sunlab-osu/ReasonBERT</p> <p>Below is a sample code snippet to load the data</p> <pre><code class="language-python">import webdataset as wds # path to the uncompressed files, should be a directory with a set of tar files url = './sentence_multi_pairs_for_pretrain_no_tokenization/{000000...000763}.tar' dataset = ( wds.Dataset(url) .shuffle(1000) # cache 1000 samples and shuffle .decode() .to_tuple("json") .batched(20) # group every 20 examples into a batch ) # Please see the documentation for WebDataset for more details about how to use it as dataloader for Pytorch # You can also iterate through all examples and dump them with your preferred data format</code></pre> <p>Below we show how the data is organized with two examples.</p> <p>Text-only</p> <pre><code>{'s1_text': 'Sils is a municipality in the comarca of Selva, in Catalonia, Spain.', # query sentence 's1_all_links': { 'Sils,_Girona': [[0, 4]], 'municipality': [[10, 22]], 'Comarques_of_Catalonia': [[30, 37]], 'Selva': [[41, 46]], 'Catalonia': [[51, 60]] }, # list of entities and their mentions in the sentence (start, end location) 'pairs': [ # other sentences that share common entity pair with the query, group by shared entity pairs { 'pair': ['Comarques_of_Catalonia', 'Selva'], # the common entity pair 's1_pair_locs': [[[30, 37]], [[41, 46]]], # mention of the entity pair in the query 's2s': [ # list of other sentences that contain the common entity pair, or evidence { 'md5': '2777e32bddd6ec414f0bc7a0b7fea331', 'text': 'Selva is a coastal comarque (county) in Catalonia, Spain, located between the mountain range known as the Serralada Transversal or Puigsacalm and the Costa Brava (part of the Mediterranean coast). Unusually, it is divided between the provinces of Girona and Barcelona, with Fogars de la Selva being part of Barcelona province and all other municipalities falling inside Girona province. Also unusually, its capital, Santa Coloma de Farners, is no longer among its larger municipalities, with the coastal towns of Blanes and Lloret de Mar having far surpassed it in size.', 's_loc': [0, 27], # in addition to the sentence containing the common entity pair, we also keep its surrounding context. 's_loc' is the start/end location of the actual evidence sentence 'pair_locs': [ # mentions of the entity pair in the evidence [[19, 27]], # mentions of entity 1 [[0, 5], [288, 293]] # mentions of entity 2 ], 'all_links': { 'Selva': [[0, 5], [288, 293]], 'Comarques_of_Catalonia': [[19, 27]], 'Catalonia': [[40, 49]] } } ,...] # there are multiple evidence sentences }, ,...] # there are multiple entity pairs in the query }</code></pre> <p>Hybrid</p> <pre><code>{'s1_text': 'The 2006 Major League Baseball All-Star Game was the 77th playing of the midseason exhibition baseball game between the all-stars of the American League (AL) and National League (NL), the two leagues comprising Major League Baseball.', 's1_all_links': {...}, # same as text-only 'sentence_pairs': [{'pair': ..., 's1_pair_locs': ..., 's2s': [...]}], # same as text-only 'table_pairs': [ 'tid': 'Major_League_Baseball-1', 'text':[ ['World Series Records', 'World Series Records', ...], ['Team', 'Number of Series won', ...], ['St. Louis Cardinals (NL)', '11', ...], ...] # table content, list of rows 'index':[ [[0, 0], [0, 1], ...], [[1, 0], [1, 1], ...], ...] # index of each cell [row_id, col_id]. we keep only a table snippet, but the index here is from the original table. 'value_ranks':[ [0, 0, ...], [0, 0, ...], [0, 10, ...], ...] # if the cell contain numeric value/date, this is its rank ordered from small to large, follow TAPAS 'value_inv_ranks': [], # inverse rank 'all_links':{ 'St._Louis_Cardinals': { '2': [ [[2, 0], [0, 19]], # [[row_id, col_id], [start, end]] ] # list of mentions in the second row, the key is row_id }, 'CARDINAL:11': {'2': [[[2, 1], [0, 2]]], '8': [[[8, 3], [0, 2]]]}, } 'name': '', # table name, if exists 'pairs': { 'pair': ['American_League', 'National_League'], 's1_pair_locs': [[[137, 152]], [[162, 177]]], # mention in the query 'table_pair_locs': { '17': [ # mention of entity pair in row 17 [ [[17, 0], [3, 18]], [[17, 1], [3, 18]], [[17, 2], [3, 18]], [[17, 3], [3, 18]] ], # mention of the first entity [ [[17, 0], [21, 36]], [[17, 1], [21, 36]], ] # mention of the second entity ] } } ] }</code></pre> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record