Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11,687
datasets available to search
ShareScore release 0.7.1
Dataset results
11,687 results for “training”
Fluorescently-labelled zebrafish pronephroi + ground truth classes (normal/cystic) + trained CNN model
<p>This upload contains :</p> <p>- <strong>images.zip: </strong> microscope images of fluorescently-labelled pronephroi in larvae of the <em>Tg(wt1b:EGFP)</em> transgenic zebrafish line showing 2 morphologies (normal vs cystic) upon injection with Co-Mo or ift172-MO, respectively. Images were obtained using an ACQUIFER Imaging Machine widefield high content screening microscope.</p> <p>Reference: </p> <p>Pandey, G., Westhoff, J., Schaefer, F. and Gehrig, J. (2019). <strong>A Smart Imaging Workflow for Organ-Specific Screening in a Cystic Kidney Zebrafish Disease Model</strong>. International Journal of Molecular Sciences <em>20</em>, 1290, doi:<a href="https://doi.org/10.3390/ijms20061290">10.3390/ijms20061290</a>.</p> <p> </p> <p>- <strong>Annotations-***.csv : </strong>Tables containing ground-truth category classes (normal vs cystic) for the images in the zip file.</p> <p>The tables contain columns with the image filename, folder and category.</p> <p>Note : <strong>the Folder column should be updated with the root folder directory once downloaded on your machine.</strong></p> <p>These files were generated with the Fiji plugin <em>single-class (button)</em> from the <em>Qualitative-Annotations</em> update site.</p> <p>The 2 files contain the same information, they only differ in the formatting of the category, the <em>singleColumn </em>file has a single category column while the <em>multiColumn</em> has 2 columns (normal/cystic) with 0/1 encoding.</p> <p>The choice of category encoding solely depends on how the table is used, i.e. in which training workflow, home-made script or software.</p> <p>- <strong>trainedModel.zip : </strong>This archive contains 2 files: <strong>(1) </strong>a h5 file corresponding to a trained deep-learning model to classify the images of the dataset in the 2 categories (normal vs cystic), and <strong>(2)</strong> a text file containing the class names. Both files are necessary to predict the category of new images similar to the one in the dataset, for instance using the published KNIME workflows.</p>
2d U-net models trained to segment human placental maternal/fetal blood volumes and blood vessels from syncrotron micro-CT data along with a sample data volume.
<p>This dataset contains a 512 x 512 x 512 pixel volume taken from an imaging dataset of human placental tissue collected at Diamond Light Source Manchester Imaging Branchline, I13-2 on visits MG23941 and MG22562 using in-line high-resolution synchrotron-sourced phase contrast micro-computed X-ray tomography. This data is saved in HDF5 format with a uint8 datatype. Alongside this are two 2d binary U-net models that have been trained to segment this data. One model segments the data into regions of maternal/fetal blood volume, the other segments the blood vessels. Both models were trained using the fastai python package, which utilises the pytorch library. These models were used to segment the data in our paper "A massively multi-scale approach to characterising tissue architecture by synchrotron micro-CT applied to the human placenta" which can be found at <a href="https://www.biorxiv.org/content/10.1101/2020.12.07.411462v1">https://www.biorxiv.org/content/10.1101/2020.12.07.411462v1</a>. The code used for training the U-net models and for predicting the segmentation of the data volume can be found at <a href="https://github.com/DiamondLightSource/placental-segmentation-2dunet">https://github.com/DiamondLightSource/placental-segmentation-2dunet</a> and is published at <a href="https://doi.org/10.5281/zenodo.4252562">https://doi.org/10.5281/zenodo.4252562</a> </p>
Trained convolutional neural network for the identification of long-duration mixed precipitation in Montréal (Canada)
<p>In this dataset the trained convolutional neural network is published that accompanies the paper "A deep learning approach for the identification of long-duration mixed precipitation in Montréal (Canada)" submitted to the special issue on "Machine-Learning Applications in the Atmospheric and Oceanic Sciences" by the journal Atmosphere&Ocean.</p> <p>The files were created using tensorflow in python. The trained network is available in .h5-format the history as numpy-file (npy).</p>
Dataset: Reinforcing Cybersecurity Hands-on Training With Adaptive Learning
<p>This repository contains supplementary materials for the following conference paper:<br> <br> Pavel Seda, Jan Vykopal, Valdemar Švábenský, Pavel Čeleda.<em><br> Reinforcing Cybersecurity Hands-on Training With Adaptive Learning. </em><br> In Proceedings of the 51st IEEE Frontiers in Education Conference (FIE 2021).<br> <a href="https://doi.org/10.1109/FIE49875.2021.9637252">https://doi.org/10.1109/FIE49875.2021.9637252</a><br> <br> Preprint available at: <a href="https://arxiv.org/abs/2201.01574">https://arxiv.org/abs/2201.01574</a></p> <p><strong>How to cite</strong></p> <p>If you use or build upon the materials, please use the BibTeX entry below to cite the original paper (not only this web link).</p> <p>Some of the linked repositories have their separate citation entry; please use that one as well, if possible.</p> <pre><code>@inproceedings{Seda2021reinforcing, author = {Seda, Pavel and Vykopal, Jan and \v{S}v\'{a}bensk\'{y}, Valdemar and \v{C}eleda, Pavel}, title = {{Reinforcing Cybersecurity Hands-on Training With Adaptive Learning}}, booktitle = {Proceedings of the 51st IEEE Frontiers in Education Conference}, series = {FIE '21}, location = {Lincoln, NE, USA}, publisher = {IEEE}, address = {New York, NY, USA}, month = {10}, year = {2021}, pages = {1--9}, numpages = {9}, isbn = {978-1-6654-3851-3}, url = {https://doi.org/10.1109/FIE49875.2021.9637252}, doi = {10.1109/FIE49875.2021.9637252}, }</code></pre> <p> </p>
Cu dataset – A copper ore labeled images dataset for segmentation training and testing
<p>This dataset is composed of 121 pairs of correlated images. Each pair contains one image of a copper ore sample acquired through reflected light microscopy (RGB, 24-bit), and the corresponding binary reference image (8-bit), in which the pixels are labeled as belonging to one of two classes: ore (0) or embedding resin (255).</p> <p>The sample came from a copper ore from Yauri Cusco (Peru) with a complex mineralogy, mainly composed of sulfides, oxides, silicates, and native copper. It was classified by size. The fraction +74-100 μm was cold mounted with epoxy resin and subsequently ground and polished.</p> <p>Correlative microscopy was employed for image acquisition. Thus, 121 fields were imaged on a reflected light microscope with a 20× (NA 0.40) objective lens and on a scanning electron microscope (SEM). In sequence, they were registered, resulting in images of 1017×753 pixels with a resolution of 0.53 µm/pixel. As matter of fact, some images (the images No. 2, 3, 24, 25, 46, 47, 69, 91, and 113) have slightly smaller sizes because they were cropped during the registration procedure to correct co-localization errors of the order of a few pixels. Finally, the images from SEM were thresholded to generate the reference images.</p> <p>Further description of this sample and its imaging procedure can be found in the work by Gomes and Paciornik (2012).</p> <p>This dataset was created for developing and testing deep learning models on semantic segmentation tasks. The paper of Filippo et al. (2021) presented a variant of the DeepLabv3+ model (Chen et al., 2018) that reached mean values of 90.56% and 92.12% for overall accuracy and F1 score, respectively, for 5 rounds of experiments (training and testing), each with a different, random initialization of network weights.</p> <p>For further questions and suggestions, please do not hesitate to contact us.</p> <p> </p> <p><strong>Contact email</strong>: ogomes@gmail.com</p> <p> </p> <p>If you use this dataset in your own work, please cite this DOI: 10.5281/zenodo.5020566</p> <p> </p> <p>Please also cite this paper, which provides additional details about the dataset:</p> <p>Michel Pedro Filippo, Otávio da Fonseca Martins Gomes, Gilson Alexandre Ostwald Pedro da Costa, Guilherme Lucio Abelha Mota. <em>Deep learning semantic segmentation of opaque and non-opaque minerals from epoxy resin in reflected light microscopy images</em>. <strong>Minerals Engineering</strong>, Volume 170, 2021, 107007, https://doi.org/10.1016/j.mineng.2021.107007.</p>
Twitter Poll: Is #OpenScience an essential rsrch skill Grad Schools should train in prep for #REF2020
<p>The Twitter Poll "Is #OpenScience an essential rsrch skill Grad Schools should train in prep for #REF2020" was run online in support of Horizon 2020 Project HEIRRI (Higher Education Insititutions & Responsible Research & Innovation) 1st Conference, 18 March 2016.</p> <p>The poll attracted 123 voters, 12,343 impressions and 517 engagements (4,2% conversion).</p> <p><em><strong>Event website:</strong></em><br /> HEIRRI 1st Conference http://heirri.eu/1st-heirri-conference/</p> <p><em><strong>CODE for EMBEDDING TWITTER POLL: </strong></em></p> <p><blockquote class="twitter-tweet" data-lang="en"><p lang="en" dir="ltr">Is <a href="https://twitter.com/hashtag/OpenScience?src=hash">#OpenScience</a> an essential rsrch skill Grad Schools should train in prep for <a href="https://twitter.com/hashtag/REF2020?src=hash">#REF2020</a> ? <a href="https://twitter.com/hashtag/OpenSci4Doc?src=hash">#OpenSci4Doc</a> <a href="https://twitter.com/HEIRRI_">@HEIRRI_</a></p>&mdash; Foster Open Science (@fosterscience) <a href="https://twitter.com/fosterscience/status/709650182800068608">March 15, 2016</a></blockquote><br /> <script async src="//platform.twitter.com/widgets.js" charset="utf-8"></script></p>
Training material for ChIP-seq analysis
<p>The data provided here are part of a Galaxy tutorial that analyzes ChIP-seq data from a study published by Wu et al., 2014 (DOI:10.1101/gr.164830.113). The goal of this study was to investigate "the dynamics of occupancy and the role in gene regulation of the transcription factor Tal1, a critical regulator of hematopoiesis, at multiple stages of hematopoietic differentiation." To this end, ChIP-seq experiments were performed in multiple mouse cell types including a G1E cell line and megakaryocytes, the two cell types represented here. The dataset contains biological replicate Tal1 ChIP-seq and input control experiments (*.fastqsanger files). Because of the long processing time for the large original files, we have downsampled the original raw data files to include only reads that align to chromosome 19 and a subset of interesting genomic loci (ChIPseq_regions_of_interest_v4.bed) pulled from the Wu et al. publication. Also included is a gene annotation file (RefSeq_gene_annotations_mm10.bed) with gene names added for viewing in a genome browser.</p>
Training material for de novo transcriptome reconstruction from RNA-seq data
<p>The data provided here are part of a Galaxy tutorial that analyzes RNA-seq data from a study published by Wu et al., 2014 (DOI:10.1101/gr.164830.113). The goal of this study was to investigate "the dynamics of occupancy and the role in gene regulation of the transcription factor Tal1, a critical regulator of hematopoiesis, at multiple stages of hematopoietic differentiation." To this end, RNA-seq libraries were constructed from multiple mouse cell types including G1E - a GATA-null immortalized cell line derived from targeted disruption of GATA-1 in mouse embryonic stem cells - and megakaryocytes. This RNA-seq data was used to determine differential gene expression between G1E and megakaryocytes and later correlated with Tal1 occupancy. This dataset (GEO Accession: GSE51338) consists of biological replicate, paired-end, polyA selected RNA-seq libraries. Because of the long processing time for the large original files, we have downsampled the original raw data files to include only reads that align to chromosome 19 and a subset of interesting genomic loci identified by Wu et al.</p>
Embodied Spatial Navigation Training in Mild Cognitive Impairment: A Proof-of-Concept Trial
<p>Raw data of included cognitive test and VR data Starting Grant Ricerca Finalizzata, code: SG-2018-12368175</p>
RDP Classifier 2.14 and the RDP bacterial and archaeal taxonomy training set No. 19
<p>RDP Classifier 2.14 (August 2023) Release Note:</p><p>The Bacteria and Archaea hierarchy model used by RDP Classifier has been updated to training set No. 19. The new version has over 600 genera and 2500 species added since last version No. 18 released in July 2020. The information that is used to update the RDP taxonomy to training set version No. 19, and RDP Classifier version 2.14 came from publicly available scientific articles and public sequence repository, mostly from International Journal of Systematic and Evolutionary Microbiology (IJSEM), the All-Species Living Tree Project (LTP) and GenBank. </p><p>It is worth noting that most of the phyla have new names, according to article "</p><p>Oren A, Garrity GM. Valid publication of the names of forty-two phyla of prokaryotes. Int J Syst Evol Microbiol. 2021 Oct;71(10). doi: 10.1099/ijsem.0.005056. PMID: 34694987."</p><p>In addition to the files to train and run the RDP Classifier, new file formats are made available to accommodate the needs of users:</p><p>1. A new file trainset19_072023_speciesrank.fa has been added to the release in RDPClassifier_16S_trainsetNo19_rawtrainingdata.zip. This file is NOT needed to train the classifier. In addition to sequences, it contains genus, species, strain, type status and taxonomy rank, which are useful for closest species identification using third-party tools (e.g. BLAST).</p><p>2. Two new files in RDPClassifier_16S_trainsetNo19_QiimeFormat.zip to retrain the RDP Classifier included in Qiime2 package.</p>
FoldingDiff CATH S40 training dataset
<p>Dataset used to develop and train FoldingDiff, a generative model for protein backbone structures. </p>
Training and test datasets for the PredictONCO tool
<p>This dataset was used for training and validating the <a href="https://loschmidt.chemi.muni.cz/predictonco/">PredictONCO </a>web tool, supporting decision-making in precision oncology by extending the bioinformatics predictions with advanced computing and machine learning. The dataset consists of 1073 single-point mutants of 42 proteins, whose effect was classified as Oncogenic (509 data points) and Benign (564 data points). All mutations were annotated with a clinically verified effect and were compiled from the ClinVar and OncoKB databases. The dataset was manually curated based on the available information in other precision oncology databases (The Clinical Knowledgebase by The Jackson Laboratory, Personalized Cancer Therapy Knowledge Base by MD Anderson Cancer Center, cBioPortal, DoCM database) or in the primary literature. To create the dataset, we also removed any possible overlaps with the data points used in the PredictSNP consensus predictor and its constituents. This was implemented to avoid any test set data leakage due to using the PredictSNP score as one of the features (see below).</p> <p>The entire dataset (<strong>SEQ</strong>) was further annotated by the pipeline of PredictONCO. Briefly, the following six features were calculated regardless of the structural information available: essentiality of the mutated residue (yes/no), the conservation of the position (the conservation grade and score), the domain where the mutation is located (cytoplasmic, extracellular, transmembrane, other), the PredictSNP score, and the number of essential residues in the protein. For approximately half of the data (<strong>STR</strong>: 377 and 76 oncogenic and benign data points, respectively), the structural information was available, and six more features were calculated: FoldX and Rosetta ddg_monomer scores, whether the residue is in the catalytic pocket (identification of residues forming the ligand-binding pocket was obtained from P2Rank), and the pKa changes (the minimum and maximum changes as well as the number of essential residues whose pKa was changed – all values obtained from PROPKA3). For both <strong>STR </strong>and <strong>SEQ </strong>datasets, 20% of the data was held out for testing. The data split was implemented at the position level to ensure that no position from the test data subset appears in the training data subset. </p> <p>For more details about the tool, please visit the <a href="https://loschmidt.chemi.muni.cz/predictonco/help">help page</a> or <a href="https://loschmidt.chemi.muni.cz/peg/contact/">get in touch with us</a>.</p> <p>14-Dec-2023 update: the file with features<em> PredictONCO-features.txt</em> now includes UniProt IDs, transcripts, PDB codes, and mutations.</p>
RCT NOVICE Surgeon training Lübeck Toolbox
<p>Surgeon residents were rated by GOALS score at their first operations by masked raters. Some had trained with the Lübeck Toolbox until sufficiently proficient in a waiting group design. </p> <p>The <em><strong>study protocol</strong></em> was published here: <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.isjp.2020.02.004" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.isjp.2020.02.004</a></p> <p>The article with the <em><strong>results</strong></em> is published here: <a href="http://dx.doi.org/10.1097/JS9.0000000000002304">http://dx.doi.org/10.1097/JS9.0000000000002304</a> </p>
Eleven years of training data for south foehn for three regions of Western Austria
<p>This south foehn training data is suited for machine learning purposes. </p> <p>It was created by applying objective foehn classification (OFC, Vergeiner 2004) on hourly data of various stations in Western Austria. Three regions (Vorarlberg, Tiroler Unterland, Tiroler Oberland) and two intensities are available, where</p> <ul> <li>0.0 means no foehn on that day,</li> <li>0.5 means localised foehn on that day (up to half the stations in the region responded to OFC),</li> <li>1.0 means widespread foehn on that day (more than half the stations in the region responded to OFC),</li> </ul> <p>provided for each region individually.</p> <p>A paper, where the process of creation is described, is in preperation and will be linked as soon as it is reviewed. </p> <p> </p>
Dataset: The effects of class balance on the training energy consumption of logistic regression models
<p>Two synthetic datasets for binary classification, generated with the Random Radial Basis Function generator from WEKA. They are the same shape and size (104.952 instances, 185 attributes), but the "balanced" dataset has 52,13% of its instances belonging to class c0, while the "unbalanced" one only has 4,04% of its instances belonging to class c0. Therefore, this set of datasets is primarily meant to study how class balance influences the behaviour of a machine learning model.</p>
XR training and gameplay 6DoF mobility dataset
<p>User mobility in extended reality (XR) can have a major impact on millimeter-wave (mmWave) links and may require dedicated mitigation strategies to ensure reliable connections and avoid service outages. The available prior art has predominantly focused on XR applications with constrained user mobility and limited impact on mmWave channels.</p> <p>We have performed dedicated experiments to extend the characterisation of relevant future XR use cases featuring a high degree of user mobility. To this end, we have carried out a tailor-made XR mobility measurement campaign, capturing the movement of the head, hands, and body in 6DoF. </p> <p>For a usage example, see the provided Jupyter Notebook in cacerumd-usage-example.zip.</p> <p>A description of the measurement campaign and a characterisation of the recorded mobility can be found in the corresponding <a href="https://ieeexplore.ieee.org/abstract/document/10634047">IEEE Magazine paper</a> or a longer version, with more details about the experiment, on <a href="https://arxiv.org/abs/2407.02636">Arxiv</a>.</p>
BTSbot v10 training set
<p>This is the production version of the BTSbot training set, limited to public (programid=1) ZTF alerts. BTSbot is a multi-modal convolutional neural network designed for real-time identification bright extragalactic transients in Zwicky Transient Facility (ZTF) data. BTSbot provides a bright transient score to individual ZTF detections using their image data and 25 extracted features. BTSbot is able to eliminate the need for daily visual inspection of new transients by automatically identifying and requesting spectroscopic follow-up observations of new bright transient candidates. </p> <p>The training data is split into two zipped files. metadata_v10.zip contains alert packet features for the alerts in the train, validation, and test splits stored a separate .csv files. images_v10.zip contains the corresponding image cutouts stored as three .npy files. The BTSbot source code contains routines for reading these files and training a model on them. They can also easily be loaded with pandas.read_csv() and numpy.load().</p> <p>This training set data is necessary for reproducing the results of the BTSbot study, although this dataset only contains ZTF public data while the production BTSbot model also trained on ZTF partnership data.</p> <p>If you use reference this data or BTSbot please cite the <a href="https://ui.adsabs.harvard.edu/abs/2024arXiv240115167R/abstract" target="_blank" rel="noopener">BTSbot paper</a>.</p> <p> </p>
Data-driven physics-based modeling of pedestrian dynamics - dataset: Pedestrian trajectories at Eindhoven train station
<p>Pedestrian trajectories measured at train station Eindhoven Centraal (the Netherlands) on platform 2 with acces to tracks 3 and 4.</p> <p>The dataset is partitioned in files containing 10 consecutive days each, recording 4 data fields:</p> <ul> <li><strong>time_ms:</strong> Passed time since start of the measurements. Unit: milliseconds.</li> <li><strong>object_identifier:</strong> unique id identifying an object.</li> <li><strong>x_position_mm: </strong>coordinates of the object along the x-axis at the given time. Unit: millimeters.</li> <li><strong>y_position_mm:</strong> coordinates of the object along the y-axis at the given time. Unit: millimeters.</li> </ul> <p>Each object resembles a pedestrian on the train platform recorded with 10 frames per second. We deliberately removed exact date and time information for privacy reasons (see additional note). The data set consists of 60 consecutive days starting at an unkown time between 00:00 AM and 01:00 AM of a random date between April 1st and May 1st 2022. An overhead image of the platform is included showing train track 3 in the bottom and train track 4 in the top of the image.</p> <p>The data set is supplemented to the paper <a title="Data-driven physics-based modeling of pedestrian dynamics" href="https://doi.org/10.48550/arXiv.2407.20794" target="_blank" rel="noopener">Data-driven physics-based modeling of pedestrian dynamics</a> and can be processed by the associated <a title="Software: Data-driven physics-based modeling of pedestrian dynamics" href="https://github.com/c-pouw/physics-based-pedestrian-modeling" target="_blank" rel="noopener">Python implementation</a> to create pedestrian models. </p>
Concentrating solar power (CSP) plants AI-training dataset for flux density measurements.
<p>In this dataset, the tools required for the training of a neural net in the context of flux density measurements in concentrating solar power (CSP) plants are included. An Excel file with 931 meteorological conditions and the positions of the power plant and the receiver is included, as well as 15928 pairs of images resulting from ray-tracing in Solarturm Juelich (STJ) each of these conditions with 17 different combinations of heliostats. <br> <br>This dataset is part of the WP1 of TOPCSP european project (funded by HORIZON MSCA Doctoral Network, Project number 101072537).</p>
Sentence/Table Pair Data from Wikipedia for Pre-training with Distant-Supervision
<p>This is the dataset used for pre-training in "<em>ReasonBERT: Pre-trained to Reason with Distant Supervision</em>", EMNLP'21.</p> <p>There are two files:</p> <p>sentence_pairs_for_pretrain_no_tokenization.tar.gz -> Contain only sentences as evidence, Text-only</p> <p>table_pairs_for_pretrain_no_tokenization.tar.gz -> At least one piece of evidence is a table, Hybrid</p> <p>The data is chunked into multiple tar files for easy loading. We use <a href="https://github.com/webdataset/webdataset">WebDataset</a>, a PyTorch Dataset (IterableDataset) implementation providing efficient sequential/streaming data access.</p> <p>For pre-training code, or if you have any questions, please check our GitHub repo https://github.com/sunlab-osu/ReasonBERT</p> <p>Below is a sample code snippet to load the data</p> <pre><code class="language-python">import webdataset as wds # path to the uncompressed files, should be a directory with a set of tar files url = './sentence_multi_pairs_for_pretrain_no_tokenization/{000000...000763}.tar' dataset = ( wds.Dataset(url) .shuffle(1000) # cache 1000 samples and shuffle .decode() .to_tuple("json") .batched(20) # group every 20 examples into a batch ) # Please see the documentation for WebDataset for more details about how to use it as dataloader for Pytorch # You can also iterate through all examples and dump them with your preferred data format</code></pre> <p>Below we show how the data is organized with two examples.</p> <p>Text-only</p> <pre><code>{'s1_text': 'Sils is a municipality in the comarca of Selva, in Catalonia, Spain.', # query sentence 's1_all_links': { 'Sils,_Girona': [[0, 4]], 'municipality': [[10, 22]], 'Comarques_of_Catalonia': [[30, 37]], 'Selva': [[41, 46]], 'Catalonia': [[51, 60]] }, # list of entities and their mentions in the sentence (start, end location) 'pairs': [ # other sentences that share common entity pair with the query, group by shared entity pairs { 'pair': ['Comarques_of_Catalonia', 'Selva'], # the common entity pair 's1_pair_locs': [[[30, 37]], [[41, 46]]], # mention of the entity pair in the query 's2s': [ # list of other sentences that contain the common entity pair, or evidence { 'md5': '2777e32bddd6ec414f0bc7a0b7fea331', 'text': 'Selva is a coastal comarque (county) in Catalonia, Spain, located between the mountain range known as the Serralada Transversal or Puigsacalm and the Costa Brava (part of the Mediterranean coast). Unusually, it is divided between the provinces of Girona and Barcelona, with Fogars de la Selva being part of Barcelona province and all other municipalities falling inside Girona province. Also unusually, its capital, Santa Coloma de Farners, is no longer among its larger municipalities, with the coastal towns of Blanes and Lloret de Mar having far surpassed it in size.', 's_loc': [0, 27], # in addition to the sentence containing the common entity pair, we also keep its surrounding context. 's_loc' is the start/end location of the actual evidence sentence 'pair_locs': [ # mentions of the entity pair in the evidence [[19, 27]], # mentions of entity 1 [[0, 5], [288, 293]] # mentions of entity 2 ], 'all_links': { 'Selva': [[0, 5], [288, 293]], 'Comarques_of_Catalonia': [[19, 27]], 'Catalonia': [[40, 49]] } } ,...] # there are multiple evidence sentences }, ,...] # there are multiple entity pairs in the query }</code></pre> <p>Hybrid</p> <pre><code>{'s1_text': 'The 2006 Major League Baseball All-Star Game was the 77th playing of the midseason exhibition baseball game between the all-stars of the American League (AL) and National League (NL), the two leagues comprising Major League Baseball.', 's1_all_links': {...}, # same as text-only 'sentence_pairs': [{'pair': ..., 's1_pair_locs': ..., 's2s': [...]}], # same as text-only 'table_pairs': [ 'tid': 'Major_League_Baseball-1', 'text':[ ['World Series Records', 'World Series Records', ...], ['Team', 'Number of Series won', ...], ['St. Louis Cardinals (NL)', '11', ...], ...] # table content, list of rows 'index':[ [[0, 0], [0, 1], ...], [[1, 0], [1, 1], ...], ...] # index of each cell [row_id, col_id]. we keep only a table snippet, but the index here is from the original table. 'value_ranks':[ [0, 0, ...], [0, 0, ...], [0, 10, ...], ...] # if the cell contain numeric value/date, this is its rank ordered from small to large, follow TAPAS 'value_inv_ranks': [], # inverse rank 'all_links':{ 'St._Louis_Cardinals': { '2': [ [[2, 0], [0, 19]], # [[row_id, col_id], [start, end]] ] # list of mentions in the second row, the key is row_id }, 'CARDINAL:11': {'2': [[[2, 1], [0, 2]]], '8': [[[8, 3], [0, 2]]]}, } 'name': '', # table name, if exists 'pairs': { 'pair': ['American_League', 'National_League'], 's1_pair_locs': [[[137, 152]], [[162, 177]]], # mention in the query 'table_pair_locs': { '17': [ # mention of entity pair in row 17 [ [[17, 0], [3, 18]], [[17, 1], [3, 18]], [[17, 2], [3, 18]], [[17, 3], [3, 18]] ], # mention of the first entity [ [[17, 0], [21, 36]], [[17, 1], [21, 36]], ] # mention of the second entity ] } } ] }</code></pre> <p> </p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.