Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,487
datasets available to search
ShareScore release 0.9.0
Dataset results
1,487 results for “Tags”
Complete tag loss in capture-recapture studies affects abundance estimates: an elephant seal case study
<p>1. In capture-recapture studies, recycled individuals occur when individuals lose all of their tags and are recaptured as though they were new individuals. Typically, the effect of these recycled individuals is assumed negligible.</p> <p>2. Through a simulation-based study of double tagging experiments, we examined the effect of recycled individuals on parameter estimates in the Jolly-Seber model with tag loss (Cowen & Schwarz, 2006). We validated the simulation framework using long-term census data of elephant seals.</p> <p>3. Including recycled individuals did not affect estimates of capture, survival, and tag-retention probabilities. However, with low tag-retention rates, high capture rates, and high survival rates, recycled individuals produced over estimates of population size. For the elephant seal case study, we found population size estimates to be between 8 and 53% larger when recycled individuals were ignored.</p> <p>4. Ignoring the effects of recycled individuals can cause large biases in population size estimates. These results are particularly noticeable in longer studies.</p>
Monitoring of cow location in barn by an open source low cost low energy Bluetooth tag system
<p>Supplementary materials for: Bloch, V., Pastell, M., 2020. Monitoring of Cow Location in a Barn by an Open-Source, Low-Cost, Low-Energy Bluetooth Tag System. Sensors 20, 3841. <a href="https://doi.org/10.3390/s20143841">https://doi.org/10.3390/s20143841</a></p>
Strong chemical tagging with APOGEE: 21 candidate star clusters that have dissolved across the Milky Way disc
<p>Two files containing the chemically tagged abundances derived by astroNN from the Apache Point Observatory Galactic Evolution Experiment. The chemical tagging procedure is described in <a href="https://arxiv.org/abs/2004.04263">https://arxiv.org/abs/2004.04263</a>. DBSCAN_labels_APOGEE_DR16.npy contains only members of groups with more than 15 members and silhouette coefficients greater than 0. DBSCAN_labels_APOGEE_DR16_all.npy contains labels for all stars in the quality controlled data set. The data are packaged as a numpy structured array. Most of the columns are described at <a href="https://data.sdss.org/datamodel/files/APOGEE_ASTRONN/apogee_astronn.html">https://data.sdss.org/datamodel/files/APOGEE_ASTRONN/apogee_astronn.html</a>. There are two additional columns: CLUSTER_ID, and SILHOUETTE_COEFF, which are respectively the label and silhouette coefficient for each star.</p>
User-aware music auto-tagging with contextual tags
<p>This is a user-aware music dataset labeled with the contextual use of each track according to each user. The dataset is composed of 10 contextual tags extracted based on user's usage through created playlists in the Deezer catalog. The tags are: " car, gym, happy, night, relax, running, sad, summer, work, workout". For each track/user pair, a contextual tag is associated with it indicating that the user listens to the track in the associated context. Additionally, the users are represented as embeddings based on their listening history computed through the matrix factorization of the user/track matrix.</p> <p>The creation of the dataset and the baseline of our auto-tagging model is described in the paper: Karim M. Ibrahim, Elena V. Epure, Geoffroy Peeters, and Gaël Richard. "Should we consider the users in contextual music auto-tagging models?" <em>21st International Society for Music Information Retrieval Conference (ISMIR)</em>. 2020. The source code of the paper is available here: <a href="https://github.com/KarimMibrahim/user-aware-music-autotagging">https://github.com/KarimMibrahim/user-aware-music-autotagging</a></p> <p>The dataset is composed of the SONG_ID which is the ID of the track in the Deezer catalog. Each track/user pair is labeled with each tag as either 1 (indicating a track's presence in the context) or 0 (indicating a track's absence). The 30 seconds track previews used to train the model in the paper can be accessed through the Deezer API: <a href="https://developers.deezer.com/api">https://developers.deezer.com/api</a>. Each user is represented with an anonymized USER_ID which is associated with the user embedding available in the user_embeddings.csv file. </p>
SONYC Urban Sound Tagging (SONYC-UST): a multilabel dataset from an urban acoustic sensor network
<p><strong>SONYC Urban Sound Tagging (SONYC-UST): a multilabel dataset from an urban acoustic sensor network</strong></p> <p>Version 2.3, September 2020</p> <p> </p> <p><strong>Created by</strong></p> <p>Mark Cartwright (1,2,3), Jason Cramer (1), Ana Elisa Mendez Mendez (1), Yu Wang (1), Ho-Hsiang Wu (1), Vincent Lostanlen (1,2,4), Magdalena Fuentes (1), Graham Dove (2), Charlie Mydlarz (1,2), Justin Salamon (5), Oded Nov (6), Juan Pablo Bello (1,2,3)</p> <ol> <li>Music and Audio Research Lab, New York University</li> <li>Center for Urban Science and Progress, New York University</li> <li>Department of Computer Science and Engineering, New York University</li> <li>Cornell Lab of Ornithology</li> <li>Adobe Research</li> <li>Department of Technology Management and Innovation, New York University</li> </ol> <p> </p> <p><strong>Publication</strong></p> <p>If using this data in an academic work, please reference the DOI and version, as well as cite the following paper, which presented the data collection procedure and the first version of the dataset:</p> <p>Cartwright, M., Cramer, J., Mendez, A.E.M., Wang, Y., Wu, H., Lostanlen, V., Fuentes, M., Dove, G., Mydlarz, C., Salamon, J., Nov, O., Bello, J.P. SONYC-UST-V2: An Urban Sound Tagging Dataset with Spatiotemporal Context. In <em>Proceedings of the Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE)</em>, 2020.<br> <a href="https://arxiv.org/abs/2009.05188">[pdf]</a></p> <p> </p> <p><strong>Description</strong></p> <p>SONYC Urban Sound Tagging (SONYC-UST) is a dataset for the development and evaluation of machine listening systems for realistic urban noise monitoring. The audio was recorded from the <a href="https://wp.nyu.edu/sonyc">SONYC</a> acoustic sensor network. Volunteers on the <a href="https://zooniverse.org">Zooniverse</a> citizen science platform tagged the presence of 23 classes that were chosen in consultation with the New York City Department of Environmental Protection. These 23 fine-grained classes can be grouped into 8 coarse-grained classes. The recordings are split into three sets: training, validation, and test. The training and validation sets are disjoint with respect to the sensor from which each recording came, and the test set is displaced in time. For increased reliability, three volunteers annotated each recording. In addition, members of the SONYC team subsequently created a subset of verified, ground-truth tags using a two-stage annotation procedure in which two annotators independently tagged and then collectively resolved any disagreements. This subset of recordings with verified annotations intersects with all three recording splits. All of the recordings in the test set have these verified annotations. In v2 version of this dataset, we have also included coarse spatiotemporal context information to aid in tag prediction when time and location is known. For more details on the motivation and creation of this dataset see the <a href="http://dcase.community/challenge2020/task-urban-sound-tagging-with-spatiotemporal-context">DCASE 2020 Urban Sound Tagging with Spatiotemporal Context Task website</a>.</p> <p> </p> <p><strong>Audio data</strong></p> <p>The provided audio has been acquired using the SONYC acoustic sensor network for urban noise pollution monitoring. Over 60 different sensors have been deployed in New York City, and these sensors have collectively gathered the equivalent of over 50 years of audio data, of which we provide a small subset. The data was sampled by selecting the nearest neighbors on VGGish features of recordings known to have classes of interest. All recordings are 10 seconds and were recorded with identical microphones at identical gain settings. To maintain privacy, we quantized the spatial information to the level of a city block, and we quantized the temporal information to the level of an hour. We also limited the occurrence of recordings with positive human voice annotations to one per hour per sensor.</p> <p> </p> <p><strong>Label taxonomy</strong></p> <p>The label taxonomy is as follows:</p> <ol> <li>engine<br> 1: small-sounding-engine<br> 2: medium-sounding-engine<br> 3: large-sounding-engine<br> X: engine-of-uncertain-size</li> <li>machinery-impact<br> 1: rock-drill<br> 2: jackhammer<br> 3: hoe-ram<br> 4: pile-driver<br> X: other-unknown-impact-machinery</li> <li>non-machinery-impact<br> 1: non-machinery-impact</li> <li>powered-saw<br> 1: chainsaw<br> 2: small-medium-rotating-saw<br> 3: large-rotating-saw<br> X: other-unknown-powered-saw</li> <li>alert-signal<br> 1: car-horn<br> 2: car-alarm<br> 3: siren<br> 4: reverse-beeper<br> X: other-unknown-alert-signal</li> <li>music<br> 1: stationary-music<br> 2: mobile-music<br> 3: ice-cream-truck<br> X: music-from-uncertain-source</li> <li>human-voice<br> 1: person-or-small-group-talking<br> 2: person-or-small-group-shouting<br> 3: large-crowd<br> 4: amplified-speech<br> X: other-unknown-human-voice</li> <li>dog<br> 1: dog-barking-whining</li> </ol> <p>The classes preceded by an <code>X</code> code indicate when an annotator was able to identify the coarse class, but couldn’t identify the fine class because either they were uncertain which fine class it was or the fine class was not included in the taxonomy. <code>dcase-ust-taxonomy.yaml</code> contains this taxonomy in an easily machine-readable form.</p> <p> </p> <p><strong>Data splits</strong></p> <p>This release contains a training subset (13538 recordings from 35 sensors), and validation subset (4308 recordings from 9 sensors), and a test subset (669 recordings from 48 sensors). The training and validation subsets are disjoint with respect to the sensor from which each recording came. The sensors in the test set will not disjoint from the training and validation subsets, but the test recordings are displaced in time, occurring after any of the recordings in the training and validation subset. The subset of recordings with verified annotations (1380 recordings) intersects with all three recording splits. All of the recordings in the test set have these verified annotations.</p> <p> </p> <p><strong>Annotation data</strong></p> <p>The annotation data are contained in <code>annotations.csv</code>, and encompass the training, validation, and test subsets. Each row in the file represents one multi-label annotation of a recording—it could be the annotation of a single citizen science volunteer, a single SONYC team member, or the agreed-upon ground truth by the SONYC team (see the <em>annotator_id</em> column description for more information). Note that since the SONYC team members annotated each class group separately, there may be multiple annotation rows by a single SONYC team annotator for a particular audio recording.</p> <p> </p> <p> </p> <p><strong>Columns</strong></p> <p><em>split</em></p> <p>The data split. (<em>train</em>, <em>validate, test</em>)</p> <p><em>sensor_id</em></p> <p>The ID of the sensor the recording is from.</p> <p><em>audio_filename</em></p> <p>The filename of the audio recording</p> <p><em>annotator_id</em></p> <p>The anonymous ID of the annotator. If this value is positive, it is a citizen science volunteer from the Zooniverse platform. If it is negative, it is a SONYC team member. If it is <code>0</code>, then it is the ground truth agreed-upon by the SONYC team.</p> <p><em>year</em></p> <p>The year the recording is from.</p> <p><em>week</em></p> <p>The week of the year the recording is from.</p> <p><em>day</em></p> <p>The day of the week the recording is from, with Monday as the start (i.e. <code>0</code>=Monday).</p> <p><em>hour</em></p> <p>The hour of the day the recording is from</p> <p><em>borough</em><br> The NYC borough in which the sensor is located (<code>1</code>=Manhattan, <code>3</code>=Brooklyn, <code>4</code>=Queens). This corresponds to the first digit in the 10-digit NYC parcel number system known as Borough, Block, Lot (BBL).</p> <p><em>block</em></p> <p>The NYC block in which the sensor is located. This corresponds to digits 2—6 digit in the 10-digit NYC parcel number system known as Borough, Block, Lot (BBL).</p> <p><em>latitude</em></p> <p>The latitude coordinate of the <strong>block</strong> in which the sensor is located.</p> <p><em>longitude</em></p> <p>The longitude coordinate of the <strong>block</strong> in which the sensor is located.</p> <p><em><coarse_id>-<fine_id>_<fine_name>_presence</em></p> <p>Columns of this form indicate the presence of fine-level class. <code>1</code> if present, <code>0</code> if not present. If <code>-1</code>, then the class was not labeled in this annotation because the annotation was performed by a SONYC team member who only annotated one coarse group of classes at a time when annotating the verified subset.</p> <p><em><coarse_id>_<coarse_name>_presence</em></p> <p>Columns of this form indicate the presence of a coarse-level class. <code>1</code> if present, <code>0</code> if not present. If <code>-1</code>, then the class was not labeled in this annotation because the annotation was performed by a SONYC team member who only annotated one coarse group of classes at a time when annotating the verified subset. These columns are computed from the fine-level class presence columns and are presented here for convenience when training on only coarse-level classes.</p> <p><em><coarse_id>-<fine_id>_<fine_name>_proximity</em></p> <p>Columns of this form indicate the proximity of a fine-level class. After indicating the presence of a fine-level class, citizen science annotators were asked to indicate the proximity of the sound event to the sensor. Only the citizen science volunteers performed this task, and therefore this data is not included in the verified annotations. This column may take on one of the following four values: (<code>near</code>, <code>far</code>, <code>notsure</code>, <code>-1</code>). If <code>-1</code>, then the proximity was not annotated because either the annotation was not performed by a citizen science volunteer, or the citizen science volunteer did not indicate the presence of the class.</p> <p> </p> <p><strong>Conditions of use</strong></p> <p>Dataset created by Mark Cartwright, Jason Cramer, Ana Elisa Mendez Mendez, Yu Wang, Ho-Hsiang Wu, Vincent Lostanlen, Magdalena Fuentes, Graham Dove, Charlie Mydlarz, Justin Salamon, Oded Nov, and Juan Pablo Bello</p> <p>The SONYC-UST dataset is offered free of charge under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) license:<br> <a href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</a></p> <p>The dataset and its contents are made available on an “as is” basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, New York University is not liable for, and expressly excludes all liability for, loss or damage however and whenever caused to anyone by any use of the SONYC-UST dataset or any part of it.</p> <p> </p> <p><strong>Feedback</strong></p> <p>Please help us improve SONYC-UST by sending your feedback to:</p> <ul> <li>Mark Cartwright: <a href="mailto:mcartwright@gmail.com">mcartwright@gmail.com</a></li> </ul> <p>In case of a problem, please include as many details as possible.</p> <p> </p> <p><strong>Acknowledgments</strong></p> <p>We would like to thank all the Zooniverse volunteers who continue to contribute to our project. This work is supported by <a href="https://www.nsf.gov/awardsearch/showAward?AWD_ID=1544753">National Science Foundation award 1544753</a>.</p> <p> </p> <p><strong>Change log</strong></p> <ul> <li>2.3 Added the ground truth annotations for the test set, and regrouped the audio files for upload to Zenodo.</li> <li>2.2 Added the audio for the test set (audio-eval.tar.gz).</li> <li>2.1 The DCASE 2020 development dataset. 14778 new recordings added along with coarse spatiotemporal context information.</li> <li>1.0 Data is the same as v0.4. Publication added to README.</li> <li>0.4 Fixed error in annotations. Previously, the coarse class "machinery-impact" was accidentally indicated as present whenever "non-machinery-impact" was present regardless of the presence of "machinery-impact". This error has been fixed.</li> <li>0.3 Test set annotations added</li> <li>0.2 Test set audio files added</li> </ul>
Dual DNA/protein tagging of open chromatin unveils dynamics of epigenomic landscapes in leukemia
<p>The architecture of chromatin specifies eukaryotic cell identity by controlling transcription factor access to sites of gene regulation. Here we describe a dual transposase/peroxidase approach, integrative DNA And Protein Tagging (iDAPT), which detects both DNA (iDAPT-seq) and protein (iDAPT-MS) associated with accessible regions of chromatin. In addition to direct identification of bound transcription factors, iDAPT enables the inference of their gene regulatory networks, protein interactors, and regulation of chromatin accessibility. We applied iDAPT to profile the epigenomic consequences of granulocytic differentiation of acute promyelocytic leukemia, yielding previously undescribed mechanistic insights with potential therapeutic implications. Our findings demonstrate the power of iDAPT as a discovery platform for both the dynamic epigenomic landscapes and their transcription factor components associated with biological phenomena and disease.</p>
GeoVectors-Antarctica-tags (v1.0)
<p><strong>Description</strong></p> <p>The GeoVectors corpus is a comprehensive large-scale linked open corpus of OpenStreetMap (https://www.openstreetmap.org/) entity embeddings that provides latent representations of over 980 million entities. The GeoVectors capture the semantic and geographic dimensions of OpenStreetMap entities and make them directly accessible to machine learning applications. The "-tags" datasets provide embeddings that capture the semantic dimension of OpenStreetMap entities. The "-location" datasets provide the geographic dimension.</p> <p><strong>Contents</strong></p> <p>This dataset was derived from an OpenStreetMap snapshot that was taken on November 10, 2020 (© OpenStreetMap contributors).</p> <p>We provide the GeoVectors in region-specific subsets. This subset contains tag-embeddings for the region "Antarctica" including the following countries:</p> <ul> <li>Antarctica</li> </ul> <p><strong>File format</strong></p> <p>The embeddings are provided in the tab-separated values (tsv) format. Each row contains the embedding of a single OpenStreetMap entity. The first column contains the OpenStreetMap type and the second column contains the OpenStreetMap ID of the respective entity. The type can either be node (n), way (w), or relation (r). The remaining columns represent the dimensions of the embedding space. (See also header.tsv)</p> <p><strong>Further information:</strong></p> <p>For further information, please visit <a href="http://geovectors.l3s.uni-hannover.de">http://geovectors.l3s.uni-hannover.de</a></p> <p><strong>Funding</strong>:</p> <p>This work was partially funded by DFG, German Research Foundation (“WorldKG", DE 2299/2-1), the Federal Ministry of Education and Research (BMBF), Germany (“Simple-ML", 01IS18054), the Federal Ministry for Economic Affairs and Energy (BMWi), Germany (“d-E-mand", 01ME19009B), and the European Commission (EU H2020, “smashHit", grant-ID 871477).</p>
GeoVectors-Germany-Ways-tags (v1.0)
<p><strong>Description</strong></p> <p>The GeoVectors corpus is a comprehensive large-scale linked open corpus of OpenStreetMap (https://www.openstreetmap.org/) entity embeddings that provides latent representations of over 980 million entities. The GeoVectors capture the semantic and geographic similarities of OpenStreetMap entities and make them directly accessible to machine learning applications. The "-tags" datasets provide embeddings that capture the semantic similarities of OpenStreetMap entities. The "-location" datasets provide the geographic similarities.</p> <p><strong>Contents</strong></p> <p>This dataset was derived from an OpenStreetMap snapshot that was taken on November 10, 2020 (© OpenStreetMap contributors).</p> <p>We provide the GeoVectors in region-specific subsets. This subset contains tag-embeddings of the ways for the region "Germany" including the following countries:</p> <ul> <li>Germany</li> </ul> <p><strong>File format</strong></p> <p>The embeddings are provided in the tab-separated values (tsv) format. Each row contains the embedding of a single OpenStreetMap entity. The first column contains the OpenStreetMap type and the second column contains the OpenStreetMap id of the respective entity. The type can either be node (n), way (w), or relation (r). The remaining columns represent the dimensions of the embedding space. (See also header.tsv)</p> <p><strong>Further information:</strong></p> <p>For further information, please visit <a href="http://geovectors.l3s.uni-hannover.de">http://geovectors.l3s.uni-hannover.de</a></p> <p><strong>Funding</strong>:</p> <p>This work was partially funded by DFG, German Research Foundation (“WorldKG", DE 2299/2-1), the Federal Ministry of Education and Research (BMBF), Germany (“Simple-ML", 01IS18054), the Federal Ministry for Economic Affairs and Energy (BMWi), Germany (“d-E-mand", 01ME19009B), and the European Commission (EU H2020, “smashHit", grant-ID 871477).</p>
GeoVectors-Germany-Nodes-Relations-tags (v1.0)
<p><strong>Description</strong></p> <p>The GeoVectors corpus is a comprehensive large-scale linked open corpus of OpenStreetMap (https://www.openstreetmap.org/) entity embeddings that provides latent representations of over 980 million entities. The GeoVectors capture the semantic and geographic similarities of OpenStreetMap entities and make them directly accessible to machine learning applications. The "-tags" datasets provide embeddings that capture the semantic similarities of OpenStreetMap entities. The "-location" datasets provide the geographic similarities.</p> <p><strong>Contents</strong></p> <p>This dataset was derived from an OpenStreetMap snapshot that was taken on November 10, 2020 (© OpenStreetMap contributors).</p> <p>We provide the GeoVectors in region-specific subsets. This subset contains tag-embeddings for the nodes and relations of the region "Germany" including the following countries:</p> <ul> <li>Germany</li> </ul> <p><strong>File format</strong></p> <p>The embeddings are provided in the tab-separated values (tsv) format. Each row contains the embedding of a single OpenStreetMap entity. The first column contains the OpenStreetMap type and the second column contains the OpenStreetMap id of the respective entity. The type can either be node (n), way (w), or relation (r). The remaining columns represent the dimensions of the embedding space. (See also header.tsv)</p> <p><strong>Further information:</strong></p> <p>For further information, please visit <a href="http://geovectors.l3s.uni-hannover.de">http://geovectors.l3s.uni-hannover.de</a></p> <p><strong>Funding</strong>:</p> <p>This work was partially funded by DFG, German Research Foundation (“WorldKG", DE 2299/2-1), the Federal Ministry of Education and Research (BMBF), Germany (“Simple-ML", 01IS18054), the Federal Ministry for Economic Affairs and Energy (BMWi), Germany (“d-E-mand", 01ME19009B), and the European Commission (EU H2020, “smashHit", grant-ID 871477).</p>
GeoVectors-Italy-tags (v1.0)
<p><strong>Description</strong></p> <p>The GeoVectors corpus is a comprehensive large-scale linked open corpus of OpenStreetMap (https://www.openstreetmap.org/) entity embeddings that provides latent representations of over 980 million entities. The GeoVectors capture the semantic and geographic similarities of OpenStreetMap entities and make them directly accessible to machine learning applications. The "-tags" datasets provide embeddings that capture the semantic similarities of OpenStreetMap entities. The "-location" datasets provide the geographic similarities.</p> <p><strong>Contents</strong></p> <p>This dataset was derived from an OpenStreetMap snapshot that was taken on November 10, 2020 (© OpenStreetMap contributors).</p> <p>We provide the GeoVectors in region-specific subsets. This subset contains tag-embeddings for the region "Italy" including the following countries:</p> <ul> <li>Italy</li> </ul> <p><strong>File format</strong></p> <p>The embeddings are provided in the tab-separated values (tsv) format. Each row contains the embedding of a single OpenStreetMap entity. The first column contains the OpenStreetMap type and the second column contains the OpenStreetMap id of the respective entity. The type can either be node (n), way (w), or relation (r). The remaining columns represent the dimensions of the embedding space. (See also header.tsv)</p> <p><strong>Further information:</strong></p> <p>For further information, please visit <a href="http://geovectors.l3s.uni-hannover.de">http://geovectors.l3s.uni-hannover.de</a></p> <p><strong>Funding</strong>:</p> <p>This work was partially funded by DFG, German Research Foundation (“WorldKG", DE 2299/2-1), the Federal Ministry of Education and Research (BMBF), Germany (“Simple-ML", 01IS18054), the Federal Ministry for Economic Affairs and Energy (BMWi), Germany (“d-E-mand", 01ME19009B), and the European Commission (EU H2020, “smashHit", grant-ID 871477).</p>
GeoVectors-US-West-tags (v1.0)
<p><strong>Description</strong></p> <p>The GeoVectors corpus is a comprehensive large-scale linked open corpus of OpenStreetMap (https://www.openstreetmap.org/) entity embeddings that provides latent representations of over 980 million entities. The GeoVectors capture the semantic and geographic dimensions of OpenStreetMap entities and make them directly accessible to machine learning applications. The "-tags" datasets provide embeddings that capture the semantic dimension of OpenStreetMap entities. The "-location" datasets provide the geographic dimension.</p> <p><strong>Contents</strong></p> <p>This dataset was derived from an OpenStreetMap snapshot that was taken on November 10, 2020 (© OpenStreetMap contributors).</p> <p>We provide the GeoVectors in region-specific subsets. This subset contains tag-embeddings for the region "US-West" including the following countries:</p> <ul> <li>US-West</li> </ul> <p><strong>File format</strong></p> <p>The embeddings are provided in the tab-separated values (tsv) format. Each row contains the embedding of a single OpenStreetMap entity. The first column contains the OpenStreetMap type and the second column contains the OpenStreetMap id of the respective entity. The type can either be node (n), way (w), or relation (r). The remaining columns represent the dimensions of the embedding space. (See also header.tsv)</p> <p><strong>Further information:</strong></p> <p>For further information, please visit <a href="http://geovectors.l3s.uni-hannover.de">http://geovectors.l3s.uni-hannover.de</a></p> <p><strong>Funding</strong>:</p> <p>This work was partially funded by DFG, German Research Foundation (“WorldKG", DE 2299/2-1), the Federal Ministry of Education and Research (BMBF), Germany (“Simple-ML", 01IS18054), the Federal Ministry for Economic Affairs and Energy (BMWi), Germany (“d-E-mand", 01ME19009B), and the European Commission (EU H2020, “smashHit", grant-ID 871477).</p>
GeoVectors-South-America-tags (v1.0)
<p><strong>Description</strong></p> <p>The GeoVectors corpus is a comprehensive large-scale linked open corpus of OpenStreetMap (https://www.openstreetmap.org/) entity embeddings that provides latent representations of over 980 million entities. The GeoVectors capture the semantic and geographic similarities of OpenStreetMap entities and make them directly accessible to machine learning applications. The "-tags" datasets provide embeddings that capture the semantic similarities of OpenStreetMap entities. The "-location" datasets provide the geographic similarities.</p> <p><strong>Contents</strong></p> <p>This dataset was derived from an OpenStreetMap snapshot that was taken on November 10, 2020 (© OpenStreetMap contributors).</p> <p>We provide the GeoVectors in region-specific subsets. This subset contains tag-embeddings for the region "South-America" including the following countries:</p> <ul> <li>Argentina</li> <li>Bolivia</li> <li>Brazil</li> <li>Chile</li> <li>Colombia</li> <li>Ecuador</li> <li>Paraguay</li> <li>Peru</li> <li>Suriname</li> <li>Uruguay</li> <li>Venezuela</li> </ul> <p><strong>File format</strong></p> <p>The embeddings are provided in the tab-separated values (tsv) format. Each row contains the embedding of a single OpenStreetMap entity. The first column contains the OpenStreetMap type and the second column contains the OpenStreetMap id of the respective entity. The type can either be node (n), way (w), or relation (r). The remaining columns represent the dimensions of the embedding space. (See also header.tsv)</p> <p><strong>Further information:</strong></p> <p>For further information, please visit <a href="http://geovectors.l3s.uni-hannover.de">http://geovectors.l3s.uni-hannover.de</a></p> <p><strong>Funding</strong>:</p> <p>This work was partially funded by DFG, German Research Foundation (“WorldKG", DE 2299/2-1), the Federal Ministry of Education and Research (BMBF), Germany (“Simple-ML", 01IS18054), the Federal Ministry for Economic Affairs and Energy (BMWi), Germany (“d-E-mand", 01ME19009B), and the European Commission (EU H2020, “smashHit", grant-ID 871477).</p>
GeoVectors-Great-Britain-tags (v1.0)
<p><strong>Description</strong></p> <p>The GeoVectors corpus is a comprehensive large-scale linked open corpus of OpenStreetMap (https://www.openstreetmap.org/) entity embeddings that provides latent representations of over 980 million entities. The GeoVectors capture the semantic and geographic similarities of OpenStreetMap entities and make them directly accessible to machine learning applications. The "-tags" datasets provide embeddings that capture the semantic similarities of OpenStreetMap entities. The "-location" datasets provide the geographic similarities.</p> <p><strong>Contents</strong></p> <p>This dataset was derived from an OpenStreetMap snapshot that was taken on November 10, 2020 (© OpenStreetMap contributors).</p> <p>We provide the GeoVectors in region-specific subsets. This subset contains tag-embeddings for the region "Great-britain" including the following countries:</p> <ul> <li>Great-Britain</li> </ul> <p><strong>File format</strong></p> <p>The embeddings are provided in the tab-separated values (tsv) format. Each row contains the embedding of a single OpenStreetMap entity. The first column contains the OpenStreetMap type and the second column contains the OpenStreetMap id of the respective entity. The type can either be node (n), way (w), or relation (r). The remaining columns represent the dimensions of the embedding space. (See also header.tsv)</p> <p><strong>Further information:</strong></p> <p>For further information, please visit <a href="http://geovectors.l3s.uni-hannover.de">http://geovectors.l3s.uni-hannover.de</a></p> <p><strong>Funding</strong>:</p> <p>This work was partially funded by DFG, German Research Foundation (“WorldKG", DE 2299/2-1), the Federal Ministry of Education and Research (BMBF), Germany (“Simple-ML", 01IS18054), the Federal Ministry for Economic Affairs and Energy (BMWi), Germany (“d-E-mand", 01ME19009B), and the European Commission (EU H2020, “smashHit", grant-ID 871477).</p>
GeoVectors-Europe-West-tags (v1.0)
<p><strong>Description</strong></p> <p>The GeoVectors corpus is a comprehensive large-scale linked open corpus of OpenStreetMap (https://www.openstreetmap.org/) entity embeddings that provides latent representations of over 980 million entities. The GeoVectors capture the semantic and geographic similarities of OpenStreetMap entities and make them directly accessible to machine learning applications. The "-tags" datasets provide embeddings that capture the semantic similarities of OpenStreetMap entities. The "-location" datasets provide the geographic similarities.</p> <p><strong>Contents</strong></p> <p>This dataset was derived from an OpenStreetMap snapshot that was taken on November 10, 2020 (© OpenStreetMap contributors).</p> <p>We provide the GeoVectors in region-specific subsets. This subset contains tag-embeddings for the region "Europe-West" including the following countries:</p> <ul> <li>Andorra</li> <li>Austria</li> <li>Azores</li> <li>Belgium</li> <li>Denmark</li> <li>Faroe-Islands</li> <li>Guernsey-Jersey</li> <li>Ireland-and-Northern-Ireland</li> <li>Isle-Of-Man</li> <li>Liechtenstein</li> <li>Luxembourg</li> <li>Malta</li> <li>Monaco</li> <li>Norway</li> <li>Portugal</li> <li>Spain</li> <li>Switzerland</li> </ul> <p><strong>File format</strong></p> <p>The embeddings are provided in the tab-separated values (tsv) format. Each row contains the embedding of a single OpenStreetMap entity. The first column contains the OpenStreetMap type and the second column contains the OpenStreetMap id of the respective entity. The type can either be node (n), way (w), or relation (r). The remaining columns represent the dimensions of the embedding space. (See also header.tsv)</p> <p><strong>Further information:</strong></p> <p>For further information, please visit <a href="http://geovectors.l3s.uni-hannover.de">http://geovectors.l3s.uni-hannover.de</a></p> <p><strong>Funding</strong>:</p> <p>This work was partially funded by DFG, German Research Foundation (“WorldKG", DE 2299/2-1), the Federal Ministry of Education and Research (BMBF), Germany (“Simple-ML", 01IS18054), the Federal Ministry for Economic Affairs and Energy (BMWi), Germany (“d-E-mand", 01ME19009B), and the European Commission (EU H2020, “smashHit", grant-ID 871477).</p>
GeoVectors-Europe-East-tags (v1.0)
<p><strong>Description</strong></p> <p>The GeoVectors corpus is a comprehensive large-scale linked open corpus of OpenStreetMap (https://www.openstreetmap.org/) entity embeddings that provides latent representations of over 980 million entities. The GeoVectors capture the semantic and geographic similarities of OpenStreetMap entities and make them directly accessible to machine learning applications. The "-tags" datasets provide embeddings that capture the semantic similarities of OpenStreetMap entities. The "-location" datasets provide the geographic similarities.</p> <p><strong>Contents</strong></p> <p>This dataset was derived from an OpenStreetMap snapshot that was taken on November 10, 2020 (© OpenStreetMap contributors).</p> <p>We provide the GeoVectors in region-specific subsets. This subset contains tag-embeddings for the region "Europe-East" including the following countries:</p> <ul> <li>Albania</li> <li>Belarus</li> <li>Bosnia-Herzegovina</li> <li>Bulgaria</li> <li>Croatia</li> <li>Cyprus</li> <li>Czech-Republic</li> <li>Estonia</li> <li>Finland</li> <li>Georgia</li> <li>Greece</li> <li>Hungary</li> <li>Iceland</li> <li>Kosovo</li> <li>Latvia</li> <li>Lithuania</li> <li>Macedonia</li> <li>Moldova</li> <li>Montenegro</li> <li>Romania</li> <li>Serbia</li> <li>Slovakia</li> <li>Slovenia</li> <li>Sweden</li> <li>Turkey</li> <li>Ukraine</li> </ul> <p><strong>File format</strong></p> <p>The embeddings are provided in the tab-separated values (tsv) format. Each row contains the embedding of a single OpenStreetMap entity. The first column contains the OpenStreetMap type and the second column contains the OpenStreetMap id of the respective entity. The type can either be node (n), way (w), or relation (r). The remaining columns represent the dimensions of the embedding space. (See also header.tsv)</p> <p><strong>Further information:</strong></p> <p>For further information, please visit <a href="http://geovectors.l3s.uni-hannover.de">http://geovectors.l3s.uni-hannover.de</a></p> <p><strong>Funding</strong>:</p> <p>This work was partially funded by DFG, German Research Foundation (“WorldKG", DE 2299/2-1), the Federal Ministry of Education and Research (BMBF), Germany (“Simple-ML", 01IS18054), the Federal Ministry for Economic Affairs and Energy (BMWi), Germany (“d-E-mand", 01ME19009B), and the European Commission (EU H2020, “smashHit", grant-ID 871477).</p>
GeoVectors-Poland-tags (v1.0)
<p><strong>Description</strong></p> <p>The GeoVectors corpus is a comprehensive large-scale linked open corpus of OpenStreetMap (https://www.openstreetmap.org/) entity embeddings that provides latent representations of over 980 million entities. The GeoVectors capture the semantic and geographic similarities of OpenStreetMap entities and make them directly accessible to machine learning applications. The "-tags" datasets provide embeddings that capture the semantic similarities of OpenStreetMap entities. The "-location" datasets provide the geographic similarities.</p> <p><strong>Contents</strong></p> <p>This dataset was derived from an OpenStreetMap snapshot that was taken on November 10, 2020 (© OpenStreetMap contributors).</p> <p>We provide the GeoVectors in region-specific subsets. This subset contains tag-embeddings for the region "Poland" including the following countries:</p> <ul> <li>Poland</li> </ul> <p><strong>File format</strong></p> <p>The embeddings are provided in the tab-separated values (tsv) format. Each row contains the embedding of a single OpenStreetMap entity. The first column contains the OpenStreetMap type and the second column contains the OpenStreetMap id of the respective entity. The type can either be node (n), way (w), or relation (r). The remaining columns represent the dimensions of the embedding space. (See also header.tsv)</p> <p><strong>Further information:</strong></p> <p>For further information, please visit <a href="http://geovectors.l3s.uni-hannover.de">http://geovectors.l3s.uni-hannover.de</a></p> <p><strong>Funding</strong>:</p> <p>This work was partially funded by DFG, German Research Foundation (“WorldKG", DE 2299/2-1), the Federal Ministry of Education and Research (BMBF), Germany (“Simple-ML", 01IS18054), the Federal Ministry for Economic Affairs and Energy (BMWi), Germany (“d-E-mand", 01ME19009B), and the European Commission (EU H2020, “smashHit", grant-ID 871477).</p>
GeoVectors-Australia-Oceania-tags (v1.0)
<p><strong>Description</strong></p> <p>The GeoVectors corpus is a comprehensive large-scale linked open corpus of OpenStreetMap (https://www.openstreetmap.org/) entity embeddings that provides latent representations of over 980 million entities. The GeoVectors capture the semantic and geographic similarities of OpenStreetMap entities and make them directly accessible to machine learning applications. The "-tags" datasets provide embeddings that capture the semantic similarities of OpenStreetMap entities. The "-location" datasets provide the geographic similarities.</p> <p><strong>Contents</strong></p> <p>This dataset was derived from an OpenStreetMap snapshot that was taken on November 10, 2020 (© OpenStreetMap contributors).</p> <p>We provide the GeoVectors in region-specific subsets. This subset contains tag-embeddings for the region "Australia-Oceania" including the following countries:</p> <ul> <li>American-Oceania</li> <li>Australia</li> <li>Cook-Islands</li> <li>Fiji</li> <li>Ile-De-Clipperton</li> <li>Kiribati</li> <li>Marshall-Islands</li> <li>Micronesia</li> <li>Nauru</li> <li>New-Caledonia</li> <li>New-Zealand</li> <li>Niue</li> <li>Palau</li> <li>Papua-New-Guinea</li> <li>Pitcairn-Islands</li> <li>Polynesie-Francaise</li> <li>Samoa</li> <li>Solomon-Islands</li> <li>Tokelau</li> <li>Tonga</li> <li>Tuvalu</li> <li>Vanuatu</li> <li>Wallis-Et-Futuna</li> </ul> <p><strong>File format</strong></p> <p>The embeddings are provided in the tab-separated values (tsv) format. Each row contains the embedding of a single OpenStreetMap entity. The first column contains the OpenStreetMap type and the second column contains the OpenStreetMap id of the respective entity. The type can either be node (n), way (w), or relation (r). The remaining columns represent the dimensions of the embedding space. (See also header.tsv)</p> <p><strong>Further information:</strong></p> <p>For further information, please visit <a href="http://geovectors.l3s.uni-hannover.de">http://geovectors.l3s.uni-hannover.de</a></p> <p><strong>Funding</strong>:</p> <p>This work was partially funded by DFG, German Research Foundation (“WorldKG", DE 2299/2-1), the Federal Ministry of Education and Research (BMBF), Germany (“Simple-ML", 01IS18054), the Federal Ministry for Economic Affairs and Energy (BMWi), Germany (“d-E-mand", 01ME19009B), and the European Commission (EU H2020, “smashHit", grant-ID 871477).</p>
GeoVectors-North-America-tags (v1.0)
<p><strong>Description</strong></p> <p>The GeoVectors corpus is a comprehensive large-scale linked open corpus of OpenStreetMap (https://www.openstreetmap.org/) entity embeddings that provides latent representations of over 980 million entities. The GeoVectors capture the semantic and geographic similarities of OpenStreetMap entities and make them directly accessible to machine learning applications. The "-tags" datasets provide embeddings that capture the semantic similarities of OpenStreetMap entities. The "-location" datasets provide the geographic similarities.</p> <p><strong>Contents</strong></p> <p>This dataset was derived from an OpenStreetMap snapshot that was taken on November 10, 2020 (© OpenStreetMap contributors).</p> <p>We provide the GeoVectors in region-specific subsets. This subset contains tag-embeddings for the region "North-America" including the following countries:</p> <ul> <li>Canada</li> <li>Greenland</li> <li>Mexico</li> </ul> <p><strong>File format</strong></p> <p>The embeddings are provided in the tab-separated values (tsv) format. Each row contains the embedding of a single OpenStreetMap entity. The first column contains the OpenStreetMap type and the second column contains the OpenStreetMap id of the respective entity. The type can either be node (n), way (w), or relation (r). The remaining columns represent the dimensions of the embedding space. (See also header.tsv)</p> <p><strong>Further information:</strong></p> <p>For further information, please visit <a href="http://geovectors.l3s.uni-hannover.de">http://geovectors.l3s.uni-hannover.de</a></p> <p><strong>Funding</strong>:</p> <p>This work was partially funded by DFG, German Research Foundation (“WorldKG", DE 2299/2-1), the Federal Ministry of Education and Research (BMBF), Germany (“Simple-ML", 01IS18054), the Federal Ministry for Economic Affairs and Energy (BMWi), Germany (“d-E-mand", 01ME19009B), and the European Commission (EU H2020, “smashHit", grant-ID 871477).</p>
GeoVectors-Asia-tags (v1.0)
<p><strong>Description</strong></p> <p>The GeoVectors corpus is a comprehensive large-scale linked open corpus of OpenStreetMap (https://www.openstreetmap.org/) entity embeddings that provides latent representations of over 980 million entities. The GeoVectors capture the semantic and geographic similarities of OpenStreetMap entities and make them directly accessible to machine learning applications. The "-tags" datasets provide embeddings that capture the semantic similarities of OpenStreetMap entities. The "-location" datasets provide the geographic similarities.</p> <p><strong>Contents</strong></p> <p>This dataset was derived from an OpenStreetMap snapshot that was taken on November 10, 2020 (© OpenStreetMap contributors).</p> <p>We provide the GeoVectors in region-specific subsets. This subset contains tag-embeddings for the region "Asia" including the following countries:</p> <ul> <li>Afghanistan</li> <li>Armenia</li> <li>Azerbaijan</li> <li>Bangladesh</li> <li>Bhutan</li> <li>Cambodia</li> <li>China</li> <li>Gcc-States</li> <li>India</li> <li>Indonesia</li> <li>Iran</li> <li>Iraq</li> <li>Israel-and-Palestine</li> <li>Japan</li> <li>Jordan</li> <li>Kazakhstan</li> <li>Kyrgyzstan</li> <li>Laos</li> <li>Lebanon</li> <li>Malaysia-Singapore-Brunei</li> <li>Maldives</li> <li>Mongolia</li> <li>Myanmar</li> <li>Nepal</li> <li>North-Korea</li> <li>Pakistan</li> <li>Philippines</li> <li>South-Korea</li> <li>Sri-Lanka</li> <li>Syria</li> <li>Taiwan</li> <li>Tajikistan</li> <li>Thailand</li> <li>Turkmenistan</li> <li>Uzbekistan</li> <li>Vietnam</li> <li>Yemen</li> </ul> <p><strong>File format</strong></p> <p>The embeddings are provided in the tab-separated values (tsv) format. Each row contains the embedding of a single OpenStreetMap entity. The first column contains the OpenStreetMap type and the second column contains the OpenStreetMap id of the respective entity. The type can either be node (n), way (w), or relation (r). The remaining columns represent the dimensions of the embedding space. (See also header.tsv)</p> <p><strong>Further information:</strong></p> <p>For further information, please visit <a href="http://geovectors.l3s.uni-hannover.de">http://geovectors.l3s.uni-hannover.de</a></p> <p><strong>Funding</strong>:</p> <p>This work was partially funded by DFG, German Research Foundation (“WorldKG", DE 2299/2-1), the Federal Ministry of Education and Research (BMBF), Germany (“Simple-ML", 01IS18054), the Federal Ministry for Economic Affairs and Energy (BMWi), Germany (“d-E-mand", 01ME19009B), and the European Commission (EU H2020, “smashHit", grant-ID 871477).</p>
GeoVectors-France-tags (v1.0)
<p><strong>Description</strong></p> <p>The GeoVectors corpus is a comprehensive large-scale linked open corpus of OpenStreetMap (https://www.openstreetmap.org/) entity embeddings that provides latent representations of over 980 million entities. The GeoVectors capture the semantic and geographic similarities of OpenStreetMap entities and make them directly accessible to machine learning applications. The "-tags" datasets provide embeddings that capture the semantic similarities of OpenStreetMap entities. The "-location" datasets provide the geographic similarities.</p> <p><strong>Contents</strong></p> <p>This dataset was derived from an OpenStreetMap snapshot that was taken on November 10, 2020 (© OpenStreetMap contributors).</p> <p>We provide the GeoVectors in region-specific subsets. This subset contains tag-embeddings for the region "France" including the following countries:</p> <ul> <li>France</li> </ul> <p><strong>File format</strong></p> <p>The embeddings are provided in the tab-separated values (tsv) format. Each row contains the embedding of a single OpenStreetMap entity. The first column contains the OpenStreetMap type and the second column contains the OpenStreetMap id of the respective entity. The type can either be node (n), way (w), or relation (r). The remaining columns represent the dimensions of the embedding space. (See also header.tsv)</p> <p><strong>Further information:</strong></p> <p>For further information, please visit <a href="http://geovectors.l3s.uni-hannover.de">http://geovectors.l3s.uni-hannover.de</a></p> <p><strong>Funding</strong>:</p> <p>This work was partially funded by DFG, German Research Foundation (“WorldKG", DE 2299/2-1), the Federal Ministry of Education and Research (BMBF), Germany (“Simple-ML", 01IS18054), the Federal Ministry for Economic Affairs and Energy (BMWi), Germany (“d-E-mand", 01ME19009B), and the European Commission (EU H2020, “smashHit", grant-ID 871477).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.