Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,487
datasets available to search
ShareScore release 0.9.0
Dataset results
1,487 results for “Tagging”
Following/Followers and Tags on 0.1 million Twitter Users
<p><strong>Abstract</strong> (our paper)</p> <p>Why does Smith follow Johnson on Twitter? In most cases, the reason why users follow other users is unavailable. In this work, we answer this question by proposing TagF, which analyzes the who-follows-whom network (matrix) and the who-tags-whom network (tensor) simultaneously. Concretely, our method decomposes a coupled tensor constructed from these matrix and tensor. The experimental results on million-scale Twitter networks show that TagF uncovers different, but explainable reasons why users follow other users.</p> <p><strong>Data</strong></p> <p>coupled_tensor:<br> The first column is the source user id (from user id), the second column is the destination user id (to user id), and the third column is the tag id.</p> <p>users.id:<br> The first column is the user id for coupled_tensor, and the second column is the user id on Twitter.</p> <p>tags.id:<br> The first column is the tag id for coupled_tensor, and the second column is the tag (<em>i.e.</em> slug or list name) on Twitter. On the tags, ###follow### and ###friend### are special tags expressing follower and following.</p> <p><strong>Publication</strong></p> <p>This dataset was created for our study. If you make use of this dataset, please cite:<br> Yuto Yamaguchi, Mitsuo Yoshida, Christos Faloutsos, Hiroyuki Kitagawa. Why Do You Follow Him? Multilinear Analysis on Twitter. <em>Proceedings of the 24th International Conference on World Wide Web (WWW '15 Companion)</em>. pp.137-138, 2015.<br> http://doi.org/10.1145/2740908.2742715</p> <p><strong>Code</strong></p> <p>Our code outputting experiment results made available at:<br> https://github.com/yamaguchiyuto/tagf</p> <p><strong>Note</strong></p> <p>If you would like to use larger dataset, the dataset on 1 million seed users made available at:<br> http://dx.doi.org/10.5281/zenodo.16267<br> (The dataset on 0.1 million seed users is not subset of the dataset on 1 million seed users.)</p>
SAA2017 TAGS Tweet Archive
<p>An open archive of Tweets from SAA2017, the Society for American Archaeology's 82nd Annual Meeting, Vancouver, BC, Canada.</p>
A part-of-speech (POS) tagged corpus of Classical Tibetan
<p>This part-of-speech (POS) tagged corpus of Classical Tibetan was prepared in the course of the research project 'Tibetan in Digital Communication' (2012-2015) hosted at SOAS, University of London and funded by the UK's Arts and Humanities Research Council (grant code: AH/J00152X/1). For a description of the tag set see Garrett et al. 2014. and Garrett et al. 2015. This corpus includes the <em>Mdzaṅs blun</em> (9th century, canonical), the <em>Bu ston chos ḥbyuṅ</em> (13th century, ecclesiastical history), the <em>Mi la ras paḥi rnam thar</em> and <em>Mar paḥi rnam thar</em> (15th century, biography).</p>
Dataset supplementing the publication Einhäuser, W., Thomassen, S., & Bendixen, A. (2017). Using binocular rivalry to tag foreground sounds: towards an objective visual measure for auditory multistability. Journal of Vision, 17:34, 1-19.
<p>These files supplement the publication Einhäuser, W., Thomassen, S., & Bendixen, A. (2017). Using binocular rivalry to tag foreground sounds: towards an objective visual measure for auditory multistability. Journal of Vision, 17:34, 1-19. The data are free for scientific use, provided this reference is appropriately cited.</p> <p>exp1_data.mat contains all the data of experiment 1 as cell arrays of size 8x16x8 (subject x block x trial) or 8x16 (subject x block). Specifically:<br> xEye: the horizontal eye position in raw (pixel coordinates)<br> gain: the OKN slow phase gain computed from the xEye data as described in the paper; in audio-visual blocks the sign is chosen such that positive gain corresponds to the direction of the grating associated with the low tone; in unambiguous visual blocks (1,16) positive sign corresponds to the direction of the grating.<br> ixLow, ixHigh, ixNone, ixBoth: indices for xEye and gain of the same subject and block for which the button corresponding to the low tone, the high tone, both buttons or no button was pressed.</p> <p>exp2_data.mat and exp3_data.mat contain the data of experiment 2 and experiment 3, respectively, and are organized analogously to exp1_data.mat.</p> <p>figure3.m through figure6.m use these data to plot the respective paper figures to exemplify usage of the data.</p> <p> </p>
Geo-tagged Tweets in Paris during Nov 2015
<p><strong>Abstract</strong></p> <p>The data sets released here have been used in our study on quantitatively evaluating the impact of disasters in the city. The study of disaster events and their impact in the urban space has been traditionally conducted through manual collections and analysis of surveys, questionnaires and authority documents. While there have been increasingly rich troves of human behavioral data related to the events of interest, the ability to obtain hindsight following a disaster event has not been scaled up. In this study, we propose a novel approach for analyzing events called PairFac. PairFac utilizes discriminant tensor analysis to automatically discover the impact of a major event from rich human behavioral data. Our method aims to (i) uncover the persistent patterns across multiple interrelated aspects of urban behavior (e.g., when, where and what citizens do in a city) and at the same time (ii) identify the salient changes following a potentially impactful event. We show the effectiveness of PairFac in comparison with previous methods through extensive experiments. We also demonstrate the advantages of our approach through case studies with real-world traffic sensor data and social media streams surrounding the 2015 terrorist attacks in Paris. Our work has both methodological contributions in studying the impact of an external stimulus on a system as well as practical implications in the area of disaster event analysis and assessment.</p> <p><strong>Dataset</strong></p> <p>There are two datasets used in this study, traffic sensor dataset and social media dataset. </p> <p>Traffic Sensor dataset was collected from open data Paris. This dataset includes all the hourly data for the flow and the occupancy rate assembled by the permanent traffic sensors installed on the Paris City network (urban network and peripheral boulevard). For interested readers, please direct to the link as: https://opendata.paris.fr/explore/dataset/comptages-routiers-permanents/</p> <p>Social media dataset was collected from Twitter API. The dataset contains geo-tagged tweets from Paris collected through Twitter API between the period of Oct 16th, 2015 and Nov 20, 2015. 75,982 geo-located tweets were extracted during the period covered.</p> <p>Duration: 2015-10-16 to 2015-11-20.</p> <p>Total number of tweets: 75,982</p> <p><strong>Publication</strong></p> <p>If you make use of this data set, please kindly cite:</p> <p>Xidao Wen, Yu-Ru Lin, and Konstantinos Pelechrinis. 2016. PairFac: Event Analytics through Discriminant Tensor Factorization. In Proceedings of the 25th ACM International on Conference on Information and Knowledge Management (CIKM '16). ACM, New York, NY, USA, 519-528. DOI: https://doi.org/10.1145/2983323.2983837</p> <p> </p>
Gujarati Movie Reviews with Tagged Sentiments (Positive/ Negative/ Neutral)
<p>This dataset encompasses around 500 entries of movie review descriptions written in the Gujarati language. Each entry is paired with a sentiment classification tag. The first column contains the actual movie review descriptions, while the second column contains sentiment tags with the following meanings:</p><ul><li>"0" denotes that the corresponding review expresses a negative sentiment.</li><li>"1" signifies a neutral sentiment.</li><li>"2" conveys a positive sentiment.</li></ul><p>To sum it up succinctly, this dataset provides a valuable collection of <strong>Gujarati movie reviews</strong>, each thoughtfully categorized as either <strong>negative</strong>, <strong>neutral</strong>, or <strong>positive</strong> in tone, offering rich insights for sentiment analysis tasks and research.</p><p>The dataset is manually tagged by native speakers with more than 20 years of experience in using the language.</p>
tags-ask-ubuntu
<h3><strong>Overview</strong></h3><p>This is a temporal hypergraph dataset, which here means a sequence of timestamped hyperedges where each hyperedge is a set of nodes. In this dataset, nodes are tags, and hyperedges are the sets of tags applied to questions on askubuntu.com. The timestamps are in ISO8601 format and are normalized to start at 0. This dataset is derived from tags on Ask Ubuntu posts.</p><h4><strong>Statistics</strong></h4><p>Some basic statistics of this dataset are:</p><ul><li>number of nodes: 3,029</li><li>number of timestamped hyperedges: 271,233</li><li>distribution of the connected components:</li></ul><p>Component Size, Number </p><ul><li>3021, 1</li><li>1, 8</li></ul><h4><strong>Source of original data</strong></h4><p>Sources:</p><ul><li><a href="https://www.cs.cornell.edu/~arb/data/tags-ask-ubuntu/">tags-ask-ubuntu dataset</a></li><li><a href="https://archive.org/details/stackexchange">StackExchange</a></li></ul><h4>References</h4><p>If you use this dataset, please cite these references:</p><ul><li><a href="https://doi.org/10.1073/pnas.1800683115">Simplicial closure and higher-order link prediction</a>. Austin R. Benson, Rediet Abebe, Michael T. Schaub, Ali Jadbabaie, and Jon Kleinberg. Proceedings of the National Academy of Sciences (PNAS), 2018.</li></ul>
tags-math-sx
<h3><strong>Overview</strong></h3><p>This is a temporal higher-order network dataset, which here means a sequence of timestamped hyperedges where each hyperedge is a set of nodes. In this dataset, nodes are tags, and hyperedges are the sets of tags applied to questions on math.stackexchange.com.</p><p>Each hyperedge corresponds to all of the tags used in a post, and each node in a hyperedge corresponds to a tag. The timestamps are normalized so that the earliest tag starts at 0 and are in millisecond resolution.</p><h4><strong>Statistics</strong></h4><p>Some basic statistics of this dataset are:</p><ul><li>number of nodes: 1,629</li><li>number of timestamped hyperedges: 822,059</li><li>number of unique hyperedges: 174,933</li></ul><h4><strong>Source of original data</strong></h4><p>Source: <a href="https://www.cs.cornell.edu/~arb/data/tags-math-sx/">tags-math-sx dataset</a></p><h4><strong>References</strong></h4><p>If you use this data, please cite the following paper:</p><ul><li><a href="https://doi.org/10.1073/pnas.1800683115">Simplicial closure and higher-order link prediction</a>. Austin R. Benson, Rediet Abebe, Michael T. Schaub, Ali Jadbabaie, and Jon Kleinberg. Proceedings of the National Academy of Sciences (PNAS), 2018.</li></ul>
Foraging behavior of tagged rock ants (Temnothorax rugatulus)
<p><span>Technological advances continue to push the boundaries of scientific inquiry in animal behavior. One such development is the emergence of automated tracking systems, which enable the collection of high-resolution spatio-temporal information for animals. Although tag-based tracking systems provide valuable insights into animal movement and collective behavior, the attachment of devices can have detrimental effects in some cases. Here, we investigated the effects of recently developed miniature tracking tags using the rock ant, </span><em><span>Temnothorax rugatulus</span></em><span>, as a model system. To do so, we compared the foraging activities of tagged ants and untagged ants (who lost their tags) within initially fully-tagged colonies. Additionally, we compared the foraging activities of these initially fully-tagged colonies with those of no-tag control colonies (no one was tagged). We found that tags did not significantly reduce individual activity, with tagged ants visiting the food source as frequently as untagged ants within initially fully-tagged colonies. However, our analysis revealed a marked difference in recruitment behavior—tagged ants were less likely to participate in tandem runs than untagged ants. Furthermore, the number of tandem runs was higher for the no-tag control colonies than the initially fully-tagged colonies, in which 69–95% of colony members had tags. Our data suggest, for the first time, that tracking tags can negatively impact ant behavior. Although tracking devices are powerful tools for understanding complex behavioral patterns, it is crucial to carefully consider their potential impact on animal behavior to ensure accurate conclusions.</span></p>
Open-Set Tagging Dataset (OST)
<p>Open-set Tagging (OST) is a synthetic dataset of 1s clips used to evaluate source-centric representation learning models in the paper <a href="https://ieeexplore.ieee.org/document/10890242">Compositional Audio Representation Learning</a>.</p> <p>Due to the size of the dataset, we only share the source files, and provide the scripts to generate the dataset are available <a href="https://github.com/sripathisridhar/moads">here.</a><br><br>The dataset generation process is as follows:<br>1. From single-source FSD50K audio files, we generate a dataset of 10s soundscapes called Open-set Soundscapes (OSS) using <a href="https://github.com/justinsalamon/scaper">Scaper</a>.</p> <p>2. We then center a 1s window around the center of each sound event in the 10s soundscapes to generate Open-set Tagging (OST), which contains ~500k clips. </p> <p>If you are not going to use OSS, you can choose to synthesize it without audio-- this will synthesize only the <a href="https://github.com/marl/jams">JAMS</a> annotation files needed for the 1s clips. Using the OSS JAMS files, OST clips can be generated deterministically.</p> <p>There are five dataset variants (~17GB each), each with a different random assignment of classes to the known and unknown class categories. For further details, refer to our previous paper <a href="https://dcase.community/documents/workshop2023/proceedings/DCASE2023Workshop_Sridhar_11.pdf">Multi-label open-set audio classification</a>. In <a href="https://ieeexplore.ieee.org/document/10890242">this work</a>, OST dataset variant 1 is referred to as OST for simplicity. <br><br>We also introduce a tiny version of the dataset called OST-Tiny, which contains ~20k clips and only 10 known classes. This is convenient for faster prototyping and to evaluate models in a more challenging open-set classification scenario.</p> <p> </p>
Animal lifestyle affects acceptable mass limits for attached tags
<p></p><p> Animal-attached devices have transformed our understanding of vertebrate ecology. To minimize any associated harm, researchers have long advocated that tag masses should not exceed 3% of carrier body mass. However, this ignores tag forces resulting from animal movement. Using data from collar-attached accelerometers on 10 diverse free-ranging terrestrial species from koalas to cheetahs, we detail a tag-based acceleration method to clarify acceptable tag mass limits. We quantify animal athleticism in terms of fractions of animal movement time devoted to different collar-recorded accelerations and convert those accelerations to forces (acceleration × tag mass) to allow derivation of any defined force limits for specified fractions of any animal's active time. Specifying that tags should exert forces that are less than 3% of the gravitational force exerted on the animal's body for 95% of the time led to corrected tag masses that should constitute between 1.6% and 2.98% of carrier mass, depending on athleticism. Strikingly, in four carnivore species encompassing two orders of magnitude in mass ( ca 2–200 kg), forces exerted by '3%' tags were equivalent to 4–19% of carrier body mass during moving, with a maximum of 54% in a hunting cheetah. This fundamentally changes how acceptable tag mass limits should be determined by ethics bodies, irrespective of the force and time limits specified. </p><p></p>
Reference data and analysis software for "Four-color single-molecule imaging with engineered tags resolves the molecular architecture of signaling complexes in the plasma membrane"
<p>Reference data set for the single molecule co-tracking analysis presented in "Four-color single-molecule imaging with engineered tags resolves the molecular architecture of signaling complexes in the plasma membrane". Corresponding author for further inquiries:</p> <p>Prof. Dr. Jacob Piehler</p> <p>University of Osnabrück, Department of Biology/Chemistry, Division of Biophysics, Barbarastr. 11, 49076 Osnabrück, Germany</p> <p>https://www.biophysik.uni-osnabrueck.de/</p>
Ca2+ activity maps of astrocytes tagged by axo-astrocytic AAV transfer
<p>Astrocytes exhibit localized Ca<sup>2+</sup> microdomain (MD) activity thought to be actively involved in information processing in the brain. However, functional organization of Ca<sup>2+</sup> MDs in space and time in relationship to behavior and neuronal activity is poorly understood. Here, we first show that Adeno-Associated Virus (AAV) particles transfer anterogradely from axons to astrocytes. Then we use this axo-astrocytic AAV transfer to express genetically encoded Ca<sup>2+</sup> indicators at high contrast circuit-specifically. In combination with two-photon microscopy and unbiased, event-based analysis we investigated cortical astrocytes embedded in the vibrissal thalamocortical circuit. We found a wide range of Ca<sup>2+</sup> MD signals, some of which were ultrafast (≤300 ms). Frequency and size of signals were extensively increased by locomotion but only subtly with sensory stimulation. The overlay of these signals resulted in behavior dependent maps with characteristic Ca<sup>2+</sup> activity hotspots, maybe representing memory engrams. These functional subdomains are stable over days, suggesting subcellular specialization.</p>
Sustainable Smart Tags with Two-Step Verification for Anticounterfeiting Triggered by the Photothermal Response of Upconverting Nanoparticles
<p>Dataset accompanying figures published in the publication DOI: https://doi.org/10.5281/zenodo.6245930</p>
Data from: Integrating tracking and resight data enables unbiased inferences about migratory connectivity and winter range survival from archival tags
<p>Archival geolocators have transformed the study of small, migratory organisms but analysis of data from these devices requires bias correction because tags are only recovered from individuals that survive and are re-captured at their tagging location. Data and code provided in this repository can be used to replicate the simulation and Painted Bunting case study results presented by Rushing et al. (2021) showing that integrating geolocator recovery data and mark–resight data enables unbiased estimates of both migratory connectivity between breeding and nonbreeding populations and region-specific survival probabilities for wintering locations.</p>
SROADEX: Dataset for binary recognition and semantic segmentation of road surface areas from high resolution Aerial Orthoimages Covering Approximately 8,650 km2 of the Spanish Territory Tagged with Road Information
<p>The data have been generated using scripts developed in Python using Open Source libraries (GDAL/OGR and MapScript) for rasterization of vector cartography representing the axes of the different types of roads (urban, interurban and rural). This cartography has been obtained from different Spanish official sources (National Geographic Institute and autonomic cartographic agencies) that we have revised and edited in a meticulous and systematic way to verify that the roads are represented on the cartography according to the orthoimages, available on January 1, 2021 in the download center of the National Center of Geographic Information (CNIG), on 16 rectangular areas (28,5 km * 18,5 km) of the Spanish territory (insular and peninsular).</p> <p>The dataset consists of 777599 images in png format of 256x256 pixels, organized in folders for the different trainings, separating those corresponding to training, testing and validation.</p> <p>The structure of the data is as follows:<br> 1-Road-Ortho and 1-Road-Mask contain the images and ground true for training the semantic segmentation networks.<br> 1-Road-Ortho and 2-NoRoad-Ortho contain aerial images containing or not containing vials, for the training of binary tessellation networks identifying tessellations with vials.<br> Moreover, in each folder the structure is the same: train, test, validation containing 90%, 5% and 5% of the total images and masks of each type.</p> <p>1-Road-Ortho</p> <p> |----Train</p> <p> |----Test</p> <p> -----Validation</p> <p>1-Road-Mask</p> <p> |----Train</p> <p> |----Test</p> <p> -----Validation</p> <p>2-NoRoad-Ortho</p> <p> |----Train</p> <p> |----Test</p> <p> -----Validation</p> <p> </p>
Data from: Pop‐off data storage tags reveal niche partitioning between native and non‐native predators in a novel ecosystem
1. Niche partitioning might be predicted to be particularly dynamic in 'novel ecosystems' characterized by human-altered environmental conditions and biological invasions. Restoration efforts for native species in such systems can be informed by detailed characterization of niche partitioning. 2. In Lake Ontario, fishery management agencies have been engaged in a long-term struggle to restore native top predators including lake trout (Salvelinus namaycush). Meanwhile, management agencies continue to stock non-native species like Chinook salmon (Oncorhynchus tshawytscha) into the lake to support a recreational fishery and to help control the abundance of a non-native forage fish, the alewife (Alosa pseudoharengus). 3. We used pop-off data storage tags to study fine scale (9.1M lines of data from 22 animals) behaviour and habitat use by lake trout (native) and Chinook salmon (non-native) in Lake Ontario in terms of depth and temperature, recorded at ≤70 s intervals for periods of up to 12 months. 4. Chinook salmon occupied warmer and shallower waters during summer than did lake trout, and their niche breadth was wider. They achieved greater niche breadth in part because they were much more active vertically, cumulatively traveling 103±1 m hour-1 during summer (model-estimated median), whereas most lake trout were relatively inactive vertically (7±1 m hour-1). In each of our analyses, there was more inter-individual variation among lake trout than among Chinook salmon, driven by some lake trout that spent considerable time making forays into warmer, shallower waters. 5. Synthesis and applications. Our results illustrate the different foraging tactics used by two species in the Great Lakes and reflect their distinct life histories. Vertical and thermal niche partitioning between Chinook salmon and lake trout helps to explain how these species can co-exist in a multi-species fishery even while having substantial overlap in diet. The diversity of behaviours exhibited here by native lake trout have likely helped them persist during dramatic changes to the forage base in recent decades; that flexibility could help underlie their long-term prospects for restoration during future changes to the ecosystem.
Mosquito Tagging Using DNA-Barcoded Nanoporous Protein Microcrystals
<p>Contains raw data for the publication titled 'Mosquito Tagging Using DNA Barcoded Nanoporous Protein Microcrystals'.</p>
MSDI: a geo-tagged drone imagery for absolute visual localization
<p>The MSDI (Manchester Surface Drone Imagery) is a geo-imagery registration dataset. <br> The dataset consists of <br> a. 446 downward-facing drone images.<br> b. 89 forward-facing(45-degree) drone images.<br> c. 64 forward-facing (0-degree) drone images. <br> d. parameter matrix of drone camera and transformation matrix.<br> e. checkerboard images for camera calibration.</p> <p>This dataset is collected by Mochuan Zhan for his MSC project: Registration of UAV Imagery to Aerial and Satellite Imagery<br> in the University of Manchester (2021/9 - 2022/9) which is supervised by Dr.Terence Patrick Morley. This project aims at <br> developing a system that could perform efficient UAV visual localization through image registration based on local feature <br> detectors and the technique of high-throughput computing.</p> <p>Notice:<br> The corresponding satellite image from Google Map and Bing Map could be obtained by my program, Link: </p> <pre>https://doi.org/10.5281/zenodo.6977652</pre> <p>A Ground Control Point (GCP) selector is provided for user to select GCP and create file with small effort.<br> By registrating corresponding images, users could evaluate the performance of their registration techniques.</p> <p>The Imagery contains images of 8 Areas in Manchester:<br> - Manchester Aquatics Center 80<br> - Manchester ASDA 76<br> - Manchester Bussiness School 37<br> - Manchester Energy Center 71<br> - Manchester Holy Name Church 58<br> - Manchester Hulme Park 47<br> - Manchester Hulme Part(0-degree) 64<br> - Manchester Hulme Part(45-degree) 89<br> - Manchester Metropolitan university 29 <br> - Manchester Museum 48</p> <p>Device Information:<br> - Drone brand: Parrot<br> - Drone model: Parrot Anafi</p> <p>Software Information:<br> - Pix4DCapture<br> - FreeFlight6</p> <p>Flight parameters:<br> - Height 100m<br> - Speed 5m/s<br> - overlap low</p> <p><br> </p>
Tagging and tracking information for radiotagged Chinook Salmon in the Copper River, Alaska 2021
<p>The first worksheet (2021 Raw Data) consists of each radiotagged fish and its relevant information including date of capture, the frequency and code of the transmitter, length (MEF) and age. Subsequent columns are Julian dates when they passed fixed tracking stations. The final 4 columns are fate columns. The last column is a general description of the general fate of each fish.</p> <p> </p> <p>The final worksheet (2021 summary) summarizes fates of all fish by tagging date. This is the primary input file for the Program R which has a code written to do the data analyses for this study.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.