Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,025
datasets available to search
ShareScore release 0.9.0
Dataset results
1,025 results for “Vision”
MCR LTER: Coral Reef: Computer Vision: Multi-annotator Comparison of Coral Photo Quadrat Analysis
This repository contains the Moorea portion of a larger data package published in conjuncture with: "Towards automated annotation of benthic survey images: variability of human experts and operational modes of automation", Beijbom et al. PLOS One, 2015. The rest of the data package is hosted at the Dryad data repository (doi:10.5061/dryad.m5pr3). The larger data package is an aggregate dataset from four Pacific coral reef monitoring projects in: Moorea (French Polynesia), the northern Line Islands, Nanwan Bay (Taiwan) and Heron Reef (Australia). It contains 5090 coral reef survey images, and 251,988 random-point annotations by coral ecology experts. The point-annotations indicate the dominant benthic substrate at 10 to 200 random point locations per image, using a label-set of 20 categories. In addition, 200 images from each location have been cross-annotated by 6 experts, for a total of 7 sets of annotations for each image. This set of cross-annotations can be used to contextualize the performance of automated annotation methods for coral reef ecology. The full data package can also be used by computer-vision and machine learning researchers to develop object classification, image segmentation, and domain transfer learning methods. These data contain a subset of the raw data from which dataset knb-lter-mcr.4 is derived.
Dreams4Cars Vision video
<p><em>This video presents the Dreams4Cars Agent sensorimotor architecture (note the video has a communication purpose, please refer to the paper for science).</em></p> <p><em>See paper: </em>M. Da Lio, R. Donà, G. P. Rosati Papini and K. Gurney, "Agent Architecture for Adaptive Behaviours in Autonomous Driving," IEEE Access, to be published. </p>
Search-and-Rescue From Drones With Computer Vision
<p>Unmanned aerial vehicles (UAVs), most commonly known as drones, are increasingly used as a technological support tool for search-and-rescue (SAR) operations (and post-disaster area explorations as well). UAVs equipped with high-resolution cameras and embedded, yet powerful GPUs, in fact, can provide an effective and efficient aid to emergency rescue operations, mainly because locating victims, which may be unconscious or injured, as much fast as possible, is crucial to improve their chance of survival. In particular, the use of drones that are able to automatically detect people in the scenes can increase detection rate, while reducing rescue time. In this repository, we provide a new dataset specifically conceived for SAR operations from drones with computer vision. As it is small-sized, the dataset is currently intended for testing and evaluation purposes only. The main aim of the repository is to encourage contributions on this intriguing topic. In particular, any contribution to make the dataset bigger is welcome.</p>
Aalborg Energy Vision scenarios
<p>EnergyPLAN scenario files</p>
Video footages captured by Matrix Vision camera and Gopro camera
<p>The two videos briefly demonstrate the video quality captured by the matrix vision camera (mb.mp4) and the GoPro camera (GH010787_Trim.mp4).</p> <p>As can be seen, severe motion blurs showed in the video captured by the matrix vision camera.</p>
Data supplementing the article "Einhäuser, W., & Nuthmann, A. (2016). Salient in space, salient in time: Fixation probability predicts fixation duration during natural scene viewing. Journal of Vision, 16(11):13, 1-17, doi:10.1167/16.11.13."
<p>These data supplement the article Einhäuser, W., & Nuthmann, A. (2016). Salient in space, salient in time: Fixation probability predicts fixation duration during natural scene viewing. Journal of Vision, 16(11):13, 1-17, doi:10.1167/16.11.13.</p> <p>The data can be used freely for academic purposes, provided the aforementioned reference is appropriately cited.</p> <p>The following files are available for experiment 2 of the article:</p> <p>allData.mat</p> <p>Includes the datamatrix allData with the following columns:</p> <p>1) Line used for analysis in the article (0 - no, 1-yes).<br> Possible reasons for exclusion:<br> a. fixation duration smaller than 50ms or larger 1000ms<br> b. fixation adjacent to a blink (preceding or following)<br> c. fixation outside the image</p> <p>2) ID of observer (1-24)</p> <p>3) ID of condition (1:grayscale, 2: reduced luminance, 3: reduced contrast, 4: equalized luminance, 5: equalized contrast, 6: phasenoise)</p> <p>4) ID of image (48 unique numbers between 1 and 135)</p> <p>5) horizontal eye position</p> <p>6) vertical eye position</p> <p>7) fixation duration in ms</p> <p>8) value of empirical map generated from search condition of experiment 1 at fixated location</p> <p>9) value of empirical map generated from preference condition of experiment 1 at fixated location</p> <p>10) value of empirical map generated from memorization condition of experiment 1 at fixated location</p> <p>11) value of empirical map generated from joining memorization and preference condition of experiment 1 at fixated location</p> <p>12) value of empirical map generated from condition 1 at fixated location</p> <p>13) value of empirical map generated from condition 2 at fixated location</p> <p>14) value of empirical map generated from condition 3 at fixated location</p> <p>15) value of empirical map generated from condition 4 at fixated location</p> <p>16) value of empirical map generated from condition 5 at fixated location</p> <p>17) value of empirical map generated from condition 6 at fixated location</p> <p>18) value of empirical map generated from condition 1 at fixated location leaving out the current observer</p> <p>19) value of empirical map generated from condition 2 at fixated location leaving out the current observer</p> <p>20) value of empirical map generated from condition 3 at fixated location leaving out the current observer</p> <p>21) value of empirical map generated from condition 4 at fixated location leaving out the current observer</p> <p>22) value of empirical map generated from condition 5 at fixated location leaving out the current observer</p> <p>23) value of empirical map generated from condition 6 at fixated location leaving out the current observer</p> <p>24) luminance at fixation</p> <p>25) luminance contrast at fixation</p> <p>26) edge density at fixation</p> <p>27) eccentricity of fixation</p> <p> </p> <p>usedData.Rdata</p> <p>- for all lines that are used for analysis (allData(:,1)==1) a field in an R dataframe is created, which contains the following fields (for details, see description of matlab file above):</p> <p>obsNum: the ID of the observer (1-24)</p> <p>condNum: the ID of the condition (1-6)</p> <p>imgNum: the ID of the image (48 unique numbers between 1 and 135)</p> <p>fixDur: fixation duration</p> <p>LUM, LCG, ED, ECC: luminance, contrast, edge density and eccentricity at fixation</p> <p>empMapFromSearch, empMapFromPref, empMapFromMem, empMapFromJoint: values of empirical maps generated from data of experiment 1 (search, preference, memorization task as well as combination of the latter two) at fixation</p> <p>empMapFromC1 through empMapFromC6: value of empirical map generated from condition 1 through 6 at fixated location</p> <p>empMapFromC1loo through empMapFromC6loo - value of empirical map generated from condition 1 through 6 at fixated location leaving out the current observer</p> <p>x,y - coordinates of fixation</p> <p> </p> <p>modelsFigure7.R - computes all models for figure 7 of the aforementioned article (Note: depending on your system, this can take substantial time; depending on the version of the lme-package results may deviate slightly from those given in the paper)</p> <p>modelsFigure8.R - computes all models for figure 8 of the aforementioned article (Note: depending on your system, this can take substantial time; depending on the version of the lme-package results may deviate slightly from those given in the paper)</p> <p> </p>
Data supplementing the article Schomaker, J., Walper, D., Wittmann, B.C., & Einhäuser, W. (2017). Attention in natural scenes: Affective-motivational factors guide gaze independently of visual salience. Vision Research, 133, 161-175.
<p>These data supplement the article Schomaker, J., Walper, D., Wittmann, B.C., & Einhäuser, W. (2017). Attention in natural scenes: Affective-motivational factors guide gaze independently of visual salience. Vision Research, 133, 161-175.</p> <p>Use is free for academic purposes, provided the aforementioned article is appropriately cited.</p> <p>The directory contains the following files</p> <p>stimuli.tar.gz - stimuli used in this study; note that this is based on the MONS database, but some deviations from the final version of the database do exist.</p> <p>ratings.mat contains the variables<br> arousal - mean arousal rating<br> valence - mean valence rating<br> valence2 - squared mean valence rating (after subtracting midpoint)<br> motivationalValue - mean motivation rating<br> motivaionalValue2 - squared mean motivation rating (after subtracting midpoint)</p> <p>All variables are 104x3, where the first dimension is the stimulus number, and the second dimension the motivation ground truth (aversive, neutral, appetitive)</p> <p><br> Experiment 1</p> <p>fixationsExperiment1.mat contains the variables fixationX, fixationY, fixationDuration, fixaitonOnset, fixationInitial, which contain for each fixation horizontal and vertical coordinate, the duration, the time of the onset relative to the trial onset and whether it is the initial fixation. All variables have dimensions 16x104x3x50, where the first dimension is the observer, the second the scene, the third the condition and the forth a counter of fixations. Whenever there are less than 50 fixations the remainder are filled with NaN.</p> <p><br> boundingBoxesExperiment1.mat contains for each critical object the bounding box coordinates x,y of upper left corner and width and height as variables boundingBoxX, boundingBoxY, boundingBoxW, boundingBoxH respectively. Note that this is relative to the eyetracker coordinates of experiment 1 (full display 1024x768, presentation in the center) and will therefore not match the coordinates of the images in the archive or the bounding box coordinates of experiment 2. Dimensions are 104x3, the dimensions representing scene number and condition, respectively.</p> <p><br> figure2.m uses these data to computes figure 2 of the article from these data</p> <p><br> dataForExperiment1.Rdata contains the data frame data, which contains for each fixation the values of the predictors used in the model of table 1. This is computed from the matlab data listed above in addition to the peak values of the AWS salience in the object.</p> <p><br> table1.R computes and prints the models for table 1</p> <p> </p> <p>Experiment 2</p> <p>fixationsExperiment2.mat contains fixation data for experiment 2. Variable names as in experiment 1. Dimensions are 18x99x3x3x50, where the first dimension is the observer, the second the image number, the third the visual condition, the third the motivational condition and the fifth the fixation count. Since only one visual condition was shown to each observer per motivational condition, there is an additional variable 'hasData', which is 1 if the image was presented to the observer in this condition and 0 otherwise. Since fixations can be outside the image and will therefore be excluded, there is also an additional variable fixationNumber to keep a correct count of the fixation number in the trial.</p> <p>boundingBoxesExperiment2.mat contains bounding box data for experiment 2 in image (and fixation) coordinates. Notation as for experiment 1, but coordinates refer to image and eyetracking coordinates used for experiment 2 and therefore can differ occasionally.</p> <p><br> figure3and4.m generates figures 3 and 4 of the article from these data files.</p> <p>dataForExperiment2.Rdata contains the data frame data, which contains for each fixation the values of the predictors used in the model of tables 2 amd 3. This is computed from the matlab data listed above in addition to the peak values of the AWS salience in the object. The fields imgMot and imgVis contain the motivational ground truth and the salience manipulation, respectively.</p> <p>table2.R uses the Rdata file to compute the models for table 2 of the article and print summary results</p> <p>table3.R uses the Rdata file to compute the models for table 3 of the article and print summary results. Note that the computation can take substantial time; results might deviate slightly depending on the exact version of R and its libraries used.</p> <p> </p>
Dataset supplementing Stoll, J., Thrun, M., Nuthmann, A., & Einhäuser, W. (2015). Overt attention in natural scenes: Objects dominate features. Vision Research, 107, 36-48. doi: 10.1016/j.visres.2014.11.006
<p>These data supplement the publication</p> <p>Stoll, J., Thrun, M., Nuthmann, A., & Einhäuser, W. (2015). Overt attention in natural scenes: Objects dominate features. Vision Research, 107, 36-48. doi: 10.1016/j.visres.2014.11.006</p> <p>and be used freely for scientific purposes provided the aforementioned paper is appropriately cited.</p> <p>Note that the image files cannot be provided on this site due to copyright restrictions.</p> <p>The dataset contains the following files:</p> <p>maps_01.mat - maps_72.mat:</p> <p>For each image the 6 maps used in the paper are contained, the maps of experiment 1 are labelled as in the paper (AWS, OOM, nOOM, PVL,UNI), AWS2 is the AWS map for the modified stimuli of experiments 2 and 3.</p> <p>exp?_fixations.mat contains all fixations of the respective experiment.</p> <p>For experiment 1, there are the variables xFix, yFix, durFix, which contain the x position, the y condition, and the fixation duration of each fixation. Dimensions are images x subjects x fixation number, where the first fixation is the 0th (initial) fixation. The variable condition (image x subject) contains the condition in which the respective image was shown to the subject. For the main analysis only the "0" condition was used, refer to the paper's appendix for the other conditions.</p> <p>For experiment 2 and 3, variables are called xFixByImage, yFixByImage, dFixByImage and the dimensions are subject x image x fixation number. In addition tFixByImage contains the start of the fixation relative to trial onset (negative for the 0th fixation).<br> In both cases, empty entries are filled with nans.</p> <p><br> computeROC.m is a helper function called by other functions.</p> <p><br> figure1.m through figure7.m reproduce the figures from the paper to exemplify data usage.</p> <p> </p>
Dataset supplementing the publication Einhäuser, W., Thomassen, S., & Bendixen, A. (2017). Using binocular rivalry to tag foreground sounds: towards an objective visual measure for auditory multistability. Journal of Vision, 17:34, 1-19.
<p>These files supplement the publication Einhäuser, W., Thomassen, S., & Bendixen, A. (2017). Using binocular rivalry to tag foreground sounds: towards an objective visual measure for auditory multistability. Journal of Vision, 17:34, 1-19. The data are free for scientific use, provided this reference is appropriately cited.</p> <p>exp1_data.mat contains all the data of experiment 1 as cell arrays of size 8x16x8 (subject x block x trial) or 8x16 (subject x block). Specifically:<br> xEye: the horizontal eye position in raw (pixel coordinates)<br> gain: the OKN slow phase gain computed from the xEye data as described in the paper; in audio-visual blocks the sign is chosen such that positive gain corresponds to the direction of the grating associated with the low tone; in unambiguous visual blocks (1,16) positive sign corresponds to the direction of the grating.<br> ixLow, ixHigh, ixNone, ixBoth: indices for xEye and gain of the same subject and block for which the button corresponding to the low tone, the high tone, both buttons or no button was pressed.</p> <p>exp2_data.mat and exp3_data.mat contain the data of experiment 2 and experiment 3, respectively, and are organized analogously to exp1_data.mat.</p> <p>figure3.m through figure6.m use these data to plot the respective paper figures to exemplify usage of the data.</p> <p> </p>
Dataset supplementing "Marx, S., & Einhäuser, W. (2015). Reward modulates perception in binocular rivalry. Journal of Vision, 15(1):11, 1–13, http://www.journalofvision.org/content/15/1/11, doi:10.1167/15.1.11."
<p>These data supplement the publication</p> <p>Marx, S., & Einhäuser, W. (2015). Reward modulates perception in binocular rivalry. Journal of Vision, 15(1):11, 1–13, http://www.journalofvision.org/content/15/1/11, doi:10.1167/15.1.11.</p> <p>and be used freely for scientific purposes provided the aforementioned paper is appropriately cited.</p> <p>exp1_data.mat contains data of experiment 1</p> <p>exp2_data.mat contains data of experiment 2</p> <p>figure2_3.m and figure4_5.m exemplify usage of the data and reproduce the figures 2-5 of the aforementioned article.</p>
Cutting the Frame: An In-Depth Look at the Hitchcock Computer Vision Dataset
<p>The Hitchcock Computer Vision Dataset is a comprehensive collection of annotated frames from fifteen Alfred Hitchcock films, created for the purpose of advancing research in film studies and digital humanities. The dataset leverages the power of Google's Vision API to provide detailed annotations such as object detection, facial recognition, web-entity analysis, explicit content filtering, and more. The dataset consists of a CSV file containing about 105,000 frames, uniformly extracted from the films using a time-based approach. Each frame is accompanied by metadata, including the film name, timestamp, and release year, allowing researchers to explore the evolution of Hitchcock's cinematic techniques and recurring themes.</p>
Red vision in animals is broadly associated with lighting environment but not types of visual task
<p>Red sensitivity is the exception rather than the norm in most animal groups. Among species with a long wavelength sensitive (LWS) photoreceptor, peak wavelength sensitivity (λ<sub>max</sub>) varies substantially and it is unclear whether this variation can be explained by visual tuning to the light environment or to visual tasks such as signalling or foraging. Here, we examine long wavelength sensitivity across a broad range of taxa showing diversity in LWS photoreceptor λ<sub>max</sub>: insects, crustaceans, arachnids, amphibians, reptiles, fish, sharks and rays. We identified 161 species with a LWS photoreceptor (λ<sub>max</sub> ≥ 550 nm). We found evidence supporting visual tuning to the light environment: terrestrial species had longer λ<sub>max</sub> than aquatic species, and of these, species from turbid shallow waters had longer λmax than those from clear or deep waters. Of the terrestrial species, diurnal species had longer λ<sub>max</sub> than nocturnal species, but we did not detect any differences across terrestrial habitats (closed, intermediate or open). We found no association with proxies for visual tasks such as having red morphological features or utilising flowers or coral reefs. These results support the emerging consensus that, in general, visual systems are adapted to broad range of tasks, rather than tuned to certain tasks.</p>
VISIONE Feature Repository for VBS: Multi-Modal Features and Detected Objects from VBSLHE Dataset
<p>This repository contains a diverse set of features extracted from the VBSLHE dataset (laparoscopic gynecology) . These features will be utilized in the VISIONE system [Amato et al. 2023, Amato et al. 2022] in the next editions of the Video Browser Showdown (VBS) competition (<a href="https://www.videobrowsershowdown.org/">https://www.videobrowsershowdown.org/</a>). </p> <p>We used a snapshot of the dataset provided by the Medical University of Vienna and Toronto that can be downloaded using the instructions provided at <a href="https://download-dbis.dmi.unibas.ch/mvk/">https://download-dbis.dmi.unibas.ch/mvk/</a>. It comprises 75 video files. We divided each video into video shots with a maximum duration of 5 seconds.</p> <p>This repository is released under a Creative Commons Attribution license. If you use it in any form for your work, please cite the following paper:</p> <blockquote> <p>@inproceedings{amato2023visione, title={VISIONE at Video Browser Showdown 2023}, author={Amato, Giuseppe and Bolettieri, Paolo and Carrara, Fabio and Falchi, Fabrizio and Gennaro, Claudio and Messina, Nicola and Vadicamo, Lucia and Vairo, Claudio}, booktitle={International Conference on Multimedia Modeling}, pages={615--621}, year={2023}, organization={Springer} } </p> </blockquote> <p> </p> <p>This repository (v2) comprises the following files:</p> <ul> <li><em><strong>msb.tar.gz </strong></em> contains tab-separated files (.tsv) for each video. Each tsv file reports, for each video segment, the timestamp and frame number marking the start/end of the video segment, along with the timestamp of the extracted middle frame and the associated identifier ("id_visione").</li> <li><em><strong>extract-keyframes-from-msb.tar.gz</strong></em> contains a Python script designed to extract the middle frame of each video segment from the MSB files. To run the script successfully, please ensure that you have the original VBSLHE videos available.</li> <li><em><strong>features-aladin.tar.gz†</strong></em><strong> </strong>contains <a href="https://github.com/mesnico/ALADIN">ALADIN</a> [Messina N. et al. 2022] features extracted for all the segment's middle frames.</li> <li><em><strong>features-clip-laion.tar.gz†</strong></em> contains <a href="https://huggingface.co/laion/CLIP-ViT-H-14-laion2B-s32B-b79K">CLIP ViT-H/14 - LAION-2B </a>[Schuhmann et al. 2022] features extracted for all the segment's middle frames.</li> <li><em><strong>features-clip-openai.tar.gz† </strong></em>contains <a href="https://huggingface.co/openai/clip-vit-large-patch14">CLIP ViT-L/14</a> [Radford et al. 2021] features extracted for all the segment's middle frames.</li> <li><em><strong>features-clip2video.tar.gz† </strong></em>contains <a href="https://github.com/CryhanFang/CLIP2Video">CLIP2Video</a> [Fang H. et al. 2021] extracted for all the video segments. <strong> </strong></li> <li><em><strong>objects-frcnn-oiv4.tar.gz* </strong></em>contains the objects detected using <a href="http://tfhub.dev/google/faster_rcnn/openimages_v4/inception_resnet_v2/1">Faster R-CNN+Inception ResNet</a> (trained on the Open Images V4 [Kuznetsova et al. 2020]).</li> <li><em><strong>objects-mrcnn-lvis.tar.gz*</strong></em> contains the objects detected using Mask R-CNN [He et al. 2017] (trained on LVIS).</li> <li><em><strong>objects-vfnet64-coco.tar.gz*</strong></em> contains the objects detected using VfNet [Zhang et al. 2021] (trained on COCO dataset).</li> </ul> <p>*Please be sure to use the <strong>v2 version </strong>of this repository, since v1 feature files may contain inconsistencies that have now been corrected</p> <p><em><strong>*Note on the object annotations:</strong></em> Within an object archive, there is a jsonl file for each video, where each row contains a record of a video segment (the <em>"_id"</em> corresponds to the <em>"id_visione"</em> used in the msb.tar.gz) . Additionally, there are three arrays representing the objects detected, the corresponding scores, and the bounding boxes. The format of these arrays is as follows:</p> <ul> <li><em>"object_class_names"</em>: vector with the class name of each detected object.</li> <li><em>"object_scores"</em>: scores corresponding to each detected object.</li> <li><em>"object_boxes_yxyx"</em>: bounding boxes of the detected objects in the format <em>(ymin, xmin, ymax, xmax).</em></li> </ul> <p> </p> <p><em><strong>†Note on the cross-modal features: </strong></em>The extracted multi-modal features (ALADIN, CLIPs, CLIP2Video) enable internal searches within the VBSLHE dataset using the query-by-image approach (features can be compared with the dot product). However, to perform searches based on free text, the text needs to be transformed into the joint embedding space according to the specific network being used (see links above). Please be aware that t<strong>he service for transforming text into features is not provided within this repository and should be developed independently using the original feature repositories linked above.</strong></p> <p>We have plans to release the code in the future, allowing the reproduction of the VISIONE system, including the instantiation of all the services to transform text into cross-modal features. However, this work is still in progress, and the code is not currently available.</p> <p> </p> <p><strong>References:</strong></p> <p>[Amato et al. 2023] Amato, G.et al., 2023, January. VISIONE at Video Browser Showdown 2023. In International Conference on Multimedia Modeling (pp. 615-621). Cham: Springer International Publishing.</p> <p>[Amato et al. 2022] Amato, G. et al. (2022). VISIONE at Video Browser Showdown 2022. In: , et al. MultiMedia Modeling. MMM 2022. Lecture Notes in Computer Science, vol 13142. Springer, Cham. </p> <p>[Fang H. et al. 2021] Fang H. et al., 2021. Clip2video: Mastering video-text retrieval via image clip. arXiv preprint arXiv:2106.11097.</p> <p>[He et al. 2017] He, K., Gkioxari, G., Dollár, P. and Girshick, R., 2017. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision (pp. 2961-2969).</p> <p>[Kuznetsova et al. 2020] Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A. and Duerig, T., 2020. The open images dataset v4. International Journal of Computer Vision, 128(7), pp.1956-1981.</p> <p>[Lin et al. 2014] Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P. and Zitnick, C.L., 2014, September. Microsoft coco: Common objects in context. In European conference on computer vision (pp. 740-755). Springer, Cham.</p> <p>[Messina et al. 2022] Messina N. et al., 2022, September. Aladin: distilling fine-grained alignment scores for efficient image-text matching and retrieval. In Proceedings of the 19th International Conference on Content-based Multimedia Indexing (pp. 64-70).</p> <p>[Radford et al. 2021] Radford A. et al., 2021, July. Learning transferable visual models from natural language supervision. In International conference on machine learning (pp. 8748-8763). PMLR.</p> <p>[Schuhmann et al. 2022] Schuhmann C. et al., 2022. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35, pp.25278-25294.</p> <p>[Zhang et al. 2021] Zhang, H., Wang, Y., Dayoub, F. and Sunderhauf, N., 2021. Varifocalnet: An iou-aware dense object detector. In Proceedings of the IEEE/CV</p>
VISIONE Feature Repository for VBS: Multi-Modal Features and Detected Objects from MVK Dataset
<p>This repository contains a diverse set of features extracted from the marine video (underwater) dataset (MVK) . These features were utilized in the VISIONE system [Amato et al. 2023, Amato et al. 2022] during the latest editions of the Video Browser Showdown (VBS) competition (<a href="https://www.videobrowsershowdown.org/">https://www.videobrowsershowdown.org/</a>). </p> <p>We used a snapshot of the MVK dataset from 2023, that can be downloaded using the instructions provided at <a href="https://download-dbis.dmi.unibas.ch/mvk/">https://download-dbis.dmi.unibas.ch/mvk/</a>. It comprises 1,372 video files. We divided each video into 1 second segments. </p> <p>This repository is released under a Creative Commons Attribution license. If you use it in any form for your work, please cite the following paper:</p> <blockquote> <pre>@inproceedings{amato2023visione, title={VISIONE at Video Browser Showdown 2023}, author={Amato, Giuseppe and Bolettieri, Paolo and Carrara, Fabio and Falchi, Fabrizio and Gennaro, Claudio and Messina, Nicola and Vadicamo, Lucia and Vairo, Claudio}, booktitle={International Conference on Multimedia Modeling}, pages={615--621}, year={2023}, organization={Springer} }</pre> </blockquote> <p> </p> <p>This repository comprises the following files:</p> <ul> <li><strong><em>msb.tar.gz </em></strong> contains tab-separated files (.tsv) for each video. Each tsv file reports, for each video segment, the timestamp and frame number marking the start/end of the video segment, along with the timestamp of the extracted middle frame and the associated identifier ("id_visione"). </li> <li><em><strong>extract-keyframes-from-msb.tar.gz</strong></em> contains a Python script designed to extract the middle frame of each video segment from the MSB files. To run the script successfully, please ensure that you have the original MVK videos available.</li> <li><strong><em>features-aladin.tar.gz<sup>†</sup></em> </strong>contains <a href="https://github.com/mesnico/ALADIN">ALADIN</a> [Messina N. et al. 2022] features extracted for all the segment's middle frames. </li> <li><em><strong>features-clip-laion.tar.gz<sup>†</sup></strong></em> contains <a href="https://huggingface.co/laion/CLIP-ViT-H-14-laion2B-s32B-b79K">CLIP ViT-H/14 - LAION-2B </a>[Schuhmann et al. 2022] features extracted for all the segment's middle frames.</li> <li><em><strong>features-clip-openai.tar.gz<sup>†</sup> </strong></em>contains <a href="https://huggingface.co/openai/clip-vit-large-patch14">CLIP ViT-L/14</a> [Radford et al. 2021] features extracted for all the segment's middle frames. </li> <li><em><strong>features-clip2video.tar.gz<sup>†</sup> </strong></em>contains <a href="https://github.com/CryhanFang/CLIP2Video">CLIP2Video</a> [Fang H. et al. 2021] extracted for all the 1s video segments. <strong> </strong></li> <li><em><strong>objects-frcnn-oiv4.tar.gz<sup>*</sup> </strong></em>contains the objects detected using <a href="http://tfhub.dev/google/faster_rcnn/openimages_v4/inception_resnet_v2/1">Faster R-CNN+Inception ResNet</a> (trained on the Open Images V4 [Kuznetsova et al. 2020]). </li> <li><em><strong>objects-mrcnn-lvis.tar.gz<sup>*</sup></strong></em> contains the objects detected using Mask R-CNN [He et al. 2017] (trained on LVIS).</li> <li><em><strong>objects-vfnet64-coco.tar.gz<sup>*</sup></strong></em> contains the objects detected using VfNet [Zhang et al. 2021] (trained on COCO dataset).</li> </ul> <p>*Please be sure to use the <strong>v2 version </strong>of this repository, since v1 feature files may contain inconsistencies that have now been corrected</p> <p><em><strong>*Note on the object annotations:</strong></em> Within an object archive, there is a jsonl file for each video, where each row contains a record of a video segment (the <em>"_id"</em> corresponds to the <em>"id_visione"</em> used in the msb.tar.gz) . Additionally, there are three arrays representing the objects detected, the corresponding scores, and the bounding boxes. The format of these arrays is as follows:</p> <ul> <li><em>"object_class_names"</em>: vector with the class name of each detected object.</li> <li><em>"object_scores"</em>: scores corresponding to each detected object.</li> <li><em>"object_boxes_yxyx"</em>: bounding boxes of the detected objects in the format <em>(ymin, xmin, ymax, xmax).</em></li> </ul> <p> </p> <p><em><strong><sup>†</sup>Note on the cross-modal features: </strong></em>The extracted multi-modal features (ALADIN, CLIPs, CLIP2Video) enable internal searches within the MVK dataset using the query-by-image approach (features can be compared with the dot product). However, to perform searches based on free text, the text needs to be transformed into the joint embedding space according to the specific network being used (see links above). Please be aware that t<strong>he service for transforming text into features is not provided within this repository and should be developed independently using the original feature repositories linked above.</strong></p> <p>We have plans to release the code in the future, allowing the reproduction of the VISIONE system, including the instantiation of all the services to transform text into cross-modal features. However, this work is still in progress, and the code is not currently available.</p> <p> </p> <p><strong>References:</strong></p> <p>[Amato et al. 2023] Amato, G.et al., 2023, January. VISIONE at Video Browser Showdown 2023. In International Conference on Multimedia Modeling (pp. 615-621). Cham: Springer International Publishing.</p> <p>[Amato et al. 2022] Amato, G. et al. (2022). VISIONE at Video Browser Showdown 2022. In: , et al. MultiMedia Modeling. MMM 2022. Lecture Notes in Computer Science, vol 13142. Springer, Cham. </p> <p>[Fang H. et al. 2021] Fang H. et al., 2021. Clip2video: Mastering video-text retrieval via image clip. arXiv preprint arXiv:2106.11097.</p> <p>[He et al. 2017] He, K., Gkioxari, G., Dollár, P. and Girshick, R., 2017. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision (pp. 2961-2969).</p> <p>[Kuznetsova et al. 2020] Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A. and Duerig, T., 2020. The open images dataset v4. International Journal of Computer Vision, 128(7), pp.1956-1981.</p> <p>[Lin et al. 2014] Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P. and Zitnick, C.L., 2014, September. Microsoft coco: Common objects in context. In European conference on computer vision (pp. 740-755). Springer, Cham.</p> <p>[Messina et al. 2022] Messina N. et al., 2022, September. Aladin: distilling fine-grained alignment scores for efficient image-text matching and retrieval. In Proceedings of the 19th International Conference on Content-based Multimedia Indexing (pp. 64-70).</p> <p>[Radford et al. 2021] Radford A. et al., 2021, July. Learning transferable visual models from natural language supervision. In International conference on machine learning (pp. 8748-8763). PMLR.</p> <p>[Schuhmann et al. 2022] Schuhmann C. et al., 2022. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35, pp.25278-25294.</p> <p>[Zhang et al. 2021] Zhang, H., Wang, Y., Dayoub, F. and Sunderhauf, N., 2021. Varifocalnet: An iou-aware dense object detector. In Proceedings of the IEEE/CV</p>
Benchmark for energy efficient obstacle detection on head mounted wearable for the vision impaired
<p>Here we present a novel benchmark dataset with the associated challenge, that is to detect obstacles based on head-mounted sensors and lightweight wearable devices to assist Blind and Visually Impaired individuals (BVIs) navigate in indoor environments. The challenge encompasses three objectives: (1) as accurately as possible to detect the obstacles on the pathway that likely lead to a collision; (2) as durably as possible on a given amount of battery power for the detection algorithm or model to run; (3) as reliably as possible to compensate natural head turns so nearby objects would not trigger false alarms. The data provided in the benchmark are collected from the following head mounted sensors: (i) nine low-cost ultrasonic sensors; (ii) one high-end ultrasonic sensor with a larger detection range but higher power consumption; (iii) a 9-Degrees of Freedom (DOF) Inertial Measurement Unit (IMU). The resulting dataset consists of more than 188,000 unique sequences obtained from multiple subjects walking in three different indoor scenarios. This benchmark is to facilitate and encourage accurate yet fast obstacle detection solutions that can really benefit BVIs. </p>
Top view of DR1/DR2 double riffle, each section contains a spawning ground made up of eight gravel-filled trays, a rest area. The "double riffle" was designed to accommodate two groups from 25 to 50 specimens of broodstock in strictly identical conditions. The spawning grounds are equipped with waterproof, motion-sensing cameras with infrared night vision, connected to a 1000 Gb recorder. The diurnal and nocturnal activities of the two groups can therefore be simultaneously recorded over a long period. in Reproduction of Zingel asper (Linnaeus, 1758) in controlled conditions: an assessment of the experiences realized since 2005 at the Besançon Natural History Museum
Top view of DR1/DR2 double riffle, each section contains a spawning ground made up of eight gravel-filled trays, a rest area. The "double riffle" was designed to accommodate two groups from 25 to 50 specimens of broodstock in strictly identical conditions. The spawning grounds are equipped with waterproof, motion-sensing cameras with infrared night vision, connected to a 1000 Gb recorder. The diurnal and nocturnal activities of the two groups can therefore be simultaneously recorded over a long period.
UNA VISION HACIA LA ECONOMIA CIRCULAR DE LOS CONSUMIDORES DE LEÑA EN LA ORINOQUIA
<p>Los consumidores presentan gustos y preferencias diferentes por el tipo de leña que utilizan en sus restaurantes. Por ello, el objetivo es identificar esos aspectos de la demanda que influyen en el comportamiento del mercado. Los datos utilizados corresponden a la información de 97 propietarios y administradores de restaurante en los departamentos de Meta, Cundinamarca y Casanare, que mediante una encuesta y técnicas multivariadas permitieron la clasificación de 6 tipos de consumidores de leña, como los mayores consumen de leña (grupo 1), los que la utilizan para la elaboración de subproductos (grupo 2 y 4), los que deben pagar costos de transporte (grupo 6), los que compran todos los días (grupo 3 y 5), entre otros aspectos. Permitiendo tomar mejores decisiones en el uso eficiente de la leña y fomento de la producción y consumo responsables de la misma.</p>
Replication Package of "Exploiting Vision-Language Models in GUI Reuse"
<p>This replication package is for the paper entitled "Exploiting Vision-Language Models in GUI Reuse". The authors remain anonymous for double-blind review purposes. The package contains six files. If the paper is accepted, then the authors will move the replication package to a public repository hosted by an institution.</p>
Österreichweites Radzielnetz VISION
<p>Datensatz zum Projekt Österreichweites Radzielnetz VISION. Der Datensatz umfasst den Endbericht, eine QGIS-Projektdatei sowie zwei Geopackage-Dateien. Die Beschreibungen der Geopackages befinden sich im Anhang des Projektberichts. </p>
Article dataset: 'How stable are visions for protected area management? Stakeholder perspectives before and during a pandemic'
<p>This dataset includes the raw data of a survey of 38 stakeholders in the region of the Sierra de Guadarrama National Park, Spain (2019). The "Dataset" sheet presents respondents' Likert-scale answers (ranging from 1 to 5) of agreement regarding values, perceived changes and perceived drivers of change of the national park landscapes. The complete methodology is described in: Lo, V.B., López-Rodríguez, M.D., Metzger, M., Oteros-Rozas, E., Cebrián-Piqueras, M. A., Ruiz-Mallén, I., March, H., Raymond, C.M. (<em>in press</em>) ‘How stable are visions for protected area management? Stakeholder perspectives before and during a pandemic.’ <em>People and Nature</em>. Accompanying supplementary data (interview script, tiles and canvasses, coding, follow-up survey questions) are available as supplementary information to this article available for open access on the <em>People and Nature</em> website.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.