Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,506
datasets available to search
ShareScore release 0.7.1
Dataset results
1,506 results for “objects”
Classifying Transients and Variable Objects with SCONE
<p>In this work, we expand the use of the Supernova Classifier with a Convolutional Neural Network (SCONE) to include classification of an additional two classes of transient objects and five classes of variable objects. SCONE has been adapted for these additional types of objects by updating its data processing pipeline to better handle variable objects, which exhibit repeated changes in brightness over time, as well as introducing class weights to ensure the model picks up on the nuances of the imbalanced training and test sets (both pulled from the PLAsTiCC dataset). Preliminary results show SCONE is capable of classifying transients with 80% accuracy and variable stars with 70% accuracy. To further improve its performance, we are now working on an original method of simulating light curves using templates found in the PLAsTiCC dataset to create a larger and more representative test set.</p>
Strawberry dataset for object detection
<p>Object detection dataset with annotated ripe and unripe strawberries and their peduncles. Data is annotated in YOLO format. Collected for: ”Real-Time CNN-based Computer Vision System for Open-Field Strawberry Harvesting Robot” <a href="https://doi.org/10.1016/j.ifacol.2022.11.109">https://doi.org/10.1016/j.ifacol.2022.11.109</a> .</p>
DoPose: dataset for object segmentation and 6D pose estimation
<p>DoPose (Dortmund Pose)is a dataset of highly cluttered and closely stacked objects. The dataset is saved in the <a href="https://github.com/thodan/bop_toolkit/blob/master/docs/bop_datasets_format.md">BOP format</a>. The dataset includes RGB images, Depth images, 6D Pose of objects, segmentation mask (all and visible), COCO Json annotation, camera transformations, and 3D model of all objects. The dataset contains 2 different types of scenes (table and bin). Each scene contains different view angles. For the bin scenes, the data contains 183 scenes with 2150 image views. In those 183 scenes 35 scenes contain 2 views, 20 contains 3 views and 128 contains 16 views. And for table scenes, the data contains 118 scenes with 1175 image views. in Those 118 scenes, 20 scenes contain 3 views, 50 scenes with 6 images, and 48 scenes with 17 images. So in total, our data contains 301 scenes and 3325 view images. Most of the scenes contain mixed objects. The dataset contains 19 objects in total.</p> <p>For more info about the dataset content and collection process please refer to our <a href="https://arxiv.org/abs/2204.13613">Arxiv preprint</a></p> <p>If you have any questions about the dataset, please contact <strong>anas.gouda@tu-dortmund.de</strong></p>
Data for 17IND08 AdvanCT "report on the traceable measurement of freeform objects"
<p>Raw and evaluated data from 17IND08 AdvanCT's case study on freeforms, "report on the traceable measurement of freeform objects".</p> <p>Contains point clouds from CT (X-ray computed tomography) and tactile point clouds of two freeform workpieces made from titanium or polymer. Also includes data from registration spheres on both parts, and histogram data from nominal-actual comparisons between tactile and CT surfaces.</p>
FLOAT: Factorized Learning of Object Attributes for Improved Multi-object Multi-part Scene Parsing
<p>Pascal-Part-201 is the most comprehensive and challenging version of the Pascal-Part dataset for Multi-object multi-part scene parsing. The dataset is the part of the publication "FLOAT: Factorized Learning of Object Attributes for Improved Multi-object Multi-part Scene Parsing" published in CVPR 2022.</p>
Oral cancer speech corpus for the paper "Objective speech outcomes after surgical treatment for oral cancer: An acoustic analysis of a spontaneous speech corpus containing 32.850 tokens"
<p>Dataset accompanying the paper "<em>Objective speech outcomes after surgical treatment for oral cancer: An acoustic analysis of a spontaneous speech corpus containing 32.850 tokens</em>"</p> <p>The zip file contains five folders:</p> <p>- <strong>Database:</strong> contains csv files for each speaker which contain the processed features</p> <p>- <strong>Recordings: </strong>the original recording from the YouTube Oral Cancer speech dataset, without further preprocessing</p> <p>- <strong>Recordings_Normalised:</strong> same as recordings but after minimal audio preprocessing (min-max scaling)</p> <p>- <strong>Textgrids: </strong>contains the textgrids which are annotated on the word-level and on phoneme-level</p> <p>- <strong>TIMIT selection: </strong>contains the textgrids for the TIMIT speakers. We unfortunately cannot share the audio date as it is not open source. More information can be found <a href="https://catalog.ldc.upenn.edu/LDC93s1">here.</a></p>
ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes
<p>Less than 35% of recyclable waste is being actually recycled in the US, which leads to increased soil and sea pollution and is one of the major concerns of environmental researchers as well as the common public. At the heart of the problem are the inefficiencies of the waste sorting process (separating paper, plastic, metal, glass, etc.) due to the extremely complex and cluttered nature of the waste stream. Recyclable waste detection poses a unique computer vision challenge as it requires detection of highly deformable and often translucent objects in cluttered scenes without the kind of context information usually present in human-centric datasets. This challenging computer vision task currently lacks suitable datasets or methods in the available literature. In this paper, we take a step towards computer-aided waste detection and present the first in-the-wild industrial-grade waste detection and segmentation dataset, ZeroWaste. We believe that ZeroWaste will catalyze research in object detection and semantic segmentation in extreme clutter as well as applications in the recycling domain.</p> <p>Our project page can be found at <a href="http://ai.bu.edu/zerowaste/">http://ai.bu.edu/zerowaste/</a></p> <p>Please use the following password to extract the zip files: UP#1VuX409z4</p>
A list of body and object concepts
<p>The list includes 784 human body parts and object concepts. The concepts are based on the available concept sets in <a href="https://concepticon.clld.org/">Concepticon</a> (List et al. <a href="https://aclanthology.org/L16-1379/">2016</a>, <a href="https://doi.org/10.5281/zenodo.596412">2021</a>) and were tagged as <em>human body part,</em> <em>animal</em>, <em>clothing</em>, <em>food</em>, <em>household</em> <em>items</em>, <em>instrument</em>, <em>landscape</em>, <em>plant</em>, <em>spatial</em> <em>relation</em>, <em>tool</em>, and <em>vehicle</em>.</p>
Optimizing parametric factors in CIELAB and CIEDE2000 color-difference formulas for 3D printed spherical objects
<p>Forty-five spherical samples were printed using a Stratasys J750 3D color printer, and 82 pairs of 3D samples were produced to investigate the human color perception of the lightness, chroma and hue differences of 3D spherical objects, and to optimize the current CIELAB and CIEDE2000 color-difference formulas. This file contains the CIELAB values of the 45 spherical samples and the calculated colour differences as well as visual colour-difference data of 82 pairs of 3D samples. Optimizations of parametric factors in CIELAB and CIEDE2000 colour-difference formulas were performed based on the colour-difference data provided.</p>
Data and original code for: A generalized approach to characterise optical properties of natural objects
<p>To understand the diversity of ways in which natural materials interact with light, it is important to consider how their reflectance changes with the angle of illumination or viewing and to consider wavelengths beyond the visible. We chose a set of existing measurements and parameters that are generalisable to any wavelength range and spectral shape and we highlight which subsets of measures are relevant to different biological questions. As a case study, we applied these measures to 30 species of Christmas beetles. Here we provide the raw spectral data of angle integrated and angle-dependent reflection by the beetle elytra. We also provide the original code used for our analysis and figures.</p>
The input data set includes 729 objects (patients) and 39 variables (clinical qualitative and quantitative descriptors).
<p>For reliable data treatment and interpretation qualitative descriptors were omitted and only numerical clinical indicators were included in the data matrix. Finally, the data set dimension was [729 x 18].</p> <p> The data were treated by hierarchical cluster analysis and factor analysis. The major goal of the data mining was to reach statistically significant partitioning of the objects and variables into similarity patterns (clusters) which helps to better understand the data structure, to assess the meaning of the partitioning achieved, thus promoting the evaluation of the health status of the patients and the role of specific descriptors for the formation of the partitioning patterns.</p> <p>3D classification Python tool.</p>
Main Sequence + Compact Object binary candidates from Gaia DR3 astrometric and spectroscopic excess noise
<p>MS+CO systems selected from Gaia DR3 via inferred periods and mass ratios derived from astrometric and spectroscopic errors.</p> <p>The sample is split into a bronze list (significant astrometric and spectroscopic RUWE, mass ratio > 1 and companion mass > 3 Msun). </p> <p>A subset of these is chosen as a silver list (propagating errors on mass ratio and companion mass to deselect systems which are not significantly above the previous criteria)</p> <p>Finally, a gold list is constructed from the subset of the silver list which shows no evidence of being significantly brighter than a single MS star and with no significant excess photometric noise.</p> <p>We include the most relevant Gaia data for the system, as well as our inferred spectroscopic and photometric errors and RUWEs, and the inferred periods and mass ratios. Gaia's DR3 source id, and the ra, dec position are included and thus other Gaia data, or data from other astronomical catalogs, can be found for these systems.</p> <p>The catalog and the underlying methods are explained in more detail in <a href="https://arxiv.org/abs/2206.04392">Andrew et al. 2022</a>.</p>
Single-Cell Autism data stored as sce object
<p>The raw autism dataset is from UCSC Cell Browser Dataset, Autism section (<a href="https://cells.ucsc.edu/">https://cells.ucsc.edu</a>). It is stored as SingleCellExperiment object for further usage. </p>
Animal Recognition Using Methods Of Fine-Grained Visual Analysis - YOLOv5 Object Detection Dataset (Oxford-IIIT Pet)
<p>Preprocessed dataset for Oxford-IIIT Pet in YOLOv5 format.. Ground truth labels for head bounding boxes, body bounding boxes (derived from segmentation mask).</p>
Animal Recognition Using Methods Of Fine-Grained Visual Analysis - YOLOv5 Object Detection Dataset (Tsinghua Dogs)
<p>Preprocessed dataset for Tsinghua Dogs in YOLOv5 format.. Ground truth labels for head bounding boxes, body bounding boxes</p>
Augmented Objects as Portals into Virtual Worlds: Using Audio to Create Immersive Experiences in Extended Realities - UMBRELLA AUDIO SPATIALIZATION DEMO
<p><strong>Technical demonstration</strong></p> <p>The results of the projection mapping system in the project are clear from the <a href="https://blog.zhdk.ch/immersivearts/dreaming-of-time-and-space/">main documentation video</a>; however, the impact of the spatial audio system in particular, is best experienced from directly underneath the umbrellas, where one can best appreciate the various levels of mixed reality. Unfortunately, it is difficult to document these effects within the artistic context of the project, and as such, we include a brief set of examples to better demonstrate the 6 degree of freedom sound spatialization capabilities of the umbrella system.</p> <p><em><strong>NOTE:</strong></em> The audio in the following examples is recorded from a fixed perspective (initially underneath the umbrella) and rendered binaurally. Unfortunately, the ambisonic microphone used does not capture directionality very well when the source (in this case, the umbrella speakers) is less than ~1 meter away, and in retrospect, a single channel of pink noise was not a wise choice as a source material, as it appears to cause additional phasing issues. Additionally, the effectiveness of binaural audio varies from listener to listener, so <em>the perceived effect in the video is not as strong as when experienced in person</em>; nonetheless, it is possible to get the basic idea of the spatialization algorithm in action from these examples.</p> <p>PLEASE WEAR HEADPHONES IN ORDER TO EXPERIENCE THE 3D EFFECT.</p> <p>In addition to the view of the entire scene from an outside perspective, several other views of the underlying software are displayed throughout the video, including:</p> <ul> <li> <p>A radar view of the scene (umbrella and sound source) as seen by the space manager software, where the:</p> <ul> <li> <p>Blue circle = umbrella</p> </li> <li> <p>Cyan triangle, yellow square = sound source</p> </li> </ul> </li> <li> <p>A view of elements of the spatialization software running on the umbrella, specifically the:</p> <ul> <li> <p>Relative gain calculations and current output levels of each speaker</p> </li> <li> <p>Results of supporting calculations (e.g. sound location after transformation from the global to local coordinate system, and scaling factors used to attenuate the overall volume of the sound as the distance from the umbrella to the sound changes)</p> </li> </ul> </li> </ul> <p><strong>Demo #1</strong></p> <p>Stationary umbrella with a moving virtual sound source (anchored to a rigid body)</p> <p><strong>Demo #2</strong></p> <p>Rotating umbrella with a stationary sound source (anchored to a rigid body)</p> <p><strong>Demo #3</strong></p> <p>Moving umbrella with a fixed sound source (anchored to a rigid body)</p> <p><strong>Demo #4</strong></p> <p>Moving umbrella with a fixed sound source (anchored to a virtual point in space, located above the microphone); as the umbrella approaches the source, the sound first fades into the room, then collapses into the umbrella, as show in Figure 7 ("Fading between umbrella and room with distance") in the main paper</p>
EDLO2ID: An Efficient-deep-learning-and-object-oriented Image Dataset for Large-scene Mapping
<p>EDLO2ID: An Efficient-deep-learning-and-object-oriented Image Dataset for Large-scene Mapping </p> <p>The dataset can be unzipped and includes an image dataset and a vector dataset, which includes nine land use/land cover categories (i.e., cropland, orchard, forestland, grassland, construction land, transportation land, water body, bare land, terrace) for object-oriented remote sensing image mapping using deep learning.</p>
GENEA Challenge 2022 objective evaluation data
<p>This Zenodo repository contains objective evaluation results for all test-set motion submitted by teams participating in the GENEA Challenge 2022.</p> <p> </p> <p>We caution the user that objective metrics are known to have poor correlation with actual, perceived motion quality. For definitions and explanations of the "Average jerk magnitude", "Average acceleration magnitude", and "Average Hellinger distance", please see the paper "Moving fast and slow: Analysis of representations and post-processing in speech-driven automatic gesture generation" by Kucherenko et al., published in the International Journal of Human–Computer Interaction in 2021. For a definition and explanation of the "Canonical correlation analysis (CCA) coefficient", see the paper "Speech-driven animation with meaningful behaviors" by <a href="https://www.sciencedirect.com/science/article/abs/pii/S0167639318300013?via%3Dihub#!">Sadoughi and Busso</a>, published in <a href="https://www.sciencedirect.com/journal/speech-communication">Speech Communication</a> in 2019. Code for computing these objective metrics is available through the challenge webpage.</p> <p> </p> <p>Attribution:</p> <p>If you use this material, please cite our latest paper on the GENEA Challenge 2022. At the time of writing (2022-08-10) this is our ACM ICMI 2022 paper:</p> <p>Youngwoo Yoon, Pieter Wolfert, Taras Kucherenko, Carla Viegas, Teodor Nikolov, Mihail Tsakov, and Gustav Eje Henter. 2022. The GENEA Challenge 2022: A large evaluation of data-driven co-speech gesture generation. In Proceedings of the ACM International Conference on Multimodal Interaction (ICMI '22). ACM.</p> <p>You can find the latest information and a BibTeX file on the project website:</p> <p><a href="https://youngwoo-yoon.github.io/GENEAchallenge2022/">https://youngwoo-yoon.github.io/GENEAchallenge2022/</a></p> <p> </p> <p>The material is available under a CC BY 4.0 international license, with the text provided in LICENSE.txt.</p> <p> </p> <p>To find more GENEA Challenge 2022 material on the web, please see:</p> <p>*<a href="https://youngwoo-yoon.github.io/GENEAchallenge2022/"> https://youngwoo-yoon.github.io/GENEAchallenge2022/</a></p> <p>*<a href="https://genea-workshop.github.io/2022/challenge/"> https://genea-workshop.github.io/2022/challenge/</a></p> <p> </p> <p>If you have any questions or comments, please contact:</p> <p>* The GENEA Challenge & Workshop organisers <genea-contact@googlegroups.com</p>
Data/Code: Objective monitoring of functional recovery after total knee and hip arthroplasty using sensor-derived gait measures
<p>Abstract</p> <p>Background: Inertial sensors hold the promise to objectively measure functional recovery after total knee (TKA) and hip arthroplasty (THA), but their value in addition to patient-reported outcome measures (PROMs) has yet to be demonstrated. This study investigated recovery of gait after TKA and THA using inertial sensors, and compared results to recovery of self-reported scores of pain and function.</p> <p>Methods: PROMs and gait parameters were assessed before and at two and fifteen months after TKA (n=24) and THA (n=24). Gait parameters were compared with healthy individuals (n=27) of similar age. Gait data were collected using inertial sensors on the feet, lower back, and trunk. Participants walked for two minutes back and forth over a 6m walkway with 180° turns. PROMs were obtained using the Knee Injury and Osteoarthritis Outcome Scores and Hip Disability and Osteoarthritis Outcome Score.</p> <p>Results: Gait parameters recovered to the level of healthy controls after both TKA and THA. Early improvements were found in gait-related trunk kinematics, while spatiotemporal gait parameters mainly improved between two and fifteen months after TKA and THA. Compared to the large and early improvements found in of PROMs, these gait parameters showed a different trajectory, with a marked discordance between the outcome of both methods at two months post-operatively.</p> <p>Conclusion: Sensor-derived gait parameters were responsive to TKA and THA, showing different recovery trajectories for spatiotemporal gait parameters and gait-related trunk kinematics. Fifteen months after TKA and THA, there were no remaining gait differences with respect to healthy controls. Given the discordance in recovery trajectories between gait parameters and PROMs, sensor-derived gait parameters seem to carry relevant information for evaluation of physical function that is not captured by self-reported scores.</p>
16S Phyloseq R object accompanying the paper: Effects of storage methods on total bacterial count and microbial composition of bovine colostrum
<p>This ready to load <strong>phyloseq</strong> R S4 object contains the ASV table, taxonomy table and sample metadata (16S V3-V4). This data was build using the DaDa2 (version 1.12.1) and phyloseq (version 1.32) R packages using our raw Illumina MiSeq PE300 sequencing data deposited at NCBI-SRA under BioProject: PRJNA872909.</p> <p>The accompanying (peer-reviewed) scientific article can be found here: </p> <ul> <li>https://www.todo</li> <li>DOI: todo</li> </ul> <p> </p> <p><strong>Study/draft abstract:</strong></p> <p>Lisa Robbers, Hannes Bijkerk, Lars Ravesloot, Alex Bossers, Mirjam Nielen, Ruurd Jorritsma, Ad Koets, Lindert Benedictus</p> <p>Neonatal calves need to acquire passive immunity through maternal colostrum, as they are immunologically naïve and the structure of the bovine placenta does not allow passage of maternal antibodies during pregnancy. Milked colostrum is not initially sterile and may even contain high bacterial counts. Minimizing total bacterial counts in colostrum is generally advised, however bacterial quality of colostrum comprises more than just bacterial quantities, but also depends on the specific bacteria present. While duration and temperature of colostrum storage are known to affect total plate counts (TPC), less is known about the effects of storage on the actual bacterial composition of the TPC. We speculated that, depending on the storage conditions, colostrum is a substrate in which certain bacterial species can thrive affecting the quality of colostrum.</p> <p>We therefore aimed to characterize the effects of different colostrum storage methods on the composition of the viable, aerobic, microbial community. Colostrum samples were stored at different temperatures and for different durations. Next, bacterial growth was assessed using the aerobe plate count culture method, followed by 16S rRNA gene amplicon sequencing. Differences in the TPC bacterial compositions of the stored colostrum samples, as determined by 16S rRNA gene amplicon sequencing, were mostly explained by the variation in bacterial composition of the colostrum sample directly after milking. In line with earlier studies, the results from our study show that the TPC increased when colostrum was stored for 24 hours at room temperature, but not when stored in a refrigerator for the same duration. Community structure of the TPC of colostrum stored at room temperature for 24 hours and stored in a refrigerator for a week was significantly different from the baseline samples. The 16S rRNA sequencing results indicate this is because of increased numbers of <em>Enterobacteriaceae.</em> <em>Enterobacteriaceae</em> abundance in refrigerated samples seemed to remain stable for the first 24 hours, but increased drastically after one week. The results indicate that microbial composition of stored colostrum is mostly influenced by the composition of the colostrum sample directly after milking, which is most probably the result of contamination or other environmental influences during the milking process. This study provides a deeper insight in the changes in the microbial composition of colostrum TPC during practical storage conditions and provides a primer for more detailed research into the determinants of bacterial composition of colostrum and the linked health effects.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.