Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
193
datasets available to search
ShareScore release 0.9.0
Dataset results
193 results for “labeled data”
Fig. 5. Data labels. A. Xylocopamoerens Perty, 1833 in The primary types of some species of Centris bees described by European entomologists in the 18 and 20 centuries (Hymenoptera: Apidae)
Fig. 5. Data labels. A. Xylocopamoerens Perty, 1833 (lectotype ♀). B. Xylocopaxanthocnemis Perty, 1833 (lectotype ♂).
Fig. 1. Data labels. A. Centris debilis Dominique, 1898 in The primary types of some species of Centris bees described by European entomologists in the 18 and 20 centuries (Hymenoptera: Apidae)
Fig. 1. Data labels. A. Centris debilis Dominique, 1898 (lectotype ♂). B. Centris dominiquella Dominique, 1898 (lectotype ♀).
Fig 2. Data labels. A. Centris zonalis Dominique, 1898 in The primary types of some species of Centris bees described by European entomologists in the 18 and 20 centuries (Hymenoptera: Apidae)
Fig 2. Data labels. A. Centris zonalis Dominique, 1898 (lectotype ♀). B. Centris lyngbyei Jensen- Haarup, 1908 (lectotype ♂).
Fig. 4. Data labels. A in The primary types of some species of Centris bees described by European entomologists in the 18 and 20 centuries (Hymenoptera: Apidae)
Fig. 4. Data labels. A. Centrisrhodophthalma Pérez, 1911 (lectotype ♀). B. Centristransversa Pérez, 1905 (lectotype ♀).
UVA laser scanning labelled las data over tropical moist forest classified as leaf or wood points
<p>UAV Laser Scanning data collected over neotropical forest (Paracou French Guiana). Four flights conducted over one ha plot in 2021 and 2022.</p> <p>Leaf wood labels were transferred from contemporaneous (2021) TLS acquisition, for which segmentation was done using LeWoS and onscreen post correction.</p> <p>Predicted values using SOUL model (see ref below) for a single UAV-LS acquisition are provided as separate file.</p> <p> </p>
Data from: Label-free timing analysis of SiPM-based modularized detectors with physics-constrained deep learning
Open the record for dataset details and reuse information.
Hourly detections of echolocation clicks in Hawaiian Island HARP data from Hawai`i, Kaua`i, and Manawai with species labels
Open the record for dataset details and reuse information.
Data from: Automated cell lineage reconstruction using label-free 4D microscopy
Open the record for dataset details and reuse information.
Data for Precursor intensity-based label-free quantification software tools for Galaxy Platform.
<p><strong>Precursor intensity-based label-free quantification software tools for proteomic and multi-omic analysis within the Galaxy Platform.</strong></p> <p> </p> <p><strong>ABRF:</strong> Data was generated through the collaborative work of the ABRF Proteomics Research Group (<a href="https://abrf.org/research-group/proteomics-research-group-prg">https://abrf.org/research-group/proteomics-research-group-prg</a>). See Reference for details: Van Riper, S. et al. ‘<a href="http://cbs.umn.edu/sites/cbs.umn.edu/files/public/downloads/2016_ABRF_PRG_Poster_for_ASMS_20160501.pdf">An ABRF-PRG study: Identification of low abundance proteins in a highly complex protein sample</a>’ at the 64th Annual Conference of American Society of Mass Spectrometry and Allied Topics" at San Antonio, TX."</p> <p><strong>UPS:</strong> <strong>MaxLFQ Cox J, Hein MY, Luber CA, Paron I, Nagaraj N, Mann M.<a href="https://www.ncbi.nlm.nih.gov/pubmed/24942700/"> Accurate proteome-wide label-free quantification by delayed normalization and maximal peptide ratio extraction, termed MaxLFQ.</a> Mol Cell Proteomics. 2014 Sep;13(9):2513-26. doi: 10.1074/mcp.M113.031591. Epub 2014 Jun 17. PubMed PMID: 24942700; PubMed Central PMCID: PMC4159666;</strong></p> <p><strong>PRIDE #5412; ProteomeXchange repository PXD000279: </strong>ftp://ftp.pride.ebi.ac.uk/pride/data/archive/2014/09/PXD000279</p>
Data set of labeled scenes in a barn in front of automatic milking system
<p>This data set includes supplementary files to our article \emph{"Deep learning image recognition of cow behavior near an automatic milking robot with an open data set"} (Leonardo Santiago Benitez Pereira, Olli Koskela, Ilpo Pölönen, Ilmo Aronen and Iivari Kunttu, in review, 2020). We acquired continuous video data over two-month period of cows in front of Automatic Milking Station (AMS). Each frame of the video is labeled as a single image belonging to one of classes described below.</p> <p>This data set includes 253 files of video data having in total \num{1526473} labeled frames. The data set used in the article consisted of 280 videos, but for privacy reasons, 27 video files including persons were removed from this data set, but their Java Script Object Notation (JSON) label files are left for, e.g., temporal analyses.</p>
Data from: Safety of single low-dose primaquine in glucose-6-phosphate dehydrogenase deficient falciparum-infected African males: two open-label, randomized, safety trials
Background: Primaquine (PQ) actively clears mature Plasmodium falciparum gametocytes but in glucose-6-phosphate dehydrogenase deficient (G6PDd) individuals can cause hemolysis. We assessed the safety of low-dose PQ in combination with artemether-lumefantrine (AL) or dihydroartemisinin-piperaquine (DP) in G6PDd African males with asymptomatic P. falciparum malaria. Methods and findings: In Burkina Faso, G6PDd adult males were randomized to treatment with AL alone (n = 10) or with PQ at 0.25 (n = 20) or 0.40 mg/kg (n = 20) dosage; G6PD-normal males received AL plus 0.25 (n = 10) or 0.40 mg/kg (n = 10) PQ. In The Gambia, G6PDd adult males and boys received DP alone (n = 10) or with 0.25 mg/kg PQ (n = 20); G6PD-normal males received DP plus 0.25 (n = 10) or 0.40 mg/kg (n = 10) PQ. The primary study endpoint was change in hemoglobin concentration during the 28-day follow-up. Cytochrome P-450 isoenzyme 2D6 (CYP2D6) metabolizer status, gametocyte carriage, haptoglobin, lactate dehydrogenase levels and reticulocyte counts were also determined. In Burkina Faso, the mean maximum absolute change in hemoglobin was -2.13 g/dL (95% confidence interval [CI], -2.78, -1.49) in G6PDd individuals randomized to 0.25 PQ mg/kg and -2.29 g/dL (95% CI, -2.79, -1.79) in those receiving 0.40 PQ mg/kg. In The Gambia, the mean maximum absolute change in hemoglobin concentration was -1.83 g/dL (95% CI, -2.19, -1.47) in G6PDd individuals receiving 0.25 PQ mg/kg. After adjustment for baseline concentrations, hemoglobin reductions in G6PDd individuals in Burkina Faso were more pronounced compared to those in G6PD-normal individuals receiving the same PQ doses (P = 0.062 and P = 0.022, respectively). Hemoglobin levels normalized during follow-up. Abnormal haptoglobin and lactate dehydrogenase levels provided additional evidence of mild transient hemolysis post-PQ. Conclusions: Single low-dose PQ in combination with AL and DP was associated with mild and transient reductions in hemoglobin. None of the study participants developed moderate or severe anemia; there were no severe adverse events. This indicates that single low-dose PQ is safe in G6PDd African males when used with artemisinin-based combination therapy.
Data from: Long-lived metabolic enzymes in the crystalline lens identified by pulse-labeling of mice and mass spectrometry
<p>The lenticular fiber cells are comprised of extremely long-lived proteins while still maintaining an active biochemical state. Dysregulation of these activities has been implicated in age-related cataracts, and other lens diseases. However, the lenticular protein dynamics underlying health and disease is unclear. We sought to measure the global protein turnover rates in the eye using dietary nitrogen-15 (15N)-labeling of mice between 3 and 15 weeks of age. By performing mass spectrometry we measured the 14N- to 15N-peptide ratios of 248 lens proteins, including Crystallin, Aquaporin, Collagen and Laminin of the lens capsule, and enzymes that catalyze glycolysis as well as oxidation and reduction reactions. Unexpectedly, like the crystallin proteins, many of these enzymes are also exceedingly long-lived. The slow replacement of these enzymes in spite of young age of the mice suggests their potential roles in age-related metabolic changes in the lens.</p>
EEG Data for: "Learning from Label Proportions in Brain-Computer Interfaces"
<p>Data about two experiments is contained in this repository. An EEG experiment utilizing visual event-related potentials (ERPs) with N=13 healthy subjects was conducted in addition to a smaller study with N=5 subjects performing both an auditory and a visual ERP paradigm. </p> <p>The dataset is used and described in the following journal article:</p> <p><em>Hübner, D., Verhoeven, T., Schmid, K., Müller, K. R., Tangermann, M., & Kindermans, P. J. (2017). Learning from label proportions in brain-computer interfaces: online unsupervised learning with guarantees. PloS one, 12(4), e0175856.</em></p> <p><strong>Please cite the above article when using the data.</strong></p> <p>The larger data set with N=13 is different to ordinary ERP datasets in the sense that the train of stimuli to spell one character (68) is divided into repetitions of two interleaved sequences with length 8 and 18, respectively. We added '#' symbols to the spelling matrix which should never be attended by the subject and hence, are non-targets by definition. The first, shorter sequence, now highlights only ordinary characters, while the second sequence also highlights '#' -- visual blank symbols. By construction, sequence 1 has a higher target ratio than sequence 2. These known, but different target and non-target proportions are then used to reconstruct the target and non-target class means. This approach which does not need explicit class labels is termed Learning from Label Proportions (LLP). It can be used to decode brain signals without prior calibration session. More details can be found in the article.</p> <p>In another study, the above data set was used to simulate a new unsupervised mixture approach which combines the mean estimation of the unsupervised expectation-maximization algorithm by Kindermans et al. (2012, PLoS One) with the means obtained with the LLP approach. This leads to an unsupervised solution for which the performance is as good as in the supervised scenario. Please find more details in the following article:</p> <p><em>Verhoeven, T., Hübner, D., Tangermann, M., Müller, K. R., Dambre, J., & Kindermans, P. J. (2017). Improving zero-training brain-computer interfaces by mixing model estimators. Journal of neural engineering, 14(3), 036021.</em></p> <p>The following files are available:</p> <p>description.pdf: <strong>Full description of the dataset</strong><br> offline_auditory.zip: <strong>Data from the auditory offline study with N=5 subjects</strong><br> offline_visual.zip: <strong>Data from the visual offline study with N=5 subjects</strong><br> online_study_1-7.zip: <strong>Data from the online study for subjects 1-7</strong><br> online_study_8-13.zip: <strong>Data from the online study for subjects 8-13</strong><br> sequence.mat: <strong>Sequence data necessary for applying LLP to the online study. It is the same for all subjects</strong></p> <p>We will create a git repository with example code soon.</p>
SEM Images of Arrays of III-V Nanowires with Labelled Data
<p>The Dataset is a .zip folder.</p><p>The Dataset contains 240 .tiff images in black and white, each paired by file name to a .json file containing labeling information. </p><p>Each image consists of an array of 66 nanowire structures. Some images contain parasitic crystals and defective wires. Wires growing from seed surfaces parallel to and misaligned with the array walls are present.</p><p>The .json label files were created and can be processed with LabelMe (<a href="https://github.com/wkentaro/labelme">https://github.com/wkentaro/labelme</a>). The explanation for the labels can be found in the associated open-access paper currently submitted for publication.</p><p>The Licensor, IBM Corp.</p>
Examples for running TraceGroomer: format and normalise your labeled metabolomics data for DIMet analysis
<p>Examples to test our tool <a href="https://github.com/cbib/TraceGroomer">TraceGroomer</a>, to prepare your files for DIMet (Differential analysis of Isotope-labeled Metabolomics data).</p> <p>Each example represents one type of input supported by TraceGroomer. Please download the entire .zip and find inside the example that best suits your case.</p> <p>The new version v2 contains four types of input:</p> <ol> <li>IsoCor .tsv generated file: <em>example-isocor_data</em></li> <li>The rule of "three .tsv files" , i.e. sampleMetadata, variableMetadata and dataMatrix: <em>example-ruletsv_data</em></li> <li>custom or generic .xlsx file: <em>example-sheet_data</em></li> <li>VIB MEC .xlsx file: <em>example-vib_data</em></li> </ol> <p>Visit the <a href="https://github.com/cbib/TraceGroomer/wiki">TraceGroomer Wiki page</a> for further information and how to run the tool on the provided examples. </p> <p><strong>For users of the Galaxy </strong>version of Tracegroomer: please only use the 'data/' folder (ignore the 'groom_files/' folder)</p>
Simple labelled data for WildFi tag lifting experiment
<p>A tiny dataset with ~2000 recordings. We used it to test the WildFi's capabilities of using a classifier to derive simple certain states of the tag from its sensor readings.</p> <div></div>
Supplementary methods and data for: Dogs with a vocabulary of object label remember labels for at least two years
<p>Long-term memory of words has a crucial role in the developing abilities of young children to acquire language. In dogs, the ability to learn object labels is present in only a small group of uniquely Gifted Word Learner (GWL) dogs. The ability of these dogs to acquire large vocabularies consisting of hundreds of names of dog toys through naturally occurring interactions in human families presents them as a valid model for studying language-related cognitive mechanisms. As they are very rare, little is known about the mechanisms through which they acquire such large vocabularies. In the current study, we tested the ability of five GWL dogs to retrieve 12 labelled objects two years after the object-label mapping acquisition. The dogs proved to remember the labels of between 3-9 objects. The results shed light on the process by which GWL dogs acquire an exceptionally large vocabulary of object names. As memory plays a crucial role in language development, these dogs supply a unique opportunity to study label retention in a non-linguistic species.</p>
The fully executable procedure of the U-Net model combined with the Multi-textRG algorithm to achieve fine ice-water classification ---- another 332 scenes of data-fused SIC labels.
<p>This data source is related to the manuscript titled "Combining the U-Net model and a Multi-textRG algorithm for fine SAR ice-water classification", which will be submitted to the journal---The Cryosphere. </p> <ul> <li>The"ready-to-train-fused_01.zip" to "ready-to-train-fused_10.zip" includes 200 scenes of data-fused SIC labels accessible with doi: 10.5281/zenodo.10973107, https://zenodo.org/records/10973107. </li> <li>The "ready-to-train-fused_11.zip" to "ready-to-train-fused_21.zip" includes another 332 scenes of data-fused SIC labels. </li> </ul>
Data associated with 'Closing the stellar labels gap: Stellar label independent evidence for [α/M] information in Gaia BP/RP spectra'
<p>We present a stellar label independent model for Gaia BP/RP (XP) spectra, which does not rely on stellar labels to simulate XP spectra. Here, we provide all of the data required to reproduce <a href="https://github.com/AlexLaroche7/xp_vae">our results</a>. All of the data files should be placed in the data folder. Below we brieflty summarize the contents of each:</p> <p>APOGEE/GAIA XP DATA:</p> <ul> <li>xp_apogee_cat.h5: Contains all Gaia XP data and APOGEE stellar labels used in this work.</li> <li>xp_corrs.npz: Contains Gaia XP covariance matrices for a subset of the test data to compute reconstruction errors.</li> </ul> <p>INTERMEDIATE DATA PRODUCTS:</p> <p><em>(All files below can be reproduced with our codebase. We simply include them to speed up runtime for reproducing our results.)</em></p> <ul> <li>lb23_xp_est_labels.npy: <a href="https://ui.adsabs.harvard.edu/abs/2024MNRAS.527.1494L/abstract">Leung & Bovy (2023)</a> XP coeffficient spectra reconstructions (stellar label dependent implementaton)</li> <li>lb23_xp_est_no_labels.npy: Leung & Bovy (2023) XP coeffficient spectra reconstructions (stellar label independent implementaton)</li> <li>xp_wavelength_space.npy: Observed Gaia XP spectra in wavelength space</li> <li>err_wavelength_space.npy: Observed Gaia XP uncertainties in wavelength space</li> <li>apogee_norm.npz: Means and standard deviations which are used to pre-process Gaia XP data before inputting into our model</li> <li>zhang_stellar_params_xmatch.npz: <a href="https://ui.adsabs.harvard.edu/abs/2023MNRAS.524.1855Z/abstract">Zhang, Green, Rix (2023)</a> stellar parameter estimates for subset of XP spectra (to compute reconstructions)</li> <li>vae_wavelength_space.npy: Our model XP coefficient reconstructions in wavelength space</li> </ul> <p>If you have any questions please reach out via email: alex.laroche@mail.utoronto.ca</p>
Data and software for "Metal Pad Sensing: exploiting the electrical double layer to improve resistance-based microfluidic cell tracking, with applications to label-free mechanophenotyping"
<p>Data and software for "Metal Pad Sensing: exploiting the electrical double layer to improve resistance-based microfluidic cell tracking, with applications to label-free mechanophenotyping"</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.