Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,139
datasets available to search
ShareScore release 0.7.1
Dataset results
2,139 results for “recognition”
Fig. 3 in A Revision Of The Portunus Pelagicus (Linnaeus, 1758) Species Complex (Crustacea: Brachyura: Portunidae), With The Recognition Of Four Species
Fig. 3. Minimum Evolution bootstrap tree incorporating all unique COI haplotypes. Haplotypes of specimens obtained from Portunus pelagicus, P. segnis, P. reticulatus and P. armatus with P. trituberculatus, P. sanguinolentus and Charydis lucifera as outgroups. '*' Indicates the dominant haplotype found in P. pelagicus shared with eight P. reticulatus individuals. '**' Denotes two individuals collected from Japan that may constitute a possible cryptic species.
Fig. 2 in A Revision Of The Portunus Pelagicus (Linnaeus, 1758) Species Complex (Crustacea: Brachyura: Portunidae), With The Recognition Of Four Species
Fig. 2. Scatter plot of canonical scores from forward stepwise discriminant function analysis. Group 1, Portunus armatus; 2, P. reticulatus; 3, P. segnis; 4, P. pelagicus.
Performance analysis of micro-expression recognition over different sample image sizes.
<p>Performance of micro-expression recognition accuracy analyzed using different sample sizes with motion and geometric features. The sample sizes analyzed are: 140x170, 280x340, 560x680 and 1120x1360 using SMIC, CASMEII, CAS(ME)^2 and SAMM. The experiments were conducted using three different feature extraction setups: (i) optimized BiWOOF, (ii) full-face graph, and (iii) full-face graph with amplitude-based emotion magnification method (A-EMM). The results were compared with the results of the original BiWOOF presented in [1] as the baseline study. These files are published under CC0 license.</p> <p> </p> <p>References</p> <p>[1] Liong, S. T., See, J., Wong, K., & Phan, R. C. W. (2018). Less is more: Micro-expression recognition from video using apex frame. <em>Signal Processing: Image Communication</em>, <em>62</em>, 82-92.</p>
Fig. 14 in A Revision Of The Portunus Pelagicus (Linnaeus, 1758) Species Complex (Crustacea: Brachyura: Portunidae), With The Recognition Of Four Species
Fig. 14. Portunus segnis (Forskål, 1775) (specimen not preserved), Doha fish market, Qatar, (photograph: H. Q. Ng).
Figure 6 in Adulthood and phylogenetic analysis in gastropods: character recognition and coding in shells of Lavigeria (Cerithioidea, Thiaridae) from Lake Tanganyika
Figure 6. Plots of length against PC1 for all eight species of Lavigeria studied with regressed lines and R2 values. Length values are log transformed.
Figure 3 in Adulthood and phylogenetic analysis in gastropods: character recognition and coding in shells of Lavigeria (Cerithioidea, Thiaridae) from Lake Tanganyika
Figure 3. Character state APPL. An adult specimen of L. n. sp. X (left) and a juvenile (right). Notice the perimetric, wrinkle-like, lines on the front surface of the apertural lip of the adult. Scale bar = 0.2 cm.
Figure 7 in Adulthood and phylogenetic analysis in gastropods: character recognition and coding in shells of Lavigeria (Cerithioidea, Thiaridae) from Lake Tanganyika
Figure 7. Fifty per cent majority-rule consensus trees of five, nine, ten and 31 trees (from top to bottom, respectively) from four matrices. Matrices are coding the data of Table 1. See Analysis for explanation of the matrices. Tree and character statistics are given in Table 3. Optimality criterion: maximum parsimony, exhaustive search. All characters binary, of equal weight and unordered. Numbers indicate percentage of topologies that include the respective branches.
Figure 5 in Adulthood and phylogenetic analysis in gastropods: character recognition and coding in shells of Lavigeria (Cerithioidea, Thiaridae) from Lake Tanganyika
Figure 5. Character states DFST, AXRB and UEPW. Adult specimens of (A) L. n. sp. W, (B) L. n. sp. F, (C) L. n. sp. K, (D) L. n. sp. J showing the aperture in side view and (E) an apertural view of an adult L. n. sp. W. In A-D the trajectory of the suture tends to deviate downwards in comparison to the trajectory of the spiral cord of the previous whorl immediately above the suture (DFST). A and D also show the loss of, or irregularities in the appearance of axial sculpture (AXRB). In E, arrowheads show the undulations formed at the edge of the parietal side of the aperture (UEPW). Scale bar = 0.2 cm.
Figure 2 in Adulthood and phylogenetic analysis in gastropods: character recognition and coding in shells of Lavigeria (Cerithioidea, Thiaridae) from Lake Tanganyika
Figure 2. Character states WGPW and APLT. An adult specimen of L. n. sp. A (left) and a juvenile (right). The adult shows a thickened (APLT) and opaque (WGPW) inner surface of the apertural lip in comparison to the juvenile. Scale bar = 0.2 cm.
Figure 1 in Adulthood and phylogenetic analysis in gastropods: character recognition and coding in shells of Lavigeria (Cerithioidea, Thiaridae) from Lake Tanganyika
Figure 1. Eight species used in this study, apertural and side views of adult specimens. (A) Lavigeria new species N. (B) L. n. sp. F. (C) L. n. sp. J. (D) L. n. sp. X. (E) L. n. sp. K. (F) L. n. sp. W. (G) L. n. sp. A. (H) L. n. sp. U. C-H all belong to the same clade. A and B belong to different clades within the genus. Scale bar = 0.2 cm.
Figure 4 in Adulthood and phylogenetic analysis in gastropods: character recognition and coding in shells of Lavigeria (Cerithioidea, Thiaridae) from Lake Tanganyika
Figure 4. Character state APDT. An adult specimen of L. n. sp. J (left) and a juvenile on the right. Notice in the adult how the parietal side of the apertural lip is completely detached from the previous whorl and a false umbilicus has developed. Scale bar = 0.2 cm.
Scaled and Translated Image Recognition (STIR)
<p><strong>Paper:</strong> <a href="https://arxiv.org/abs/2211.10288">[2211.10288] Just a Matter of Scale? Reevaluating Scale Equivariance in Convolutional Neural Networks (arxiv.org)</a><br> <strong>Code:</strong> <a href="https://github.com/taltstidl/scale-equivariant-cnn">taltstidl/scale-equivariant-cnn: Official code for "Just a Matter of Scale? Reevaluating Scale Equivariance in Convolutional Neural Networks" (github.com)</a></p> <p>While convolutions are known to be invariant to (discrete) translations, scaling continues to be a challenge and most image recognition networks are not invariant to them. To explore these effects, we have created the Scaled and Translated Image Recognition (STIR) dataset. This dataset contains objects of size <span class="math-tex">\(s \in [17,64]\)</span>, each randomly placed in a <span class="math-tex">\(64 \times 64\)</span> pixel image.</p> <p><strong>Using the dataset</strong></p> <p>Depending on which data you are planning to use, download one or more of the following files. Data is stored in compressed <code>.npz</code> format and can be loaded as documented <a href="https://numpy.org/doc/stable/reference/generated/numpy.load.html">here</a>.</p> <table> <thead> <tr> <th>File</th> <th>Description</th> </tr> </thead> <tbody> <tr> <td><code>emoji.npz</code></td> <td>Emoji vector icons rendered as white icon on black background</td> </tr> <tr> <td><code>mnist.npz</code></td> <td>Classic MNIST handwritten digits rescaled to varying sizes</td> </tr> <tr> <td><code>trafficsign.npz</code></td> <td>Traffic signs from street imagery downscaled to varying sizes</td> </tr> <tr> <td><code>aerial.npz</code></td> <td>Objects in aerial imagery downscaled to varying sizes</td> </tr> </tbody> </table> <p>Each file contains multiple arrays that can be accessed in a dictionary-like fashion. The keys are documented below, where <code>n</code> is the number of classes for a given file and <code>m</code> is the number of instances for each class. Both <code>emoji.npz</code> (36 classes, 1 instance) and <code>mnist.npz</code> (10 classes, 50 instances) are in black & white while <code>trafficsign.npz</code> (16 classes, 25 instances) and <code>aerial.npz</code> (9 classes, 25 instances) are in color.</p> <table> <thead> <tr> <th>Key</th> <th>Shape</th> <th>Description</th> </tr> </thead> <tbody> <tr> <td><code>imgs</code></td> <td><code>(3, 48, n, m, 64, 64)</code> black & white, <code>(3, 48, n, 64, 64, 3)</code> color</td> <td>Images grouped into 3 sets (training, validation, testing) and 48 different scales. Values will be in range <code>0</code> to <code>255</code>.</td> </tr> <tr> <td><code>lbls</code></td> <td><code>(3, 48, n, m)</code></td> <td>Indices referencing ground truth labels. See <code>lbldata</code> for descriptive names. Values will be in range <code>0</code> to <code>n - 1</code>.</td> </tr> <tr> <td><code>scls</code></td> <td><code>(3, 48, n, m)</code></td> <td>Known scales as given by bounding box size. Values will be in range <code>17</code> to <code>64</code>.</td> </tr> <tr> <td><code>psts</code></td> <td><code>(3, 48, n, m, 2)</code></td> <td>Known position of bounding box. First value is distance to left edge, second value distance to top edge.</td> </tr> <tr> <td><code>metadata</code></td> <td><code>(6, 2)</code></td> <td>Metadata on title, description, author, license, version and date.</td> </tr> <tr> <td><code>lbldata</code></td> <td><code>(n,)</code></td> <td>Descriptive names for each ground truth labels.</td> </tr> </tbody> </table> <p>For use in Python a dataset class is provided that implements the basic functionality for loading a certain split and scale selection, as illustrated in the code below. It ensures shuffling is done in a consistent manner such that ground truth scales and positions can be retrieved. Metadata and label descriptions can be retrieved via <code>metadata</code> and <code>labeldata</code>, respectively.</p> <pre><code class="language-python">from data.dataset import STIRDataset dataset = STIRDataset('data/emoji.npz') # Obtain images and labels for training images, labels = dataset.to_torch(split='train', scales=[32, 64], shuffle=True) # Obtain known scales and positions for above scales, positions = dataset.get_latents(split='train', scales=[32, 64], shuffle=True) # Get metadata and label descriptions metadata = dataset.metadata label_descriptions = dataset.labeldata</code></pre> <p><strong>License and Attribution</strong></p> <p>When using this dataset for your own research, please respect the individual licenses of the original data. These are distributed within the data files' metadata. For attribution in papers, we recommend the following citations.</p> <ol> <li>D. Gandy, J. Otero, E. Emanuel, F. Botsford, J. Lundien, K. Jackson, M. Wilkerson, R. Madole, J. Raphael, T. Chase, G. Taglialatela, B. Talbot, and T. Chase. Font Awesome. https://fontawesome.com/v5/download, Nov. 2022.</li> <li>Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. <em>Proc. IEEE</em>, 86(11):2278–2324, Nov. 1998.</li> <li> C. Ertler, J. Mislej, T. Ollmann, L. Porzi, G. Neuhold, and Y. Kuang. The Mapillary Traffic Sign Dataset for Detection and Classification on a Global Scale. In <em>2020 16th Eur. Conf. Comput. Vision (ECCV)</em>, Glasgow, UK, Aug. 2020.</li> <li>G.-S. Xia, X. Bai, J. Ding, Z. Zhu, S. Belongie, J. Luo, M. Datcu, M. Pelillo, and L. Zhang. DOTA: A Large-Scale Dataset for Object Detection in Aerial Images. In <em>2018 IEEE/CVF Conf. Comput. Vision and Pattern Recognition (CVPR)</em>, pages 3974–3983, Salt Lake City, UT, USA, June 2018.</li> </ol>
Synthetic Hand Recognition Dataset
<p>This is the synthetic hand recognition dataset generated with the Synthetic Human Dataset Generator (available on github). It contains 90,000 images of 10 different virtual humans (5 female, 5 male) in various poses. All images are annotated with the color coded segmented hands.</p>
Rapid resource depletion on coral reefs disrupts competitor recognition processes among butterflyfish species
<p>Avoiding costly fights can help conserve energy needed to survive rapid environmental change. Competitor recognition processes help resolve contests without escalating to attack, yet we have limited understanding of how they are affected by resource depletion and potential effects on species coexistence. Using a mass coral mortality event as a natural experiment and 3,770 field observations of butterflyfish encounters, we test how rapid resource depletion could disrupt recognition processes in butterflyfishes. Following resource loss, heterospecifics approached each other more closely before initiating aggression, fewer contests were resolved by signalling, and the energy invested in attacks was greater. In contrast, behaviour towards conspecifics did not change. As predicted by theory, conspecifics approached one another more closely and were more consistent in attack intensity yet, contrary to expectations, resolution of contests via signalling was more common among heterospecifics. Phylogenetic relatedness or body size did not predict these outcomes. Our results suggest that competitor recognition processes for heterospecifics became less accurate after mass coral mortality, which we hypothesise is due to altered resource overlaps following dietary shifts. Our work implies that competitor recognition is common among heterospecifics, and disruption of this system could lead to suboptimal decision-making, exacerbating sublethal impacts of food scarcity.</p>
Dataset for Named Entity Recognition and Entity Linking from Greek Wikipedia Events
<p>An automated benchmark dataset for (Named Entity Recognition) NER and (Named Entity Linking) NEL tools, based on Greek Wikipedia events pages.</p> <p>Note: This data includes data from the following sources:<br> - Wikipedia el.wikipedia.org</p> <p><strong>Description</strong></p> <p>The dataset is provided in the form of three JSON-formatted subsets i.e., train, validation and test in an analogy of 70-20-10. The current version of the dataset contains 18,617 events annotated with 40,798 entity mentions and 36,189 links to elWikipedia (and wikidata ids). The dataset contains annotations belonging to 8 entity types: person, organization, location, gpe, event, facility, product and work of art.</p> <table> <caption>Overall dataset statistics</caption> <thead> <tr> <th scope="col"> </th> <th scope="col">Docs</th> <th scope="col">Tokens</th> <th scope="col">Sentences</th> <th scope="col">Surface Mentions</th> <th scope="col">Valid Links</th> <th scope="col">Red Links</th> </tr> </thead> <tbody> <tr> <td><strong>Train</strong></td> <td>13,031</td> <td>332,077</td> <td>16,927</td> <td>28,593</td> <td>25,365</td> <td>3,228</td> </tr> <tr> <td><strong>Validation</strong></td> <td>3,722</td> <td>94,746</td> <td>4,844</td> <td>8,168</td> <td>7,240</td> <td>928</td> </tr> <tr> <td><strong>Test</strong></td> <td>1,862</td> <td>47,450</td> <td>2,427</td> <td>4,037</td> <td>3,584</td> <td>453</td> </tr> <tr> <td><strong>Total</strong></td> <td>18,617</td> <td>474,361</td> <td>24,200</td> <td>40,798</td> <td>36,189</td> <td>4,609</td> </tr> </tbody> </table> <p><strong>Example</strong></p> <p>A record example is given below.</p> <p>{</p> <p>"json_file": "February 2012_39_0 events",<br> "text": "Sudan and South Sudan sign non-aggression pact.",<br> "ground_truth_mentions": [<br> {"start": 0, "end": 4, "surface_mention": "Sudan", "mention_type": "GPE"},<br> {"start": 10, "end": 20, "surface_mention": "South Sudan", "mention_type": "GPE"}<br> ],<br> "ground_truth_links": [<br> {"enwiki": "Sudan","wikidata": "Q1049"},<br> {"enwiki": "South_Sudan", "wikidata": "Q958"}<br> ]<br> }</p> <p><strong>Code</strong></p> <p><a href="https://gitlab.isl.ics.forth.gr/debatelab/elwiki_events_benchmark">https://gitlab.isl.ics.forth.gr/debatelab/elwiki_events_benchmark</a></p> <p><strong>Acknowledgments</strong></p> <p>This work has received funding from the Hellenic Foundation for Research and Innovation (HFRI) and the General Secretariat for Research and Technology (GSRT), under grant agreement No 4195.</p>
Not so weak-PICO: Leveraging weak supervision for Participants, Interventions, and Outcomes recognition for systematic review automation
<p><strong>Objective: </strong>PICO (Participants, Interventions, Comparators, Outcomes) analysis is vital but time-consuming for conducting systematic reviews (SRs). Supervised machine learning can help fully automate it, but a lack of large annotated corpora limits the quality of automated PICO recognition systems. The largest currently available PICO corpus is manually annotated, which is an approach that is often too expensive for the scientific community to apply. Depending on the specific SR question, PICO criteria are extended to PICOC (C-Context), PICOT (T-timeframe), and PIBOSO (B-Background, S-Study design, O-Other) meaning the static hand-labelled corpora need to undergo costly re-annotation as per the downstream requirements. We aim to test the feasibility of designing a weak supervision system to extract these entities without hand-labelled data.</p> <p><strong>Methodology:</strong> We decompose PICO spans into its constituent entities and re-purpose multiple medical and non-medical ontologies and expert-generated rules to obtain multiple noisy labels for these entities. These labels obtained using several sources are then aggregated using simple majority voting and generative modelling approaches. The resulting programmatic labels are used as weak signals to train a weakly-supervised discriminative model and observe performance changes. We explore mistakes in the currently available PICO corpus that could have led to inaccurate evaluation of several automation methods.</p> <p><strong>Results: </strong>We present Weak-PICO, a weakly-supervised PICO entity recognition approach using medical and non-medical ontologies, dictionaries and expert-generated rules. Our approach does not use hand-labelled data.</p> <p><strong>Conclusion: </strong>Weak supervision using weak-PICO for PICO entity recognition has encouraging results, and the approach can potentially extend to more clinical entities readily. </p>
Data for: High-throughput profiling of sequence recognition by tyrosine kinases and SH2 domains using bacterial peptide display
<p>Tyrosine kinases and SH2 (phosphotyrosine recognition) domains have binding specificities that depend on the amino acid sequence surrounding the target (phospho)tyrosine residue. Although the preferred recognition motifs of many kinases and SH2 domains are known, we lack a quantitative description of sequence specificity that could guide predictions about signaling pathways or be used to design sequences for biomedical applications. Here, we present a platform that combines genetically-encoded peptide libraries and deep sequencing to profile sequence recognition by tyrosine kinases and SH2 domains. We screened several tyrosine kinases against a million-peptide random library and used the resulting profiles to design high-activity sequences. We also screened several kinases against a library containing thousands of human proteome-derived peptides and their naturally-occurring variants. These screens recapitulated independently measured phosphorylation rates and revealed hundreds of phosphosite-proximal mutations that impact phosphosite recognition by tyrosine kinases. We extended this platform to the analysis of SH2 domains and showed that screens could predict relative binding affinities. Finally, we expanded our method to assess the impact of non-canonical and post-translationally modified amino acids on sequence recognition. This specificity profiling platform will shed new light on phosphotyrosine signaling and could readily be adapted to other protein modification/recognition domains.</p>
FIG. 7 in A synopsis of Ptisana Murdock ferns (Marattiaceae) in New Caledonia based on sequence data and morphology with the recognition of a new vulnerable species, P. soluta (Compton) Murdock & Perrie, comb. nov., stat. nov.
FIG. 7. — Median-joining networks based on trnSGG and rps4-trnS sequences: A, the Ptisana attenuata clade; B, the P. salicina/P. soluta comb. nov., stat. nov./P. smithii clade. The size of each circle is proportional to the haplotype frequency. Undetected intermediate haplotypes on nodes are shown as black circles and hatch marks represent mutational steps separating haplotypes.
FIG. 6 in A synopsis of Ptisana Murdock ferns (Marattiaceae) in New Caledonia based on sequence data and morphology with the recognition of a new vulnerable species, P. soluta (Compton) Murdock & Perrie, comb. nov., stat. nov.
FIG. 6. — Phylogram from the Bayesian phylogenetic analysis of the chloroplast DNA sequence data for Ptisana Murdock. Support values for branches are given in the order of Bayesian inference posterior probability; maximum parsimony bootstrap support; and maximum likelihood bootstrap support. Only values>0.80 PP and 60% BS are shown.
FIG. 4 in A synopsis of Ptisana Murdock ferns (Marattiaceae) in New Caledonia based on sequence data and morphology with the recognition of a new vulnerable species, P. soluta (Compton) Murdock & Perrie, comb. nov., stat. nov.
FIG. 4. — Distribution map for the New Caledonian endemic species of Ptisana attenuata (Labill.) Murdock (), P. rolandi-principis (Rosenst.) Christenh. (Δ), and P. soluta (Compton) Murdock & Perrie, comb. nov., stat. nov. (, with unvouchered field observations indicated by a broken outline). Shaded areas are ultramafic substrates. The collecting sites of the sequenced P. attenuata samples are indicated.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.