Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
47
datasets available to search
ShareScore release 0.9.0
Dataset results
47 results for “structural annotation”
Cross-phyla protein annotation by structural prediction and alignment
<p><strong>Background:</strong> Protein annotation is a major goal in molecular biology, yet experimentally determined knowledge is typically limited to a few model organisms. In non-model species, the sequence-based prediction of gene orthology can be used to infer protein identity, however this approach loses predictive power at longer evolutionary distances. Here we propose a workflow for protein annotation using structural similarity, exploiting the fact that similar protein structures often reflect homology and are more conserved than protein sequences.</p> <p><strong>Results:</strong> We propose a workflow of openly available tools for the functional annotation of proteins via structural similarity (MorF: <strong>Mor</strong>pholog<strong>F</strong>inder) and use it to annotate the complete proteome of a sponge. Sponges are highly relevant for inferring the early history of animals, yet their proteomes remain sparsely annotated. MorF accurately predicts the functions of proteins with known homology in >90% cases, and annotates an additional 50% of the proteome beyond standard sequence-based methods. We uncover new functions for sponge cell types, including extensive FGF, TGF and Ephrin signalling in sponge epithelia, and redox metabolism and control in myopeptidocytes. Notably, we also annotate genes specific to the enigmatic sponge mesocytes, proposing they function to digest cell walls.</p> <p><strong>Conclusions:</strong> Our work demonstrates that structural similarity is a powerful approach that complements and extends sequence similarity searches to identify homologous proteins over long evolutionary distances. We anticipate this to be a powerful approach that boosts discovery in numerous -omics datasets, especially for non-model organisms.</p>
VoroCrack3d: An annotated data set of 3d CT concrete images with synthetic crack structures
<p>VoroCrack3d is an annotated data set of 3d CT images of concrete with synthetic crack structures. Its main purpose is the training and testing of machine learning models for 3d crack segmentation. The data set comprises 1344 images together with their corresponding ground truths. The concrete backgrounds are cropped out sections of size 400x400x400 voxels of CT images of concrete. To this end, several different concrete samples were scanned (normal concrete (NC), high-performance concrete (HPC), ultra-high-performance concrete (UHPC), air pore concrete; without and with reinforcements (straight steel fibers, crimped steel fibers, hooked-end steel fibers, polypropylene fibers, fibers made of glass fiber-reinforced polymer). The original concrete images have a resolution between 2.8 and 106 micrometers.</p> <p>The crack structures are modeled via minimum-weight surfaces in Voronoi diagrams according to the paper</p> <p>[1] C. Jung, C. Redenbach, Crack Modeling via Minimum-Weight Surfaces in 3d Voronoi Diagrams, Journal of Mathematics in Industry, 13, 10 (2023). https://doi.org/10.1186/s13362-023-00138-1.</p> <p>The surfaces are discretized, dilated and superimposed on the concrete backgrounds.</p> <p>The data set offers a high variety regarding concrete types, noise levels and crack widths, shapes, regularity and branching. This makes it suitable for studying the generalizability and robustness of 3d crack segmentation methods.</p> <p>______________________________________________________________________________________________</p> <p>The folder 'data' contains seven subfolders, each containing the data generated from a specific concrete type (NC, HPC, air pore concrete, polypropylene fiber-reinforced concrete, steel fiber-reinforced concrete (straight, crimped and hooked-end steel fibers)).</p> <p>Each subfolder again contains four subfolders according to the point process model that was used for generating the 3d Voronoi diagrams. The point processes and Voronoi diagrams are restricted to windows of size 400x150x400. </p> <p>- 'hc': Hard core point process with 60% volume density and intensity 0.000025 obtained from force-biased sphere packing.<br>- 'matclust': Matérn cluster process with parent intensity 0.0002/50, offspring intensity 50 and cluster radius 20.<br>- 'ppp': Poisson point process with intensity 0.0002.<br>- 'ppp-scaled': Poisson point process with intensity 0.0002 (but inside 200x150x200 window). The resulting Voronoi diagram is stretched in x- and z- direction by a factor of 2.</p> <p>Each of these contains five subfolders: one for the 3d input images, two for the corresponding labels (ground truths; one with and one without pores/fibers), one for the input and label previews (slice z=200 for each of the images) and a misc folder containing the concrete background without crack and, if applicable, the pore/fiber segmentation image.</p> <p>The data itself then contains 48 images:<br>1a-1d: crack with up to seven branches; fixed crack width (~1 voxel).<br>2a-2d: crack with up to four branches; fixed crack width (~1 voxel).<br>3a-3d: crack with up to one branch; fixed crack width (~1 voxel).<br>4a-4d: crack with no branches; fixed crack width (~1 voxel).<br>5a-5d: crack with no branches; fixed crack width (~3 voxels).<br>6a-6d: crack with no branches; fixed crack width (~5 voxels).<br>7a-7d: crack with no branches; fixed crack width (~7 voxels).<br>8a-8d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.01);<br>9a-9d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.02);<br>10a-10d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.05);<br>11a-11d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.1);<br>12a-12d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.2);</p> <p>The names 'a'-'d' indicate level of added noise added to the image:<br>a: None.<br>b: Uniformly on [-sigma,sigma] <br>c: Uniformly on [-2*sigma,2*sigma] <br>d: Uniformly on [-4*sigma,4*sigma] <br>Negative values are mapped to 0. <br>For inputs of type int, noise values are rounded to the nearest integer.<br>(sigma = standard deviation of voxel greyvalues in image)</p> <p>Note that the grey values in the ground truths correspond to the local crack width. They can be thresholded to obtain binary masks.</p> <p>For more details, we refer to [1].</p>
Structure Annotations of Assessment and Plan Sections from MIMIC-III
<p>Physicians record their detailed thought-processes about diagnoses and treatments as unstructured text in a section of a clinical note called the "assessment and plan". This information is more clinically rich than structured billing codes assigned for an encounter but harder to reliably extract given the complexity of clinical language and documentation habits. To structure these sections we collected a dataset of annotations over assessment and plan sections from the publicly available and de-identified MIMIC-III dataset, and developed deep-learning based models to perform this task, described in the associated paper available as a pre-print at: <a href="https://www.medrxiv.org/content/10.1101/2022.04.13.22273438v1">https://www.medrxiv.org/content/10.1101/2022.04.13.22273438v1</a></p> <p>When using this data please cite our paper:</p> <pre><code>@article {Stupp2022.04.13.22273438, author = {Stupp, Doron and Barequet, Ronnie and Lee, I-Ching and Oren, Eyal and Feder, Amir and Benjamini, Ayelet and Hassidim, Avinatan and Matias, Yossi and Ofek, Eran and Rajkomar, Alvin}, title = {Structured Understanding of Assessment and Plans in Clinical Documentation}, year = {2022}, doi = {10.1101/2022.04.13.22273438}, publisher = {Cold Spring Harbor Laboratory Press}, URL = {https://www.medrxiv.org/content/early/2022/04/17/2022.04.13.22273438}, journal = {medRxiv} }</code></pre> <p>The dataset, presented here, contains annotations of assessment and plan sections of notes from the publicly available and de-identified MIMIC-III dataset, marking the active problems, their assessment description, and plan action items. Action items are additionally marked as one of 8 categories (listed below). The dataset contains over 30,000 annotations of 579 notes from distinct patients, annotated by 6 medical residents and students. </p> <p>The dataset is divided into 4 partitions - a training set (481 notes), validation set (50 notes), test set (48 notes) and an inter-rater set. The inter-rater set contains the annotations of each of the raters over the test set. Rater 1 in the inter-rater set should be regarded as an intra-rater comparison (details in the paper). The labels underwent automatic normalization to capture entire word boundaries and remove flanking non-alphanumeric characters.</p> <p>Code for transforming labels into TensorFlow examples and training models as described in the paper will be made available at GitHub: <a href="https://github.com/google-research/google-research/tree/master/assessment_plan_modeling">https://github.com/google-research/google-research/tree/master/assessment_plan_modeling</a></p> <p>In order to use these annotations, the user additionally needs to obtain the text of the notes which is found in the NOTE_EVENTS table from MIMIC-III, access to which is to be acquired independently (<a href="http://mimic.mit.edu">https://mimic.mit.edu/</a>)</p> <p>Annotations are given as character spans in a CSV file with the following schema:</p> <table> <tbody> <tr> <td>Field</td> <td>Type</td> <td>Semantics</td> </tr> <tr> <td>partition</td> <td>categorical (one of [train, val, test, interrater]</td> <td>The set of ratings the span belongs to.</td> </tr> <tr> <td>rater_id</td> <td>int</td> <td>Unique id for each the raters</td> </tr> <tr> <td>note_id</td> <td>int</td> <td>The note’s unique note_id, links to the MIMIC-III notes table (as ROW-ID).</td> </tr> <tr> <td>span_type</td> <td>categorical (one of [PROBLEM_TITLE,<br> PROBLEM_DESCRIPTION, ACTION_ITEM]</td> <td>Type of the span as annotated by raters.</td> </tr> <tr> <td>char_start</td> <td>int</td> <td>Character offsets from note start</td> </tr> <tr> <td>char_end</td> <td>int</td> </tr> <tr> <td>action_item_type</td> <td>categorical (one of [MEDICATIONS, IMAGING, OBSERVATIONS_LABS, CONSULTS, NUTRITION, THERAPEUTIC_PROCEDURES, OTHER_DIAGNOSTIC_PROCEDURES, OTHER])</td> <td>Type of action item if the span is an action item (empty otherwise) as annotated by raters.</td> </tr> </tbody> </table>
PDEStrIAn: A phosphodiesterase structure and ligand interaction annotated database as a tool for structure-based drug design
<p>A systematic analysis is presented of the 220 phosphodiesterase (PDE) catalytic domain crystal structures present in the Protein Data Bank (PDB) with a focus on PDE-ligand interactions. The consistent structural alignment of 57 PDE ligand binding site residues enables the systematic analysis of PDE-ligand Interaction FingerPrints (IFPs), the identification of subtype-specific PDE-ligand interaction features, and the classification of ligands according to their binding modes. We illustrate how systematic mining of this phosphodiesterase structure and ligand interaction annotated (PDEStrIAn) database provides new insights into how conserved and selective PDE interaction hot spots can accommodate the large diversity of chemical scaffolds in PDE ligands. A substructure analysis of the co-crystalized PDE ligands in combination with those in the ChEMBL database provides a toolbox for scaffold hopping and ligand design. These analyses lead to an improved understanding of the structural requirements of PDE binding that will be useful in future drug discovery studies.</p>
Annotation dataset for the article titled "On the Emerging Supremacy of Structured Digital Data in Archaeology: A Preliminary Assessment of Information, Knowledge and Wisdom Left Behind"
<p>This is the resulting dataset from the text annotation exercise in the article titled "<strong>On the Emerging Supremacy of Structured Digital Data in Archaeology: A Preliminary Assessment of Information, Knowledge and Wisdom Left Behind</strong>" that will appear in the journal Open Archeology in a special issue titled Archaeological Practice on Shifting Grounds (edited by Åsa Berggren and Antonia Davidovic-Walther). The article is accepted for publication and the annotations are final. CIDOC CRM is used for text annotations.</p>
Dataset: "Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data"
<p>Dataset used in the experiments of the publication: "Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data" by Bach et al.</p> <p><strong>File description:</strong></p> <ul> <li> <p>cfmid4.tar: MS² spectra simulated using <a href="https://bitbucket.org/wishartlab/cfm-id-code/src/CFM-ID_4.0.7/">CFM-ID (v4.0.7)</a> for all molecular candidate structures</p> </li> <li> <p>db_layout.png: Visualization of the SQLite database (DB) layout</p> </li> <li> <p>massbank.sqlite.gz: DB containing all needed data to (re-)run the experiments shown in the paper. Please read "DB_README.md" for further details. The database file can be unpacked using gzip.</p> </li> <li> <p>metfrag.tar: MetFrag input files and MS² scores for all candidate sets computed using the <a href="https://ipb-halle.github.io/MetFrag/projects/metfragcl/">MetFrag software</a>.</p> </li> <li> <p>sirius_scores.tar: MS² scores for all candidates and measured spectra using the <a href="https://bio.informatik.uni-jena.de/software/sirius/">SIRIUS software</a>.</p> </li> <li> <p>sirius_inputs.tar: Input (ms-files) for the SIRIUS software.</p> </li> <li> <p>DB_README.md: Description of each table in the "massbank.sqlite" SQLite DB.</p> </li> <li> <p>db_processing_scripts.tar: Scripts to re-produce the "massbank.sqlite" and a README.md providing further information on the process.</p> </li> <li> <p>massbank__2020.11__v0.6.1.sqlite: Base SQLite DB from which the "massbank.sqlite" was build up. It was created using the "<a href="https://github.com/bachi55/massbank2db">massbank2db</a>" (v0.6.1) Python package using the <a href="https://github.com/bachi55/MassBank-data/tree/2020.11-branch">MassBank release 2020.11</a>.</p> </li> <li> <p>substructure_fingerprints.tar: Pre-computed substructure counting fingerprints for all candidates related to our experiments.</p> </li> </ul> <p><strong>Instructions:</strong></p> <p>The "massbank.sqlite" can be directly used with the Structure Support Vector Machine Model (SSVM) described in the manuscript and implemented in the "<a href="https://github.com/aalto-ics-kepaco/msms_rt_ssvm">ssvm</a>" Python package.</p> <p>If desired, the database can be re-produced using the scripts provided in "db_processing_scripts.tar":</p> <ol> <li>Create a directory for all data</li> <li>Download and extract the ... <ol> <li>Processing scripts</li> <li>MS² scorer outputs (e.g. metfrag.tar)</li> <li>Pre-computed substructure fingerprints</li> </ol> </li> <li>Follow the instructions given in the "README.md" of the "db_processing_scripts.tar"</li> </ol>
FAPM: Functional annotation of proteins using multi-modal models beyond structural modeling
<p>Assigning accurate property labels to proteins, like functional terms and catalytic activity, is challenging, especially for proteins without homologs and "tail labels" with few known examples. Unlike previous methods that mainly focused on protein sequence features, we use a pretrained large natural language model to understand the semantic meaning of protein labels. Specifically, we introduce FAPM, a contrastive multi-modal model that links natural language with protein sequence language. This model combines a pretrained protein sequence model with a pretrained large language model to generate labels, such as Gene Ontology (GO) functional terms and catalytic activity predictions, in natural language. Our results show that FAPM excels in understanding protein properties, outperforming models based solely on protein sequences or structures. It achieves state-of-the-art performance on public benchmarks and in-house experimentally annotated phage proteins, which often have few known homologs. Additionally, FAPM's flexibility allows it to incorporate extra text prompts, like taxonomy information, enhancing both its predictive performance and explainability. This novel approach offers a promising alternative to current methods that rely on multiple sequence alignment for protein annotation.</p>
Figure 4 in Riodinid butterfly fauna (Lepidoptera) of the Cosñipata Region, Peru: Annotated checklist, community structure, and contrast with Lycaenidae
Figure 4. Proportion of species recorded from January to the given month for Riodinidae (398 species) and Lycaenidae (342 species). Using a two-sample Kolmogorov–Smirnov test for cumulative distributions differences, D-stat = 0.101443, D-crit = 0.0992, p = 0.042.
Figure 3 in Riodinid butterfly fauna (Lepidoptera) of the Cosñipata Region, Peru: Annotated checklist, community structure, and contrast with Lycaenidae
Figure 3. Proportion of species recorded below a given elevation for Riodinidae (398 species) and Lycaenidae (342 species). Using a two-sample Kolmogorov-Smirnov test for cumulative distributions differences, D-stat = 0.16650504, D-crit = 0.099199506, p = 6.13109E-05.
FAPM: Functional annotation of proteins using multi-modal models beyond structural modeling
Open the record for dataset details and reuse information.
Dataset associated to "Annotation matters: the effect of structural gene annotation on orthology"
<p>Dataset including the input and output files in "Annotation matters: the effect of structural gene annotation on orthology".</p> <ul> <li>Input proteomes for OMA and their corresponding splice files are in the OMAproteomes zipped folder. The OMA results for each annotation method are in the zipped folders with the method name (e.g. Augustus.zip).</li> <li>The fasta files (proteomes) are the same for OrthoFinder input in the cases of UniProt and Augustus (as they only have one isoform per gene). In these cases, the OrthoFinder folders (e.g. OFUniProt.zip), include the OrthoFinder output for that proteomes set. For Ensembl and NCBI, given the different approach each orthology method follows, the specific orthofinder proteomes are also included in the OrthoFinder (OF) zipped folder (e.g. OFtopNCBI.zip), in their corresponding primary_transcripts subfolder. </li> <li>The folder GSTDBenchmarOutput.zip contains the results from the Generalized Species Tree Discordance Benchmark.</li> <li>topNCBI/topEnsembl correspond to the original proteomes downloaded from the databases.</li> <li>priNCBI/primEnsembl correspond to the proteomes sets which include only the genes found on the primary assembly (reference sequences).</li> <li>For the species code to species name correspondance, please check the Code-Species.csv file.</li> </ul>
Result files (ONLYSTEREO): "Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data"
<p>Result files associated with the publication: "<strong>Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data</strong>" by Bach et al.</p> <p>The following files are included in the archive:</p> <ul> <li>Raw max-marginal predictions using LC-MS²Struct for all LC-MS² experiments of the ONLYSTEREO setup</li> <li>Averaged max-marginals for the LC-MS²Struct over all SSVM models</li> <li>Ranks for the ground-truth structures predicted by Only MS² and LC-MS²Struct (molecule class analysis)</li> </ul> <p>Instructions:</p> <ul> <li>clone the repository containing the experimental scripts and analysis notebooks: <a href="https://github.com/aalto-ics-kepaco/lcms2struct_exp">https://github.com/aalto-ics-kepaco/lcms2struct_exp</a></li> <li>download the archive in this repository</li> <li>unpack the archive in the git-repository root directory</li> <li>follow the instructions given in the <a href="https://github.com/aalto-ics-kepaco/lcms2struct_exp/blob/main/README.md">README.md</a> of the git-repository to reproduce the figures, etc.</li> </ul>
Result files (ALLDATA): "Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data with LC-MS²Struct"
<p>Result files associated with the publication: "<strong>Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data</strong>" by Bach et al.</p> <p>The following files are included in the archive:</p> <ul> <li>Raw max-marginal predictions using LC-MS²Struct for all LC-MS² experiments of the ALLDATA setup</li> <li>Averaged max-marginals for the LC-MS²Struct over all SSVM models</li> <li>Top-k accuracies for the comparison methods (MS²+RT, ...)</li> <li>Ranks for the ground-truth structures predicted by Only MS² and LC-MS²Struct (molecule class analysis)</li> </ul> <p>Instructions:</p> <ul> <li>clone the repository containing the experimental scripts and analysis notebooks: <a href="https://github.com/aalto-ics-kepaco/lcms2struct_exp">https://github.com/aalto-ics-kepaco/lcms2struct_exp</a></li> <li>download the archive in this repository</li> <li>unpack the archive in the git-repository root directory</li> <li>follow the instructions given in the <a href="https://github.com/aalto-ics-kepaco/lcms2struct_exp/blob/main/README.md">README.md</a> of the git-repository to reproduce the figures, etc.</li> </ul> <p>Version history:</p> <ul> <li><strong>Version 1</strong>: Experimental results for "Method comparison" and "Molecule classification analysis" where performed with <strong>2D fingerprints</strong> (<a href="https://www.biorxiv.org/content/10.1101/2022.02.11.480137v1">preprint v1</a>)</li> <li><strong>Version 2</strong> <em>(this version)</em>: Experimental results for "Method comparison" and "Molecule classification analysis" where performed with <strong>3D fingerprints</strong></li> </ul>
FURNA: a database of functional annotations of RNA structures (part 2)
<p>A copy of the FURNA database curated on June 9, 2024 (part 2). Download both part 1 (10.5281/zenodo.11664059) and part 2 (10.5281/zenodo.11672037) to combine them by:</p> <p>$ cat xaa xab > furna.tar.bz2</p> <p>$ rm xaa xab</p> <p>$ tar -xvf furna.tar.bz2</p>
FURNA: a database of functional annotations of RNA structures (part 1)
<p>A copy of the FURNA database curated on June 9, 2024 (part 1). Download both part 1 (10.5281/zenodo.11664059) and part 2 (10.5281/zenodo.11672037) to combine them by:</p> <p>$ cat xaa xab > furna.tar.bz2</p> <p>$ rm xaa xab</p> <p>$ tar -xvf furna.tar.bz2</p>
Reducing the structure bias of RNA-Seq reveals a large number of non-annotated non-coding RNA Data Archive
<p>[This repository contains the source data for the workflow presented in the manuscript "<strong>Reducing the structure bias of RNA-Seq reveals a large number of non-annotated non-coding RNA</strong>". The workflow can be found here: http://gitlabscottgroup.med.usherbrooke.ca/gaspard/snakemake_blockbuster ]</p> <p>The study of RNA expression is the fastest growing area of genomic research. However, despite the dramatic increase in the number of sequenced transcriptomes, we still do not have accurate estimates of the number and expression levels of non-coding RNA genes. Non-coding transcripts are often overlooked due to incomplete genome annotation. In this study, we use annotation-independent detection of RNA reads generated using a reverse transcriptase with low structure bias to identify non-coding RNA. Transcripts between 20 and 500 nucleotides were filtered and crosschecked with non-coding RNA annotations revealing 115 non-annotated non-coding RNAs expressed in different cell lines and tissues. Inspecting the sequence and structural features of these transcripts indicated that 60% of these transcripts correspond to new tRNA and snoRNA genes. The identified genes exhibited features of their respective families in terms of structure, expression, conservation and response to depletion of interacting proteins. Together, our data reveal a new group of RNA that are difficult to detect using standard gene prediction and RNA sequencing techniques, suggesting that reliance on actual gene annotation and sequencing techniques distort the perceived architecture of the human transcriptome.</p>
Figure 1 in Riodinid butterfly fauna (Lepidoptera) of the Cosñipata Region, Peru: Annotated checklist, community structure, and contrast with Lycaenidae
Figure 1. Location of the Cosñipata Region (yellow box) in southeast Peru. © Amazonia Lodge.
Figure 2 in Riodinid butterfly fauna (Lepidoptera) of the Cosñipata Region, Peru: Annotated checklist, community structure, and contrast with Lycaenidae
Figure 2. The Cosñipata Valley at 1,200 m looking towards the northeast. © Loran D. Gibson.
FIGURES 54–71. Morphological structures. 54—tergite 7 in An annotated checklist and key to the Bulgarian cockroaches (Dictyoptera: Blattodea)
FIGURES 54–71. Morphological structures. 54—tergite 7 with glandular pit of Ectobius balcani (BG, Alibotoush Mt., Livade place); 55—tergite 7 with glandular pit of E. burri (BG, Zemen gorge, Zemen town); 56—tergite 7 with glandular pit of E. erythronotus erythronotus (BG, Lozen Mt., Lozen village); 57—tergite 7 with glandular pit of E. lapponicus (BG, Western Predbalkan range, Montana town); 58—tergite 7 with glandular pit of E. sylvestris (BG, Pirin Mts., Bansko town); 59—tergite 7 with glandular pit Ectobius vittiventris (BG, Strandzha Mt., Byala voda village); 60—tergite 7 with glandular pit of E. punctatissimus (Montenegro, Durmitor Mt., Tara gorge); 61—tergite 7 with glandular pit of Phyllodromica brevipennis (BG, Vitosha Mt., Yarlovo village); 62—tergite 7 with glandular pit of Ph. carniolica (BG, Vitosha Mt., Bosnek village); 63—tergite 7 with glandular pit of Ph. pallida (R Macedonia, Plačkovica Mt., Kozbunar village); 64—tergite 7 with glandular pit of Ph. marginata (BG, Eastern Rhodope Mts., Madzharovo town); 65—tergite 7 with glandular pit of Ph. pulcherrima (BG, Chepan Mt., Dragoman village); 66—helmet sclerite of E. sylvestris (BG, Pirin Mts., Bansko town); 67—helmet sclerite of Ph. brevipennis (BG, Vitosha Mt., Yarlovo village); 68—helmet sclerite of Ph. carniolica (BG, Vitosha Mt., Bosnek village); 69— helmet sclerite of Ph. pallida (R Macedonia, Plačkovica Mt., Kozbunar village); 70—tegmina of male Ph. carniolica (BG, Western Rhodope Mts., Ribnovo village); 71—tegmina of male Ph. pallida (Albania, Galičica Mt., Dolna Gorica village). Scale (figures 54–65) = 500 µm; Scale (figures 66–69) = 200 µm; Scale (figures 70, 71) = 2 mm.
FIGURE 9 in The reef fish assemblage of the Laje de Santos Marine State Park, Southwestern Atlantic: annotated checklist with comments on abundance, distribution, trophic structure, symbiotic associations, and conservation
FIGURE 9. Selected examples of symbiotic associations between reef fishes recorded at the Laje de Santos Marine State Park. The barber goby Elacatinus figaro cleans the head of the jubauna reeffish Chromis jubauna hovering close to the goby's cleaning station (a); the same cleaner species inspects the back of the nocturnal squirrelfish Holocentrus adscensionis that approached its cleaning station (b); juvenile spotfin hogfish Bodianus pulchellus cleans the mouth of the spotted moray Gymnothorax moringa (c); adult of the same hogfish species cleans the head of the jubauna reeffish (d); the wrasse Halichoeres sp. n. follows a group of the white trevally Pseudocaranx dentex, which stir sediment clouds while feeding on the sandy bottom (e); the dusky grouper Mycteroperca marginata closely follows the goldspotted snake eel Myrichthys ocellatus that nudges its head in rocky crevices (f); the spotfin hogfish follows the flying gurnard Dactylopterus volitans moving close to the bottom (g); two diskfish Remora remora attached near the mouth of the Atlantic manta Manta birostris (h).Photos: M. Andrade (h); A. Carvalho Filho (e-f); J. P. Krajewski (a); O.J. Luiz Jr. (c-d, g); A. de Luca Jr. (b).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.