Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

36

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

36 results for “Structural classification”

Learn how ShareScore rates datasets ↗
zenodo44/100

Taxnonomic classifications for all structure in the QM9 dataset

<p>The classification of molecules according to ClassyFire [1] for the QM9 dataset [2].</p> <p>The QM9 dataset&nbsp;is a set of nearly 140k organic molecules with no more than 9 C, N, O, and F atoms optimized to a stable structure with DFT.&nbsp;</p> <p>ClassyFire is a tool and taxonomic library for the labeling of molecules.</p> <p>1. Djoumbou Feunang, Y. <em>et al.</em> ClassyFire: automated chemical classification with a comprehensive, computable taxonomy. <em>J. Cheminform.</em> <strong>8</strong>, 1&ndash;20 (2016).</p> <p>2. Ramakrishnan, R., Dral, P. O., Rupp, M. &amp; Von Lilienfeld, O. A. Quantum chemistry structures and properties of 134 kilo molecules. <em>Sci. Data</em> <strong>1</strong>, 1&ndash;7 (2014).</p> <p>&nbsp;</p> <p>The data directory (&#39;QM9_jsons_classified.tar.gz&#39;) contains a `json` file for each structure in the QM9 dataset. The name of the file is the same identifier as from QM9. Data fields include:</p> <p>- `cf_alternative_parents` : classifications describing the compound that do not fall in the given ancestry</p> <p>- `cf_ancestors` :&nbsp;classes along the taxonomic branch&nbsp;for the structure&nbsp;</p> <p>- `cf_class` : ClassyFire given class</p> <p>- `cf_superclass` : ClassyFire given super class</p> <p>- `cf_subclass` : ClassyFire given subclass</p> <p>- `cf_direct_parent` : Class one level above this structure on the taxonomic branch</p> <p>- `cf_description` : Exposition on the given class</p> <p>- `cf_identifier` : identifier for the structure in the ClassyFire database</p> <p>- `cf_intermediate_nodes` : classes connecting branches on taxonomic tree</p> <p>- `cf_kingdom` : ClassyFire given kingdom</p> <p>- `cf_molecular_framework` : describes aromaticity and number of cycles</p> <p>- `cf_predicted_chebi_terms` : terms describing the molecule in the ChEBI framework&nbsp; &nbsp;</p> <p>- `cf_predicted_lipidmaps_terms` : terms describing the molecule in LIPID MAPS framework</p> <p>- `cf_smiles` : smiles string given by ClassyFire</p> <p>- `cf_substituents` : substituent groups in the structure&nbsp;</p> <p>&nbsp;&nbsp;</p> <p>Many fields contain subfields, seen in the example below for molecule with QM9 id 000123:</p> <p>{&quot;cf_alternative_parents&quot;:[{&quot;name&quot;:&quot;Dialkylamines&quot;,&quot;description&quot;:&quot;Organic compounds containing a dialkylamine group, characterized by two alkyl groups bonded to the amino nitrogen.&quot;,&quot;chemont_id&quot;:&quot;CHEMONTID:0002228&quot;,&quot;url&quot;:&quot;http:\/\/classyfire.wishartlab.com\/tax_nodes\/C0002228&quot;},{&quot;name&quot;:&quot;Organopnictogen compounds&quot;,&quot;description&quot;:&quot;Compounds containing a bond between carbon a pnictogen atom. Pnictogens are p-block element atoms that are in the group 15 of the periodic table.&quot;,&quot;chemont_id&quot;:&quot;CHEMONTID:0004557&quot;,&quot;url&quot;:&quot;http:\/\/classyfire.wishartlab.com\/tax_nodes\/C0004557&quot;},{&quot;name&quot;:&quot;Hydrocarbon derivatives&quot;,&quot;description&quot;:&quot;Derivatives of hydrocarbons obtained by substituting one or more carbon atoms by an heteroatom. They contain at least one carbon atom and heteroatom.&quot;,&quot;chemont_id&quot;:&quot;CHEMONTID:0004150&quot;,&quot;url&quot;:&quot;http:\/\/classyfire.wishartlab.com\/tax_nodes\/C0004150&quot;}],&quot;cf_ancestors&quot;:[&quot;Alpha-aminonitriles&quot;,&quot;Amines&quot;,&quot;Chemical entities&quot;,&quot;Dialkylamines&quot;,&quot;Hydrocarbon derivatives&quot;,&quot;Nitriles&quot;,&quot;Organic compounds&quot;,&quot;Organic cyanides&quot;,&quot;Organic nitrogen compounds&quot;,&quot;Organonitrogen compounds&quot;,&quot;Organopnictogen compounds&quot;,&quot;Secondary amines&quot;],&quot;cf_class&quot;:&quot;Organonitrogen compounds&quot;,&quot;cf_classification_version&quot;:&quot;2.1&quot;,&quot;cf_description&quot;:&quot;This compound belongs to the class of organic compounds known as alpha-aminonitriles. These are organonitrogen compounds that contain an amino group located on the carbon at the position alpha to a carbonitrile group. &nbsp;They have the general formula RC(NH2)C#N, where the amine group can be substituted.&quot;,&quot;cf_direct_parent&quot;:{&quot;name&quot;:&quot;Alpha-aminonitriles&quot;,&quot;description&quot;:&quot;Organonitrogen compounds that contain an amino group located on the carbon at the position alpha to a carbonitrile group. &nbsp;They have the general formula RC(NH2)C#N, where the amine group can be substituted.&quot;,&quot;chemont_id&quot;:&quot;CHEMONTID:0004453&quot;,&quot;url&quot;:&quot;http:\/\/classyfire.wishartlab.com\/tax_nodes\/C0004453&quot;},&quot;cf_external_descriptors&quot;:[],&quot;cf_identifier&quot;:&quot;Q5198051-1&quot;,&quot;cf_inchikey&quot;:&quot;InChIKey=PVVRRUUMHFWFQV-UHFFFAOYSA-N&quot;,&quot;cf_intermediate_nodes&quot;:[{&quot;name&quot;:&quot;Nitriles&quot;,&quot;description&quot;:&quot;Compounds having the structure RC#N; thus C-substituted derivatives of hydrocyanic acid, HC#N.&quot;,&quot;chemont_id&quot;:&quot;CHEMONTID:0000362&quot;,&quot;url&quot;:&quot;http:\/\/classyfire.wishartlab.com\/tax_nodes\/C0000362&quot;}],&quot;cf_kingdom&quot;:&quot;Organic compounds&quot;,&quot;cf_molecular_framework&quot;:&quot;Aliphatic acyclic compounds&quot;,&quot;cf_predicted_chebi_terms&quot;:[&quot;chemical entity (CHEBI:24431)&quot;,&quot;organic molecular entity (CHEBI:50860)&quot;,&quot;organonitrogen compound (CHEBI:35352)&quot;,&quot;secondary amino compound (CHEBI:50995)&quot;,&quot;nitrile (CHEBI:18379)&quot;,&quot;amine (CHEBI:32952)&quot;,&quot;secondary amine (CHEBI:32863)&quot;,&quot;cyanides (CHEBI:23424)&quot;,&quot;organic molecule (CHEBI:72695)&quot;,&quot;pnictogen molecular entity (CHEBI:33302)&quot;,&quot;nitrogen molecular entity (CHEBI:51143)&quot;],&quot;cf_predicted_lipidmaps_terms&quot;:[],&quot;cf_smiles&quot;:&quot;CNCC#N&quot;,&quot;cf_subclass&quot;:&quot;Organic cyanides&quot;,&quot;cf_substituents&quot;:[&quot;Alpha-aminonitrile&quot;,&quot;Secondary amine&quot;,&quot;Secondary aliphatic amine&quot;,&quot;Organopnictogen compound&quot;,&quot;Hydrocarbon derivative&quot;,&quot;Amine&quot;,&quot;Aliphatic acyclic compound&quot;],&quot;cf_superclass&quot;:&quot;Organic nitrogen compounds&quot;}</p> <p>&nbsp;</p> <p>A visualization &#39;&#39;qm9_pie_labeled.png&quot; is given of a fracturization of superclasses within qm9 down to subclass.</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Classification of Matching Molecular Series on the Basis of SAR Phenotypes and Structural Relationships

<p>A database comprising a total of 13,236 pairs of MMS&nbsp;with different SAR characteristics is provided. For each pair the corresponding MMS-cores are provided &nbsp;as SMILES. In addition, for each MMS-core&nbsp;the number of compounds and the SAR phenotype are given.&nbsp;ChEMBL target IDs (CHEMBLID_Target) designate target sets from which the MMS pairs originate. &nbsp;</p>

opencc-zeroJan 2016View details →
zenodo40/100

→ Fig. 10. FESEM images of the test structure in lagenid foraminifers from Recent, Admiralty Bay, King George Island, West Antarctica (A) and from the Jurassic of Gnaszyn, Poland (B, C). A. Unilocular Procerolagena gracilis Williamson, 1848, MWGUW ZI/67/44/02. B. Unilocular Lagena globosa Montagu, 1803, MWGUW ZI/67/61/09. C. Uniserial Nodosaria pulchra Franke, 1936, MWGUW ZI/67/61/26. Oblique cross-sectional views (A1, A2, A4, B1, B2, C); transverse cross-sectional views, showing single-crystal interlocked bundle structures, inner pores which extend along the entire length of the bundles as well as prominent calcite cleavage (A3, B3). Abbreviations: c, prominent calcite cleavage; ip, inner pore. in Chamber arrangement versus wall structure in the high-rank phylogenetic classification of Foraminifera

→ Fig. 10. FESEM images of the test structure in lagenid foraminifers from Recent, Admiralty Bay, King George Island, West Antarctica (A) and from the Jurassic of Gnaszyn, Poland (B, C). A. Unilocular Procerolagena gracilis Williamson, 1848, MWGUW ZI/67/44/02. B. Unilocular Lagena globosa Montagu, 1803, MWGUW ZI/67/61/09. C. Uniserial Nodosaria pulchra Franke, 1936, MWGUW ZI/67/61/26. Oblique cross-sectional views (A1, A2, A4, B1, B2, C); transverse cross-sectional views, showing single-crystal interlocked bundle structures, inner pores which extend along the entire length of the bundles as well as prominent calcite cleavage (A3, B3). Abbreviations: c, prominent calcite cleavage; ip, inner pore.

opencc-by-4.0Jan 2019View details →
zenodo40/100

Fig. 9 in Chamber arrangement versus wall structure in the high-rank phylogenetic classification of Foraminifera

Fig. 9. FESEM images of "monocrystalline" test structure in Spirillinata → foraminifers from the Jurassic of Gnaszyn, Poland (A) and Recent from Ronsard Bay, Western Australia (B). A. Paalzowella pazdroe Bielecka and Styk, 1969, MWGUW ZI/67/61/27; view of the test cross-section (A1); significantly magnified view of the test cross-section (A2, A4, A5); oblique cross-sectional view of the test showing "monocrystalline" test structure (A3); oblique cross sections of the test showing test composed of a few layers (A6, A7). B. Patellina sp., MWGUW ZI/67/61/22; oblique cross sections of the test showing prominent calcite cleavage (B1, B2).

opencc-by-4.0Jan 2019View details →
zenodo40/100

Fig. 8 in Chamber arrangement versus wall structure in the high-rank phylogenetic classification of Foraminifera

Fig. 8. FESEM images of the test structure in Tubothalamea from the Jurassic of Gnaszyn, Poland. A. Ophthalmidium carinatum Pazdro, 1958, MWGUW → ZI/67/08/5.03; front views of the abraded test surface, showing the extrados and porcelain (A1, A2). B.?Cornuspira radiata (Terquem, 1886), MWGUW ZI/67/55/11; front view of the test surface (B1, B3); oblique view of the test cross section, showing the test as being entirely composed of needle-shaped crystallites (B2, B4). C. Planiinvoluta sp., MWGUW ZI/67/57/13; view of the inner test surface (C1); side view of the test cross section, showing irregular meshwork of needle-shaped crystallites (C2). Abbreviations: e, extrados; p, porcelain.

opencc-by-4.0Jan 2019View details →
zenodo40/100

Fig. 7 in Chamber arrangement versus wall structure in the high-rank phylogenetic classification of Foraminifera

Fig. 7. FESEM images of the test structure in Recent calcareous cemented agglutinated textulariid (Globothalamea; A, B) and miliolid (Tubothalamea; C) → foraminifers from Ronsard Bay, Western Australia. A. Textularia sp., MWGUW ZI/67/55/24, front view of the test, showing agglutinated grains and the calcareous nanogranular matrix (A1); details of test wall (A2, A3). B. Gaudryina sp., MWGUW ZI/67/61/16, front view of the test, showing agglutinated grains and the matrix (B1); details of nanogranular matrix (B2, B3). C. Quinqueloculina arenata Said, 1949, MWGUW ZI/67/57/02, front view of the test (C1); oblique cross-sectional view of test showing foreign particle partially embedded in the irregular meshwork of needle-shaped crystallites (C2). Abbreviations: g, foreign particle; m, calcareous matrix. Arrows indicate pores.

opencc-by-4.0Jan 2019View details →
zenodo40/100

Figure 1. Brain Structure-Classification of Human Emotion from Deap EEG Signal Using Hybrid Improved Neural Networks with Cuckoo Search

<p>EEG data have collected from<br> desirable subjects. Each and every EEG signal has different kind of bands like Alpha, Beta,<br> Gamma, Theta, and Delta. Each band stores the particular information about the emotions. Alpha<br> band (8-13 Hz) which located in Frontal Occipital, Beta band (13-30 Hz) which located in Frontal<br> Central, Gamma band (30-100 Hz), Theta band (4- 7 Hz) which located in Midline Temp, Delta<br> band (0-4Hz) which located in Frontal Lobe. Before processing the EEG signal and extracting these<br> bands, preprocess the signal and reduce the noise. The basic brain figure is shown in below.</p>

opencc-by-4.0Aug 2015View details →
zenodo40/100

DPCstruct Classification of AlphaFold2-Predicted Protein Structures

<p>This dataset contains DPCstruct domain classifications for protein structures predicted by AlphaFold2, as presented in the paper "Unsupervised Domain Classification of AlphaFold2-Predicted Protein Structures."</p> <p>DPCstruct was applied to a non-redundant set of the AlphaFold Database v4.0, known as Foldseek Clusters, which includes approximately 15 million representative proteins, as described in the work by <a href="https://doi.org/10.1038/s41586-023-06510-w">Barrio-Hernandez et al.</a></p> <p>This repository provides the results of our classification, along with all the data related to the analyses presented in our study. DPCstruct algorithm can be found at <a href="https://github.com/RitAreaSciencePark/DPCstruct">https://github.com/RitAreaSciencePark/DPCstruct</a> together with examples on how to use it.</p> <p><strong>FILES DESCRIPTION:</strong></p> <ul> <li><strong>dpcstruct_classification.tsv: </strong>List of domains identified by DPCstruct and their corresponding metacluster. Columns: Metacluster ID, Protein Uniprot ID, domain start, domain end.</li> <li><strong>mcs_reps.fasta:</strong> For each metacluster, two representative domains were selected: one representing the center of the cluster and the other being the domain with the highest pLDDT score. If these are the same, only one domain is included as the representative. This file contains the list of representative domains and their sequences in FASTA format.</li> <li><strong><span>mcs_reps_pdbs.zip: </span></strong>Contains a PDB file for each representative domain. The filename is structured as 'proteinID_metacluster.pdb'.</li> <li><strong>mcs_properties.tsv:</strong> Set of properties per metacluster, including: <ul> <li><strong>mcID:</strong> Metacluster ID.</li> <li><strong>size:</strong> Number of domains.</li> <li><strong>len_aa:</strong> Average length of domains (number of amino acids).</li> <li><strong>len_std:</strong> Standard deviation of domain lengths.</li> <li><strong>len_ratio:</strong> Ratio of len_std to len_aa.</li> <li><strong>plddt:</strong> Average predicted LDDT as reported by AlphaFold2.</li> <li><strong>disorder:</strong> Average intrinsic disorder score calculated with AIUPred.</li> <li><strong>alntmscore:</strong> Pairwise alignment TM-score between domains, averaged over all pairs.</li> <li><strong>tmscore:</strong> Pairwise alignment TM-score between domains, averaged over all pairs, using the maximum between TM-score normalized by query or target.</li> <li><strong>lddt:</strong> Pairwise LDDT score, averaged over all pairs.</li> <li><strong>prob:</strong> Pairwise probability of homology according to SCOPe, as reported by Foldseek.</li> <li><strong>pident:</strong> Pairwise percentage identity, averaged over all pairs.</li> </ul> </li> <li><span><strong>annotated_[cath|scop]_qc[x]_t[x]_l[x].tsv:</strong>&nbsp;</span>For each fold in [CATH|SCOP], we provide the best matching DPCstruct domain, if available, along with the structural alignment information as reported by Foldseek. A fold is considered annotated if its alignment values meet or exceed the following thresholds: <ul> <li>qc: query coverage.</li> <li>t: template modelling score of the alignment.</li> <li>l: lddt score of the alignment.</li> </ul> </li> <li><strong>dpcstruct_consistency.tsv:</strong> Consistency of DPCstruct metaclusters with respect to Pfam 36.0 labels. Note that we consider a Pfam label to overlap with a DPCstruct domain even if it shares just one amino acid, which is why some metaclusters have many labels. In such cases, we only display 5 representative labels.</li> <li><strong>pfam_consistency.tsv:</strong> Consistency of Pfam Clans with respecto to DPCstruct labels.</li> </ul> <p><strong>Note:</strong> All 'tsv' files contain a header as the first row.</p> <p>If there is any doubt regarding the data or there is something missing please contact us:&nbsp;</p> <p>federico.barone@areasciencepark.it</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Figs. 72–77. Male genitalic structures. 72 in Classification Of The Bee Tribe Augochlorini (Hymenoptera: Halictidae)

Figs. 72–77. Male genitalic structures. 72. Genital capsule of Thectochlora alaris (Vachal), ventral aspect. 73. Genital capsule of Megalopta (Megalopta) genalis Meade-Waldo, ventral aspect. 74. Genital capsule of Rhinocorynura briseis (Smith), ventral aspect. 75. Gonocoxite-gonostylus junction in R. briseis depicting basal gonostylar process on ventral surface, process with strong setae at apex. 76. Penis valve of Temnosoma metallicum Smith, profile, with large dorsal processes. 77. Basal gonostylar process of Corynura (Callistochlora) prothysteres (Vachal) lacking setae, partially hidden by setose ventral

opencc-by-4.0Apr 2000View details →
zenodo40/100

Figs. 33–38. Anatomical structures. 33 in Classification, Natural History, and Evolution of Epiphloeinae (Coleoptera: Cleridae). Part V. Decorosa Opitz, a New Genus of Checkered Beetles from Hispaniola with Description of Its Four New Species

Figs. 33–38. Anatomical structures. 33. forebody of Madoniella dislocata (Say). 34. Aedeagus of Decorosa aladecoris Opitz. 35–36. Madoniella dislocata (35, terminal labial palpomere; 36, terminal maxillary palpomere). 37–38. D. iviei Opitz (37, terminal maxillary palpomere; 38, terminal labial palpomere.

opencc-by-4.0Sep 2008View details →
zenodo40/100

Figs. 28–32. Anatomical structures and body outlines. 28 in Classification, Natural History, and Evolution of Epiphloeinae (Coleoptera: Cleridae). Part V. Decorosa Opitz, a New Genus of Checkered Beetles from Hispaniola with Description of Its Four New Species

Figs. 28–32. Anatomical structures and body outlines. 28. Amboakis nova (Opitz) antenna. 29. Madoniella dislocata (Say) antenna. 30. A. nova body outline. 31. Decorosa iviei Opitz spicular fork. 32. M. dislocata body outline.

opencc-by-4.0Sep 2008View details →
zenodo40/100

NCCD-PF - A pre-failure narrow concrete cracks dataset for engineering structures damage classification and semantic segmentation

<p>The&nbsp;NCCD-PF dataset was developed for the classification and semantic segmentation of narrow concrete cracks in engineering structures elements at the pre-failure state. It only includes cracks whose width is narrower than 0.3 mm, i.e. the limit value specified in EC 1992-1-1 for typical elements of engineering structures and environmental conditions.</p> <p>This dataset is dedicated to the early crack detection at a stage when the serviceability limit state has not yet been exceeded and the failure of a structural element has not occurred. By implementing the early crack detection approach, it is possible to protect cracks in order to stop or slow down their propagation and thus to extend the structure's lifespan.</p> <p>This dataset contains images of cracks appearing on various elements of engineering structures (bridges, viaducts, tunnels) made of reinforced concrete (including abutments, tunnel walls, concrete barriers, pillars). The images were captured on construction sites and during inspections of engineering structures, at different stages of the reinforced concrete structure's working conditions - from the construction stage (when the elements are loaded only by their own weight) to the structure's use stage (when the elements are loaded by most of the design loads). The images are also differentiated by the cause of the cracking (ex., thermal and shrinkage stresses in young concrete, excessive stresses). The images were acquired using fixed-focus cameras without prior conditioning in order to represent the real working conditions of a bridge engineer during structural inspections. The images are characterised by a high degree of complexity due to the quality of the concrete surface finish (e.g. presence of formwork marks, concrete trowel marks), which could potentially be recognised&nbsp;as cracks.</p> <p>This dataset is dedicated to researchers working in the fields of computer vision, machine learning and deep learning. In particular, it contains domain knowledge in structural health monitoring, so that it can support the work of engineers in detecting cracks of concrete elements in a pre-failure state.</p> <p>A detailed description of the dataset is presented in <a href="https://www.nature.com/articles/s41597-023-02839-z" target="_blank" rel="noopener">A pre-failure narrow concrete cracks dataset for engineering structures damage classification and segmentation</a> (DOI: 10.1038/s41597-023-02839-z).</p>

opencc-by-4.0Aug 2023View details →
zenodo36/100

Code and data for: Network-based protein structural classification

<p>Code and data related to the research article titled, &quot;Network-based protein structural classification&quot;.</p> <p>More information about the code and the data is available at https://nd.edu/~cone/NETPCLASS/</p>

opencc-by-4.0May 2020View details →
dryad36/100

Population genetic structure and classification of cultivated and wild pea (Pisum sp.) based on morphological traits and SSR markers

<p>Pea (<em>Pisum</em> <em>sativum</em> L.) is an important legume crop that is widely grown worldwide for human consumption and livestock feed. Despite extensive studies, the population genetic structure and classification of cultivated and wild pea (<em>Pisum</em> sp.) are remaining controversial. To characterize patterns of genetic and morphological variation and investigate the classification of <em>Pisum</em>, we conducted comprehensive population genetic analyses for 323 accessions from cultivated and wild pea representing three species of <em>Pisum</em> utilizing 34 morphological traits and 87 polymorphic SSR markers. First, we identified three distinct genetic groups among all samples. Group I was primarily composed of <em>Pisum fulvum</em>, <em>Pisum</em> <em>abyssinicum</em> and some wild <em>P. sativum</em> accessions, whereas groups II and III consisted of the two genetic groups under <em>P. sativum </em>representing different geographic distributions of cultivated pea. Analyses of morphological variation revealed significant differences among the three species. Second, among pea germplasms representing eight taxa of <em>Pisum</em>, <em>P. fulvum</em> and <em>P. abyssinicum</em> possessed unique genetic backgrounds and morphological characteristics, corroborating their independent species status. The intraspecific subdivisions of <em>P. sativum</em> described by some authors were not supported in this study, with the exception of several genotypes of <em>P. sativum</em> subsp. <em>elatius</em> that were clustered with <em>P. fulvum</em> and <em>P. abyssinicum</em>. Finally, we confirmed that the Chinese pea germplasm was genetically distinct and could be divided into two genetic groups, each of which included both spring-sowing and autumn-sowing ecotypes. These results provide a robust foundation for understanding pea domestication and the utilization of wild genetic resources of pea.</p>

opencc-zeroNov 2020View details →
zenodo36/100

Classification and structural analysis of value chain contracts for biodiversity conservation in the European Union

<p><span>The data provided by this dataset are the raw data published in the paper "<strong>Classification and structural analysis of value chain contracts for biodiversity conservation in the European Union</strong>" (<a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.sftr.2024.100372" target="_blank" rel="noopener"><span><span>https://doi.org/10.1016/j.sftr.2024.100372</span></span></a>).&nbsp;</span></p>

opencc-by-4.0Nov 2024View details →
dryad36/100

IIb-RAD-seq coupled with random forest classification indicates regional population structuring and sex-specific differentiation in salmon lice (Lepeophtheirus salmonis)

<p><span>The aquaculture industry has been dealing with salmon lice problems forming serious threats to salmonid farming. Several treatment approaches have been used to control the parasite. Treatment effectiveness must be optimized, and the systematic genetic differences between sub-populations must be studied to monitor louse species and enhance targeted control measures. We have used IIb-RAD sequencing in tandem with a random forest classification algorithm to detect the regional genetic structure of the Norwegian salmon lice and identify important markers for sex differentiation of this species. We identified 19428 single nucleotide polymorphisms (SNPs) from 95 individuals of salmon lice. These SNPs, however, were not able to distinguish differential structure of lice populations. Using the random forest algorithm, we selected 91 SNPs important for geographical classification and 14 SNPs important for sex classification. The geographically important SNP data substantially improved the genetic understanding of the population structure and classified regional demographic clusters along the Norwegian coast. </span><span>We also uncovered SNP markers that could help determine the sex of the salmon louse. </span><span>A large portion of the SNPs identified to be under directional selection were also ranked highly important by random forest. According to our findings, there is a regional population structure of salmon lice associated with the geographical location along the Norwegian coastline.</span></p>

opencc-zeroApr 2022View details →
zenodo36/100

cldf-datasets/normansinitic: Structural and lexical data for the paper by Norman (2013) on Chinese dialect classification

<p><strong>Norman, J. (2003): Chinese dialects. Phonology. In: Thurgood, G. &amp; LaPolla, R.: The Sino-Tibetan Languages. Routledge: London and New York. 72-83.</strong></p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

Data, code, models for "Weakly Supervised Semantic Segmentation for Joint Key Local Structure Localization and Classification of Aurora Image"

<p>Data, code and models for https://ieeexplore.ieee.org/document/8410588/</p>

opencc-by-4.0Jul 2018View details →
zenodo36/100

cldf-datasets/normansinitic: Structural and lexical data for the paper by Norman (2013) on Chinese dialect classification

<p>Original source of the data:</p> <blockquote> <p>Norman, J. (2003): Chinese dialects. Phonology. In: Thurgood, G. &amp; LaPolla, R.: The Sino-Tibetan Languages. Routledge: London and New York. 72-83.</p> </blockquote>

opencc-by-4.0Nov 2019View details →
dryad36/100

Population genetic structure and classification of cultivated and wild pea (Pisum sp.) based on morphological traits and SSR markers

Open the record for dataset details and reuse information.

publicNov 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record