Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
14,447
datasets available to search
ShareScore release 0.7.1
Dataset results
14,447 results for “Identification”
Plum Island LTER phytoplankton identification using HPLC and Chem Taxonomy along transects in the Plum Island Sound estuary, Massachusetts. (Reformatted to the ecocomDP Design Pattern)
This data package is formatted as an ecocomDP (Ecological Community Data Pattern). For more information on ecocomDP see https://github.com/EDIorg/ecocomDP. This Level 1 data package was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-pie/404/4. The abstract below was extracted from the Level 0 data package and is included for context: Water column samples are collected along an estuarine salinity gradient as part of our monitoring surveys of the Parker River estuary each spring and late summer (typically high vs low freshwater input). Samples are filtered, and stored frozen for later pigment analyses by HPLC. Pigment data are then analyzed by CHEMTAX, calibrated to a matrix of pigment ratios based on taxonomy and enumeration of selected subsamples by microspcopy. Data are presented in terms of chlorophyll a concentrations partitionaed among the major phytoplankton groups as determined by CHEMTAX. For 2003-2006, sampling stations along the Plum Island Sound-Parker River were at fixed geographic locations at specific "Bends" in the river. In 2008, we began sampling the water column in salinity space rather than at specific geographic locations along the river. This sampling approach was adopted in order to follow particular water masses in this macrotidal estuary. In practical terms, it means that sampling locations, or stations, are not static. Therefore, we have mapped the 11 sampling locations (latitude and longitude are logged at each station) from each transect along the mainstem of the estuary, so each station may be placed along the river (to the nearest 0.5km) as well as in salinity space. We have also used the km marker to assign the sampling locations from each survey to one of four bounding boxes : the Sound (Plum Island Sound; EST-PR-SoundBND) which encompasses approximatly the first 9.5 km or the transect, with Okm at the mouth of the sound; the Lower Parker River (EST-PR-LowerParkerBND) , ~9.5 - 1
Identification of an altitudinal migration pattern of Abies pinsapo in the Baetic Mountains through the presence of its life stages
<p>This data set is used to explore the altitudinal shift of <em>Abies pinsapo</em> Boiss. in the Baetic System. We analysed the potential distribution of the realised and reproductive niches of <em>A. pinsapo</em> populations in the Ronda Mountains (Southern Spain) by using species distribution models (SDMs) for two life stages within the current populations. The realised and reproductive niches of <em>A. pinsapo</em> are different to one another, which may indicate a displacement in its altitudinal distribution.</p>
Cirrus formation regimes - Data driven identification and quantification of mineral dust effect
<p>This repository contains the data for the paper: </p> <p>Authors: Kai Jeggle , David Neubauer , Hanin Binder and Ulrike Lohmann<br>Titel: Cirrus formation regimes - Data driven identification and quantification of mineral dust effect<br>Date: 2024</p> <p>Note that the scripts can be found in the accompanying code repository (https://github.com/tabularaza27/cloud_clustering)<br><br>Contents:<br><br>├── cirrus_cloud_trajectories.ftr<br>├── cluster_input_data.ftr<br>├── cluster_models<br>│ └── temperature_clustering_k4_12<br>│ ├── cloud_ids.npy<br>│ ├── model_params.json<br>│ └── trained_model.hdf5</p> <p>│ └── temperature_clustering_k4_24<br>│ ├── cloud_ids.npy<br>│ ├── model_params.json<br>│ └── trained_model.hdf5</p> <p>├── cluster_predictions.ftr<br>└── readme.txt<br><br>For more info, please have a look at the <em>readme.txt</em><br><br>This is an updated version of the data, containing updated models and predictions based on the Journal revisions</p>
LIAS light – A Database for Rapid Identification of Lichens – Subset Switzerland p. pte.
<p>This subset of the LIAS light database focuses on lichens found in Switzerland, providing comprehensive data for ecological research and taxon identification purposes. Not all taxa and only categorical (but 2 numerical) characters (descriptors) of the recorded taxa are covered. Updates and additional data will be published subsequently.</p>
Interlaboratory study: Testing reproducibility of solid biofuels component identification using reflected light microscopy
<p><strong>Submitted data was used to write an article: </strong>Drobniak, A., Mastalerz, M., Jelonek, Z., Jelonek, I., Adsul, T., Andolšek, N., Ardakani, O.H., Congo, T., Demberelsuren, B., Donohoe, B.S., Douds, A., Flores, D., Ganzorig, R., Ghosh, S., Gize, A., Goncalves, P.A., Hackely, P., Hatcherian, J., Hower, J.C., Kalaitzidis, S., Kędzior, S., Knowles, W., Kuś, J., Lis, K., Lis, G., Liu, B., Luo, Q., Du, M., Mishra, D., Misz-Kennan, M., Mugerwa, T., O'Keefe, J., Park, J., Pearson, R., Petersen, H., Reyes, J., Ribeiro, J., Niedzwiedzkas, J.L., de la Rosa Rodriguez, G., Sosnowski, P., Valentine, B., Varma, A., Wojtaszek-Kalaitzidi, M., Xu, Z., Zdravkov, A., Ziemianin, K., Interlaboratory study: Testing reproducibility of biomass fuels component identification using reflected light microscopy. International Journal of Coal Geology 277, 104331. <a href="https://doi.org/10.1016/j.coal.2023.104331">https://doi.org/10.1016/j.coal.2023.104331</a>.</p> <p> </p> <p><strong>Funding acknowledgments: </strong>The project is co-financed by the Polish National Agency for Academic Exchange within the Polish Returns Programme (BPN/PPO/2021/1/00005/DEC/1), the National Science Center, Poland (2022/01/1/ST10/00024), and the research activities co-financed by the funds granted under the Research Excellence Initiative of the University of Silesia in Katowice, Poland. </p> <p> </p> <p><strong>Article Abstract: </strong>Considering global market trends and concerns about climate change and sustainability, increased biomass use for energy is expected to continue. As more diverse materials are being utilized to manufacture solid biomass fuels, it is critical to implement quality assessment methods to analyze these fuels thoroughly. One such method is reflected light microscopy (RLM), which has the potential to complement and enhance current standard testing, leading to improving fuel quality assessment and, ultimately, preventing avoidable air pollution. An interlaboratory study (ILS) was conducted to test the reproducibility of biomass fuels component identification using a reflected light microscopy technique. The exercise was conducted on thirty photomicrographs showing biomass and various undesired components (like plastics or mineral matter), which were purposely added (by the ILS organizers) to contaminate wood pellets and charcoal-based grilling fuels. Forty-six participants had various levels of difficulty identifying the marked components, and as a result, the percentage of correct answers ranged from 52.2 to 94.4%. Among the most difficult components to distinguish were petroleum products and inorganic matter. Various reasons led to the misidentification, including insufficient morphological descriptions of the components provided to participants, ambiguities of the nomenclature, limitations of the analytical and exercise method, and insufficient experience of the participants. Overall, the results indicate that RLM has the potential to enhance the quality assessment of biomass fuels. However, they also demonstrate that the petrographic classification used in this exercise requires further refinement before it can be standardized. While a new simplified classification of solid biomass fuels components was created as an outcome of this study, future research is necessary to refine the nomenclature, develop a microscopic morphological description of the components, and verify the accuracy of component identification with a follow-up ILS.</p>
SJR Dolphin SCA: Degradation Scores, OL Length and Identifications of Otoliths Collected from Bottlenose Dolphin Stomachs
Otoliths were collected from stomach contents of stranded bottlenose dolphins (Tursiops erebennus) in the St. Johns River in Jacksonville, Florida. Otoliths were analyzed by a panel of 3 reviewers to determine the level of otolith degradation that occurred during digestive processes. Otoliths with scores ≤ 3 were measured. Otolith length measurements were used to estimate the size of most species by applying standard regression equations developed from fish species collected from a nearby water system, the Indian River Lagoon, and for one species, equations developed from violet gobies collected from the St. Johns River. These equations enabled estimation of the mass of each prey species in each dolphin’s stomach, then the calculation of their relative proportions of reconstructed mass across all stomachs. For otolith identification purposes, a panel of 4 reviewers assigned each otolith with a family-level and species-level identification. Each identification was given a confidence code ranging from 1 (no confidence) to 4 (certainty). When the average code for all reviewers was < 3, the otolith was considered unidentified. If two of the reviewers agreed with the “weight” reviewer and all gave scores ≥ 3, the score of the outlying reviewer was discarded. Identification was assigned when the average confidence code was ≥ 3. The minimum number of species per dolphin stomach was determined by counting the left and right otoliths for each species separately, using the higher count as the minimum prey number. Unidentified species were counted, and half of their sum was considered the minimum prey number. The frequency of occurrence (%FO, or proportion of stomachs in which a species was detected) and numerical proportion (%N, or proportion of a given species pooled across all stomach samples) of each prey species were then calculated.
Plum Island LTER phytoplankton identification using HPLC and Chem Taxonomy along transects in the Plum Island Sound estuary, Massachusetts. (Reformatted to a Darwin Core Archive)
This data package is formatted as a Darwin Core Archive (DwC-A, event core). For more information on Darwin Core see https://www.tdwg.org/standards/dwc/. This Level 2 data package was derived from the Level 1 data package found here: https://pasta.lternet.edu/package/metadata/eml/edi/338/2, which was derived from the Level 0 data package found here: https://pasta.lternet.edu/package/metadata/eml/knb-lter-pie/404/4. The abstract below was extracted from the Level 0 data package and is included for context: Water column samples are collected along an estuarine salinity gradient as part of our monitoring surveys of the Parker River estuary each spring and late summer (typically high vs low freshwater input). Samples are filtered, and stored frozen for later pigment analyses by HPLC. Pigment data are then analyzed by CHEMTAX, calibrated to a matrix of pigment ratios based on taxonomy and enumeration of selected subsamples by microspcopy. Data are presented in terms of chlorophyll a concentrations partitionaed among the major phytoplankton groups as determined by CHEMTAX. For 2003-2006, sampling stations along the Plum Island Sound-Parker River were at fixed geographic locations at specific "Bends" in the river. In 2008, we began sampling the water column in salinity space rather than at specific geographic locations along the river. This sampling approach was adopted in order to follow particular water masses in this macrotidal estuary. In practical terms, it means that sampling locations, or stations, are not static. Therefore, we have mapped the 11 sampling locations (latitude and longitude are logged at each station) from each transect along the mainstem of the estuary, so each station may be placed along the river (to the nearest 0.5km) as well as in salinity space. We have also used the km marker to assign the sampling locations from each survey to one of four bounding boxes : the Sound (Plum Island Sound; EST-PR-SoundBND) which encompasses approximatly the first 9
High-Voltage Disconnector State Identification: synthetic and real images of substation disconnectors
<p>This dataset contains the training and test images used in the work detailed in the article: Barpp Gomes, V., Marchesi, B., Gruber, Y.A. <em>et al.</em> Exploring Synthetic Data for Training Deep Learning Models for High-Voltage Disconnector State Identification. <em>J Control Autom Electr Syst</em> (2025). <a href="https://doi.org/10.1007/s40313-025-01204-2">https://doi.org/10.1007/s40313-025-01204-2</a></p> <p>Contais about 940,000 synthetic (CGI-rendered) and 60,000 real (camera-captured) samples of four types of substation disconnectors, on both open and closed states:</p> <ul> <li>230 kV center break (type 1, as indicated in the article);</li> <li>230 kV center break (type 2);</li> <li>230 kV double side break</li> <li>525 kV horizontal semi-pantograph</li> </ul> <p>Each zip file contains images of one type of substation disconnector. Images are sized 320x128 and are organized in folders, as follows:</p> <ul> <li>00_train_synth: Synthetic training images.</li> <li>01_train_real: A small set of real training images, as indicated in the article.</li> <li>02_test_real_normal1: One set of real test images.</li> <li>03_test_real_normal2: Another set of real test images, from a different time period.</li> <li>04_test_real_maneuvers: A special set of real test images in which the switches have been operated (are in different states).</li> </ul>
BWILD: Beach seagrass Wrack Identification Labelled Dataset
<h1>Training dataset</h1> <p>BWILD is a dataset tailored to train Artificial Intelligence applications to automate beach seagrass wrack detection in RGB images. It includes oblique RGB images captured by SIRENA beach video-monitoring systems, along with corresponding annotations, auxiliary data and a README file. BWILD encompasses data from two microtidal sandy beaches in the Balearic Islands, Spain. The dataset consists of images with varying fields of view (9 cameras), beach wrack abundance, degrees of occupation, and diverse meteoceanic and lighting conditions. The annotations categorise image pixels into five classes: i) Landwards, ii) Seawards, iii) Diffuse wrack, iv) Intermediate wrack, and v) Dense wrack.</p> <h1>Technical details</h1> <p>The BWILD version 1.1.0 is packaged in a compressed file (BWILD_v1.1.0.zip). A total of 3286 RGB images are shared in PNG format, corresponding annotations and masks in various formats (PNG, XML, JSON,TXT), and the README file in PDF format.</p> <h2>Data preprocessing</h2> <p>The BWILD dataset utilizes snapshot images from two SIRENA beach video-monitoring systems. To facilitate annotation while maintaining a diverse range of scenarios, the original 1280x960 pixel images were cropped to smaller regions, with a uniform resolution of 640x480 pixels. A subset of images was carefully curated to minimize annotation workload while ensuring representation of various time periods, distances to camera, and environmental conditions. Image selection involved filtering for quality, clustering for diversity, and prioritizing scenes containing beach seagrass wracks. Further details are available in the README file. </p> <h2>Data splitting</h2> <p>Data splitting requirements may vary depending on the chosen Artificial Intelligence approach (e.g., splitting by entire images or by image patches). Researchers should use a consistent method and document the approach and splits used in publications, enabling reproducible results and facilitating comparisons between studies. </p> <h2>Classes, labels and annotations</h2> <p>The BWILD dataset has been labelled manually using the 'Computer Vision Annotation Tool' (CVAT), categorising pixels into five labels of interest using polygon annotations.</p> <table> <tbody> <tr> <td><strong> Label</strong></td> <td><strong> Description</strong></td> </tr> <tr> <td>landwards</td> <td>Pixels that are towards the landside with respect to the shoreline</td> </tr> <tr> <td>seawards</td> <td>Pixels that are towards the seaside with respect to the shoreline</td> </tr> <tr> <td>diffuse wrack</td> <td>Pixels that potentially resembled beach wracks based on colour and shape, yet the annotator could not confirm this with certainty, were denoted as ‘diffuse wrack’</td> </tr> <tr> <td>Intermediate wrack</td> <td>Pixels with low-density beach wracks or mixed beach wracks and sand surfaces</td> </tr> <tr> <td>Dense wrack</td> <td>Pixels with high-density beach wracks</td> </tr> </tbody> </table> <p>Annotations were exported from CVAT in four different formats: (i) CVAT for images (XML); (ii) Segmentation Mask 1.0 (PNG); (iii) COCO (JSON); (iv) Ultralytics YOLO Segmentation 1.0 (TXT). These diverse annotation formats can be used for various applications including object detection and segmentation, and simplify the interaction with the dataset, making it more user-friendly. Further details are available in the README file. </p> <h2>Parameters</h2> <p>RGB values or any transformation in the colour space can be used as parameters.</p> <h2>Data sources</h2> <p>A SIRENA system consists of a set of RGB cameras mounted at the top of buildings on the beachfront. These cameras take oblique pictures of the beach, with overlapping sights, at 7.5 FPS during the first 10 minutes of each hour in daylight hours. From these pictures, different products are generated, including snapshots, which correspond to the frame of the video at the 5th minute. In the Balearic Islands, SIRENA stations are managed by the Balearic Islands Coastal Observing and Forecasting System (SOCIB), and are mounted at the top of hotels located in front of the coastline. The present dataset includes snapshots from the SIRENA systems operating since 2011 at Cala Millor (5 cameras) and Son Bou (4 cameras) beaches, located in Mallorca and Menorca islands (Balearic Islands, Spain), respectively. All latest and historical SIRENA images are available at the Beamon app viewer (https://apps.socib.es/beamon). </p> <h2>Data quality</h2> <p>All images included in BWILD have been supervised by the authors of the dataset. However, variable presence of beach segrass wracks across different beach segments and seasons impose a variable distribution of images across different SIRENA stations and cameras. Users of BWILD dataset must be aware of this variance. Further details are available in the README file. </p> <h2>Image resolution</h2> <p>The resolution of the images in BWILD is of 640x480 pixels.</p> <h2>Spatial coverage</h2> <p>The BWILD version 1.1.0 contains data from two SIRENA beach video-monitoring stations, encompassing two microtidal sandy beaches in the Balearic Islands, Spain. These are: Cala Millor (<em>clm</em>) and Son Bou (<em>snb</em>). </p> <table> <tbody> <tr> <td><strong>SIRENA station</strong></td> <td><strong> Longitude</strong></td> <td><strong> Latitude</strong></td> </tr> <tr> <td><em>clm</em></td> <td>3.383</td> <td>39.596</td> </tr> <tr> <td><em>snb</em></td> <td>4.077</td> <td>39.898</td> </tr> </tbody> </table> <h2>Contact information</h2> <p>For further technical inquiries or additional information about the annotated dataset, please contact jsoriano@socib.es.</p>
xPore: Identification of differential RNA modifications from nanopore direct RNA sequencing
<p>xPore is a Python package for identification and quantification of differential RNA modifications from direct RNA sequencing.</p> <p>The detailed usage is documented at <a href="https://xpore.readthedocs.io/en/latest/">https://xpore.readthedocs.io/en/latest</a>, while all scripts and source code are available at <a href="https://github.com/GoekeLab/xpore">https://github.com/GoekeLab/xpore</a>.</p> <p>All the preprocessed datasets used in the paper are provided here. </p> <p>Please cite our paper below when using these data.<br> Ploy N. Pratanwanich et al. "Detection of differential RNA modifications from direct RNA sequencing of human cell lines." bioRxiv (2020).</p>
Protein structure files for the paper "Multiplexed identification of RAS paralog imbalance as a driver of lung cancer growth" in Nature Cell Biology by Tang et al.
<p>This archive contains models of HRAS, KRAS, and NRAS homo- and heterodimers with various mutations discussed in the paper, "Multiplexed identification of RAS paralog imbalance as a driver of lung cancer growth" in Nature Cell Biology by Tang et al.<br> as well as crystallographic dimers of these proteins as identified by the ProtCAD database, http://dunbrack2.fccc.edu/ProtCAD/Results/PfamArchClusterInfo.aspx?GroupId=8 (cluster 5). Several of the models are shown in Supp. Figure 11b and the crystallographic dimers of RAS that provide evidence for the possible biological relevance of these models are shown in Supp. Figure 11a.</p> <p>The crystallographic dimers were identified by clustering all possible interfaces generated by symmetry operators in crystals of HRAS, KRAS, and NRAS as described in the paper: Xu, Q., Dunbrack, R.L. ProtCID: a data resource for structural information on protein interactions. <em>Nat Commun</em> <strong>11</strong>, 711 (2020). https://doi.org/10.1038/s41467-020-14301-4.</p> <p>The models were created by superposing monomers of HRAS, KRAS, or NRAS onto the alpha4-alpha5 dimer present in the crystal of PDB entry 3k8y. Mutations were made in PyMOL. The structures were relaxed with the FastRelax protocol and the Ref2015 scoring function in the program Rosetta, which uses the backbone-dependent rotamer library of Shapovalov and Dunbrack to repack side chains.</p> <p>The crystallographic dimers are contained in a zipped PyMOL session. The mmCIF format for all the structures is present in a zip file, Tang_et_al_crystallographic_and_modeled_RAS_dimer_ciffiles.zip. The PyMOL session and zip file contains 87 HRAS dimers, 14 KRAS dimers, and 1 NRAS dimer, all having the interface consisting of the alpha4 and alpha5 helices. The PyMOL session also contains the modeled structures. Only Mg ions and GTP/GNP/GDP ligands are shown. Others are present but hidden and may be displayed by PyMOL ("show sticks, het").</p> <p> </p>
DIPROMATS 2024 - Shared Task 2: testing data for narrative identification
<p>Narratives are causally connected sequences of events that are selected and evaluated as meaningful for a particular audience. They make sense of the world by identifying the significance of people, places, objects, and events in time. In international relations, international actors create strategic narratives to “construct a shared meaning of the past, present, and future of international politics to shape the behavior of domestic and international actors”</p> <p>DIPROMATS 2024 Task 2 is a multiclass multilabel classification problem. Given a series of predefined narratives of each international actor, systems must determine which narrative the tweets belong to. Systems will receive the description of each narrative and a few examples of tweets in both languages (English and Spanish) that belong to each of them (few-shot learning). A tweet may be associated with one, several or none of the narratives.</p> <p>The few-shot training data can be found here: <a href="https://doi.org/10.5281/zenodo.10820961" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.10820961</a></p> <p>These are the testing datasets for Englsih and Spanish. They are provided without the keys so the large language models can't be contaminated. If you are interested on testing your system, write anselmo@lsi.uned.es for details on submission and leaderboards.</p>
BioDeep/metabolomics-report-standards: BioDeep LC-MS Metabolite Identification Demo Report
<p><em>A Metabolomics unknown feature identification report industry standards from <a href="http://www.bionovogene.com/">BioNovoGene</a> corporation.</em></p> <p>2019.08.16# at Suzhou, China</p> <p>There is a general consensus that supports the need for standardized reporting of metadata or information describing large-scale metabolomics data sets. Reporting of standard metadata provides a biological and empirical context for the data, enables the reinterrogation and comparison of data by others, which is also could let us interpret the result in a more clearly way.</p> <p>This article is mainly address at the unknown metabolite identification in LC-MS experiment, and proposes the reporting standards related to the chemical analysis aspects of metabolomics experiments its metabolite identification.</p> <p>Some terms in this article that address to:</p> <ul> <li>feature, the term feature in this article is refer to a parent ion in LC-MS experiment result raw data. Where a parent ion feature is a peak in chromatography data, which is consist of mass to charge ratio in ms1 level and its retention time (with a range of lower bound and upper bound) in chromatography experiment result.</li> <li>annotation, the term annotation in this article is refer to the multidimensional information about the metabolite that assigned to a unknown feature, which such multidimensional information consist with the metabolite its cross reference id in different database, common name, basic chemical data like mass and formula composition and its molecule structure information, etc.</li> <li>alignment, the term alignment means a kind of operation that use to compare the similarity of the mass spectrum data between user sample and the reference standard library. Such similarity comparison result is the most important evidence that use for unknown feature its identification.</li> <li>score, the term score is a kind of numeric value that produced by the alignment comparison calculation. Literally, the higher score the alignment it produce, the better the result it is.</li> </ul> <p>Our metabolite identification report consist with two parts of data which present to our user:</p> <ol> <li>Report excel table that contains the raw sample information and the meta annotation information of the metabolite.</li> <li>Data visual plot for the mass spectrum alignment details.</li> </ol>
Datasets for phylogenetic analyses and phylogenetic trees for: Genetic barcodes for species identification and phylogenetic estimation in ghost spiders (Araneae: Anyphaenidae: Amaurobioidinae). Invertebrate Systematics, 2024
<p>We combined the COI sequence data with legacy multigene sequence data to create a new, taxon-rich phylogeny for the Amaurobioidinae. We used sequences for four loci that have been used in previous studies on the subfamily: two mitochondrial loci, COI (658bp) and ribosomal subunit 16S (16S, 410bp); and two nuclear loci, Histone H3 (H3, 327bp) and ribosomal subunit 28S (28S, 839bp). We complemented the Amaurobioidinae data with sequences from several non-amaurobioidine anyphaenids and two clubionids as outgroups. Sequence alignment was performed using the MAFFT (ver. 7.308) plugin in Geneious, allowing MAFFT to automatically select an appropriate alignment strategy based on the properties of each locus, or with the online MAFFT server (https://mafft.cbrc.jp), which consistently selected the L-INS-i algorithm. Finally, alignments of the four loci were concatenated to construct a 2234 bp multigene sequence matrix containing 692 taxa, with about 55% missing/gap data (“full” matrix henceforth). To ensure that excessive missing data did not affect the resulting topology, we also constructed a reduced matrix by removing additional COI-only specimens so that each species and morphotype was represented by just one or two specimens for which all loci were available (where possible). After realignment, this reduced matrix was 2235 bp long, included 167 taxa, and had about 22% missing/gap data (“reduced” matrix henceforth). Phylogenetic analyses under maximum likelihood, including model selection, were then conducted with IQ-TREE 2. We performed phylogenetic analyses on both concatenated matrices (the full matrix and the reduced matrix) and on each individual locus. For model selection, we provided an initial scheme that partitioned the matrix by locus, and further partitioned the protein-coding loci (COI and H3) by codon position. We used ModelFinder and searched for the best partition scheme, all in IQ-TREE. The best models (partitions) for the full dataset were: GTR+F+I+G4 (16S), GTR+F+I+I+R4 (28S), TVM+F+I+I+R2 (COI-1), TIM2+F+R4 (COI-2), GTR+F+R5 (COI-3), TVMe+G4 (H3-1-H3-2), SYM+G4 (H3-3); and for the reduced dataset: GTR+F+I+G4 (16S), GTR+F+I+G4: (28S), GTR+F+I+G4: (COI-2), GTR+F+I+G4: (COI-3), TVM+F+I+G4: (COI-1, H3-2), GTR+F+I+G4: (H3-1), GTR+F+I+G4: (H3-3). For each dataset, once the best models and partitions were defined, we executed 10 independent replicates of tree calculations followed by 1000 ultrafast bootstrap replicates, and the replicate reaching the maximum likelihood was chosen. Phylogenetic analyses under parsimony were made with TNT, under equal weights, using the “new technology” search with default values, asking for 10 independent hits to the minimal length, and submitting the resulting trees to a round of TBR branch swapping. </p>
Plant image identification application demonstrates high accuracy in Northern Europe dataset
<p><strong>Images and data for the study "Plant image identification application demonstrates high accuracy in Northern Europe"</strong></p> <p><strong>Details: Jaak Pärtel, Meelis Pärtel, Jana Wäldchen, Plant image identification application demonstrates high accuracy in Northern Europe, <em>AoB PLANTS</em>, Volume 13, Issue 4, August 2021, plab050, <a href="https://doi.org/10.1093/aobpla/plab050">https://doi.org/10.1093/aobpla/plab050</a></strong></p> <p>The data table displays Flora Incognita's identification results together with species and observations characteristics. All (3199) used images are included.</p> <p>The study was conducted in two parts: database and field study.</p> <p>Database study images have been taken from eBiodiversity database (https://elurikkus.ee/en) under Creative Commons Attribution 4.0 International (CC BY 4.0) licence (https://creativecommons.org/licenses/by/4.0/). Please cite the original source for the images as well when using the dataset.</p> <p>Field study images were taken by Jaak Pärtel in 2020 in field conditions from different habitats across Estonia.</p>
Identification and characterization of the cell division protein MapZ of Streptococcus suis
<p>Supplementary data and code related to the manuscript "Identification and characterization of the cell division protein MapZ of <em>Streptococcus suis</em>".</p>
Genus Erica: An Identification Aid Version 4.03
<p>Genus <em>Erica</em>: An Identification Aid is a tool to help both amateurs and professionals identify (using a limited number of accessible characteristics) and find information about the 851 species and many subspecific taxa of the flowering plant genus <em>Erica</em>. Version 4.00 includes new features such as integrating distribution data from GBIF and iNaturalist, links to taxonomic resources through World Flora Online, and a probability function for identifications; V. 4.01 also includes eFlora links and an optional research data window (see below). It is freely available for PCs.</p> <p>You can install the <em>Erica </em>ID aid whether you have MS office or not. If you have a current version of MS office installed, use 'GenusEricaAnIdentificationAid400 normal.exe'. V.4.00.02 onwards also includes alternative installation kits for in case you have specific older versions of MS office installed. If you do not have MS office, any of these should work.</p> <p>V.4.00.04 is an incremental update following peer review. The ms. is currently in press at the journal PhytoKeys [published 24th April 2024: <a href="https://doi.org/10.3897/phytokeys.241.117604" target="_blank" rel="noopener">https://doi.org/10.3897/phytokeys.241.117604</a>]</p> <p>V.4.01: Through work with WFO and SANBI, unique taxon identifiers have been linked across WFO (<a href="https://www.worldfloraonline.org/" target="_blank" rel="noopener">https://www.worldfloraonline.org/</a>), the SANBI eFlora SA (<a href="https://www.sanbi.org/biodiversity/foundations/biosystematics-collections/e-flora/" target="_blank" rel="noopener">https://www.sanbi.org/biodiversity/foundations/biosystematics-collections/e-flora/</a>), and GBiF (<a href="https://www.gbif.org/" target="_blank" rel="noopener">https://www.gbif.org/</a>). This means the <em>Erica </em>ID aid can now refer to and receive data unambiguously from these sources. SANBI e-Flora references have been added and can be reached by clicking on the e-Flora button on the <em>Erica </em>Characters window. An additional Research Data window has been added. This will be mostly of use (and largely self-explanatory) to the systematics/conservation research community, and can be accessed through a button on the ribbon/an option in the settings. Updates to GBIF distribution data; improvement of data for Madagascan taxa along with other minor taxonomic changes and improvements to the data and helpfile.</p> <p>V.4.02: Incorporates incremental improvements in WFO data released in June 2025 and newly includes depreciated names (now searchable, for reference, with other published names that do not correspond to known taxa). All currently accepted subspecific taxa have been added (but in most cases not yet character-coded). New species described in 2025 added. iNaturalist taxon profile images available on CC licenses in April 2025 are included with license data embedded in the image metadata. Further development and updating of the research and other data.</p> <p>V.4.03: Updated GBIF, South African Red List and E-Flora and IUCN threat status data, with extended and updated research data/summaries used in/documenting the 'Conservation gap analysis for Erica (Ericaceae)' ms. currently in review (<span><a href="https://doi.org/10.3897/arphapreprints.e176100">https://doi.org/10.3897/arphapreprints.e176100</a></span>).</p>
Identification of microbial exopolymer producers in sandy and muddy intertidal sediments by compound-specific isotope analysis.
<p>This dataset supports the version 2 of the paper entitled <em>Identification of microbial exopolymer producers in sandy and muddy intertidal sediments by compound-specific isotope analysis :</em></p> <p><em>Hubas, Cédric; Gaubert-Boussarie, Julie; D’Hondt, An-Sofie; Jesus, Bruno; Lamy, Dominique; Meleder, Vona; Prins, Antoine; Rosa, Philippe; Stock, Willem; Sabbe, Koen. Identification of microbial exopolymer producers in sandy and muddy intertidal sediments by compound-specific isotope analysis. Peer Community Journal, Volume 3 (2023), article no. e104. doi : <a href="https://doi.org/10.24072/pcjournal.336">10.24072/pcjournal.336</a>. <a href="https://peercommunityjournal.org/articles/10.24072/pcjournal.336/">https://peercommunityjournal.org/articles/10.24072/pcjournal.336/</a></em></p>
Lots for greening: Identification of metropolitan vacant land and its potential use for cooling and agriculture in Phoenix, Arizona, USA
This project provides the first systematic assessment of non-governmental vacant parcels for potential greening (VPPG) the Phoenix metropolitan area—land parcels that are or can be privately owned but which contain no buildings, are unpaved, have no apparent use, and are potential candidates for urban greening. To achieve the data, a new method for the identification of vacant lands was employed that combines remote sensing techniques and cadastral data and trains the computer to distinguish different forms of vacant land. The classification result proved to be an effective approach for open land identification and identified approximately 19500 ha of open land in the metro area. The model achieved an average accuracy of 90.67%. This dataset only includes VPPG and does not include other vacant land determined to be inappropriate for potential greening (developed/abandoned or impervious surface). (Overall accuracy for all classes was 87.20%).
Plum Island LTER phytoplankton identification using HPLC and Chem Taxonomy along transects in the Plum Island Sound estuary, Massachusetts.
Water column samples are collected along an estuarine salinity gradient as part of our monitoring surveys of the Parker River estuary each spring and late summer (typically high vs low freshwater input). Samples are filtered, and stored frozen for later pigment analyses by HPLC. Pigment data are then analyzed by CHEMTAX, calibrated to a matrix of pigment ratios based on taxonomy and enumeration of selected subsamples by microspcopy. Data are presented in terms of chlorophyll a concentrations partitionaed among the major phytoplankton groups as determined by CHEMTAX. For 2003-2006, sampling stations along the Plum Island Sound-Parker River were at fixed geographic locations at specific "Bends" in the river. In 2008, we began sampling the water column in salinity space rather than at specific geographic locations along the river. This sampling approach was adopted in order to follow particular water masses in this macrotidal estuary. In practical terms, it means that sampling locations, or stations, are not static. Therefore, we have mapped the 11 sampling locations (latitude and longitude are logged at each station) from each transect along the mainstem of the estuary, so each station may be placed along the river (to the nearest 0.5km) as well as in salinity space. We have also used the km marker to assign the sampling locations from each survey to one of four bounding boxes : the Sound (Plum Island Sound; EST-PR-SoundBND) which encompasses approximatly the first 9.5 km or the transect, with Okm at the mouth of the sound; the Lower Parker River (EST-PR-LowerParkerBND) , ~9.5 - 14.5 km; the Middle Parker River (EST-PRMiddleParkerBND), ~14.5 - 18.75 km, and the Upper Parker River (EST-PR-UpperParker BND)., ~18.75 to 24.25 km (the Parker R. Dam).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.