Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
40
datasets available to search
ShareScore release 0.9.0
Dataset results
40 results for “Automatic Identification”
InsectSet32: Dataset for automatic acoustic identification of insects (Orthoptera and Cicadidae)
<p>This dataset contains recordings of 32 sound producing insect species with a total 335 files and a length of 57 minutes. The dataset was compiled for training neural networks to automatically identify insect species while comparing adaptive, waveform-based frontends to conventional mel-spectrogram frontends for audio feature extraction. This work was <a href="https://doi.org/10.1371/journal.pcbi.1011541">published</a> in PLOS Computational Biology and this dataset can be used to replicate the results, as well as other uses. The scripts for audio processing and the machine learning implementations are published on <a href="https://github.com/mariusfaiss/InsectSet32-Adaptive-Representations-of-Sound-for-Automatic-Insect-Recognition">Github</a>.</p> <p>The recordings are split into two datasets. Roughly half of the recordings (147) are of nine species belonging to the order Orthoptera. These recordings stem from a dataset that was originally compiled by <a href="https://orcid.org/0000-0002-8929-2737">Baudewijn Odé</a> (unpublished). </p> <p>The remaining recordings (188) are of 23 species in the family Cicadidae. These recordings were selected from the Global Cicada Sound Collection hosted on <a href="https://bio.acousti.ca/">Bioacoustica</a> (<a href="https://doi.org/10.1093/database/bav054">doi.org/10.1093/database/bav054</a>), including recordings published in <a href="https://doi.org/10.3897/BDJ.3.e5792">doi.org/10.3897/BDJ.3.e5792</a> & <a href="https://doi.org/10.11646/zootaxa.4340.1">doi.org/10.11646/zootaxa.4340.1</a>. Many recordings from this collection included speech annotations in the beginning of the recordings, therefore the last ten seconds of audio were extracted and used in this dataset. </p> <p>All files were manually inspected and files with strong noise interference or with sounds of multiple species were removed. Between species, the number of files ranges from four to 22 files and the length from 40 seconds to almost nine minutes of audio material for a single species. The files range in length from less than one second to several minutes. All original files were available with sample rates of at least 44.1 kHz or higher but were resampled to 44.1 kHz mono WAV files for consistency. The annotation files contain information for each recording, including the file name, species name and identifier, as well as the data subset they were included in for training the neural network (training, test, validation).</p>
A software for automatic identification of oyster species
<p>The files includes all data and final analysis of the work done in CS8 - oysters, task 8.2, CSTP8.2.2_A new software for automatic identification of oyster species. This includes the data management descriptor document (DataSheet_oyster_image_classification.docx), images used (oyster_classification_images.zip), the code developed (oyster_classification_code_package.zip) and different models evaluated (oyster_classification_models.zip), the genetics data produced (oyster_classification_biometrics and PCR.xlsx) and the project report (C639_ostronklassificering.pdf). The content of the files is described briefely below. The data is used in deliverables D1.2, D1.4, D1.5 and D1.6 in the AquaVitae project.</p> <p>oyster_classification_images.zip</p> <p>The data set contains the images used for training the classification models that are capable of classifying images of oysters as either Ostrea edulis or Magallana gigas. The images are sorted in folders named “train” (training data) and “validation” (validation data) with both folders containing sub-folders called “mg” (images of Magallana gigas) and “oe” (images of Ostrea edulis).</p> <p>oyster_classification_code_package.zip</p> <p>The data set contains the code for training a neural network for classifying oyster species based on images. The code also includes localization of oyster within an image and inference of the classification along with the trained models.</p> <p>oyster_classification_models.zip</p> <p>The data set contains the trained classification models that are capable of classifying images of oysters as either Ostrea edulis or Magallana gigas.</p> <p>oyster_classification_biometrics and PCR.xlsx</p> <p>The data set contains biometric information for a subset of 240 Ostrea edulis, 240 Magallana gigas and 204 oysters of unsure species denotation sampled as a start pool for the image analysis project and for genetic evaluation of species belonging.</p>
Dataset for Automatic Refactoring Candidate Identification Leveraging Effective Code Representation
<p>The dataset consists of positive and negative case methods for Extract Method refactoring for selected GitHub repositories.<br> <br> <em>Each sample format - </em></p> <pre><code class="language-json">{ "repo_name": "...", "repo_url": "...", "positive_case_methods": ["...", "...", ...], "negative_case_methods": ["...", "...", ...] }</code></pre> <p> </p>
InsectSet47 & InsectSet66: Expanded datasets for automatic acoustic identification of insects (Orthoptera and Cicadidae)
<p><strong>Updated full version with training, validation and test sets.</strong></p> <p>Two newly compiled datasets for training neural networks to automatically identify insect species while comparing adaptive, waveform-based frontends to conventional mel-spectrogram frontends for audio feature extraction. This work was <a href="https://doi.org/10.1371/journal.pcbi.1011541">published in PLOS</a> Computational Biology and the machine learning implementations were published on <a href="https://github.com/mariusfaiss/InsectSet47-InsectSet66-Adaptive-Representations-of-Sound-for-Automatic-Insect-Recognition">Github</a>.</p> <p>These datasets expand on the previously published <a href="https://doi.org/10.5281/zenodo.7072196">InsectSet32</a> by including recently published collections of insect recordings by citizen scientists from around the world. Recordings from <a href="https://bio.acousti.ca/">BioAcoustica</a>, <a href="http://xeno-canto.org/">xeno-canto</a> and <a href="http://inaturalist.org/">iNaturalist</a>, as well as private collections by <a href="https://orcid.org/0000-0002-8929-2737">Baudewijn Odé</a> were downloaded and manually inspected. Files with strong noise interference or intense filtering, as well as files containing sounds of multiple species were removed to compile these datasets. The files were standardised to 44.1 kHz mono WAV files ranging in length from less than one second to several minutes. Files containing long periods without insect sounds were edited into multiple smaller files with silent periods no longer than 5 seconds. These files are marked as edits in the annotation file and should be assigned together into train/validation/test sets to prevent data leakage. The annotation files contain information for each recording, including the file name, species name and identifier, as well as the data subset they were included in for training the neural network (training, test, validation).</p> <p>InsectSet47 expands on <a href="https://doi.org/10.5281/zenodo.7072196">InsectSet32</a> with recordings from <a href="http://xeno-canto.org/">xeno-canto</a> and contains 1006 original recordings from 47 species, with at least ten files per species. The total length of InsectSet47 is 22 hours. InsectSet66 further expands on InsectSet47 by adding research-grade audio observations from <a href="http://inaturalist.org/">iNaturalist</a>, with a total of 1554 recordings from 66 species, a total length of over 24 hours and a minimum of ten files per species.</p> <p>The datasets were split into the training, validation and test sets while ensuring a roughly equal distribution of audio files and audio material for every species in all three subsets. This resulted in a 60/20/20 split (train/validation/test) by file number and a 64/19.5/16.5 split by file length.</p>
Automatic taxonomic identification based on the Fossil Image Dataset (>415,000 images) and deep convolutional neural networks
<p>This is a Fossil Image Dataset, which contains >415000 images. A total of 50 clades were labeled, with a final 90% accuracy. We used the web crawler to download fossil images from the Internet. We declare that all the collected images are used for academic purposes only. If anyone wants to use this dataset, please agree on the Terms of access for the Fossil Image Dataset (FID). We uploaded two datasets: FID (contains 0.415 million images) and reduced-FID (60 thousand images, 1200 for each clade). Requirements of necessary preinstalled Python libraries, algorithms for analysis, and the model weights are available at <a href="https://github.com/XiaokangLiuCUG/Fossil_Image_Dataset">https://github.com/XiaokangLiuCUG/Fossil_Image_Dataset</a>.</p>
Figure 2 in Using deep-learning for automatic identification of images of marine benthic macro-invertebrate bycatch: a proof of concept
Figure 2. – Three images of organisms obtained by cropping images of lots; from left to right: Chalinidae (Porifera), Polyclinidae (Chordata), Hormatidae (Cnidaria).
Figure 3 in Using deep-learning for automatic identification of images of marine benthic macro-invertebrate bycatch: a proof of concept
Figure 3. – Image of a batch of macro-invertebrate bycatch organisms from Kerguelen Exclusive Economical Zone (Poker 4 survey, 2017), including corals, a crinoïd, an ophiurid, a sea urchin and a brachiopoda; organisms are incomplete and have been quickly spread out over a small plate to take the picture.
Figure 6 in Using deep-learning for automatic identification of images of marine benthic macro-invertebrate bycatch: a proof of concept
Figure 6. – Example of detection and classification obtained with an image including an Ophiuroid, a piece of coral and a sea star with network 2; red squares and annotations have been provided by the computer with no human action.
Figure 5 in Using deep-learning for automatic identification of images of marine benthic macro-invertebrate bycatch: a proof of concept
Figure 5. – Example of detection and classification obtained with an image including Ascidians and a sea star with network 2; red squares and annotations have been provided by the computer with no human action.
Datasets for automatic acoustic identification of individual birds
<p>Bird individual audio recordings (foreground and background) to accompany the work:</p> <p><em><strong>"Automatic acoustic identification of individuals: Improving generalisation across species and recording conditions"</strong></em><br> by Dan Stowell, Tereza Petrusková, Martin Šálek, Pavel Linhart</p> <p><a href="https://royalsocietypublishing.org/doi/10.1098/rsif.2018.0940">https://royalsocietypublishing.org/doi/10.1098/rsif.2018.0940</a></p> <p><br> This dataset contains labelled recordings of individuals from three different bird species:</p> <ul> <li>Little owl</li> <li>Chiffchaff</li> <li>Tree Pipit</li> </ul> <p>For more information, please see the README.txt file, and the research article.</p> <p>The dataset takes approx 11 GB of disk space after the ZIP files have been uncompressed.</p> <p> </p>
A comparison of automatic cell identification methods for single-cell RNA-sequencing data
<p>Benchmark datasets used to evaluate the performance of 22 classifiers for cell type classification for scRNA-seq data</p>
Automatic Identification of Kidney Cell Types in scRNA-seq and snRNA-seq Data Using Machine Learning Algorithms - Datasets
<p>Datasets for reproducibility of the results found in Automatic Identification of Kidney Cell Types in scRNA-seq and snRNA-seq Data Using Machine Learning Algorithms. This study utilized data from the following 4 journals:</p> <p>Lake, B.B. et al. A single-nucleus RNA-sequencing pipeline to decipher the molecular anatomy and pathophysiology of human kidneys. Nat Commun 10, 2832 (2019).</p> <p>Liao, J., Yu, Z., Chen, Y. et al. Single-cell RNA sequencing of human kidney. Sci Data 7, 4 (2020).</p> <p>Menon, R. et al. Single cell transcriptomics identifies focal segmental glomerulosclerosis remission endothelial biomarker. JCI Insight 5, e133267 (2020).</p> <p>Wu, H. et al. Single-cell transcriptomics of a human kidney allograft biopsy specimen defines a diverse inflammatory response. J Am Soc Nephrol 29: 2069–2080 (2018).</p> <p>Young, M. D. et al. Single-cell transcriptomes from human kidneys reveal the cellular identity of renal tumors. Science 361, 594–599 (2018).</p>
Replication Package for the Paper: "Will Data Influence the Experiment Results?: A Replication Study of Automatic Identification of Decisions"
<p>This is the replication package for the paper: "Will Data Influence the Experiment Results?: A Replication Study of Automatic Identification of Decisions". It contains the source code and dataset of our experiment for the replication by other researchers. In the meanwhile, we provide brief description of the files in the replication package below.</p> <p><strong>1. main_code folder</strong></p> <ul> <li><em>automatic_approach.py </em>contains the main source code of the automatic approach for identifying decisions in our experiment, which is conducted on MacOs and Python 3.7.9. <strong>Note that you may get slightly</strong> <strong>different experiment results when conducting the experiments on different environment configurations.</strong></li> <li><em>requirement.txt</em> records all the installation packages and their version numbers needed for the current program to run. You can use "<em>pip install -r requirement.txt</em>" to rebuild the project and install all dependencies. <strong>Note that you may get slightly different experiment results when using different packages or versions. </strong></li> </ul> <p><strong>2. dataset folder</strong></p> <ul> <li><em>EASE2020 - 650 decisions.xlsx </em>contains 650 decision sentences from our previous work (EASE2020)</li> <li><em>EASE2020 - 650 non-decisions.xlsx </em>contains 650 non-decision sentences from our previous work (EASE2020)</li> <li><em>Our 844 relabeled decisions.xlsx</em> contains 844 relabeled decisions in this work.</li> <li><em>Our 750 assumptions.xlsx</em> contains 750 assumptions from our previous work (APSEC2019)</li> </ul> <p><strong>3. RQ1 folder</strong></p> <ul> <li><em>experiment_RQ1.py</em> contains the main source code of the experiment for answering RQ1, which is conducted on the same environment configuration as the <em>automatic_approach.py.</em></li> </ul> <p><strong>4. RQ2 folder</strong></p> <ul> <li><em>experiment_RQ2.py</em> contains the main source code of the experiment for answering RQ2, which is conducted on the same environment configuration as the <em>automatic_approach.py.</em></li> </ul> <p><strong>5. RQ3 folder</strong></p> <ul> <li><em>experiment_RQ3.py</em> contains the main source code of the experiment for answering RQ3, which is conducted on the same environment configuration as the <em>automatic_approach.py.</em></li> </ul>
Dataset and additional files/softwares required for the paper "LeSICiN: A Heterogeneous Graph-based Approach for Automatic Legal Statute Identification from Indian Legal Documents"
<p>This dump contains all files and softwares required for running the codes for the paper "LeSICiN: A Heterogeneous Graph-based Approach for Automatic Legal Statute Identification from Indian Legal Documents". Specifically, these codes are available at https://github.com/Law-AI/LeSICiN.</p> <p>LeSICiN is a deep neural network for the task of Legal Statute Identification which also uses graphical properties of the document-statute citation network for training and predictions.</p> <p>We have three datasets --- train, dev and test. These are all .jsonl files with each instance dict per line; each instance dict contains the unique id, list of sentences and cited labels of the particular instance. Also, there is a fourth file --- secs.jsonl, which stores the text of all the statutes in similar format.</p> <p>schemas.json list out the metapath schemas for fact and section type nodes, while type_map.json maps the id of each node to its type (Act/Chapter/Topic/Section/Fact). </p> <p>label_tree.json and citation_network.json list out the edges for the two parts of the network in the format of a 3-tuple ('source id', 'relationship type', 'target id')</p> <p>"ils2v.bin" is the pretrained sent2vec vectorizer that can generate a 200-dim vector for each sentence</p>
Data for publication 'Recreational vessels without Automatic Identification System (AIS) dominate anthropogenic noise contributions to a shallow water soundscape' (Scientific Reports 2019)
<p>Data on vessel tracks and underwater noise levels presented in the publication Hermannsen, L., Mikkelsen, L., Tougaard, J., Beedholm, K., Johnson, M. and P. T. Madsen, "Recreational vessels without Automatic Identification System (AIS) dominate anthropogenic noise contributions to a shallow water soundscape", Scientific Reports 9:15477 (<a href="https://doi.org/10.1038/s41598-019-51222-9">https://doi.org/10.1038/s41598-019-51222-9</a>).</p>
LeaData: a novel reference data of digital microscopic leather images for automatic species identification
Open the record for dataset details and reuse information.
Limitations of Human Identification of Automatically Generated Text: Dataset
<p>This repository contains the data for the paper:</p> <p><span>Nadège Alavoine, Maximin Coavoux, Emmanuelle Esperança-Rodier, Romane Gallienne, Carlos-Emiliano González-Gallardo, Jérôme Goulian, Jose G. Moreno, Aurélie Névéol, Didier Schwab, Vincent Segonne, and Johanna Simoens. 2024. <a href="https://aclanthology.org/2024.lrec-main.919">Limitations of Human Identification of Automatically Generated Text</a>. In <em>Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)</em>, pages 10511–10516, Torino, Italia. ELRA and ICCL.</span></p>
Automatic taxonomic identification based on the Fossil Image Dataset (>415,000 images) and deep convolutional neural networks
<p>The rapid and accurate taxonomic identification of fossils is of great significance in paleontology, biostratigraphy, and other fields. However, taxonomic identification is often labor-intensive and tedious, and the requisition of extensive prior knowledge about a taxonomic group also requires long-term training. Moreover, identification results are often inconsistent across researchers and communities. Accordingly, in this study, we used deep learning to support taxonomic identification. We used web crawlers to collect the Fossil Image Dataset (FID) via the Internet, obtaining 415,339 images belonging to 50 fossil clades. Then we trained three powerful convolutional neural networks on a high-performance workstation. The Inception ResNet v2 architecture achieved an average accuracy of 0.90 in the test dataset when transfer learning was applied. The clades of microfossils and vertebrate fossils exhibited the highest identification accuracies of 0.95 and 0.90, respectively. In contrast, clades of sponges, bryozoans, and trace fossils with various morphologies or with few samples in the dataset exhibited a performance below 0.80. Visual explanation methods further highlighted the discrepancies among different fossil clades and suggested similarities between the identifications made by machine classifiers and taxonomists. Collecting large paleontological datasets from various sources, such as the literature, digitization of dark data, citizen-science data, and public data from the Internet may further enhance deep learning methods and their adoption. Such developments will also possibly lead to image-based systematic taxonomy to be replaced by machine-aided classification in the future. Pioneering studies can include microfossils and some invertebrate fossils. To contribute to this development, we deployed our model on a server for public access at www.ai-fossil.com.</p>
Figure 3 in Automatic identification of bird females using egg phenotype
Figure 3. Correlation between spectral (A), pattern/luminance (B) and shape (C) distances, respectively and genetic distances. Individual phenotypic distances of average eggs laid by nine genotyped common cuckoo females: spectral (D), pattern/luminance (E) and shape (F) distances.
Figure 1 in Automatic identification of bird females using egg phenotype
Figure 1. Values for individual eggs on the two most important PC variables (according to the random forest model), grouped by cuckoo female ID based on the genetic assignment. PCA2 pattern variable indicates egg skew and PC2 spectra variable indicates blueness/greenness of eggs (for details, see Table 2).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.