Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
345
datasets available to search
ShareScore release 0.9.0
Dataset results
345 results for “classification analysis”
Meta-analysis and gender classification of 914 national and international surveys in six European countries (2000-2023)
<p><span>This data frame presents the results of a quan</span><span>ti</span><span>ta</span><span>ti</span><span>ve content analysis of the occurrence of gender‐based concepts, themes, issues, and solu</span><span>ti</span><span>ons within large‐scale poli</span><span>ti</span><span>cal and sociological survey ques</span><span>ti</span><span>onnaires fielded cross‐na</span><span>ti</span><span>onally in Europe and in six European countries: Denmark, Germany, Hungary, Switzerland and the UK, spanning 2000‐2023. Data was collected by teams from each country between September 2023‐January 2024. Teams collected ques</span><span>ti</span><span>ons in the original language and provided a transla</span><span>ti</span><span>on into English. Analysis was conducted using the translated text. The unit of analysis (‘CODING_UNIT_TEXT’) was the individual 'gender‐related argument' within a survey ques</span><span>ti</span><span>on. This could be the en</span><span>ti</span><span>re survey ques</span><span>ti</span><span>on, a sub‐ques</span><span>ti</span><span>on (in the case of matrix ques</span><span>ti</span><span>ons), or a singular response op</span><span>ti</span><span>on (for mul</span><span>ti</span><span>ple choice ques</span><span>ti</span><span>ons). Coding units were coded in three key domains:(1) Gender concepts, (2) Themes/issues, and (3) Solu</span><span>ti</span><span>ons. Up to two Themes/Issues and Solu</span><span>ti</span><span>ons could be coded per coding unit. Several coding categories within the Themes/Issues and Solu</span><span>ti</span><span>ons domains func</span><span>ti</span><span>on hierarchically, where a coder first assigned a higher‐level category and then as many subcategories as applicable. For example, a ques</span><span>ti</span><span>on concerning government‐funded childcare is coded as B1_Economy ‐> B1_4_LabourMarket ‐> B1_4_1_CareWork ‐> B1_4_1_3_Childcare. The corresponding codebook presents the uni</span><span>ti</span><span>sa</span><span>ti</span><span>on process and coding categories in full detail.</span></p>
Data of the article Analysis of the self-archiving policies of journals in the highest rank category of the Finnish journal classification system within computer science, physics and electronic engineering
<p>The publication forum level three journals representing the three fields of science of computer science, computer science and electrical engineering were identified by utilizing the MinEdu field search filter while searching for the top-ranked journals from the publication channel search (https://www.tsv.fi/julkaisufoorumi/haku.php?lang=en), which is based on Field of Science, Statistics Finland classification (https://www.stat.fi/meta/luokitukset/tieteenala/001-2010/index_en.html). The data were extracted during august 2017 consists of total of 127 individual journals. It is worth noting that circa 30 journals were classified into more than one fields of sciences under scrutiny. First, the journals were divided into representing gold and hybrid model journals. Second, green open access policies of the identified hybrid journals were analyzed using Laakso’s (2014) publisher policy coding framework. Also publishers of the individual journals were identified and subsequently added to the data.</p> <p>NOTE! The data includes the shortest embargo to either institutional or subject repositories. For example, Elsevier had no embargo to opening accepted manuscripts from arXiv subject repository and thus no embargoes to Elsevier's journals are included within this datasheet.</p> <p>Data is in CSV. format</p> <p> </p> <p> </p>
Datasets and results from: "Random Forest Classification and Solar Flares Data: Analysis and Validation"
<p><strong>Instructions for the data and code repository</strong></p> <p>Results, post-processing workflow, and datasets for the research paper titled "Random Forest Classification and Solar Flares Data: Analysis and Validation".</p> <p>The folder contains three .csv files: the complete dataset (dataset.csv), the balanced training dataset (train_dataset.csv), and the testing dataset (test_dataset.csv).</p> <p>The folder also contains the result files from the research (.csv output files with predictions and .html files with evaluation metrics, etc.) exported from the JASP software. The number in each file name corresponds to the number of trees utilized in Random Forest modelling.</p> <p>In addition, the Python script for the post-processing workflow is provided, with comments located in the script.</p> <p>The soft range X-ray irradiance and VLF amplitude data were obtained from:<br> National Centers for Environmental Information (NCEI) Available online: https://www.ncei.noaa.gov/. Accessed on: 24th June 2023. <br> Worldwide archive of low-frequency data and observations (WALDO) Available online: https://waldo.world/. Accessed on: 24th June 2023.</p>
Data for paper titled : Comparing Clothing-Mounted Sensors with Wearable Sensors for Movement Analysis and Activity Classification (published in Sensors (MDPI))
<p>Data for paper titled : Comparing Clothing-Mounted Sensors with Wearable Sensors for Movement Analysis and Activity Classification (published in Sensors (MDPI))</p>
Animal Recognition Using Methods Of Fine-Grained Visual Analysis - YOLOv5 Breed Classification Dataset (Oxford-IIIT Pet)
<p>Oxford-IIIT Pet Dataset with ground truth labels for breeds (from https://public.roboflow.com/object-detection/oxford-pets).</p>
Sentiment Analysis of RUU PDP with Naive Bayes, Support Vector Machine, and Random Forest Classification Algorithm
<p>Dataset from the results of data crawling via Twitter which discusses the Rancangan Undang Undang Pelindungan Data Pribadi to be used in the sentiment analysis process. The dataset is divided into several parts according to the process executed on RapidMiner.</p>
Assessment of risk and vulnerability to disasters by field protocols and classification trees: an analysis of the Municipal Risk Reduction Plans (MRRP) of São Bernardo do Campo and Franco da Rocha (2020-2021), in Brazil
<p><span>Most disaster risk assessment methodologies were developed by natural science professionals. The UN Sendai Framework confirms the need to include these methods as interweaving complex networks of socially vulnerable financial processes, especially in developing countries. This work aimed to analyze the experience of including social vulnerability information in field protocols of the Municipal Risk Reduction Plans (MRRP) of São Bernardo do Campo and Franco da Rocha (2020-2021). A georeferenced database was structured with data from field protocols on risk sectors and then the relationship between vulnerability and risk was analyzed using the classification tree method. The results show that the geotechnical and social vulnerability variables relevant to risk attribution are largely overlapping over the sectors, but that the integrated analysis of the two dimensions allows a better interpretation of the risk level. However, the models with information from field protocols cannot replace a broader holistic assessment by the specialist in the field. Finally, qualitative reflections are made on the limitations and potential of including aspects of vulnerability in the specific case study.</span></p>
FIGURE 6. Classification Tree results and predictions for the five fossil localities. A in What are the best modern analogs for ancient South American mammal communities? Evidence from ecological diversity analysis (EDA)
FIGURE 6. Classification Tree results and predictions for the five fossil localities. A) Results and predictions for CT1, vegetative cover. B) Results and predictions for CT2, biogeographic realm. Abbreviations: LV, La Venta; QH, Quebrada Honda; RU, Rümikon; SC, Santa Cruz; TG, Tinguiririca.
Fig. 3 in Phylogenetic analysis of the tribe Neanurini questions tribal classification of the subfamily Neanurinae (Collembola: Neanuridae)
Fig. 3 Unambiguous morphological character optimisation obtained from an analysis of the data (Appendix 2) under implied weights (k = 17). The numbers above and below circles on the branches show character numbers and states, respectively. White and black circles represent homoplasious and non-homoplasious character state transformations, respectively
Fig. 1 Strict consensus cladogram obtained from 45 in Phylogenetic analysis of the tribe Neanurini questions tribal classification of the subfamily Neanurinae (Collembola: Neanuridae)
Fig. 1 Strict consensus cladogram obtained from 45 most parsimonious trees under equal weights. Values of Jackknife support and symmetric resampling are indicated on and below branches, respectively. Only values above 40 are indicated to facilitate the visualisation of the most internal branches. The main clades are indicated with letters (a–e) on branches
Fig. 3 Trees obtained under the implied weighting using three concavity values k in First phylogenetic analysis of the tribe Oligaphorurini (Collembola: Onychiuridae) inferred from morphological data, with implications for generic classification
Fig. 3 Trees obtained under the implied weighting using three concavity values k = 6 (a), 9 (b), and 12 (c)
Fig. 4 Abdominal sternite IV in First phylogenetic analysis of the tribe Oligaphorurini (Collembola: Onychiuridae) inferred from morphological data, with implications for generic classification
Fig. 4 Abdominal sternite IV showing organization of furcal remnant. a, b Oligaphorura ursi Fjellberg, 1984; c Micraphorura gamae Buşmachiu and Weiner, 2013; d Oligaphorura groenlandica (Tullberg, 1876); e Dimorphaphorura inya Weiner and Kaprus, 2014; f Protaphorura eichhorni (Gisin, 1954)
Figure 3 in Revised classification design of the Anatolian species of Nannospalax (Rodentia: Spalacidae) using RFLP analysis
Figure 3. Neighbor-joining and span tree showing genetic relationships among populations, based on Nei's genetic distance measure.
Figure 1 in Revised classification design of the Anatolian species of Nannospalax (Rodentia: Spalacidae) using RFLP analysis
Figure 1. Sampling localities of Nannospalax xanthodon and Nannospalax ehrenbergi from Turkey for molecular studies. Names for numbered localities indicated in Table 1.
Figure 2 in Cladistic analysis and a revised classification of fossil and recent mysticetes
Figure 2. Previously proposed phylogenies of extinct mysticetes previously referred to Cetotheriidae. A, cladogram from Geisler & Sanders (2003). B, proposed phylogeny from Sanders & Barnes (2002a). C, cladogram from Geisler & Luo (1996). D, cladogram from Kimura & Ozawa (2002). E, cladogram from Bouetel & de Muizon (2006). In the case where the taxon is listed as [Odontoceti], the order was represented by two to several species/genera of odontocetes in the original analyses.
Arc fault detection and appliances classification in AC home electrical networks using Recurrence Quantification Plots and Image Analysis
<p>The data provided can be used for the development of methods for the detection of arcing faults in a domestic low-voltage electrical networks (230V - 50 Hz). The data files are current and voltage signatures experimentally measured.</p> <p>The ReadMe file describes :</p> <p>- the test set up and the the procedure followed to make the measurements</p> <p>- the list of household appliances and their main characteristics.</p> <p>- the name of the data files</p> <p>- the type of arcing faults</p> <p> </p>
MedMNIST Classification Decathlon: A Lightweight AutoML Benchmark for Medical Image Analysis
<p>This data repository for MedMNIST v1 is out of date! Please check the <a href="http://medmnist.github.io">latest version</a> of MedMNIST v2. </p> <p> </p> <p><strong>Abstract</strong></p> <p>We present MedMNIST, a collection of 10 pre-processed medical open datasets. MedMNIST is standardized to perform classification tasks on lightweight 28x28 images, which requires no background knowledge. Covering the primary data modalities in medical image analysis, it is diverse on data scale (from 100 to 100,000) and tasks (binary/multi-class, ordinal regression and multi-label). MedMNIST could be used for educational purpose, rapid prototyping, multi-modal machine learning or AutoML in medical image analysis. Moreover, MedMNIST Classification Decathlon is designed to benchmark AutoML algorithms on all 10 datasets; We have compared several baseline methods, including open-source or commercial AutoML tools. The datasets, evaluation code and baseline methods for MedMNIST are publicly available at <a href="https://medmnist.github.io/">https://medmnist.github.io/</a>.</p> <p> </p> <p>Please note that this dataset is <strong>NOT</strong> intended for clinical use.</p> <p> </p> <p>We recommend our official <a href="https://github.com/MedMNIST/MedMNIST">code</a> to download, parse and use the MedMNIST dataset:</p> <blockquote> <pre>pip install medmnist</pre> </blockquote> <p> </p> <p><strong>Citation and Licenses</strong></p> <p>If you find this project useful, please cite our ISBI'21 paper as:<br> <em> Jiancheng Yang, Rui Shi, Bingbing Ni. "MedMNIST Classification Decathlon: A Lightweight AutoML Benchmark for Medical Image Analysis," arXiv preprint arXiv:2010.14925, 2020.</em><br> <br> or using bibtex:<br> <em> @article{medmnist,<br> title={MedMNIST Classification Decathlon: A Lightweight AutoML Benchmark for Medical Image Analysis},<br> author={Yang, Jiancheng and Shi, Rui and Ni, Bingbing},<br> journal={arXiv preprint arXiv:2010.14925},<br> year={2020}<br> }</em></p> <p>Besides, please cite the corresponding paper if you use any subset of MedMNIST. Each subset uses the <strong>same license</strong> as that of the source dataset.</p> <p> </p> <p><strong>PathMNIST</strong></p> <p>Jakob Nikolas Kather, Johannes Krisam, et al., "Predicting survival from colorectal cancer histology slides using deep learning: A retrospective multicenter study," PLOS Medicine, vol. 16, no. 1, pp. 1–22, 01 2019.</p> <p><em><strong>License</strong>: <a href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a></em></p> <p> </p> <p><strong>ChestMNIST</strong></p> <p>Xiaosong Wang, Yifan Peng, et al., "Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases," in CVPR, 2017, pp. 3462–3471.</p> <p><em><strong>License</strong>: <a href="https://creativecommons.org/publicdomain/zero/1.0/">CC0 1.0</a></em></p> <p> </p> <p><strong>DermaMNIST</strong></p> <p>Philipp Tschandl, Cliff Rosendahl, and Harald Kittler, "The ham10000 dataset, a large collection of multisource dermatoscopic images of common pigmented skin lesions," Scientific data, vol. 5, pp. 180161, 2018.</p> <p>Noel Codella, Veronica Rotemberg, Philipp Tschandl, M. Emre Celebi, Stephen Dusza, David Gutman, Brian Helba, Aadi Kalloo, Konstantinos Liopyris, Michael Marchetti, Harald Kittler, and Allan Halpern: “Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC)”, 2018; arXiv:1902.03368.</p> <p><em><strong>License</strong>: <a href="https://creativecommons.org/licenses/by-nc/4.0/">CC BY-NC 4.0</a></em></p> <p> </p> <p><strong>OCTMNIST/PneumoniaMNIST</strong></p> <p>Daniel S. Kermany, Michael Goldbaum, et al., "Identifying medical diagnoses and treatable diseases by image-based deep learning," Cell, vol. 172, no. 5, pp. 1122 – 1131.e9, 2018.</p> <p><em><strong>License</strong>: <a href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a></em></p> <p> </p> <p><strong>RetinaMNIST</strong></p> <p>DeepDR Diabetic Retinopathy Image Dataset (DeepDRiD), "The 2nd diabetic retinopathy – grading and image quality estimation challenge," https://isbi.deepdr.org/data.html, 2020.</p> <p><em><strong>License</strong>: <a href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a></em></p> <p> </p> <p><strong>BreastMNIST</strong></p> <p>Walid Al-Dhabyani, Mohammed Gomaa, Hussien Khaled, and Aly Fahmy, "Dataset of breast ultrasound images," Data in Brief, vol. 28, pp. 104863, 2020.</p> <p><em><strong>License</strong>: <a href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a></em></p> <p> </p> <p><strong>OrganMNIST_{Axial,Coronal,Sagittal}</strong></p> <p>Patrick Bilic, Patrick Ferdinand Christ, et al., "The liver tumor segmentation benchmark (lits)," arXiv preprint arXiv:1901.04056, 2019.</p> <p>Xuanang Xu, Fugen Zhou, et al., "Efficient multiple organ localization in ct image using 3d region proposal network," IEEE Transactions on Medical Imaging, vol. 38, no. 8, pp. 1885–1898, 2019.</p> <p><em><strong>License</strong>: <a href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a></em></p> <div> <div class="gtx-trans-icon"> </div> </div>
Classification and structural analysis of value chain contracts for biodiversity conservation in the European Union
<p><span>The data provided by this dataset are the raw data published in the paper "<strong>Classification and structural analysis of value chain contracts for biodiversity conservation in the European Union</strong>" (<a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.sftr.2024.100372" target="_blank" rel="noopener"><span><span>https://doi.org/10.1016/j.sftr.2024.100372</span></span></a>). </span></p>
Data Regarding Classification of Infrasonic Atmospheric Events Using Electromagnetic Pulse Analysis
<p> </p> <div>The following data files were used for the analysis presented in the paper </div> <div>"Classification of Infrasonic Atmospheric Events Using Electromagnetic Pulse Analysis"</div> <div> </div> <div>The files include details of the infrasonicly detected evnents, and features extracted from electromagnetic signals, as explined in the README file.</div>
Supplementary material for "Classification of Eurasian Watermilfoil (Myriophyllum spicatum) Using Drone-enabled Multispectral Imagery Analysis" paper
<p>This material has the unsummarized versions of the water chemistry and light data collected for each site, by date. It also includes all error matrices for each classification. It also includes details of each classification site and result. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.