Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2
datasets available to search
ShareScore release 0.9.0
Dataset results
2 results for “Imbalanced Data”
Supporting datasets PubFig05 for: "Heterogeneous Ensemble Combination Search using Genetic Algorithm for Class Imbalanced Data Classification"
<p><strong>Faces Dataset: PubFig05</strong></p> <p>This is a subset of the ''PubFig83'' dataset [1] which provides 100 images each of 5 most difficult celebrities to recognise (referred as class in the classification problem). For each celebrity persons, we took 100 images and separated them into training and testing sets of 90 and 10 images, respectively:</p> <p><strong>Person: </strong>Jenifer Lopez; Katherine Heigl; Scarlett Johansson; Mariah Carey; Jessica Alba</p> <p> </p> <p><strong>Feature Extraction</strong></p> <p>To extract features from images, we have applied the HT-L3-model as described in [2] and obtained 25600 features.</p> <p><strong>Feature Selection</strong></p> <p>Details about feature selection followed in brief as follows:</p> <ol> <li> <p><strong>Entropy Filtering:</strong> First we apply an implementation of Fayyad and Irani's [3] entropy base heuristic to discretise the dataset and discarded features using the minimum description length (MDL) principle and only 4878 passed this entropy based filtering method.</p> </li> <li> <p><strong>Class-Distribution Balancing:</strong> Next, we have converted the dataset to binary-class problem by separating into 5 binary-class datasets using one-vs-all setup. Hence, these datasets became <em>imbalanced</em> at a ratio of 1:4. Then we converted them into <em>balanced binary-class</em> datasets using random sub-sampled method. Further processing of the dataset has been described in the paper.</p> </li> <li> <p><strong>(alpha,beta)-k Feature selection:</strong> To get a good feature set for training the classifier, we select the features using the approach based on the (alpha,beta)-k feature selection [4] problem. It selects a minimum subset of features that maximise both within class similarity and dissimilarity in different classes. We applied the entropy filtering and (alpha,beta)-k feature subset selection methods in three ways and obtained different numbers of features (in the Table below) after consolidating them into binary class dataset.</p> </li> </ol> <ul> <li> <p><strong>UAB:</strong> We applied (alpha,beta)-k feature set method on each of the balanced binary-class datasets and we took the <em>union</em> of selected features for each binary-class datasets. Finally, we applied the (alpha,beta)-k feature set selection method on each of the binary-class datasets and get a set of features.</p> </li> <li> <p><strong>IAB:</strong> We applied (alpha,beta)-k feature set method on each of the balanced binary-class datasets and we took the <em>intersection</em> of selected features for each binary-class datasets. Finally, we applied the (alpha,beta)-k feature set selection method on each of the binary-class datasets and get a set of features.</p> </li> <li> <p><strong>UEAB:</strong> We applied (alpha,beta)-k feature set method on each of the balanced binary-class datasets. Then, we applied the entropy filtering and (alpha,beta)-k feature set selection method on each of the balanced binary-class datasets. Finally, we took the <em>union</em> of selected features for each <em>balanced binary-class</em> datasets and get a set of features.</p> </li> </ul> <p>All of these datasets are inside the compressed folder. It also contains the document describing the process detail.</p> <p> </p> <p><strong>References</strong></p> <p>[1] Pinto, N., Stone, Z., Zickler, T., & Cox, D. (2011). Scaling up biologically-inspired computer vision: A case study in unconstrained face recognition on facebook. In Computer Vision and Pattern Recognition Workshops (CVPRW), 2011 IEEE Computer Society Conference on (pp. 35–42).</p> <p>[2] Cox, D., & Pinto, N. (2011). Beyond simple features: A large-scale feature search approach to unconstrained face recognition. In Automatic Face Gesture Recognition and Workshops (FG 2011), 2011 IEEE International Conference on (pp. 8–15).</p> <p>[3] Fayyad, U. M., & Irani, K. B. (1993). Multi-Interval Discretization of Continuous-Valued Attributes for Classification Learning. In International Joint Conference on Artificial Intelligence (pp. 1022–1029).</p> <p>[4] Berretta, R., Mendes, A., & Moscato, P. (2005). Integer programming models and algorithms for molecular classification of cancer from microarray data. In Proceedings of the Twenty-eighth Australasian conference on Computer Science - Volume 38 (pp. 361–370). 1082201: Australian Computer Society, Inc.</p> <p> </p>
Data from: The influence of balanced and imbalanced resource supply on biodiversity-functioning relationship across ecosystems
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.