This work presents a simple and effective oversampling method based on k-means clustering and SMOTE (synthetic minority oversampled technique), which avoids the generation of noise and effectively overcomes imbalances between and within classes.
Authors
F. Bação
2 papers
F. Last
1 papers
Georgios Douzas
1 papers
References52 items
1
Clustering-based undersampling in class-imbalanced data
2
Noise Reduction A Priori Synthetic Over-Sampling for class imbalanced data sets
3
An empirical comparison of techniques for the class imbalance problem in churn prediction
4
Self-Organizing Map Oversampling (SOMO) for imbalanced data set learning
5
CURE-SMOTE algorithm and hybrid algorithm for feature selection and parameter optimization based on random forests
6
Imbalanced-learn: A Python Toolbox to Tackle the Curse of Imbalanced Datasets in Machine Learning
7
Ordering-based pruning for improving the performance of ensembles of classifiers in the framework of imbalanced datasets
8
A bi-directional sampling based on K-means method for imbalance text classification
9
Adaptive semi-unsupervised weighted oversampling (A-SUWO) for imbalanced datasets
10
Diversity techniques improve the performance of the best imbalance learning ensembles
11
A new cluster-based oversampling method for improving survival prediction of hepatocellular carcinoma patients
12
SMOTE-IPF: Addressing the noisy and borderline examples problem in imbalanced classification by a re-sampling method with filtering
13
Assessment Metrics for Imbalanced Learning
14
Analysing the classification of imbalanced data-sets with multiple classes: Binarization techniques and ad-hoc approaches
15
A Review on Ensembles for the Class Imbalance Problem: Bagging-, Boosting-, and Hybrid-Based Approaches
16
Scikit-learn: Machine Learning in Python
17
Web-scale k-means clustering
18
COG: local decomposition for rare class analysis
19
Learning from Imbalanced Data
20
Safe-Level-SMOTE: Safe-Level-Synthetic Minority Over-Sampling TEchnique for Handling the Class Imbalanced Problem
21
Machine Learning from Imbalanced Data Sets 101
22
Automatically countering imbalance and its empirical relationship to cost
23
Robustness of learning techniques in handling class noise in imbalanced datasets
24
Statistical Comparisons of Classifiers over Multiple Data Sets
25
Combating imbalance in network intrusion datasets
26
Borderline-SMOTE: A New Over-Sampling Method in Imbalanced Data Sets Learning
27
Learning with Class Skews and Small Disjuncts
28
Class imbalances versus small disjuncts
29
Editorial: special issue on learning from imbalanced data sets
30
A study of the behavior of several methods for balancing machine learning training data
31
The Problem of Overfitting
32
A Multiple Resampling Method for Learning from Imbalanced Data Sets
33
Greedy function approximation: A gradient boosting machine.
34
MetaCost: a general method for making classifiers cost-sensitive
Concept Learning and the Problem of Small Disjuncts
37
Generalized Linear Models
38
The Use of Ranks to Avoid the Assumption of Normality Implicit in the Analysis of Variance
39
THE DARWIN CELEBRATION AT CAMBRIDGE.
40
Imbalanced classification in sparse and large behaviour datasets
41
Hellinger distance decision trees are robust and skew-insensitive
42
KEEL Data-Mining Software Tool: Data Set Repository, Integration of Algorithms and Experimental Analysis Framework
43
Cost-Sensitive Learning vs. Sampling: Which is Best for Handling Unbalanced Classes with Unequal Error Costs?
44
Handling imbalanced datasets: A review
45
Data Mining for Imbalanced Datasets: An Overview
46
Design of experiments for the NIPS 2003 variable selection benchmark
47
SMOTE: Synthetic Minority Over-sampling Technique
48
Using Unsupervised Learning to Guide Resampling in Imbalanced Data Sets
49
A Simple Sequentially Rejective Multiple Test Procedure
50
Some methods for classification and analysis of multivariate observations
51
Some methods for classi cation and analysis of multivariate observations
52
To obtain a measure of density, divide each cluster’s number of minority instances by its average minority distance raised to the power of the number of features