Using semi‐supervised classifiers for credit scoring

0.00 Avg rating—0 Votes

Article ID:	iaor20132534
Volume:	64
Issue:	4
Start Page Number:	513
End Page Number:	529
Publication Date:	Apr 2013
Journal:	Journal of the Operational Research Society
Authors:	Kennedy K, Namee B Mac, Delany S J
Keywords:	classification, credit scoring, portfolio analysis

Abstract:

In credit scoring, low‐default portfolios (LDPs) are those for which very little default history exists. This makes it problematic for financial institutions to estimate a reliable probability of a customer defaulting on a loan. Banking regulation (Basel II Capital Accord), and best practice, however, necessitate an accurate and valid estimate of the probability of default. In this article the suitability of semi‐supervised one‐class classification (OCC) algorithms as a solution to the LDP problem is evaluated. The performance of OCC algorithms is compared with the performance of supervised two‐class classification algorithms. This study also investigates the suitability of over sampling, which is a common approach to dealing with LDPs. Assessment of the performance of one‐ and two‐class classification algorithms using nine real‐world banking data sets, which have been modified to replicate LDPs, is provided. Our results demonstrate that only in the near or complete absence of defaulters should semi‐supervised OCC algorithms be used instead of supervised two‐class classification algorithms. Furthermore, we demonstrate for data sets whose class labels are unevenly distributed that optimising the threshold value on classifier output yields, in many cases, an improvement in classification performance. Finally, our results suggest that oversampling produces no overall improvement to the best performing two‐class classification algorithms.

Reviews

Required fields are marked *. Your email address will not be published.