Interpretable Benchmarking of Classical Machine Learning Models for Phishing Website Detection
Résumé fourni par la source
This preprint presents an interpretable benchmarking study of classical machine-learning models for phishing website detection. Using the UCI Phishing Websites dataset, which contains 11,055 observations and 30 features, several classification approaches were evaluated through stratified cross-validation using ROC-AUC and complementary performance metrics. The results indicate that the HistGradientBoosting classifier achieved the strongest predictive performance, with a mean cross-validated ROC-AUC of 0.9966. The study also examines feature importance and discusses the practical value and limitations of automated phishing detection. The analysis is intended to provide a transparent and reproducible baseline for future cybersecurity research. This document is a preprint and has not undergone peer review.