COMPARATIVE ANALYSIS AND OPTIMISATION OF SUPERVISED MACHINE LEARNING ALGORITHMS FOR REAL-TIME PHISHING URL DETECTION
Keywords:
Phishing URL, supervised machine learning, real-time detection, algorithm optimisation, inference latency, Random Forest, XGBoost, SVM, feature selectionAbstract
Real-time phishing URL detection differs from offline classification because a useful detector must make accurate decisions within a strict computational budget. This analytical paper compares supervised machine learning algorithms from the joint perspectives of classification quality, feature-processing cost, inference latency, calibration, memory demand, interpretability, and operational false-positive risk. The analysis deliberately differs from feature-family-centred phishing studies by treating the model-selection and optimisation problem as the main research object. Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, k-Nearest Neighbour, Naive Bayes, AdaBoost and gradient-boosted trees are examined as candidate classifiers. Recent public URL datasets demonstrate that large-scale supervised benchmarking is feasible: LegitPhish contains more than 100,000 labelled URLs with engineered structural and lexical attributes, while another 2022 dataset contains 111,660 URLs with 22 numerical lexical and structural features. Recent optimisation research also shows that reducing feature counts can lower training complexity while preserving useful predictive performance. The paper proposes a latency-aware evaluation protocol in which feature extraction, preprocessing and model inference are timed separately; predictive metrics are assessed alongside false-positive rate and throughput; and model optimisation uses feature pruning, hyperparameter search, class weighting, calibration and threshold tuning.




