A COMPREHENSIVE MACHINE LEARNING FRAMEWORK FOR PHISHING URL DETECTION USING LEXICAL, HOST-BASED AND DOMAIN-BASED FEATURES

Authors

  • Shipra Goyal Research Scholar, Department of Computer Science, Sabarmati University, Ahmedabad, Gujarat Author
  • Dr. Rajeev Yadav (Associate Professor) Research Supervisor, Department of Computer Science, Sabarmati University, Ahmedabad, Gujarat Author

Keywords:

Phishing URL, machine learning, lexical features, host-based features, domain-based features, cybersecurity, Random Forest, XGBoost, feature engineering

Abstract

Phishing URLs remain a persistent cybersecurity threat because attackers can create new domains, manipulate URL strings, imitate trusted brands, and rapidly rotate hosting infrastructure. This analytical study develops a comprehensive machine learning framework for phishing URL detection by integrating three complementary feature families: lexical features derived directly from the URL string, host-based features describing network and server characteristics, and domain-based features describing registration, DNS, top-level-domain, age, and reputation signals. The study follows a literature-grounded analytical design rather than claiming newly collected primary observations. Recent public datasets and peer-reviewed studies are examined to identify robust feature groups, suitable preprocessing procedures, model families, evaluation criteria, and deployment constraints. The analysis compares Logistic Regression, Support Vector Machine, Decision Tree, Random Forest, Gradient Boosting/XGBoost, k-Nearest Neighbour, and Naive Bayes from the perspectives of predictive capability, computational cost, interpretability, and real-time suitability. Evidence from recent datasets shows that lexical features are fast and broadly available, while host- and domain-based features add contextual information that can reduce ambiguity in structurally deceptive URLs. Tree ensembles are especially well suited to mixed numerical and categorical security features, but high reported accuracy on a single dataset should not be treated as proof of generalization. Cross-dataset validation, temporal splitting, class-imbalance control, probability calibration, explainability, and false-positive analysis are therefore incorporated into the proposed framework.

0 0

Downloads

Published

2023-07-28

How to Cite

Shipra Goyal, & Dr. Rajeev Yadav (Associate Professor). (2023). A COMPREHENSIVE MACHINE LEARNING FRAMEWORK FOR PHISHING URL DETECTION USING LEXICAL, HOST-BASED AND DOMAIN-BASED FEATURES. International Journal of Arts, Commerce & Education, 11(2), 21-32. https://avjournals.com/index.php/files/article/view/75