An effective detection approach for phishing websites using URL and HTML features

Ali Aljofey; Qingshan Jiang; Abdur Rasool; Hui Chen; Wenyin Liu; Qiang Qu; Yang Wang

doi:10.1038/s41598-022-10841-5

An effective detection approach for phishing websites using URL and HTML features

Sci Rep. 2022 May 25;12(1):8842. doi: 10.1038/s41598-022-10841-5.

Authors

Ali Aljofey^{1

2}, Qingshan Jiang³, Abdur Rasool^{1

2}, Hui Chen^{1

2}, Wenyin Liu⁴, Qiang Qu¹, Yang Wang⁵

Affiliations

¹ Shenzhen Key Laboratory for High Performance Data Mining, Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences, Shenzhen, 518055, China.
² Shenzhen College of Advanced Technology, University of Chinese Academy of Sciences, Beijing, 100049, China.
³ Shenzhen Key Laboratory for High Performance Data Mining, Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences, Shenzhen, 518055, China. qs.jiang@siat.ac.cn.
⁴ Department of Computer Science, Guangdong University of Technology, Guangzhou, China.
⁵ Cloud Computing Center, Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences, Shenzhen, 518055, China.

Abstract

Today's growing phishing websites pose significant threats due to their extremely undetectable risk. They anticipate internet users to mistake them as genuine ones in order to reveal user information and privacy, such as login ids, pass-words, credit card numbers, etc. without notice. This paper proposes a new approach to solve the anti-phishing problem. The new features of this approach can be represented by URL character sequence without phishing prior knowledge, various hyperlink information, and textual content of the webpage, which are combined and fed to train the XGBoost classifier. One of the major contributions of this paper is the selection of different new features, which are capable enough to detect 0-h attacks, and these features do not depend on any third-party services. In particular, we extract character level Term Frequency-Inverse Document Frequency (TF-IDF) features from noisy parts of HTML and plaintext of the given webpage. Moreover, our proposed hyperlink features determine the relationship between the content and the URL of a webpage. Due to the absence of publicly available large phishing data sets, we needed to create our own data set with 60,252 webpages to validate the proposed solution. This data contains 32,972 benign webpages and 27,280 phishing webpages. For evaluations, the performance of each category of the proposed feature set is evaluated, and various classification algorithms are employed. From the empirical results, it was observed that the proposed individual features are valuable for phishing detection. However, the integration of all the features improves the detection of phishing sites with significant accuracy. The proposed approach achieved an accuracy of 96.76% with only 1.39% false-positive rate on our dataset, and an accuracy of 98.48% with 2.09% false-positive rate on benchmark dataset, which outperforms the existing baseline approaches.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Algorithms*
Benchmarking
Computer Security*
Data Collection
Privacy