Detecting Phishing Websites from Character-Level URL Sequences: A Bidirectional LSTM Approach with Class-Imbalance Mitigation

Authors

  • Ilobekemen Perpetual Oladoja * The Federal University of Technology, Akure, Nigeria. https://orcid.org/0000-0002-1717-4985
  • Chukwuemeka C. Ugwu The Federal University of Technology, Akure, Nigeria.
  • Anuoluwapo P. Ajibade The Federal University of Technology, Akure, Nigeria.
  • Tolulope A. Ugwu The Federal University of Technology, Akure, Nigeria.
  • Olugbenga Ayomide Madamidola The Federal University of Technology, Akure, Nigeria. https://orcid.org/0000-0001-5991-4067

https://doi.org/10.48314/isti.vi.55

Abstract

Phishing websites remain a significant cybersecurity threat, exploiting deceptive Uniform Resource Locator (URLs) to trick users into revealing sensitive information. Existing detection approaches, particularly blacklist-based and traditional Machine Learning (ML) methods, struggle to generalize to newly crafted Phishing URLs and often fail to capture the sequential patterns inherent in URL structures. To address these limitations, this study proposes a Deep Learning (DL)–based Phishing website detection model using a Bidirectional Long Short-Term Memory (Bi-LSTM) network that operates on character-level URL representations. A large-scale dataset comprising 450,176 URLs (Phishing and legitimate) sourced from the Mendeley repository was employed. The URLs were preprocessed through normalization, character-level tokenization, and embedding, while class imbalance was mitigated using Synthetic Minority Over-Sampling Technique (SMOTE). The proposed Bi-LSTM model was trained and evaluated using standard performance metrics, including accuracy, precision, recall, F1-score, and Receiver Operating Characteristic (ROC)–Area Under the Curve (AUC), and was compared against a baseline Convolutional Neural Network (CNN). Experimental results demonstrate that the Bi-LSTM model achieves an accuracy of 97.81%, precision of 94.75%, recall of 95.97%, and an F1-score of 95.36%, outperforming the CNN baseline in terms of accuracy and precision while maintaining comparable recall.

Keywords:

Deep learning, Phishing uniform resource locators, Legitimate uniform resource locators, Detection model, Machine learning, Bidirectional long short-term memory, Convolutional neural network

References

  1. [1] Al-Ahmadi, S., & Alharbi, Y. (2020). A deep learning technique for web Phishing detection combined URL features and visual similarity. International journal of computer networks and communications, 12(5), 41–54. https://doi.org/10.5121/ijcnc.2020.12503

  2. [2] Devan, K. P. K., Nagaraj, R., Kamaleshwaran, P., & Karthick, S. (2023). Detection of Phishing websites using machine learning. Proceedings of the 2022 international conference on intelligent communication, control and devices (ICCCD) (pp. 1–4). IEEE. https://doi.org/10.2139/ssrn.4366981

  3. [3] Raja, A. S., Pradeepa, G., Mahalakshmi, S., & Jayakumar, M. S. (2023). Natural language based malicious domain detection using machine learning and deep learning. Scientific and technical journal of information technologies, mechanics and optics, 32(2), 304–312. https://doi.org/10.17586/2226-1494-2023-23-2-304-312

  4. [4] Alsaedi, M., Ghaleb, F. A., Saeed, F., Ahmad, J., & Alasli, M. (2022). Cyber threat intelligence-based malicious URL detection model using ensemble learning. Sensors, 22(9), 3373. https://doi.org/10.3390/s22093373

  5. [5] Kadiyala, P., KV, S. S., Shashank, B. S., & Kumar, K. A. (2021). Phishing website detection using emerging machine learning models. Proceedings of the international conference on smart data intelligence (ICSMDI) (pp. 1–5). IEEE. https://doi.org/10.2139/ssrn.3853016

  6. [6] Sahingoz, O. K., Buber, E., Demir, O., & Diri, B. (2019). Machine learning based Phishing detection from URLs. Expert systems with applications, 117, 345-357. https://doi.org/10.1016/j.eswa.2018.09.029

  7. [7] Rayalla, P. S., Vivek, S. V., Golla, H., Katakam, G., & AN, M. Z. (2023). Detecting Phishing websites using deep learning. 1st-international conference on recent innovations in computing, science & technology (pp. 1–4). IEEE. https://doi.org/10.2139/ssrn.4527759

  8. [8] Barik, K., Misra, S., & Mohan, R. (2025). Web-based Phishing URL detection model using deep learning optimization techniques. International journal of data science and analytics, 20(5), 4449–4471. https://doi.org/10.1007/s41060-025-00728-9

  9. [9] Alsabri, A. A., & Al-Hadi, M. A. (2025). A hybrid CNN-BLSTM model for Phishing attack detection using deep learning to strengthen internet security. Sana’a university journal of applied sciences and technology, 3(4), 964–972. https://doi.org/10.59628/jast.v3i4.1822

  10. [10] Aljofey, A., Jiang, Q., Rasool, A., Chen, H., Liu, W., Qu, Q., & Wang, Y. (2022). An effective detection approach for Phishing websites using URL and HTML features. Scientific reports, 12(1), 8842. https://doi.org/10.1038/s41598-022-10841-5

  11. [11] Murhej, M., & Nallasivan, G. (2025). Component features based enhanced Phishing website detection system using EfficientNet, FH-BERT, and SELU-CRNN methods. Frontiers in computer science, 7, 1582206. https://doi.org/10.3389/fcomp.2025.1582206

  12. [12] Fang, C., Yin, X., Teng, S., Zhao, C., & Huang, D. (2025). Contrastive multimodal network for Phishing detection via enhanced website authenticity analysis. International journal of digital crime and forensics, 17(1), 1–23. https://dx.doi.org/10.2139/ssrn.5036126

  13. [13] Xie, L., Zhang, H., Yang, H., Hu, Z., & Cheng, X. (2025). A scalable Phishing website detection model based on dual-branch TCN and mask attention. Computer networks, 263, 111230. https://doi.org/10.1016/j.comnet.2025.111230

  14. [14] Abid, I., & Abid, M. (2025). AI-based Phishing detection for emails and URLs. https://dx.doi.org/10.2139/ssrn.5336075

  15. [15] Kaitholikkal, J. K. S., & Arthi, B. (2024). Phishing URL dataset. Mendeley data. https://doi.org/10.17632/vfszbj9b36.1

Published

2026-12-09

How to Cite

Oladoja, I. P. ., Ugwu, C. C. ., Ajibade, A. P. ., Ugwu, T. A. ., & Madamidola, O. A. . (2026). Detecting Phishing Websites from Character-Level URL Sequences: A Bidirectional LSTM Approach with Class-Imbalance Mitigation. Information Sciences and Technological Innovations, 2(4), 288-302. https://doi.org/10.48314/isti.vi.55

Similar Articles

1-10 of 18

You may also start an advanced similarity search for this article.