Interpretable Phishing Email Detection based on CNN-Word2vec Models and LIME Explanations

Authors

  • Bakr Thamer Hammoudi Informatics Institute for Postgraduate Studies, University of Information Technology and Communications (UoITC), Baghdad, Iraq
  • Omar Z. Akif Department of Computer Science, College of Education for Pure Science (Ibn al-Haitham), University of Baghdad, Iraq

DOI:

https://doi.org/10.56714/bjrs.52.1.13

Keywords:

Cybersecurity, Phishing Email Detection, Convolutional Neural Network, Word2Vec, Word Embedding, Fasttext

Abstract

phishing emails are a major cybersecurity threat which primarily attempt to steal private information or lose money. Attackers counterfeit well-known organizations and devise sophisticated tricks to remain unnoticed, so it remains difficult for detection systems even with ample academic research. This study focuses on assessing several methods of vectorizing text and deep learning architecture in classifying emails. The main contribution of this study is to enhance the CNN with Word2vec embedding for phishing email classification. To ensure the robustness, we built a variety of benchmark dataset by merging eight public sources and promoted the feature extraction by some specific modifications of convolutional filters. Extensive preprocessing like data augmentation to handle class imbalance was used. The proposed model was longitudinally compared against TF-IDF, fastText, Word Embedding and traditional machine learning methods. Experimental results show that our proposed model achieved the highest accuracy of 99.05%, outperforming word embedding (98.72%), TF-IDF (98.85%) and fastText (98.99%). The Convolutional Neural Network (CNN) with word2vec proved highly effective in classifying phishing emails. To enhance transparency and trust, the LIME framework was incorporated, offering interpretability via contribution scores. This clarifies the models decision-making process while reducing methodological inconsistencies

Downloads

Download data is not yet available.

References

[1] A. Ozcan, C. Catal, E. Donmez, and B. Senturk, “A hybrid DNN–LSTM model for detecting phishing URLs,” Neural Comput. Appl., vol. 35, no. 7, pp. 4957–4973, 2023, DOI: 10.1007/s00521-021-06401-z.

[2]Available:https://www.simplilearn.com/ice9/free_resources_article_thumb/phishing_working_2-What_Is_Phishing.PNG

[3] Anti-Phishing Working Group, “Phishing Activity Trends Report 2nd Quarter 2025,” Most. Accessed: 30. Aug, 2025. [online]. Available: https://docs.apwg.org/reports/apwg_trends_report_q1_2025.pdf

[4] Anti-Phishing Working Group, “Phishing Activity Trends Report 2nd Quarter 2025,” Most. Accessed: 1. Sep, 2025. [online]. Available: https://docs.apwg.org/reports/apwg_trends_report_q2_2025.pdf

[5] S. Bagui, D. Nandi, S. Bagui, and R. J. White, “Machine Learning and Deep Learning for Phishing Email Classification using One-Hot Encoding,” J. Comput. Sci., vol. 17, no. 7, pp. 610–623, 2021, DOI: 10.3844/jcssp.2021.610.623. DOI: https://doi.org/10.3844/jcssp.2021.610.623

[6] S. Atawneh and H. Aljehani, “Phishing Email Detection Model Using Deep Learning,” Electron., vol. 12, no. 20, 2023, DOI: 10.3390/electronics12204261. DOI: https://doi.org/10.3390/electronics12204261

[7] R. Brindha, S. Nandagopal, H. Azath, V. Sathana, G. P. Joshi, and S. W. Kim, “Intelligent Deep Learning Based Cybersecurity Phishing Email Detection and Classification,” Comput. Mater. Contin., vol. 74, no. 3, pp. 5901–5914, 2023, DOI: 10.32604/cmc.2023.030784. DOI: https://doi.org/10.32604/cmc.2023.030784

[8] P. Krishnamoorthy, M. Sathiyanarayanan, and H. P. Proença, “A novel and secured email classification and emotion detection using hybrid deep neural network,” Int. J. Cogn. Comput. Eng., vol. 5, no. December 2023, pp. 44–57, 2024, DOI: 10.1016/j.ijcce.2024.01.002. DOI: https://doi.org/10.1016/j.ijcce.2024.01.002

[9] N. Altwaijry, I. Al-Turaiki, R. Alotaibi, and F. Alakeel, “Advancing Phishing Email Detection: A Comparative Study of Deep Learning Models,” Sensors, vol. 24, no. 7, pp. 1–19, 2024, DOI: 10.3390/s24072077. DOI: https://doi.org/10.3390/s24072077

[10] A. A. Orunsolu, A. S. Sodiya, and A. T. Akinwale, “A predictive model for phishing detection,” J. King Saud Univ. - Comput. Inf. Sci., vol. 34, no. 2, pp. 232–247, 2022, DOI: 10.1016/j.jksuci.2019.12.005. DOI: https://doi.org/10.1016/j.jksuci.2019.12.005

[11] M. Somesha and A. R. Pais, “Classification of Phishing Email Using Word Embedding and Machine Learning Techniques,” J. Cyber Secur. Mobil., vol. 11, no. 3, pp. 279–320, 2022, DOI: 10.13052/jcsm2245-1439.1131. DOI: https://doi.org/10.13052/jcsm2245-1439.1131

[12] A. Ozcan, C. Catal, E. Donmez, and B. Senturk, “A hybrid DNN–LSTM model for detecting phishing URLs,” Neural Comput. Appl., vol. 35, no. 7, pp. 4957–4973, 2023, DOI: 10.1007/s00521-021-06401-z. DOI: https://doi.org/10.1007/s00521-021-06401-z

[13] A. Al-Subaiey, M. Al-Thani, N. Abdullah Alam, K. F. Antora, A. Khandakar, and S. A. Uz Zaman, “Novel interpretable and robust web-based AI platform for phishing email detection,” Comput. Electr. Eng., vol. 120, no. Ml, 2024, DOI: 10.1016/j.compeleceng.2024.109625. DOI: https://doi.org/10.1016/j.compeleceng.2024.109625

[14] A. Alhuzali, A. Alloqmani, M. Aljabri, and F. Alharbi, “In-Depth Analysis of Phishing Email Detection: Evaluating the Performance of Machine Learning and Deep Learning Models Across Multiple Datasets,” Appl. Sci., vol. 15, no. 6, pp. 1–30, 2025, DOI: 10.3390/app15063396. DOI: https://doi.org/10.3390/app15063396

[15] Ö. Ş. Akçam, A. Tekerek, and M. Tekerek, “Development of BiLSTM deep learning model to detect URL-based phishing attacks,” Comput. Electr. Eng., vol. 123, no. September 2024, 2025, DOI: 10.1016/j.compeleceng.2025.110212. DOI: https://doi.org/10.1016/j.compeleceng.2025.110212

[16] CEAS 2008 Spam Corpus, “CEAS 2008 Spam Corpus,” Conference on Email and Anti-Spam (CEAS). Accessed: Aug. 20, 2025. [Online]. Available: https://www.kaggle.com/datasets/naserabdullahalam/phishing-email-dataset?select=CEAS_08.csv

[17] J. Nazario, “Nazario Spam and Phishing Corpus,” University of California, Irvine. Accessed: Aug. 20, 2025. [Online]. Available: https://www.kaggle.com/datasets/naserabdullahalam/phishing-email-dataset?select=Nazario.csv

[18] Nigerian Fraud, “Nigerian Fraud Email Dataset (419 Scam Corpus).” Accessed: Aug. 20, 2025. [Online]. Available: https://www.kaggle.com/datasets/naserabdullahalam/phishing-email-dataset?select=Nigerian_Fraud.csv

[19] The Apache Software Foundation, “SpamAssassin Public Mail Corpus,” The Apache Software Foundation. Accessed: Aug. 20, 2025. [Online]. Available: https://www.kaggle.com/datasets/naserabdullahalam/phishing-email-dataset?select=SpamAssasin.csv

[20] T. R. Cormack, Gordon V. and Lynam, “TREC 2005 Public Spam Corpus,” National Institute of Standards and Technology (NIST). Accessed: Aug. 20, 2025. [Online]. Available: https://zenodo.org/records/8339691/files/TREC_05.csv?download=1

[21] T. R. Cormack, Gordon V. and Lynam, “TREC 2006 Public Corpus,” National Institute of Standards and Technology (NIST). Accessed: Aug. 20, 2025. [Online]. Available: https://zenodo.org/records/8339691/files/TREC_06.csv?download=1

[22] T. R. Cormack, Gordon V. and Lynam, “TREC 2007 Public Corpus,” National Institute of Standards and Technology (NIST). Accessed: Aug. 20, 2025. [Online]. Available: https://zenodo.org/records/8339691/files/TREC_07.csv?download=1

[23] OpenDNS, “PhishTank: An anti-phishing site,” Cisco Systems, Inc. Accessed: Sep. 18, 2025. [Online]. Available: https://phishtank.org/developer_info.php%0A

[24] J. Koushik, “Understanding Convolutional Neural Networks,” no. 3, pp. 1–6, 2016.

[25] P. Bojanowski, E. Grave, A. Joulin, and T. Mikolov, “Enriching Word Vectors with Subword Information,” 1996.

[26] R. Eckhardt, “Convolutional Neural Networks and Long Short Term Memory for Phishing Email Classification,” vol. 19, no. 5, pp. 27–35, 2021.

[27] M. Korkmaz, E. Kocyigit, O. K. Sahingoz, and B. Diri, “A Hybrid Phishing Detection System Using Deep Learning-based URL and Content Analysis,” 2022. DOI: https://doi.org/10.5755/j02.eie.31197

Downloads

Published

30-06-2026

Issue

Section

Articles

How to Cite

Interpretable Phishing Email Detection based on CNN-Word2vec Models and LIME Explanations. (2026). Journal of Basrah Researches (Sciences), 52(1), 170-190. https://doi.org/10.56714/bjrs.52.1.13