Interpretable Phishing Email Detection based on CNN-Word2vec Models and LIME Explanations
DOI:
https://doi.org/10.56714/bjrs.52.1.13Keywords:
Cybersecurity, Phishing Email Detection, Convolutional Neural Network, Word2Vec, Word Embedding, FasttextAbstract
phishing emails are a major cybersecurity threat which primarily attempt to steal private information or lose money. Attackers counterfeit well-known organizations and devise sophisticated tricks to remain unnoticed, so it remains difficult for detection systems even with ample academic research. This study focuses on assessing several methods of vectorizing text and deep learning architecture in classifying emails. The main contribution of this study is to enhance the CNN with Word2vec embedding for phishing email classification. To ensure the robustness, we built a variety of benchmark dataset by merging eight public sources and promoted the feature extraction by some specific modifications of convolutional filters. Extensive preprocessing like data augmentation to handle class imbalance was used. The proposed model was longitudinally compared against TF-IDF, fastText, Word Embedding and traditional machine learning methods. Experimental results show that our proposed model achieved the highest accuracy of 99.05%, outperforming word embedding (98.72%), TF-IDF (98.85%) and fastText (98.99%). The Convolutional Neural Network (CNN) with word2vec proved highly effective in classifying phishing emails. To enhance transparency and trust, the LIME framework was incorporated, offering interpretability via contribution scores. This clarifies the models decision-making process while reducing methodological inconsistencies
Downloads
References
[1] A. Ozcan, C. Catal, E. Donmez, and B. Senturk, “A hybrid DNN–LSTM model for detecting phishing URLs,” Neural Comput. Appl., vol. 35, no. 7, pp. 4957–4973, 2023, DOI: 10.1007/s00521-021-06401-z.
[2]Available:https://www.simplilearn.com/ice9/free_resources_article_thumb/phishing_working_2-What_Is_Phishing.PNG
[3] Anti-Phishing Working Group, “Phishing Activity Trends Report 2nd Quarter 2025,” Most. Accessed: 30. Aug, 2025. [online]. Available: https://docs.apwg.org/reports/apwg_trends_report_q1_2025.pdf
[4] Anti-Phishing Working Group, “Phishing Activity Trends Report 2nd Quarter 2025,” Most. Accessed: 1. Sep, 2025. [online]. Available: https://docs.apwg.org/reports/apwg_trends_report_q2_2025.pdf
[5] S. Bagui, D. Nandi, S. Bagui, and R. J. White, “Machine Learning and Deep Learning for Phishing Email Classification using One-Hot Encoding,” J. Comput. Sci., vol. 17, no. 7, pp. 610–623, 2021, DOI: 10.3844/jcssp.2021.610.623. DOI: https://doi.org/10.3844/jcssp.2021.610.623
[6] S. Atawneh and H. Aljehani, “Phishing Email Detection Model Using Deep Learning,” Electron., vol. 12, no. 20, 2023, DOI: 10.3390/electronics12204261. DOI: https://doi.org/10.3390/electronics12204261
[7] R. Brindha, S. Nandagopal, H. Azath, V. Sathana, G. P. Joshi, and S. W. Kim, “Intelligent Deep Learning Based Cybersecurity Phishing Email Detection and Classification,” Comput. Mater. Contin., vol. 74, no. 3, pp. 5901–5914, 2023, DOI: 10.32604/cmc.2023.030784. DOI: https://doi.org/10.32604/cmc.2023.030784
[8] P. Krishnamoorthy, M. Sathiyanarayanan, and H. P. Proença, “A novel and secured email classification and emotion detection using hybrid deep neural network,” Int. J. Cogn. Comput. Eng., vol. 5, no. December 2023, pp. 44–57, 2024, DOI: 10.1016/j.ijcce.2024.01.002. DOI: https://doi.org/10.1016/j.ijcce.2024.01.002
[9] N. Altwaijry, I. Al-Turaiki, R. Alotaibi, and F. Alakeel, “Advancing Phishing Email Detection: A Comparative Study of Deep Learning Models,” Sensors, vol. 24, no. 7, pp. 1–19, 2024, DOI: 10.3390/s24072077. DOI: https://doi.org/10.3390/s24072077
[10] A. A. Orunsolu, A. S. Sodiya, and A. T. Akinwale, “A predictive model for phishing detection,” J. King Saud Univ. - Comput. Inf. Sci., vol. 34, no. 2, pp. 232–247, 2022, DOI: 10.1016/j.jksuci.2019.12.005. DOI: https://doi.org/10.1016/j.jksuci.2019.12.005
[11] M. Somesha and A. R. Pais, “Classification of Phishing Email Using Word Embedding and Machine Learning Techniques,” J. Cyber Secur. Mobil., vol. 11, no. 3, pp. 279–320, 2022, DOI: 10.13052/jcsm2245-1439.1131. DOI: https://doi.org/10.13052/jcsm2245-1439.1131
[12] A. Ozcan, C. Catal, E. Donmez, and B. Senturk, “A hybrid DNN–LSTM model for detecting phishing URLs,” Neural Comput. Appl., vol. 35, no. 7, pp. 4957–4973, 2023, DOI: 10.1007/s00521-021-06401-z. DOI: https://doi.org/10.1007/s00521-021-06401-z
[13] A. Al-Subaiey, M. Al-Thani, N. Abdullah Alam, K. F. Antora, A. Khandakar, and S. A. Uz Zaman, “Novel interpretable and robust web-based AI platform for phishing email detection,” Comput. Electr. Eng., vol. 120, no. Ml, 2024, DOI: 10.1016/j.compeleceng.2024.109625. DOI: https://doi.org/10.1016/j.compeleceng.2024.109625
[14] A. Alhuzali, A. Alloqmani, M. Aljabri, and F. Alharbi, “In-Depth Analysis of Phishing Email Detection: Evaluating the Performance of Machine Learning and Deep Learning Models Across Multiple Datasets,” Appl. Sci., vol. 15, no. 6, pp. 1–30, 2025, DOI: 10.3390/app15063396. DOI: https://doi.org/10.3390/app15063396
[15] Ö. Ş. Akçam, A. Tekerek, and M. Tekerek, “Development of BiLSTM deep learning model to detect URL-based phishing attacks,” Comput. Electr. Eng., vol. 123, no. September 2024, 2025, DOI: 10.1016/j.compeleceng.2025.110212. DOI: https://doi.org/10.1016/j.compeleceng.2025.110212
[16] CEAS 2008 Spam Corpus, “CEAS 2008 Spam Corpus,” Conference on Email and Anti-Spam (CEAS). Accessed: Aug. 20, 2025. [Online]. Available: https://www.kaggle.com/datasets/naserabdullahalam/phishing-email-dataset?select=CEAS_08.csv
[17] J. Nazario, “Nazario Spam and Phishing Corpus,” University of California, Irvine. Accessed: Aug. 20, 2025. [Online]. Available: https://www.kaggle.com/datasets/naserabdullahalam/phishing-email-dataset?select=Nazario.csv
[18] Nigerian Fraud, “Nigerian Fraud Email Dataset (419 Scam Corpus).” Accessed: Aug. 20, 2025. [Online]. Available: https://www.kaggle.com/datasets/naserabdullahalam/phishing-email-dataset?select=Nigerian_Fraud.csv
[19] The Apache Software Foundation, “SpamAssassin Public Mail Corpus,” The Apache Software Foundation. Accessed: Aug. 20, 2025. [Online]. Available: https://www.kaggle.com/datasets/naserabdullahalam/phishing-email-dataset?select=SpamAssasin.csv
[20] T. R. Cormack, Gordon V. and Lynam, “TREC 2005 Public Spam Corpus,” National Institute of Standards and Technology (NIST). Accessed: Aug. 20, 2025. [Online]. Available: https://zenodo.org/records/8339691/files/TREC_05.csv?download=1
[21] T. R. Cormack, Gordon V. and Lynam, “TREC 2006 Public Corpus,” National Institute of Standards and Technology (NIST). Accessed: Aug. 20, 2025. [Online]. Available: https://zenodo.org/records/8339691/files/TREC_06.csv?download=1
[22] T. R. Cormack, Gordon V. and Lynam, “TREC 2007 Public Corpus,” National Institute of Standards and Technology (NIST). Accessed: Aug. 20, 2025. [Online]. Available: https://zenodo.org/records/8339691/files/TREC_07.csv?download=1
[23] OpenDNS, “PhishTank: An anti-phishing site,” Cisco Systems, Inc. Accessed: Sep. 18, 2025. [Online]. Available: https://phishtank.org/developer_info.php%0A
[24] J. Koushik, “Understanding Convolutional Neural Networks,” no. 3, pp. 1–6, 2016.
[25] P. Bojanowski, E. Grave, A. Joulin, and T. Mikolov, “Enriching Word Vectors with Subword Information,” 1996.
[26] R. Eckhardt, “Convolutional Neural Networks and Long Short Term Memory for Phishing Email Classification,” vol. 19, no. 5, pp. 27–35, 2021.
[27] M. Korkmaz, E. Kocyigit, O. K. Sahingoz, and B. Diri, “A Hybrid Phishing Detection System Using Deep Learning-based URL and Content Analysis,” 2022. DOI: https://doi.org/10.5755/j02.eie.31197
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Basrah Researches (Sciences)

This work is licensed under a Creative Commons Attribution 4.0 International License.
This journal is licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
Under this license, users are permitted to read, download, copy, distribute, print, search, link to the full texts of articles, and create derivative works, including for commercial purposes, provided that appropriate credit is given to the original author(s) and the source.
Authors retain the copyright of their published work, while granting the Journal of Basrah Researches Sciences (JBRS) the right of first publication. Proper attribution must include the article title, author name(s), journal name, DOI, and a link to the Creative Commons license.





