An Author’s Framework for Applying Machine Learning in Multi-Layer Enterprise Cybersecurity Systems Pre-emptive Architecture, Combined Labeling, and Ensemble Integration
Abstract
This book sets out a framework for applying machine learning in corporate cybersecurity that the author developed across fifteen years of building and operating defensive systems. The framework did not begin as a theory. It began as a set of recurring engineering decisions, made and remade across several security products, that turned out to matter more than the choice of any individual model: when a model is allowed to act, what evidence it is taught from, and where it sits among the other controls in a defense. Naming those decisions as principles, and showing that their combination can be verified, is the purpose of the book.
The intended reader is a practitioner. The book is addressed to the engineers, architects, and security leaders who design machine-learning systems for enterprise defense and who must make these decisions whether or not they are made deliberately. For that reader the book aims to be useful rather than exhaustive: it states each principle, gives the architectural requirements it imposes, and grounds it in a system that was built and measured. The quantitative core of the book is an email security platform whose results were established in a peer-reviewed study, and the book is careful throughout to separate what has been measured on that system from what is projected for the adjacent domains where the same principles apply.
The book treats the published literature as a working tool rather than as a survey to be completed. Sources are cited where they establish a mechanism, corroborate a claim, or mark a limit, and the reader who wishes to go deeper will find the trail. The aim is a book that reads as the distilled judgment of practice, supported by the literature, rather than as a compilation of it.
Downloads
References
- Aboaoja, F. A., Zainal, A. B., Ali, A. M., Ghaleb, F. A., Alsolami, F. J., & Rassam, M. A. (2023). Dynamic extraction of initial behavior for evasive malware detection. Mathematics, 11(2), 416. https://doi.org/10.3390/math11020416
- Adnan, M., Imam, M. O., Javed, M. F., & Murtza, I. (2024). Improving spam email classification accuracy using ensemble techniques: A stacking approach. International Journal of Information Security, 23, 505–517. https://doi.org/10.1007/s10207-023-00756-1
- Afianian, A., Niksefat, S., Sadeghiyan, B., & Baptiste, D. (2019). Malware dynamic analysis evasion techniques: A survey. ACM Computing Surveys, 52(6), 126. https://doi.org/10.1145/3365001
- Ahmad, Z., Khan, A. S., Shiang, C. W., Abdullah, J., & Ahmad, F. (2020). Network intrusion detection system: A systematic study of machine learning and deep learning approaches. Transactions on Emerging Telecommunications Technologies, 32(1). https://doi.org/10.1002/ett.4150
- Ahmed, N., Amin, R., Aldabbas, H., Koundal, D., Alouffi, B., & Shah, T. (2022). Machine learning techniques for spam detection in email and IoT platforms: Analysis and research challenges. Security and Communication Networks, 2022, 1862888. https://doi.org/10.1155/2022/1862888
- Alabdan, R. (2020). Phishing Attacks Survey: Types, Vectors, and Technical Approaches. Future Internet, 12(10), 168. https://doi.org/10.3390/fi12100168
- Aldweesh, A., Derhab, A., & Emam, A. (2019). Deep learning approaches for anomaly-based intrusion detection systems: A survey, taxonomy, and open issues. Knowledge-Based Systems, 189, 105124. https://doi.org/10.1016/j.knosys.2019.105124
- Aljofey, A., Jiang, Q., Qu, Q., Huang, M., & Niyigena, J. (2020). An Effective Phishing Detection Model Based on Character Level Convolutional Neural Network from URL. Electronics, 9(9), 1514. https://doi.org/10.3390/electronics9091514
- Almujahid, N. F., Haq, M. A., & Alshehri, M. (2024). Comparative evaluation of machine learning algorithms for phishing site detection. PeerJ Computer Science, 10, e2131. https://doi.org/10.7717/peerj-cs.2131
- Alnemari, S., & Alshammari, M. (2023). Detecting phishing domains using machine learning. Applied Sciences, 13(8), 4649. https://doi.org/10.3390/app13084649
- Altwaijry, N., Al-Turaiki, I., Alotaibi, R., & Alakeel, F. (2024). Advancing phishing email detection: A comparative study of deep learning models. Sensors, 24(7), 2077. https://doi.org/10.3390/s24072077
- Alzubaidi, L., Zhang, J., Humaidi, A. J., Al-Dujaili, A. Q., Duan, Y., Al-Shamma, O., Santamaría, J., Fadhel, M. A., Al‐Amidie, M., & Farhan, L. (2021). Review of deep learning: concepts, CNN architectures, challenges, applications, future directions. Journal Of Big Data, 8(1), 53. https://doi.org/10.1186/s40537-021-00444-8
- Anti-Phishing Working Group. (2023). Phishing activity trends report: 4th quarter 2023. APWG. https://docs.apwg.org/reports/apwg_trends_report_q4_2023.pdf
- Apruzzese, G., Laskov, P., de, E. M., Mallouli, W., Rapa, L. B., Grammatopoulos, A. V., & Franco, F. D. (2022). The Role of Machine Learning in Cybersecurity. Digital Threats Research and Practice, 4(1), 1–38. https://doi.org/10.1145/3545574
- Atlam, H. F., & Oluwatimilehin, O. (2023). Business email compromise phishing detection based on machine learning: A systematic literature review. Electronics, 12(1), 42. https://doi.org/10.3390/electronics12010042
- Bagui, S., & Li, K. (2021). Resampling imbalanced data for network intrusion detection datasets. Journal Of Big Data, 8(1). https://doi.org/10.1186/s40537-020-00390-x
- Ban, T., Takahashi, T., Ndichu, S., & Inoue, D. (2023). Breaking alert fatigue: AI-assisted SIEM framework for effective incident response. Applied Sciences, 13(11), 6610. https://doi.org/10.3390/app13116610
- Basit, A., Zafar, M., Liu, X., Javed, A. R., Jalil, Z., & Kifayat, K. (2020). A comprehensive survey of AI-enabled phishing attacks detection techniques. Telecommunication Systems, 76(1), 139–154. https://doi.org/10.1007/s11235-020-00733-2
- Bensaoud, A., Kalita, J., & Bensaoud, M. (2024). A survey of malware detection using deep learning. Machine Learning with Applications, 16, 100546. https://doi.org/10.1016/j.mlwa.2024.100546
- Breiman, L. (2001). Random Forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/a:1010933404324
- Cao, Y., & Pokhrel, S. R. (2024). Automation and orchestration of zero trust architecture: Potential solutions and challenges. Machine Intelligence Research, 21, 294–317. https://doi.org/10.1007/s11633-023-1456-2
- Cervantes, J., García‐Lamont, F., Rodríguez-Mazahua, L., & López‐Chau, A. (2020). A comprehensive survey on support vector machine classification: Applications, challenges and trends. Neurocomputing, 408, 189–215. https://doi.org/10.1016/j.neucom.2019.10.118
- Chakraborty, A., Alam, M., Dey, V., Chattopadhyay, A., & Mukhopadhyay, D. (2021). A survey on adversarial attacks and defences. CAAI Transactions on Intelligence Technology, 6(1), 25–45. https://doi.org/10.1049/cit2.12028
- Chawla, N. V., Bowyer, K. W., Hall, L., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence Research, 16, 321–357. https://doi.org/10.1613/jair.953
- Chen, T., & Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System. KDD '16: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. Pages 785 - 794. https://doi.org/10.1145/2939672.2939785
- Devlin, J., Chang, M., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 4171–4186. https://doi.org/10.18653/v1/n19-1423
- Doshi, J., Parmar, K., Sanghavi, R., & Shekokar, N. M. (2023). A comprehensive dual-layer architecture for phishing and spam email detection. Computers & Security, 133, 103378. https://doi.org/10.1016/j.cose.2023.103378
- Elsadig, M., Ibrahim, A. O., Basheer, S., Alohali, M. A., Alshunaifi, S., Alqahtani, H. M., Alharbi, N., & Nagmeldin, W. (2022). Intelligent Deep Machine Learning Cyber Phishing URL Detection Based on BERT Features Extraction. Electronics, 11(22), 3647. https://doi.org/10.3390/electronics11223647
- Engelen, J. E. v., & Hoos, H. H. (2019). A survey on semi-supervised learning. Machine Learning, 109(2), 373–440. https://doi.org/10.1007/s10994-019-05855-6
- Fang, Y., Liu, Y., Huang, C., & Liu, L. (2020). FastEmbed: Predicting vulnerability exploitation possibility based on ensemble machine learning algorithm. PLoS ONE, 15(2), e0228439. https://doi.org/10.1371/journal.pone.0228439
- Gamage, S., & Samarabandu, J. (2020). Deep learning methods in network intrusion detection: A survey and an objective comparison. Journal of Network and Computer Applications, 169, 102767. https://doi.org/10.1016/j.jnca.2020.102767
- Gambo, L., & Almulhem, A. (2025). Zero trust architecture: A systematic literature review. Journal of Network and Systems Management. https://doi.org/10.1007/s10922-025-09998-x
- Gholampour, P. M., & Verma, R. M. (2023). Adversarial robustness of phishing email detection models. In Proceedings of the 9th ACM International Workshop on Security and Privacy Analytics (IWSPA ’23) (pp. 67–76). ACM. https://doi.org/10.1145/3579987.3586567
- Goodfellow, I., Shlens, J., & Szegedy, C. (2014). Explaining and Harnessing Adversarial Examples. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1412.6572
- Gopinath, M. A., & Sethuraman, S. C. (2022). A comprehensive survey on deep learning based malware detection techniques. Computer Science Review, 47, 100529. https://doi.org/10.1016/j.cosrev.2022.100529
- Han, X., & Zhang, T. (2022). Text adversarial attacks and defenses: Issues, taxonomy, and perspectives. Security and Communication Networks, 2022, 6458488. https://doi.org/10.1155/2022/6458488
- Hao, W., Tran, V., Rideout, V., Wang, Z., Dasbach-Prisk, A., Afifi, M. H., Yang, J., Katz-Bassett, E., Ho, G., & Cidon, A. (2025). Do spammers dream of electric sheep? Characterizing the prevalence of LLM-generated malicious emails. In Proceedings of the 2025 ACM Internet Measurement Conference (IMC ’25) (pp. 192–203). ACM. https://doi.org/10.1145/3730567.3732922
- Hazell, J. (2023). Spear phishing with large language models. arXiv preprint arXiv:2305.06972. https://doi.org/10.48550/arXiv.2305.06972
- Hemalatha, J., Roseline, S. A., Geetha, S., Kadry, S., & Damaševičius, R. (2021). An Efficient DenseNet-Based Deep Learning Model for Malware Detection. Entropy, 23(3), 344. https://doi.org/10.3390/e23030344
- Hureau, O., Bayer, J., Duda, A., & Korczyński, M. (2024). Spoofed emails: An analysis of the issues hindering a larger deployment of DMARC. In Passive and Active Measurement (PAM 2024), Lecture Notes in Computer Science (Vol. 14537, pp. 232–261). Springer. https://doi.org/10.1007/978-3-031-56249-5_10
- Innab, N., Osman, A. A. F., Ataelfadiel, M. A. M., Abu-Zanona, M., Elzaghmouri, B. M., Zawaideh, F. H., & Alawneh, M. F. (2024). Phishing attacks detection using ensemble machine learning algorithms. Computers, Materials & Continua, 80(1), 1325–1345. https://doi.org/10.32604/cmc.2024.051778
- Ismail, I., Kurnia, R., Brata, Z. A., Nelistiani, G. A., & Kim, H. (2025). Toward robust security orchestration and automated response in security operations centers. Information, 16(5), 365. https://doi.org/10.3390/info16050365
- Jacobs, J., Romanosky, S., Suciu, O., Edwards, B., & Sarkar, A. (2023). Enhancing vulnerability prioritization: Data-driven exploit predictions with community-driven insights. In 2023 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW) (pp. 194–206). IEEE. https://doi.org/10.1109/eurospw59978.2023.00027
- Jamal, S., Wimmer, H., & Sarker, I. H. (2024). An improved transformer-based model for detecting phishing, spam, and ham emails: A large language model approach. Security and Privacy, 7(3), e402. https://doi.org/10.1002/spy2.402
- Jáñez-Martino, F., Alaiz-Rodríguez, R., González-Castro, V., Fidalgo, E., & Alegre, E. (2022). A review of spam email detection: Analysis of spammer strategies and the dataset shift problem. Artificial Intelligence Review, 56, 1145–1173. https://doi.org/10.1007/s10462-022-10195-4
- Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., … Zhao, S. (2021). Advances and open problems in federated learning. Foundations and Trends in Machine Learning, 14(1–2), 1–210. https://doi.org/10.1561/2200000083
- Kang, H., Liu, G., Wang, Q., Meng, L., & Liu, J. (2023). Theory and application of zero trust security: A brief survey. Entropy, 25(12), 1595. https://doi.org/10.3390/e25121595
- Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), 100804. https://doi.org/10.1016/j.patter.2023.100804
- Karataş, G., Demir, Ö., & Şahingöz, Ö. K. (2020). Increasing the Performance of Machine Learning-Based IDSs on an Imbalanced and Up-to-Date Dataset. IEEE Access, 8, 32150–32162. https://doi.org/10.1109/access.2020.2973219
- Karim, A., Azam, S., Shanmugam, B., Kannoorpatti, K., & Alazab, M. (2019). A Comprehensive Survey for Intelligent Spam Email Detection. IEEE Access, 7, 168261–168295. https://doi.org/10.1109/access.2019.2954791
- Karim, A., Shahroz, M., Mustofa, K., Belhaouari, S. B., & Joga, S. R. K. (2023). Phishing Detection System Through Hybrid Machine Learning Based on URL. IEEE Access, 11, 36805–36822. https://doi.org/10.1109/access.2023.3252366
- Khan, N., Lee, K., & Houng, A. Y. (2022). Anomaly detection and enterprise security using user and entity behavior analytics (UEBA). In 2022 3rd International Conference on Next Generation Computing Applications (NextComp). IEEE. https://doi.org/10.1109/iconics56716.2022.10100596
- Khraisat, A., Gondal, I., Vamplew, P., & Kamruzzaman, J. (2019). Survey of intrusion detection systems: techniques, datasets and challenges. Cybersecurity, 2(1). https://doi.org/10.1186/s42400-019-0038-7
- Kolchin, R. (2025). Development and implementation of the Mail Security Guardian (MSG) system for multi-layer proactive email protection against spam, phishing and malware. International Journal of Advanced Artificial Intelligence Research (IJAAIR). [Author’s published study; full volume/issue/pages/DOI to be confirmed at submission.]
- Kumar, P., Antony, K., Banga, D., & Sohal, A. S. (2024). PhishNet: A phishing website detection tool using XGBoost. arXiv preprint arXiv:2407.04732. https://doi.org/10.48550/arXiv.2407.04732
- Liu, H., & Lang, B. (2019). Machine Learning and Deep Learning Methods for Intrusion Detection Systems: A Survey. Applied Sciences, 9(20), 4396. https://doi.org/10.3390/app9204396
- Liu, S., Feng, P., Wang, S., Sun, K., & Cao, J. (2022). Enhancing malware analysis sandboxes with emulated user behavior. Computers & Security, 115, 102613. https://doi.org/10.1016/j.cose.2022.102613
- Lu, J., Liu, A., Dong, F., Gu, F., Gama, J., & Zhang, G. (2018). Learning under Concept Drift: A Review. IEEE Transactions on Knowledge and Data Engineering, 1. https://doi.org/10.1109/tkde.2018.2876857
- Lundberg, S., & Lee, S. (2017). A Unified Approach to Interpreting Model Predictions. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1705.07874
- Macas, M., Wu, C., & Fuertes, W. (2023). Adversarial examples: A survey of attacks and defenses in deep learning-enabled cybersecurity. Expert Systems with Applications, 238, 122223. https://doi.org/10.1016/j.eswa.2023.122223
- Mahdavifar, S., & Ghorbani, A. A. (2019). Application of deep learning to cybersecurity: A survey. Neurocomputing, 347, 149–176. https://doi.org/10.1016/j.neucom.2019.02.056
- McMahan, H. B., Moore, E., Ramage, D., Hampson, S., & Arcas, B. A. y. (2016). Communication-Efficient Learning of Deep Networks from DecentralizedData. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1602.05629
- Mienye, I. D., & Sun, Y. (2022). A Survey of Ensemble Learning: Concepts, Algorithms, Applications, and Prospects. IEEE Access, 10, 99129–99149. https://doi.org/10.1109/access.2022.3207287
- Otieno, D. O., Namin, A. S., & Jones, K. S. (2023). The application of the BERT transformer model for phishing email classification. In Proceedings of the 2023 IEEE 47th Annual Computers, Software, and Applications Conference (COMPSAC) (pp. 1303–1308). IEEE. https://doi.org/10.1109/COMPSAC57700.2023.00198
- Qiu, S., Liu, Q., Zhou, S., & Wu, C. (2019). Review of Artificial Intelligence Adversarial Attack and Defense Technologies. Applied Sciences, 9(5), 909. https://doi.org/10.3390/app9050909
- Ragheb, M., Elmedany, W., & Sharif, M. (2023). The effectiveness of DKIM and SPF in strengthening email security. In Proceedings of the 10th International Conference on Future Internet of Things and Cloud (FiCloud 2023) (pp. 422–426). IEEE. https://doi.org/10.1109/FiCloud58648.2023.00068
- Rose, S., Borchert, O., Mitchell, S., & Connelly, S. (2020). Zero Trust Architecture. National Institute of Standards and Technology Special Publication 800-207. https://doi.org/10.6028/nist.sp.800-207
- Roy, S. S., Thota, P., Naragam, K. V., & Nilizadeh, S. (2024). From chatbots to phishbots? Phishing scam generation in commercial large language models. In Proceedings of the 2024 IEEE Symposium on Security and Privacy (SP) (pp. 36–54). IEEE. https://doi.org/10.1109/SP54263.2024.00182
- Salloum, S. A., Gaber, T., Vadera, S., & Shaalan, K. (2022). A systematic literature review on phishing email detection using natural language processing techniques. IEEE Access, 10, 65703–65727. https://doi.org/10.1109/ACCESS.2022.3183083
- Sarker, I. H., Kayes, A. S. M., Badsha, S., Alqahtani, H., Watters, P., & Ng, A. (2020). Cybersecurity data science: an overview from machine learning perspective. Journal Of Big Data, 7(1). https://doi.org/10.1186/s40537-020-00318-5
- Saxena, N., Hayes, E., Bertino, E., Ojo, P., Choo, K. R., & Burnap, P. (2020). Impact and Key Challenges of Insider Threats on Organizations and Critical Businesses. Electronics, 9(9), 1460. https://doi.org/10.3390/electronics9091460
- Shaukat, K., Luo, S., Varadharajan, V., Hameed, I. A., & Xu, M. (2020). A Survey on Machine Learning Techniques for Cyber Security in the Last Decade. IEEE Access, 8, 222310–222354. https://doi.org/10.1109/access.2020.3041951
- Shorten, C., & Khoshgoftaar, T. M. (2019). A survey on Image Data Augmentation for Deep Learning. Journal Of Big Data, 6(1). https://doi.org/10.1186/s40537-019-0197-0
- Singh, J., & Singh, J. (2020). A survey on machine learning-based malware detection in executable files. Journal of Systems Architecture, 112, 101861. https://doi.org/10.1016/j.sysarc.2020.101861
- Sun, N., Ding, M., Jiang, J., Xu, W., Mo, X., Tai, Y., & Zhang, J. (2023). Cyber Threat Intelligence Mining for Proactive Cybersecurity Defense: A Survey and New Perspectives. IEEE Communications Surveys & Tutorials, 25(3), 1748–1774. https://doi.org/10.1109/comst.2023.3273282
- Syed, N. F., Shah, S. W., Shaghaghi, A., Anwar, A., Baig, Z. A., & Doss, R. R. M. (2022). Zero trust architecture (ZTA): A comprehensive survey. IEEE Access, 10, 57143–57179. https://doi.org/10.1109/ACCESS.2022.3174679
- Vielberth, M., Böhm, F., Fichtinger, I., & Pernul, G. (2020). Security operations center: A systematic study and open challenges. IEEE Access, 8, 227756–227779. https://doi.org/10.1109/ACCESS.2020.3045514
- Vinayakumar, R., Alazab, M., Soman, K. P., Poornachandran, P., Al-Nemrat, A., & Venkatraman, S. (2019). Deep Learning Approach for Intelligent Intrusion Detection System. IEEE Access, 7, 41525–41550. https://doi.org/10.1109/access.2019.2895334
- Wen, J., Zhang, Z., Lan, Y., Cui, Z., Cai, J., & Zhang, W. (2022). A survey on federated learning: Challenges and applications. International Journal of Machine Learning and Cybernetics, 14, 513–535. https://doi.org/10.1007/s13042-022-01647-y
- Yang, P., Zhao, G., & Zeng, P. (2019). Phishing Website Detection Based on Multidimensional Features Driven by Deep Learning. IEEE Access, 7, 15196–15209. https://doi.org/10.1109/access.2019.2892066
- Zhang, Z., Hamadi, H. A., Damiani, E., Yeun, C. Y., & Taher, F. (2022). Explainable Artificial Intelligence Applications in Cyber Security: State-of-the-Art in Research. IEEE Access, 10, 93104–93139. https://doi.org/10.1109/access.2022.3204051
- Zieni, R., Massari, L., & Calzarossa, M. C. (2023). Phishing or Not Phishing? A Survey on the Detection of Phishing Websites. IEEE Access, 11, 18499–18519. https://doi.org/10.1109/access.2023.3247135