★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★Open research. Global perspectives.
Engineering and Technology Book OPEN ACCESS

An Author’s Framework for Applying Machine Learning in Multi-Layer Enterprise Cybersecurity Systems Pre-emptive Architecture, Combined Labeling, and Ensemble Integration

Rustam Kolchin
Head of Cybersecurity Software Product Development SoftLine PJSC, Almaty, Kazakhstan
tajet 2026
JOURNAL PUBLICATION JANUARY
VOLUME —
ISSUE —
YEAR 2026
PAGES 01-59

Abstract

This book sets out a framework for applying machine learning in corporate cybersecurity that the author developed across fifteen years of building and operating defensive systems. The framework did not begin as a theory. It began as a set of recurring engineering decisions, made and remade across several security products, that turned out to matter more than the choice of any individual model: when a model is allowed to act, what evidence it is taught from, and where it sits among the other controls in a defense. Naming those decisions as principles, and showing that their combination can be verified, is the purpose of the book.

The intended reader is a practitioner. The book is addressed to the engineers, architects, and security leaders who design machine-learning systems for enterprise defense and who must make these decisions whether or not they are made deliberately. For that reader the book aims to be useful rather than exhaustive: it states each principle, gives the architectural requirements it imposes, and grounds it in a system that was built and measured. The quantitative core of the book is an email security platform whose results were established in a peer-reviewed study, and the book is careful throughout to separate what has been measured on that system from what is projected for the adjacent domains where the same principles apply.

The book treats the published literature as a working tool rather than as a survey to be completed. Sources are cited where they establish a mechanism, corroborate a claim, or mark a limit, and the reader who wishes to go deeper will find the trail. The aim is a book that reads as the distilled judgment of practice, supported by the literature, rather than as a compilation of it.

Downloads

Download data is not yet available.

References

  1. Aboaoja, F. A., Zainal, A. B., Ali, A. M., Ghaleb, F. A., Alsolami, F. J., & Rassam, M. A. (2023). Dynamic extraction of initial behavior for evasive malware detection. Mathematics, 11(2), 416. https://doi.org/10.3390/math11020416
  2. Adnan, M., Imam, M. O., Javed, M. F., & Murtza, I. (2024). Improving spam email classification accuracy using ensemble techniques: A stacking approach. International Journal of Information Security, 23, 505–517. https://doi.org/10.1007/s10207-023-00756-1
  3. Afianian, A., Niksefat, S., Sadeghiyan, B., & Baptiste, D. (2019). Malware dynamic analysis evasion techniques: A survey. ACM Computing Surveys, 52(6), 126. https://doi.org/10.1145/3365001
  4. Ahmad, Z., Khan, A. S., Shiang, C. W., Abdullah, J., & Ahmad, F. (2020). Network intrusion detection system: A systematic study of machine learning and deep learning approaches. Transactions on Emerging Telecommunications Technologies, 32(1). https://doi.org/10.1002/ett.4150
  5. Ahmed, N., Amin, R., Aldabbas, H., Koundal, D., Alouffi, B., & Shah, T. (2022). Machine learning techniques for spam detection in email and IoT platforms: Analysis and research challenges. Security and Communication Networks, 2022, 1862888. https://doi.org/10.1155/2022/1862888
  6. Alabdan, R. (2020). Phishing Attacks Survey: Types, Vectors, and Technical Approaches. Future Internet, 12(10), 168. https://doi.org/10.3390/fi12100168
  7. Aldweesh, A., Derhab, A., & Emam, A. (2019). Deep learning approaches for anomaly-based intrusion detection systems: A survey, taxonomy, and open issues. Knowledge-Based Systems, 189, 105124. https://doi.org/10.1016/j.knosys.2019.105124
  8. Aljofey, A., Jiang, Q., Qu, Q., Huang, M., & Niyigena, J. (2020). An Effective Phishing Detection Model Based on Character Level Convolutional Neural Network from URL. Electronics, 9(9), 1514. https://doi.org/10.3390/electronics9091514
  9. Almujahid, N. F., Haq, M. A., & Alshehri, M. (2024). Comparative evaluation of machine learning algorithms for phishing site detection. PeerJ Computer Science, 10, e2131. https://doi.org/10.7717/peerj-cs.2131
  10. Alnemari, S., & Alshammari, M. (2023). Detecting phishing domains using machine learning. Applied Sciences, 13(8), 4649. https://doi.org/10.3390/app13084649
  11. Altwaijry, N., Al-Turaiki, I., Alotaibi, R., & Alakeel, F. (2024). Advancing phishing email detection: A comparative study of deep learning models. Sensors, 24(7), 2077. https://doi.org/10.3390/s24072077
  12. Alzubaidi, L., Zhang, J., Humaidi, A. J., Al-Dujaili, A. Q., Duan, Y., Al-Shamma, O., Santamaría, J., Fadhel, M. A., Al‐Amidie, M., & Farhan, L. (2021). Review of deep learning: concepts, CNN architectures, challenges, applications, future directions. Journal Of Big Data, 8(1), 53. https://doi.org/10.1186/s40537-021-00444-8
  13. Anti-Phishing Working Group. (2023). Phishing activity trends report: 4th quarter 2023. APWG. https://docs.apwg.org/reports/apwg_trends_report_q4_2023.pdf
  14. Apruzzese, G., Laskov, P., de, E. M., Mallouli, W., Rapa, L. B., Grammatopoulos, A. V., & Franco, F. D. (2022). The Role of Machine Learning in Cybersecurity. Digital Threats Research and Practice, 4(1), 1–38. https://doi.org/10.1145/3545574
  15. Atlam, H. F., & Oluwatimilehin, O. (2023). Business email compromise phishing detection based on machine learning: A systematic literature review. Electronics, 12(1), 42. https://doi.org/10.3390/electronics12010042
  16. Bagui, S., & Li, K. (2021). Resampling imbalanced data for network intrusion detection datasets. Journal Of Big Data, 8(1). https://doi.org/10.1186/s40537-020-00390-x
  17. Ban, T., Takahashi, T., Ndichu, S., & Inoue, D. (2023). Breaking alert fatigue: AI-assisted SIEM framework for effective incident response. Applied Sciences, 13(11), 6610. https://doi.org/10.3390/app13116610
  18. Basit, A., Zafar, M., Liu, X., Javed, A. R., Jalil, Z., & Kifayat, K. (2020). A comprehensive survey of AI-enabled phishing attacks detection techniques. Telecommunication Systems, 76(1), 139–154. https://doi.org/10.1007/s11235-020-00733-2
  19. Bensaoud, A., Kalita, J., & Bensaoud, M. (2024). A survey of malware detection using deep learning. Machine Learning with Applications, 16, 100546. https://doi.org/10.1016/j.mlwa.2024.100546
  20. Breiman, L. (2001). Random Forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/a:1010933404324
  21. Cao, Y., & Pokhrel, S. R. (2024). Automation and orchestration of zero trust architecture: Potential solutions and challenges. Machine Intelligence Research, 21, 294–317. https://doi.org/10.1007/s11633-023-1456-2
  22. Cervantes, J., García‐Lamont, F., Rodríguez-Mazahua, L., & López‐Chau, A. (2020). A comprehensive survey on support vector machine classification: Applications, challenges and trends. Neurocomputing, 408, 189–215. https://doi.org/10.1016/j.neucom.2019.10.118
  23. Chakraborty, A., Alam, M., Dey, V., Chattopadhyay, A., & Mukhopadhyay, D. (2021). A survey on adversarial attacks and defences. CAAI Transactions on Intelligence Technology, 6(1), 25–45. https://doi.org/10.1049/cit2.12028
  24. Chawla, N. V., Bowyer, K. W., Hall, L., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence Research, 16, 321–357. https://doi.org/10.1613/jair.953
  25. Chen, T., & Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System. KDD '16: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. Pages 785 - 794. https://doi.org/10.1145/2939672.2939785
  26. Devlin, J., Chang, M., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 4171–4186. https://doi.org/10.18653/v1/n19-1423
  27. Doshi, J., Parmar, K., Sanghavi, R., & Shekokar, N. M. (2023). A comprehensive dual-layer architecture for phishing and spam email detection. Computers & Security, 133, 103378. https://doi.org/10.1016/j.cose.2023.103378
  28. Elsadig, M., Ibrahim, A. O., Basheer, S., Alohali, M. A., Alshunaifi, S., Alqahtani, H. M., Alharbi, N., & Nagmeldin, W. (2022). Intelligent Deep Machine Learning Cyber Phishing URL Detection Based on BERT Features Extraction. Electronics, 11(22), 3647. https://doi.org/10.3390/electronics11223647
  29. Engelen, J. E. v., & Hoos, H. H. (2019). A survey on semi-supervised learning. Machine Learning, 109(2), 373–440. https://doi.org/10.1007/s10994-019-05855-6
  30. Fang, Y., Liu, Y., Huang, C., & Liu, L. (2020). FastEmbed: Predicting vulnerability exploitation possibility based on ensemble machine learning algorithm. PLoS ONE, 15(2), e0228439. https://doi.org/10.1371/journal.pone.0228439
  31. Gamage, S., & Samarabandu, J. (2020). Deep learning methods in network intrusion detection: A survey and an objective comparison. Journal of Network and Computer Applications, 169, 102767. https://doi.org/10.1016/j.jnca.2020.102767
  32. Gambo, L., & Almulhem, A. (2025). Zero trust architecture: A systematic literature review. Journal of Network and Systems Management. https://doi.org/10.1007/s10922-025-09998-x
  33. Gholampour, P. M., & Verma, R. M. (2023). Adversarial robustness of phishing email detection models. In Proceedings of the 9th ACM International Workshop on Security and Privacy Analytics (IWSPA ’23) (pp. 67–76). ACM. https://doi.org/10.1145/3579987.3586567
  34. Goodfellow, I., Shlens, J., & Szegedy, C. (2014). Explaining and Harnessing Adversarial Examples. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1412.6572
  35. Gopinath, M. A., & Sethuraman, S. C. (2022). A comprehensive survey on deep learning based malware detection techniques. Computer Science Review, 47, 100529. https://doi.org/10.1016/j.cosrev.2022.100529
  36. Han, X., & Zhang, T. (2022). Text adversarial attacks and defenses: Issues, taxonomy, and perspectives. Security and Communication Networks, 2022, 6458488. https://doi.org/10.1155/2022/6458488
  37. Hao, W., Tran, V., Rideout, V., Wang, Z., Dasbach-Prisk, A., Afifi, M. H., Yang, J., Katz-Bassett, E., Ho, G., & Cidon, A. (2025). Do spammers dream of electric sheep? Characterizing the prevalence of LLM-generated malicious emails. In Proceedings of the 2025 ACM Internet Measurement Conference (IMC ’25) (pp. 192–203). ACM. https://doi.org/10.1145/3730567.3732922
  38. Hazell, J. (2023). Spear phishing with large language models. arXiv preprint arXiv:2305.06972. https://doi.org/10.48550/arXiv.2305.06972
  39. Hemalatha, J., Roseline, S. A., Geetha, S., Kadry, S., & Damaševičius, R. (2021). An Efficient DenseNet-Based Deep Learning Model for Malware Detection. Entropy, 23(3), 344. https://doi.org/10.3390/e23030344
  40. Hureau, O., Bayer, J., Duda, A., & Korczyński, M. (2024). Spoofed emails: An analysis of the issues hindering a larger deployment of DMARC. In Passive and Active Measurement (PAM 2024), Lecture Notes in Computer Science (Vol. 14537, pp. 232–261). Springer. https://doi.org/10.1007/978-3-031-56249-5_10
  41. Innab, N., Osman, A. A. F., Ataelfadiel, M. A. M., Abu-Zanona, M., Elzaghmouri, B. M., Zawaideh, F. H., & Alawneh, M. F. (2024). Phishing attacks detection using ensemble machine learning algorithms. Computers, Materials & Continua, 80(1), 1325–1345. https://doi.org/10.32604/cmc.2024.051778
  42. Ismail, I., Kurnia, R., Brata, Z. A., Nelistiani, G. A., & Kim, H. (2025). Toward robust security orchestration and automated response in security operations centers. Information, 16(5), 365. https://doi.org/10.3390/info16050365
  43. Jacobs, J., Romanosky, S., Suciu, O., Edwards, B., & Sarkar, A. (2023). Enhancing vulnerability prioritization: Data-driven exploit predictions with community-driven insights. In 2023 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW) (pp. 194–206). IEEE. https://doi.org/10.1109/eurospw59978.2023.00027
  44. Jamal, S., Wimmer, H., & Sarker, I. H. (2024). An improved transformer-based model for detecting phishing, spam, and ham emails: A large language model approach. Security and Privacy, 7(3), e402. https://doi.org/10.1002/spy2.402
  45. Jáñez-Martino, F., Alaiz-Rodríguez, R., González-Castro, V., Fidalgo, E., & Alegre, E. (2022). A review of spam email detection: Analysis of spammer strategies and the dataset shift problem. Artificial Intelligence Review, 56, 1145–1173. https://doi.org/10.1007/s10462-022-10195-4
  46. Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., … Zhao, S. (2021). Advances and open problems in federated learning. Foundations and Trends in Machine Learning, 14(1–2), 1–210. https://doi.org/10.1561/2200000083
  47. Kang, H., Liu, G., Wang, Q., Meng, L., & Liu, J. (2023). Theory and application of zero trust security: A brief survey. Entropy, 25(12), 1595. https://doi.org/10.3390/e25121595
  48. Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), 100804. https://doi.org/10.1016/j.patter.2023.100804
  49. Karataş, G., Demir, Ö., & Şahingöz, Ö. K. (2020). Increasing the Performance of Machine Learning-Based IDSs on an Imbalanced and Up-to-Date Dataset. IEEE Access, 8, 32150–32162. https://doi.org/10.1109/access.2020.2973219
  50. Karim, A., Azam, S., Shanmugam, B., Kannoorpatti, K., & Alazab, M. (2019). A Comprehensive Survey for Intelligent Spam Email Detection. IEEE Access, 7, 168261–168295. https://doi.org/10.1109/access.2019.2954791
  51. Karim, A., Shahroz, M., Mustofa, K., Belhaouari, S. B., & Joga, S. R. K. (2023). Phishing Detection System Through Hybrid Machine Learning Based on URL. IEEE Access, 11, 36805–36822. https://doi.org/10.1109/access.2023.3252366
  52. Khan, N., Lee, K., & Houng, A. Y. (2022). Anomaly detection and enterprise security using user and entity behavior analytics (UEBA). In 2022 3rd International Conference on Next Generation Computing Applications (NextComp). IEEE. https://doi.org/10.1109/iconics56716.2022.10100596
  53. Khraisat, A., Gondal, I., Vamplew, P., & Kamruzzaman, J. (2019). Survey of intrusion detection systems: techniques, datasets and challenges. Cybersecurity, 2(1). https://doi.org/10.1186/s42400-019-0038-7
  54. Kolchin, R. (2025). Development and implementation of the Mail Security Guardian (MSG) system for multi-layer proactive email protection against spam, phishing and malware. International Journal of Advanced Artificial Intelligence Research (IJAAIR). [Author’s published study; full volume/issue/pages/DOI to be confirmed at submission.]
  55. Kumar, P., Antony, K., Banga, D., & Sohal, A. S. (2024). PhishNet: A phishing website detection tool using XGBoost. arXiv preprint arXiv:2407.04732. https://doi.org/10.48550/arXiv.2407.04732
  56. Liu, H., & Lang, B. (2019). Machine Learning and Deep Learning Methods for Intrusion Detection Systems: A Survey. Applied Sciences, 9(20), 4396. https://doi.org/10.3390/app9204396
  57. Liu, S., Feng, P., Wang, S., Sun, K., & Cao, J. (2022). Enhancing malware analysis sandboxes with emulated user behavior. Computers & Security, 115, 102613. https://doi.org/10.1016/j.cose.2022.102613
  58. Lu, J., Liu, A., Dong, F., Gu, F., Gama, J., & Zhang, G. (2018). Learning under Concept Drift: A Review. IEEE Transactions on Knowledge and Data Engineering, 1. https://doi.org/10.1109/tkde.2018.2876857
  59. Lundberg, S., & Lee, S. (2017). A Unified Approach to Interpreting Model Predictions. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1705.07874
  60. Macas, M., Wu, C., & Fuertes, W. (2023). Adversarial examples: A survey of attacks and defenses in deep learning-enabled cybersecurity. Expert Systems with Applications, 238, 122223. https://doi.org/10.1016/j.eswa.2023.122223
  61. Mahdavifar, S., & Ghorbani, A. A. (2019). Application of deep learning to cybersecurity: A survey. Neurocomputing, 347, 149–176. https://doi.org/10.1016/j.neucom.2019.02.056
  62. McMahan, H. B., Moore, E., Ramage, D., Hampson, S., & Arcas, B. A. y. (2016). Communication-Efficient Learning of Deep Networks from DecentralizedData. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1602.05629
  63. Mienye, I. D., & Sun, Y. (2022). A Survey of Ensemble Learning: Concepts, Algorithms, Applications, and Prospects. IEEE Access, 10, 99129–99149. https://doi.org/10.1109/access.2022.3207287
  64. Otieno, D. O., Namin, A. S., & Jones, K. S. (2023). The application of the BERT transformer model for phishing email classification. In Proceedings of the 2023 IEEE 47th Annual Computers, Software, and Applications Conference (COMPSAC) (pp. 1303–1308). IEEE. https://doi.org/10.1109/COMPSAC57700.2023.00198
  65. Qiu, S., Liu, Q., Zhou, S., & Wu, C. (2019). Review of Artificial Intelligence Adversarial Attack and Defense Technologies. Applied Sciences, 9(5), 909. https://doi.org/10.3390/app9050909
  66. Ragheb, M., Elmedany, W., & Sharif, M. (2023). The effectiveness of DKIM and SPF in strengthening email security. In Proceedings of the 10th International Conference on Future Internet of Things and Cloud (FiCloud 2023) (pp. 422–426). IEEE. https://doi.org/10.1109/FiCloud58648.2023.00068
  67. Rose, S., Borchert, O., Mitchell, S., & Connelly, S. (2020). Zero Trust Architecture. National Institute of Standards and Technology Special Publication 800-207. https://doi.org/10.6028/nist.sp.800-207
  68. Roy, S. S., Thota, P., Naragam, K. V., & Nilizadeh, S. (2024). From chatbots to phishbots? Phishing scam generation in commercial large language models. In Proceedings of the 2024 IEEE Symposium on Security and Privacy (SP) (pp. 36–54). IEEE. https://doi.org/10.1109/SP54263.2024.00182
  69. Salloum, S. A., Gaber, T., Vadera, S., & Shaalan, K. (2022). A systematic literature review on phishing email detection using natural language processing techniques. IEEE Access, 10, 65703–65727. https://doi.org/10.1109/ACCESS.2022.3183083
  70. Sarker, I. H., Kayes, A. S. M., Badsha, S., Alqahtani, H., Watters, P., & Ng, A. (2020). Cybersecurity data science: an overview from machine learning perspective. Journal Of Big Data, 7(1). https://doi.org/10.1186/s40537-020-00318-5
  71. Saxena, N., Hayes, E., Bertino, E., Ojo, P., Choo, K. R., & Burnap, P. (2020). Impact and Key Challenges of Insider Threats on Organizations and Critical Businesses. Electronics, 9(9), 1460. https://doi.org/10.3390/electronics9091460
  72. Shaukat, K., Luo, S., Varadharajan, V., Hameed, I. A., & Xu, M. (2020). A Survey on Machine Learning Techniques for Cyber Security in the Last Decade. IEEE Access, 8, 222310–222354. https://doi.org/10.1109/access.2020.3041951
  73. Shorten, C., & Khoshgoftaar, T. M. (2019). A survey on Image Data Augmentation for Deep Learning. Journal Of Big Data, 6(1). https://doi.org/10.1186/s40537-019-0197-0
  74. Singh, J., & Singh, J. (2020). A survey on machine learning-based malware detection in executable files. Journal of Systems Architecture, 112, 101861. https://doi.org/10.1016/j.sysarc.2020.101861
  75. Sun, N., Ding, M., Jiang, J., Xu, W., Mo, X., Tai, Y., & Zhang, J. (2023). Cyber Threat Intelligence Mining for Proactive Cybersecurity Defense: A Survey and New Perspectives. IEEE Communications Surveys & Tutorials, 25(3), 1748–1774. https://doi.org/10.1109/comst.2023.3273282
  76. Syed, N. F., Shah, S. W., Shaghaghi, A., Anwar, A., Baig, Z. A., & Doss, R. R. M. (2022). Zero trust architecture (ZTA): A comprehensive survey. IEEE Access, 10, 57143–57179. https://doi.org/10.1109/ACCESS.2022.3174679
  77. Vielberth, M., Böhm, F., Fichtinger, I., & Pernul, G. (2020). Security operations center: A systematic study and open challenges. IEEE Access, 8, 227756–227779. https://doi.org/10.1109/ACCESS.2020.3045514
  78. Vinayakumar, R., Alazab, M., Soman, K. P., Poornachandran, P., Al-Nemrat, A., & Venkatraman, S. (2019). Deep Learning Approach for Intelligent Intrusion Detection System. IEEE Access, 7, 41525–41550. https://doi.org/10.1109/access.2019.2895334
  79. Wen, J., Zhang, Z., Lan, Y., Cui, Z., Cai, J., & Zhang, W. (2022). A survey on federated learning: Challenges and applications. International Journal of Machine Learning and Cybernetics, 14, 513–535. https://doi.org/10.1007/s13042-022-01647-y
  80. Yang, P., Zhao, G., & Zeng, P. (2019). Phishing Website Detection Based on Multidimensional Features Driven by Deep Learning. IEEE Access, 7, 15196–15209. https://doi.org/10.1109/access.2019.2892066
  81. Zhang, Z., Hamadi, H. A., Damiani, E., Yeun, C. Y., & Taher, F. (2022). Explainable Artificial Intelligence Applications in Cyber Security: State-of-the-Art in Research. IEEE Access, 10, 93104–93139. https://doi.org/10.1109/access.2022.3204051
  82. Zieni, R., Massari, L., & Calzarossa, M. C. (2023). Phishing or Not Phishing? A Survey on the Detection of Phishing Websites. IEEE Access, 11, 18499–18519. https://doi.org/10.1109/access.2023.3247135