Engineering and Technology Book | Open Access | DOI: https://doi.org/10.37547/tajet/book-26-04

An Author’s Framework for Applying Machine Learning in Multi-Layer Enterprise Cybersecurity Systems Pre-emptive Architecture, Combined Labeling, and Ensemble Integration

Abstract

This book sets out a framework for applying machine learning in corporate cybersecurity that the author developed across fifteen years of building and operating defensive systems. The framework did not begin as a theory. It began as a set of recurring engineering decisions, made and remade across several security products, that turned out to matter more than the choice of any individual model: when a model is allowed to act, what evidence it is taught from, and where it sits among the other controls in a defense. Naming those decisions as principles, and showing that their combination can be verified, is the purpose of the book.

The intended reader is a practitioner. The book is addressed to the engineers, architects, and security leaders who design machine-learning systems for enterprise defense and who must make these decisions whether or not they are made deliberately. For that reader the book aims to be useful rather than exhaustive: it states each principle, gives the architectural requirements it imposes, and grounds it in a system that was built and measured. The quantitative core of the book is an email security platform whose results were established in a peer-reviewed study, and the book is careful throughout to separate what has been measured on that system from what is projected for the adjacent domains where the same principles apply.

The book treats the published literature as a working tool rather than as a survey to be completed. Sources are cited where they establish a mechanism, corroborate a claim, or mark a limit, and the reader who wishes to go deeper will find the trail. The aim is a book that reads as the distilled judgment of practice, supported by the literature, rather than as a compilation of it.

Keywords

References

Aboaoja, F. A., Zainal, A. B., Ali, A. M., Ghaleb, F. A., Alsolami, F. J., & Rassam, M. A. (2023). Dynamic extraction of initial behavior for evasive malware detection. Mathematics, 11(2), 416. https://doi.org/10.3390/math11020416

Adnan, M., Imam, M. O., Javed, M. F., & Murtza, I. (2024). Improving spam email classification accuracy using ensemble techniques: A stacking approach. International Journal of Information Security, 23, 505–517. https://doi.org/10.1007/s10207-023-00756-1

Afianian, A., Niksefat, S., Sadeghiyan, B., & Baptiste, D. (2019). Malware dynamic analysis evasion techniques: A survey. ACM Computing Surveys, 52(6), 126. https://doi.org/10.1145/3365001

Ahmad, Z., Khan, A. S., Shiang, C. W., Abdullah, J., & Ahmad, F. (2020). Network intrusion detection system: A systematic study of machine learning and deep learning approaches. Transactions on Emerging Telecommunications Technologies, 32(1). https://doi.org/10.1002/ett.4150

Ahmed, N., Amin, R., Aldabbas, H., Koundal, D., Alouffi, B., & Shah, T. (2022). Machine learning techniques for spam detection in email and IoT platforms: Analysis and research challenges. Security and Communication Networks, 2022, 1862888. https://doi.org/10.1155/2022/1862888

Alabdan, R. (2020). Phishing Attacks Survey: Types, Vectors, and Technical Approaches. Future Internet, 12(10), 168. https://doi.org/10.3390/fi12100168

Aldweesh, A., Derhab, A., & Emam, A. (2019). Deep learning approaches for anomaly-based intrusion detection systems: A survey, taxonomy, and open issues. Knowledge-Based Systems, 189, 105124. https://doi.org/10.1016/j.knosys.2019.105124

Aljofey, A., Jiang, Q., Qu, Q., Huang, M., & Niyigena, J. (2020). An Effective Phishing Detection Model Based on Character Level Convolutional Neural Network from URL. Electronics, 9(9), 1514. https://doi.org/10.3390/electronics9091514

Almujahid, N. F., Haq, M. A., & Alshehri, M. (2024). Comparative evaluation of machine learning algorithms for phishing site detection. PeerJ Computer Science, 10, e2131. https://doi.org/10.7717/peerj-cs.2131

Alnemari, S., & Alshammari, M. (2023). Detecting phishing domains using machine learning. Applied Sciences, 13(8), 4649. https://doi.org/10.3390/app13084649

Altwaijry, N., Al-Turaiki, I., Alotaibi, R., & Alakeel, F. (2024). Advancing phishing email detection: A comparative study of deep learning models. Sensors, 24(7), 2077. https://doi.org/10.3390/s24072077

Alzubaidi, L., Zhang, J., Humaidi, A. J., Al-Dujaili, A. Q., Duan, Y., Al-Shamma, O., Santamaría, J., Fadhel, M. A., Al‐Amidie, M., & Farhan, L. (2021). Review of deep learning: concepts, CNN architectures, challenges, applications, future directions. Journal Of Big Data, 8(1), 53. https://doi.org/10.1186/s40537-021-00444-8

Anti-Phishing Working Group. (2023). Phishing activity trends report: 4th quarter 2023. APWG. https://docs.apwg.org/reports/apwg_trends_report_q4_2023.pdf

Apruzzese, G., Laskov, P., de, E. M., Mallouli, W., Rapa, L. B., Grammatopoulos, A. V., & Franco, F. D. (2022). The Role of Machine Learning in Cybersecurity. Digital Threats Research and Practice, 4(1), 1–38. https://doi.org/10.1145/3545574

Atlam, H. F., & Oluwatimilehin, O. (2023). Business email compromise phishing detection based on machine learning: A systematic literature review. Electronics, 12(1), 42. https://doi.org/10.3390/electronics12010042

Bagui, S., & Li, K. (2021). Resampling imbalanced data for network intrusion detection datasets. Journal Of Big Data, 8(1). https://doi.org/10.1186/s40537-020-00390-x

Ban, T., Takahashi, T., Ndichu, S., & Inoue, D. (2023). Breaking alert fatigue: AI-assisted SIEM framework for effective incident response. Applied Sciences, 13(11), 6610. https://doi.org/10.3390/app13116610

Basit, A., Zafar, M., Liu, X., Javed, A. R., Jalil, Z., & Kifayat, K. (2020). A comprehensive survey of AI-enabled phishing attacks detection techniques. Telecommunication Systems, 76(1), 139–154. https://doi.org/10.1007/s11235-020-00733-2

Bensaoud, A., Kalita, J., & Bensaoud, M. (2024). A survey of malware detection using deep learning. Machine Learning with Applications, 16, 100546. https://doi.org/10.1016/j.mlwa.2024.100546

Breiman, L. (2001). Random Forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/a:1010933404324

Cao, Y., & Pokhrel, S. R. (2024). Automation and orchestration of zero trust architecture: Potential solutions and challenges. Machine Intelligence Research, 21, 294–317. https://doi.org/10.1007/s11633-023-1456-2

Cervantes, J., García‐Lamont, F., Rodríguez-Mazahua, L., & López‐Chau, A. (2020). A comprehensive survey on support vector machine classification: Applications, challenges and trends. Neurocomputing, 408, 189–215. https://doi.org/10.1016/j.neucom.2019.10.118

Chakraborty, A., Alam, M., Dey, V., Chattopadhyay, A., & Mukhopadhyay, D. (2021). A survey on adversarial attacks and defences. CAAI Transactions on Intelligence Technology, 6(1), 25–45. https://doi.org/10.1049/cit2.12028

Chawla, N. V., Bowyer, K. W., Hall, L., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence Research, 16, 321–357. https://doi.org/10.1613/jair.953

Chen, T., & Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System. KDD '16: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. Pages 785 - 794. https://doi.org/10.1145/2939672.2939785

Devlin, J., Chang, M., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 4171–4186. https://doi.org/10.18653/v1/n19-1423

Doshi, J., Parmar, K., Sanghavi, R., & Shekokar, N. M. (2023). A comprehensive dual-layer architecture for phishing and spam email detection. Computers & Security, 133, 103378. https://doi.org/10.1016/j.cose.2023.103378

Elsadig, M., Ibrahim, A. O., Basheer, S., Alohali, M. A., Alshunaifi, S., Alqahtani, H. M., Alharbi, N., & Nagmeldin, W. (2022). Intelligent Deep Machine Learning Cyber Phishing URL Detection Based on BERT Features Extraction. Electronics, 11(22), 3647. https://doi.org/10.3390/electronics11223647

Engelen, J. E. v., & Hoos, H. H. (2019). A survey on semi-supervised learning. Machine Learning, 109(2), 373–440. https://doi.org/10.1007/s10994-019-05855-6

Fang, Y., Liu, Y., Huang, C., & Liu, L. (2020). FastEmbed: Predicting vulnerability exploitation possibility based on ensemble machine learning algorithm. PLoS ONE, 15(2), e0228439. https://doi.org/10.1371/journal.pone.0228439

Gamage, S., & Samarabandu, J. (2020). Deep learning methods in network intrusion detection: A survey and an objective comparison. Journal of Network and Computer Applications, 169, 102767. https://doi.org/10.1016/j.jnca.2020.102767

Gambo, L., & Almulhem, A. (2025). Zero trust architecture: A systematic literature review. Journal of Network and Systems Management. https://doi.org/10.1007/s10922-025-09998-x

Gholampour, P. M., & Verma, R. M. (2023). Adversarial robustness of phishing email detection models. In Proceedings of the 9th ACM International Workshop on Security and Privacy Analytics (IWSPA ’23) (pp. 67–76). ACM. https://doi.org/10.1145/3579987.3586567

Goodfellow, I., Shlens, J., & Szegedy, C. (2014). Explaining and Harnessing Adversarial Examples. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1412.6572

Gopinath, M. A., & Sethuraman, S. C. (2022). A comprehensive survey on deep learning based malware detection techniques. Computer Science Review, 47, 100529. https://doi.org/10.1016/j.cosrev.2022.100529

Han, X., & Zhang, T. (2022). Text adversarial attacks and defenses: Issues, taxonomy, and perspectives. Security and Communication Networks, 2022, 6458488. https://doi.org/10.1155/2022/6458488

Hao, W., Tran, V., Rideout, V., Wang, Z., Dasbach-Prisk, A., Afifi, M. H., Yang, J., Katz-Bassett, E., Ho, G., & Cidon, A. (2025). Do spammers dream of electric sheep? Characterizing the prevalence of LLM-generated malicious emails. In Proceedings of the 2025 ACM Internet Measurement Conference (IMC ’25) (pp. 192–203). ACM. https://doi.org/10.1145/3730567.3732922

Hazell, J. (2023). Spear phishing with large language models. arXiv preprint arXiv:2305.06972. https://doi.org/10.48550/arXiv.2305.06972

Hemalatha, J., Roseline, S. A., Geetha, S., Kadry, S., & Damaševičius, R. (2021). An Efficient DenseNet-Based Deep Learning Model for Malware Detection. Entropy, 23(3), 344. https://doi.org/10.3390/e23030344

Hureau, O., Bayer, J., Duda, A., & Korczyński, M. (2024). Spoofed emails: An analysis of the issues hindering a larger deployment of DMARC. In Passive and Active Measurement (PAM 2024), Lecture Notes in Computer Science (Vol. 14537, pp. 232–261). Springer. https://doi.org/10.1007/978-3-031-56249-5_10

Innab, N., Osman, A. A. F., Ataelfadiel, M. A. M., Abu-Zanona, M., Elzaghmouri, B. M., Zawaideh, F. H., & Alawneh, M. F. (2024). Phishing attacks detection using ensemble machine learning algorithms. Computers, Materials & Continua, 80(1), 1325–1345. https://doi.org/10.32604/cmc.2024.051778

Ismail, I., Kurnia, R., Brata, Z. A., Nelistiani, G. A., & Kim, H. (2025). Toward robust security orchestration and automated response in security operations centers. Information, 16(5), 365. https://doi.org/10.3390/info16050365

Jacobs, J., Romanosky, S., Suciu, O., Edwards, B., & Sarkar, A. (2023). Enhancing vulnerability prioritization: Data-driven exploit predictions with community-driven insights. In 2023 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW) (pp. 194–206). IEEE. https://doi.org/10.1109/eurospw59978.2023.00027

Jamal, S., Wimmer, H., & Sarker, I. H. (2024). An improved transformer-based model for detecting phishing, spam, and ham emails: A large language model approach. Security and Privacy, 7(3), e402. https://doi.org/10.1002/spy2.402

Jáñez-Martino, F., Alaiz-Rodríguez, R., González-Castro, V., Fidalgo, E., & Alegre, E. (2022). A review of spam email detection: Analysis of spammer strategies and the dataset shift problem. Artificial Intelligence Review, 56, 1145–1173. https://doi.org/10.1007/s10462-022-10195-4

Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., … Zhao, S. (2021). Advances and open problems in federated learning. Foundations and Trends in Machine Learning, 14(1–2), 1–210. https://doi.org/10.1561/2200000083

Kang, H., Liu, G., Wang, Q., Meng, L., & Liu, J. (2023). Theory and application of zero trust security: A brief survey. Entropy, 25(12), 1595. https://doi.org/10.3390/e25121595

Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), 100804. https://doi.org/10.1016/j.patter.2023.100804

Karataş, G., Demir, Ö., & Şahingöz, Ö. K. (2020). Increasing the Performance of Machine Learning-Based IDSs on an Imbalanced and Up-to-Date Dataset. IEEE Access, 8, 32150–32162. https://doi.org/10.1109/access.2020.2973219

Karim, A., Azam, S., Shanmugam, B., Kannoorpatti, K., & Alazab, M. (2019). A Comprehensive Survey for Intelligent Spam Email Detection. IEEE Access, 7, 168261–168295. https://doi.org/10.1109/access.2019.2954791

Karim, A., Shahroz, M., Mustofa, K., Belhaouari, S. B., & Joga, S. R. K. (2023). Phishing Detection System Through Hybrid Machine Learning Based on URL. IEEE Access, 11, 36805–36822. https://doi.org/10.1109/access.2023.3252366

Khan, N., Lee, K., & Houng, A. Y. (2022). Anomaly detection and enterprise security using user and entity behavior analytics (UEBA). In 2022 3rd International Conference on Next Generation Computing Applications (NextComp). IEEE. https://doi.org/10.1109/iconics56716.2022.10100596

Khraisat, A., Gondal, I., Vamplew, P., & Kamruzzaman, J. (2019). Survey of intrusion detection systems: techniques, datasets and challenges. Cybersecurity, 2(1). https://doi.org/10.1186/s42400-019-0038-7

Kolchin, R. (2025). Development and implementation of the Mail Security Guardian (MSG) system for multi-layer proactive email protection against spam, phishing and malware. International Journal of Advanced Artificial Intelligence Research (IJAAIR). [Author’s published study; full volume/issue/pages/DOI to be confirmed at submission.]

Kumar, P., Antony, K., Banga, D., & Sohal, A. S. (2024). PhishNet: A phishing website detection tool using XGBoost. arXiv preprint arXiv:2407.04732. https://doi.org/10.48550/arXiv.2407.04732

Liu, H., & Lang, B. (2019). Machine Learning and Deep Learning Methods for Intrusion Detection Systems: A Survey. Applied Sciences, 9(20), 4396. https://doi.org/10.3390/app9204396

Liu, S., Feng, P., Wang, S., Sun, K., & Cao, J. (2022). Enhancing malware analysis sandboxes with emulated user behavior. Computers & Security, 115, 102613. https://doi.org/10.1016/j.cose.2022.102613

Lu, J., Liu, A., Dong, F., Gu, F., Gama, J., & Zhang, G. (2018). Learning under Concept Drift: A Review. IEEE Transactions on Knowledge and Data Engineering, 1. https://doi.org/10.1109/tkde.2018.2876857

Lundberg, S., & Lee, S. (2017). A Unified Approach to Interpreting Model Predictions. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1705.07874

Macas, M., Wu, C., & Fuertes, W. (2023). Adversarial examples: A survey of attacks and defenses in deep learning-enabled cybersecurity. Expert Systems with Applications, 238, 122223. https://doi.org/10.1016/j.eswa.2023.122223

Mahdavifar, S., & Ghorbani, A. A. (2019). Application of deep learning to cybersecurity: A survey. Neurocomputing, 347, 149–176. https://doi.org/10.1016/j.neucom.2019.02.056

McMahan, H. B., Moore, E., Ramage, D., Hampson, S., & Arcas, B. A. y. (2016). Communication-Efficient Learning of Deep Networks from DecentralizedData. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1602.05629

Mienye, I. D., & Sun, Y. (2022). A Survey of Ensemble Learning: Concepts, Algorithms, Applications, and Prospects. IEEE Access, 10, 99129–99149. https://doi.org/10.1109/access.2022.3207287

Otieno, D. O., Namin, A. S., & Jones, K. S. (2023). The application of the BERT transformer model for phishing email classification. In Proceedings of the 2023 IEEE 47th Annual Computers, Software, and Applications Conference (COMPSAC) (pp. 1303–1308). IEEE. https://doi.org/10.1109/COMPSAC57700.2023.00198

Qiu, S., Liu, Q., Zhou, S., & Wu, C. (2019). Review of Artificial Intelligence Adversarial Attack and Defense Technologies. Applied Sciences, 9(5), 909. https://doi.org/10.3390/app9050909

Ragheb, M., Elmedany, W., & Sharif, M. (2023). The effectiveness of DKIM and SPF in strengthening email security. In Proceedings of the 10th International Conference on Future Internet of Things and Cloud (FiCloud 2023) (pp. 422–426). IEEE. https://doi.org/10.1109/FiCloud58648.2023.00068

Rose, S., Borchert, O., Mitchell, S., & Connelly, S. (2020). Zero Trust Architecture. National Institute of Standards and Technology Special Publication 800-207. https://doi.org/10.6028/nist.sp.800-207

Roy, S. S., Thota, P., Naragam, K. V., & Nilizadeh, S. (2024). From chatbots to phishbots? Phishing scam generation in commercial large language models. In Proceedings of the 2024 IEEE Symposium on Security and Privacy (SP) (pp. 36–54). IEEE. https://doi.org/10.1109/SP54263.2024.00182

Salloum, S. A., Gaber, T., Vadera, S., & Shaalan, K. (2022). A systematic literature review on phishing email detection using natural language processing techniques. IEEE Access, 10, 65703–65727. https://doi.org/10.1109/ACCESS.2022.3183083

Sarker, I. H., Kayes, A. S. M., Badsha, S., Alqahtani, H., Watters, P., & Ng, A. (2020). Cybersecurity data science: an overview from machine learning perspective. Journal Of Big Data, 7(1). https://doi.org/10.1186/s40537-020-00318-5

Saxena, N., Hayes, E., Bertino, E., Ojo, P., Choo, K. R., & Burnap, P. (2020). Impact and Key Challenges of Insider Threats on Organizations and Critical Businesses. Electronics, 9(9), 1460. https://doi.org/10.3390/electronics9091460

Shaukat, K., Luo, S., Varadharajan, V., Hameed, I. A., & Xu, M. (2020). A Survey on Machine Learning Techniques for Cyber Security in the Last Decade. IEEE Access, 8, 222310–222354. https://doi.org/10.1109/access.2020.3041951

Shorten, C., & Khoshgoftaar, T. M. (2019). A survey on Image Data Augmentation for Deep Learning. Journal Of Big Data, 6(1). https://doi.org/10.1186/s40537-019-0197-0

Singh, J., & Singh, J. (2020). A survey on machine learning-based malware detection in executable files. Journal of Systems Architecture, 112, 101861. https://doi.org/10.1016/j.sysarc.2020.101861

Sun, N., Ding, M., Jiang, J., Xu, W., Mo, X., Tai, Y., & Zhang, J. (2023). Cyber Threat Intelligence Mining for Proactive Cybersecurity Defense: A Survey and New Perspectives. IEEE Communications Surveys & Tutorials, 25(3), 1748–1774. https://doi.org/10.1109/comst.2023.3273282

Syed, N. F., Shah, S. W., Shaghaghi, A., Anwar, A., Baig, Z. A., & Doss, R. R. M. (2022). Zero trust architecture (ZTA): A comprehensive survey. IEEE Access, 10, 57143–57179. https://doi.org/10.1109/ACCESS.2022.3174679

Vielberth, M., Böhm, F., Fichtinger, I., & Pernul, G. (2020). Security operations center: A systematic study and open challenges. IEEE Access, 8, 227756–227779. https://doi.org/10.1109/ACCESS.2020.3045514

Vinayakumar, R., Alazab, M., Soman, K. P., Poornachandran, P., Al-Nemrat, A., & Venkatraman, S. (2019). Deep Learning Approach for Intelligent Intrusion Detection System. IEEE Access, 7, 41525–41550. https://doi.org/10.1109/access.2019.2895334

Wen, J., Zhang, Z., Lan, Y., Cui, Z., Cai, J., & Zhang, W. (2022). A survey on federated learning: Challenges and applications. International Journal of Machine Learning and Cybernetics, 14, 513–535. https://doi.org/10.1007/s13042-022-01647-y

Yang, P., Zhao, G., & Zeng, P. (2019). Phishing Website Detection Based on Multidimensional Features Driven by Deep Learning. IEEE Access, 7, 15196–15209. https://doi.org/10.1109/access.2019.2892066

Zhang, Z., Hamadi, H. A., Damiani, E., Yeun, C. Y., & Taher, F. (2022). Explainable Artificial Intelligence Applications in Cyber Security: State-of-the-Art in Research. IEEE Access, 10, 93104–93139. https://doi.org/10.1109/access.2022.3204051

Zieni, R., Massari, L., & Calzarossa, M. C. (2023). Phishing or Not Phishing? A Survey on the Detection of Phishing Websites. IEEE Access, 11, 18499–18519. https://doi.org/10.1109/access.2023.3247135

Download and View Statistics

Views: 0   |   Downloads: 0

Copyright License

Download Citations

How to Cite

Kolchin, R. (2026). An Author’s Framework for Applying Machine Learning in Multi-Layer Enterprise Cybersecurity Systems Pre-emptive Architecture, Combined Labeling, and Ensemble Integration. The American Journal of Engineering and Technology, 01–59. https://doi.org/10.37547/tajet/book-26-04