Methodological Foundations for Merging Structured and Unstructured Sources in ML Pipelines
Read Abstract
The article presents a theoretical and applied analysis of the methodological foundations for merging structured and unstructured data sources within machine learning systems. The study is based on an interdisciplinary approach that integrates architectural design of ML pipelines, data representation theory, and practices of heterogeneous format integration. Particular attention is paid to the analysis of recent scientific publications highlighting the application of Retrieve–Merge–Predict architectures, agent-based discovery systems, and multimodal frameworks involving large language models. Four stable strategies for data merging are identified, ranging from static unification to end-to-end processing within a unified training loop. The importance of selecting matching metrics and adaptation schemes when dealing with unstable data streams is demonstrated using experiments from Cappuzzo and Eltabakh. Special emphasis is placed on the methodological limitations of universal solutions, including the generalization paradox, sensitivity to structural evolution, and the lack of formalized testing scenarios in agent-oriented pipelines. It is shown that sustainable development of ML architectures requires a shift from linear ETL pipelines to coherent, iterative systems with internal adaptation and feedback from the model to the data. The article will be of interest to researchers in data preparation automation, developers of multimodal ML systems, data engineering specialists, and digital platform architects working with multi-format sources.