Full Article: PDF
Scientific Object Identifier: http://s-o-i.org/1.1/TAS-01-153-16
DOI: https://dx.doi.org/10.15863/TAS.2026.01.153.16
Language: English
Citation: Mukayev, T. (2026). Modern methods of monitoring and logging in distributed machine learning systems. ISJ Theoretical & Applied Science, 01 (153), 144-151. Soi: https://s-o-i.org/1.1/TAS-01-153-16 Doi: https://dx.doi.org/10.15863/TAS.2026.01.153.16 |
Pages: 144-151
Published: 30.01.2026
Abstract: The article examines modern approaches to implementing observability in distributed machine learning systems, including the configuration of metrics, event logging, and request tracing. The role of tools such as Prometheus, Grafana, the ELK stack, Fluentd, Loki, and OpenTelemetry in ensuring the resilience, transparency, and controllability of machine learning system infrastructures is analysed. The importance of integrating observability into model lifecycle management pipelines at all stages – from data preparation to operation in the production environment – is emphasised. Typical constraints related to scalability, caching, and the complexity of diagnostics in microservice architectures are identified. It is concluded that the deployment of an effective observability architecture requires a comprehensive and systematic approach.
Key words: observability, monitoring, logging, tracing, MLOps, distributed systems.
|