Machine Learning-Driven Performance Anomaly Detection and Auto-Tuning in Distributed Java Full-Stack Systems: A Comprehensive Review
Main Article Content
Abstract
Contemporary distributed Java full-stack applications encounter greater performance complexity with microservices, dynamic workloads, and heterogeneous components. Static monitoring and manual tuning are inadequate anymore. In this review, we discuss the use of machine learning (ML) methods—supervised, unsupervised, and reinforcement learning—towards detecting performance anomalies and auto-tuning. We discuss existing tools, methods, and real-world applications for JVM, databases, and microservices. Major challenges such as noisy data, interpretability, and real-time are brought up, with future research directions focusing on federated learning and self-healing systems. This paper is focused on helping academia and industry build intelligent, adaptive Java-based distributed systems.
Article Details
Section
How to Cite
References
1. Gao et al., Reinforcement Learning for Cloud Resource Management, IEEE TNSM
2. Xu et al., DeepLog: Anomaly Detection from System Logs, IEEE TDSC
3. Wang et al., Monitoring and Diagnosing Distributed Systems Using Time-Series Analysis, IEEE Access
4. Chen et al., Resource Management with Static Thresholds in Microservices, ACM/IEEE SEC
5. Zhang et al., Proactive Failure Prediction in Distributed Systems, IEEE IoT Journal
6. Liu et al., Auto-tuning in Cloud Platforms with Machine Learning, IEEE TCC
7. Huang et al., Self-Adaptive Systems in Cloud Using Online Learning, IEEE TSC
8. Zhao et al., Trust and Explainability in Machine Learning Systems, IEEE Software
9. Netflix Tech Blog, Uber Engineering, LinkedIn Engineering (industry sources referenced for examples)
10. Xu, W., Huang, L., Fox, A., Patterson, D., & Jordan, M. I., “Detecting Large-Scale System Problems by Mining Console Logs,” IEEE Transactions on Dependable and Secure Computing, vol. 10, no. 1, pp. 1–14, Jan.–Feb. 2013.
11. Gao, Y., Liu, C., Xu, Z., & Jin, H., “A Reinforcement Learning Approach for Auto-Scaling Microservices in Cloud Applications,” IEEE Transactions on Network and Service Management, vol. 18, no. 3, pp. 3187–3202, Sep. 2021.
12. Zhang, X., Wang, Y., Li, X., & Zhou, M., “Anomaly Detection in Cloud Computing Systems Using Behavioral Patterns,” IEEE Internet of Things Journal, vol. 7, no. 9, pp. 8611–8622, Sep. 2020.
13. Wang, J., Chen, Z., Zheng, Z., & Lyu, M. R., “Performance Monitoring and Diagnosis for Microservice Systems via Metric Correlation Analysis,” IEEE Access, vol. 7, pp. 77846–77857, 2019.
14. Liu, F., Meng, D., Zhang, W., & Tan, Y., “Intelligent Auto-Tuning of Cloud Services Configuration Parameters Using Machine Learning,” IEEE Transactions on Cloud Computing, vol. 10, no. 2, pp. 1058–1070, Apr.–Jun. 2022.
15. Huang, Y., He, Q., & Jin, H., “Online Learning for Self-Adaptive Cloud Systems: A Survey,” IEEE Transactions on Services Computing, vol. 15, no. 1, pp. 219–237, Jan.–Feb. 2022.
16. Chen, J., Liu, Z., He, Q., & Chan, W. K., “Learning-Based Threshold Configuration for Performance Anomaly Detection in Microservices,” Proc. ACM/IEEE Symposium on Edge Computing (SEC), pp. 78–91, 2021.
17. Zhao, H., Xu, Y., Lu, Z., & Liu, F., “Explainable Machine Learning for Anomaly Detection in Large-Scale Systems,” IEEE Software, vol. 39, no. 2, pp. 80–88, Mar.–Apr. 2022.
18. Netflix Technology Blog, “Auto-Tuning at Scale: Netflix’s Approach with Scryer,” [Online]. Available: https://netflixtechblog.com
19. LinkedIn Engineering Blog, “PerfBot: Machine Learning for JVM Optimization at Scale,” [Online]. Available: https://engineering.linkedin.com
20. Uber Engineering Blog, “Michelangelo: Machine Learning Platform at Uber,” [Online]. Available: https://eng.uber.com/michelangelo
21. Twitter Engineering, “Adaptive Systems for Performance Management,” [Online]. Available: https://blog.twitter.com
22. Kim, S., Park, J., & Han, S., “Log-based Performance Anomaly Detection for Cloud Applications Using LSTM and Attention Mechanism,” IEEE Access, vol. 9, pp. 112789–112801, 2021.
23. Breier, J., & Hudec, L., “Anomaly Detection from System Logs Using Machine Learning: A Survey,” IEEE Access, vol. 9, pp. 120379–120396, 2021.
24. Duan, R., Huang, Q., Xie, Q., & Liu, D., “AIOps-Based Predictive Auto-Scaling in Cloud Environments Using Adaptive Reinforcement Learning,” IEEE Transactions on Network and Service Management, vol. 19, no. 1, pp. 410–422, Mar. 2022.
25. Lin, Q., Ding, Y., & Xu, W., “MicroRCA: Root Cause Localization of Performance Issues in Microservices,” IEEE Transactions on Cloud Computing, vol. 11, no. 1, pp. 37–49, Jan.–Mar. 2023.
26. Dai, T., Gao, Y., Hu, L., & Pan, G., “Detecting Microservice Performance Anomalies Through Ensemble Learning,” IEEE Transactions on Services Computing, vol. 15, no. 3, pp. 1192–1206, May–Jun. 2022.