Modernizing Legacy Big Data Platforms Using Cloud-Native Lakehouse Architectures

Main Article Content

Karthikeyan Selvarajan

Abstract

The capacity to scale up the infrastructure, intricacy of the infrastructure, utilization of resources, governance of data, and compatibility with new analytics are just a few issues with the old models of the Hadoop-based big data systems of the past. This research paper suggests the implementation of cloud-based lakehouse architecture to upgrade the outdated big data systems implemented on Hadoop that has been integrated into the existing data and analytics without affecting the existing data resources. The suggested solution will result in a scaled and expandable data management space in cloud object storing methods, Apache Spark, Databricks, distributed data processing, an open table format, and centralized governance. A measured mode of modernization, which includes an assessment of the workload, data migration, transformation of the layers of processing, the addition of metadata, governance, and optimization of the workload, is planned. It also allows the storage and compute capabilities to be independently scaled as well as having the ability of running batch processing, streaming analytics, machine learning, and business intelligence on the same platform. The framework adds controls on data quality and metadata, security policies, lineage Garnier, cost-sensitive deployment of resources, to the functionality of the traditional Hadoop deployments. Architectural considerations show that the proposed lakehouse model has the ability to ease the management of the infrastructure, enhance resource elasticity, reduce platform fragmentation, and allow various analytical workloads. The paper demonstrates how the ideas of the lakehouse that are cloud native provide an organized guideline to organizations that are planning on upgrading their past big data infrastructures without sacrificing the interoperability, governance and operational continuity.

Article Details

Section

Articles

How to Cite

Modernizing Legacy Big Data Platforms Using Cloud-Native Lakehouse Architectures. (2024). International Journal of Research Publications in Engineering, Technology and Management (IJRPETM), 7(2), 10377-10385. https://doi.org/10.15662/IJRPETM.2024.0702008

References

[1] M. Armbrust et al., “Lakehouse: A new generation of open platforms that unify data warehousing and advanced analytics,” Proceedings of the VLDB Endowment, 2021.

[2] M. Armbrust et al., “Delta Lake: High-performance ACID table storage over cloud object stores,” Proceedings of the VLDB Endowment, vol. 13, no. 12, pp. 3411–3424, 2020.

[3] M. R. Llave, “Data lakes in business intelligence: Reporting from the trenches,” Procedia Computer Science, vol. 138, pp. 516–524, 2018.

[4] H. Decker et al., “Data lakes: Trends and perspectives,” Proceedings of the VLDB Endowment, 2019.

[5] P. P. Khine and Z. S. Wang, “Data lake: A new ideology in big data era,” Proceedings of the IEEE International Conference on Smart Cloud (SmartCloud), 2018.

[6] A. Behm et al., “Photon: A fast query engine for lakehouse systems,” Proceedings of the VLDB Endowment, 2022.

[7] T. Y. Chen et al., “On construction of a power data lake platform using Spark,” International Journal of Computer Science and Information Security, 2019.

[8] R. Hai, S. Geisler, and C. Quix, “Constance: An intelligent data lake system,” Proceedings of the ACM SIGMOD International Conference on Management of Data, 2016.

[9] A. Laplante et al., “Architecting data lakes,” IEEE Software, 2016.

[10] C. Walker et al., “Personal data lake with data gravity pull,” Future Generation Computer Systems, 2015.