Resilient and Scalable Data Lake Management Using Intelligent Orchestration for AI and Big Data

Authors

  • Dr. Rafi Pratama Department of Artificial Intelligence Indonesian Institute of Intelligent Technologies Jakarta, Indonesia

DOI:

https://doi.org/10.37547/ibast/Volume06Issue08-05

Keywords:

Data Lakes, Intelligent Orchestration, Artificial Intelligence, Big Data

Abstract

The rapid convergence of artificial intelligence (AI), multimodal analytics, and large-scale data processing has intensified the need for data lake environments that can accommodate heterogeneous workloads while maintaining scalability, resilience, fairness, and operational efficiency. Conventional data lake management approaches often treat storage, processing, workload scheduling, and governance as relatively independent functions, which can create bottlenecks when AI workloads become dynamic and resource-intensive. This paper proposes a conceptual framework for resilient and scalable data lake management based on intelligent orchestration, in which workload characteristics, data-processing requirements, fairness constraints, and system resilience are jointly considered during resource and task coordination. The study develops its theoretical foundation by synthesizing the provided literature on fairness-aware learning, multimodal datasets, bias reduction, and representation learning, while connecting these principles to intelligent data-lake orchestration. Particular attention is given to multitenant environments in which concurrent AI and big-data workloads compete for computational and storage resources. The proposed approach emphasizes adaptive workload classification, fairness-aware resource allocation, resilience-oriented task management, and continuous orchestration. The analysis indicates that intelligent orchestration can provide a stronger conceptual basis for balancing scalability with reliability and responsible AI requirements. The findings further suggest that fairness should not be treated exclusively as a model-level concern but should also influence data and workload management decisions. The resulting framework offers a research direction for building adaptive data-lake infrastructures capable of supporting heterogeneous AI workloads without sacrificing operational resilience or responsible data utilization.

Downloads

Download data is not yet available.

References

1. Xu, T., White, J., Kalkan, S., & Gunes, H. (2020). Investigating Bias And Fairness In Facialexpression Recognition. Incomputer Vision–Eccv 2020 Workshops: Glasgow, Uk,August 23–28, 2020, Proceedings, Part Vi 16, Pp. 506–523. Springer.

2. Yoon, J., Kang, C., Kim, S., & Han, J. (2022). D-Vlog: Multimodal Vlog Dataset For Depres-Sion Detection.Proceedings Of The Aaai Conference On Artificial Intelligence,36(11),12226–12234.

3. Zafar, M. B., Valera, I., Gomez Rodriguez, M., & Gummadi, K. P. (2017). Fairness Beyonddisparate Treatment & Disparate Impact: Learning Classification Without Disparatemistreatment. Inproceedings Of The 26th International Conference On World Wide Web,Pp. 1171–1180.

4. Zanna, K., Sridhar, K., Yu, H., & Sano, A. (2022). Bias Reducing Multitask Learning Onmental Health Prediction..

5. Zemel, R., Wu, Y., Swersky, K., Pitassi, T., & Dwork, C. (2013). Learning Fair Representa-Tions. Ininternational Conference On Machine Learning, Pp. 325–333. Pmlr.

6. K. K. Goyal, "Scalable Data Lakes For Ai Workloads: A Multitenant Architecture For Big Data Orchestration," 2025 Ieee International Conference On Computing (Icoco), Kuching, Malaysia, 2025, Pp. 266-271, Doi: 10.1109/Icoco67189.2025.11334100.

Published

2026-08-21

How to Cite

Resilient and Scalable Data Lake Management Using Intelligent Orchestration for AI and Big Data. (2026). International Bulletin of Applied Science and Technology, 6(8), 52-59. https://doi.org/10.37547/ibast/Volume06Issue08-05

Similar Articles

11-20 of 301

You may also start an advanced similarity search for this article.