Articles
| Open Access |
Cloud-Native Multi-Tenant Data Lakes for Intelligent AI Workload Orchestration and Scalability
Dr. Bekzod Karimov , Department of Artificial Intelligence Central Asian Institute of Intelligent Computing Tashkent, Uzbekistan Dr. Dilnoza Ismailova , Department of Computer Science and AI Uzbekistan Center for Digital Intelligence Samarkand, UzbekistanAbstract
The rapid expansion of artificial intelligence (AI), machine learning (ML), and data-intensive applications has increased the need for data-lake architectures capable of supporting heterogeneous workloads, multiple organizational tenants, elastic resource demands, and increasingly complex analytical pipelines. Conventional data-lake implementations frequently treat storage, computation, workload scheduling, and tenant management as loosely connected concerns, limiting their ability to provide predictable scalability and efficient resource utilization. This research and review paper develops a conceptual framework for cloud-native multi-tenant data lakes in which intelligent workload orchestration is integrated with elastic infrastructure management, tenant-aware resource allocation, explainability, and data-quality considerations. The study synthesizes the provided literature on explainable AI, human-centered AI systems, analogy-based reasoning, cognitive bias, crowdsourced knowledge generation, and collaborative decision-making, while positioning these concepts within an intelligent data-lake orchestration model. The framework conceptualizes workload orchestration as a feedback-driven process involving workload characterization, tenant-aware prioritization, resource prediction, execution monitoring, and adaptive rescheduling. The analysis indicates that scalability cannot be treated solely as a capacity problem; it is also a decision-quality, observability, and governance problem. The proposed architecture therefore combines cloud-native elasticity with explainable orchestration decisions and multi-tenant isolation. The study further identifies challenges involving resource contention, fairness, explainability, workload unpredictability, and evaluation validity. The resulting framework provides a research-oriented foundation for designing scalable data lakes capable of supporting intelligent AI workloads across dynamically changing multi-tenant environments.
Keywords
Cloud-native computing, Multi-tenant data lakes, AI workload orchestration, Big data scalability
References
Abdul, A., von der Weth, C., Kankanhalli, M., & Lim, B. Y. (2020). Cogam: measuring and mod-erating cognitive load in machine learning model explanations. InProceedings of the 2020CHI Conference on Human Factors in Computing Systems, pp. 1–14.
Adams, T. L., Li, Y., & Liu, H. (2020). A replication of beyond the turk: Alternative platforms forcrowdsourcing behavioral research–sometimes preferable to student groups.AIS Transactionson Replication Research,6(1), 15.
Aroyo, L., & Welty, C. (2015). Truth is a lie: Crowd truth and the seven myths of human annotation.AI Magazine,36(1), 15–24.
Arrieta, A. B., D ́ıaz-Rodr ́ıguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., Garc ́ıa, S.,Gil-L ́opez, S., Molina, D., Benjamins, R., et al. (2020). Explainable artificial intelligence(xai): Concepts, taxonomies, opportunities and challenges toward responsible ai.Informationfusion,58, 82–115.
Balayn, A., He, G., Hu, A., Yang, J., & Gadiraju, U. (2022a). Ready player one! eliciting diverseknowledge using A configurable game. In Laforest, F., Troncy, R., Simperl, E., Agarwal, D.,Gionis, A., Herman, I., & M ́edini, L. (Eds.),WWW ’22: The ACM Web Conference 2022,Virtual Event, Lyon, France, April 25 - 29, 2022, pp. 1709–1719. ACM.
Balayn, A., Rikalo, N., Lofi, C., Yang, J., & Bozzon, A. (2022b). How can explainability methodsbe used to support bug identification in computer vision models?. InCHI Conference onHuman Factors in Computing Systems, pp. 1–16.
Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., & Weld, D. (2021). Doesthe whole exceed its parts? the effect of ai explanations on complementary team performance.InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems, pp.1–16.
Bartha, P. (2022). Analogy and Analogical Reasoning. In Zalta, E. N. (Ed.),The Stanford Encyclo-pedia of Philosophy(Summer 2022 edition). Metaphysics Research Lab, Stanford University.
Bertrand, A., Belloum, R., Eagan, J. R., & Maxwell, W. (2022). How cognitive biases affect xai-assisted decision-making: A systematic review. InProceedings of the 2022 AAAI/ACM con-ference on AI, ethics, and society, pp. 78–91.
Bounhas, M., Pirlot, M., Prade, H., & Sobrie, O. (2019). Comparison of analogy-based methodsfor predicting preferences. In Amor, N. B., Quost, B., & Theobald, M. (Eds.),ScalableUncertainty Management - 13th International Conference, SUM 2019, Compi`egne, France,December 16-18, 2019, Proceedings, Vol. 11940 ofLecture Notes in Computer Science, pp.339–354. Springer.
K. K. Goyal, "Scalable Data Lakes for AI Workloads: A Multitenant Architecture for Big Data Orchestration," 2025 IEEE International Conference on Computing (ICOCO), Kuching, Malaysia, 2025, pp. 266-271, doi: 10.1109/ICOCO67189.2025.11334100.
Article Statistics
Downloads
Copyright License
Copyright (c) 2026 Dr. Bekzod Karimov, Dr. Dilnoza Ismailova

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Copyright and Ethics:
- Authors are responsible for obtaining permission to use any copyrighted materials included in their manuscript.
- Authors are also responsible for ensuring that their research was conducted in an ethical manner and in compliance with institutional and national guidelines for the care and use of animals or human subjects.
- By submitting a manuscript to International Journal of Computer Science & Information System (IJCSIS), authors agree to transfer copyright to the journal if the manuscript is accepted for publication.