Cloud-Native Architectures for Deploying Generative Intelligence Systems

Authors

  • Matteo Garcia Associate Professor, Institute of Intelligent Systems, Western Europe Data Science University, Madrid, Spain Author https://orcid.org/6913-1181-8514-5804
  • Eva Klein Assistant Professor, Institute of Intelligent Systems, European Institute of AI, Berlin, Germany Author

Keywords:

cloud-native, generative AI, deployment architecture, Kubernetes, auto-scaling, multi-tenancy, GPU serving, microservices

Abstract

The deployment of large generative AI systems at production scale requires cloud infrastructure architectures that can handle highly variable request loads, provide horizontal scalability, ensure high availability, and integrate with existing enterprise data and security infrastructure -- challenges that differ fundamentally from traditional software deployment due to the compute intensity, memory requirements, and latency sensitivity of generative AI inference. Cloud-native architectures -- systems built on containerisation, microservices, orchestration platforms, and auto-scaling -- provide the infrastructure foundation for this deployment at scale, but standard cloud-native patterns require substantial adaptation to accommodate the specific requirements of generative AI workloads. This paper proposes the Generative AI Cloud-Native Deployment (GACND) framework, a comprehensive reference architecture for deploying large generative AI systems on cloud infrastructure, comprising five architectural layers: a generative model serving layer (GMSL) with GPU-optimised container orchestration; an intelligent request routing layer (IRRL) that matches requests to appropriate model variants based on complexity and latency requirements; an adaptive auto-scaling layer (AASL) that responds to GPU memory pressure and request queue depth; a multi-tenancy and isolation layer (MTIL) that provides secure model serving across multiple tenant organisations; and a monitoring and observability layer (MOL) with generative AI-specific metrics. GACND is evaluated through a 90-day production deployment study on a cloud infrastructure serving 12 enterprise tenants with diverse generative AI workloads, measuring availability, latency, cost efficiency, and security isolation. GACND achieves 99.94% availability (SD = 0.02%), mean TTFT = 318ms (SD = 68ms) at peak load, 42.8% infrastructure cost reduction versus naive GPU reservation, and zero cross-tenant data leakage events. The study contributes the GACND reference architecture, a Deployment Efficiency Index (DEI), and empirical evidence for the substantial operational benefits of cloud-native principles adapted for generative AI workloads.

Author Biographies

  • Matteo Garcia, Associate Professor, Institute of Intelligent Systems, Western Europe Data Science University, Madrid, Spain

    Associate Professor, Institute of Intelligent Systems, Western Europe Data Science University, Madrid, Spain

  • Eva Klein, Assistant Professor, Institute of Intelligent Systems, European Institute of AI, Berlin, Germany

    Assistant Professor, Institute of Intelligent Systems, European Institute of AI, Berlin, Germany

Downloads

Published

2025-10-25

How to Cite

Cloud-Native Architectures for Deploying Generative Intelligence Systems. (2025). Journal of Generative Intelligence E: 3117-6429 P: 3117-6437, 2(4), 17-24. https://galaxiauniverse.com/index.php/JGI/article/view/315