Asia/Kolkata
Projects

Building a Cloud-Native AI/ML Platform on AWS EKS with Kubernetes, GitOps, and FinOps

Architected a production AI/ML platform on AWS EKS serving 10+ data science teams at 99.9% availability, and delivered a 45% reduction in cloud infrastructure costs through FinOps practices.
image
On this page
June 1, 2024
Modern machine learning teams move fast. They need the ability to run experiments, train models, and deploy inference services without waiting for infrastructure. Traditional VM-based environments often struggle to provide the flexibility and scalability required for modern AI workloads. To solve this challenge, we designed and built a cloud-native AI/ML platform on AWS EKS using Kubernetes, enabling 10+ data science teams to run GPU-accelerated training workloads and scalable inference services in a self-service environment. The platform provides:
  • Kubernetes-based orchestration for AI workloads
  • Dynamic node provisioning using Karpenter
  • GitOps-driven infrastructure management
  • Cost visibility through FinOps practices
  • High reliability with a 99.9% availability SLA
By adopting a Kubernetes-based platform engineering approach, the infrastructure team was able to provide a standardized and scalable environment for machine learning development while maintaining strong cost control.
Before the platform was introduced, data science teams relied heavily on manually provisioned EC2 instances for training models and running experiments. This approach created several operational and cost challenges. First, infrastructure provisioning was inconsistent. Each team configured their own environments, which led to configuration drift and unpredictable performance. Second, GPU resources were often over-provisioned. Teams would launch large instances to ensure training jobs could run, but many of these instances remained underutilized. Third, there was no reliable isolation between teams. A heavy training workload from one team could consume the majority of cluster resources, affecting other teams' workloads. Finally, cost visibility was limited. Without proper monitoring and cost attribution, it was difficult to understand which workloads were driving cloud spending. The goal was clear: create a cloud-native AI/ML platform that provides scalable infrastructure, strong workload isolation, and efficient resource utilization.
The platform is built on Amazon EKS, which provides a managed Kubernetes control plane while allowing full flexibility for workload orchestration. The architecture separates compute capacity into different node pools based on workload characteristics.
The Kubernetes cluster uses three primary node pools: CPU Node Pools These pools run general workloads such as inference services, experiment tracking systems, and lightweight data processing tasks. They are dynamically managed by Karpenter, which provisions instances based on real-time scheduling requirements. GPU Node Pools GPU nodes are dedicated to model training workloads. These nodes are configured with NVIDIA GPUs and primarily run on spot instances to reduce cost. To maintain reliability, the platform uses on-demand fallback instances when spot capacity becomes unavailable. System Node Pool A small set of fixed nodes runs platform services such as:
  • Argo CD
  • Prometheus
  • Grafana
  • cluster networking components
Separating system workloads ensures that platform services remain stable even when training workloads spike.
One of the most important architectural decisions was choosing Karpenter instead of the traditional Kubernetes Cluster Autoscaler. Cluster Autoscaler relies on predefined node groups, which often leads to inefficient resource usage. Workloads must fit into existing instance sizes, forcing teams to over-provision capacity. Karpenter takes a different approach. It provisions nodes dynamically based on the exact resource requirements of pending pods. This provides several benefits:
  • faster pod scheduling
  • better instance selection
  • reduced infrastructure waste
  • improved GPU utilization
Training workloads that previously required large on-demand instances can now run on right-sized spot instances, dramatically reducing infrastructure cost. When GPU workloads finish, unused nodes automatically terminate, allowing the cluster to scale down to zero GPU nodes when idle.
To ensure fairness and stability across teams, the platform implements namespace-based multi-tenancy. Each data science team receives its own Kubernetes namespace with:
  • Resource quotas
  • Network policies
  • RBAC permissions
  • isolated service accounts
This approach prevents a single team's training workload from exhausting cluster resources. If a training job attempts to exceed its resource allocation, Kubernetes enforces the defined limits, protecting the rest of the platform. This structure also makes it easier to implement cost attribution and monitoring per team.
Machine learning workloads often produce large model checkpoints and require access to shared datasets. To support this requirement, the platform uses Amazon EFS integrated with Kubernetes through the EFS CSI driver. EFS provides several advantages for AI workloads:
  • shared filesystem access across pods
  • persistent training checkpoints
  • simplified dataset management
  • ReadWriteMany access mode
When a training job is interrupted — for example due to a spot instance termination — the checkpoint remains safely stored on EFS. The training job can resume without losing progress.
To manage the growing complexity of the platform, all infrastructure and Kubernetes configurations follow a GitOps workflow using Argo CD. GitOps ensures that the cluster state always matches the desired configuration stored in Git repositories. Key benefits include:
  • consistent environment configuration
  • automated deployments
  • improved auditability
  • safer rollbacks
Platform components such as monitoring stacks, networking configurations, and namespace policies are all defined declaratively and deployed through Argo CD. This approach allows the platform team to maintain reliability while enabling faster iteration.
Operating an AI/ML platform at scale requires deep visibility into system performance and infrastructure cost. The platform uses a Prometheus and Thanos based monitoring architecture with Grafana dashboards providing real-time insights. Metrics include:
  • GPU utilization per team
  • training job queue depth
  • node provisioning latency
  • spot interruption frequency
  • inference service performance
To support FinOps practices, cost attribution metrics are exposed at the namespace level. Teams can see exactly how much compute their workloads consume. This transparency encourages teams to optimize their workloads and avoid unnecessary resource usage.
After deploying the platform, the organization saw significant improvements in both developer productivity and infrastructure efficiency. Key results include:
  • 99.9% platform availability maintained over a 12-month period
  • 45% reduction in GPU compute costs through spot instance strategy and right-sized provisioning
  • ~80% faster onboarding for new data science team members
  • consistent infrastructure environments across teams
  • zero data loss incidents despite regular spot interruptions
Most importantly, data science teams were able to focus on building and deploying machine learning models instead of managing infrastructure.
Building a Kubernetes platform for AI workloads revealed several key lessons. First, dynamic node provisioning is essential for GPU workloads. Static node groups quickly lead to resource waste. Second, cost visibility must be built into the platform from the start. Without FinOps practices, GPU infrastructure costs can grow rapidly. Third, GitOps significantly improves operational stability. Managing cluster configuration through Git creates a clear and auditable deployment process. Finally, platform engineering plays a critical role in enabling data science productivity. By abstracting infrastructure complexity, teams can focus on experimentation and model development.
Building a cloud-native AI/ML platform on AWS EKS using Kubernetes enabled the organization to scale machine learning workloads while maintaining reliability and cost efficiency. By combining Kubernetes orchestration, Karpenter autoscaling, GitOps automation, and FinOps cost visibility, the platform now supports GPU-accelerated workloads for multiple data science teams without operational overhead. This architecture demonstrates how platform engineering principles can transform machine learning infrastructure, providing both scalability and developer productivity. As AI adoption continues to grow, cloud-native platforms like this will play an increasingly important role in enabling organizations to run machine learning workloads at scale.
Technologies: AWS EKS · Kubernetes · Karpenter · Terraform · Helm · Argo CD · Prometheus · Thanos · Grafana · Amazon EFS · NVIDIA Device Plugin · External Secrets Operator