Overview
The Problem
- manual kubectl apply commands
- Helm releases triggered from CI pipelines
- direct cluster modifications during troubleshooting
- Inconsistent deployment workflows across teams
- Configuration drift between environments
- Limited visibility into the actual state of clusters
- Difficult rollbacks during production incidents
- Slow incident response due to uncertainty about deployed versions
The GitOps Model
- Desired cluster configuration is stored in Git.
- Argo CD continuously monitors the repository.
- The controller reconciles the Kubernetes cluster state with the declared configuration.
- Any configuration drift is automatically corrected or flagged for review.
- reliable and repeatable deployments
- improved traceability of changes
- automatic drift detection
- simplified rollbacks through Git history
Repository Structure
gitops-platform/
├── charts/ # Packaged application templates (Helm)
│ └── core-services/
├── values/ # Environment-specific configurations
│ ├── dev.yaml # Development values
│ ├── staging.yaml # Staging values
│ └── production.yaml # Production values
└── clusters/ # Argo CD Root Application definitions
The charts/ directory contains reusable Helm charts that define the desired state of our applications.
Environment-specific overrides are managed through Helm values files located in the values/ directory. This allows us to maintain a single chart while varying configurations like replica counts, resource limits, and ingress hosts across different clusters.
By integrating Helm directly with Argo CD, the platform can automatically render manifests and detect drift based on the combination of the base chart and the specific environment's values file.
This approach simplifies the management of complex deployments while ensuring that environment differences are transparent and version-controlled.
Multi-Cluster GitOps Architecture
Management Cluster
- Argo CD
- monitoring stack
- shared infrastructure components
Development Cluster
- Automatic synchronization with Git
- Deployments triggered on merges to the main branch
- Fast feedback cycles for developers
Staging Cluster
Production Cluster
Observability and Deployment Metrics
- synchronization success and failure rates
- drift detection events
- deployment frequency per team
- lead time from commit to production
- rollback frequency
Secret Management Strategy
- Secrets are created and managed in AWS Secrets Manager.
- ExternalSecret resources are defined in Git and synced via Argo CD.
- The ESO controller fetches the sensitive data from AWS and dynamically creates native Kubernetes Secrets.
Yaml
apiVersion: external-secrets.io/v1beta1
kind: ExternalSecret
metadata:
name: database-credentials
spec:
refreshInterval: 1h
secretStoreRef:
name: aws-secrets-manager
kind: ClusterSecretStore
target:
name: db-secret
data:
- secretKey: password
remoteRef:
key: prod/database/credentials
property: password
Results and Impact
- ~80% reduction in deployment errors
- ~60% faster incident detection through automated drift alerts
- 100% audit trail for infrastructure and application changes via Git history
- Deployment frequency increased from weekly releases to multiple deployments per day
- Mean time to rollback reduced from 30+ minutes to under 5 minutes
Conclusion
- consistent deployment workflows
- automated drift detection
- faster recovery during incidents
- improved visibility into system changes
Technologies: Argo CD · Kubernetes · Helm · External Secrets Operator · Prometheus · Grafana · GitHub Actions · AWS Secrets Manager · AWS EKS
