Multi-Cloud Kubernetes Consolidation, Observability and FinOps
Challenge
Migrating live workloads without downtime and changing team habits around resource requests. We used phased namespace migrations, automated request recommendations and SLOs agreed with each product team.
Approach
We standardised on a cluster blueprint (Terraform + Cluster API) for AWS and Azure, with Karpenter/cluster-autoscaler, shared ingress, cert management and secrets via External Secrets and a vault. Observability uses OpenTelemetry, Prometheus, Loki and Tempo with Grafana dashboards and SLO-based alerts. Kubecost-style cost allocation by team and namespace feeds monthly showback reports. Right-sizing and spot/preemptible node pools were introduced for stateless workloads.
Outcome
Illustratively ~25–40% lower compute spend through right-sizing and autoscaling Fewer, more actionable alerts tied to SLOs Team-level cost showback driving better decisions One blueprint for new clusters across clouds Faster incident triage with correlated logs, metrics and traces