Kubernetes has definitively won the container orchestration war. However, the marketing pitch—"just use EKS/AKS and you don't have to worry about the control plane"—dangerously understates the operational complexity of running K8s in production.
The Day 2 Operations Nightmare
Getting a cluster up is easy. Day 2 operations—upgrades, security, ingress, and observability—are where engineering teams drown. Here is our prescriptive stack for taming enterprise Kubernetes.
1. GitOps is Mandatory (Flux / ArgoCD)
If your developers are running kubectl apply from their laptops, you are doing it wrong. We enforce strict GitOps. The entire cluster state is declared in a Git repository. Flux or ArgoCD continuously monitors Git and synchronizes the cluster state. This provides an immutable audit trail and instant disaster recovery.
2. Policy as Code (Kyverno)
You cannot trust 500 developers to write secure YAML. We deploy Kyverno to enforce cluster-wide policies:
- Reject any pod requesting
privileged: true. - Mandate that every pod has CPU and Memory
requestsandlimitsdefined. - Ensure images are only pulled from the trusted internal Azure Container Registry (ACR).
3. Observability and Cost
K8s abstracts compute, which means developers will consume infinite resources if allowed. We deploy KubeCost to map namespace resource usage back to specific business units, providing showback reports. For telemetry, we standardize on Prometheus for metrics and Promtail/Loki for log aggregation.