Uber streamlines Kubernetes scaling with new failover system

Uber has launched a new framework to streamline how its Kubernetes clusters handle scaling decisions, separating the decision-making process from its execution. The company’s ServiceScale controller now enables multiple orchestrators to work together on scaling the same workloads without requiring idle capacity to be held in every region. This approach supports active-active data centers by allowing Uber to reroute traffic during disruptions without over-allocating resources across its infrastructure.
This change follows an internal review by Uber’s Container Platform team, which found inefficiencies in the company’s previous failover strategy. Before the update, Uber maintained reserved capacity across all regions to handle unexpected traffic surges, which led to wasted resources. The new system instead scales down lower-priority workloads and ramps up critical ones during failovers, reusing available capacity more efficiently.
Uber’s internal platform, Up, already manages a Kubernetes fleet of over 100 clusters spread across data centers and cloud providers such as Oracle and Google. The platform supports roughly 4,000 services running on 3 million CPU cores and launches 1.5 million pods daily. Service owners define scaling requirements through Up, while the Uber Deployment Controller (UDC) converts these into Kubernetes-native configurations. The newly introduced ServiceScale Controller (SSC) adds an additional layer, allowing failover orchestrators to influence scaling choices without altering the core functions of UDC.
Read Also: Building adaptive recommendation systems
Why Uber split scaling decisions from execution
Initially, Uber’s engineers explored adding failover logic directly into UDC. However, they determined that embedding failover-specific behavior into a controller responsible for critical workflows would increase complexity. A failure in failover handling could disrupt normal deployments across the entire fleet. As an alternative, they developed a custom ServiceScale resource definition and a dedicated controller to reconcile scaling instructions from multiple sources.
The architecture avoids typical challenges like external databases or separate coordination services. By embedding scale intent directly into Kubernetes, engineers gain clear visibility into which orchestrator is driving each scaling decision. The system also simplifies failback procedures, as both steady-state and temporary scaling configurations are stored within the resource definition. This eliminates the need to reconstruct logs during recovery operations.
During testing, three key production issues emerged. The first involved stale informer caches, a recurring problem in Kubernetes controllers. Uber’s controllers rely on caches that can delay updates by seconds. If a status update appeared successful but the cache had not refreshed, it could trigger irreversible workflow steps. To address this, the team implemented a read-your-own-write consistency guardrail, where controllers verify that cached data aligns with their latest updates before proceeding. This method aligns with a fix introduced in Kubernetes v1.36, released in April 2026, which included staleness mitigation for controllers using similar patterns.
Read Also: Vercel Labs releases native TypeScript compiler
Fixing race conditions and stale data flaws
A second challenge arose from multi-writer systems. When UDC and SSC updated the same Kubernetes resource simultaneously, race conditions caused ReplicaSet metadata to diverge from its intended configuration. This disrupted proportional scaling during rolling updates and sometimes left workloads in an unstable state. Uber responded by adding fleet-wide monitoring to detect metadata-spec mismatches, integrating an automated repair mechanism in UDC to correct affected ReplicaSets, and pursuing a long-term solution in the scaling pipeline. Research published in January 2026 on Uber’s Unified Failover Architecture in arXiv showed the system reduced steady-state provisioning from 2x to 1.3x and removed over 1 million CPU cores from reserved capacity.
Uber’s academic paper on the Unified Failover Architecture is available on arXiv.
