Uber has published a detailed account of its new ServiceScale controller, which allows multiple orchestrators to safely manage the scaling of the same Kubernetes workloads. The blog post, written by senior software engineers Egor Grishechko and Srikar Paruchuru, describes how the company separated scaling intent from execution to support regional failover without carrying reserved idle capacity.
Uber's Container Platform team manages over 100 compute clusters across data centres and cloud providers including Oracle and Google, running roughly 4,000 services on 3 million cores with 1.5 million daily pod launches. The team's internal platform, called Up, acts as a federation layer for the Kubernetes fleet. Service owners use Up to deploy builds and set scaling expectations, and a dedicated controller, the Uber Deployment Controller (UDC), reconciles that intent into Kubernetes primitives. InfoQ previously covered Uber's migration to Up and the subsequent completion of its Kubernetes migration.

The motivation came from a change to how Uber handles regional failovers. Uber runs active-active data centres across different regions. When an outage occurs, traffic is rerouted to a surviving region, which needs enough idle compute to handle the increased load. Historically, Uber kept reserved idle capacity in all data centres. Engineers wanted to reuse capacity from low-tier workloads instead, scaling them down and scaling up high-tier workloads during a failover.
This created a new source of scaling intent. Up and UDC still needed to own the normal desired state of services, but a failover orchestrator now needed to influence scaling decisions as well. Engineers considered extending UDC with failover logic but decided against it. UDC already sat on the hot path for service lifecycle operations, and adding failover-specific behaviour would increase the complexity of a controller that powered the most critical workflows. Grishechko and Paruchuru note that "a regression in failover handling wouldn't stay isolated to failover" and could affect normal deployments across the fleet.
Instead, the team introduced a new custom resource definition called ServiceScale and a new Service Scale Controller (SSC). Each orchestrator can express its own scaling desire through ServiceScale, and SSC reconciles the combined intent into Kubernetes objects. Grishechko and Paruchuru explain that they deliberately kept the model simple.
"We didn't want an additional external database, a separate coordination service, or a control plane that'd become harder to debug under incident pressure."
Egor Grishechko and Srikar Paruchuru
The materialisation of scale intent directly in Kubernetes made the system easier to inspect. When something looked wrong, engineers could check ServiceScale to see which orchestrator wanted what. Failback also became simpler because both steady-state and temporary failover information were saved in the CRD spec, meaning recovery did not require reconstructing state from logs.
The blog post describes three production lessons. Stale informer caches presented the first problem. Like most Kubernetes controllers, Uber's consume resources through informer caches that can lag reality by a few seconds. Up treated a status field as a terminal input, where a success signal from UDC could trigger the next irreversible step in a workflow. Engineers implemented a read-your-own-write consistency guardrail: when a controller updates downstream resources, it attaches its current generation as an annotation, then verifies that its cached data reflects at least that generation before reporting status.
This problem is not unique to Uber. Kubernetes v1.36, released in April 2026, introduced staleness mitigation for controllers using a comparable approach. If the cache's latest resource version is lower than what the controller has written, the controller does not take action. The Kubernetes project is also working with controller-runtime to bring read-your-own-write semantics to all controllers built with that tooling.
Writing on LinkedIn, software engineer Prasad M K described the read-your-own-writes gap as "an API contract problem, not a backend cache problem" and recommended version tokens to validate reads against writes. His analysis frames the issue at the application level, while Uber's solution operates within the Kubernetes controller pattern.
Multi-writer systems presented another challenge. When UDC and SSC began simultaneously updating the same Kubernetes resource, specific timing conditions caused the ReplicaSet to become inconsistent. Its metadata drifted from its spec, which broke proportional scaling for rolling updates and sometimes caused workloads to become stuck. Engineers added fleet-wide observability to detect metadata-spec drift, built an automated healer in UDC to patch affected ReplicaSets, and pursued a long-term fix in the scaling path. An academic paper on Uber's failover architecture, published on arXiv in January 2026, reports that the broader Unified Failover Architecture reduced steady-state provisioning from 2x to 1.3x and eliminated over one million CPU cores.
The rollout took a year. Engineers used staging environments and canary deployments, and invested in integration testing with the kind testing tool to simulate real controller interactions and catch race conditions. The rollout supported native Kubernetes Deployments and OpenKruise CloneSets, and completed without customer-impacting outages. InfoQ also covered Uber's continuous deployment optimisation, which shares the same emphasis on gradual, safe rollouts.
"Multi-orchestrator systems aren't hard because of the APIs. They're hard because of everything that happens between writes."
Egor Grishechko and Srikar Paruchuru
Uber's experience highlights the distributed systems problems that multi-orchestrator Kubernetes introduces in production. The CNCF recently graduated Karmada, a multi-cluster orchestration project that addresses failover across clusters through a different architecture. The Kubernetes project's own staleness mitigation in v1.36 suggests these problems are gaining recognition at the platform level.