Skip to content

Configure capacity policy

Capacity policy reads each celld Pod’s private /state endpoint and Kubernetes Metrics Server. Install Metrics Server and permit operator access to port 8081 before using it. You provision nodes, IAM roles, the CSI driver, and Metrics Server. Adding PersistentFleet replicas requests new PVCs through the configured StorageClass, except where an ordinal reattaches the disk it kept after an earlier scale-in. Start with Shadow to see recommendations without changing the workload.

On my-fleet in namespace fleets, apply the shadow sample after adapting its bounds to your zones. A minimal patch is:

Terminal window
kubectl --context YOUR_CONTEXT -n fleets patch celldfleet my-fleet \
--type merge -p '{"spec":{"capacity":{"mode":"Shadow","minReplicas":3,"maxReplicas":10}}}'
kubectl --context YOUR_CONTEXT -n fleets get celldfleet my-fleet \
-o json | jq '.status.capacity'

Set minReplicas at least as high as placement.azCount. Defaults and tuning fields are in the capacity policy fields. A Shadow desired count is advice only. Confirm that status.capacity has fresh, complete observations from every expected replica, and compare the recommendation with actual workload demand before enabling actions.

Once the recommendation is useful, change only the mode:

Terminal window
kubectl --context YOUR_CONTEXT -n fleets patch celldfleet my-fleet \
--type merge -p '{"spec":{"capacity":{"mode":"ScaleOut"}}}'
kubectl --context YOUR_CONTEXT -n fleets get celldfleet my-fleet \
-o json | jq '.status.capacity'

ScaleOut can add bounded capacity after a stable high-demand window. spec.replicas remains your manual target; the applied count can exceed it. Automatic also requests reductions. Requests use the same path as manual scaling: reductions remove one member per step, each after the previous change has rolled out, and automatic reductions are also gated on survivor capacity for every possible victim. External instead assigns spec.replicas to one /scale writer and disables built-in demand collection.

PendingCapacity means a requested addition has not become useful. IneffectiveCapacity means it passed its provisioning deadline; inspect Pods, PVCs, and node capacity. IncompleteMetrics and RepeatedSamples reset stabilization. ObservingRedistribution waits to see whether additions relieve incumbents; LoadNotRedistributed holds further pressure-driven batches. The operator does not count an idle new Pod as useful redistribution. An addition whose effect stays mixed, such as a new Pod whose load keeps crossing cpuLowMillicores, counts as ineffective after twice redistributionObservationSeconds. Neither hold needs you to clear it. If a member restarts or the fleet contracts before an addition is judged, the operator drops that addition without counting it either way, and the next addition still waits for a new stable window and the cooldown. These reasons are explained in conditions and capacity troubleshooting.

Pause workload changes with spec.maintenance.paused: true. Removing the policy returns toward the manual replica target and can request a blocked reduction. Check the resulting target and scaling procedure first.

Experimental software for evaluation.Capabilities and limitations· Contribute