Skip to content

Implementation detail: Capacity policy

The built-in collector reads celld’s typed /state and Kubernetes Metrics Server. Incomplete observations are invalid, never zero demand. CPU, memory and readiness do not establish PersistentFleet removal safety; a removed member keeps its disk (PersistentFleet lifecycle). Manual and policy requests share one path: growth in one step, contraction one member per step. A PersistentFleet grows one run of kept or fresh disks at a time (where a fresh disk goes); a policy addition that spans both is cut at the end of its first run and recorded as the step taken.

Mode Replica ownership
Omitted Manual spec.replicas.
Shadow Observe and report recommendations only.
ScaleOut Apply stable bounded additions.
Automatic Request additions and contractions.
External One external /scale writer owns spec.replicas; no built-in demand collection.

Contraction removes one member per step after the previous change has rolled out. PersistentFleet removes the highest ordinal and keeps its disk. Automatic and External steps require survivor capacity for every possible victim, else CapacityUncertain. Test records describe the capacity and lifecycle scenarios exercised.

Field under capacity Default Meaning
mode Shadow Shadow, ScaleOut (explicit additions), Automatic (requests both directions), External (one /scale writer owns spec.replicas)
minReplicas / maxReplicas 3 / 10 Automatic bounds; minimum must cover the explicit AZ count
scaleOutStep 1 At most this many additions per completed stable decision; 1–10
sampleIntervalSeconds 15 Minimum interval between counted observations
maxAgeSeconds 45 Maximum source/receipt age and gap between observations
minWindowSeconds / maxWindowSeconds 5 / 60 Allowed Metrics Server CPU averaging window
minSamples 3 Distinct advancing observations required in each stable window
scaleOutStabilizationSeconds 30 Continuous high-demand window
scaleInStabilizationSeconds 600 Continuous low-demand window
scaleOutCooldownSeconds 300 Minimum time from last durable action to another addition
scaleInCooldownSeconds 900 Minimum time from last durable action to a removal request
provisioningTimeoutSeconds 600 Deadline for useful capacity, after which additions remain blocked
redistributionObservationSeconds 120 Continuous complete observation window for judging an addition, 30–3600 seconds; evidence still mixed after twice this counts as ineffective
cpuHighMillicores / cpuLowMillicores 200 / 80 Absolute per-container CPU thresholds
memoryHighMiB / memoryLowMiB 768 / 384 Absolute per-container memory thresholds

These defaults are starting points, not production capacity guarantees. Source samples must advance and fit configured age/window limits. Built-in contraction requires complete fresh low-demand observations and ready survivors.

spec.replicas is never rewritten by built-in automation. A manual edit takes precedence for the next operation. Policy mode and bounds do not retarget issued work. Removing a policy returns toward the manual target. Shadow pauses automatic actions at the applied count.

Current capacity state retains stabilization timestamps, cooldown and bounded redistribution observations. PendingCapacity and IneffectiveCapacity identify additions not yet useful. IncompleteMetrics and RepeatedSamples reset stable windows. ObservingRedistribution waits for newcomer activity or measured relief; LoadNotRedistributed holds repeated ineffective additions. An idle ready Pod alone is not evidence of useful redistribution. Explicit minimum capacity and manual replica requests retain their documented precedence.

An addition is judged against the members it grew from, once one outcome, effective or ineffective, holds for a whole redistributionObservationSeconds window of at least minSamples observations. Evidence that has not settled after twice that window and twice minSamples of continuous complete observation is mixed: the addition counts as ineffective, never as relief. Anything that resets stable windows restarts that count.

If one of those members restarts or is replaced, or the fleet contracts below the size the addition reached, before the addition is judged, it can no longer be judged and is dropped. It counts as neither relief nor an ineffective addition. The next addition still waits for a new stable window, the cooldown and ready capacity. If two additions have been judged ineffective since the last effective one, the next reports LoadNotRedistributed while it is observed, and the hold resumes if it is ineffective too.

Maintenance pause stops new workload changes and invalidates positive demand windows. A membership change, such as a restarted member, restarts stable windows. Controller restart or status clearing does not reset policy state. Holds are released by later observations. Never edit reservation annotations to remove a policy hold.

External autoscalers target CelldFleet using its /scale subresource, never the managed workload. The subresource reports nonterminal observed Pods, including terminating ones, with the fleet’s exact label selector. A desired target that cannot pass lifecycle checks remains visible as desired versus applied count. See the external sample.

Experimental software for evaluation.Capabilities and limitations· Contribute