Implementation detail: Capacity policy
The built-in collector reads celld’s typed /state and Kubernetes Metrics Server.
Incomplete observations are invalid, never zero demand. CPU, memory and readiness
do not establish PersistentFleet removal safety; a removed member keeps its disk
(PersistentFleet lifecycle). Manual and policy requests
share one path: growth in one step, contraction one member per step. A
PersistentFleet grows one run of kept or fresh disks at a time
(where a fresh disk goes); a
policy addition that spans both is cut at the end of its first run and
recorded as the step taken.
| Mode | Replica ownership |
|---|---|
| Omitted | Manual spec.replicas. |
| Shadow | Observe and report recommendations only. |
| ScaleOut | Apply stable bounded additions. |
| Automatic | Request additions and contractions. |
| External | One external /scale writer owns spec.replicas; no built-in demand collection. |
Contraction removes one member per step after the previous change has rolled
out. PersistentFleet removes the highest ordinal and keeps its disk. Automatic
and External steps require survivor capacity for every possible victim, else
CapacityUncertain. Test records describe the
capacity and lifecycle scenarios exercised.
Field under capacity |
Default | Meaning |
|---|---|---|
| mode | Shadow | Shadow, ScaleOut (explicit additions), Automatic (requests both directions), External (one /scale writer owns spec.replicas) |
| minReplicas / maxReplicas | 3 / 10 | Automatic bounds; minimum must cover the explicit AZ count |
| scaleOutStep | 1 | At most this many additions per completed stable decision; 1–10 |
| sampleIntervalSeconds | 15 | Minimum interval between counted observations |
| maxAgeSeconds | 45 | Maximum source/receipt age and gap between observations |
| minWindowSeconds / maxWindowSeconds | 5 / 60 | Allowed Metrics Server CPU averaging window |
| minSamples | 3 | Distinct advancing observations required in each stable window |
| scaleOutStabilizationSeconds | 30 | Continuous high-demand window |
| scaleInStabilizationSeconds | 600 | Continuous low-demand window |
| scaleOutCooldownSeconds | 300 | Minimum time from last durable action to another addition |
| scaleInCooldownSeconds | 900 | Minimum time from last durable action to a removal request |
| provisioningTimeoutSeconds | 600 | Deadline for useful capacity, after which additions remain blocked |
| redistributionObservationSeconds | 120 | Continuous complete observation window for judging an addition, 30–3600 seconds; evidence still mixed after twice this counts as ineffective |
| cpuHighMillicores / cpuLowMillicores | 200 / 80 | Absolute per-container CPU thresholds |
| memoryHighMiB / memoryLowMiB | 768 / 384 | Absolute per-container memory thresholds |
These defaults are starting points, not production capacity guarantees. Source samples must advance and fit configured age/window limits. Built-in contraction requires complete fresh low-demand observations and ready survivors.
spec.replicas is never rewritten by built-in automation. A manual edit takes
precedence for the next operation. Policy mode and bounds do not retarget issued
work. Removing a policy returns toward the manual target. Shadow pauses automatic
actions at the applied count.
Current capacity state retains stabilization timestamps, cooldown and bounded
redistribution observations. PendingCapacity and IneffectiveCapacity identify
additions not yet useful. IncompleteMetrics and RepeatedSamples reset stable
windows. ObservingRedistribution waits for newcomer activity or measured relief;
LoadNotRedistributed holds repeated ineffective additions. An idle ready Pod
alone is not evidence of useful redistribution. Explicit minimum capacity and
manual replica requests retain their documented precedence.
An addition is judged against the members it grew from, once one outcome,
effective or ineffective, holds for a whole redistributionObservationSeconds
window of at least minSamples observations. Evidence that has not settled
after twice that window and twice minSamples of continuous complete
observation is mixed: the addition counts as ineffective, never as relief.
Anything that resets stable windows restarts that count.
If one of those members restarts or is replaced, or the fleet contracts below
the size the addition reached, before the addition is judged, it can no longer
be judged and is dropped. It counts as neither relief nor an ineffective
addition. The next addition still waits for a new stable window, the cooldown
and ready capacity. If two additions have been judged ineffective since the
last effective one, the next reports LoadNotRedistributed while it is
observed, and the hold resumes if it is ineffective too.
Maintenance pause stops new workload changes and invalidates positive demand windows. A membership change, such as a restarted member, restarts stable windows. Controller restart or status clearing does not reset policy state. Holds are released by later observations. Never edit reservation annotations to remove a policy hold.
External autoscalers target CelldFleet using its /scale subresource, never the
managed workload. The subresource reports nonterminal observed Pods, including
terminating ones, with the fleet’s exact label selector. A desired target that
cannot pass lifecycle checks remains visible as desired versus applied count.
See the external sample.
Experimental software for evaluation.Capabilities and limitations· Contribute