PersistentFleet lifecycle
PersistentFleet runs CELLD_DURABILITY=fleet on a StatefulSet. Each member
keeps its disk for the life of the fleet. The StatefulSet controller performs
every restart and scaling step, and celld recovers every member. The operator
renders and applies the objects and replaces a member that cannot come back on
its own, so the fleet heals without an administrator. It keeps no lifecycle
state of its own. Bucket fleets roll the same way through their Deployment or
StatefulSet; their disks are caches.
What celld does
Section titled “What celld does”- On SIGTERM, a member hands its cells to the other members.
- On start, a member first serves its follower data to its peers. It then recovers its previous session and takes its lease. Several members, or all of them, can restart at once.
- The readiness probe is celld’s health endpoint. It returns 200 only after that recovery. A new member also holds its first 200 until the fleet has absorbed the change: no other member is handing off cells, and every live member has memory headroom.
- Leaders stop using a departed member on their own. Until their ensembles re-form, usually within seconds, they acknowledge writes through the bucket.
Changes
Section titled “Changes”| Change | Behavior |
|---|---|
| Restart or upgrade | A new maintenance.restartToken or runtimeImage changes the Pod template. The StatefulSet uses RollingUpdate: Kubernetes restarts one member at a time, highest ordinal first, and waits for each to be Ready. Each member reattaches its disk. An operator release that changes the template rolls the same way. |
| Scale-in | One member per step, the highest ordinal, and only after the previous change has rolled out. The removed member keeps its PVC. Automatic and External contraction also require survivor-capacity evidence (CapacityUncertain). |
| Growth | Applied one run of ordinals at a time, so kept disks are placed before fresh ones; see where a fresh disk goes. An ordinal that had a member before reattaches its kept PVC, and celld treats that as a restart. New ordinals get new claims. |
| Node drain | The PodDisruptionBudget allows maxUnavailable: 1. A member that is not Ready counts against it, so drains proceed one member at a time. |
| Pause | maintenance.paused stops the operator from writing workload changes and from replacing members. A rollout the StatefulSet controller has already started continues. |
| Full stop | A runtime pair that cannot run together, such as 0.5.1-ewhauser to 0.6.0-ewhauser, is upgraded by pausing the fleet, scaling its StatefulSet to zero, then setting the new image and resuming in one change. Every member returns on its own disk at the declared count, whatever the capacity policy. See upgrade the runtime. |
| Deletion | The StatefulSet is deleted in the foreground and members drain on SIGTERM. The operator then deletes every fleet PVC, including those of removed members, and releases the finalizer. The bucket reservation is permanent. |
While the fleet exists, the operator deletes a Pod or a claim only to heal a
member; see self-healing. Before it creates or grows the
StatefulSet, it checks each member’s claim name: a claim without the fleet’s
UID label reports StorageIdentityConflict. A disk from another fleet is never
adopted.
Self-healing
Section titled “Self-healing”No administrator is expected to intervene. On every reconcile the operator reads each member’s Pod and claim, takes at most one of these actions, and records nothing:
| Condition | Action | Event |
|---|---|---|
| A member Pod is still present more than two minutes after its termination grace ended, because its node no longer answers | Force-delete the Pod. The StatefulSet recreates it and the member keeps its disk. Also applies to Ordered Bucket fleets. | MemberForceDeleted |
A member’s claim is Lost, because its volume no longer exists |
Delete the claim and the Pod. The StatefulSet recreates both. | MemberDiskLost |
| One member has been down for the replacement delay while every other member has been ready for five minutes | Delete the claim and the Pod. The member returns on a fresh disk. | MemberReplaced |
celld’s lease fences a process that survives on a lost node, and a
ReadWriteOncePod disk attaches to one node at a time.
The replacement delay defaults to 10 minutes. Set it with the chart value
memberReplacementDelay or the operator flag --member-replacement-delay; it
must be at least one minute. The delay is patience, not the safety condition:
it outlasts the roughly six minutes Kubernetes takes to force-detach a volume
from a lost node, and typical node provisioning, so a member that can return
usually does so on its own disk. The delay covers a disk stranded in an unavailable
zone, a volume that no longer attaches, and a corrupt disk that keeps celld
from starting. While a member waits for it, the fleet reports Provisioning
with the time the member will be replaced. Two cases are left alone because a
new disk would not help: a member waiting on its image or configuration, and a
member already on a disk created for its current Pod. A claim that was only
deleted, while its volume still existed, is recreated by the StatefulSet as
soon as the member’s next Pod is created.
Leaders stop using a departed member within seconds, and every member sweeps dead leaders every 30 seconds. Once the rest of the fleet has been ready for five minutes, no session depends on the down member’s disk. A member on a fresh disk answers recovery for its old disk with a conclusive “no fragment”, so celld seals any session whose only complete copy was on the old disk and records a bounded loss. With the rest of the fleet ready, no such session remains unless a second failure happened first.
Lost disks describes what you see while this happens.
Where a fresh disk goes
Section titled “Where a fresh disk goes”A member’s disk pins it to the zone where the disk was created, and the zone spread counts only Pods that are on nodes. A member whose Pod is pending leaves its zone looking empty, so a fresh disk scheduled at that moment can take the zone that member’s disk needs. The pending member then has no zone it may run in. The operator therefore gives a member a fresh disk only when:
- the StatefulSet already runs the operator’s current template and replica count and has observed them, so no rollout is about to recreate another member; and
- the scheduler has placed every other member’s Pod on a node, or found no node for it.
The fresh disk then goes to the zone the fleet is missing. A lost volume is usually replaced within seconds. During an operator upgrade, a member that an earlier release left down is restarted by the rollout on its own disk before the operator considers replacing it.
Growth follows the same rule. The StatefulSet creates every new Pod at once,
so the operator grows one run of ordinals at a time: consecutive ordinals that
all kept their claims, or consecutive ordinals that all get new ones. Growing
back to removed members reattaches their disks first. Each run waits until the
scheduler has placed, or found no node for, every current member’s Pod, and
the fleet reports Provisioning while it waits:
Waiting for the scheduler to place member my-fleet-2, or find no node for it, before adding members my-fleet-3 to my-fleet-4 on fresh disksA manual or External target is reached over as many runs as it takes. A capacity-policy addition that spans kept and fresh ordinals is cut at the end of its first run, and the policy adds more after its next stable window if pressure remains.
A member whose Pod the scheduler cannot place is named in the fleet’s status
with the scheduler’s reason. If it is the one member down, it is replaced after
the replacement delay like any other. If the operator does not replace it,
because another member is also down or because it has no disk yet, the fleet
reports Blocked with reason MemberUnschedulable once the member has gone
unscheduled for the replacement delay. See
scheduling problems.
Limits
Section titled “Limits”- Two members that cannot come back wait. When two members are down at
once, neither is replaced until one returns, unless its volume is lost.
Replacing either could lose writes that only their disks hold. celld’s
guarantee covers the loss of one member. If one of them cannot be scheduled,
the fleet reports
MemberUnschedulableafter the replacement delay. - Hand-deleted claims do not wait for placement. The StatefulSet gives a claim you delete a fresh disk as soon as it creates the member’s Pod. If that disk takes the zone a pending member’s disk needs, the pending member is replaced after the replacement delay, once the rest of the fleet is ready.
- Growth is paced by scheduling. Each run of new members waits until the scheduler has decided every current member’s Pod. A member whose Pod never reaches the scheduler holds growth back; one the scheduler cannot place does not.
- Pacing is readiness. Kubernetes does not wait for celld to finish recovering other sessions between restarts. With disks kept, a restart destroys nothing.
- A rollout waits only for the member it restarted. It does not wait for
another member that is already down, so two members can be down at once.
Both keep their disks, so this costs availability, not acknowledged writes.
Kubernetes’
MaxUnavailableStatefulSetfeature gate would count every unavailable member, but it is off by default. - Removed members’ disks cost storage. They stay until growth reattaches them or the fleet is deleted. Delete one by hand only if you don’t plan to grow back to that ordinal. A claim deleted before the leaders have re-formed their ensembles can still hold the only copy of recent writes.
- A hand edit to the StatefulSet template starts rolling at once. The operator restores its template on the next reconcile, and the member that started rolling returns to that template on its own disk.
What the operator stores
Section titled “What the operator stores”Nothing about the lifecycle. The storage reservation’s
celld.eric.dev/current-operation annotation holds only capacity-policy state,
and is absent unless the fleet has set spec.capacity. status.lifecycle is
always empty. Editing status authorizes nothing.
Fleets from earlier releases
Section titled “Fleets from earlier releases”For a fleet created by an earlier release, the operator switches the
StatefulSet from OnDelete to RollingUpdate and writes the current template.
Kubernetes then replaces each earlier Pod, highest ordinal first, including
Pods held by the launcher scheduling gate. Claims keep their names and
identities. An in-flight strict operation is dropped with an
OperationSuperseded event. See upgrade the operator.
The lifecycle contract and retained disks describe the implementation.
Experimental software for evaluation.Capabilities and limitations· Contribute