Skip to content

PersistentFleet lifecycle

PersistentFleet runs CELLD_DURABILITY=fleet on a StatefulSet. Each member keeps its disk for the life of the fleet. The StatefulSet controller performs every restart and scaling step, and celld recovers every member. The operator renders and applies the objects and replaces a member that cannot come back on its own, so the fleet heals without an administrator. It keeps no lifecycle state of its own. Bucket fleets roll the same way through their Deployment or StatefulSet; their disks are caches.

  • On SIGTERM, a member hands its cells to the other members.
  • On start, a member first serves its follower data to its peers. It then recovers its previous session and takes its lease. Several members, or all of them, can restart at once.
  • The readiness probe is celld’s health endpoint. It returns 200 only after that recovery. A new member also holds its first 200 until the fleet has absorbed the change: no other member is handing off cells, and every live member has memory headroom.
  • Leaders stop using a departed member on their own. Until their ensembles re-form, usually within seconds, they acknowledge writes through the bucket.
Change Behavior
Restart or upgrade A new maintenance.restartToken or runtimeImage changes the Pod template. The StatefulSet uses RollingUpdate: Kubernetes restarts one member at a time, highest ordinal first, and waits for each to be Ready. Each member reattaches its disk. An operator release that changes the template rolls the same way.
Scale-in One member per step, the highest ordinal, and only after the previous change has rolled out. The removed member keeps its PVC. Automatic and External contraction also require survivor-capacity evidence (CapacityUncertain).
Growth Applied one run of ordinals at a time, so kept disks are placed before fresh ones; see where a fresh disk goes. An ordinal that had a member before reattaches its kept PVC, and celld treats that as a restart. New ordinals get new claims.
Node drain The PodDisruptionBudget allows maxUnavailable: 1. A member that is not Ready counts against it, so drains proceed one member at a time.
Pause maintenance.paused stops the operator from writing workload changes and from replacing members. A rollout the StatefulSet controller has already started continues.
Full stop A runtime pair that cannot run together, such as 0.5.1-ewhauser to 0.6.0-ewhauser, is upgraded by pausing the fleet, scaling its StatefulSet to zero, then setting the new image and resuming in one change. Every member returns on its own disk at the declared count, whatever the capacity policy. See upgrade the runtime.
Deletion The StatefulSet is deleted in the foreground and members drain on SIGTERM. The operator then deletes every fleet PVC, including those of removed members, and releases the finalizer. The bucket reservation is permanent.

While the fleet exists, the operator deletes a Pod or a claim only to heal a member; see self-healing. Before it creates or grows the StatefulSet, it checks each member’s claim name: a claim without the fleet’s UID label reports StorageIdentityConflict. A disk from another fleet is never adopted.

No administrator is expected to intervene. On every reconcile the operator reads each member’s Pod and claim, takes at most one of these actions, and records nothing:

Condition Action Event
A member Pod is still present more than two minutes after its termination grace ended, because its node no longer answers Force-delete the Pod. The StatefulSet recreates it and the member keeps its disk. Also applies to Ordered Bucket fleets. MemberForceDeleted
A member’s claim is Lost, because its volume no longer exists Delete the claim and the Pod. The StatefulSet recreates both. MemberDiskLost
One member has been down for the replacement delay while every other member has been ready for five minutes Delete the claim and the Pod. The member returns on a fresh disk. MemberReplaced

celld’s lease fences a process that survives on a lost node, and a ReadWriteOncePod disk attaches to one node at a time.

The replacement delay defaults to 10 minutes. Set it with the chart value memberReplacementDelay or the operator flag --member-replacement-delay; it must be at least one minute. The delay is patience, not the safety condition: it outlasts the roughly six minutes Kubernetes takes to force-detach a volume from a lost node, and typical node provisioning, so a member that can return usually does so on its own disk. The delay covers a disk stranded in an unavailable zone, a volume that no longer attaches, and a corrupt disk that keeps celld from starting. While a member waits for it, the fleet reports Provisioning with the time the member will be replaced. Two cases are left alone because a new disk would not help: a member waiting on its image or configuration, and a member already on a disk created for its current Pod. A claim that was only deleted, while its volume still existed, is recreated by the StatefulSet as soon as the member’s next Pod is created.

Leaders stop using a departed member within seconds, and every member sweeps dead leaders every 30 seconds. Once the rest of the fleet has been ready for five minutes, no session depends on the down member’s disk. A member on a fresh disk answers recovery for its old disk with a conclusive “no fragment”, so celld seals any session whose only complete copy was on the old disk and records a bounded loss. With the rest of the fleet ready, no such session remains unless a second failure happened first.

Lost disks describes what you see while this happens.

A member’s disk pins it to the zone where the disk was created, and the zone spread counts only Pods that are on nodes. A member whose Pod is pending leaves its zone looking empty, so a fresh disk scheduled at that moment can take the zone that member’s disk needs. The pending member then has no zone it may run in. The operator therefore gives a member a fresh disk only when:

  • the StatefulSet already runs the operator’s current template and replica count and has observed them, so no rollout is about to recreate another member; and
  • the scheduler has placed every other member’s Pod on a node, or found no node for it.

The fresh disk then goes to the zone the fleet is missing. A lost volume is usually replaced within seconds. During an operator upgrade, a member that an earlier release left down is restarted by the rollout on its own disk before the operator considers replacing it.

Growth follows the same rule. The StatefulSet creates every new Pod at once, so the operator grows one run of ordinals at a time: consecutive ordinals that all kept their claims, or consecutive ordinals that all get new ones. Growing back to removed members reattaches their disks first. Each run waits until the scheduler has placed, or found no node for, every current member’s Pod, and the fleet reports Provisioning while it waits:

Waiting for the scheduler to place member my-fleet-2, or find no node for it, before adding members my-fleet-3 to my-fleet-4 on fresh disks

A manual or External target is reached over as many runs as it takes. A capacity-policy addition that spans kept and fresh ordinals is cut at the end of its first run, and the policy adds more after its next stable window if pressure remains.

A member whose Pod the scheduler cannot place is named in the fleet’s status with the scheduler’s reason. If it is the one member down, it is replaced after the replacement delay like any other. If the operator does not replace it, because another member is also down or because it has no disk yet, the fleet reports Blocked with reason MemberUnschedulable once the member has gone unscheduled for the replacement delay. See scheduling problems.

  • Two members that cannot come back wait. When two members are down at once, neither is replaced until one returns, unless its volume is lost. Replacing either could lose writes that only their disks hold. celld’s guarantee covers the loss of one member. If one of them cannot be scheduled, the fleet reports MemberUnschedulable after the replacement delay.
  • Hand-deleted claims do not wait for placement. The StatefulSet gives a claim you delete a fresh disk as soon as it creates the member’s Pod. If that disk takes the zone a pending member’s disk needs, the pending member is replaced after the replacement delay, once the rest of the fleet is ready.
  • Growth is paced by scheduling. Each run of new members waits until the scheduler has decided every current member’s Pod. A member whose Pod never reaches the scheduler holds growth back; one the scheduler cannot place does not.
  • Pacing is readiness. Kubernetes does not wait for celld to finish recovering other sessions between restarts. With disks kept, a restart destroys nothing.
  • A rollout waits only for the member it restarted. It does not wait for another member that is already down, so two members can be down at once. Both keep their disks, so this costs availability, not acknowledged writes. Kubernetes’ MaxUnavailableStatefulSet feature gate would count every unavailable member, but it is off by default.
  • Removed members’ disks cost storage. They stay until growth reattaches them or the fleet is deleted. Delete one by hand only if you don’t plan to grow back to that ordinal. A claim deleted before the leaders have re-formed their ensembles can still hold the only copy of recent writes.
  • A hand edit to the StatefulSet template starts rolling at once. The operator restores its template on the next reconcile, and the member that started rolling returns to that template on its own disk.

Nothing about the lifecycle. The storage reservation’s celld.eric.dev/current-operation annotation holds only capacity-policy state, and is absent unless the fleet has set spec.capacity. status.lifecycle is always empty. Editing status authorizes nothing.

For a fleet created by an earlier release, the operator switches the StatefulSet from OnDelete to RollingUpdate and writes the current template. Kubernetes then replaces each earlier Pod, highest ordinal first, including Pods held by the launcher scheduling gate. Claims keep their names and identities. An in-flight strict operation is dropped with an OperationSuperseded event. See upgrade the operator.

The lifecycle contract and retained disks describe the implementation.

Experimental software for evaluation.Capabilities and limitations· Contribute