Design record: 0024: PersistentFleet is a StatefulSet
Status: accepted and implemented. Builds on 0023, and replaces the PersistentFleet lifecycle that 0023 first proposed.
Context
Section titled “Context”0023 as first accepted kept PersistentFleet disks across restart and upgrade, but it gave the operator a lifecycle of its own:
- It read celld’s
/state.node_logfrom every member on each reconcile. It derived a fleet-widesettledvalue from it and gated each restart on it. - It deleted outdated Pods itself, one at a time. It lowered the PodDisruptionBudget to zero while the fleet was not settled.
- It deleted a removed member’s disk only when no celld session still needed it, and it held growth until that disk was gone.
- It inspected each member’s claim and volume, replaced members whose disk it
judged lost, and acted on a
replace-memberannotation. - It stored the last disruption time, applied count, image, restart token and workload identity in an annotation on the storage reservation.
This is about 600 lines of controller code plus a settlement package and a node-log client. It restates what celld already does.
celld’s own documentation (upstream v0.5.1) describes how to operate it:
- Roll out with the orchestrator’s rolling update. Stop a node with SIGTERM, wait for its replacement to report healthy, then move to the next node. A replacement holds its first healthy response until the fleet has absorbed the change: no active donor, and memory headroom on every live node. The docs say a deployer needs no separate fleet-level gate.
- celld does not scale itself. An external system starts and stops nodes, and celld hands their cells off.
- Local state persists across restarts. A restarting node serves its follower data to its peers before it recovers its own previous session, so nodes that restart together recover from each other’s disks.
- Node failure is a normal input, not a recovery procedure.
Nothing in celld asks an orchestrator to delete a node’s disk, to wait for node-log settlement, or to replace members.
Decision
Section titled “Decision”The operator renders a StatefulSet and its prerequisites and applies them. The StatefulSet controller performs every restart and every scaling step, and celld recovers every node. The operator replaces a member that cannot come back on its own, so the fleet heals without an administrator. Ordinary rollout and scaling need no operator lifecycle state. A disk replacement keeps a bounded handoff record on its victim Pod until that Pod can be released.
- Rollout.
RollingUpdatewith partition 0. Kubernetes restarts one member at a time, highest ordinal first, and waits for each one to be Ready. The readiness probe is celld’s health endpoint, which returns 200 only after the node has recovered its predecessor session and passed celld’s first-readiness gate. Restart, runtime upgrade and operator upgrade all take this path. - Disks. A member keeps its disk until the fleet is deleted or the member
is replaced because it cannot come back (see below). Restart, upgrade and
scale-in keep the claim. Scaling back out reattaches it, which celld treats as
a restart on the same disk. StatefulSet claim retention stays
Retainfor scale-in and deletion. - Disruption budget. A fixed
maxUnavailable: 1. A Pod that is not Ready counts against it, so node drains proceed one member at a time. - Replica count. Growth is applied one run of ordinals at a time, kept disks and fresh disks never in the same run (see below). Contraction removes one member per step, and only after the previous change has rolled out. This is the same contraction rule Bucket fleets follow. Automatic and External contraction still require survivor-capacity evidence.
- Storage identity. Before creating or growing the StatefulSet, the operator refuses to proceed if a claim at a member’s name does not carry the fleet’s UID label. A disk from another fleet is never adopted.
- Reservation state. The reservation annotation holds only capacity-policy
state, which exists once a fleet sets
spec.capacity. The operator removes the annotation when there is none. Fields written by earlier releases are dropped on the first write, including claim identities, strict operations and disruption times. A selected disk replacement uses its own Pod annotation, index label and finalizer, described below; it adds no fleet-wide rollout state.
The operator does not read /state.node_log and does not release the disks of
removed members.
Self-healing
Section titled “Self-healing”No administrator is expected to intervene. The operator judges each member from what the cluster reports on each reconcile and replaces at most one member at a time:
- A Pod on a node that no longer answers. A member Pod still present more
than two minutes after its termination grace has ended is force-deleted, so
the StatefulSet can recreate it. The member keeps its disk. A process that
survives on the lost node is fenced by celld’s lease, and a
ReadWriteOncePoddisk attaches to one node at a time. This applies to every StatefulSet fleet, including Ordered Bucket. - A lost volume. When Kubernetes marks a member’s claim
Lost, the volume behind it no longer exists. The operator commits deletion of the selected Pod before deleting its original claim, using the handoff below. The StatefulSet recreates both. - A member that cannot come back. A member that has stayed down for the
replacement delay is treated as lost if every other member has been ready for
five minutes. The delay defaults to 10 minutes (
--member-replacement-delay). It is patience, not the safety condition. It outlasts the roughly six minutes Kubernetes takes to force-detach a volume from a lost node, and typical node provisioning, so a member that can return usually does so on its own disk. This covers a disk stranded in an unavailable zone, a volume that no longer attaches, and a corrupt disk that keeps celld from starting. Leaders stop using a departed member within seconds, and every node sweeps dead leaders every 30 seconds. So once the rest of the fleet has been ready that long, no session depends on the down member’s disk. The operator commits deletion of the selected Pod before deleting its original claim, and the member returns on a fresh disk. Two cases are left alone because a new disk would not help: a member waiting on its image or configuration, and a member already on a disk created for its current Pod. Only one down member is ever replaced this way. When two are down, replacing either could lose writes that only their disks hold. - Where a fresh disk goes. A zonal disk pins its member to one zone, and
the zone spread counts only Pods that are on nodes. A fresh disk scheduled
while another member’s Pod is pending can take the zone that member’s disk
needs, and that member then cannot be scheduled at all
(#79). The operator
gives a member a fresh disk only when the StatefulSet already runs and has
observed the operator’s spec, so no rollout is about to recreate another
member, and the scheduler has placed, or found no node for, every other
member’s Pod. The fresh disk then goes to the zone the fleet is missing.
Growth follows the same rule: the StatefulSet creates every new Pod at once,
so the operator adds one run of consecutive ordinals at a time, all
reattaching kept claims or all getting new ones, and each run waits until the
StatefulSet has observed its spec and the scheduler has decided every current
member’s Pod. The runs are read from the claims on each reconcile. A
capacity-policy addition that spans both kinds is cut at the end of its first
run and recorded as that step. A member the scheduler cannot place is named in the fleet’s status. If the
operator has not replaced it by the end of the replacement delay, the fleet
reports
Blockedwith reasonMemberUnschedulable.
Replacement handoff
Section titled “Replacement handoff”A Pod read and a PVC delete are separate API operations. Rereading the Pod before deleting its claim cannot close the race: the same Pod can recover after that read, or a new Pod can take its StatefulSet name. Replacement instead uses the Pod delete as the commitment point:
- The operator prepares the selected, non-terminating Pod with its own
finalizer and a record of the fleet UID and original claim’s UID and
resourceVersion. A private
celld.eric.dev/replacement-fleetlabel indexes that hold by its original fleet UID. - It deletes that exact Pod with UID and resourceVersion preconditions. A recovery or metadata change before the delete makes it conflict, so the disk is retained. The finalizer holds the terminating Pod’s name and prevents the StatefulSet from creating a successor against the old claim.
- Only an acknowledged Pod delete allows the record to be marked committed. An interrupted prepared handoff always aborts and keeps the disk, including an externally deleted Pod or a successful delete whose response was lost.
- A committed handoff deletes only the recorded claim, with both UID and resourceVersion preconditions. A live claim changed since preparation is retained. A claim at the same name with a different UID is never deleted.
- Once deletion of the original claim is durably accepted, or that claim UID is absent, the operator removes its own finalizer, handoff record and index label. It does not wait for the PVC to disappear: Kubernetes’ PVC protection needs the old Pod to finish before it can finish deleting an attached claim.
A new reconciler aborts prepared records and resumes committed ones from these persisted objects. This cleanup runs before runtime validation and normal workload gates, and includes indexed Pods outside the current replica range. The private index still finds a hold after the ordinary fleet label or owner reference changes; that drift aborts disk deletion and releases the hold. Pausing or deleting the fleet releases its holds without starting another disk delete. An already accepted deletion cannot be undone. Ordinary rollouts, scale-in and Pod restarts do not use this handoff and retain their disks.
For a replaced disk, the new member answers recovery for its old disk with a
conclusive “no fragment” (0.5.1-ewhauser.7; earlier builds refuse until the
new member publishes its own lease). celld then records a bounded loss for any
session whose only complete copy was on the old disk, and seals it. With the
rest of the fleet ready, no such session remains unless a second failure
happened first.
Consequences
Section titled “Consequences”- PersistentFleet reconciles like Bucket.
persistent.gokeeps the profile’s storage checks, self-healing and fleet deletion.internal/fleethealth, the node-log client, thereplace-memberannotation and the settlement-driven budget are removed. - The operator no longer depends on
/state.node_log. Any runtime restarts under the same rolling update, including runtimes that predate node-log state. - Pacing is readiness only. The operator does not wait for celld to finish recovering other sessions before the next restart. With disks retained, a restart destroys nothing. Exposure to a second, unplanned disk loss during a roll is celld’s replication factor, as 0023 already noted.
- A rolling update waits for the member it just restarted, not for an
unrelated member that is already down. Two members can then be down at once.
Both keep their disks, so the cost is availability, not acknowledged writes.
Kubernetes’
MaxUnavailableStatefulSetgate would count every unavailable member, but it is off by default. - A removed member’s disk costs storage until the fleet grows back or is
deleted. An administrator may delete such a claim by hand. celld stops
depending on a departed member once its leaders re-form their ensembles,
within seconds on
0.5.1-ewhauser.6or later. A claim deleted before then can still hold the only copy of recent writes. - A member that cannot come back holds up a rollout or a contraction for at most the replacement delay. A lost volume is replaced once the scheduler has decided the other members’ Pods, usually within seconds. The fleet keeps serving on the other members throughout.
- Growth reattaches kept disks before it adds fresh ones, and adds each run only once every current member’s Pod has been placed or found unschedulable. Growth across a mix of kept and fresh ordinals therefore takes several scheduling rounds instead of one. A capacity-policy addition cut short this way adds the rest only if the policy asks again.
- A claim an administrator deletes gets a fresh disk from the StatefulSet without that wait. If it takes the zone a pending member’s disk needs, the pending member cannot be scheduled and is replaced after the replacement delay, as a member that cannot come back.
- The operator deletes an existing disk on its own only when that disk’s member has been down for the replacement delay while the rest of the fleet was ready. This is a deliberate change from 0023’s first lifecycle, where only an administrator could do that.
- Two members that cannot come back at the same time wait until one returns, unless their volumes are lost. celld’s guarantee covers the loss of one node.
- The fleet rebuild that 0023’s first lifecycle left as an explicit
administrative operation is not needed. A fleet recovers from any outage on its members’ own disks, and a
rolling restart is the only lever an administrator needs. Replacing every
disk at once would record every session whose writes were only on those
disks as lost, and on runtimes before
0.5.1-ewhauser.7it stalls recovery. - The ClusterRole no longer needs PersistentVolume, Node or VolumeAttachment reads, and the namespaced Role no longer needs to create claims.
- Disk replacement uses the existing namespaced Pod update permission for its handoff record and finalizer. Before rolling back to an operator that does not recognize those records, let current replacements finish and release their holds.
Migration
Section titled “Migration”Existing fleets are adopted in place. For a StatefulSet created by v0.0.5 or
by the unreleased 0023 controller, the operator changes the update strategy
from OnDelete to RollingUpdate and writes the current template. Kubernetes
then replaces each earlier Pod, highest ordinal first, including Pods still
held by the launcher scheduling gate. Retained claims keep their names and
identities. The storage reservation annotation is rewritten to the capacity
history alone, or removed.
Qualification
Section titled “Qualification”The kind suites run the real fork under continuous write load and require every acknowledged write to be readable afterwards:
- a rolling restart and a runtime upgrade on retained disks;
- a graceful Pod delete, a forced Pod delete,
SIGKILLof celld, a node drain held by the budget, and an operator restart in the middle of a rollout and of a contraction; - a deleted claim, which the StatefulSet replaces with a fresh disk;
- a lost volume, which the operator replaces;
- a node whose kubelet stops while celld keeps running on it. The operator force-deletes the stranded Pod and, after the replacement delay, replaces the member’s disk.
- scale-in that keeps the removed member’s disk, and scale-out that reattaches it;
- fleet deletion that removes compute, then disks, and keeps the bucket reservation;
SIGKILLof celld on every member at once, and every member Pod deleted at once;- an upgrade from v0.0.5 after its launcher retired a member’s disk. The member rejoins on that disk (#60).
Experimental software for evaluation.Capabilities and limitations· Contribute