Upgrade the runtime
Use a compatible digest-pinned fork and qualify its recovery/storage-format
compatibility with the currently deployed version. Every possible recovery
participant must understand the fork’s native bucket_complete proof. A
matching image-name pattern does not establish compatibility.
The recorded .2 → .3 upgrade passes
for both profiles, with 12/12 acknowledged values preserved in each fleet.
That is graceful maintenance evidence, not permission to use .2 for crash
recovery or proof that other source/target pairs are compatible. Use the
current artifact for new fleets.
Most runtime changes roll. Some pairs cannot run together in one fleet and need a full stop instead. The compatibility reference says which.
Rolling upgrade
Section titled “Rolling upgrade”Change only the image:
kubectl --context YOUR_CONTEXT -n fleets patch celldfleet my-fleet --type merge -p '{"spec":{"runtimeImage":"ghcr.io/ewhauser/celld@sha256:YOUR_VERIFIED_DIGEST"}}'Replace YOUR_VERIFIED_DIGEST with the actual 64-character lowercase digest
before running the command. Members are replaced one at a time, so old and new
versions run together until the rollout completes; qualify that mix. No
downtime permission is needed.
A PersistentFleet member keeps its disk. The StatefulSet restarts one member at a time, highest ordinal first, and waits for each to be Ready before the next; see the PersistentFleet lifecycle.
Full-stop upgrade
Section titled “Full-stop upgrade”A PersistentFleet moving from a 0.5.1-ewhauser build to a 0.6.0-ewhauser
build needs a full stop. A 0.6.0 member recovers its previous session only
from followers that return the ranged tail format, which 0.5.1 members cannot
send. A rolling update would stall on its first member. Bucket fleets have no
followers and use the rolling upgrade above.
The fleet serves nothing from step 2 until its members are Ready again. Each
member keeps its disk, drains on SIGTERM and recovers from that disk.
CONTEXT=YOUR_CONTEXTNAMESPACE=fleetsFLEET=my-fleetIMAGE=ghcr.io/ewhauser/celld@sha256:YOUR_VERIFIED_DIGESTFLEET_UID=$(kubectl --context "$CONTEXT" -n "$NAMESPACE" get celldfleet "$FLEET" -o jsonpath='{.metadata.uid}')
# 1. Pause the fleet. The operator stops changing its workload and replacing members.kubectl --context "$CONTEXT" -n "$NAMESPACE" patch celldfleet "$FLEET" --type merge \ -p '{"spec":{"maintenance":{"paused":true}}}'kubectl --context "$CONTEXT" -n "$NAMESPACE" wait "celldfleet/$FLEET" \ --for=condition=MaintenancePaused --timeout=2m
# 2. Stop every member, and wait until no member Pod remains.kubectl --context "$CONTEXT" -n "$NAMESPACE" scale statefulset "$FLEET" --replicas=0while kubectl --context "$CONTEXT" -n "$NAMESPACE" get pods \ -l "celld.eric.dev/fleet-uid=$FLEET_UID" -o name | grep -q .; do sleep 5; done
# 3. Set the new image and resume in one change. Every member starts on it at once.kubectl --context "$CONTEXT" -n "$NAMESPACE" patch celldfleet "$FLEET" --type merge \ -p "{\"spec\":{\"runtimeImage\":\"$IMAGE\",\"maintenance\":{\"paused\":false}}}"kubectl --context "$CONTEXT" -n "$NAMESPACE" wait "celldfleet/$FLEET" \ --for=condition=Ready --timeout=20mDo not change the image before every member Pod is gone, and do not resume
without the new image: either starts some members on each version. Step 3
restores the fleet’s declared member count, including when spec.capacity is
set. The maintenance Kind suite runs these commands against a fleet under
write load and requires every acknowledged write to be readable afterwards.
Neither profile has an old release adapter or automatic rollback. Watch the conditions and troubleshoot waits.
Experimental software for evaluation.Capabilities and limitations· Contribute