Skip to content

Implementation detail: Fleet API

celld.eric.dev/v1alpha1 defines a namespaced CelldFleet and a cluster-scoped CelldStorageReservation. The generated CRD is the field and admission reference.

A fleet supplies an explicit compatible runtime digest, an existing runtime ServiceAccount, a dedicated bucket, region and zone allowlist. The operator creates workloads, private Services, NetworkPolicies and a PodDisruptionBudget. Pods run celld directly: Bucket with CELLD_DURABILITY=bucket, PersistentFleet with CELLD_DURABILITY=fleet on retained CSI claims.

Bucket has immutable Deployment or Ordered layout; both contract one member at a time, and Ordered removes the highest ordinal. PersistentFleet uses a StatefulSet. Neither layout is converted in place. Infrastructure outside Kubernetes, including bucket, IAM roles, worker nodes and EBS CSI, remains administrator-owned.

The mutable request fields are replicas, capacity, runtimeImage, maintenance, export, routing, mesh, execution, lifecycle, env and telemetry. Profile, service account, storage, layout and placement are fixed at creation. maintenance.paused stops new workload changes. A new restartToken requests a same-version restart. Restarts and upgrades are one-member rolling updates for both profiles; allowCoordinatedDowntime is accepted and ignored.

PersistentFleet has no member-replacement API. The operator replaces a member that cannot come back on its own: when its claim is Lost, or after the replacement delay when it is the one member down and the rest of the fleet is ready. The StatefulSet creates a fresh claim. See self-healing.

spec.env adds up to 32 CELLD_ variables with either a literal value or a same-namespace secretKeyRef (name and key). Operator-owned identity, network, durability and storage settings and reserved prefixes such as CELLD_REEXEC_ and CELLD_STRICT_ cannot be overridden. The OpenTelemetry namespace is reserved for its dedicated API. Changing environment entries rolls the fleet one member at a time. Secret values never enter the fleet API, status, or operator logs; Kubernetes resolves references when a Pod starts. Rotating a Secret does not update running processes: change maintenance.restartToken to restart the fleet after a rotation.

spec.telemetry enables OTLP export to collectorURL; omission leaves it disabled. The operator emits CELLD_OTEL=<collector URL>, optional sampler and flush settings, and optional OTEL_EXPORTER_OTLP_HEADERS from a same-namespace Secret key. It also emits OTEL_RESOURCE_ATTRIBUTES with k8s.namespace.name, k8s.pod.name (from POD_NAME, set by the downward API) and celld.fleet.name, so celld’s OTLP signals join the operator’s celld_fleet_* metrics on namespace and fleet; celld reads it from 0.6.0-ewhauser.2. It never emits the removed CELLD_OTEL_SINK. egress is required and adds one TCP rule for the collector URL’s port, scoped to one IP address (cidr /32 or /128) or labeled collector Pods (podLabels, with optional namespace). The administrator must ensure the URL resolves to the selected destination. The existing broad HTTPS rule for S3 and STS still permits other port 443 destinations; NetworkPolicy alone does not make HTTPS collector traffic exclusive to the telemetry egress selector. A collector using a different port receives only its scoped rule. Changing telemetry rolls the fleet one member at a time. The collector egress rule changes before the rollout starts, so members still running the old settings lose egress to a removed collector. Rotate the headers Secret with a restart so new Pods read it.

spec.export turns on celld’s change export, which streams every committed change to the cells’ SQLite databases out of celld; omission leaves it disabled. The operator emits CELLD_EXPORT=1 and the CELLD_EXPORT_* settings below, and reserves the CELLD_EXPORT prefix in spec.env. Omitted fields keep celld’s defaults.

field celld setting
sink CELLD_EXPORT_SINK: Bucket (default) or Kafka
classes CELLD_EXPORT_CLASSES, the classes to export
excludeTables CELLD_EXPORT_TABLES, Class.table entries never exported
maxTransactionBytes, maxRecordBytes, queueBytes CELLD_EXPORT_MAX_TX_BYTES, CELLD_EXPORT_MAX_RECORD_BYTES, CELLD_EXPORT_QUEUE_BYTES
bucket.name CELLD_EXPORT_BUCKET, an export bucket other than the fleet bucket
bucket.flushMilliseconds, bucket.flushBytes CELLD_EXPORT_FLUSH_MS, CELLD_EXPORT_FLUSH_BYTES
bucket.retentionDays CELLD_EXPORT_RETENTION=<n>d
kafka.brokers, kafka.topic, kafka.retryMilliseconds CELLD_EXPORT_KAFKA_BROKERS, CELLD_EXPORT_TOPIC, CELLD_EXPORT_RETRY_MS
kafka.propertiesSecretKeyRef CELLD_EXPORT_KAFKA_PROPERTIES=/etc/celld/export/<key>; the whole Secret is mounted read-only there

The Bucket sink writes Parquet objects under export/changes/ in the fleet bucket, or in bucket.name on the same endpoint and credentials, so the fleet’s ServiceAccount must be able to write there. It needs no new egress. The Kafka sink needs kafka.egress, which adds one TCP rule for the brokers’ ports to labeled broker Pods (podLabels, with optional namespace) or to a cidr, which may cover a whole network such as a managed cluster’s subnets. Create the topic before the fleet; the sink never creates it. librdkafka settings for TLS, SASL, compression and batching go in the properties key, one name=value per line, so SASL credentials stay in the Secret. Every key of the Secret is mounted, so a private CA or client certificate rides in the same Secret and the properties name its file, for example ssl.ca.location=/etc/celld/export/ca.crt. celld refuses properties that weaken delivery: acks below all, message.timeout.ms, delivery.timeout.ms, and turning on delivery.report.only.error or allow.auto.create.topics. Kafka needs a celld built with the export-kafka feature. Starting with v0.6.1-ewhauser.3, the fork publishes a -kafka image variant with this feature; pin its verified manifest digest in runtimeImage (see runtime requirements). The default image includes only the bucket sink. The operator does not support celld’s blob-stream sink, and the reserved CELLD_EXPORT prefix keeps spec.env from selecting it.

The export queue (queueBytes, 256 MiB by default) is held in the celld process, so size execution.memoryLimit for it. Export starts with the cells a node activates once it is on; backfill earlier cells and schedule the reconciler with the celld export commands. celld reports export gauges (celld.export.*) through spec.telemetry. Change export needs a celld build that has it: see runtime requirements. Adding, changing or removing export rolls members one at a time. The egress policy follows the new spec at the start of the rollout, so a member that still runs the old export can lose its sink until it is replaced; its queue fills and drops records as gaps rather than blocking writes. Cells activated on members that did not export yet are not exported: run the backfill once the rollout completes. A fleet created with export by v0.0.8 or v0.0.9 keeps its reservation only while that export is unchanged; recreate it to change export.

A bucket reservation permanently binds the bucket to the fleet UID. Recreating a fleet with the same name does not transfer ownership. Its celld.eric.dev/current-operation annotation holds only capacity-policy history and is removed when there is none; fleet status is a reconstructible projection. The reservation hash covers the fleet’s creation-time execution, lifecycle, env and telemetry, which the operator records in the celld.eric.dev/reserved-tuning annotation so the hash still verifies after they change. The record is trusted only when it reproduces the hash. A fleet created before this release gets its record on the first reconcile while those fields are unchanged, so upgrade the operator and let it reconcile before changing them. Otherwise the fleet reports StorageScopeConflict until the original values are restored. Preview fleets do not get a record, so their tuning remains tied to the reservation. Preview configuration lives on an existing fleet under spec.previews. It reserves a separate preview bucket to that parent fleet UID, then binds each child fleet to a disjoint storage.prefix and previewFleetRef. A generated preview fleet may set a Secret-backed custom storage.endpoint and separate storage.scratch.request/limit for its disk-backed emptyDir. These fields are immutable and do not change ordinary dedicated fleet defaults. spec.previews may be added once; it is immutable thereafter and excluded from the parent runtime reservation hash. Initialization requests and executor status use the existing CelldStorageReservation; its current-operation annotation remains operator-owned, while its status subresource holds the retained seed result. See PersistentFleet lifecycle before changing lifecycle code.

The /scale subresource maps desired replicas to spec.replicas and observed nonterminal Pods to status.replicas. External autoscalers target the CelldFleet, never the managed Deployment or StatefulSet. They gain no deletion authority.

Install with namespace-scoped RBAC for every managed namespace and attest CNI NetworkPolicy enforcement. See installation and security boundaries.

spec.mesh.istio places members in an Istio sidecar mesh that may run STRICT mTLS and AuthorizationPolicies. The Pod template gains the sidecar.istio.io/inject: "true" and celld.eric.dev/fleet: FLEET labels and asks for a native sidecar (sidecar.istio.io/nativeSidecar: "true"), which starts before celld and stops after it, so peer RPC works through startup recovery and the SIGTERM handoff. The EKS Pod Identity agent address bypasses the proxy. The fleet NetworkPolicy adds egress to istiod Pods (app: istiod) in controlPlaneNamespace (default istio-system) on TCP 15012. The operator owns an ALLOW AuthorizationPolicy named FLEET-mesh that admits port 8081 only from the fleet’s ServiceAccount and the operator’s (--operator-service-account), matched in any trust domain. With applicationAccess: AllowAll (the default) it also allows port 8080 from any caller; with Policies, port 8080 callers need an AuthorizationPolicy you write. The policy is created before the workload; a missing security.istio.io/v1 API or a same-named policy the fleet does not own reports InfrastructureBlocked. controlPlaneNamespace and applicationAccess change only policies and roll nothing. Adding or removing mesh rolls members one at a time. Until every member Pod matches, the operator keeps a PeerAuthentication FLEET-mesh that makes port 8081 PERMISSIVE and drops the identity requirement from the peer-port rule, since an unmeshed member’s plaintext peer RPC has no identity; the fleet NetworkPolicy still limits 8081 to members and the operator. It checks actual proxy containers, including native sidecars and terminating Pods, before restoring the strict rule and deleting the PeerAuthentication. Missing injection keeps the fleet in Provisioning. Departure retains istiod egress until the final proxy exits, then deletes both mesh objects and closes that egress. Replacement Pods disable injection explicitly; the retained fleet annotation celld.eric.dev/istio-control-plane keeps this opt-out after operator restarts, even in a namespace configured for injection. Fleets that never opted in keep their previous template. The operator reads each member’s /state on 8081 directly, so under STRICT mTLS the operator Pods must be in the mesh too. See Istio networking.

spec.routing optionally creates one Gateway API v1 HTTPRoute or a standard Ingress, named FLEET-routing. Both forward / for explicit hostnames to FLEET:8080. A required source namespace and Pod label selector admit only the selected data-plane Pods through a separate NetworkPolicy. Routing never exposes 8081. Gateway mode references an existing Gateway; Ingress mode specifies an IngressClass and optionally a same-namespace TLS Secret and annotations.

Routing can be added, edited, switched or removed without changing the workload or storage reservation. The operator updates only resources controlled by the current fleet UID, deletes with UID/resourceVersion preconditions and waits for obsolete routes to disappear before creating a replacement. Before removing an existing route, it checks the selected API, Service, isolation setting and resource ownership; a failed preflight preserves the existing route and policy. Removed managed annotations are pruned; unrelated controller annotations are preserved. Removing routing deletes the owned route, then its ingress policy. Routing also closes when fleet deletion starts; it does not delay runtime shutdown.

Reserved or oversized Ingress annotations are rejected at admission. Invalid routing stored by older schemas or admission bypasses reports RoutingReady=False without blocking runtime operations.

RoutingReady is separate from fleet readiness. HTTPRoute acceptance requires current-generation Accepted=True and ResolvedRefs=True for the configured parent. Ingress has no portable acceptance condition: an assigned address is reported, or AwaitingAddress remains Unknown. Neither observation proves DNS, TLS, Gateway listener programming or application connectivity. Missing optional Gateway CRDs report GatewayAPIUnavailable without blocking runtime operations. The operator polls routing status on normal reconciliation; it does not watch or provision Gateway, IngressClass, certificate or DNS resources. See the networking guide and the routing samples.

status.application and the ApplicationConverged condition describe application code loaded by the runtime. They are independent of the container image digest, Kubernetes object generation and fleet lifecycle readiness. Observation never publishes code, calls /reload, changes replicas or authorizes disk deletion.

The operator reads the existing private GET /state endpoint at most once every 15 seconds per fleet, with eight concurrent requests, a 10-second collection deadline and a 1 MiB limit per response. This requires no new runtime patch or image update. Missing deployment fields, oversized responses and unavailable nodes yield Unknown while normal lifecycle operations continue.

Status contains a collection timestamp, expected and freshly observed node counts, unavailable coverage, loaded versions grouped by version and artifact prefix, pending/swapping cell totals and at most 100 Pod diagnostics. Responses must be at most 90 seconds old. Pod UID, container incarnation, IP, readiness and workload ownership are checked against opening and closing membership inventories. Counts omit unavailable samples and are not complete fleet totals when coverage is incomplete. observedVersion is present only with complete fresh coverage and agreement on the loaded version and prefix; cells may still be transitioning.

ApplicationConverged Meaning
True / Converged Every current Pod is ready and freshly observed, all report the same loaded version and prefix, and no resident cells are observed pending or swapping.
False / MixedVersions Coverage is complete, but loaded versions or artifact prefixes differ.
False / Converging Nodes agree on the loaded version, but resident cells are still transitioning.
Unknown Observations are stale, unsupported, malformed, incomplete or interrupted by membership changes.

This reports observed agreement, not adoption of the latest published release. The existing API exposes neither the deployment pointer’s freshness nor failed adoption attempts. All nodes may agree on an old version even when a newer release cannot be loaded or the pointer is unavailable. A pipeline waiting for a particular release must compare status.application.observedVersion.version and prefix with its expected artifacts, require ApplicationConverged=True and check a fresh observedAt and condition observedGeneration.

Numeric application generations are process-local and never compared across nodes. A resident cell is pending when its generation differs from its own node’s current generation; rollback is supported. Celld samples its actor census and current generation separately, so this is an observation rather than an atomic rollout-completion guarantee. It does not prove application health or that every long-lived request from an earlier generation has ended. The condition does not gate Ready, scaling, maintenance or deletion. If the operator stops, persisted observations age; clients must check timestamps.

Experimental software for evaluation.Capabilities and limitations· Contribute