Hi Lorenzo,

Thanks for the thorough revision of FIP-41. I think the proposal is in great 
shape now and we can start a VOTE thread.

One suggestion: it would be good to address Anton's latest feedback (Jun 22) 
first before starting vote.

Yours,
Yunhong

On 2026/07/01 16:52:16 Giannis Polyzos wrote:
> Hi Lorenzo and all,
> 
> I think this proposal should be in a good shape now to start a vote thread.
> 
> One last item I think it’s worth discussing is where the operator will live.
> 
> Looking into other projects and following the approach we took with the
> clients, maybe we should create a separate repo - similar to how Flink does
> it, to host the operator.
> 
> Looking forward for your thoughts on this.
> 
> Best,
> Giannis
> 
> On Sun, 28 Jun 2026 at 5:05 AM, Forward Xu <[email protected]> wrote:
> 
> > Hi Lorenzo,
> >
> > Thanks for kicking off this discussion — big +1 from me on introducing a
> > Fluss Kubernetes Operator.
> >
> > I fully agree with the motivation: a Helm chart is great for the initial
> > deployment, but it's a one-shot templating tool with no awareness of
> > Fluss's runtime state. The operational events you listed — safe rolling
> > restarts with tablet leadership migration, ordered version upgrades,
> > scale-in with tablet draining, partial-failure recovery (PVC reattachment,
> > pod identity preservation), and dynamic config updates without unnecessary
> > restarts — are exactly the kind of Fluss-aware reconciliation that a
> > controller should own.
> >
> > I'd also like to add that this will be especially valuable on managed cloud
> > Kubernetes services. For example, on *Tencent Cloud TKE*, an Operator that
> > understands Fluss's state would let users:
> >
> >    - run Fluss as a first-class, declarative FlussCluster resource and
> >    integrate naturally with cloud-native tooling (autoscaling, monitoring,
> >    etc.);
> >    - safely leverage elastic node scaling — TKE node pools can scale in/out
> >    frequently, so having the Operator drain TabletServers and migrate
> >    leadership before pod termination is critical to avoid data loss;
> >    - handle storage lifecycle cleanly with cloud block storage (CBS-backed
> >    PVCs), ensuring PVC reattachment and pod identity are preserved across
> >    rescheduling.
> >
> > On the runtime side, it's also worth noting that the JDK choice is
> > flexible. In addition to common OpenJDK distributions, *Tencent Kona JDK*
> > is a well-proven, production-grade option (it's an OpenJDK downstream
> > that's widely used at scale and works well in containerized/Kubernetes
> > environments), so users on Tencent Cloud can run Fluss on a fully
> > supported, battle-tested JDK.
> >
> > Following the proven pattern of the Apache Flink and Spark Kubernetes
> > Operators is the right direction, and it would make Fluss much easier to
> > operate in production on the cloud.
> >
> > Looking forward to the detailed design. Thanks again for driving this!
> >
> > Best,
> >
> > ForwardXu
> >
> > Anton Borisov <[email protected]> 于2026年6月22日周一 11:42写道:
> >
> > > Hi Lorenzo,
> > > Thank you for the update, read through the updated proposal.
> > >
> > > A few clarifications:
> > > 1. External clients. listeners.client.serviceType offers
> > > LoadBalancer/NodePort, but advertised.listeners is derived from the
> > > headless Service DNS, i.e. in-cluster names, and a client connects
> > > directly to each bucket leader's advertised address. Is out-of-cluster
> > > client access in scope for v1alpha1? If so, how does the operator
> > > derive a per-server external advertised address? LoadBalancer opens
> > > the port, but metadata still hands clients in-cluster DNS, similar to
> > > the per-broker case Strimzi handles. If not, should the client
> > > listener expose LoadBalancer/NodePort yet?
> > > 2. Per-step upgrade gate. It waits for getClusterHealth() == GREEN,
> > > which is all-or-nothing across every bucket via computeClusterHealth.
> > > How does the gate avoid stalling until recoveryTimeout when an
> > > unrelated bucket is briefly out of ISR and keeps the whole cluster
> > > non-GREEN, even though the just-restarted server is fine? Is GREEN
> > > meant to be scoped to the restarted server here, or is the per-server
> > > recovery predicate, currently filed under the observability open
> > > question as "not required for correctness", actually needed for
> > > upgrade liveness?
> > > 3. Storage resize. StatefulSet volumeClaimTemplates are immutable, so
> > > a change to tabletServer.storage.size cannot be applied to the
> > > StatefulSet template. Existing PVCs can be expanded in place if the
> > > StorageClass allows it, but new pods keep the old size until the
> > > StatefulSet is recreated. How does the operator handle a storage-size
> > > change, or is resize out of scope for v1alpha1?
> > >
> > > One smaller question:
> > > Server tags are keyed by server id in ZK, and ids are ordinal-reused
> > > on scale down/up. Does the operator clear PERMANENT_OFFLINE on
> > > scale-in/out, so scaling back up does not reuse an ordinal id that is
> > > still tagged and gets silently fenced from placement by
> > > CoordinatorMetadataCache.getLiveServers?
> > >
> > > Let me know what you think
> > >
> > > -- Anton
> > >
> > > чт, 11 июн. 2026 г. в 16:09, Lorenzo Affetti via dev <
> > [email protected]
> > > >:
> > > >
> > > > Subject: Re: [DISCUSS] FIP-41: Fluss Kubernetes Operator -- revised
> > > >
> > > > Hi all,
> > > >
> > > > Thanks Anton, Yunhong, Giannis, Keith and Michael -- the review
> > > materially
> > > > improved the FIP. I've pushed a revised version; the full diff is
> > > attached for
> > > > the details. Summary of the decisions:
> > > >
> > > > Rebalance / scale-in / scale-out (Anton, Giannis)
> > > >   - Corrected two wrong claims: a decommission API DOES already exist,
> > > and the
> > > >     coordinator does NOT auto-rebalance.
> > > >   - Split on the read/drive line:
> > > >       * v1alpha1 ENFORCES the scale-in safety gate -- it refuses to
> > > remove a
> > > >         TabletServer that still hosts replicas -- backed by one small
> > > read API we
> > > >         commit to (Admin.describeTabletServers(), same shape as #3400).
> > > >         Evacuation itself stays a manual admin action.
> > > >       * The operator does NOT drive data movement in v1alpha1
> > (rebalance
> > > is
> > > >         cluster-wide/unscoped, no listServerTags, not min-ISR-aware).
> > > >       * v1beta1 drives it behind an opt-in, once listServerTags /
> > > >         decommissionServer land.
> > > >   - Scale-in and scale-out share one coherent (non-driving) posture.
> > > >
> > > > Recovery & upgrades (Yunhong)
> > > >   - Readiness is anchored on the new getClusterHealth() GREEN signal --
> > > GREEN
> > > >     already means "KV-restored + caught-up + in-ISR + leaders-active".
> > > >   - The operator gates each rolling-upgrade step on GREEN (held for a
> > > >     stabilization window).
> > > >   - Split controlledShutdownTimeout from recoveryTimeout; on timeout
> > the
> > > upgrade
> > > >     stalls (no auto-rollback).
> > > >   - Kept TabletServer-first ordering, per the official upgrade docs.
> > > >
> > > > Disruption & PVCs (Keith, Giannis, Anton)
> > > >   - Operator rolls pods via direct delete + GREEN gate.
> > > >   - PDB defaults to maxUnavailable: 0.
> > > >   - PVCs use persistentVolumeClaimRetentionPolicy and are never
> > > auto-deleted.
> > > >   - Server identity is ordinal-derived and guarded by disk.properties.
> > > >   - Dropped status.phase in favour of conditions only.
> > > >
> > > > Implementation (Michael)
> > > >   - Java Operator SDK (built on fabric8).
> > > >   - Proposed minimum Kubernetes version: 1.29.
> > > >
> > > > A few things I deliberately scoped out and would value your read on:
> > > >
> > > >   - Observability (Yunhong): GREEN plus its aggregate counts are the
> > > operator's
> > > >     control contract; richer per-server / KV-restore / lag signals are
> > > >     diagnosis-grade (and largely exist as JMX metrics already), so
> > > they're an
> > > >     optional v1beta1+ enhancement rather than frozen here.
> > > >
> > > >   - Secrets (Keith): plaintext config is a pre-existing Fluss/chart
> > > limitation,
> > > >     not operator-specific. The operator materializes config in a Secret
> > > >     (defense-in-depth, admittedly cosmetic); the real fix (config
> > > interpolation /
> > > >     secret references) is tracked at the Fluss/chart level.
> > > >
> > > >   - Lake tiering: an explicit v1alpha1 non-goal.
> > > >
> > > > Details and exact wording are in the attached diff. Happy to discuss
> > any
> > > of the
> > > > above.
> > > >
> > > > Best,
> > > > Lorenzo
> > > >
> > > > On Fri, May 29, 2026 at 9:52 AM Keith Lee <[email protected]> wrote:
> > > >>
> > > >> Hi Lorenzo.
> > > >>
> > > >> Thank you for the detailed proposal. This will improve the operational
> > > >> experience of maintaining Fluss clusters.
> > > >>
> > > >> 1. Since configurationOverrides is rendered straight into a ConfigMap,
> > > any
> > > >> credential-bearing keys would end up in plaintext. Could we add a
> > > >> sensitive-overrides path that resolves from a Secret, even if the
> > > >> non-sensitive map stays as-is for v1alpha1? This is likely to be an
> > > >> important factor in application security reviews of users considering
> > to
> > > >> adopt Fluss operator.
> > > >>
> > > >> 2. Can you clarify how the operator will terminate/roll pods e.g.
> > > delete or
> > > >> eviction API? I believe using direct delete with pod disruption budget
> > > of
> > > >> n-1 might cause unavailability in TOCTOU scenarios as direct deletion
> > > bypasses
> > > >> PDB. For reference, strimzi sets maxUnavailable to zero (effectively
> > > >> minAvailable N) [1] and direct deletion, because they recognise that
> > > >> readiness probe is a weak signal of ISR status: “
> > > >> the readiness probe doesn't guarantee that all replicas are in-sync
> > > again”.
> > > >>
> > > >> Best regards
> > > >> Keith
> > > >>
> > > >> [1]
> > > >>
> > >
> > https://github.com/strimzi/proposals/blob/main/029-adopt-the-drain-cleaner-utility.md#disruptions
> > > >>
> > > >> On Wed, 27 May 2026 at 15:13, Giannis Polyzos <[email protected]>
> > > wrote:
> > > >>
> > > >> > Hi Lorenzo
> > > >> > And thank you for the proposal.
> > > >> >
> > > >> > I'm sharing similar thoughts to the previous comments. Overall, the
> > > >> > proposal is great, but there are a few things I think will help
> > > improve.
> > > >> >
> > > >> > Cluster rebalance is an important operation, and Fluss already
> > > supports it,
> > > >> > so I think it would be great to address:
> > > >> > 1. Scale-out without a rebalance trigger leaves clusters unbalanced,
> > > so the
> > > >> > operator must call rebalance() post scale-out and poll progress
> > > >> > 2. Scale-in with no drain path is unusable, so the operator must do
> > > >> > TEMPORARY_OFFLINE tag + rebalance + verify zero-bucket must precede
> > > replica
> > > >> > reduction
> > > >> > 3. Server ID reuse on PVC recycle corrupts ZK state, so the operator
> > > design
> > > >> > should have ID assignment; so I would like to see the PVC lifecycle
> > > >> > ownership
> > > >> >
> > > >> > I think it might also be useful to have a
> > > >> > terminationGracePeriodSeconds/preStop hook design
> > > >> >
> > > >> > For PVCs, it's also not clear to me about the model, and I think it
> > > should
> > > >> > be clear that we never auto-delete PVCs
> > > >> >
> > > >> > With the above, we can set a good foundation for the operator, and
> > > with
> > > >> > these addressed its also a +1 for me
> > > >> >
> > > >> > Best,
> > > >> > Giannis
> > > >> >
> > > >> > On Thu, May 14, 2026 at 9:51 AM Yunhong Zheng <[email protected]>
> > > wrote:
> > > >> >
> > > >> > > Hi Lorenzo, and thanks Anton for the detailed review.
> > > >> > >
> > > >> > > +1 to the overall direction. I'd like to add one more dimension to
> > > the
> > > >> > > rolling-upgrade story that I think the FIP currently
> > > under-specifies: the
> > > >> > > operator must gate progression on replica recovery on the way in,
> > > not
> > > >> > only
> > > >> > > on leader migration on the way out.
> > > >> > >
> > > >> > > Today the FIP talks about integrating with controlled shutdown,
> > > which
> > > >> > > protects the outgoing pod (leader is migrated before termination).
> > > But
> > > >> > once
> > > >> > > the new pod comes up, a Fluss process can become listener-ready
> > > within
> > > >> > > seconds, while:
> > > >> > >  - log replicas may still be catching up to the leader,
> > > >> > >  - PK/KV replicas may still be downloading the snapshot from
> > remote
> > > >> > > storage and replaying the changelog up to the high-watermark,
> > > >> > >  - none of the buckets hosted on that server have rejoined the ISR
> > > yet.
> > > >> > >
> > > >> > > If Pod.Ready flips to true at process-start time, a StatefulSet
> > > >> > > RollingUpdate will happily move on to the next ordinal, and the
> > > operator
> > > >> > > can find itself restarting pod N+1 while pod N is still
> > recovering.
> > > For
> > > >> > PK
> > > >> > > tables with non-trivial RocksDB state this window can be minutes.
> > > >> > >
> > > >> > > This connects to Anton's point that "TCP probes are too weak for
> > > operator
> > > >> > > readiness". I'd push it one step further:
> > > >> > >
> > > >> > > 1. Define Pod.Ready as "fully recovered", not just "process up".
> > > >> > > Concretely: all replicas hosted on the server are LogCaughtUp &&
> > > >> > > (KvRestored if PK) && InSyncWithLeader. The cleanest
> > implementation
> > > is a
> > > >> > > Pod readiness gate (e.g. fluss.apache.org/replicas-recovered=True
> > )
> > > >> > > flipped by the operator after polling the cluster, rather than
> > > >> > overloading
> > > >> > > the HTTP/TCP probe.
> > > >> > >
> > > >> > > 2. Make the rolling-upgrade gate a two-condition predicate:
> > > >> > > (a) the previously-restarted pod is fully recovered, AND
> > > >> > > (b) cluster-level invariants under-min-isr buckets == 0 and
> > > kv-recovering
> > > >> > > replicas == 0 held for a configurable stabilization window.
> > > >> > >
> > > >> > > 3. Freeze the observability contract here, even if drain stays in
> > a
> > > >> > > companion FIP. Path A as written cannot actually answer "is
> > replica
> > > >> > > recovery complete?" — replicasOnTabletServer() being internal
> > > (Anton's
> > > >> > > point) is one example, but the bigger gap is that there is no
> > > >> > > cluster-visible signal for KV restore progress or follower lag
> > > aggregated
> > > >> > > per server. Concretely I'd like the FIP to commit to exposing:
> > > >> > >   - per-bucket / per-server in-sync predicate
> > > >> > >   - kv.snapshot.restore.in-progress and restore.bytes-remaining
> > > >> > >   - log.follower.lag-bytes / lag-time-ms
> > > >> > >   - cluster-level under-min-isr-bucket-count and
> > > offline-bucket-count
> > > >> > >
> > > >> > > 4. Extend status.tabletServer.pods[] accordingly: recovered,
> > > >> > > underReplicatedBuckets, kvRestoreProgress, logCatchupLagBytes;
> > and a
> > > >> > > cluster-level summary (underMinIsrBucketCount,
> > offlineBucketCount).
> > > This
> > > >> > > also gives operators a SRE-grade observable surface during
> > > incidents.
> > > >> > >
> > > >> > > 5. Distinguish log-only vs PK workloads in defaults. A PK
> > > TabletServer
> > > >> > can
> > > >> > > take orders of magnitude longer to recover than a log-only one. A
> > > single
> > > >> > > perPodTimeout is misleading; consider splitting
> > > controlledShutdownTimeout
> > > >> > > from recoveryTimeout. On timeout the upgrade should stall with a
> > > Stalled
> > > >> > > condition, not auto-rollback — auto-rollback for stateful systems
> > is
> > > >> > > dangerous and the operator should require explicit human input
> > > (e.g. an
> > > >> > > annotation) to resume or proceed.
> > > >> > >
> > > >> > > 6. Coordinator/TabletServer ordering during upgrade. With
> > > >> > > coordinator.replicas=1, there is a Coordinator-unavailable window
> > > during
> > > >> > > its own upgrade. ControlledShutdown, ISR changes and KV-standby
> > > >> > > coordination all depend on the Coordinator being up. The FIP
> > should
> > > state
> > > >> > > that TabletServer rolling only begins after the Coordinator is
> > > upgraded
> > > >> > and
> > > >> > > has stabilized for the same window, and that each TS step
> > re-checks
> > > >> > > Coordinator health.
> > > >> > >
> > > >> > > Otherwise +1 from me. Looking forward to the poc code.
> > > >> > >
> > > >> > > Yours,
> > > >> > > Yunhong Zheng (Swuferhong)
> > > >> > >
> > > >> > > On 2026/05/12 16:17:20 Anton Borisov wrote:
> > > >> > > > Hi Lorenzo,
> > > >> > > >
> > > >> > > > Thanks for writing this up. I like the direction. An operator is
> > > the
> > > >> > > > right next step after the Helm chart, and the proposed shape is
> > > broadly
> > > >> > > > reasonable.
> > > >> > > >
> > > >> > > > I checked the FIP against the current code and I think a few
> > > points
> > > >> > > > are worth discussing/considering:
> > > >> > > >
> > > >> > > > 1. Drain / scale-in
> > > >> > > >
> > > >> > > > The FIP says Path B needs a new drain primitive, while Path A
> > only
> > > >> > blocks
> > > >> > > > scale-in if the target TabletServer still has replicas.
> > > >> > > >
> > > >> > > > I think we already have most of the low-level pieces:
> > > >> > > >
> > > >> > > > - ServerTag.PERMANENT_OFFLINE / TEMPORARY_OFFLINE
> > > >> > > > - AddServerTagRequest / RemoveServerTagRequest
> > > >> > > > - RebalanceRequest
> > > >> > > > - ListRebalanceProgressRequest
> > > >> > > > - RebalanceStatus
> > > >> > > >
> > > >> > > > From the rebalance code, offline-tagged servers are not just
> > > excluded
> > > >> > > > from new placements. ReplicaDistributionGoal moves replicas out
> > of
> > > >> > them,
> > > >> > > > including followers, and the target replica count is effectively
> > > zero.
> > > >> > > >
> > > >> > > > So scale-in could be:
> > > >> > > >
> > > >> > > > AddServerTag(PERMANENT_OFFLINE)
> > > >> > > > -> Rebalance
> > > >> > > > -> wait for COMPLETED
> > > >> > > > -> scale down StatefulSet
> > > >> > > >
> > > >> > > > TEMPORARY_OFFLINE may also be usable before rolling restart.
> > > >> > > >
> > > >> > > > Path A also needs server work as written. The operator needs to
> > > know
> > > >> > > > whether a TabletServer still has replicas, but that count is
> > > internal
> > > >> > in
> > > >> > > > CoordinatorContext.replicasOnTabletServer() and is not exposed
> > > through
> > > >> > > > Admin.getServerNodes().
> > > >> > > >
> > > >> > > > So I do not think the choice is “Path A without server work” vs
> > > “Path B
> > > >> > > > with server work”. Both need a small server/API decision.
> > > >> > > >
> > > >> > > > The important caveats are:
> > > >> > > >
> > > >> > > > - rebalance can fail if the remaining TabletServers cannot
> > absorb
> > > the
> > > >> > > >   replicas;
> > > >> > > > - I do not see min-ISR being considered by the rebalance
> > > executor, so
> > > >> > > >   decommission close to min-ISR may create write-availability
> > > windows;
> > > >> > > > - tag + rebalance is currently multi-step, so the operator needs
> > > crash
> > > >> > > >   recovery if it dies between the two calls;
> > > >> > > > - the operator would depend on @PublicEvolving Admin/rebalance
> > > APIs.
> > > >> > > >
> > > >> > > > Given that, I think a small wrapper may be cleaner:
> > > >> > > >
> > > >> > > > DecommissionServer(serverId) -> rebalanceId
> > > >> > > >
> > > >> > > > The server would own tag + rebalance as one operation. The
> > > operator
> > > >> > would
> > > >> > > > start it and poll progress.
> > > >> > > >
> > > >> > > > 2. Helm naming / adoption
> > > >> > > >
> > > >> > > > The FIP says operator naming matches the Helm chart to support
> > > in-place
> > > >> > > > adoption. The current chart uses fixed names such as
> > > tablet-server,
> > > >> > > > tablet-server-hs, coordinator-server and coordinator-server-hs.
> > > >> > > >
> > > >> > > > That means two Helm installs in one namespace collide. If the
> > > operator
> > > >> > > > keeps these names, v1alpha1 is effectively one FlussCluster per
> > > >> > > > namespace.
> > > >> > > >
> > > >> > > > I think the FIP should either state that constraint, or move
> > both
> > > chart
> > > >> > > > and operator resources to cluster-prefixed names. Otherwise
> > > adoption
> > > >> > and
> > > >> > > > multi-cluster UX are pulling in different directions.
> > > >> > > >
> > > >> > > > Smaller points
> > > >> > > >
> > > >> > > > - I would not use both status.phase and conditions. Conditions
> > are
> > > >> > enough
> > > >> > > >   and avoid forcing mutually-overlapping states into one enum.
> > > >> > > > - TCP probes are too weak for operator readiness. Readiness
> > > should mean
> > > >> > > >   the server is actually usable from Fluss’ point of view:
> > > registered
> > > >> > > >   with the Coordinator, ZK connected, etc.
> > > >> > > > - PVC retention should use StatefulSet
> > > >> > > >   persistentVolumeClaimRetentionPolicy. I would not add custom
> > PVC
> > > >> > > >   finalizer logic unless we really need it.
> > > >> > > > - Lake tiering is not mentioned. It is fine to keep it out of
> > > v1alpha1,
> > > >> > > >   but the FIP should say that explicitly.
> > > >> > > > - Your "Configuration Updates" section asserts the registry is
> > > sourced
> > > >> > > > from Fluss, but your Open Questions section
> > > >> > > > correctly flags this as unresolved - I'd align the body with
> > Open
> > > >> > > > Questions. For v1alpha1 a small client-side allowlist mirrored
> > > >> > > > from DynamicServerConfig.ALLOWED_CONFIG_KEYS is probably fine I
> > > reckon
> > > >> > > >
> > > >> > > >
> > > >> > > > Overall, I like the proposal. The main thing I would clarify is
> > > the
> > > >> > > > operator/server boundary: what the operator can safely
> > orchestrate
> > > >> > today,
> > > >> > > > what needs a small server API, and what should be left out of
> > > v1alpha1.
> > > >> > > >
> > > >> > > > -- Anton
> > > >> > > >
> > > >> > > > чт, 7 мая 2026 г. в 09:57, Lorenzo Affetti via dev <
> > > >> > [email protected]
> > > >> > > >:
> > > >> > > > >
> > > >> > > > > Hi Michael,
> > > >> > > > >
> > > >> > > > > Thanks for the careful read.
> > > >> > > > >
> > > >> > > > > *Framework*: Java Operator SDK, not plain fabric8. JOSDK is
> > > built on
> > > >> > > > > fabric8 anyway, so we keep the option to drop down whether
> > > needed.
> > > >> > > > >
> > > >> > > > > *Minimum Kubernetes version*: not currently fixed in the FIP.
> > My
> > > >> > > proposal
> > > >> > > > > is 1.29 as the floor.
> > > >> > > > >
> > > >> > > > > Reasoning for framework:
> > > >> > > > >
> > > >> > > > > For a Fluss operator that has to coordinate rolling upgrades,
> > > >> > scale-in
> > > >> > > > > safety gates, dynamic vs restart-inducing config diffs, and a
> > > >> > migration
> > > >> > > > > state machine, JOSDK's dependent-resource and workflow
> > > primitives are
> > > >> > > > > well-aligned. Fabric8 alone would push us toward
> > reimplementing
> > > them
> > > >> > > > > ourselves.
> > > >> > > > >
> > > >> > > > > The clearest signal is Strimzi. Strimzi predates JOSDK and was
> > > built
> > > >> > > > > directly on fabric8—but their newer components (Access
> > operator,
> > > >> > Schema
> > > >> > > > > Registry operator) use JOSDK. The most mature ASF
> > > distributed-system
> > > >> > > > > operator effectively says: if we were starting today, we'd use
> > > JOSDK.
> > > >> > > > >
> > > >> > > > > For the version floor, the constraints are JOSDK 5.x (Java
> > 17+)
> > > and
> > > >> > the
> > > >> > > > > Kubernetes APIs the operator uses — all GA well before 1.27:
> > > CRD v1
> > > >> > > (1.16),
> > > >> > > > > admission webhook v1 (1.16), PDB v1 (1.21), Lease v1 (1.14).
> > > >> > > > > The useful-but-optional StatefulSet
> > > >> > > persistentVolumeClaimRetentionPolicy
> > > >> > > > > was alpha in 1.23, beta in 1.27, GA in 1.32.
> > > >> > > > > Kubernetes upstream currently patches 1.31–1.33.
> > > >> > > > >
> > > >> > > > > *1.29 keeps us within a defensible distance of upstream while
> > > leaving
> > > >> > > room
> > > >> > > > > for users on enterprise distributions a release or two behind.
> > > 1.27
> > > >> > is
> > > >> > > more
> > > >> > > > > permissive; 1.30+ stricter.*Open to community input on what
> > > Fluss
> > > >> > users
> > > >> > > > > actually run.
> > > >> > > > >
> > > >> > > > > I'll fold both into the FIP once we converge.
> > > >> > > > >
> > > >> > > > > Thank you!
> > > >> > > > >
> > > >> > > > > On Wed, May 6, 2026 at 3:51 PM Michael Koepf <
> > > >> > [email protected]>
> > > >> > > > > wrote:
> > > >> > > > >
> > > >> > > > > > Hi Lorenzo,
> > > >> > > > > >
> > > >> > > > > > Thanks for the FIP. I believe a dedicated Fluss Kubernetes
> > > Operator
> > > >> > > will
> > > >> > > > > > further simplify deployment and operations in large-scale
> > > >> > production
> > > >> > > > > > environments.
> > > >> > > > > >
> > > >> > > > > > I skimmed over the FIP.
> > > >> > > > > >
> > > >> > > > > > > We propose introducing a Fluss Kubernetes Operator,
> > > implemented
> > > >> > in
> > > >> > > Java
> > > >> > > > > > [...]
> > > >> > > > > >
> > > >> > > > > > 1. The first question that comes to my mind; are there
> > already
> > > >> > > detailed
> > > >> > > > > > plans regarding implementation? Do you plan to use the Java
> > > >> > Operator
> > > >> > > SDK
> > > >> > > > > > framework (https://javaoperatorsdk.io/)? Or just the plain
> > > fabric8
> > > >> > > Java
> > > >> > > > > > Kubernetes client (
> > > https://github.com/fabric8io/kubernetes-client
> > > >> > )?
> > > >> > > > > >
> > > >> > > > > > 2. What will be the minimum supported Kubernetes version?
> > > >> > > > > >
> > > >> > > > > > Looking forward to this.
> > > >> > > > > >
> > > >> > > > > > Thanks.
> > > >> > > > > > --
> > > >> > > > > > Best,
> > > >> > > > > > Michael
> > > >> > > > > >
> > > >> > > > > > On 2026/05/05 09:05:45 Lorenzo Affetti via dev wrote:
> > > >> > > > > > > Hello community!
> > > >> > > > > > >
> > > >> > > > > > > I would like to start a discussion about FIP-41: Fluss
> > > Kubernetes
> > > >> > > > > > Operator.
> > > >> > > > > > > Here is the motivation:
> > > >> > > > > > >
> > > >> > > > > > > Fluss 0.8 introduced a Helm chart
> > > >> > > > > > > <https://github.com/apache/fluss/issues/779> that
> > > simplifies the
> > > >> > > initial
> > > >> > > > > > > deployment of a Fluss cluster on Kubernetes by packaging
> > > >> > manifests,
> > > >> > > > > > > configuration, and dependencies into a versioned release.
> > > While
> > > >> > > this is a
> > > >> > > > > > > good foundation, a Helm chart is fundamentally a one-shot
> > > >> > > templating
> > > >> > > > > > tool.
> > > >> > > > > > > It has no awareness of Fluss's runtime state and cannot
> > > react to
> > > >> > > > > > > operational events such as pod failures, rolling upgrades,
> > > or
> > > >> > > scale-in
> > > >> > > > > > > operations that risk data loss.
> > > >> > > > > > >
> > > >> > > > > > > Running Fluss in production on Kubernetes today requires
> > > users to
> > > >> > > > > > manually
> > > >> > > > > > > coordinate:
> > > >> > > > > > >
> > > >> > > > > > >    - Safe rolling restarts of TabletServers, ensuring
> > tablet
> > > >> > > leadership
> > > >> > > > > > is
> > > >> > > > > > >    migrated before each pod terminates (Fluss 0.8
> > > introduced the
> > > >> > > graceful
> > > >> > > > > > >    shutdown
> > > >> > > > > > >    <
> > > >> > > > > >
> > > >> > >
> > > https://fluss.apache.org/docs/maintenance/operations/graceful-shutdown/>
> > > >> > > > > > > primitive
> > > >> > > > > > >    that makes this possible, but does not orchestrate it
> > > across
> > > >> > > pods)
> > > >> > > > > > >    - Version upgrades that must sequence CoordinatorServer
> > > and
> > > >> > > > > > TabletServer
> > > >> > > > > > >    updates correctly
> > > >> > > > > > >    - Scale-in operations where a TabletServer must be
> > > drained of
> > > >> > > tablets
> > > >> > > > > > >    before its pod is terminated
> > > >> > > > > > >    - Recovery from partial failures (e.g., PVC
> > > reattachment, pod
> > > >> > > identity
> > > >> > > > > > >    preservation across restarts)
> > > >> > > > > > >    - Leveraging Fluss 0.8's dynamic configuration updates
> > > >> > > > > > >    <
> > > >> > > > > >
> > > >> > >
> > > https://fluss.apache.org/docs/maintenance/operations/updating-configs/>
> > > >> > > > > > for
> > > >> > > > > > >    keys that do not require restart, instead of
> > triggering a
> > > >> > > rolling
> > > >> > > > > > restart
> > > >> > > > > > >    for every config change
> > > >> > > > > > >
> > > >> > > > > > > This gap between deployment and operations is best
> > > addressed by a
> > > >> > > > > > *Kubernetes
> > > >> > > > > > > Operator* — a controller that continuously reconciles the
> > > desired
> > > >> > > state
> > > >> > > > > > of
> > > >> > > > > > > a FlussCluster resource against the actual state of the
> > > cluster,
> > > >> > > and
> > > >> > > > > > > executes Fluss-aware transitions safely.
> > > >> > > > > > >
> > > >> > > > > > > The Flink and Spark ecosystems have established this
> > pattern
> > > >> > > successfully
> > > >> > > > > > > with the Apache Flink Kubernetes Operator
> > > >> > > > > > > <https://github.com/apache/flink-kubernetes-operator> and
> > > the
> > > >> > > Apache
> > > >> > > > > > Spark
> > > >> > > > > > > Kubernetes Operator <
> > > >> > > https://github.com/apache/spark-kubernetes-operator
> > > >> > > > > > >.
> > > >> > > > > > > Fluss should follow suit.
> > > >> > > > > > >
> > > >> > > > > > > --
> > > >> > > > > > > Lorenzo Affetti
> > > >> > > > > > > Team Leader of Stream Storage
> > > >> > > > > > > [email protected]
> > > >> > > > > > > www.ververica.com
> > > >> > > > > > > ------------------------------
> > > >> > > > > > >
> > > >> > > > > > > <https://www.ververica.com/>
> > > >> > > > > > > Ververica GmbH | Herzogspitalstrasse 24 | 80331 München |
> > > Germany
> > > >> > > > > > >
> > > >> > > > > > > Follow us:
> > > >> > > > > > > <
> > > https://www.linkedin.com/company/ververica/posts/?feedView=all>
> > > >> > > > > > > <https://www.youtube.com/@ververica>
> > > >> > > > > > > <
> > > >> > > > > >
> > > >> > >
> > > https://open.spotify.com/show/2XME9h8iBOyr6YupqM99ir?si=87b064644add42a1
> > > >> > > > > > >Available
> > > >> > > > > > > on:  <
> > > >> > https://aws.amazon.com/marketplace/pp/prodview-luvmqd6leha4i
> > > >> > > >
> > > >> > > > > > > <
> > > >> > > > > >
> > > >> > >
> > > >> >
> > >
> > https://marketplace.microsoft.com/en-us/product/saas/ververica.vvc_managed?tab=Overview
> > > >> > > > > > >
> > > >> > > > > > >
> > > >> > > > > > > Pflichtangaben/Mandatory Information
> > > >> > > > > > > <https://www.ververica.com/mandatory-information>
> > > >> > > > > > >
> > > >> > > > > >
> > > >> > > > >
> > > >> > > > >
> > > >> > > > > --
> > > >> > > > > Lorenzo Affetti
> > > >> > > > > Senior Software Engineer @ Flink Team
> > > >> > > > > Ververica <http://www.ververica.com>
> > > >> > > >
> > > >> > >
> > > >> >
> > > >
> > > >
> > > >
> > > > --
> > > > Lorenzo Affetti
> > > > Senior Software Engineer @ Flink Team
> > > > Ververica
> > >
> >
> 

Reply via email to