Hi Jark,

Thanks a lot for your valuable suggestion! I will build and test the latest
main branch
following the official guide to verify whether these stability issues are
resolved. I will
report any new findings to the community in time.

Best regards,

Zuo Wei
Email: [email protected]

Jark Wu <[email protected]> 于2026年7月24日周五 23:03写道:

> Hi Zuo,
>
> Thank you for the detailed feedback. Results from real-world testing
> are very valuable to the Fluss community.
>
> Fluss 0.9.1 was released nearly three months ago, so it does not
> include fixes for some issues discovered since then. Since you are
> currently testing Fluss, I recommend building the latest main branch
> and checking whether these issues can still be reproduced:
>
> https://fluss.apache.org/community/dev/building/
>
> Testing the main branch now can help uncover remaining issues early
> and give the community time to address them before 1.0 is finalized.
> This will help ensure that the official 1.0 release can be reliably
> deployed in production environments such as Xiaomi’s.
>
> Thanks again for your support!
>
> Best regards,
> Jark
>
> On Fri, 24 Jul 2026 at 15:43, zuo wei <[email protected]> wrote:
> >
> > Hi Giannis,
> >
> > Thanks a lot for your clear and detailed reply! It helps me fully
> > understand
> > the root causes of the stability issues we met in Fluss 0.9.1. I will
> apply
> > your
> > recommended solutions on our 0.9.1 test cluster immediately. I'm looking
> > forward to the 1.0 release with the related PR fixes.
> >
> > I really appreciate the community’s efficient problem solving and
> technical
> > support!
> > During the follow-up testing and optimization process, if I encounter any
> > new problems or have relevant verification results, I will actively
> > synchronize
> > with the community in a timely manner.Thanks again for your great help!
> >
> > Best regards,
> >
> > Zuo Wei
> > Email: [email protected]
> >
> >
> >
> > Giannis Polyzos <[email protected]> 于2026年7月24日周五 13:52写道:
> >
> > > Hi Zuo,
> > >
> > > Thank you for the detailed report. Let me see if I can provide a
> breakdown
> > > here.
> > >
> > > 1. NotifyLeaderAndIsrRequest and UpdateMetadataRequest from the
> Coordinator
> > > are both handled directly on the Netty/RPC worker thread pool in
> > > TabletService. The becomeLeaderOrFollower path inside those handlers
> can be
> > > slow (opening RocksDB, replaying logs, updating metadata cache), which
> > > causes the worker thread to block. Because the worker pool is shared,
> all
> > > concurrent RPC requests, including subsequent metadata updates, queue
> > > behind it. Under leader-election storms (e.g., rolling restarts,
> failover),
> > > this can stall the entire tablet server.
> > > This PR should help with this:
> https://github.com/apache/fluss/pull/3647
> > > along with: https://github.com/apache/fluss/pull/3654 .. Until a new
> > > release is cut, maybe increasing *rpc.worker.threads* helps with this
> > > issue.
> > >
> > > 2. This is most likely a metadata propagation timing issue. When
> > > auto-partition creates a new partition in ZooKeeper, there is a window
> > > before the TabletServer's in-memory metadata cache receives the
> > > UpdateMetadata from the Coordinator. If the client resolves the
> partition
> > > and attempts to write during that window, the client's retry logic
> does not
> > > always recover cleanly. This PR probably helps with this:
> > > https://github.com/apache/fluss/pull/3511 .
> > > Some other configuration options that might help:
> > > - Increase *auto-partition.num-precreate* so partitions are pre-created
> > > ahead of time, reducing the race window.
> > > - Add retry logic in your producer application, for example, a
> transient
> > > PartitionNotExistException during auto-partition creation should be
> > > retried; the partition will eventually appear.
> > > - Ensure your auto-partition *time-interval *is not too aggressive
> relative
> > > to the time it takes for metadata to propagate.
> > >
> > > 3. When a TabletServer restarts abnormally, the Coordinator triggers
> leader
> > > elections for all KV buckets that were on that node. Each
> > > NotifyLeaderAndIsrRequest results in becomeLeaderOrFollower → RocksDB
> > > open() call (which acquires the LOCK file, reads SST files, replays
> WAL).
> > > This is inherently slow and happens on the RPC worker thread in 0.9.1,
> > > serializing all subsequent RPCs. PR #3654 should probably help with
> this
> > > issue as well, and PR https://github.com/apache/fluss/pull/3651
> further
> > > reduces shutdown time so restarts are faster overall.
> > >
> > > So I believe the fixes are most likely there and targeted for the
> upcoming
> > > 1.0 release. So if your timeline allows, the safest path is to for the
> 1.0
> > > release. PRs #3647, #3654, #3511, and #3651 will all be included.
> > >
> > > The goal is to release within the next month.
> > >
> > > Let me know if the above helps.
> > >
> > > Best,
> > > Giannis
> > >
> > > On Thu, Jul 23, 2026 at 3:22 PM zuo wei <[email protected]> wrote:
> > >
> > > > Dear Apache Fluss Community Developers,
> > > >
> > > > I’m a R&D engineer from Xiaomi. We’re currently testing *Apache Fluss
> > > > 0.9.1*
> > > > for our internal business scenarios, and ran into several severe
> unstable
> > > > issues during our tests. I’d like to share these problems and ask for
> > > your
> > > > guidance.
> > > >
> > > > Here are the some main issues we encountered:
> > > > *1. RPC stuck issue*: RPC requests (e.g. update metadata request)
> from
> > > > Coordinator Server to Tablet Server often get stuck for a long time,
> > > > causing service blocking.
> > > > *2. Partition not exist error*: Data writing throws
> partition-not-exist
> > > > exceptions even when the target partitions exist normally.
> > > > *3. RocksDB lock blocking issue*: After abnormal restart of Tablet
> Server
> > > > nodes, a lot of KV buckets leader switching occurs. The RocksDB open
> > > > process holds locks for a long time, seriously blocking other worker
> > > > threads.
> > > >
> > > > I'm not sure whether these problems are caused by our improper usage,
> > > > configuration mistakes, or existing known bugs of Fluss 0.9.1. And
> I’d
> > > like
> > > > to consult the community whether Fluss 0.9.1 currently is
> recommended for
> > > > production environment deployment?
> > > > We really want to adopt Fluss in our business environment. Thanks a
> lot
> > > for
> > > > your help! Looking forward to your reply.
> > > >
> > > > Best regards,
> > > >
> > > > Zuo Wei
> > > > Email: [email protected]
> > > >
> > >
>

Reply via email to