Xiaobing Fang created FLINK-40966:
-------------------------------------
Summary: Reuse Fluss metadata clients across schema operations
Key: FLINK-40966
URL: https://issues.apache.org/jira/browse/FLINK-40966
Project: Flink
Issue Type: Improvement
Components: Flink CDC
Reporter: Xiaobing Fang
h2. Description
{\{FlussMetaDataApplier}} creates and closes a \{{Connection}} and \{{Admin}}
for every \{{CreateTable}}, \{{AddColumn}} and \{{DropTable}} operation, and
for every existing-table schema lookup. Closing the connection waits for the
Netty event loop to shut down. With the default quiet period, this adds roughly
two seconds per operation even when the metadata RPC itself is fast.
Schema changes are processed serially. As a result, initializing many tables
accumulates this shutdown cost. Existing-table expansion adds a separate schema
lookup before \{{CreateTable}} and can incur the cost twice per table, even
when no columns need to be added.
h2. Steps to reproduce
# Start a pipeline with a Fluss sink and multiple tables.
# Observe the intervals between schema operations, or capture the schema worker
stack during initialization.
# The worker repeatedly waits in \{{FlussConnection.close}} /
\{{NettyClient.close}}.
h2. Expected behavior
Reuse one \{{Connection}}/\{{Admin}} pair per metadata applier and close it
when the coordinator is disposed. Schema operations should not repeatedly shut
down the network client.
h2. Suggested fix
Lazily initialize transient clients, reuse them for schema lookups and DDL, and
close them together. Stop schema workers before releasing the metadata applier
so shutdown cannot race with an in-flight operation. Preserve client isolation
and serialization behavior.
h2. Description
{\{FlussMetaDataApplier}} creates and closes a \{{Connection}} and \{{Admin}}
for every \{{CreateTable}}, \{{AddColumn}} and \{{DropTable}} operation, and
for every existing-table schema lookup. Closing the connection waits for the
Netty event loop to shut down. With the default quiet period, this adds roughly
two seconds per operation even when the metadata RPC itself is fast.
Schema changes are processed serially. As a result, initializing many tables
accumulates this shutdown cost. Existing-table expansion adds a separate schema
lookup before \{{CreateTable}} and can incur the cost twice per table, even
when no columns need to be added.
h2. Steps to reproduce
# Start a pipeline with a Fluss sink and multiple tables.
# Observe the intervals between schema operations, or capture the schema worker
stack during initialization.
# The worker repeatedly waits in \{{FlussConnection.close}} /
\{{NettyClient.close}}.
h2. Expected behavior
Reuse one \{{Connection}}/\{{Admin}} pair per metadata applier and close it
when the coordinator is disposed. Schema operations should not repeatedly shut
down the network client.
h2. Suggested fix
Lazily initialize transient clients, reuse them for schema lookups and DDL, and
close them together. Stop schema workers before releasing the metadata applier
so shutdown cannot race with an in-flight operation. Preserve client isolation
and serialization behavior.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)