[
https://issues.apache.org/jira/browse/SPARK-59032?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated SPARK-59032:
-----------------------------------
Labels: pull-request-available (was: )
> Spark Connect client can hang on a wedged stream: gRPC keepalive is parsed
> but never applied
> --------------------------------------------------------------------------------------------
>
> Key: SPARK-59032
> URL: https://issues.apache.org/jira/browse/SPARK-59032
> Project: Spark
> Issue Type: Bug
> Components: Project Infra
> Affects Versions: 5.0.0
> Reporter: Hyukjin Kwon
> Priority: Major
> Labels: pull-request-available
>
> The Python Connect parity gate intermittently hangs on
> test_parity_arrow_python_udf.py: the file times out with p=0 (not one test
> completed) and is killed only by the per-file cap. It was observed once on
> master at file 5/84 on a fresh server, while all 83 other files (including
> the sibling arrow/UDF files) passed.
> It is NOT deterministic and NOT a slowness/timeout problem:
> - the file passes standalone in ~122s (matching the reference client);
> - it passes even with the suite's reattach-forcing server limits
> (senderMaxStreamDuration=1s, senderMaxStreamSize=123);
> - it did NOT reproduce in 8 back-to-back runs under the exact gate invocation;
> - a temporary CI diagnostic (client py-spy + faulthandler, server jstack,
> Python-worker stacks) found no deterministic hang.
> Root cause: the client never applies gRPC keepalive to its channel. The
> connection string's grpc_keepalive_* options (default: enabled, 60s interval
> / 20s timeout, keep-alive-while-idle) are parsed by ChannelBuilder but were
> never applied to the tonic channel, and no HTTP/2 keepalive was set
> otherwise. So a server-streaming RPC whose connection wedges (a half-open
> TCP, or a server-side stall) blocks in stream.message() indefinitely: nothing
> pings the peer to surface the break, the reattachable-execute iterator never
> gets the transport error it needs to reattach, and the stream only unblocks
> when the client is closed (at CI teardown).
> Fix (PR: https://github.com/apache/spark-connect-rust/pull/58):
> - Apply the parsed keepalive settings to the tonic endpoint (TLS and
> plaintext paths), so an unacked PING within the timeout fails the stream with
> a transport error that the reattachable iterator treats as retriable and
> recovers from.
> - Retry the file in the parity gate (FLAKY_FILES) as a belt-and-suspenders
> safety net for any residual rarity, instead of skipping the whole file or
> letting a rare hang burn the cap and fail.
> - Reverted the earlier, wrong senderMaxStreamDuration=1200s server-conf
> attempt (a larger cap only makes a hang wait longer), and removed the
> temporary diagnostic workflow.
> --- Update (review of PR #58) ---
> The keepalive fix (applying the parsed grpc_keepalive_* to the channel) makes
> a wedged transport RECOVERABLE — an expired PING surfaces as UNAVAILABLE,
> which the reattachable-execute iterator retries — but it is NOT confirmed to
> be the root cause. The captured evidence (509b9e6) shows the client blocked
> on the very FIRST next() of the reattachable stream while the server produced
> nothing at all, whereas the reference client gets the error in <1s. Keepalive
> explains a wedged transport; it does not by itself explain "server produced
> nothing." So this ticket stays OPEN for the true root cause; keepalive + the
> FLAKY_FILES retry make the gate green by making the symptom recoverable
> without necessarily fixing the underlying cause. A server-side jstack was
> added to the parity gate's timeout path to pin down whether the server's
> execute thread is deadlocked on the (dead) Python worker or blocked writing
> to us. A separate, related "documented-but-unapplied option" bug (gRPC max
> message length) is tracked as SPARK-59037.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]