DanielLeens commented on issue #12339: URL: https://github.com/apache/seatunnel/issues/12339#issuecomment-5760331453
Classification: D / Zeta checkpoint-scheduler design review. Thank you for correcting the utilization conclusion and for separating the measured saturation result from the unproven production scenario. The benchmark establishes a useful result: the shared path is flat through the measured scale, but an overloaded bounded dispatcher turns its unbounded queue into potentially unbounded trigger delay. I also re-checked the relevant code. On current `dev`, `startSavepoint()` holds the coordinator `lock` while it waits in its 500 ms sleep loop; #12165 sends periodic `tryTriggerPendingCheckpoint()` work through the member-wide bounded dispatcher before that work acquires the same lock. Therefore the cross-pipeline coupling you describe is a valid source-based design concern, even though it is not yet a reproduced incident. It means the claim that dispatcher bodies cannot block is not a sufficient merge argument by itself. Decision for this STIP: do not add a scheduler-thread-count option now. Keep the proposal in review, and before #12165 is accepted, make the dispatcher contract explicit and prove it: shared dispatcher workers must not remain blocked waiting for a coordinator lock while a savepoint is pending. A non-blocking lock/reschedule design or moving the savepoint wait outside that lock are possible directions; the selected direction needs a regression that fills the dispatch capacity with savepoint-contended coordinators and demonstrates that an unrelated pipeline still receives its trigger/watchdog work. Scheduling-delay observability should be prioritized ahead of any future tuning option. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
