yunfengzhou-hub opened a new issue, #1213: URL: https://github.com/apache/flink-agents/issues/1213
### Search before asking - [x] I searched in the [issues](https://github.com/apache/flink-agents/issues) and found nothing similar. ### Description Calling an internal sub-agent makes the parent block in `awaitSubagentCall(...)` on `responseFuture.get()` (`runtime/.../subagent/InternalSubagentSetup.java`). That wait runs inside a `DurableCallable` dispatched off the mailbox thread onto the operator's shared async worker pool, whose size is `num-async-threads` (`runtime/.../operator/ActionExecutionOperator.java` → `ContinuationActionExecutor`). A waiting parent therefore holds one pool worker for the whole child call, while the child's own durable async operations are dispatched onto the same finite pool. Expected: awaiting a sub-agent should suspend without holding a worker its child needs, so nested or concurrent sub-agent calls make progress under any `num-async-threads >= 1`. Actual: with `num-async-threads=1`, a single parent awaiting a child that performs any async call can deadlock — the parent holds the only worker while blocking on the child, and the child can never obtain a worker to progress. A larger pool only raises the threshold; enough concurrent or nested awaits reach the same self-deadlock, and nesting (a child awaiting its own sub-agent) multiplies the workers consumed by pure waiting. Raising `num-async-threads` does not remove the dependency. The fix should let the parent's await yield its worker (continuation / async-completion style) and resume on the child's result future, without breaking the mailbox suspension the wait already requires (`requireMailboxSuspension`) or durable-call replay/recovery of the awaited result. Split out from #1137, which lists this under "Semantics to define" ("a wait must not hold a resource the child needs to make progress — otherwise even a single nested call can deadlock"); raised by @wenjin272 while reviewing #1138. ### How to reproduce On the #1138 branch (internal sub-agent support), set `num-async-threads=1`: 1. A root agent action submits to an internal sub-agent (registered as an `AGENT` resource) and awaits its result. 2. The child performs any durable async operation (e.g. an async chat-model or tool call) before emitting its OutputEvent. 3. The parent's await occupies the single async worker; the child's async operation queues behind it for a worker that never frees. The job then hangs with no progress and no failure — a self-deadlock between the waiting parent and the worker-starved child. With a larger `num-async-threads`, proportionally more concurrent or nested awaits trigger the same condition. ### Version and environment `subagent-framework-internal` (PR #1138) on top of `main`. JDK 21 (continuation-based suspension). Flink 2.x. The deadlock is in the Java runtime async pool and applies to both Java- and Python-authored children. ### Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
