mrhhsg opened a new pull request, #68610:
URL: https://github.com/apache/doris/pull/68610

   ### What problem does this PR solve?
   
   Issue Number: None
   
   Related PR: #51690, #58144, #58273, #65814, #67070
   
   Problem Summary: Workload groups pick the scan scheduler for internal and
   external table scans from `enable_task_executor_in_internal_table` and
   `enable_task_executor_in_external_table`. Both configs currently default to
   `true` on master, so every workload group creates
   `TaskExecutorSimplifiedScanScheduler` (the time-sharing task executor) for
   local (`ls_*`) and remote (`rs_*`) scans. The time-sharing executor is still
   maturing (for example the self-deadlock fixed in #58273; the internal-table
   default was already reverted once in #58144), so the default should be the
   `ThreadPoolSimplifiedScanScheduler` for both internal and external tables.
   
   This change sets both configs' default values to `false`. The task executor
   can still be enabled explicitly through `be.conf`.
   
   Making the thread pool path the default exposed one scheduling gap between
   the two admission paths. #65814 stops the TaskExecutor path
   (`ScannerContext::_pull_next_scan_task`) from submitting pending scanners
   once the shared scan LIMIT is exhausted while a completed or in-flight task
   can still wake the operator. The thread pool admission rewritten in #67070
   (`ScannerContext::can_admit_scan_task`) does not apply this rule, so for a
   `SELECT ... LIMIT n` over many tablets every remaining pending scanner would
   still be admitted and run `prepare()`/`open()` (tablet reader init) only to
   report EOS right away. The same rule is now applied in
   `can_admit_scan_task`, including the existing escape that admits one pending
   scanner when nothing else can wake the pipeline task.
   
   ### Release note
   
   The BE configs `enable_task_executor_in_internal_table` and
   `enable_task_executor_in_external_table` now default to `false`; internal
   and external table scans use the thread pool scan scheduler by default.
   
   ### Check List (For Author)
   
   - Test:
       - Unit Test: added 
`WorkloadGroupManagerTest.DefaultScanSchedulersUseThreadPool`
         (registered defaults are `false` and the workload group creates
         thread pool schedulers for local and remote scans) and
         `ScannerContextTest.thread_pool_admission_stops_after_shared_limit`
         (fails without the admission fix); ran `WorkloadGroupManagerTest.*`,
         `ScannerContextTest.*` and other scanner/scheduler/task-executor UTs
       - Regression test: ran existing suites on a local cluster with the new
         default: correctness_p0/test_shared_scan_limit_pending_tasks,
         query_p0 limit pushdown suites, query_p0/scan_range, and stream load
         CSV suites
   - Behavior changed: Yes (default scan scheduler for internal and external
     table scans changes from the task executor to the thread pool; it can be
     switched back via the two BE configs)
   - Does this need documentation: No
   
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to