u70b3 commented on issue #66497:
URL: https://github.com/apache/doris/issues/66497#issuecomment-6054343315

   > After reconsidering the implementation against the execution model of 
Iceberg `rewrite_data_files`, please simplify the design and implementation for 
this issue. This changes the previously accepted v5.1 direction: a durable 
asynchronous Lance index job framework should no longer be a requirement for 
the initial implementation.
   > 
   > ### 1. Use synchronous SQL with bounded worker execution
   > The SQL statement should validate the request, dispatch resource-isolated 
work to BE, wait for completion, refresh metadata after a confirmed commit, and 
return the result. Internal parallel execution does not require an asynchronous 
SQL/job interface.
   > 
   > Keep heavy Lance work outside FE and the BE main process, with hard worker 
resource limits, bounded concurrency/queueing, deadlines, and reliable child 
cleanup. These protections are independent of durable job management and should 
remain.
   > 
   > ### 2. Revert the job infrastructure already introduced
   > Please prepare an explicit cleanup/revert on `branch-4.1`, rather than 
leaving the old framework disabled or maintaining two execution paths:
   > 
   > * Revert the durable Lance index job state machine, manager, persistent 
same-name fences, unresolved-job quotas, and job replay infrastructure 
introduced by [branch-4.1: [feature](lance) Durable Lance index job 
infrastructure with fence, quota and replay 
#67235](https://github.com/apache/doris/pull/67235).
   > * Remove the job-specific parts of [branch-4.1: [feature](lance) Add Lance 
index admission, job inspection SQL and catalog DDL guard 
#67630](https://github.com/apache/doris/pull/67630): durable job creation and 
JobId results, `SHOW LANCE INDEX JOB(S)`, and catalog guards tied to unresolved 
jobs. Preserve and adapt the useful privilege, schema, parameter, snapshot, and 
IF-condition validation for synchronous execution.
   > * Remove the associated unused configuration, serialization/RPC fields, 
tests, and documentation as appropriate. Review journal/image compatibility 
before removing persisted readers or operation codes; do not make existing 
metadata unreadable or reuse old codes. Any necessary compatibility handling 
should be minimal and should not preserve an active job framework.
   > * Stop pursuing [branch-4.1: [feature](lance) Add Lance index job 
dispatcher with thrift dispatch boundary 
#67978](https://github.com/apache/doris/pull/67978) and [branch-4.1: 
[feature](lance) Add RESOLVE LANCE INDEX JOB (FORCE_RELEASE) and job retention 
GC #67754](https://github.com/apache/doris/pull/67754) in their current form. 
Rework [branch-4.1: [feature](lance) Add the cgroup-isolated one-shot worker 
for Lance index mutations 
#68668](https://github.com/apache/doris/pull/68668)/[branch-4.1: [test](lance) 
Prove the IVF_PQ tracer and worker-fault convergence through the isolated 
worker #68669](https://github.com/apache/doris/pull/68669) around the 
synchronous worker lifecycle and its failure semantics, retaining useful 
isolation and query-consumption tests.
   > 
   > Keep the read-only `SHOW INDEX` and physical-index inspection 
functionality, and the reusable target-aware DDL validation. Please update the 
issue design and PR roadmap to reflect this scope change.
   > 
   > ### 3. State the simpler failure contract explicitly
   > The initial version does not promise persistent job history, FE failover 
recovery, resumable builds, or durable same-name serialization across failures.
   > 
   > A timeout, disconnect, worker loss, or lost response after dispatch must 
not be presented as proof that nothing committed. Return an explicit 
indeterminate-outcome error when applicable, do not automatically retry a 
potentially committed mutation, and let the user inspect authoritative index 
metadata. Observing a same-name index does not prove which request created it.
   > 
   > Cancellation cannot undo an existing commit. Worker capacity must remain 
accounted for until termination is confirmed. A confirmed commit followed by 
metadata-refresh failure must be reported distinctly from a failed build.
   > 
   > ### 4. Keep a path to distributed construction
   > Synchronous SQL does not preclude distributed index construction. The 
desired extension is:
   > 
   > `pin snapshot -> partition fragments -> parallel BE builds of uncommitted 
segments -> validate/collect outputs -> one coordinated final commit -> refresh 
-> return`
   > 
   > Please keep build and commit responsibilities separable. Verify the APIs 
available in the pinned Lance bindings and add the necessary adapter support 
before implementing this distributed path; do not run several complete, 
independently committing `create_index` calls against the same logical index 
name.
   > 
   > The initial implementation may use one worker. Multi-BE construction can 
follow without introducing durable jobs; persistent recovery should be a 
separate, justified follow-up requirement rather than a prerequisite for index 
lifecycle support.
   
    Agreed — accepted, I'll rework this to the synchronous model.
   
   The synchronous shape was my starting point when drafting the original 
design, simply because it is easier to implement and verify. The durable async 
framework came in only to keep post-dispatch ambiguity safe across failover, 
and its complexity — fences, quotas, replay, force-release — has outgrown that 
value. The synchronous path is also a much smaller diff to review and leaves 
fewer corners for hard-to-debug failure modes.
   
   Concretely:
   
   - revert the #67235 job infrastructure and the job-specific parts of #67630 
on branch-4.1, keeping the privilege/schema/parameter/snapshot/IF validation, 
SHOW INDEX, and the inspection TVF; journal/image compatibility reviewed before 
any persisted reader or op code is removed
   - close the open dispatcher and force-release PRs, and rework the worker PR 
around the synchronous lifecycle and its failure semantics
   - update the design and roadmap with the explicit failure contract, and keep 
build and commit responsibilities separable for the later distributed path
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to