[ 
https://issues.apache.org/jira/browse/HDDS-16823?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Siyao Meng updated HDDS-16823:
------------------------------
    Attachment: CR-1-move-recover-lease-scm-call-out-of-apply.patch

> OMRecoverLeaseRequest calls SCM and generates block tokens in 
> validateAndUpdateCache, so applying RecoverLease depends on SCM
> -----------------------------------------------------------------------------------------------------------------------------
>
>                 Key: HDDS-16823
>                 URL: https://issues.apache.org/jira/browse/HDDS-16823
>             Project: Apache Ozone
>          Issue Type: Bug
>            Reporter: Siyao Meng
>            Priority: Critical
>         Attachments: CR-1-move-recover-lease-scm-call-out-of-apply.patch
>
>
> h3. Mechanism
> {{OMRecoverLeaseRequest.validateAndUpdateCache}} runs on the OM state machine 
> apply thread of every OM. Through {{doWork}} and {{updateBlockInfo}} it calls 
> {{getContainerWithPipeline}} on SCM, up to three times (last block of the 
> file, last two blocks of the open file), while holding the bucket write lock. 
> The pipelines, and the block tokens generated next to them, are only put into 
> the reply for the recovering client. They are never written to the OM DB.
> An {{IOException}} from the SCM call is not an {{OMException}}, so 
> {{OzoneManagerRatisUtils.exceptionToResponseStatus}} turns it into 
> {{INTERNAL_ERROR}}, which {{OzoneManagerStateMachine.processResponse}} treats 
> as an unrecoverable error of the OM. The request is a replicated log entry, 
> so every OM that applies it makes the same call. {{RecoverLease}} is the only 
> OM write request that calls SCM from {{validateAndUpdateCache}}. Block 
> allocation and upgrade finalization call it from {{preExecute}} on the leader.
> HDDS-10626 was an earlier OM shutdown from the same method (block token 
> generation while a {{RecoverLease}} entry was applied) and was fixed by 
> changing the startup order, which left both calls in the apply path.
> h3. Trigger
> A {{RecoverLease}} entry is applied while {{getContainerWithPipeline}} fails 
> with an exception that is not an {{OMException}}.
> h3. Impact
> * The entry is answered with {{INTERNAL_ERROR}} and the state machine 
> terminates the OM, on each OM that applies it. The status is shown by the 
> regression test below. What the state machine does with that status is from 
> reading {{processResponse}}.
> * Without any failure, each OM's single apply thread waits for up to three 
> SCM calls with the bucket write lock held, so no other write is applied on 
> that OM in the meantime.
> h3. Reproduction
> There is no standalone reproduction for this issue. The regression test in 
> the attached patch covers it: on unmodified source the apply level part of 
> {{TestOMRecoverLeaseRequest.testScmFailureDoesNotFailApply}} (a mocked SCM 
> client that throws {{IOException}}) fails with "expected: <OK> but was: 
> <INTERNAL_ERROR>".
> h3. Patch
> [^CR-1-move-recover-lease-scm-call-out-of-apply.patch], against 
> ea69b4a7d9abd040e5259dbcbb7c6ce52eb5d199. It also applies to master at 
> 84f5594a6a2.
> {{validateAndUpdateCache}} no longer talks to SCM or generates block tokens. 
> The leader adds both to the reply in 
> {{OzoneManagerProtocolServerSideTranslatorPB.internalProcessRequest}}, after 
> the request has been applied and outside any OM lock, in the same place where 
> the S3 derived key is already added to a {{CreateKey}} reply. If the SCM call 
> fails there, the caller gets {{SCM_GET_PIPELINE_EXCEPTION}} (the code 
> {{KeyManagerImpl}} already uses for a failed pipeline refresh). The file is 
> then already marked as under recovery, and a later {{recoverLease}} call 
> continues from there, as a second recovery of the same file does today.
> There is no wire or protocol change, and entries already in the log are 
> applied as before, without the SCM call. Two side effects. The block token is 
> now issued for the calling user, as it is for block allocation, instead of 
> the OM's own user (the apply thread has no RPC caller). And the SCM call 
> still blocks its caller for as long as the SCM client takes, but on the RPC 
> handler thread of the leader only, not on the apply thread and not under the 
> bucket lock.
> Covered by {{testPipelinesAreAddedAfterApply}} and 
> {{testScmFailureDoesNotFailApply}} in the existing 
> {{TestOMRecoverLeaseRequest}}. With the patch {{TestOMRecoverLeaseRequest}}, 
> {{TestOpenKeyCleanupService}} and {{TestOMKeyCommitRequestWithFSO}} (48 
> tests) pass and checkstyle is clean. On a mini cluster {{TestHSync}} (42), 
> {{TestSecureOzoneRpcClient}} (241, block tokens enabled) and 
> {{TestOFSWithFSO}} (73, 9 skipped) pass.
> Found by code review of the lease recovery paths for hsynced files, as part 
> of the TLA+ verification effort under HDDS-15926, on commit 
> ea69b4a7d9abd040e5259dbcbb7c6ce52eb5d199. Checked against HDDS issues and 
> apache/ozone pull requests for duplicates before filing. Related but not 
> duplicates: HDDS-10626 (earlier shutdown from the same method), HDDS-16767 
> (wall clock read while the same request is applied). The attached patch is a 
> proposal for review. Generated with Specula (Claude Opus 5.5).



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to