[ 
https://issues.apache.org/jira/browse/HDDS-16432?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Siyao Meng reassigned HDDS-16432:
---------------------------------

    Assignee: Siyao Meng

> Formal verification for Ozone SCM HA Ratis follower catch-up and checkpoint 
> install with TLA+
> ---------------------------------------------------------------------------------------------
>
>                 Key: HDDS-16432
>                 URL: https://issues.apache.org/jira/browse/HDDS-16432
>             Project: Apache Ozone
>          Issue Type: Sub-task
>            Reporter: Siyao Meng
>            Assignee: Siyao Meng
>            Priority: Major
>
> Use TLA+ (via the Specula pipeline) to model and verify SCM HA follower catch 
> up: the inter SCM checkpoint download and install during an apply window, the 
> SCM HA transaction buffer flush versus durable index, and follower catch up 
> after leader transfer. Model check the specification and validate real SCM 
> traces against it. Bugs found by this effort are linked under this issue. 
> Adjacent to HDDS-16141 (SCM delete-block) and HDDS-16079 (bootstrap 
> checkpoint install).
> h3. Run environment
> {noformat}
> Ozone commit: 9fbf9ee0cb1bd2f5f5d437b6719ebbe5309351fb
> Specula:      v1.1.0 (commit c6aa3dfa)
> Agent/model:  claude-code, Claude Opus 4.8 (1M context)
> {noformat}
> h3. Coverage
> ||Target||Effort||Guidance||Result||
> |scm-ha-tx-buffer|medium|{{campaigns/ozone-9fbf9ee/targets/018-scm-ha-tx-buffer/.prompt-extra.md}}|1
>  reproduced Critical, 1 masked Medium|
> |scm-follower-catchup|high|{{campaigns/ozone-9fbf9ee/targets/019-scm-follower-catchup/.prompt-extra.md}}|1
>  masked Medium|
> h3. Reproduce
> {code:none}
> specula run --agent=claude-code --effort=<effort> --keep-original 
> --max-parallel=2 \
>   --enable-reviews --confirm-debate --tlc-memory-limit=28G 
> --tlc-worker-limit=8 \
>   "<target>|apache/ozone|Java|Use the target-specific .prompt-extra.md"
> {code}
> h3. Bugs (link via "Testing discovered")
> * (REPRODUCED, Critical) HDDS-16452: an InterSCMGrpcService.download 
> checkpoint served during an apply window ships data at index i under the 
> stale durable index i minus 1, so a follower installs it and durably holds 
> data ahead of its recorded index. [scm-ha-tx-buffer MC-2]
> * (MASKED, Medium) HDDS-16453: an unguarded transaction buffer flush commits 
> data under a stale index (DurableIndexMatchesData violated), masked by SCM 
> idempotent replay, with divergence risk the moment a non idempotent 
> replicated transaction is added. [scm-ha-tx-buffer MC-1]
> * (MASKED, Medium) HDDS-16454: a new SCM leader whose apply lags commit drops 
> reports for committed but unapplied containers as CONTAINER_NOT_FOUND, 
> leaving replica location accounting incomplete, masked by the WARN only 
> default and replication manager readiness gating. [scm-follower-catchup MC-2]
> Generated with Specula (Claude Opus 4.8).



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to