Siyao Meng created HDDS-16463:
---------------------------------

             Summary: Recon power loss losing an un synced derived write 
permanently drops an applied update
                 Key: HDDS-16463
                 URL: https://issues.apache.org/jira/browse/HDDS-16463
             Project: Apache Ozone
          Issue Type: Bug
            Reporter: Siyao Meng


h3. Finding
A power-loss crash that loses the un-synced derived-store write while the Derby 
cursor survives makes startup reconciliation skip reprocess, permanently 
dropping an applied update so every Recon derived-table query (container-key 
API/UI) serves stale/incomplete data with no automatic recovery. 
Production-reachable via the normal delta-sync path plus a differential 
power-loss within Ozone's fault model. ENV_LIMITED: the environment limit is 
the absence of a crash-consistency harness to truncate one RocksDB WAL tail 
relative to another.

h3. Classification
* Verdict: ENV_LIMITED
* Severity: Critical
* Source: Specula TLA+ model checking and confirmation debate, finding MC-1

h3. Reproduce
{noformat}
Ozone commit: 9fbf9ee0cb1bd2f5f5d437b6719ebbe5309351fb
Specula:      v1.1.0 (commit c6aa3dfa)
Target:       recon-om-sync
Guidance:     campaigns/ozone-9fbf9ee/targets/023-recon-om-sync/.prompt-extra.md
{noformat}
{code:none}
specula run --agent=claude-code --effort=medium --keep-original 
--max-parallel=2 \
  --enable-reviews --confirm-debate --tlc-memory-limit=28G --tlc-worker-limit=8 
\
  "recon-om-sync|apache/ozone|Java|Use the target-specific .prompt-extra.md"
{code}
Discovered under HDDS-16434 (Specula TLA+ verification effort). The TLA+ 
specification, counterexample, and confirmation debate live in the Specula run 
artifacts.

Generated with Specula (Claude Opus 4.8).



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to