Siyao Meng created HDDS-16463:
---------------------------------
Summary: Recon power loss losing an un synced derived write
permanently drops an applied update
Key: HDDS-16463
URL: https://issues.apache.org/jira/browse/HDDS-16463
Project: Apache Ozone
Issue Type: Bug
Reporter: Siyao Meng
h3. Finding
A power-loss crash that loses the un-synced derived-store write while the Derby
cursor survives makes startup reconciliation skip reprocess, permanently
dropping an applied update so every Recon derived-table query (container-key
API/UI) serves stale/incomplete data with no automatic recovery.
Production-reachable via the normal delta-sync path plus a differential
power-loss within Ozone's fault model. ENV_LIMITED: the environment limit is
the absence of a crash-consistency harness to truncate one RocksDB WAL tail
relative to another.
h3. Classification
* Verdict: ENV_LIMITED
* Severity: Critical
* Source: Specula TLA+ model checking and confirmation debate, finding MC-1
h3. Reproduce
{noformat}
Ozone commit: 9fbf9ee0cb1bd2f5f5d437b6719ebbe5309351fb
Specula: v1.1.0 (commit c6aa3dfa)
Target: recon-om-sync
Guidance: campaigns/ozone-9fbf9ee/targets/023-recon-om-sync/.prompt-extra.md
{noformat}
{code:none}
specula run --agent=claude-code --effort=medium --keep-original
--max-parallel=2 \
--enable-reviews --confirm-debate --tlc-memory-limit=28G --tlc-worker-limit=8
\
"recon-om-sync|apache/ozone|Java|Use the target-specific .prompt-extra.md"
{code}
Discovered under HDDS-16434 (Specula TLA+ verification effort). The TLA+
specification, counterexample, and confirmation debate live in the Specula run
artifacts.
Generated with Specula (Claude Opus 4.8).
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]