jsedding commented on code in PR #3137:
URL: https://github.com/apache/jackrabbit-oak/pull/3137#discussion_r4091866797
##########
oak-segment-tar/src/main/java/org/apache/jackrabbit/oak/segment/file/GCJournal.java:
##########
@@ -68,7 +68,15 @@ public synchronized void persist(long reclaimedSize, long
repoSize,
@NotNull GCGeneration gcGeneration, long nodes, @NotNull String
root
) {
GCJournalEntry current = read();
- if (current.getGcGeneration().equals(gcGeneration)) {
+ GCGeneration currentGeneration = current.getGcGeneration();
+ // Compare generation and fullGeneration only, not isCompacted: the
on-disk journal
+ // format never persists isCompacted (GCJournalEntry.fromString always
deserializes it
+ // as false), so comparing the full GCGeneration would wrongly treat
every persist() of
+ // an already-journaled, still-compacted generation as new after a
restart (e.g. a
+ // standalone cleanup() on a store re-opened with an already-compacted
head), causing a
+ // duplicate entry to be written on every such restart.
+ if (currentGeneration.getGeneration() == gcGeneration.getGeneration()
+ && currentGeneration.getFullGeneration() ==
gcGeneration.getFullGeneration()) {
Review Comment:
I am wondering if the "more correct fix" would be to have
`GCJournalEntry.fromString` always return a `GCGeneration` with
`isCompacted=true`?
I have very little knowledge about the cleanup and GCJournal parts. But it
seems to me that we are logging successful compactions, and thus, intuitively,
it would make sense to assume that the GCJournal contains only compacted
generations.
Furthermore, I assume that `GCJournal.persist` is only ever called for
compacted `GCGenerations`.
Thoughts?
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]