Shivam-Agg commented on code in PR #3137:
URL: https://github.com/apache/jackrabbit-oak/pull/3137#discussion_r4092641367
##########
oak-segment-tar/src/main/java/org/apache/jackrabbit/oak/segment/file/GCJournal.java:
##########
@@ -68,7 +68,15 @@ public synchronized void persist(long reclaimedSize, long
repoSize,
@NotNull GCGeneration gcGeneration, long nodes, @NotNull String
root
) {
GCJournalEntry current = read();
- if (current.getGcGeneration().equals(gcGeneration)) {
+ GCGeneration currentGeneration = current.getGcGeneration();
+ // Compare generation and fullGeneration only, not isCompacted: the
on-disk journal
+ // format never persists isCompacted (GCJournalEntry.fromString always
deserializes it
+ // as false), so comparing the full GCGeneration would wrongly treat
every persist() of
+ // an already-journaled, still-compacted generation as new after a
restart (e.g. a
+ // standalone cleanup() on a store re-opened with an already-compacted
head), causing a
+ // duplicate entry to be written on every such restart.
+ if (currentGeneration.getGeneration() == gcGeneration.getGeneration()
+ && currentGeneration.getFullGeneration() ==
gcGeneration.getFullGeneration()) {
Review Comment:
> I am wondering if the "more correct fix" would be to have
GCJournalEntry.fromString always return a GCGeneration with isCompacted=true?
This looks promising and simpler to me but I am not very confident with the
regression it may cause with other consumers.
And provided we call `GCJournal.persist` for compacted generations only,
then `isCompacted()` check is anyways redundant and safe to remove while
comparing.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]