linliu-code commented on issue #20057: URL: https://github.com/apache/hudi/issues/20057#issuecomment-5860012827
Linking prior art I should have found before filing this: #16359 describes the same underlying behaviour — Spark caches the resolved relation, so the same `HoodieFileIndex` instance is reused for the life of the session — and it names `cachedAllPartitionPaths` and `cachedAllInputFileSlices` directly. The two are not duplicates: #16359 is about the read consequence (a query-only session never seeing later commits), this issue is about the memory consequence (those maps grow with every partition touched and are never evicted). But they share a root cause, and a fix for one is likely to constrain the design of the other, so they are probably best considered together. Also related: #20104, where the same reused-then-rebuilt index causes Spark's CacheManager to strand `CACHE TABLE` entries. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
