linliu-code commented on issue #20057:
URL: https://github.com/apache/hudi/issues/20057#issuecomment-5860012827

   Linking prior art I should have found before filing this: #16359 describes 
the same underlying behaviour — Spark caches the resolved relation, so the same 
`HoodieFileIndex` instance is reused for the life of the session — and it names 
`cachedAllPartitionPaths` and `cachedAllInputFileSlices` directly.
   
   The two are not duplicates: #16359 is about the read consequence (a 
query-only session never seeing later commits), this issue is about the memory 
consequence (those maps grow with every partition touched and are never 
evicted). But they share a root cause, and a fix for one is likely to constrain 
the design of the other, so they are probably best considered together.
   
   Also related: #20104, where the same reused-then-rebuilt index causes 
Spark's CacheManager to strand `CACHE TABLE` entries.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to