yuqi1129 opened a new issue, #13580:
URL: https://github.com/apache/gravitino/issues/13580

   ## Version
   
   main branch, `ab0785324c5cb391ae63e4501d349f4336b8d3aa`; Java 17.0.8.
   
   ## Describe what's wrong
   
   The unmodified main cache implementation deadlocks under insertion/eviction 
pressure with weighted eviction disabled and `gravitino.cache.maxEntries=64`. A 
standalone component probe stalls in `CaffeineEntityCache.put`; repeated `jcmd 
Thread.print -l` reports a Java-level deadlock between the writer and 
`CaffeineEntityCache-Cleanup`.
   
   The writer holds a segment lock while entering `cacheData.put`, then waits 
for Caffeine's eviction lock. The cleanup thread holds the eviction lock, fills 
the shared bounded cleanup queue, and executes a removal callback inline via 
`CallerRunsPolicy`. That callback calls `invalidateExpiredItem` and waits for 
the writer's segment lock. Increasing queue capacity alone does not fix the 
lock-order cycle.
   
   ## Error message and/or stacktrace
   
   ```
   main holds segment lock -> BoundedLocalCache.performCleanUp -> waits 
eviction lock
   CaffeineEntityCache-Cleanup holds eviction lock
     -> BoundedLocalCache.notifyRemoval
     -> ThreadPoolExecutor.CallerRunsPolicy.rejectedExecution
     -> CaffeineEntityCache.invalidateExpiredItem
     -> SegmentedLock.withLock -> waits segment lock
   Found one Java-level deadlock
   ```
   
   ## How to reproduce
   
   Use exact main cache classes in a Java 17 JVM with 512 MiB heap. Instantiate 
`CaffeineEntityCache` with `CACHE_WEIGHER_ENABLED=false`, 
`CACHE_MAX_ENTRIES=64`, stats disabled and default lock segments. Insert many 
distinct `BaseMetalake` entities with 8 KiB unique properties using 
`cache.put`. The retained count is small; the test stresses eviction/callback 
scheduling, not intentional unbounded residency.
   
   The saved `TestCacheRetentionProbe` first performs three 20,000-entry 
fill/reclaim cycles in weighted mode, then switches to count-bounded mode in 
the same JVM; the count-mode first fill deadlocks in the confirmation run. Its 
incremental JSON confirms the preceding weighted cycles reclaimed both data and 
prefix index. Capture writer/cleanup lock ownership with `jcmd PID Thread.print 
-l` while stuck. A second eight-writer eviction stress checks the same 
mechanism. Timing is schedule-sensitive: an isolated sequential count-mode run 
also completed successfully.
   
   ## Additional context
   
   Evidence: `cache-retention-both-threads-1.txt` and subsequent repeated dumps 
include the exact two-lock cycle and JVM deadlock detection. This was a 
component test with explicit count-bounded configuration; the small default 
weighted HTTP benchmark did not encounter a deadlock. Default weighted 
eviction/expiration reaches the same cleanup mechanism, but its saturation was 
not reproduced here. Fix lock ordering/callback execution so eviction cannot 
synchronously reacquire entity segment locks while holding Caffeine's eviction 
lock; retain reliable index cleanup and add bounded-time concurrent eviction 
regression tests.
   
   Related but different: closed 
[#13462](https://github.com/apache/gravitino/issues/13462) and 
[#13465](https://github.com/apache/gravitino/issues/13465) concern a reentrant 
global read-to-write gate upgrade. This reproduction uses ordinary puts and an 
eviction-lock/segment-lock cycle without that upgrade.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to