yuqi1129 opened a new issue, #13580:
URL: https://github.com/apache/gravitino/issues/13580
## Version
main branch, `ab0785324c5cb391ae63e4501d349f4336b8d3aa`; Java 17.0.8.
## Describe what's wrong
The unmodified main cache implementation deadlocks under insertion/eviction
pressure with weighted eviction disabled and `gravitino.cache.maxEntries=64`. A
standalone component probe stalls in `CaffeineEntityCache.put`; repeated `jcmd
Thread.print -l` reports a Java-level deadlock between the writer and
`CaffeineEntityCache-Cleanup`.
The writer holds a segment lock while entering `cacheData.put`, then waits
for Caffeine's eviction lock. The cleanup thread holds the eviction lock, fills
the shared bounded cleanup queue, and executes a removal callback inline via
`CallerRunsPolicy`. That callback calls `invalidateExpiredItem` and waits for
the writer's segment lock. Increasing queue capacity alone does not fix the
lock-order cycle.
## Error message and/or stacktrace
```
main holds segment lock -> BoundedLocalCache.performCleanUp -> waits
eviction lock
CaffeineEntityCache-Cleanup holds eviction lock
-> BoundedLocalCache.notifyRemoval
-> ThreadPoolExecutor.CallerRunsPolicy.rejectedExecution
-> CaffeineEntityCache.invalidateExpiredItem
-> SegmentedLock.withLock -> waits segment lock
Found one Java-level deadlock
```
## How to reproduce
Use exact main cache classes in a Java 17 JVM with 512 MiB heap. Instantiate
`CaffeineEntityCache` with `CACHE_WEIGHER_ENABLED=false`,
`CACHE_MAX_ENTRIES=64`, stats disabled and default lock segments. Insert many
distinct `BaseMetalake` entities with 8 KiB unique properties using
`cache.put`. The retained count is small; the test stresses eviction/callback
scheduling, not intentional unbounded residency.
The saved `TestCacheRetentionProbe` first performs three 20,000-entry
fill/reclaim cycles in weighted mode, then switches to count-bounded mode in
the same JVM; the count-mode first fill deadlocks in the confirmation run. Its
incremental JSON confirms the preceding weighted cycles reclaimed both data and
prefix index. Capture writer/cleanup lock ownership with `jcmd PID Thread.print
-l` while stuck. A second eight-writer eviction stress checks the same
mechanism. Timing is schedule-sensitive: an isolated sequential count-mode run
also completed successfully.
## Additional context
Evidence: `cache-retention-both-threads-1.txt` and subsequent repeated dumps
include the exact two-lock cycle and JVM deadlock detection. This was a
component test with explicit count-bounded configuration; the small default
weighted HTTP benchmark did not encounter a deadlock. Default weighted
eviction/expiration reaches the same cleanup mechanism, but its saturation was
not reproduced here. Fix lock ordering/callback execution so eviction cannot
synchronously reacquire entity segment locks while holding Caffeine's eviction
lock; retain reliable index cleanup and add bounded-time concurrent eviction
regression tests.
Related but different: closed
[#13462](https://github.com/apache/gravitino/issues/13462) and
[#13465](https://github.com/apache/gravitino/issues/13465) concern a reentrant
global read-to-write gate upgrade. This reproduction uses ordinary puts and an
eviction-lock/segment-lock cycle without that upgrade.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]