Hi,
actually you only need to track the main IndexInput with a Arena. All
clones should throw a "AlreadyClosed" Exception once the main one is
closed. Look at MMapDirectory and MemorySegmentIndexInput for more
details, it also has an implementation using MemorySegments. All clones
and the master indexinput share a common Arena (shared) with the mian
one. When the close or slice is used after the main one is closed, the
clones throw AlreadyClosedExcepotion (because the shared Arena was
closed). This is managed by catching the "IllegalStateException" when
accessing the MemorySegment.
In short: When you manage the "off heap" buffers using arenas this is
all easy to do. In addition I'd suggest to use a Caffeine cache for the
global buffer management (e.g., when something is borrowed from the buffer).
Small tip: The people at Opensearch already built something like this:
https://github.com/opensearch-project/opensearch-storage-encryption --
The main idea here was to buffer the "decrypted index contents". It also
uses DirectIO to access the underlying index files, decrypts them to a
Caffeine-based buffer cache and have MemorySegmentIndexInput like index
inputs on top of it. It uses ahred arenas to handle lifetime management.
Uwe
Am 03.09.2026 um 05:07 schrieb Trevor McCulloch:
I had looked at something similar (LRU cache in native code) and ran
into the same issue. AFAICT there is no good work around for this --
you can either attach your native memory to an auto arena and have it
freed at some arbitrary point in the future, or you can copy the data
on to the JVM heap. Making slice/cloned inputs closeable is a big API
change. I worry that fixing it within Lucene could be very difficult
and the API changes would be very disruptive to consumers of the library.
Also, BufferedIndexInput is very tempting but is not really designed
for the kind of aligned reads you need to do with DirectIO.
On Wed, Sep 2, 2026 at 5:04 PM Sagar Upadhyay
<[email protected]> wrote:
Hello folks,
I was trying to build a custom off-heap buffer cache on top of
Lucene's DirectIODirectory. To keep memory bounded and avoid
allocating (and freeing via Cleaner/GC) an off-heap buffer on
every miss, I was looking into recycling block buffers from a
freelist instead of allocating on every miss.
To do that I need to know that a buffer is no longer referenced by
any live IndexInput(clones and slices). So was looking at a
reference-counting based mechanism for that, but it would need
close() to be called on cloned/sliced IndexInputs, and I am not
sure that's possible today.
Has anyone looked into this i.e. a refcounted or recycled buffer
cache on top of DirectIODirectory? Any help/advice is appreciated.
Thanks,
Sagar
--
Uwe Schindler
Achterdiek 19, D-28357 Bremen
https://www.thetaphi.de
eMail:[email protected]