[
https://issues.apache.org/jira/browse/TIKA-4901?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18116563#comment-18116563
]
ASF GitHub Bot commented on TIKA-4901:
--------------------------------------
tballison merged PR #3191:
URL: https://github.com/apache/tika/pull/3191
> Don't inflate an embedded document twice to digest it
> -----------------------------------------------------
>
> Key: TIKA-4901
> URL: https://issues.apache.org/jira/browse/TIKA-4901
> Project: Tika
> Issue Type: Improvement
> Reporter: Tim Allison
> Priority: Minor
>
> :robot: description
> {noformat}
> InputStreamDigester reads a stream to hash it, then rewinds so the parse
> can read it again. For a file that costs nothing — the page cache serves
> the second read at ~6 GB/s. For an embedded document it is not a file:
> every ReopenableSource is a re-openable supplier, and re-opening a zip entry
> means inflating it again, measured at 480–670 MB/s. The page cache holds
> the compressed container, so it does not help.
>
> TikaInputSource gains an advisory tryRetainInMemory(), default false.
> ReopenableSource implements it by reusing its existing in-memory drain and
> budget reservation, refusing when the length is unknown, when the content
> exceeds the 1 MB floor with no budget, or when the budget cannot cover it
> — so it never spills and never starts a drain that cannot finish.
> ensureOpen() and seekTo() now serve from the retained content when there is
> some. InputStreamDigester asks for retention after enableRewind;
> FileSource and CachingSource take the default and are unaffected.
>
> Full recursive parse of 120 zips × 16 × 128 KB entries with embedded
> digesting on, only tika-core swapped, 3 reps × 2 rounds: wall 3.875–4.081 s →
> 3.047–3.141 s, CPU 4.59–5.64 s → 3.78–4.83 s. Output identical (digests
> and extracted lengths checksum the same in every run). The durable number
> is the absolute one — about 3.4 ms per MB of embedded content, one inflate
> — since the ratio depends on how expensive the embedded parser is; text
> entries make the inflate share look large.
>
> Limits: entries above the 1 MB floor need budget headroom or retention is
> refused, and zip is the best case — an OLE2 document stream or a PST item
> re-opens more cheaply.
> Also fixes a latent bug in the same method: ensureOpen() at a non-zero
> position re-read from byte 0 while continuing to report positions as if it
> had not.
> {noformat}
--
This message was sent by Atlassian Jira
(v8.20.10#820010)