[ 
https://issues.apache.org/jira/browse/TIKA-4795?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18101375#comment-18101375
 ] 

Kristian Rickert edited comment on TIKA-4795 at 8/3/26 1:04 PM:
----------------------------------------------------------------

Do you think using the disk is the best route?  Can that be extracted to be 
in-memory or an in-memory filesystem?  I know it feels wasteful, but tika loads 
the full bytes into memory via an input stream anyway - soo adding it to disk 
is probably overkill.  I work with secure files, and this is something that 
would make or break if I can use it as serializing the input to disk can be a 
security risk and even flag some security scanners.

 

Yeah we should make it configurable.  I have a 60GB server and it chews through 
100MB files all the time and it's NBD - but I do bring the memory up per 
message to 1 or 2GB often.


was (Author: kristian):
Do you think using the disk is the best route?  Can that be extracted to be 
in-memory or an in-memory filesystem?  I know it feels wasteful, but tika loads 
the full bytes into memory via an input stream anyway - soo adding it to disk 
is probably overkill.  I work with secure files, and this is something that 
would make or break if I can use it as serializing the input to disk can be a 
security risk and even flag some security scanners.

> ParseBytes: parse-only entrypoint for callers that already hold the document 
> bytes
> ----------------------------------------------------------------------------------
>
>                 Key: TIKA-4795
>                 URL: https://issues.apache.org/jira/browse/TIKA-4795
>             Project: Tika
>          Issue Type: New Feature
>          Components: tika-pipes
>    Affects Versions: 4.0.0
>            Reporter: Davide Polato
>            Priority: Major
>              Labels: grpc, pipes, protobuf
>
> h4. Goal
> Add a parse-only tika-grpc entrypoint that parses the exact bytes supplied by 
> the caller, without asking Tika to fetch or re-fetch the resource.
> This was split from the typed {{Document}} work in 
> [TIKA-4766|https://issues.apache.org/jira/browse/TIKA-4766] / [PR 
> #2961|https://github.com/apache/tika/pull/2961], whose description states 
> that "{{ParseBytes}} is deferred to its own JIRA."
> h4. Motivation
> {{FetchAndParse}} obtains content through a registered fetcher. Callers such 
> as web crawlers have already acquired the resource bytes. Requiring another 
> acquisition path can duplicate the transfer or parse a different 
> representation from the one the caller observed because of redirects, 
> cookies, authentication, robots and rate-limiting decisions, or transient 
> content.
> The parse result should correspond to the exact representation captured by 
> the caller.
> Apache StormCrawler is a concrete consumer. Its parsing bolts already receive 
> the acquired bytes together with URL and protocol metadata. StormCrawler can 
> contribute crawler requirements and input fixtures including truncated 
> payloads, declared-vs-actual charset mismatches, malformed archives, 
> documents containing embedded resources, HTML with {{<base>}}, and distinct 
> URLs with identical content.
> h4. Proposed observable behavior
> This issue defines observable behavior, not the internal buffering, fetcher, 
> or Pipes implementation.
> The request carries the exact byte sequence. It may also carry:
> * an optional opaque caller correlation id, echoed but never interpreted
> * source URI and effective URI when available, plus a base URI for link 
> resolution; these are provenance metadata and must never be dereferenced by 
> the service
> * resource name
> * declared media type and charset hints
> * optional declared content length and digest
> * a truncation flag
> Input must be bounded, and deadline or cancellation must terminate the 
> associated parse work.
> The reply reuses the typed {{Document}} contract introduced by #2961 rather 
> than defining a second parse-result model.
> The current direction discussed in #2961 is to place this on the experimental 
> v2 surface so that it does not gate the 4.0.0 release.
> h4. Open design decisions
> * bounded unary versus client streaming for the request bytes
> * exact service and package placement within v2
> * which provenance fields and parsing hints are necessary
> * whether and how caller-supplied length and digest are validated
> * maximum payload size and behavior for truncated payloads
> h4. Non-goals
> This issue does not define the structured content tree or its Markdown 
> projection, the embedded-document representation, extension payloads such as 
> {{google.protobuf.Any}}, or downstream NLP and embedding enrichment. Those 
> remain separate TIKA-4766 stages and follow-up issues.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to