[
https://issues.apache.org/jira/browse/TIKA-4795?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18101375#comment-18101375
]
Kristian Rickert commented on TIKA-4795:
----------------------------------------
Do you think using the disk is the best route? Can that be extracted to be
in-memory or an in-memory filesystem? I know it feels wasteful, but tika loads
the full bytes into memory via an input stream anyway - soo adding it to disk
is probably overkill. I work with secure files, and this is something that
would make or break if I can use it as serializing the input to disk can be a
security risk and even flag some security scanners.
> ParseBytes: parse-only entrypoint for callers that already hold the document
> bytes
> ----------------------------------------------------------------------------------
>
> Key: TIKA-4795
> URL: https://issues.apache.org/jira/browse/TIKA-4795
> Project: Tika
> Issue Type: New Feature
> Components: tika-pipes
> Affects Versions: 4.0.0
> Reporter: Davide Polato
> Priority: Major
> Labels: grpc, pipes, protobuf
>
> h4. Goal
> Add a parse-only tika-grpc entrypoint that parses the exact bytes supplied by
> the caller, without asking Tika to fetch or re-fetch the resource.
> This was split from the typed {{Document}} work in
> [TIKA-4766|https://issues.apache.org/jira/browse/TIKA-4766] / [PR
> #2961|https://github.com/apache/tika/pull/2961], whose description states
> that "{{ParseBytes}} is deferred to its own JIRA."
> h4. Motivation
> {{FetchAndParse}} obtains content through a registered fetcher. Callers such
> as web crawlers have already acquired the resource bytes. Requiring another
> acquisition path can duplicate the transfer or parse a different
> representation from the one the caller observed because of redirects,
> cookies, authentication, robots and rate-limiting decisions, or transient
> content.
> The parse result should correspond to the exact representation captured by
> the caller.
> Apache StormCrawler is a concrete consumer. Its parsing bolts already receive
> the acquired bytes together with URL and protocol metadata. StormCrawler can
> contribute crawler requirements and input fixtures including truncated
> payloads, declared-vs-actual charset mismatches, malformed archives,
> documents containing embedded resources, HTML with {{<base>}}, and distinct
> URLs with identical content.
> h4. Proposed observable behavior
> This issue defines observable behavior, not the internal buffering, fetcher,
> or Pipes implementation.
> The request carries the exact byte sequence. It may also carry:
> * an optional opaque caller correlation id, echoed but never interpreted
> * source URI and effective URI when available, plus a base URI for link
> resolution; these are provenance metadata and must never be dereferenced by
> the service
> * resource name
> * declared media type and charset hints
> * optional declared content length and digest
> * a truncation flag
> Input must be bounded, and deadline or cancellation must terminate the
> associated parse work.
> The reply reuses the typed {{Document}} contract introduced by #2961 rather
> than defining a second parse-result model.
> The current direction discussed in #2961 is to place this on the experimental
> v2 surface so that it does not gate the 4.0.0 release.
> h4. Open design decisions
> * bounded unary versus client streaming for the request bytes
> * exact service and package placement within v2
> * which provenance fields and parsing hints are necessary
> * whether and how caller-supplied length and digest are validated
> * maximum payload size and behavior for truncated payloads
> h4. Non-goals
> This issue does not define the structured content tree or its Markdown
> projection, the embedded-document representation, extension payloads such as
> {{google.protobuf.Any}}, or downstream NLP and embedding enrichment. Those
> remain separate TIKA-4766 stages and follow-up issues.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)