[
https://issues.apache.org/jira/browse/SPARK-59804?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
L. C. Hsieh reassigned SPARK-59804:
-----------------------------------
Assignee: L. C. Hsieh
> Fix UTF-8 byte limits and physical line numbers during event log replay
> -----------------------------------------------------------------------
>
> Key: SPARK-59804
> URL: https://issues.apache.org/jira/browse/SPARK-59804
> Project: Spark
> Issue Type: Bug
> Components: Web UI
> Affects Versions: 5.0.0
> Reporter: L. C. Hsieh
> Assignee: L. C. Hsieh
> Priority: Major
>
> SPARK-59407 (commit 38258cbea90) introduced bounded event-log line reading in
> ReplayListenerBus. There are two related correctness issues:
> 1. spark.history.fs.eventLog.maxLineLength is a byte-valued configuration,
> but boundedLines compares StringBuilder.length() * 2 with the configured
> limit. This measures UTF-16 storage rather than UTF-8 input length. With an 8
> KiB limit, a valid ASCII SparkListenerApplicationStart event containing a
> 6,000-character application name (approximately 6 KiB of JSON) is incorrectly
> skipped. Multibyte UTF-8 input also has inconsistent boundaries. A trailing
> CR in CRLF is counted before being removed.
> 2. boundedLines removes oversized lines before replay applies zipWithIndex.
> Consequently, subsequent parse diagnostics report renumbered positions
> instead of physical file lines. For example, a valid event on line 1,
> oversized lines 2 and 3, and an invalid event {} on line 4 produce "Malformed
> line #2" instead of "Malformed line #4".
> The proposed fix counts UTF-8 content bytes, excludes line endings from the
> limit, and carries original line indices through skipped lines. It preserves
> strict rejection of malformed UTF-8, filtered iterator behavior, and
> truncated-log handling. Buffer growth remains bounded by the configured line
> limit; the limit specifies encoded content length rather than an exact JVM
> heap budget.
> Local regression coverage includes ASCII content, 1- through 4-byte UTF-8
> characters, exact and over-limit boundaries, decoder-buffer crossings,
> LF/CRLF/unterminated final lines, physical line numbers after multiple
> skipped lines, filtered iterators, truncated logs, and malformed UTF-8 in
> oversized lines.
> Validation: the original implementation fails three new regression tests.
> With the fix, all 11 tests in ReplayListenerSuite pass under SBT, and both
> production and test scalastyle checks pass.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]