[ 
https://issues.apache.org/jira/browse/SPARK-59804?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

L. C. Hsieh reassigned SPARK-59804:
-----------------------------------

    Assignee: L. C. Hsieh

> Fix UTF-8 byte limits and physical line numbers during event log replay
> -----------------------------------------------------------------------
>
>                 Key: SPARK-59804
>                 URL: https://issues.apache.org/jira/browse/SPARK-59804
>             Project: Spark
>          Issue Type: Bug
>          Components: Web UI
>    Affects Versions: 5.0.0
>            Reporter: L. C. Hsieh
>            Assignee: L. C. Hsieh
>            Priority: Major
>
> SPARK-59407 (commit 38258cbea90) introduced bounded event-log line reading in 
> ReplayListenerBus. There are two related correctness issues:
> 1. spark.history.fs.eventLog.maxLineLength is a byte-valued configuration, 
> but boundedLines compares StringBuilder.length() * 2 with the configured 
> limit. This measures UTF-16 storage rather than UTF-8 input length. With an 8 
> KiB limit, a valid ASCII SparkListenerApplicationStart event containing a 
> 6,000-character application name (approximately 6 KiB of JSON) is incorrectly 
> skipped. Multibyte UTF-8 input also has inconsistent boundaries. A trailing 
> CR in CRLF is counted before being removed.
> 2. boundedLines removes oversized lines before replay applies zipWithIndex. 
> Consequently, subsequent parse diagnostics report renumbered positions 
> instead of physical file lines. For example, a valid event on line 1, 
> oversized lines 2 and 3, and an invalid event {} on line 4 produce "Malformed 
> line #2" instead of "Malformed line #4".
> The proposed fix counts UTF-8 content bytes, excludes line endings from the 
> limit, and carries original line indices through skipped lines. It preserves 
> strict rejection of malformed UTF-8, filtered iterator behavior, and 
> truncated-log handling. Buffer growth remains bounded by the configured line 
> limit; the limit specifies encoded content length rather than an exact JVM 
> heap budget.
> Local regression coverage includes ASCII content, 1- through 4-byte UTF-8 
> characters, exact and over-limit boundaries, decoder-buffer crossings, 
> LF/CRLF/unterminated final lines, physical line numbers after multiple 
> skipped lines, filtered iterators, truncated logs, and malformed UTF-8 in 
> oversized lines.
> Validation: the original implementation fails three new regression tests. 
> With the fix, all 11 tests in ReplayListenerSuite pass under SBT, and both 
> production and test scalastyle checks pass.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to