rkhachatryan commented on code in PR #28661:
URL: https://github.com/apache/flink/pull/28661#discussion_r3724924102


##########
flink-runtime/src/main/java/org/apache/flink/runtime/checkpoint/channel/FetchedChannelStateReader.java:
##########
@@ -0,0 +1,127 @@
+/*
+ * Licensed to the Apache Software Foundation (ASF) under one or more
+ * contributor license agreements.  See the NOTICE file distributed with
+ * this work for additional information regarding copyright ownership.
+ * The ASF licenses this file to You under the Apache License, Version 2.0
+ * (the "License"); you may not use this file except in compliance with
+ * the License.  You may obtain a copy of the License at
+ *
+ *    http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ */
+
+package org.apache.flink.runtime.checkpoint.channel;
+
+import org.apache.flink.annotation.Internal;
+
+import java.io.Closeable;
+import java.io.InputStream;
+import java.util.Collections;
+import java.util.Optional;
+
+/**
+ * Forward reader over a {@link FetchedChannelState}'s spill files. This is 
our own segment reader,

Review Comment:
   I would prefer the code to be self-documenting (so that there's no need to 
read the javadoc).
   How about `advanceAndGetNextSegment`?



##########
flink-runtime/src/main/java/org/apache/flink/runtime/checkpoint/channel/FetchedChannelStateReader.java:
##########
@@ -0,0 +1,127 @@
+/*
+ * Licensed to the Apache Software Foundation (ASF) under one or more
+ * contributor license agreements.  See the NOTICE file distributed with
+ * this work for additional information regarding copyright ownership.
+ * The ASF licenses this file to You under the Apache License, Version 2.0
+ * (the "License"); you may not use this file except in compliance with
+ * the License.  You may obtain a copy of the License at
+ *
+ *    http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ */
+
+package org.apache.flink.runtime.checkpoint.channel;
+
+import org.apache.flink.annotation.Internal;
+
+import java.io.Closeable;
+import java.io.InputStream;
+import java.util.Collections;
+import java.util.Optional;
+
+/**
+ * Forward reader over a {@link FetchedChannelState}'s spill files. This is 
our own segment reader,
+ * on purpose <em>not</em> a Java {@link java.util.Iterator}: our access 
pattern ("a body must be
+ * fully read before the next segment", "body ownership is handed to the 
consumer", "consume and
+ * commit are separate steps") does not fit the {@code hasNext/next} contract.
+ *
+ * <p>This interface is the contract callers depend on; {@link 
FetchedChannelStateReaderImpl} holds
+ * the implementation (the live file stream, the two progress positions, the 
bounded body view).
+ *
+ * <p>Reading is strictly sequential: a reader is positioned once (offset 0 
for the root reader, or
+ * the committed position for a {@link #snapshot()}), then consumes forward 
only via {@link
+ * #nextSegment()}. It never seeks backward and never re-positions 
mid-iteration.
+ *
+ * <p>The drain thread reads the root reader front to back and commits via 
{@link

Review Comment:
   TBH, drain reader is also not very clear to me :) 
   I'm struggling to come up with a better name though. Maybe main reader?
   
   (this is very NIT anyways)



##########
flink-runtime/src/main/java/org/apache/flink/runtime/checkpoint/channel/FetchedChannelStateReader.java:
##########
@@ -0,0 +1,127 @@
+/*
+ * Licensed to the Apache Software Foundation (ASF) under one or more
+ * contributor license agreements.  See the NOTICE file distributed with
+ * this work for additional information regarding copyright ownership.
+ * The ASF licenses this file to You under the Apache License, Version 2.0
+ * (the "License"); you may not use this file except in compliance with
+ * the License.  You may obtain a copy of the License at
+ *
+ *    http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ */
+
+package org.apache.flink.runtime.checkpoint.channel;
+
+import org.apache.flink.annotation.Internal;
+
+import java.io.Closeable;
+import java.io.InputStream;
+import java.util.Collections;
+import java.util.Optional;
+
+/**
+ * Forward reader over a {@link FetchedChannelState}'s spill files. This is 
our own segment reader,
+ * on purpose <em>not</em> a Java {@link java.util.Iterator}: our access 
pattern ("a body must be
+ * fully read before the next segment", "body ownership is handed to the 
consumer", "consume and
+ * commit are separate steps") does not fit the {@code hasNext/next} contract.
+ *
+ * <p>This interface is the contract callers depend on; {@link 
FetchedChannelStateReaderImpl} holds
+ * the implementation (the live file stream, the two progress positions, the 
bounded body view).
+ *
+ * <p>Reading is strictly sequential: a reader is positioned once (offset 0 
for the root reader, or
+ * the committed position for a {@link #snapshot()}), then consumes forward 
only via {@link
+ * #nextSegment()}. It never seeks backward and never re-positions 
mid-iteration.
+ *
+ * <p>The drain thread reads the root reader front to back and commits via 
{@link
+ * SpillSegment#commit()}; each checkpoint derives a fresh {@link #snapshot()} 
that resumes from the
+ * committed position. {@link #snapshot()} and {@link SpillSegment#commit()} 
must be called under
+ * the drainer lock; disk reads happen outside it.
+ */
+@Internal
+public interface FetchedChannelStateReader extends Closeable {
+
+    /**
+     * Advances to the next segment and returns it, or {@link 
Optional#empty()} when no segment
+     * remains. Advancing and probing are one step; there is no separate 
{@code hasNext}.
+     *
+     * <p>Entry rule (the first call is exempt): the previous segment's body 
must be fully read,
+     * otherwise this is a contract violation and fails loud (no skip-ahead).
+     */
+    Optional<SpillSegment> nextSegment();
+
+    /**
+     * Derives an independent resume point starting from the committed 
position. The snapshot holds

Review Comment:
   Sorry but it doesn't tell me much.
   Could you clarify in the javadoc - who calls commit and when?
   
   > before anything is committed it's simply the reader's start position; now 
defined in the interface javadoc.
   
   👍 



##########
flink-runtime/src/main/java/org/apache/flink/runtime/checkpoint/channel/SequentialChannelStateReaderImpl.java:
##########
@@ -110,7 +111,7 @@ public Optional<FetchedChannelState> readInputData(
             // only signals "there is state to recover". The spilling backend 
returns a real,
             // file-backed container here.

Review Comment:
   > added the TODO about returning getProducedChannelState() once the 
Spilling* handlers are selected.
   
   Can't we return it already? We already return - just the wrong value:
   ```
               return filterContext.isCheckpointingDuringRecoveryEnabled() && 
readAny
                       ? Optional.of(new 
FetchedChannelState(java.util.Collections.emptyList()))
   ```



##########
flink-runtime/src/main/java/org/apache/flink/runtime/checkpoint/channel/FetchedChannelStateReader.java:
##########
@@ -0,0 +1,127 @@
+/*
+ * Licensed to the Apache Software Foundation (ASF) under one or more
+ * contributor license agreements.  See the NOTICE file distributed with
+ * this work for additional information regarding copyright ownership.
+ * The ASF licenses this file to You under the Apache License, Version 2.0
+ * (the "License"); you may not use this file except in compliance with
+ * the License.  You may obtain a copy of the License at
+ *
+ *    http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ */
+
+package org.apache.flink.runtime.checkpoint.channel;
+
+import org.apache.flink.annotation.Internal;
+
+import java.io.Closeable;
+import java.io.InputStream;
+import java.util.Collections;
+import java.util.Optional;
+
+/**
+ * Forward reader over a {@link FetchedChannelState}'s spill files. This is 
our own segment reader,
+ * on purpose <em>not</em> a Java {@link java.util.Iterator}: our access 
pattern ("a body must be
+ * fully read before the next segment", "body ownership is handed to the 
consumer", "consume and
+ * commit are separate steps") does not fit the {@code hasNext/next} contract.
+ *
+ * <p>This interface is the contract callers depend on; {@link 
FetchedChannelStateReaderImpl} holds
+ * the implementation (the live file stream, the two progress positions, the 
bounded body view).
+ *
+ * <p>Reading is strictly sequential: a reader is positioned once (offset 0 
for the root reader, or
+ * the committed position for a {@link #snapshot()}), then consumes forward 
only via {@link
+ * #nextSegment()}. It never seeks backward and never re-positions 
mid-iteration.
+ *
+ * <p>The drain thread reads the root reader front to back and commits via 
{@link
+ * SpillSegment#commit()}; each checkpoint derives a fresh {@link #snapshot()} 
that resumes from the
+ * committed position. {@link #snapshot()} and {@link SpillSegment#commit()} 
must be called under
+ * the drainer lock; disk reads happen outside it.
+ */
+@Internal
+public interface FetchedChannelStateReader extends Closeable {
+
+    /**
+     * Advances to the next segment and returns it, or {@link 
Optional#empty()} when no segment
+     * remains. Advancing and probing are one step; there is no separate 
{@code hasNext}.
+     *
+     * <p>Entry rule (the first call is exempt): the previous segment's body 
must be fully read,
+     * otherwise this is a contract violation and fails loud (no skip-ahead).
+     */
+    Optional<SpillSegment> nextSegment();
+
+    /**
+     * Derives an independent resume point starting from the committed 
position. The snapshot holds
+     * its own {@link FetchedChannelState} lifecycle grant; the caller must 
open a reader from it
+     * via {@link FetchedChannelStateSnapshot#reader()} and close that reader 
when done.
+     *
+     * <p>Must be called under the drainer lock so that the copied position 
reflects the latest
+     * committed state.
+     *
+     * @return a snapshot capturing the current committed position; caller 
must open and close a
+     *     reader from it
+     */
+    FetchedChannelStateSnapshot snapshot();
+
+    /**
+     * Returns a reader with no segments — its first {@link #nextSegment()} is 
empty. Each call
+     * hands out a fresh instance: readers have independent lifecycle and 
{@link #close()} is
+     * single-use, so a shared instance would let one consumer's close break 
later consumers. Used
+     * wherever there is nothing to snapshot (e.g. after drain finished, or 
the no-op recovery
+     * trigger).
+     */
+    static FetchedChannelStateReader emptyReader() {
+        return new FetchedChannelState(Collections.emptyList()).reader();
+    }
+
+    /**
+     * One per-channel segment produced by {@link #nextSegment()}.
+     *
+     * <p>The segment body bytes are opaque to the reader; record framing is 
handled by the
+     * consumer's deserializer. A consumer reads {@link #bodyStream()} to EOF 
(after {@link
+     * #length()} bytes), and the drain consumer additionally calls {@link 
#commit()} under the
+     * drainer lock after each delivery so that a later {@link 
FetchedChannelStateReader#snapshot()}
+     * resumes from the delivered boundary.
+     *
+     * <p>Ownership of {@link #bodyStream()} passes to the consumer: the 
reader no longer tracks how
+     * far it has been read. The "previous body must be fully read" rule (no 
skip-ahead) is enforced
+     * at the next {@link FetchedChannelStateReader#nextSegment()} call, not 
here.
+     *
+     * <p>A segment is valid only until the next {@code nextSegment()} call on 
the parent reader.
+     */
+    interface SpillSegment {
+
+        /** The channel whose data this segment contains. */
+        InputChannelInfo channelInfo();
+
+        /**
+         * Returns an {@link InputStream} bounded to this segment's body. 
Reading returns {@code -1}
+         * (EOF) after {@link #length()} bytes; it never reads into the next 
segment or the next
+         * file.
+         *
+         * <p>The stream is single-use, not thread-safe, and must be fully 
consumed before the next
+         * {@link FetchedChannelStateReader#nextSegment()}.
+         */
+        InputStream bodyStream();
+
+        /**
+         * Number of body bytes this segment hands out before EOF. For the 
snapshot path this is the
+         * not-yet-delivered remainder used as the length prefix when writing 
to the checkpoint
+         * stream. Bounded by the spill file size limit, so it always fits in 
an {@code int}.
+         */
+        int length();
+
+        /**
+         * Advances the reader's committed position to match how many body 
bytes have been read from
+         * {@link #bodyStream()} so far. Must be called under the drainer lock 
after each buffer
+         * delivery so that a subsequent {@link 
FetchedChannelStateReader#snapshot()} sees the
+         * correct delivered boundary. Only the drain (root) reader commits; 
the snapshot reader

Review Comment:
   I think this interface split already makes sense - it makes it clear how 
each call site uses the interface.
   I don't understand why the implementation is the same? Can't we split it as 
well?
   
   > What does bring real simplification is your suggestion in #28661 
(comment): once the snapshot carries the resume segment, most of the 
contractual comments this split would enforce disappear anyway.
   
   Probably we should discuss it offline.



##########
flink-runtime/src/main/java/org/apache/flink/runtime/checkpoint/channel/FetchedChannelStateReaderImpl.java:
##########
@@ -0,0 +1,533 @@
+/*
+ * Licensed to the Apache Software Foundation (ASF) under one or more
+ * contributor license agreements.  See the NOTICE file distributed with
+ * this work for additional information regarding copyright ownership.
+ * The ASF licenses this file to You under the Apache License, Version 2.0
+ * (the "License"); you may not use this file except in compliance with
+ * the License.  You may obtain a copy of the License at
+ *
+ *    http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ */
+
+package org.apache.flink.runtime.checkpoint.channel;
+
+import org.apache.flink.annotation.Internal;
+import 
org.apache.flink.runtime.checkpoint.channel.FetchedChannelStateReader.SpillSegment;
+
+import javax.annotation.Nullable;
+
+import java.io.ByteArrayInputStream;
+import java.io.DataInputStream;
+import java.io.EOFException;
+import java.io.IOException;
+import java.io.InputStream;
+import java.nio.file.Files;
+import java.nio.file.Path;
+import java.util.List;
+import java.util.Optional;
+
+import static 
org.apache.flink.runtime.checkpoint.channel.AbstractSpillingHandler.SEGMENT_HEADER_BYTES;
+import static org.apache.flink.util.Preconditions.checkState;
+
+/**
+ * The single {@link FetchedChannelStateReader} implementation over a {@link 
FetchedChannelState}'s
+ * spill files.
+ *
+ * <p>Reading is strictly sequential and never seeks mid-iteration. There is 
exactly one place that
+ * skips bytes: the very first {@link #nextSegment()} call, where a snapshot 
reader started mid-body
+ * discards the already-delivered prefix to land on the not-yet-delivered 
remainder. Every later
+ * call does no skipping at all — the previous body was read to its end, so 
the stream already sits
+ * on the next segment's header. This "skip only on first positioning" rule is 
what keeps the
+ * steady-state path free of any seek/skip.
+ *
+ * <p>The reader holds <b>two</b> {@link Position}s and nothing else 
duplicates them:
+ *
+ * <ul>
+ *   <li>{@code current} — the live read position; its {@code readOffset} is 
exactly where the open
+ *       file stream sits, advancing as the header and the consumer's body 
reads consume bytes (the
+ *       latter outside the drainer lock).
+ *   <li>{@code committed} — the delivered boundary; {@link 
SpillSegment#commit()} advances it from
+ *       {@code current} (under the drainer lock). {@link #snapshot()} derives 
a new reader from it.
+ * </ul>
+ *
+ * <p>The "previous body fully read before advancing" rule is checked at the 
{@link #nextSegment()}
+ * entry (the first call is exempt — there is no previous segment). Body 
ownership is handed to the
+ * consumer, so the reader does not track body progress except through {@code 
current}.
+ */
+@Internal
+final class FetchedChannelStateReaderImpl implements FetchedChannelStateReader 
{
+
+    private final FetchedChannelStateSnapshot snapshot;
+    private final FetchedChannelState channelState;
+    private final List<Path> files;
+
+    /** Live read position; {@code readOffset} is where the open stream 
physically sits. */
+    private final Position current;
+
+    /** Delivered boundary; {@link SpillSegment#commit()} advances it from 
{@link #current}. */
+    private final Position committed;
+
+    /** Open stream over {@code current.fileIndex}, or {@code null} before the 
first read. */
+    @Nullable private InputStream fileStream;
+
+    /** Size of the file currently open. */
+    private long currentFileSize;
+
+    /** Body view of the segment returned by the last {@link #nextSegment()}, 
or {@code null}. */
+    @Nullable private BoundedSegmentStream currentBody;
+
+    private boolean positioned;
+    private boolean closed;
+
+    FetchedChannelStateReaderImpl(FetchedChannelStateSnapshot snapshot) {
+        this.snapshot = snapshot;
+        this.channelState = snapshot.channelState();
+        this.files = channelState.files();
+        // Must copy the position so that this reader's commits do not mutate 
the snapshot's state.
+        this.committed = snapshot.position().copy();
+        this.current = committed.copy();
+    }
+
+    @Override
+    public Optional<SpillSegment> nextSegment() {
+        checkState(!closed, "FetchedChannelStateReader is closed");
+        checkState(
+                currentBody == null || currentBody.remaining() == 0,
+                "Previous segment body not fully consumed before advancing: %s 
bytes left",
+                currentBody == null ? 0 : currentBody.remaining());
+        try {
+            if (!positioned) {
+                positioned = true;
+                return firstSegment();
+            }
+            return followingSegment();
+        } catch (IOException e) {
+            throw new RuntimeException("Failed to read segment", e);
+        }
+    }
+
+    /**
+     * First positioning — the only path that may skip bytes. Opens the file 
at the committed header
+     * offset and reads the header. A snapshot may resume in the middle of a 
segment: the committed
+     * {@code readOffset} says how many body bytes were already delivered, and 
that prefix is
+     * skipped so the returned body starts at the not-yet-delivered remainder. 
If the segment was
+     * already fully delivered (prefix == whole body), it is exhausted here 
and we move on to the
+     * next one.
+     */
+    private Optional<SpillSegment> firstSegment() throws IOException {
+        // The committed read offset may sit mid-body (after a partial 
commit), but the header lives
+        // at segmentStartOffset. Capture how much was already delivered, then 
rewind the live read
+        // offset to the header so we open the file there and read the header, 
not mid-body.
+        int deliveredPrefix = (int) current.deliveredBodyBytes();
+        current.rewindToSegmentStart();
+
+        if (!openCurrentFile()) {
+            return Optional.empty();
+        }
+        SegmentHeader header = readHeaderAtCurrent();
+        checkState(
+                deliveredPrefix <= header.bufferLength,
+                "Delivered offset %s exceeds segment length %s",
+                deliveredPrefix,
+                header.bufferLength);
+
+        if (deliveredPrefix == header.bufferLength) {
+            // This segment was already fully delivered before the snapshot; 
nothing remains in it.
+            // Skip its whole body to reach the next segment's header, then 
take the steady path.
+            skipBody(header.bufferLength);
+            return followingSegment();
+        }
+
+        // Discard the already-delivered prefix (the one and only skip in this 
class), then hand out
+        // the remainder. alreadyDelivered is carried so commit() records the 
boundary from the
+        // head.
+        skipBody(deliveredPrefix);
+        currentBody =
+                new BoundedSegmentStream(header.bufferLength - 
deliveredPrefix, deliveredPrefix);
+        return Optional.of(new Segment(header.channelInfo, currentBody));
+    }
+
+    /**
+     * Steady-state path — no skipping. The previous body was read to its end, 
so the stream sits
+     * exactly on this segment's header (or at the current file's end, in 
which case we roll to the
+     * next file). Reads the header and returns the whole-body view.
+     */
+    private Optional<SpillSegment> followingSegment() throws IOException {
+        if (!openCurrentFile()) {
+            return Optional.empty();
+        }
+        SegmentHeader header = readHeaderAtCurrent();
+        currentBody = new BoundedSegmentStream(header.bufferLength);
+        return Optional.of(new Segment(header.channelInfo, currentBody));
+    }
+
+    @Override
+    public FetchedChannelStateSnapshot snapshot() {
+        checkState(!closed, "FetchedChannelStateReader is closed");
+        return new FetchedChannelStateSnapshot(channelState, committed.copy());
+    }
+
+    @Override
+    public void close() throws IOException {
+        if (closed) {
+            return;
+        }
+        closed = true;
+        try {
+            closeFileStream();
+        } finally {
+            snapshot.release();
+        }
+    }
+
+    // 
-------------------------------------------------------------------------------------------
+    // Sequential IO over the spill files; all of it advances 
current.readOffset / current.fileIndex
+    // 
-------------------------------------------------------------------------------------------
+
+    /**
+     * Ensures a file is open with the stream positioned at {@code current}'s 
read offset, ready to
+     * read this segment's header. Rolls to the next file when the current one 
is exhausted. Returns
+     * false when no segment remains.
+     *
+     * <p>{@code current.segmentStartOffset} is set to where the header 
begins, so a later {@link
+     * SpillSegment#commit()} records the right segment for the snapshot to 
resume from.
+     */
+    private boolean openCurrentFile() throws IOException {
+        if (current.fileIndex >= files.size()) {
+            return false;
+        }
+        openFileAndSeek();
+        if (current.readOffset < currentFileSize) {
+            current.startSegmentHere();
+            return true;
+        }
+        // Current file fully read: move to the next file's first segment.
+        closeFileStream();
+        current.rollToNextFile();
+        if (current.fileIndex >= files.size()) {
+            return false;
+        }
+        openFileAndSeek();
+        if (current.readOffset < currentFileSize) {
+            current.startSegmentHere();
+            return true;
+        }
+        return false;
+    }
+
+    /** Reads the 12-byte header at the current read offset; advances past it. 
*/
+    private SegmentHeader readHeaderAtCurrent() throws IOException {
+        byte[] headerBytes = new byte[SEGMENT_HEADER_BYTES];
+        readFully(headerBytes);
+        DataInputStream h = new DataInputStream(new 
ByteArrayInputStream(headerBytes));
+        int gateIdx = h.readInt();
+        int channelIdx = h.readInt();
+        int bufferLength = h.readInt();
+        checkState(bufferLength >= 0, "negative segment length: %s", 
bufferLength);
+        return new SegmentHeader(new InputChannelInfo(gateIdx, channelIdx), 
bufferLength);
+    }
+
+    /**
+     * Ensures the file at {@code current.fileIndex} is open with the stream 
positioned at {@code
+     * current.readOffset}. If a stream is already open it is left as-is: 
sequential reading
+     * guarantees it is already there.
+     */
+    private void openFileAndSeek() throws IOException {
+        if (fileStream != null) {
+            return;
+        }
+        Path path = files.get(current.fileIndex);
+        currentFileSize = Files.size(path);
+        InputStream in = Files.newInputStream(path);
+        try {
+            skipOnStream(in, current.readOffset, path);
+        } catch (IOException e) {
+            in.close();
+            throw e;
+        }
+        fileStream = in;
+    }
+
+    /** Skips {@code count} body bytes on the open stream, advancing the read 
offset. */
+    private void skipBody(long count) throws IOException {
+        if (count > 0) {
+            skipOnStream(fileStream, count, files.get(current.fileIndex));
+            current.advanceReadOffset(count);
+        }
+    }
+
+    /** Skips exactly {@code count} bytes on {@code in}, failing loud if the 
file ends early. */
+    private void skipOnStream(InputStream in, long count, Path path) throws 
IOException {

Review Comment:
   Yes, this comment is about `skipOnStream`, sorry for misplacing.
   I think `skipOnStream` can be replaced by `IOUtils#skipFully`.
   
   `readFully` as well - it was the 
[next](https://github.com/apache/flink/pull/28661#discussion_r3693955986) 
comment



##########
flink-runtime/src/main/java/org/apache/flink/runtime/checkpoint/channel/FetchedChannelStateReaderImpl.java:
##########
@@ -0,0 +1,533 @@
+/*
+ * Licensed to the Apache Software Foundation (ASF) under one or more
+ * contributor license agreements.  See the NOTICE file distributed with
+ * this work for additional information regarding copyright ownership.
+ * The ASF licenses this file to You under the Apache License, Version 2.0
+ * (the "License"); you may not use this file except in compliance with
+ * the License.  You may obtain a copy of the License at
+ *
+ *    http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ */
+
+package org.apache.flink.runtime.checkpoint.channel;
+
+import org.apache.flink.annotation.Internal;
+import 
org.apache.flink.runtime.checkpoint.channel.FetchedChannelStateReader.SpillSegment;
+
+import javax.annotation.Nullable;
+
+import java.io.ByteArrayInputStream;
+import java.io.DataInputStream;
+import java.io.EOFException;
+import java.io.IOException;
+import java.io.InputStream;
+import java.nio.file.Files;
+import java.nio.file.Path;
+import java.util.List;
+import java.util.Optional;
+
+import static 
org.apache.flink.runtime.checkpoint.channel.AbstractSpillingHandler.SEGMENT_HEADER_BYTES;
+import static org.apache.flink.util.Preconditions.checkState;
+
+/**
+ * The single {@link FetchedChannelStateReader} implementation over a {@link 
FetchedChannelState}'s
+ * spill files.
+ *
+ * <p>Reading is strictly sequential and never seeks mid-iteration. There is 
exactly one place that
+ * skips bytes: the very first {@link #nextSegment()} call, where a snapshot 
reader started mid-body
+ * discards the already-delivered prefix to land on the not-yet-delivered 
remainder. Every later
+ * call does no skipping at all — the previous body was read to its end, so 
the stream already sits
+ * on the next segment's header. This "skip only on first positioning" rule is 
what keeps the
+ * steady-state path free of any seek/skip.
+ *
+ * <p>The reader holds <b>two</b> {@link Position}s and nothing else 
duplicates them:
+ *
+ * <ul>
+ *   <li>{@code current} — the live read position; its {@code readOffset} is 
exactly where the open
+ *       file stream sits, advancing as the header and the consumer's body 
reads consume bytes (the
+ *       latter outside the drainer lock).
+ *   <li>{@code committed} — the delivered boundary; {@link 
SpillSegment#commit()} advances it from
+ *       {@code current} (under the drainer lock). {@link #snapshot()} derives 
a new reader from it.
+ * </ul>
+ *
+ * <p>The "previous body fully read before advancing" rule is checked at the 
{@link #nextSegment()}
+ * entry (the first call is exempt — there is no previous segment). Body 
ownership is handed to the
+ * consumer, so the reader does not track body progress except through {@code 
current}.
+ */
+@Internal
+final class FetchedChannelStateReaderImpl implements FetchedChannelStateReader 
{
+
+    private final FetchedChannelStateSnapshot snapshot;
+    private final FetchedChannelState channelState;
+    private final List<Path> files;
+
+    /** Live read position; {@code readOffset} is where the open stream 
physically sits. */
+    private final Position current;
+
+    /** Delivered boundary; {@link SpillSegment#commit()} advances it from 
{@link #current}. */
+    private final Position committed;
+
+    /** Open stream over {@code current.fileIndex}, or {@code null} before the 
first read. */
+    @Nullable private InputStream fileStream;
+
+    /** Size of the file currently open. */
+    private long currentFileSize;
+
+    /** Body view of the segment returned by the last {@link #nextSegment()}, 
or {@code null}. */
+    @Nullable private BoundedSegmentStream currentBody;
+
+    private boolean positioned;
+    private boolean closed;
+
+    FetchedChannelStateReaderImpl(FetchedChannelStateSnapshot snapshot) {
+        this.snapshot = snapshot;
+        this.channelState = snapshot.channelState();
+        this.files = channelState.files();
+        // Must copy the position so that this reader's commits do not mutate 
the snapshot's state.
+        this.committed = snapshot.position().copy();
+        this.current = committed.copy();
+    }
+
+    @Override
+    public Optional<SpillSegment> nextSegment() {
+        checkState(!closed, "FetchedChannelStateReader is closed");
+        checkState(
+                currentBody == null || currentBody.remaining() == 0,
+                "Previous segment body not fully consumed before advancing: %s 
bytes left",
+                currentBody == null ? 0 : currentBody.remaining());
+        try {
+            if (!positioned) {
+                positioned = true;
+                return firstSegment();
+            }
+            return followingSegment();
+        } catch (IOException e) {
+            throw new RuntimeException("Failed to read segment", e);
+        }
+    }
+
+    /**
+     * First positioning — the only path that may skip bytes. Opens the file 
at the committed header
+     * offset and reads the header. A snapshot may resume in the middle of a 
segment: the committed
+     * {@code readOffset} says how many body bytes were already delivered, and 
that prefix is
+     * skipped so the returned body starts at the not-yet-delivered remainder. 
If the segment was
+     * already fully delivered (prefix == whole body), it is exhausted here 
and we move on to the
+     * next one.
+     */
+    private Optional<SpillSegment> firstSegment() throws IOException {
+        // The committed read offset may sit mid-body (after a partial 
commit), but the header lives
+        // at segmentStartOffset. Capture how much was already delivered, then 
rewind the live read
+        // offset to the header so we open the file there and read the header, 
not mid-body.
+        int deliveredPrefix = (int) current.deliveredBodyBytes();
+        current.rewindToSegmentStart();
+
+        if (!openCurrentFile()) {
+            return Optional.empty();
+        }
+        SegmentHeader header = readHeaderAtCurrent();
+        checkState(
+                deliveredPrefix <= header.bufferLength,
+                "Delivered offset %s exceeds segment length %s",
+                deliveredPrefix,
+                header.bufferLength);
+
+        if (deliveredPrefix == header.bufferLength) {
+            // This segment was already fully delivered before the snapshot; 
nothing remains in it.
+            // Skip its whole body to reach the next segment's header, then 
take the steady path.
+            skipBody(header.bufferLength);
+            return followingSegment();
+        }
+
+        // Discard the already-delivered prefix (the one and only skip in this 
class), then hand out
+        // the remainder. alreadyDelivered is carried so commit() records the 
boundary from the
+        // head.
+        skipBody(deliveredPrefix);
+        currentBody =
+                new BoundedSegmentStream(header.bufferLength - 
deliveredPrefix, deliveredPrefix);
+        return Optional.of(new Segment(header.channelInfo, currentBody));
+    }
+
+    /**
+     * Steady-state path — no skipping. The previous body was read to its end, 
so the stream sits
+     * exactly on this segment's header (or at the current file's end, in 
which case we roll to the
+     * next file). Reads the header and returns the whole-body view.
+     */
+    private Optional<SpillSegment> followingSegment() throws IOException {
+        if (!openCurrentFile()) {
+            return Optional.empty();
+        }
+        SegmentHeader header = readHeaderAtCurrent();
+        currentBody = new BoundedSegmentStream(header.bufferLength);
+        return Optional.of(new Segment(header.channelInfo, currentBody));
+    }
+
+    @Override
+    public FetchedChannelStateSnapshot snapshot() {
+        checkState(!closed, "FetchedChannelStateReader is closed");
+        return new FetchedChannelStateSnapshot(channelState, committed.copy());
+    }
+
+    @Override
+    public void close() throws IOException {
+        if (closed) {
+            return;
+        }
+        closed = true;
+        try {
+            closeFileStream();
+        } finally {
+            snapshot.release();
+        }
+    }
+
+    // 
-------------------------------------------------------------------------------------------
+    // Sequential IO over the spill files; all of it advances 
current.readOffset / current.fileIndex
+    // 
-------------------------------------------------------------------------------------------
+
+    /**
+     * Ensures a file is open with the stream positioned at {@code current}'s 
read offset, ready to
+     * read this segment's header. Rolls to the next file when the current one 
is exhausted. Returns
+     * false when no segment remains.
+     *
+     * <p>{@code current.segmentStartOffset} is set to where the header 
begins, so a later {@link
+     * SpillSegment#commit()} records the right segment for the snapshot to 
resume from.
+     */
+    private boolean openCurrentFile() throws IOException {
+        if (current.fileIndex >= files.size()) {
+            return false;
+        }
+        openFileAndSeek();
+        if (current.readOffset < currentFileSize) {
+            current.startSegmentHere();
+            return true;
+        }
+        // Current file fully read: move to the next file's first segment.
+        closeFileStream();
+        current.rollToNextFile();
+        if (current.fileIndex >= files.size()) {
+            return false;
+        }
+        openFileAndSeek();
+        if (current.readOffset < currentFileSize) {
+            current.startSegmentHere();
+            return true;
+        }
+        return false;
+    }
+
+    /** Reads the 12-byte header at the current read offset; advances past it. 
*/
+    private SegmentHeader readHeaderAtCurrent() throws IOException {
+        byte[] headerBytes = new byte[SEGMENT_HEADER_BYTES];
+        readFully(headerBytes);
+        DataInputStream h = new DataInputStream(new 
ByteArrayInputStream(headerBytes));
+        int gateIdx = h.readInt();
+        int channelIdx = h.readInt();
+        int bufferLength = h.readInt();
+        checkState(bufferLength >= 0, "negative segment length: %s", 
bufferLength);
+        return new SegmentHeader(new InputChannelInfo(gateIdx, channelIdx), 
bufferLength);
+    }
+
+    /**
+     * Ensures the file at {@code current.fileIndex} is open with the stream 
positioned at {@code
+     * current.readOffset}. If a stream is already open it is left as-is: 
sequential reading
+     * guarantees it is already there.
+     */
+    private void openFileAndSeek() throws IOException {
+        if (fileStream != null) {
+            return;
+        }
+        Path path = files.get(current.fileIndex);
+        currentFileSize = Files.size(path);
+        InputStream in = Files.newInputStream(path);
+        try {
+            skipOnStream(in, current.readOffset, path);
+        } catch (IOException e) {
+            in.close();
+            throw e;
+        }
+        fileStream = in;
+    }
+
+    /** Skips {@code count} body bytes on the open stream, advancing the read 
offset. */
+    private void skipBody(long count) throws IOException {
+        if (count > 0) {
+            skipOnStream(fileStream, count, files.get(current.fileIndex));
+            current.advanceReadOffset(count);
+        }
+    }
+
+    /** Skips exactly {@code count} bytes on {@code in}, failing loud if the 
file ends early. */
+    private void skipOnStream(InputStream in, long count, Path path) throws 
IOException {
+        long skipped = 0;
+        while (skipped < count) {
+            long s = in.skip(count - skipped);
+            if (s <= 0) {
+                // skip can return 0 near EOF; read-and-discard as a fallback.
+                if (in.read() < 0) {
+                    throw new EOFException(
+                            "Cannot position to offset " + count + " in spill 
file " + path);
+                }
+                skipped++;
+            } else {
+                skipped += s;
+            }
+        }
+    }
+
+    private void readFully(byte[] buf) throws IOException {

Review Comment:
   Do we still need to catch the IOException after switching to 
IOUtils.readFully?



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to