slachiewicz commented on code in PR #289: URL: https://github.com/apache/flink-connector-kafka/pull/289#discussion_r3961081834
########## flink-connector-kafka/src/main/java/org/apache/flink/connector/kafka/dynamic/source/enumerator/ReaderRecoveryGate.java: ########## @@ -0,0 +1,112 @@ +/* + * Licensed to the Apache Software Foundation (ASF) under one + * or more contributor license agreements. See the NOTICE file + * distributed with this work for additional information + * regarding copyright ownership. The ASF licenses this file + * to you under the Apache License, Version 2.0 (the + * "License"); you may not use this file except in compliance + * with the License. You may obtain a copy of the License at + * + * http://www.apache.org/licenses/LICENSE-2.0 + * + * Unless required by applicable law or agreed to in writing, software + * distributed under the License is distributed on an "AS IS" BASIS, + * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. + * See the License for the specific language governing permissions and + * limitations under the License. + */ + +package org.apache.flink.connector.kafka.dynamic.source.enumerator; + +import org.apache.flink.annotation.Internal; +import org.apache.flink.connector.kafka.dynamic.source.split.DynamicKafkaSourceSplit; + +import java.util.ArrayList; +import java.util.Collection; +import java.util.Collections; +import java.util.HashMap; +import java.util.HashSet; +import java.util.List; +import java.util.Map; +import java.util.NavigableMap; +import java.util.Set; +import java.util.TreeMap; + +/** + * Tracks the recovery-time reader registration state of the {@link DynamicKafkaSourceEnumerator}. + * + * <p>When the enumerator is restored from checkpointed state, split assignment and metadata update + * events must be deferred until the first metadata discovery has completed and every reader has + * (re-)registered, so that restored reader splits can be redistributed consistently. This class + * owns that gating state; the enumerator remains responsible for acting on it. + */ +@Internal +class ReaderRecoveryGate { Review Comment: Added in `586295b1` as a Javadoc paragraph on `ReaderRecoveryGate`: not thread-safe, coordinator-thread only, naming the `SourceCoordinator` and `runInCoordinatorThread` paths it arrives through. *This comment was created with AI assistance.* ########## flink-connector-kafka/src/main/java/org/apache/flink/connector/kafka/dynamic/source/enumerator/ReaderRecoveryGate.java: ########## @@ -0,0 +1,112 @@ +/* + * Licensed to the Apache Software Foundation (ASF) under one + * or more contributor license agreements. See the NOTICE file + * distributed with this work for additional information + * regarding copyright ownership. The ASF licenses this file + * to you under the Apache License, Version 2.0 (the + * "License"); you may not use this file except in compliance + * with the License. You may obtain a copy of the License at + * + * http://www.apache.org/licenses/LICENSE-2.0 + * + * Unless required by applicable law or agreed to in writing, software + * distributed under the License is distributed on an "AS IS" BASIS, + * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. + * See the License for the specific language governing permissions and + * limitations under the License. + */ + +package org.apache.flink.connector.kafka.dynamic.source.enumerator; + +import org.apache.flink.annotation.Internal; +import org.apache.flink.connector.kafka.dynamic.source.split.DynamicKafkaSourceSplit; + +import java.util.ArrayList; +import java.util.Collection; +import java.util.Collections; +import java.util.HashMap; +import java.util.HashSet; +import java.util.List; +import java.util.Map; +import java.util.NavigableMap; +import java.util.Set; +import java.util.TreeMap; + +/** + * Tracks the recovery-time reader registration state of the {@link DynamicKafkaSourceEnumerator}. + * + * <p>When the enumerator is restored from checkpointed state, split assignment and metadata update + * events must be deferred until the first metadata discovery has completed and every reader has + * (re-)registered, so that restored reader splits can be redistributed consistently. This class + * owns that gating state; the enumerator remains responsible for acting on it. + */ +@Internal +class ReaderRecoveryGate { + + /** Set on restore; cleared once all readers have registered after the first discovery. */ + private boolean initialReaderRegistrationPending; + + /** Splits reported by readers on registration, pending redistribution. */ + private final Map<Integer, List<DynamicKafkaSourceSplit>> pendingReportedSplitsByReader = + new HashMap<>(); + + /** Readers whose metadata update events were deferred during recovery. */ + private final Set<Integer> pendingMetadataUpdateReaders = new HashSet<>(); + + ReaderRecoveryGate(boolean restoredFromCheckpoint) { + this.initialReaderRegistrationPending = restoredFromCheckpoint; + } + + /** Records splits a reader reported on registration; an empty report is ignored. */ + void recordReportedSplits(int subtaskId, List<DynamicKafkaSourceSplit> reportedSplits) { Review Comment: Correct, no issue. The reports are held in a map keyed by subtask id, so a second `addReader` for the same subtask replaces that reader entry instead of accumulating, and the duplicate-owner check in `reassignReportedSplits` cannot be tripped by a reader's own re-report. There was no test, so `586295b1` adds two: `testRepeatedReportForSameReaderReplacesPreviousReport` and `testEmptyRepeatedReportRetainsPreviousReport`. The second pins a pre-existing edge — an empty re-report is ignored, so an earlier non-empty report survives until the drain. `addReader` had the same guard around the same `put` before this extraction, so that is unchanged behaviour, not something introduced here. *This comment was created with AI assistance.* -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
