MartijnVisser commented on code in PR #289: URL: https://github.com/apache/flink-connector-kafka/pull/289#discussion_r3967174427
########## flink-connector-kafka/src/main/java/org/apache/flink/connector/kafka/dynamic/source/enumerator/ReaderRecoveryGate.java: ########## @@ -0,0 +1,116 @@ +/* + * Licensed to the Apache Software Foundation (ASF) under one + * or more contributor license agreements. See the NOTICE file + * distributed with this work for additional information + * regarding copyright ownership. The ASF licenses this file + * to you under the Apache License, Version 2.0 (the + * "License"); you may not use this file except in compliance + * with the License. You may obtain a copy of the License at + * + * http://www.apache.org/licenses/LICENSE-2.0 + * + * Unless required by applicable law or agreed to in writing, software + * distributed under the License is distributed on an "AS IS" BASIS, + * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. + * See the License for the specific language governing permissions and + * limitations under the License. + */ + +package org.apache.flink.connector.kafka.dynamic.source.enumerator; + +import org.apache.flink.annotation.Internal; +import org.apache.flink.connector.kafka.dynamic.source.split.DynamicKafkaSourceSplit; + +import java.util.ArrayList; +import java.util.Collection; +import java.util.Collections; +import java.util.HashMap; +import java.util.HashSet; +import java.util.List; +import java.util.Map; +import java.util.NavigableMap; +import java.util.Set; +import java.util.TreeMap; + +/** + * Tracks the recovery-time reader registration state of the {@link DynamicKafkaSourceEnumerator}. + * + * <p>When the enumerator is restored from checkpointed state, split assignment and metadata update + * events must be deferred until the first metadata discovery has completed and every reader has + * (re-)registered, so that restored reader splits can be redistributed consistently. This class + * owns that gating state; the enumerator remains responsible for acting on it. Review Comment: This describes the restore trigger only. The gate is also armed on a running enumerator when a reader re-registers after a partial failover and reports its checkpointed splits (`recordReportedSplits` with a non-empty list while `initialReaderRegistrationPending` is already false). Since making the gating explicit is the point of this class, please name both triggers and what each waits for: a restore waits for the first discovery plus all readers; a reported-splits registration waits for all readers to be registered again. ########## flink-connector-kafka/src/main/java/org/apache/flink/connector/kafka/dynamic/source/enumerator/DynamicKafkaSourceEnumerator.java: ########## @@ -830,7 +819,7 @@ private void reassignReportedSplits() { long currentTimeMillis = System.currentTimeMillis(); for (Entry<Integer, List<DynamicKafkaSourceSplit>> readerSplits : - new TreeMap<>(pendingReportedSplitsByReader).entrySet()) { + readerRecoveryGate.drainReportedSplits().entrySet()) { Review Comment: This is the one spot where the extraction is not a pure move: the pending map used to be cleared at the end of `reassignReportedSplits`, now it is cleared before the loop runs. Not observable today, nothing between here and the end of the method reads the gate, and an exception in here fails the job through the coordinator anyway. It does mean that anyone who later consults the gate from `handleNoMoreSplits`, which runs re-entrantly from inside this loop via the sub-enumerator's no-more-splits callback, sees an empty map. A one-line note on `drainReportedSplits` that the state is cleared eagerly and must not be consulted while reassigning is enough. ########## flink-connector-kafka/src/test/java/org/apache/flink/connector/kafka/dynamic/source/enumerator/ReaderRecoveryGateTest.java: ########## @@ -0,0 +1,138 @@ +/* + * Licensed to the Apache Software Foundation (ASF) under one + * or more contributor license agreements. See the NOTICE file + * distributed with this work for additional information + * regarding copyright ownership. The ASF licenses this file + * to you under the Apache License, Version 2.0 (the + * "License"); you may not use this file except in compliance + * with the License. You may obtain a copy of the License at + * + * http://www.apache.org/licenses/LICENSE-2.0 + * + * Unless required by applicable law or agreed to in writing, software + * distributed under the License is distributed on an "AS IS" BASIS, + * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. + * See the License for the specific language governing permissions and + * limitations under the License. + */ + +package org.apache.flink.connector.kafka.dynamic.source.enumerator; + +import org.apache.flink.connector.kafka.dynamic.source.split.DynamicKafkaSourceSplit; +import org.apache.flink.connector.kafka.source.split.KafkaPartitionSplit; + +import org.apache.kafka.common.TopicPartition; +import org.junit.jupiter.api.Test; + +import java.util.Arrays; +import java.util.Collections; +import java.util.List; +import java.util.NavigableMap; + +import static org.assertj.core.api.Assertions.assertThat; + +/** Tests for {@link ReaderRecoveryGate}. */ +class ReaderRecoveryGateTest { + + @Test + void testFreshStartHasNoPendingRecovery() { + ReaderRecoveryGate gate = new ReaderRecoveryGate(false); + + assertThat(gate.hasPendingRecovery()).isFalse(); + assertThat(gate.shouldDeferMetadataUpdateEvents(false)).isFalse(); + assertThat(gate.shouldDeferMetadataUpdateEvents(true)).isFalse(); + assertThat(gate.hasReportedSplits()).isFalse(); + } + + @Test + void testRestoredStartGatesUntilInitialRegistrationCompletes() { + ReaderRecoveryGate gate = new ReaderRecoveryGate(true); + + assertThat(gate.hasPendingRecovery()).isTrue(); + assertThat(gate.shouldDeferMetadataUpdateEvents(true)).isTrue(); + assertThat(gate.shouldDeferMetadataUpdateEvents(false)).isTrue(); + + gate.markInitialRegistrationComplete(); + + assertThat(gate.hasPendingRecovery()).isFalse(); + assertThat(gate.shouldDeferMetadataUpdateEvents(true)).isFalse(); + } + + @Test + void testReportedSplitsGateUntilAllReadersRegistered() { + ReaderRecoveryGate gate = new ReaderRecoveryGate(false); + gate.recordReportedSplits(1, Collections.singletonList(split("topic", 0))); + + assertThat(gate.hasPendingRecovery()).isTrue(); + assertThat(gate.hasReportedSplits()).isTrue(); + assertThat(gate.shouldDeferMetadataUpdateEvents(false)).isTrue(); + assertThat(gate.shouldDeferMetadataUpdateEvents(true)).isFalse(); + } Review Comment: These cover each field on its own but not the interaction between the restore flag and the reported splits, which is the part the description calls hard to follow. One more test: `new ReaderRecoveryGate(true)`, record splits for one reader, `markInitialRegistrationComplete()`; then `hasPendingRecovery()` is still true, `shouldDeferMetadataUpdateEvents(false)` is true and `shouldDeferMetadataUpdateEvents(true)` is false; after `drainReportedSplits()` all three are false. That pins that completing the initial registration does not release the gate while reported splits are still pending. ########## flink-connector-kafka/src/main/java/org/apache/flink/connector/kafka/dynamic/source/enumerator/ReaderRecoveryGate.java: ########## @@ -0,0 +1,112 @@ +/* + * Licensed to the Apache Software Foundation (ASF) under one + * or more contributor license agreements. See the NOTICE file + * distributed with this work for additional information + * regarding copyright ownership. The ASF licenses this file + * to you under the Apache License, Version 2.0 (the + * "License"); you may not use this file except in compliance + * with the License. You may obtain a copy of the License at + * + * http://www.apache.org/licenses/LICENSE-2.0 + * + * Unless required by applicable law or agreed to in writing, software + * distributed under the License is distributed on an "AS IS" BASIS, + * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. + * See the License for the specific language governing permissions and + * limitations under the License. + */ + +package org.apache.flink.connector.kafka.dynamic.source.enumerator; + +import org.apache.flink.annotation.Internal; +import org.apache.flink.connector.kafka.dynamic.source.split.DynamicKafkaSourceSplit; + +import java.util.ArrayList; +import java.util.Collection; +import java.util.Collections; +import java.util.HashMap; +import java.util.HashSet; +import java.util.List; +import java.util.Map; +import java.util.NavigableMap; +import java.util.Set; +import java.util.TreeMap; + +/** + * Tracks the recovery-time reader registration state of the {@link DynamicKafkaSourceEnumerator}. + * + * <p>When the enumerator is restored from checkpointed state, split assignment and metadata update + * events must be deferred until the first metadata discovery has completed and every reader has + * (re-)registered, so that restored reader splits can be redistributed consistently. This class + * owns that gating state; the enumerator remains responsible for acting on it. + */ +@Internal +class ReaderRecoveryGate { + + /** Set on restore; cleared once all readers have registered after the first discovery. */ + private boolean initialReaderRegistrationPending; + + /** Splits reported by readers on registration, pending redistribution. */ + private final Map<Integer, List<DynamicKafkaSourceSplit>> pendingReportedSplitsByReader = + new HashMap<>(); + + /** Readers whose metadata update events were deferred during recovery. */ + private final Set<Integer> pendingMetadataUpdateReaders = new HashSet<>(); + + ReaderRecoveryGate(boolean restoredFromCheckpoint) { + this.initialReaderRegistrationPending = restoredFromCheckpoint; + } + + /** Records splits a reader reported on registration; an empty report is ignored. */ + void recordReportedSplits(int subtaskId, List<DynamicKafkaSourceSplit> reportedSplits) { Review Comment: Settled, thanks for the tests. For the record: on a subtask failover Flink runs `subtaskReset` (which calls `addSplitsBack` and drops the reader from `registeredReaders`) and then a fresh `addReader`, and, while the gate is still pending, the reader reports the same checkpointed split set both times, so replacing the entry is the right semantics and the duplicate-owner check only ever compares different readers. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
