[ 
https://issues.apache.org/jira/browse/FLINK-40584?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ASF GitHub Bot updated FLINK-40584:
-----------------------------------
    Labels: pull-request-available  (was: )

> Bound RocksDB prefix iterators to avoid scanning unrelated tombstones
> ---------------------------------------------------------------------
>
>                 Key: FLINK-40584
>                 URL: https://issues.apache.org/jira/browse/FLINK-40584
>             Project: Flink
>          Issue Type: Bug
>          Components: Runtime / State Backends
>            Reporter: Aleksandr Savonin
>            Assignee: Aleksandr Savonin
>            Priority: Major
>              Labels: pull-request-available
>
> The RocksDB state backend creates all prefix iterators with a plain {{new 
> ReadOptions()}} and never sets {{iterate_upper_bound}}. A RocksDB 
> {{Seek(prefix)}} returns the first *visible* key at or after the prefix. 
> Without an upper bound it keeps stepping over deletion markers and 
> overwritten entries until it finds a visible key anywhere later in the column 
> family, or reaches the end. Flink checks the key prefix in Java only after 
> the seek has returned, so that work cannot be avoided.
> Restore after a downscale: the timer service constructor 
> ({{KeyGroupPartitionedPriorityQueue}} peeking every key group) issues one 
> seek per key group, i.e. tens of thousands of seeks at a typical max 
> parallelism. If the timer column family holds tens of megabytes of tombstones 
> and no live timers, which is the normal state right after a large batch of 
> timers has fired (for example after a bulk load of the versioned table), 
> every seek scans to the end of the column family. 



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to