LGTM +1 on the design draft.
Best Regards, Yuepeng Pan Xuyang <[email protected]> 于2026年7月24日周五 10:04写道: > Hi, Weiqing. > Thanks for driving this FLIP forward again, +1 > > > > > > -- > > Best! > Xuyang > > > > 在 2026-07-17 14:21:21,"Weiqing Yang" <[email protected]> 写道: > >Hi all, > > > >Gentle bump on this. > > > >The proposal is unchanged at a high level. It's been re-aligned with the > >current master, and the earlier overhead concern is addressed (details in > >my previous message). > > > >Shengkai, Alan, Xuyang, Zakelly — thank you for the thorough review last > >time. Yuepeng and Shengkai, thank you for your votes on the earlier round. > >Since you already know this proposal well, I'd really appreciate a quick > >check that the points you raised still sit right in the current version. > > > >If there are no other concerns over the next few days, I'll start a fresh > >[VOTE] thread so the votes reflect the current design. > > > >FLIP doc: > > > https://docs.google.com/document/d/1ZTN_kSxTMXKyJcrtmP6I9wlZmfPkK8748_nA6EVuVA0/edit > > > >Best, > >Weiqing > > > >On Thu, Jul 9, 2026 at 7:20 PM Weiqing Yang <[email protected]> > >wrote: > > > >> Hi all, > >> > >> I'd like to revive this discussion. This FLIP was originally proposed > and > >> went to a vote last year (it received two +1s) [1]. Since the original > >> proposal and vote were about a year ago, I want to revisit the > discussion > >> before starting a fresh vote to ensure the design remains fully aligned > >> with the current master branch. > >> > >> At a high level, the proposal is unchanged from the previous round. The > >> one substantive refinement since the earlier votes is the sampling > approach > >> detailed below, which directly addresses the performance overhead > question > >> Zakelly raised (the last open item on the thread). > >> > >> Overhead (Zakelly's question): Processing time is no longer measured on > >> every invocation. It now uses the same counter-based sampling as Flink's > >> state latency tracking (FLINK-21736). A new > >> `table.exec.udf-metric.sample-interval` option (default 100) means only > >> every Nth invocation is timed. The non-sampled fast path is a single > >> integer increment, and `udfProcessingTime` is now a Histogram > >> (p50/p75/p95/p99) backed by a bounded 128-entry circular buffer. > Combined > >> with the existing `table.exec.udf-metric-enabled` gate (off by default, > >> meaning nothing is registered when disabled), this provides two robust > >> layers of protection against performance overhead. > >> > >> - Proposal doc (updated): [2] > >> > >> - cwiki FLIP-485: [3] > >> > >> - Draft PR: [4] (implements the proposal; kept in draft until we > converge) > >> > >> I welcome your renewed feedback and any questions. If there are no > further > >> concerns, I'll start a fresh [VOTE] thread on the updated proposal so > the > >> final votes reflect the current design. > >> > >> Thanks, > >> > >> Weiqing > >> > >> [1] Previous vote thread: > >> https://lists.apache.org/thread/d0sv36839p5h03t3okv89pco2jy6vbg3 > >> > >> [2] > >> > https://docs.google.com/document/d/1ZTN_kSxTMXKyJcrtmP6I9wlZmfPkK8748_nA6EVuVA0/edit > >> > >> [3] > >> > https://cwiki.apache.org/confluence/spaces/FLINK/pages/373885706/FLIP-485+Add+UDF+Metrics > >> > >> [4] https://github.com/apache/flink/pull/28692 > >> > >> > >> On Wed, Jun 3, 2026 at 4:59 PM Weiqing Yang <[email protected]> > >> wrote: > >> > >>> Hi Zakelly, > >>> > >>> Kindly pinging here to see if you had any remaining concerns regarding > >>> the FLIP. > >>> > >>> If there are no further questions or concerns from anyone, I plan to > >>> close this discussion thread and proceed with the vote thread. > >>> > >>> Thanks, > >>> Weiqing > >>> > >>> On Tue, Mar 3, 2026 at 2:58 PM Weiqing Yang <[email protected]> > >>> wrote: > >>> > >>>> Hi Zakelly, > >>>> > >>>> > >>>> Thanks for the feedback and sorry for the late response - I am now > >>>> picking it back up. > >>>> > >>>> You raised a great point about the performance overhead, referencing > >>>> FLINK-16444 <https://issues.apache.org/jira/browse/FLINK-16444>. I've > >>>> updated the FLIP to adopt the same counter-based sampling approach > used by > >>>> Flink's state latency tracking (FLINK-21736 > >>>> <https://issues.apache.org/jira/browse/FLINK-21736>). Specifically: > >>>> > >>>> 1. New config: table.exec.udf-metric.sample-interval (default: 100 > >>>> [1]) - only every Nth invocation is measured > >>>> 2. Fast path: Non-sampled invocations are a single integer > increment - > >>>> negligible overhead > >>>> 3. Sampled path: System.nanoTime() around the UDF call, stored in a > >>>> DescriptiveStatisticsHistogram with a bounded 128-entry circular > buffer [2] > >>>> 4. Metric type change: udfProcessingTime is now a Histogram (reports > >>>> p50/p75/p95/p99/mean/min/max) instead of the original Gauge > >>>> 5. Exception counting: Not sampled, since exceptions are rare events > >>>> and counting each one has negligible cost > >>>> > >>>> Combined with the existing feature gate (table.exec.udf-metric-enabled > >>>> defaulting to false), users have two layers of protection: the > feature is > >>>> off by default, and when enabled, sampling keeps overhead minimal. > >>>> The updated FLIP is here: link > >>>> < > https://docs.google.com/document/d/1ZTN_kSxTMXKyJcrtmP6I9wlZmfPkK8748_nA6EVuVA0/edit?tab=t.0#heading=h.ljww281maxj1 > > > >>>> > >>>> Would this address your concern? If so, it would be great to have your > >>>> vote on the vote thread [3]. > >>>> > >>>> [1] 100: state.latency-track.sample-interval default value > >>>> > >>>> [2] 128: state.latency-track.history-size default value (line 55), > which > >>>> is the circular buffer size for the DescriptiveStatisticsHistogram > >>>> [3] https://lists.apache.org/thread/d0sv36839p5h03t3okv89pco2jy6vbg3 > >>>> > >>>> Thanks, > >>>> Weiqing > >>>> > >>>> On Thu, Aug 21, 2025 at 12:24 AM Zakelly Lan <[email protected]> > >>>> wrote: > >>>> > >>>>> Hi Weiqing, > >>>>> > >>>>> Sorry for the late reply. And I have one question: > >>>>> > >>>>> I'm wondering whether the UDF processing time is measured for every > >>>>> individual UDF invocation, with the average then reported, or if > >>>>> sampling > >>>>> is used instead? I'm concerned about the potential overhead if we > >>>>> measure > >>>>> every single invocation. We've encountered similar performance issues > >>>>> when > >>>>> implementing state latency tracking [1]. > >>>>> > >>>>> > >>>>> [1] https://issues.apache.org/jira/browse/FLINK-16444 > >>>>> > >>>>> Best, > >>>>> Zakelly > >>>>> > >>>>> On Fri, Aug 15, 2025 at 5:04 AM Weiqing Yang < > [email protected]> > >>>>> wrote: > >>>>> > >>>>> > Cool - I’ll proceed to start the VOTE. > >>>>> > Thanks! > >>>>> > > >>>>> > Weiqing > >>>>> > > >>>>> > On Thu, Aug 14, 2025 at 12:53 AM Shengkai Fang <[email protected]> > >>>>> wrote: > >>>>> > > >>>>> > > I don't have any more comments. > >>>>> > > > >>>>> > > Best, > >>>>> > > Shengkai > >>>>> > > > >>>>> > > Weiqing Yang <[email protected]> 于2025年8月14日周四 14:47写道: > >>>>> > > > >>>>> > > > Thanks, Shengkai. I’ve updated the proposal doc with the > >>>>> recommended > >>>>> > > > configuration name. Please let me know if you have any > additional > >>>>> > > feedback. > >>>>> > > > > >>>>> > > > Best, > >>>>> > > > Weiqing > >>>>> > > > > >>>>> > > > On Wed, Aug 13, 2025 at 6:58 PM Shengkai Fang < > [email protected]> > >>>>> > wrote: > >>>>> > > > > >>>>> > > > > Sorry for the late response. I prefer to use > >>>>> > > > > `table.exec.udf-metric-enabled` as the option name. > >>>>> > > > > > >>>>> > > > > Best, > >>>>> > > > > Shengkai > >>>>> > > > > > >>>>> > > > > Weiqing Yang <[email protected]> 于2025年8月13日周三 > 23:54写道: > >>>>> > > > > > >>>>> > > > > > Hi Shengkai, Alan, Xuyang, and all, > >>>>> > > > > > > >>>>> > > > > > Since there have been no further objections, I’ll proceed > to > >>>>> start > >>>>> > > the > >>>>> > > > > VOTE > >>>>> > > > > > on this proposal shortly. > >>>>> > > > > > > >>>>> > > > > > Thanks, > >>>>> > > > > > Weiqing > >>>>> > > > > > > >>>>> > > > > > On Thu, Jul 31, 2025 at 10:26 PM Weiqing Yang < > >>>>> > > > [email protected]> > >>>>> > > > > > wrote: > >>>>> > > > > > > >>>>> > > > > > > Hi Shengkai, Alan and Xuyang, > >>>>> > > > > > > > >>>>> > > > > > > Just checking in - do you have any concerns or feedback? > >>>>> > > > > > > > >>>>> > > > > > > If there are no further objections from anyone, I’ll mark > >>>>> the > >>>>> > FLIP > >>>>> > > as > >>>>> > > > > > > ready for voting. > >>>>> > > > > > > > >>>>> > > > > > > > >>>>> > > > > > > Best, > >>>>> > > > > > > Weiqing > >>>>> > > > > > > > >>>>> > > > > > > > >>>>> > > > > > > On Mon, Jul 14, 2025 at 9:10 PM Weiqing Yang < > >>>>> > > > [email protected] > >>>>> > > > > > > >>>>> > > > > > > wrote: > >>>>> > > > > > > > >>>>> > > > > > >> Hi Xuyang, > >>>>> > > > > > >> > >>>>> > > > > > >> Thank you for reviewing the proposal! > >>>>> > > > > > >> > >>>>> > > > > > >> I’m planning to use: *udf.metrics.process-time* and > >>>>> > > > > > >> *udf.metrics.exception-count*. These follow the naming > >>>>> > convention > >>>>> > > > used > >>>>> > > > > > >> in Flink (e.g., RocksDB native metrics > >>>>> > > > > > >> < > >>>>> > > > > > > >>>>> > > > > > >>>>> > > > > >>>>> > > > >>>>> > > >>>>> > https://nightlies.apache.org/flink/flink-docs-master/docs/deployment/config/#rocksdb-native-metrics > >>>>> > > > > > >). > >>>>> > > > > > >> I’ve added these names to the proposal doc. > >>>>> > > > > > >> > >>>>> > > > > > >> Alternatively, I also considered: > >>>>> > > *metrics.udf.process-time.enabled* > >>>>> > > > > and > >>>>> > > > > > >> *metrics.udf.exception-count.enabled. * > >>>>> > > > > > >> > >>>>> > > > > > >> Happy to hear any feedback on which style might be more > >>>>> > > appropriate. > >>>>> > > > > > >> > >>>>> > > > > > >> > >>>>> > > > > > >> Best, > >>>>> > > > > > >> Weiqing > >>>>> > > > > > >> > >>>>> > > > > > >> On Mon, Jul 14, 2025 at 2:55 AM Xuyang < > [email protected] > >>>>> > > >>>>> > > wrote: > >>>>> > > > > > >> > >>>>> > > > > > >>> Hi, Weiqing. > >>>>> > > > > > >>> > >>>>> > > > > > >>> Thanks for driving to improve this. I just have one > >>>>> question. I > >>>>> > > > > notice > >>>>> > > > > > a > >>>>> > > > > > >>> new configuration is introduced in this flip. I just > >>>>> wonder > >>>>> > what > >>>>> > > > the > >>>>> > > > > > >>> configuration name is. Could you please include the > full > >>>>> name > >>>>> > of > >>>>> > > > this > >>>>> > > > > > >>> configuration? (just similar to the other names in > >>>>> > > MetricOptions?) > >>>>> > > > > > >>> > >>>>> > > > > > >>> > >>>>> > > > > > >>> > >>>>> > > > > > >>> > >>>>> > > > > > >>> -- > >>>>> > > > > > >>> > >>>>> > > > > > >>> Best! > >>>>> > > > > > >>> Xuyang > >>>>> > > > > > >>> > >>>>> > > > > > >>> > >>>>> > > > > > >>> > >>>>> > > > > > >>> > >>>>> > > > > > >>> > >>>>> > > > > > >>> 在 2025-07-13 12:03:59,"Weiqing Yang" < > >>>>> [email protected] > >>>>> > > > >>>>> > > > 写道: > >>>>> > > > > > >>> >Hi Alan, > >>>>> > > > > > >>> > > >>>>> > > > > > >>> >Thanks for reviewing the proposal and for highlighting > >>>>> the > >>>>> > > > > ASYNC_TABLE > >>>>> > > > > > >>> work. > >>>>> > > > > > >>> > > >>>>> > > > > > >>> >Yes, I’ve updated the proposal to cover both > >>>>> ASYNC_SCALAR and > >>>>> > > > > > >>> ASYNC_TABLE. > >>>>> > > > > > >>> >For async UDFs, the plan is to instrument both the > >>>>> > invokeAsync() > >>>>> > > > > call > >>>>> > > > > > >>> and > >>>>> > > > > > >>> >the async callback handler to measure the full > end-to-end > >>>>> > > latency > >>>>> > > > > > until > >>>>> > > > > > >>> the > >>>>> > > > > > >>> >result or error is returned from the future. > >>>>> > > > > > >>> > > >>>>> > > > > > >>> >Let me know if you have any further questions or > >>>>> suggestions. > >>>>> > > > > > >>> > > >>>>> > > > > > >>> >Best, > >>>>> > > > > > >>> >Weiqing > >>>>> > > > > > >>> > > >>>>> > > > > > >>> >On Thu, Jul 10, 2025 at 4:15 PM Alan Sheinberg > >>>>> > > > > > >>> ><[email protected]> wrote: > >>>>> > > > > > >>> > > >>>>> > > > > > >>> >> Hi Weiqing, > >>>>> > > > > > >>> >> > >>>>> > > > > > >>> >> From your doc, the entrypoint for UDF calls in the > >>>>> codegen > >>>>> > is > >>>>> > > > > > >>> >> ExprCodeGenerator which should invoke > >>>>> > > > BridgingSqlFunctionCallGen, > >>>>> > > > > > >>> which > >>>>> > > > > > >>> >> could be instrumented with metrics. This works well > >>>>> for > >>>>> > > > > synchronous > >>>>> > > > > > >>> calls, > >>>>> > > > > > >>> >> but what about ASYNC_SCALAR and the soon to be > merged > >>>>> > > > ASYNC_TABLE > >>>>> > > > > ( > >>>>> > > > > > >>> >> https://github.com/apache/flink/pull/26567)? > Timing > >>>>> > metrics > >>>>> > > > > would > >>>>> > > > > > >>> only > >>>>> > > > > > >>> >> account for what it takes to call invokeAsync, not > for > >>>>> the > >>>>> > > > result > >>>>> > > > > to > >>>>> > > > > > >>> >> complete (with a result or error from the future > >>>>> object). > >>>>> > > > > > >>> >> > >>>>> > > > > > >>> >> There are appropriate places which can handle the > async > >>>>> > > > callbacks, > >>>>> > > > > > >>> but they > >>>>> > > > > > >>> >> are in other locations. Will you be able to support > >>>>> those > >>>>> > as > >>>>> > > > > well? > >>>>> > > > > > >>> >> > >>>>> > > > > > >>> >> Thanks, > >>>>> > > > > > >>> >> Alan > >>>>> > > > > > >>> >> > >>>>> > > > > > >>> >> On Wed, Jul 9, 2025 at 7:52 PM Shengkai Fang < > >>>>> > > [email protected] > >>>>> > > > > > >>>>> > > > > > >>> wrote: > >>>>> > > > > > >>> >> > >>>>> > > > > > >>> >> > I just have some questions: > >>>>> > > > > > >>> >> > > >>>>> > > > > > >>> >> > 1. The current metrics hierarchy shows that the > UDF > >>>>> metric > >>>>> > > > group > >>>>> > > > > > >>> belongs > >>>>> > > > > > >>> >> to > >>>>> > > > > > >>> >> > the TaskMetricGroup. I think it would be better > for > >>>>> the > >>>>> > UDF > >>>>> > > > > metric > >>>>> > > > > > >>> group > >>>>> > > > > > >>> >> to > >>>>> > > > > > >>> >> > belong to the OperatorMetricGroup instead, > because a > >>>>> UDF > >>>>> > > might > >>>>> > > > > be > >>>>> > > > > > >>> used by > >>>>> > > > > > >>> >> > multiple operators. > >>>>> > > > > > >>> >> > 2. What are the naming conventions for UDF > metrics? > >>>>> Could > >>>>> > > you > >>>>> > > > > > >>> provide an > >>>>> > > > > > >>> >> > example? Do the metric name contains the UDF name? > >>>>> > > > > > >>> >> > 3. Why is the UDFExceptionCount metric introduced? > >>>>> If a > >>>>> > UDF > >>>>> > > > > throws > >>>>> > > > > > >>> an > >>>>> > > > > > >>> >> > exception, the job fails immediately. Why do we > need > >>>>> to > >>>>> > > track > >>>>> > > > > this > >>>>> > > > > > >>> value? > >>>>> > > > > > >>> >> > > >>>>> > > > > > >>> >> > Best > >>>>> > > > > > >>> >> > Shengkai > >>>>> > > > > > >>> >> > > >>>>> > > > > > >>> >> > > >>>>> > > > > > >>> >> > Weiqing Yang <[email protected]> > 于2025年7月9日周三 > >>>>> > > 12:59写道: > >>>>> > > > > > >>> >> > > >>>>> > > > > > >>> >> > > Hi all, > >>>>> > > > > > >>> >> > > > >>>>> > > > > > >>> >> > > I’d like to initiate a discussion about adding > UDF > >>>>> > > metrics. > >>>>> > > > > > >>> >> > > > >>>>> > > > > > >>> >> > > *Motivation* > >>>>> > > > > > >>> >> > > > >>>>> > > > > > >>> >> > > User-defined functions (UDFs) are essential for > >>>>> custom > >>>>> > > logic > >>>>> > > > > in > >>>>> > > > > > >>> Flink > >>>>> > > > > > >>> >> > jobs > >>>>> > > > > > >>> >> > > but often act as black boxes, making debugging > and > >>>>> > > > performance > >>>>> > > > > > >>> tuning > >>>>> > > > > > >>> >> > > difficult. When issues like high latency or > >>>>> frequent > >>>>> > > > > exceptions > >>>>> > > > > > >>> occur, > >>>>> > > > > > >>> >> > it's > >>>>> > > > > > >>> >> > > hard to pinpoint the root cause inside UDFs. > >>>>> > > > > > >>> >> > > > >>>>> > > > > > >>> >> > > Flink currently lacks built-in metrics for key > UDF > >>>>> > aspects > >>>>> > > > > such > >>>>> > > > > > as > >>>>> > > > > > >>> >> > > per-record processing time or exception count. > This > >>>>> > limits > >>>>> > > > > > >>> >> observability > >>>>> > > > > > >>> >> > > and complicates: > >>>>> > > > > > >>> >> > > > >>>>> > > > > > >>> >> > > - Debugging production issues > >>>>> > > > > > >>> >> > > - Performance tuning and resource allocation > >>>>> > > > > > >>> >> > > - Supplying reliable signals to autoscaling > >>>>> systems > >>>>> > > > > > >>> >> > > > >>>>> > > > > > >>> >> > > Introducing standard, opt-in UDF metrics will > >>>>> improve > >>>>> > > > platform > >>>>> > > > > > >>> >> > > observability and overall health. > >>>>> > > > > > >>> >> > > Here’s the proposal document: Link > >>>>> > > > > > >>> >> > > < > >>>>> > > > > > >>> >> > > > >>>>> > > > > > >>> >> > > >>>>> > > > > > >>> >> > >>>>> > > > > > >>> > >>>>> > > > > > > >>>>> > > > > > >>>>> > > > > >>>>> > > > >>>>> > > >>>>> > https://docs.google.com/document/d/1ZTN_kSxTMXKyJcrtmP6I9wlZmfPkK8748_nA6EVuVA0/edit?tab=t.0#heading=h.ljww281maxj1 > >>>>> > > > > > >>> >> > > > > >>>>> > > > > > >>> >> > > > >>>>> > > > > > >>> >> > > Your feedback and ideas are welcome to refine > this > >>>>> > > feature. > >>>>> > > > > > >>> >> > > > >>>>> > > > > > >>> >> > > > >>>>> > > > > > >>> >> > > Thanks, > >>>>> > > > > > >>> >> > > Weiqing > >>>>> > > > > > >>> >> > > > >>>>> > > > > > >>> >> > > >>>>> > > > > > >>> >> > >>>>> > > > > > >>> > >>>>> > > > > > >> > >>>>> > > > > > > >>>>> > > > > > >>>>> > > > > >>>>> > > > >>>>> > > >>>>> > >>>> >
