[jira] [Commented] (FLINK-5803) Add [partitioned] processing time OVER RANGE BETWEEN UNBOUNDED PRECEDING aggregation to SQL

ASF GitHub Bot (JIRA) Fri, 24 Feb 2017 08:12:57 -0800

    [ 
https://issues.apache.org/jira/browse/FLINK-5803?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15882982#comment-15882982
 ]


ASF GitHub Bot commented on FLINK-5803:
---------------------------------------

Github user fhueske commented on a diff in the pull request:

    https://github.com/apache/flink/pull/3397#discussion_r102968906
  
    --- Diff: 
flink-libraries/flink-table/src/main/scala/org/apache/flink/table/runtime/aggregate/UnboundedProcessingOverProcessFunction.scala
 ---
    @@ -0,0 +1,105 @@
    +/*
    + * Licensed to the Apache Software Foundation (ASF) under one
    + * or more contributor license agreements.  See the NOTICE file
    + * distributed with this work for additional information
    + * regarding copyright ownership.  The ASF licenses this file
    + * to you under the Apache License, Version 2.0 (the
    + * "License"); you may not use this file except in compliance
    + * with the License.  You may obtain a copy of the License at
    + *
    + *     http://www.apache.org/licenses/LICENSE-2.0
    + *
    + * Unless required by applicable law or agreed to in writing, software
    + * distributed under the License is distributed on an "AS IS" BASIS,
    + * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
    + * See the License for the specific language governing permissions and
    + * limitations under the License.
    + */
    +package org.apache.flink.table.runtime.aggregate
    +
    +import org.apache.flink.api.common.typeinfo.TypeInformation
    +import org.apache.flink.configuration.Configuration
    +import org.apache.flink.streaming.api.functions.ProcessFunction.{Context, 
OnTimerContext}
    +import org.apache.flink.types.Row
    +import org.apache.flink.streaming.api.functions.RichProcessFunction
    +import org.apache.flink.util.{Collector, Preconditions}
    +import org.apache.flink.api.common.typeutils.TypeSerializer
    +import org.apache.flink.api.common.state.ValueStateDescriptor
    +import org.apache.flink.api.java.typeutils.RowTypeInfo
    +import org.apache.flink.api.common.state.ValueState
    +import org.apache.flink.api.java.typeutils.runtime.RowSerializer
    +
    +class UnboundedProcessingOverProcessFunction(
    +    private val aggregates: Array[Aggregate[_]],
    +    private val projectionsMapping: Array[(Int, Int)],
    +    private val aggregateMapping: Array[(Int, Int)],
    +    private val  intermediateRowType: RowTypeInfo,
    +  @transient private val returnType: TypeInformation[Row])
    +  extends RichProcessFunction[Row, Row]{
    +
    +  protected var stateSerializer: TypeSerializer[Row] = _
    +  protected var stateDescriptor: ValueStateDescriptor[Row] = _
    +
    +  private var output: Row = _
    +  private var state: ValueState[Row] = _
    +
    +  override def open(config: Configuration) {
    +    Preconditions.checkNotNull(aggregates)
    +    Preconditions.checkNotNull(projectionsMapping)
    +    Preconditions.checkNotNull(aggregateMapping)
    +    Preconditions.checkArgument(aggregates.length == 
aggregateMapping.length)
    +
    +    val finalRowLength: Int = projectionsMapping.length + 
aggregateMapping.length
    +    output = new Row(finalRowLength)
    +    stateSerializer = 
intermediateRowType.createSerializer(getRuntimeContext.getExecutionConfig)
    +    stateDescriptor = new ValueStateDescriptor[Row]("overState", 
stateSerializer)
    +
    +  }
    +
    +  override def processElement(
    +    value2: Row,
    +    ctx: Context,
    +    out: Collector[Row]): Unit = {
    +    state = getRuntimeContext.getState(stateDescriptor)
    --- End diff --
    
    We can do the `getState()` in the `open()` method and store the state in a 
`transitive` member variable.


> Add [partitioned] processing time OVER RANGE BETWEEN UNBOUNDED PRECEDING 
> aggregation to SQL
> -------------------------------------------------------------------------------------------
>
>                 Key: FLINK-5803
>                 URL: https://issues.apache.org/jira/browse/FLINK-5803
>             Project: Flink
>          Issue Type: Sub-task
>          Components: Table API & SQL
>            Reporter: sunjincheng
>            Assignee: sunjincheng
>
> The goal of this issue is to add support for OVER RANGE aggregations on 
> processing time streams to the SQL interface.
> Queries similar to the following should be supported:
> {code}
> SELECT 
>   a, 
>   SUM(b) OVER (PARTITION BY c ORDER BY procTime() RANGE BETWEEN UNBOUNDED 
> PRECEDING AND CURRENT ROW) AS sumB,
>   MIN(b) OVER (PARTITION BY c ORDER BY procTime() RANGE BETWEEN UNBOUNDED 
> PRECEDING AND CURRENT ROW) AS minB
> FROM myStream
> {code}
> The following restrictions should initially apply:
> - All OVER clauses in the same SELECT clause must be exactly the same.
> - The ORDER BY clause may only have procTime() as parameter. procTime() is a 
> parameterless scalar function that just indicates processing time mode.
> - bounded PRECEDING is not supported (see FLINK-5654)
> - FOLLOWING is not supported.
> The restrictions will be resolved in follow up issues. If we find that some 
> of the restrictions are trivial to address, we can add the functionality in 
> this issue as well.
> This issue includes:
> - Design of the DataStream operator to compute OVER ROW aggregates
> - Translation from Calcite's RelNode representation (LogicalProject with 
> RexOver expression).



--
This message was sent by Atlassian JIRA
(v6.3.15#6346)

[jira] [Commented] (FLINK-5803) Add [partitioned] processing time OVER RANGE BETWEEN UNBOUNDED PRECEDING aggregation to SQL

Reply via email to