On the producer side, I make sure data for a specific user lands on the same partition. On the consumer side, I use a regular Spark kafka readstream and read the data. I also use a console write stream to print out the spark kafka DataFrame. What I observer is, the data for a specific user (even though in the same partition) arrives out of order in the console.
I also verified the data ordering by running a simple Kafka consumer in Java and the data seems to be ordered. What am I missing here ? Thanks, JK -- Sent from: http://apache-spark-user-list.1001560.n3.nabble.com/ --------------------------------------------------------------------- To unsubscribe e-mail: user-unsubscr...@spark.apache.org