zhangyue19921010 commented on code in PR #10180:
URL: https://github.com/apache/hudi/pull/10180#discussion_r1407409304


##########
hudi-flink-datasource/hudi-flink/src/main/java/org/apache/hudi/configuration/FlinkOptions.java:
##########
@@ -125,6 +125,20 @@ private FlinkOptions() {
       .withDescription("Payload class used. Override this, if you like to roll 
your own merge logic, when upserting/inserting.\n"
           + "This will render any value set for the option in-effective");
 
+  public static final ConfigOption<String> INSERT_PARTITIONER_CLASS_NAME = 
ConfigOptions
+      .key("write.insert.partitioner.class.name")
+      .stringType()
+      .defaultValue("")
+      .withDescription("Insert partitioner to use aiming to re-balance records 
and reducing small file number "
+          + "in the scenario of multi-level partitioning. For example 
dt/hour/eventID"
+          + "Currently support 
org.apache.hudi.sink.partitioner.DefaultInsertPartitioner");
+
+  public static final ConfigOption<Integer> DEFAULT_PARALLELISM_PER_PARTITION 
= ConfigOptions
+      .key("write.insert.partitioner.parallelism.per.partition")
+      .intType()
+      .defaultValue(30)
+      .withDescription("The parallelism to use in each partition when using 
DefaultInsertPartitioner.");

Review Comment:
   `Does this option controls the max number of output files per-partition ?`
   Exactly, unless the file size is larger than `max parquet file size` and 
auto split out.
   
   `And does the default 30 comes from production experience ?`
   Yes, one of our hudi append job, 100,000,000/min, using 950 cores/slots. 
Also set this config to 30



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to