[ https://issues.apache.org/jira/browse/HIVE-22661?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17009047#comment-17009047 ]
Hive QA commented on HIVE-22661: -------------------------------- | (/) *{color:green}+1 overall{color}* | \\ \\ || Vote || Subsystem || Runtime || Comment || || || || || {color:brown} Prechecks {color} || | {color:green}+1{color} | {color:green} @author {color} | {color:green} 0m 0s{color} | {color:green} The patch does not contain any @author tags. {color} | || || || || {color:brown} master Compile Tests {color} || | {color:blue}0{color} | {color:blue} mvndep {color} | {color:blue} 1m 33s{color} | {color:blue} Maven dependency ordering for branch {color} | | {color:green}+1{color} | {color:green} mvninstall {color} | {color:green} 7m 31s{color} | {color:green} master passed {color} | | {color:green}+1{color} | {color:green} compile {color} | {color:green} 1m 45s{color} | {color:green} master passed {color} | | {color:green}+1{color} | {color:green} checkstyle {color} | {color:green} 0m 56s{color} | {color:green} master passed {color} | | {color:blue}0{color} | {color:blue} findbugs {color} | {color:blue} 3m 55s{color} | {color:blue} ql in master has 1531 extant Findbugs warnings. {color} | | {color:blue}0{color} | {color:blue} findbugs {color} | {color:blue} 0m 40s{color} | {color:blue} itests/hive-unit in master has 2 extant Findbugs warnings. {color} | | {color:green}+1{color} | {color:green} javadoc {color} | {color:green} 1m 26s{color} | {color:green} master passed {color} | || || || || {color:brown} Patch Compile Tests {color} || | {color:blue}0{color} | {color:blue} mvndep {color} | {color:blue} 0m 26s{color} | {color:blue} Maven dependency ordering for patch {color} | | {color:green}+1{color} | {color:green} mvninstall {color} | {color:green} 2m 10s{color} | {color:green} the patch passed {color} | | {color:green}+1{color} | {color:green} compile {color} | {color:green} 1m 47s{color} | {color:green} the patch passed {color} | | {color:green}+1{color} | {color:green} javac {color} | {color:green} 1m 47s{color} | {color:green} the patch passed {color} | | {color:green}+1{color} | {color:green} checkstyle {color} | {color:green} 0m 58s{color} | {color:green} the patch passed {color} | | {color:green}+1{color} | {color:green} whitespace {color} | {color:green} 0m 0s{color} | {color:green} The patch has no whitespace issues. {color} | | {color:green}+1{color} | {color:green} findbugs {color} | {color:green} 4m 44s{color} | {color:green} the patch passed {color} | | {color:green}+1{color} | {color:green} javadoc {color} | {color:green} 1m 25s{color} | {color:green} the patch passed {color} | || || || || {color:brown} Other Tests {color} || | {color:green}+1{color} | {color:green} asflicense {color} | {color:green} 0m 13s{color} | {color:green} The patch does not generate ASF License warnings. {color} | | {color:black}{color} | {color:black} {color} | {color:black} 30m 10s{color} | {color:black} {color} | \\ \\ || Subsystem || Report/Notes || | Optional Tests | asflicense javac javadoc findbugs checkstyle compile | | uname | Linux hiveptest-server-upstream 3.16.0-4-amd64 #1 SMP Debian 3.16.43-2+deb8u5 (2017-09-19) x86_64 GNU/Linux | | Build tool | maven | | Personality | /data/hiveptest/working/yetus_PreCommit-HIVE-Build-20081/dev-support/hive-personality.sh | | git revision | master / c6e27ee | | Default Java | 1.8.0_111 | | findbugs | v3.0.1 | | modules | C: ql itests/hive-unit U: . | | Console output | http://104.198.109.242/logs//PreCommit-HIVE-Build-20081/yetus.txt | | Powered by | Apache Yetus http://yetus.apache.org | This message was automatically generated. > Compaction fails on non bucketed table with data loaded inpath > -------------------------------------------------------------- > > Key: HIVE-22661 > URL: https://issues.apache.org/jira/browse/HIVE-22661 > Project: Hive > Issue Type: Bug > Reporter: Ádám Szita > Assignee: Ádám Szita > Priority: Major > Attachments: HIVE-22661.0.patch, HIVE-22661.1.patch, > HIVE-22661.2.patch > > > Compaction cannot handle situations where: > * data was ingested with {{LOAD DATA INPATH}} > * this ingest method is run multiple times, and > ** with different number of files getting created in the delta directories > Therefore, for file/dir structures such as: > {code:java} > /warehouse/tablespace/managed/hive/comp3/delta_0000001_0000001_0000 > /warehouse/tablespace/managed/hive/comp3/delta_0000001_0000001_0000/000000_0 > /warehouse/tablespace/managed/hive/comp3/delta_0000001_0000001_0000/000001_0 > /warehouse/tablespace/managed/hive/comp3/delta_0000002_0000002_0000 > /warehouse/tablespace/managed/hive/comp3/delta_0000002_0000002_0000/000000_0 > /warehouse/tablespace/managed/hive/comp3/delta_0000002_0000002_0000/000001_0 > /warehouse/tablespace/managed/hive/comp3/delta_0000002_0000002_0000/000002_0 > {code} > Although the table is not bucketed, bucket is calculated from the (raw) > files' names. Compaction in the above case will fail on delta1-1 not having > data for 'bucket' 2. > Steps to repro using small dataset: > {code:java} > set tez.grouping.min-size=8; > set tez.grouping.max-size=8; > set mapreduce.input.fileinputformat.split.minsize=8; > set mapreduce.input.fileinputformat.split.minsize=8; > create external table comp0 (a string); > insert into comp0 values ("qwertyuiopasdfghjklzxcvbnm"); > insert into comp0 values ("qwertyuiopasdfghjklzxcvbnm"); > create external table comp1 stored as orc as select * from comp0; > insert into comp0 values ("qwertyuiopasdfghjklzxcvbnm"); > create external table comp2 stored as orc as select * from comp0; > create table comp3 (a string); > load data inpath '/warehouse/tablespace/external/hive/comp1' into table comp3; > load data inpath '/warehouse/tablespace/external/hive/comp2' into table > comp3;{code} -- This message was sent by Atlassian Jira (v8.3.4#803005)