sql query orc slow

patcharee Thu, 08 Oct 2015 01:44:18 -0700

Hi,

I am using spark sql 1.5 to query a hive table stored as partitioned orcfile. We have the total files is about 6000 files and each file size isabout 245MB.


What is the difference between these two query methods below:

1. Using query on hive table directly

hiveContext.sql("select col1, col2 from table1")

2. Reading from orc file, register temp table and query from the temp table

val c = hiveContext.read.format("orc").load("/apps/hive/warehouse/table1")
c.registerTempTable("regTable")
hiveContext.sql("select col1, col2 from regTable")

When the number of files is large (query all from the total 6000 files), the second case is much slower then the first one. Any ideas why?


BR,




---------------------------------------------------------------------
To unsubscribe, e-mail: user-unsubscr...@spark.apache.org
For additional commands, e-mail: user-h...@spark.apache.org

sql query orc slow

Reply via email to