Michael Gummelt created SPARK-20328:
---------------------------------------
Summary: HadoopRDDs create a MapReduce JobConf, but are not
MapReduce jobs
Key: SPARK-20328
URL: https://issues.apache.org/jira/browse/SPARK-20328
Project: Spark
Issue Type: Bug
Components: Spark Core
Affects Versions: 2.1.0, 2.1.1, 2.1.2
Reporter: Michael Gummelt
In order to obtain {{InputSplit}} information, {{HadoopRDD}} creates a
MapReduce {{JobConf}} out of the Hadoop {{Configuration}}:
https://github.com/apache/spark/blob/master/core/src/main/scala/org/apache/spark/rdd/HadoopRDD.scala#L138
Semantically, this is a problem because a HadoopRDD does not represent a Hadoop
MapReduce job. Practically, this is a problem because this line:
https://github.com/apache/spark/blob/master/core/src/main/scala/org/apache/spark/rdd/HadoopRDD.scala#L194
results in this MapReduce-specific security code being called:
https://github.com/apache/hadoop/blob/trunk/hadoop-mapreduce-project/hadoop-mapreduce-client/hadoop-mapreduce-client-core/src/main/java/org/apache/hadoop/mapreduce/security/TokenCache.java#L130,
which assumes the MapReduce master is configured. If it isn't, an exception
is thrown.
So I'm seeing this exception thrown as I'm trying to add Kerberos support for
the Spark Mesos scheduler. I have a workaround where I set a YARN-specific
configuration variable to trick {{TokenCache}} into thinking YARN is
configured, but this is obviously suboptimal.
The proper fix to this would likely require significant {{hadoop}} refactoring
to make split information available without going through {{JobConf}}, so I'm
not yet sure what the best course of action is.
--
This message was sent by Atlassian JIRA
(v6.3.15#6346)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]