[
https://issues.apache.org/jira/browse/SPARK-20328?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15968388#comment-15968388
]
Michael Gummelt edited comment on SPARK-20328 at 4/13/17 11:27 PM:
-------------------------------------------------------------------
Hey [~vanzin], thanks for the response.
Everything you said is correct, but I want to clarify one thing:
> You just need to make the Mesos backend in Spark do that automatically for
> the submitting user.
The problem can't be solved in the Mesos backend. When I fetch delegation
tokens for transmission to Executors in the Mesos backend, there's no problem.
I can set whatever renewer I want.
The problem is that there's a second location where delegation tokens are
fetched: {{HadoopRDD}}. This is entirely separate from the fetching that the
scheduler backends do (either Mesos or YARN). {{HadoopRDD}} tries to fetch
split data, and ultimately calls into {{TokenCache}} in the hadoop library,
which fetches delegation tokens with the renewer set to the YARN
ResourceManager's principal:
https://github.com/apache/hadoop/blob/trunk/hadoop-mapreduce-project/hadoop-mapreduce-client/hadoop-mapreduce-client-core/src/main/java/org/apache/hadoop/mapred/FileInputFormat.java#L213.
I'm currently solving this by setting that config var in {{SparkSubmit}}.
The big question I have, which I suppose is more for the {{hadoop}} team, is
why in the world is {{FileInputFormat}} fetching delegation tokens? AFAICT,
they're not sending those tokens to any other process. They're just fetching
split data directly from the Name Nodes, and there should be no delegation
required.
was (Author: mgummelt):
Hey [~vanzin], thanks for the response.
Everything you said is correct, but I want to clarify one thing:
> You just need to make the Mesos backend in Spark do that automatically for
> the submitting user.
The problem can't be solved in the Mesos backend. When I fetch delegation
tokens for transmission to Executors in the Mesos backend, there's no problem.
I can set whatever renewer I want.
The problem is that there's a second location where delegation tokens are
fetched: {{HadoopRDD}}. This is entirely separate from the fetching that the
scheduler backends do (either Mesos or YARN). {{HadoopRDD}} tries to fetch
split data, and ultimately calls into {{TokenCache}} in the hadoop library,
which fetches delegation tokens with the renewer set to the YARN
ResourceManager's principal:
https://github.com/apache/hadoop/blob/trunk/hadoop-mapreduce-project/hadoop-mapreduce-client/hadoop-mapreduce-client-core/src/main/java/org/apache/hadoop/mapred/FileInputFormat.java#L213.
I'm currently solving this by setting that config var in {{SparkSubmit}}.
The big question I have, which I suppose is more for the {{hadoop}} team, is
why in the world is {{FileInputFormat}} fetching delegation tokens. AFAICT,
they're not sending those tokens to any other process. They're just fetching
split data directly from the Name Nodes, and there should be no delegation
required.
> HadoopRDDs create a MapReduce JobConf, but are not MapReduce jobs
> -----------------------------------------------------------------
>
> Key: SPARK-20328
> URL: https://issues.apache.org/jira/browse/SPARK-20328
> Project: Spark
> Issue Type: Bug
> Components: Spark Core
> Affects Versions: 2.1.0, 2.1.1, 2.1.2
> Reporter: Michael Gummelt
>
> In order to obtain {{InputSplit}} information, {{HadoopRDD}} creates a
> MapReduce {{JobConf}} out of the Hadoop {{Configuration}}:
> https://github.com/apache/spark/blob/master/core/src/main/scala/org/apache/spark/rdd/HadoopRDD.scala#L138
> Semantically, this is a problem because a HadoopRDD does not represent a
> Hadoop MapReduce job. Practically, this is a problem because this line:
> https://github.com/apache/spark/blob/master/core/src/main/scala/org/apache/spark/rdd/HadoopRDD.scala#L194
> results in this MapReduce-specific security code being called:
> https://github.com/apache/hadoop/blob/trunk/hadoop-mapreduce-project/hadoop-mapreduce-client/hadoop-mapreduce-client-core/src/main/java/org/apache/hadoop/mapreduce/security/TokenCache.java#L130,
> which assumes the MapReduce master is configured (e.g. via
> {{yarn.resourcemanager.*}}). If it isn't, an exception is thrown.
> So I'm seeing this exception thrown as I'm trying to add Kerberos support for
> the Spark Mesos scheduler:
> {code}
> Exception in thread "main" java.io.IOException: Can't get Master Kerberos
> principal for use as renewer
> at
> org.apache.hadoop.mapreduce.security.TokenCache.obtainTokensForNamenodesInternal(TokenCache.java:116)
> at
> org.apache.hadoop.mapreduce.security.TokenCache.obtainTokensForNamenodesInternal(TokenCache.java:100)
> at
> org.apache.hadoop.mapreduce.security.TokenCache.obtainTokensForNamenodes(TokenCache.java:80)
> at
> org.apache.hadoop.mapred.FileInputFormat.listStatus(FileInputFormat.java:205)
> at
> org.apache.hadoop.mapred.FileInputFormat.getSplits(FileInputFormat.java:313)
> at org.apache.spark.rdd.HadoopRDD.getPartitions(HadoopRDD.scala:202)
> {code}
> I have a workaround where I set a YARN-specific configuration variable to
> trick {{TokenCache}} into thinking YARN is configured, but this is obviously
> suboptimal.
> The proper fix to this would likely require significant {{hadoop}}
> refactoring to make split information available without going through
> {{JobConf}}, so I'm not yet sure what the best course of action is.
--
This message was sent by Atlassian JIRA
(v6.3.15#6346)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]