pithecuse527 opened a new issue, #13454: URL: https://github.com/apache/gravitino/issues/13454
### Version main branch ### Describe what's wrong Naming the fileset in spark.yarn.access.hadoopFileSystems should let the Spark job read it successfully. <img width="2036" height="867" alt="Image" src="https://github.com/user-attachments/assets/9a716abe-0ab5-4d84-b404-9bf70f9c9c04" /> ### Root cause - Spark collects delegation tokens before the job is submitted. `BaseGVFSOperations.addDelegationTokensForAllFS` iterates the file system cache, which is only populated when a file is actually accessed. - At collection time the cache is empty, so no token is returned, and the executors later reach the non-default filesystem with no credentials. ### Error message and/or stacktrace ``` java.io.IOException: DestHost:destPort <namenode-host>:8020 , LocalHost:localPort <executor-host>:0. Failed on local exception: java.io.IOException: org.apache.hadoop.security.AccessControlException: Client cannot authenticate via:[TOKEN, KERBEROS] at org.apache.hadoop.net.NetUtils.wrapWithMessage(NetUtils.java:930) at org.apache.hadoop.net.NetUtils.wrapException(NetUtils.java:905) at org.apache.hadoop.ipc.Client.getRpcResponse(Client.java:1571) at org.apache.hadoop.ipc.Client.call(Client.java:1513) at org.apache.hadoop.ipc.Client.call(Client.java:1410) at org.apache.hadoop.ipc.ProtobufRpcEngine2$Invoker.invoke(ProtobufRpcEngine2.java:258) at jdk.proxy2/jdk.proxy2.$Proxy35.getBlockLocations(Unknown Source) at org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolTranslatorPB.getBlockLocations(ClientNamenodeProtocolTranslatorPB.java:335) ... ``` The job then fails with: ``` org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 0.0 failed 4 times ``` ### How to reproduce - Gravitino (1.3.0 or master version) - Spark 3.x on YARN client mode - Kerberized HDFS. `fs.defaultFS` is a nameservice - The fileset is an EXTERNAL fileset whose storage location points at a namenode directly (`hdfs://<namenode-host>:8020/...`), so its authority **differs from the default filesystem.** | `spark.yarn.access.hadoopFileSystems` | Result | | --- | --- | | (unset) | FAIL — `java.io.IOException: DestHost:destPort <namenode-host>:8020` | | `gvfs://fileset/<catalog>/<schema>/<fileset>` | FAIL — same error | | `hdfs://<namenode-host>:8020` | OK | The second row is the defect: naming the Fileset behaves exactly as if the property were unset. ### Additional context If it's okay, i'd like to propose the patch -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
