pithecuse527 opened a new issue, #13454:
URL: https://github.com/apache/gravitino/issues/13454

   ### Version
   
   main branch
   
   ### Describe what's wrong
   
   Naming the fileset in spark.yarn.access.hadoopFileSystems should let the 
Spark job read it successfully.
   
   <img width="2036" height="867" alt="Image" 
src="https://github.com/user-attachments/assets/9a716abe-0ab5-4d84-b404-9bf70f9c9c04";
 />
   
   ### Root cause
   - Spark collects delegation tokens before the job is submitted. 
`BaseGVFSOperations.addDelegationTokensForAllFS` iterates the file system 
cache, which is only populated when a file is actually accessed.
   - At collection time the cache is empty, so no token is returned, and the 
executors later reach the non-default filesystem with no credentials.
   
   
   
   ### Error message and/or stacktrace
   
   ```
   java.io.IOException: DestHost:destPort <namenode-host>:8020 ,
     LocalHost:localPort <executor-host>:0.
     Failed on local exception: java.io.IOException:
     org.apache.hadoop.security.AccessControlException:
     Client cannot authenticate via:[TOKEN, KERBEROS]
   
         at org.apache.hadoop.net.NetUtils.wrapWithMessage(NetUtils.java:930)
         at org.apache.hadoop.net.NetUtils.wrapException(NetUtils.java:905)
         at org.apache.hadoop.ipc.Client.getRpcResponse(Client.java:1571)
         at org.apache.hadoop.ipc.Client.call(Client.java:1513)
         at org.apache.hadoop.ipc.Client.call(Client.java:1410)
         at 
org.apache.hadoop.ipc.ProtobufRpcEngine2$Invoker.invoke(ProtobufRpcEngine2.java:258)
         at jdk.proxy2/jdk.proxy2.$Proxy35.getBlockLocations(Unknown Source)
         at 
org.apache.hadoop.hdfs.protocolPB.ClientNamenodeProtocolTranslatorPB.getBlockLocations(ClientNamenodeProtocolTranslatorPB.java:335)
         ...
     ```
   
     The job then fails with:
   
     ```
     org.apache.spark.SparkException: Job aborted due to stage failure:
       Task 0 in stage 0.0 failed 4 times
     ```
   
   ### How to reproduce
   
   - Gravitino (1.3.0 or master version)
   - Spark 3.x on YARN client mode
   - Kerberized HDFS. `fs.defaultFS` is a nameservice
   - The fileset is an EXTERNAL fileset whose storage location points at a 
namenode directly (`hdfs://<namenode-host>:8020/...`), so its authority 
**differs from the default filesystem.**
   
     | `spark.yarn.access.hadoopFileSystems` | Result |
     | --- | --- |
     | (unset) | FAIL — `java.io.IOException: DestHost:destPort 
<namenode-host>:8020` |
     | `gvfs://fileset/<catalog>/<schema>/<fileset>` | FAIL — same error |
     | `hdfs://<namenode-host>:8020` | OK |
   
   The second row is the defect: naming the Fileset behaves exactly as if the 
property were unset.
   
   ### Additional context
   
   If it's okay, i'd like to propose the patch


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to