Copilot commented on code in PR #13455:
URL: https://github.com/apache/gravitino/pull/13455#discussion_r4077822878
##########
clients/filesystem-hadoop3/src/main/java/org/apache/gravitino/filesystem/hadoop/BaseGVFSOperations.java:
##########
@@ -449,6 +449,30 @@ public abstract FSDataOutputStream append(Path gvfsPath,
int bufferSize, Progres
*/
public abstract Token<?>[] addDelegationTokens(String renewer, Credentials
credentials);
+ /**
+ * Resolve the given fileset path so that its actual file system is created
before delegation
+ * tokens are collected.
+ *
+ * <p>Delegation tokens are collected by iterating the file system cache,
which is only populated
+ * when a file is actually accessed. A token consumer such as Spark collects
tokens before the job
+ * is submitted, when nothing has been accessed yet, so without this step no
token is returned and
+ * executors fail to reach a fileset that lives outside the default file
system.
+ *
+ * @param filesetPath the virtual path this file system was initialized with.
+ */
+ public void prepareDelegationTokens(Path filesetPath) {
+ if (!fileSystemCache.asMap().isEmpty()) {
+ return;
+ }
Review Comment:
This guard treats any cached filesystem as sufficient, but `fileSystemCache`
is shared across fileset paths and keyed only by the actual storage
scheme/authority/user. After another fileset has been accessed, the cache can
be non-empty while `filesetPath` still has no corresponding filesystem, so
token collection skips this path and Spark still receives no token. Resolve the
requested location or check for its matching cache entry instead of checking
global emptiness.
This issue also appears on line 469 of the same file.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]