jarredhj0214 opened a new issue, #13561: URL: https://github.com/apache/gravitino/issues/13561
### What would you like to be improved? The Spark connector must discover visible relational catalog names during driver startup because Spark requires every top-level V2 catalog to be registered through `spark.sql.catalog.<catalog_name>`. The connector currently calls `listCatalogsInfo()` with full details. This resolves properties and initializes server-side catalog wrappers for every visible catalog, even though startup registration only needs the catalog name, type, and provider. With a large number of catalogs, this significantly increases Spark application startup time and Gravitino server load. This is related to Apache Spark’s lack of dynamic discovery for unknown V2 catalog names: https://github.com/apache/spark/issues/58444 ### How should we improve? Add an option to list catalog information without resolving properties. During Spark startup, retrieve only lightweight catalog descriptors containing the name, type, and provider. Do not place these descriptors in the complete catalog cache. Load and cache the complete catalog information only when the catalog is first used. Preserve the existing behavior by including properties by default. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
