jarredhj0214 opened a new issue, #13561:
URL: https://github.com/apache/gravitino/issues/13561

   ### What would you like to be improved?
   
   The Spark connector must discover visible relational catalog names during 
driver startup because Spark requires every top-level V2 catalog to be 
registered through `spark.sql.catalog.<catalog_name>`.
   
   The connector currently calls `listCatalogsInfo()` with full details. This 
resolves properties and initializes server-side catalog wrappers for every 
visible catalog, even though startup registration only needs the catalog name, 
type, and provider.
   
   With a large number of catalogs, this significantly increases Spark 
application startup time and Gravitino server load.
   
   This is related to Apache Spark’s lack of dynamic discovery for unknown V2 
catalog names: https://github.com/apache/spark/issues/58444
   
   ### How should we improve?
   
   Add an option to list catalog information without resolving properties.
   
   During Spark startup, retrieve only lightweight catalog descriptors 
containing the name, type, and provider. Do not place these descriptors in the 
complete catalog cache. Load and cache the complete catalog information only 
when the catalog is first used.
   
   Preserve the existing behavior by including properties by default.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to