jarredhj0214 commented on PR #13562:
URL: https://github.com/apache/gravitino/pull/13562#issuecomment-5881865240

   I measured the catalog listing path affected by this PR locally, because the 
test environment does not have many Spark jobs/catalogs and cannot reflect the 
production impact well. In production we have 1,300+ catalogs, so eager loading 
catalog properties during Spark driver initialization has a much larger impact.
   
   Local benchmark with 1,300 catalogs, cold catalog cache, 20 iterations:
   
   - `listCatalogsInfo(false)`: 0.231 ms avg
   - `listCatalogsInfo(true)`: 142.867 ms avg
   
   So the catalog listing part used during Spark driver initialization is about 
617.57x faster, reducing this part by about 99.84%.
   
   The reason is that Spark startup now only fetches lightweight catalog 
descriptors with `includeProperties=false`, and skips eager property resolution 
/ full catalog wrapper loading. Full catalog information is still loaded later 
when the catalog is actually used.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to