yuqi1129 opened a new issue, #13581:
URL: https://github.com/apache/gravitino/issues/13581

   ## What would you like to be improved?
   
   On main `ab0785324c5cb391ae63e4501d349f4336b8d3aa`, the default change-log 
poller reads at most 2,000 records and waits three seconds after each cycle. 
Its maximum consumption rate is therefore below 667 records/s, including when a 
backlog remains.
   
   In a local two-JVM/MySQL 8.0.35 experiment, 32 independent-object writers 
sustained about 1,165 successful updates/s. Eight separately warmed metalakes 
were each updated once on A and read from B. Five were still stale at the 
15-second test deadline; all eight acknowledged values were verified directly 
in MySQL. Poll failures were zero and the lag gauge grew substantially.
   
   With only `gravitino.entityChangeLog.pollIntervalSecs=1` changed after 
restarting both nodes with cold caches, all eight trials succeeded, with 
maximum visibility delay about 1.26 seconds. The 15-second deadline is an 
experiment threshold, not a claimed product SLA. This is propagation-capacity 
evidence, not evidence of permanent data loss.
   
   ## How should we improve?
   
   When a full batch is fetched, continue draining with a bounded work/time 
budget rather than sleeping for the full normal interval. Keep the ordinary 
delay for an empty or caught-up poll and preserve failure backoff. Consider a 
configurable batch size and document propagation capacity in records/s. Add 
backlog and distinct-entity visibility tests under sustained writes.
   
   Reproduction: the `main-occ-cache-20260928` benchmark's 
`probe_backlog.compare_distinct()` uses two 512 MiB Java 17 JVMs, simple 
authentication, authorization disabled, a dedicated four-CPU MySQL container, 
32 independent writers and eight distinct marker entities. Evidence: 
`probe-visibility-distinct-poll3s.json`, 
`probe-visibility-distinct-poll1s.json`, the corresponding durable TSVs, SQL 
snapshots and Prometheus samples.
   
   Minimal reproduction: start two main servers against the same MySQL with 
cache enabled and default polling; create 64 independent metalakes plus eight 
marker metalakes; run 32 concurrent property-update clients alternating nodes 
and sustain more than 667 updates/s; for each marker, GET on B to warm it, PUT 
a unique property value on A once, then poll B every 50 ms for up to 15 seconds 
while writers continue. Verify the acknowledged marker values directly in 
`metalake_meta`. Restart both servers and repeat with the one-second interval.
   
   Related prior tracking: 
[#12377](https://github.com/apache/gravitino/issues/12377) already proposed 
draining more than one batch per poll cycle. This measured follow-up shows the 
fixed-batch capacity limitation remains on the tested main SHA; it does not 
claim the idea is new.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to