nevzheng commented on PR #13553:
URL: https://github.com/apache/gravitino/pull/13553#issuecomment-5881785368

   **Shared-container tests are now faster, not slower — tuning the DB, not the 
architecture, was the missing piece.**
   
   Pushed a follow-up commit (`4b0626ba8`) that tunes the shared 
MySQL/PostgreSQL test containers instead of leaving them on production 
defaults. Result: this PR now wins on **both** axes — half the containers, and 
dramatically faster.
   
   ## The numbers (8vCPU/31GB, matches this repo's CI runner class)
   
   | Backend | Before: 2 containers/fork (old design) | After: 1 shared 
container + tuning (this PR) | Time cut | Speedup |
   |---|---|---|---|---|
   | MySQL | 26m 25s | **7m 05s** | 73.2% | 3.73x |
   | PostgreSQL | 7m 10s | **2m 11s** | 69.4% | 3.27x |
   | H2 (control -- no container, in-memory) | 1m 27s | 1m 21s | 6.5% | 1.07x |
   
   **Column meaning:** "Before" = the pre-PR design, one dedicated container 
per Gradle test fork (2 forks = 2 containers), untuned defaults. "After" = this 
PR's single shared container across both forks, with the tuning below applied. 
"Time cut" = `(before - after) / before * 100`, the percentage of wall-clock 
time removed. "Speedup" = `before / after`, how many times faster (3.73x means 
the tuned run finishes in about a quarter of the original time).
   
   H2 is included as a control: it never used a container (always in-memory), 
so its near-flat time (6.5% cut, within run-to-run noise) confirms the 
MySQL/PostgreSQL wins come from the container tuning itself, not from the host 
machine just being faster this run.
   
   520/520 tests pass on every run. 0 deadlocks, 0 lock-waits. PostgreSQL's 2 
skips are pre-existing -- identical in the "before" run, unrelated to this 
change.
   
   ## What changed
   
   Earlier in this thread we found the shared container was ~5.6% *slower* than 
the old per-fork design -- not an architecture problem, a config problem. MySQL 
and PostgreSQL ship tuned for production crash durability, which costs real I/O 
on every commit. Under 2 forks hammering one shared instance, that cost 
compounds.
   
   Fix: an externalized, checked-in config 
(`buildSrc/.../shared-db-tuning/{mysql,postgresql}.args`, plain text, one flag 
per line, commented) that relaxes durability guarantees for this container only:
   - **MySQL**: `innodb_flush_log_at_trx_commit=0`, `sync_binlog=0`, 
`skip-log-bin`, `innodb_doublewrite=0`, larger buffer pool/log buffer, 
`performance_schema=OFF`
   - **PostgreSQL**: `fsync=off`, `synchronous_commit=off`, 
`full_page_writes=off`, `autovacuum=off`, `wal_level=minimal` (+ 
`max_wal_senders=0`, required alongside it), larger 
`shared_buffers`/`max_wal_size`
   - **Both**: data directory mounted on `tmpfs` -- the single biggest lever, 
removing real disk I/O entirely
   
   Kept as a separate resource file rather than hardcoded flags so it's 
reviewable and adjustable independent of the Java diff.
   
   ## Why this is safe
   
   These settings would be reckless in production. They're fine here because 
**this container tests whether Gravitino's code talks correctly to 
MySQL/PostgreSQL -- schema migrations, query correctness, transaction semantics 
-- not whether the database itself survives a crash.** Durability and 
correctness are orthogonal properties; disabling `fsync` doesn't change what a 
query returns while the process is alive.
   
   Confirmed, not assumed: grepped the full test suite for anything asserting 
crash-recovery/WAL-replay semantics. Nothing does. The container is destroyed 
at the end of every run regardless, so there's no crash to recover from either 
way.
   
   ## Recommendation
   
   No longer "hold" -- ship. Fewer containers and a large, verified speedup, 
with the durability tradeoff scoped to exactly where it's safe. `#13518` 
(fixture-reuse, still unmerged upstream, 4.15x on H2) remains the bigger 
separate opportunity -- worth prioritizing next, but doesn't block this.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to