[ 
https://issues.apache.org/jira/browse/CASSANDRA-21655?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Jon Haddad updated CASSANDRA-21655:
-----------------------------------
    Attachment: flame-cpu-forward.html
                flame-cpu-reverse.html

> async compaction pipeline
> -------------------------
>
>                 Key: CASSANDRA-21655
>                 URL: https://issues.apache.org/jira/browse/CASSANDRA-21655
>             Project: Apache Cassandra
>          Issue Type: Improvement
>            Reporter: Jon Haddad
>            Priority: Normal
>         Attachments: flame-cpu-forward.html, flame-cpu-reverse.html
>
>
> Compaction compresses each chunk, writes it, and updates its CRC serially on 
> the compaction thread, inside \{{SequentialWriter.doFlush}}.  A CPU profile 
> of a 4.5 GiB compaction puts that write side at 37.2% of the total, so the 
> thread that merges rows spends over a third of its time not merging.  
> \{{trickle_fsync}} compounds it: the 10 MiB byte interval fires about 450 
> times over that compaction, and each \{{fdatasync}} forces the whole growing 
> file rather than a range.
> {noformat}
> compaction, 100%
> |<---------------- merge side, 62.8% ---------------->|<---- doFlush, 37.2% 
> ---->|
>                                                       |<-LZ4 21.1%->|<CRC 
> 7.6%>|io|
>                                                                               
>  8.5%
> {noformat}
> Compression plus CRC is 77% of the write side.  Moving the whole 37.2% off 
> the compaction thread makes the merge side the slowest stage, a ceiling of 
> 1.6x.
> The patch gives each SSTable writer its own writer thread and hands chunks to 
> it through a pool of reusable chunk-sized off-heap slots, 4 MiB by default; 
> the producer takes a fresh slot instead of waiting for the one it just 
> filled.  Compression, the write and the checksum all run on that thread.  The 
> byte-interval fsync is replaced by a periodic force, about six per compaction 
> instead of hundreds, which removes work rather than merely overlapping it.  
> The pipeline is a collaborator both \{{CompressedSequentialWriter}} and 
> \{{DirectCompressedSequentialWriter}} own, so direct IO gets it too, and 
> \{{DataComponent.buildWriter}} routes flush, streaming and index builds 
> through the same path.  It sits behind \{{async_compaction_writer_enabled}}.
> Measured on a 4543 MiB dataset across 16 SSTables, 16 cores, JDK 21, 
> throttling off:
> ||path||before||after||change||
> |cursor|15809 ms, 287 MiB/s|5952 ms, 763 MiB/s|+166%|
> |iterator|18476 ms, 246 MiB/s|9129 ms, 498 MiB/s|+102%|
> The result beats the 1.6x ceiling because that ceiling assumed a fixed total; 
> the redundant fsyncs were part of the problem.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to