[
https://issues.apache.org/jira/browse/SOLR-18415?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18112646#comment-18112646
]
Andrzej Bialecki edited comment on SOLR-18415 at 9/8/26 10:23 AM:
------------------------------------------------------------------
After further analysis this looks somewhat better than I thought, especially
for updates that are less than 8192 byte in size (serialized), because for
these the array is resized to the actual size of the payload (as accumulated in
the {{FastOutputStream}}'s buffer that the {{JavaBinCodec}} uses). In this case
no extra padding bytes are wasted.
Since the producer side ({{MirroringUpdateProcessor / KafkaMirroringSink}})
doesn't accumulate individual requests in batches and in a typical scenario
most documents are smaller than 8kB serialized, this optimization will mostly
hold.
This is not the case when sending larger documents or when sending
MirroredConfigSetRequest-s that have multiple content streams. Here's the
breakdown:
{noformat}
actualSize=61 byteArraySize=64 waste=3 (4.7% of
transmitted bytes)
actualSize=153 byteArraySize=153 waste=0 (0.0% of
transmitted bytes)
actualSize=604 byteArraySize=604 waste=0 (0.0% of
transmitted bytes)
actualSize=4104 byteArraySize=4104 waste=0 (0.0% of
transmitted bytes)
actualSize=6104 byteArraySize=6104 waste=0 (0.0% of
transmitted bytes)
actualSize=8104 byteArraySize=8104 waste=0 (0.0% of
transmitted bytes)
actualSize=8304 byteArraySize=16464 waste=8160 (49.6% of
transmitted bytes)
actualSize=10104 byteArraySize=20064 waste=9960 (49.6% of
transmitted bytes)
actualSize=15104 byteArraySize=30064 waste=14960 (49.8% of
transmitted bytes)
actualSize=20105 byteArraySize=40066 waste=19961 (49.8% of
transmitted bytes)
actualSize=30105 byteArraySize=32768 waste=2663 (8.1% of
transmitted bytes)
actualSize=50105 byteArraySize=65536 waste=15431 (23.5% of
transmitted bytes)
actualSize=100105 byteArraySize=131072 waste=30967 (23.6% of
transmitted bytes)
{noformat}
The {{actualSize}} represents the size of serialized data, {{byteArraySize}} is
the size of the returned byte array that contains the serialized data. Up until
the boundary of 8192 bytes the optimization works very well, but above this
threshold it produces a lot of waste.
I propose to add a conditional in {{ExposedByteArrayOutputStream.byteArray()}}
to return the raw buffer only if it's at most 10% larger than the size of
serialized data. This will cover all small-doc cases and avoid an array copy
for some less egregious cases for larger docs.
Here's the breakdown with the proposed optimization:
{noformat}
actualSize=61 returnedSize=64 (allocatedBufferSize=64) waste=3 (4.7%
of transmitted bytes)
actualSize=153 returnedSize=153 (allocatedBufferSize=153) waste=0
(0.0% of transmitted bytes)
actualSize=462 returnedSize=462 (allocatedBufferSize=462) waste=0
(0.0% of transmitted bytes)
actualSize=1205 returnedSize=1205 (allocatedBufferSize=1205) waste=0
(0.0% of transmitted bytes)
actualSize=1544 returnedSize=1544 (allocatedBufferSize=1544) waste=0
(0.0% of transmitted bytes)
actualSize=1771 returnedSize=1771 (allocatedBufferSize=1771) waste=0
(0.0% of transmitted bytes)
actualSize=2007 returnedSize=2007 (allocatedBufferSize=2007) waste=0
(0.0% of transmitted bytes)
actualSize=2264 returnedSize=2264 (allocatedBufferSize=2264) waste=0
(0.0% of transmitted bytes)
actualSize=2855 returnedSize=2855 (allocatedBufferSize=2855) waste=0
(0.0% of transmitted bytes)
actualSize=3921 returnedSize=3921 (allocatedBufferSize=3921) waste=0
(0.0% of transmitted bytes)
actualSize=4865 returnedSize=4865 (allocatedBufferSize=4865) waste=0
(0.0% of transmitted bytes)
actualSize=5835 returnedSize=5835 (allocatedBufferSize=5835) waste=0
(0.0% of transmitted bytes)
actualSize=6192 returnedSize=6192 (allocatedBufferSize=6192) waste=0
(0.0% of transmitted bytes)
actualSize=6346 returnedSize=6346 (allocatedBufferSize=6346) waste=0
(0.0% of transmitted bytes)
actualSize=7330 returnedSize=7330 (allocatedBufferSize=7330) waste=0
(0.0% of transmitted bytes)
actualSize=7931 returnedSize=7931 (allocatedBufferSize=7931) waste=0
(0.0% of transmitted bytes)
actualSize=8296 returnedSize=8296 (allocatedBufferSize=16384) waste=0
(0.0% of transmitted bytes)
actualSize=13570 returnedSize=13570 (allocatedBufferSize=26996)
waste=0 (0.0% of transmitted bytes)
actualSize=17776 returnedSize=17776 (allocatedBufferSize=35408)
waste=0 (0.0% of transmitted bytes)
actualSize=26031 returnedSize=26031 (allocatedBufferSize=32768)
waste=0 (0.0% of transmitted bytes)
actualSize=29673 returnedSize=32768 (allocatedBufferSize=32768)
waste=3095 (9.4% of transmitted bytes)
actualSize=33733 returnedSize=33733 (allocatedBufferSize=65536)
waste=0 (0.0% of transmitted bytes)
actualSize=35836 returnedSize=35836 (allocatedBufferSize=65536)
waste=0 (0.0% of transmitted bytes)
actualSize=37350 returnedSize=37350 (allocatedBufferSize=65536)
waste=0 (0.0% of transmitted bytes)
actualSize=38449 returnedSize=38449 (allocatedBufferSize=65536)
waste=0 (0.0% of transmitted bytes)
actualSize=46211 returnedSize=46211 (allocatedBufferSize=65536)
waste=0 (0.0% of transmitted bytes)
actualSize=54614 returnedSize=54614 (allocatedBufferSize=65536)
waste=0 (0.0% of transmitted bytes)
actualSize=60514 returnedSize=65536 (allocatedBufferSize=65536)
waste=5022 (7.7% of transmitted bytes)
actualSize=69125 returnedSize=69125 (allocatedBufferSize=131072)
waste=0 (0.0% of transmitted bytes)
actualSize=74878 returnedSize=74878 (allocatedBufferSize=131072)
waste=0 (0.0% of transmitted bytes)
actualSize=76876 returnedSize=76876 (allocatedBufferSize=131072)
waste=0 (0.0% of transmitted bytes)
actualSize=79056 returnedSize=79056 (allocatedBufferSize=131072)
waste=0 (0.0% of transmitted bytes)
actualSize=80375 returnedSize=80375 (allocatedBufferSize=131072)
waste=0 (0.0% of transmitted bytes)
actualSize=83160 returnedSize=83160 (allocatedBufferSize=131072)
waste=0 (0.0% of transmitted bytes)
actualSize=89323 returnedSize=89323 (allocatedBufferSize=131072)
waste=0 (0.0% of transmitted bytes)
actualSize=92109 returnedSize=92109 (allocatedBufferSize=131072)
waste=0 (0.0% of transmitted bytes)
actualSize=99649 returnedSize=99649 (allocatedBufferSize=131072)
waste=0 (0.0% of transmitted bytes)
{noformat}
was (Author: ab):
After further analysis this looks somewhat better than I thought, especially
for updates that are less than 8192 byte in size (serialized), because for
these the array is resized to the actual size of the payload (as accumulated in
the {{FastOutputStream}}'s buffer that the {{JavaBinCodec}} uses). In this case
no extra padding bytes are wasted.
Since the producer side ({{MirroringUpdateProcessor / KafkaMirroringSink}})
doesn't accumulate individual requests in batches and in a typical scenario
most documents are smaller than 8kB serialized, this optimization will mostly
hold.
This is not the case when sending larger documents or when sending
MirroredConfigSetRequest-s that have multiple content streams. Here's the
breakdown:
{noformat}
actualSize=61 byteArraySize=64 waste=3 (4.7% of
transmitted bytes)
actualSize=153 byteArraySize=153 waste=0 (0.0% of
transmitted bytes)
actualSize=604 byteArraySize=604 waste=0 (0.0% of
transmitted bytes)
actualSize=4104 byteArraySize=4104 waste=0 (0.0% of
transmitted bytes)
actualSize=6104 byteArraySize=6104 waste=0 (0.0% of
transmitted bytes)
actualSize=8104 byteArraySize=8104 waste=0 (0.0% of
transmitted bytes)
actualSize=8304 byteArraySize=16464 waste=8160 (49.6% of
transmitted bytes)
actualSize=10104 byteArraySize=20064 waste=9960 (49.6% of
transmitted bytes)
actualSize=15104 byteArraySize=30064 waste=14960 (49.8% of
transmitted bytes)
actualSize=20105 byteArraySize=40066 waste=19961 (49.8% of
transmitted bytes)
actualSize=30105 byteArraySize=32768 waste=2663 (8.1% of
transmitted bytes)
actualSize=50105 byteArraySize=65536 waste=15431 (23.5% of
transmitted bytes)
actualSize=100105 byteArraySize=131072 waste=30967 (23.6% of
transmitted bytes)
{noformat}
The {{actualSize}} represents the size of serialized data, {{byteArraySize}} is
the size of the returned byte array that contains the serialized data. Up until
the boundary of 8192 bytes the optimization works very well, but above this
threshold it produces a lot of waste.
I propose to add a conditional in {{ExposedByteArrayOutputStream.byteArray()}}
to return the raw buffer only if it's at most 10% larger than the size of
serialized data. This will cover all small-doc cases and avoid an array copy
for some less egregious cases for larger docs.
Here's the breakdown with the proposed optimization:
{noformat}
actualSize=61 byteArraySize=64 waste=3 (4.7% of
transmitted bytes)
actualSize=153 byteArraySize=153 waste=0 (0.0% of
transmitted bytes)
actualSize=604 byteArraySize=604 waste=0 (0.0% of
transmitted bytes)
actualSize=4104 byteArraySize=4104 waste=0 (0.0% of
transmitted bytes)
actualSize=6104 byteArraySize=6104 waste=0 (0.0% of
transmitted bytes)
actualSize=8104 byteArraySize=8104 waste=0 (0.0% of
transmitted bytes)
actualSize=8304 byteArraySize=8304 waste=0 (0.0% of
transmitted bytes)
actualSize=17386 byteArraySize=17386 waste=0 (0.0% of
transmitted bytes)
actualSize=24700 byteArraySize=24700 waste=0 (0.0% of
transmitted bytes)
actualSize=28550 byteArraySize=28550 waste=0 (0.0% of
transmitted bytes)
actualSize=31026 byteArraySize=32768 waste=1742 (5.3% of
transmitted bytes)
actualSize=36877 byteArraySize=36877 waste=0 (0.0% of
transmitted bytes)
actualSize=43667 byteArraySize=43667 waste=0 (0.0% of
transmitted bytes)
actualSize=47067 byteArraySize=47067 waste=0 (0.0% of
transmitted bytes)
actualSize=50696 byteArraySize=50696 waste=0 (0.0% of
transmitted bytes)
actualSize=57516 byteArraySize=57516 waste=0 (0.0% of
transmitted bytes)
actualSize=61110 byteArraySize=65536 waste=4426 (6.8% of
transmitted bytes)
actualSize=69145 byteArraySize=69145 waste=0 (0.0% of
transmitted bytes)
actualSize=79692 byteArraySize=79692 waste=0 (0.0% of
transmitted bytes)
actualSize=89342 byteArraySize=89342 waste=0 (0.0% of
transmitted bytes)
actualSize=91734 byteArraySize=91734 waste=0 (0.0% of
transmitted bytes)
actualSize=93769 byteArraySize=93769 waste=0 (0.0% of
transmitted bytes)
{noformat}
> CrossDC Common: ExposedByteArrayOutputStream optimizes wrong aspect
> -------------------------------------------------------------------
>
> Key: SOLR-18415
> URL: https://issues.apache.org/jira/browse/SOLR-18415
> Project: Solr
> Issue Type: Bug
> Components: module - crossDC
> Affects Versions: 10.0, 9.10.1
> Reporter: Andrzej Bialecki
> Assignee: Andrzej Bialecki
> Priority: Major
> Fix For: main(11.0), 10.1
>
>
> {{MirroredSolrRequestSerializer}} uses a subclass of
> {{ByteArrayOutputStream}} in an attempt to optimize data transfer by avoiding
> array copy - instead it passes the whole buffer directly as the output bytes
> from {{{}serialize(){}}}.
> This sounds like a plausible optimization. However, the initial size of the
> buffer (32 bytes) is doubled on each overflow as the codec writes to the
> stream, which means that on average the actual serialized data occupies
> around half of the buffer, the rest are zeroes. Deserialization will handle
> these trailing zeros just fine, it's the cost of transmitting roughly twice
> as much data over the wire that is questionable - this buffer is then passed
> as-is to the output stream and eventually ends up as network traffic.
> This "optimization" avoids one array copy cost (which reduces GC) but
> generates multiple costs on the way - not only the empty data has to be sent
> over the network but also it needs to be stored in source Kafka, in
> MirrorMaker and in target Kafka, to be finally retrieved by the Consumer
> (which has to allocate a buffer sufficiently large to fit also the zeroes),
> only to discard the zeroes during deserialization.
> I propose to get rid of this optimization and use the plain
> {{ByteArrayOutputStream}} instead.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]