On Thu, Aug 27, 2026 at 7:59 AM Jia Jia <[email protected]> wrote:
>
> vhost_transport_do_send_pkt() calls vhost_add_used() for every Guest RX
> buffer even though it delays the Guest signal until the worker finishes.
> Each call publishes one used entry and updates the used index separately.
>
> Collect the completed buffer heads in the arrays already allocated for the
> virtqueue and publish them with vhost_add_used_n(). Bound the batch by the
> ring size, array capacity, and worker packet budget. Flush before
> re-enabling notifications or leaving the worker.
>
> Each used entry describes one completed RX buffer and keeps its actual used
> length, so set nheads to 1 for every entry. This patch does not change
> negotiated features or compress multiple buffers into one used entry.
>
> This patch is limited to the current skb-based vhost-vsock RX path.
>
> Performance:
>
> Tested with a QEMU/KVM guest on a host with 4 online CPUs, using 2 vCPUs
> pinned to host CPUs 2 and 3, QEMU 10.2.1, q35, 1536 MiB, and Linux
> 7.2.0-rc3-next-20260713-next-debug-kasan. The vhost-vsock source is based
> on linux-next master at 49362394dad7df66c274c867a271394c10ca2bb8.
>
> Current vhost-vsock does not implement VIRTIO_F_IN_ORDER or
> VIRTIO_F_RING_PACKED, so both configurations used packed=off and
> in_order=off:
>
> baseline: RX batching=off
> vhost-vsock RX batching: RX batching=on
>
> The test used vsock_perf. The Guest receiver was started with:
>
> vsock_perf --port PORT --buf-size 64M --vsk-size 64M --rcvlowat 1
>
> The Host sender was started with:
>
> vsock_perf --sender 3 --port PORT --bytes BYTES \
> --buf-size SEND_BUF --vsk-size 64M
>
> Each workload transferred BYTES=1 GiB. The SEND_BUF values were 256 B
> (SEND_BUF=256), 512 B (SEND_BUF=512), 4 KiB (SEND_BUF=4K), and 64 KiB
> (SEND_BUF=64K).
> Each state used a fresh Guest. Each workload uses 20 paired runs, with 10
> runs in each order. The reported values are
> Guest RX throughput in Gbits/s. The baseline and batching columns are the
> geometric means over the 20 runs; change is batching / baseline - 1,
> computed from the unrounded values:
>
> workload baseline RX batching RX change faster
> 256 B 0.0795724 0.0831509 +4.497% 20/20
> 512 B 0.1194885 0.1210297 +1.290% 14/20
> 4 KiB 0.7208273 0.7242053 +0.469% 11/20
> 64 KiB 2.1712797 2.1951941 +1.101% 13/20
>
> For reference, the table below gives the 95% normal-approximation intervals
> obtained from the 20 paired log(batching / baseline) values:
>
> workload paired 95% interval
> 256 B +3.985% to +5.011%
> 512 B +0.206% to +2.385%
> 4 KiB -1.442% to +2.416%
> 64 KiB -1.474% to +3.745%
>
> All transfers passed byte-count checks, and no kernel errors were observed
> in the logs. The 256-byte workload improved in every pair. The 512 B
> workload was faster in 14 of 20 pairs, with a small gain. The 4 KiB and
> 64 KiB workloads showed no material throughput change; the difference
> between their results may be due to scheduling and execution variation.
>
> Link: https://lore.kernel.org/r/[email protected]
> Link: https://lore.kernel.org/r/[email protected]
> Signed-off-by: Jia Jia <[email protected]>
> ---
> drivers/vhost/vsock.c | 41 +++++++++++++++++++++++++++++++++++++++--
> 1 file changed, 39 insertions(+), 2 deletions(-)
>
> diff --git a/drivers/vhost/vsock.c b/drivers/vhost/vsock.c
> index 9aaab6bb8..d8050cca2 100644
> --- a/drivers/vhost/vsock.c
> +++ b/drivers/vhost/vsock.c
> @@ -103,12 +103,37 @@ static bool vhost_transport_has_remote_cid(struct
> vsock_sock *vsk, u32 cid)
> return found;
> }
>
> +static bool vhost_vsock_flush_used(struct vhost_virtqueue *vq,
> + unsigned int *used_count)
> +{
> + if (!*used_count)
> + return false;
> +
> + vhost_add_used_n(vq, vq->heads, vq->nheads, *used_count);
> + *used_count = 0;
> + return true;
> +}
> +
I think that the used_count should be updated by the caller based on
the boolean returned by this function.
> +static void vhost_vsock_add_used(struct vhost_virtqueue *vq,
> + unsigned int *used_count,
> + unsigned int head, unsigned int len)
> +{
> + struct vring_used_elem *used = &vq->heads[*used_count];
> +
> + used->id = cpu_to_vhost32(vq, head);
> + used->len = cpu_to_vhost32(vq, len);
> + vq->nheads[*used_count] = 1;
> + (*used_count)++;
> +}
Same here, why pass it as an argument? The caller knows way better.
Other than that, the patch looks good to me,
Acked-by: Eugenio Pérez <[email protected]>
Thanks!
> +
> static void
> vhost_transport_do_send_pkt(struct vhost_vsock *vsock,
> struct vhost_virtqueue *vq)
> {
> struct vhost_virtqueue *tx_vq = &vsock->vqs[VSOCK_VQ_TX];
> int pkts = 0, total_len = 0;
> + unsigned int used_count = 0;
> + unsigned int used_limit;
> bool added = false;
> bool restart_tx = false;
>
> @@ -120,6 +145,12 @@ vhost_transport_do_send_pkt(struct vhost_vsock *vsock,
> if (!vq_meta_prefetch(vq))
> goto out;
>
> + used_limit = min_t(unsigned int, vq->num,
> + min_t(unsigned int, vq->dev->iov_limit,
> + vq->dev->weight));
> + if (unlikely(!used_limit))
> + goto out;
> +
> /* Avoid further vmexits, we're already processing the virtqueue */
> vhost_disable_notify(&vsock->dev, vq);
>
> @@ -134,9 +165,13 @@ vhost_transport_do_send_pkt(struct vhost_vsock *vsock,
> u32 offset;
> int head;
>
> + if (used_count == used_limit)
> + added |= vhost_vsock_flush_used(vq, &used_count);
> +
> skb = virtio_vsock_skb_dequeue(&vsock->send_pkt_queue);
>
> if (!skb) {
> + added |= vhost_vsock_flush_used(vq, &used_count);
> vhost_enable_notify(&vsock->dev, vq);
> break;
> }
> @@ -153,6 +188,7 @@ vhost_transport_do_send_pkt(struct vhost_vsock *vsock,
> /* We cannot finish yet if more buffers snuck in while
> * re-enabling notify.
> */
> + added |= vhost_vsock_flush_used(vq, &used_count);
> if (unlikely(vhost_enable_notify(&vsock->dev, vq))) {
> vhost_disable_notify(&vsock->dev, vq);
> continue;
> @@ -230,8 +266,8 @@ vhost_transport_do_send_pkt(struct vhost_vsock *vsock,
> */
> virtio_transport_deliver_tap_pkt(skb);
>
> - vhost_add_used(vq, head, sizeof(*hdr) + payload_len);
> - added = true;
> + vhost_vsock_add_used(vq, &used_count, head,
> + sizeof(*hdr) + payload_len);
>
> VIRTIO_VSOCK_SKB_CB(skb)->offset += payload_len;
> total_len += payload_len;
> @@ -264,6 +300,7 @@ vhost_transport_do_send_pkt(struct vhost_vsock *vsock,
> virtio_transport_consume_skb_sent(skb, true);
> }
> } while(likely(!vhost_exceeds_weight(vq, ++pkts, total_len)));
> + added |= vhost_vsock_flush_used(vq, &used_count);
> if (added)
> vhost_signal(&vsock->dev, vq);
>
> --
> 2.53.0
>