On 9/22/2026 8:12 PM, Michael Kelley wrote:
From: Naman Jain <[email protected]> Sent: Tuesday, September 22,
2026 2:09 AM
On 9/22/2026 2:39 AM, Michael Kelley wrote:
From: Naman Jain <[email protected]> Sent: Sunday, September 6, 2026
10:48 PM
On Hyper-V guests each virtual PCI bus is enumerated by its own
hv_pci_probe() call. The probe performs several synchronous host
request/response exchanges while negotiating the protocol, querying bus
relations, entering D0, and reporting allocated resources. These waits
are latency-bound rather than CPU-bound.
hv_pci registers as an ordinary VMBus driver, so driver_register() walks
matching vPCI buses and probes them sequentially while the driver's
initcall runs. On guests that expose several devices, each through its
own vPCI bus, this serialization adds the host round-trip latencies to
device initialization.
Each bus is described by its own struct hv_pcibus_device, so
independent buses can be probed concurrently. Request asynchronous
probing via PROBE_PREFER_ASYNCHRONOUS, causing the driver core to
schedule matching buses for asynchronous probe work.
On an Azure Standard_L32s_v3 guest with five vPCI targets (four NVMe
controllers and one Mellanox VF), Linux 7.2.3 was tested with one warm-up
and three measured boots per variant. The median interval from the first
hv_pci_probe() entry to the last return decreased from 2847.968 ms to
2786.709 ms, a 61.259 ms (2.15%) improvement.
The elapsed time improvement is rather disappointing given the
complexity of the probing sequence and the number of interactions
with the Hyper-V host. Do you have any insight into why there isn't a
larger reduction? Is something mostly serializing the work even though
PROBE_PREFER_ASYNCHRONOUS is specified?
Michael
I can see these reasons for not seeing great improvements:
1. Timing of device offers from the host is beyond the control of guest
and the Hyper-V host may also be serializing the requests from the host.
Ah, right. This is probably the key factor.
2. Shared locks that needs to be handled separately:
* hyperv_mmio_lock during VMBus MMIO allocation.
This probably has minimal impact. While there's a decent amount
of code protected by the lock, I don't think there's any interaction
with the host (even via traps), so it should run quickly.
* pci_rescan_remove_lock during PCI resource assignment and device
addition.
OK. I don’t know about this one.
I digged more into it, and it is indeed because of late offers from
Hyper-V. I was considering the start of first probe to the last return,
for time calculations.
For the 4 PCI devices on my setup whose offers were delivered together,
the performance improvement was about 24%. However with the last offer
coming late for MLX PCI device, overall improvement in time was lesser
in terms of percentage.
For traditional configs where the VF NIC is paired with a netvsc
instance, the host doesn't offer the VF NIC to the guest until the
corresponding netvsc instance has been probed and the netvsc
driver has told the host it will accept a VF NIC. In these cases, the
Mellanox or MANA VF NIC device is always "late" and probably
shouldn't be included when determining the speed-up of doing
hv_pci probing asynchronously.
Dexuan had removed pci_rescan_remove_lock in his previous upstream
attempt, but I ommitted it intentionally this time because from AI
review, I saw a potential race condition that we would introduce if we
remove it. Secondly, I did not observe any benefits of removing this
lock. But I am going to revisit it again.
Thanks for the discussion. Using PROBE_PREFER_ASYNCHRONOUS
is still a good thing to use, but the benefit will accrue the most when
there are a significant number of NVMe devices and when Hyper-V
is quick about offering them. As I described above, it will be harder
to get parallelism with the VF NIC.
Michael
Thanks Michael, this makes sense now. I'm glad that you asked.
Regards,
Naman