From: "Huang, Honglei" <honghuan@amd.com>
To: Akihiko Odaki <odaki@rsg.ci.i.u-tokyo.ac.jp>
Cc: qemu-devel@nongnu.org, virtio-comment@lists.oasis-open.org,
dri-devel@lists.freedesktop.org, virtualization@lists.linux.dev,
"Honglei Huang" <honglei1.huang@amd.com>,
"Huang Rui" <Ray.Huang@amd.com>,
"Michael S. Tsirkin" <mst@redhat.com>,
"Alex Bennée" <alex.bennee@linaro.org>,
"Dmitry Osipenko" <dmitry.osipenko@collabora.com>,
"Marc-André Lureau" <marcandre.lureau@redhat.com>,
"Stefano Garzarella" <sgarzare@redhat.com>,
"Gerd Hoffmann" <kraxel@redhat.com>,
"David Airlie" <airlied@redhat.com>,
"Peter Maydell" <peter.maydell@linaro.org>
Subject: Re: About new backend for GPU compute ROCm in qemu
Date: Tue, 18 Aug 2026 12:26:10 +0800 [thread overview]
Message-ID: <63d485c3-a418-493d-aad0-26b3c115f5b8@amd.com> (raw)
In-Reply-To: <33bf6649-c1b9-43f6-94dd-09167d8e5db6@rsg.ci.i.u-tokyo.ac.jp>
On 8/18/2026 12:05 PM, Akihiko Odaki wrote:
> On 2026/08/18 11:50, Huang, Honglei wrote:
>>
>>
>> On 8/18/2026 12:29 AM, Akihiko Odaki wrote:
>>> On 2026/08/17 22:44, Huang, Honglei wrote:
>>>>
>>>>
>>>> On 8/17/2026 7:44 PM, Akihiko Odaki wrote:
>>>>> On 2026/08/17 12:19, Huang, Honglei wrote:
>>>>>>
>>>>>> Hi Michael, Alex, Dmitry, Akihiko,
>>>>>
>>>>> Hi Honglei,
>>>>>
>>>>>>
>>>>>> I'm bringing AMD GPU compute ROCm based on virtio. I posted a ROCm
>>>>>> over virtio
>>>>>> implementation to virglrenderer nine months ago (MR !1568 [1]).
>>>>>> The ROCm side has
>>>>>> been supportted by ROCm offical.
>>>>>>
>>>>>> Current implementation is a virtio gpu context type capset handled
>>>>>> inside
>>>>>> virglrenderer, sharing the display path. That's an awkward fit, many
>>>>>> compute GPUs have no display engine at all.
>>>>>
>>>>> I think "sharing the display path" conflates several layers and
>>>>> makes the problem difficult to assess. It would help to identify
>>>>> the concrete constraint behind "awkward fit."
>>>>>
>>>>> End-to-end, there are four relevant layers:
>>>>>
>>>>> 1. Host GPU stack: hardware, host kernel, and host userspace
>>>>> 2. Paravirtualization stack: virglrenderer and QEMU
>>>>
>>>> Yes we are asking can we add a new file like virtio-gpu specific for
>>>> compute, but maybe we can only add a new backend like virglrenderer
>>>> specific for compute.
>>>>
>>>>> 3. Host/guest interface: virtio and the capset-specific command stream
>>>>
>>>> In this plan we may need just add a capset id.
>>>>
>>>>> 4. Guest GPU stack: guest kernel and guest userspace
>>>>
>>>> Won't modify the guest kernel in this plan, this email list.
>>>>
>>>>>
>>>>> Orthogonally, acceleration is separate from display and scanout. A
>>>>> physical device may provide both, but acceleration does not require a
>>>>> display engine. Linux likewise exposes render and compute interfaces
>>>>> separately from modesetting. The userspace interface virglrenderer
>>>>> uses is messy; there is Vulkan, EGL, OpenGL, and now you are adding
>>>>> ROCm. But there is one thing I must note is that acceleration and
>>>>> display is decoupled, and acceleration does not require display.
>>>>
>>>> Yes totally agreed.
>>>>
>>>>>
>>>>> At the protocol layer, context command buffers are carried by
>>>>> VIRTIO_GPU_CMD_SUBMIT_3D. Scanout uses separate core virtio-gpu
>>>>> commands, and VIRTIO_GPU_CMD_GET_DISPLAY_INFO may report no enabled
>>>>> displays. At the implementation layer, QEMU handles scanout
>>>>> presentation. virgl_cmd_set_scanout() obtains resource information
>>>>> through virgl_renderer_resource_get_info() or
>>>>> virgl_renderer_resource_get_info_ext(). That does not make scanout
>>>>> a virglrenderer-owned display path.
>>>>
>>>> Yes, agreed.
>>>>
>>>>>
>>>>> Therefore, if "sharing the display path" means sharing the same
>>>>> device, control queue, and QEMU execution context, that identifies
>>>>> a possible source of contention. If it means that capsets or
>>>>> virglrenderer are inherently tied to display, I do not think that
>>>>> is accurate. Vulkan compute is already used through Venus with
>>>>> libkrun [2], and VCL proposes OpenCL support through virglrenderer
>>>>> [3].
>>>>
>>>> Yes,but the vulkan is for GFX originally, and for some formal AI
>>>> frame work like pytorch, it's support is limited, and it performance
>>>> is lower than ROCm, and vulkan also lacks many AI infrastructure,
>>>> like composable kernel.
>>>> And for virCL, actually it is came from same project with ROCm
>>>> native context, but the original author didn't continue to support
>>>> it, they handed it over to someone else to take over. And in the
>>>> first version of
>>>> virCL, it didn't pass the test of actual projects.
>>>>
>>>> And it seems like virCL didn't upstream into virglrenderer also,
>>>> correct me if I am wrong.
>>>
>>> I cited Venus and VCL only as examples showing that virtio-gpu and
>>> virglrenderer are not intrinsically tied to display. I did not suggest
>>> either as a substitute for ROCm.
>>>
>>>>
>>>>>> Beyond that, sharing the display path is increasingly painful:
>>>>>>
>>>>>> - Compute hammers the queues more than graphics, so sharing
>>>>>> virtio gpu's single control queue with display/virgl causes
>>>>>> contention
>>>>>> and display stutter.
>>>>>
>>>>> All non-cursor commands do share one control queue, but a fence
>>>>> avoids serialization.
>>>>
>>>> yes, agreed.
>>>>
>>>>>
>>>>> There may still be implementation-level contention, and it is not
>>>>> necessarily specific to compute. A sufficiently busy graphics workload
>>>>> could expose the same bottlenecks. Possible contributors in current
>>>>> QEMU
>>>>> include:
>>>>>
>>>>> a) qemu_console_hw_gl_block() blocks the entire queue when QEMU only
>>>>> needs to fence scanout commands.
>>>>
>>>> Yes, agreed.
>>>>
>>>>>
>>>>> b) virtio_gpu_virgl_unmap_resource_blob() may also block the entire
>>>>> queue just to delay one command.
>>>>
>>>> Yes, but it is seems like it is must, someone else in AMD tried to
>>>> use async method to relase blob, but it failed to consistency issue,
>>>> then
>>>> reverted to sync version.
>>>
>>> Queue-wide suspension is not inherently required. Commit 4eb0aace85f5
>>> ("virtio-gpu: Support mapping hostmem blobs with map_fixed") added a
>>> path that avoids per-blob MemoryRegion teardown when
>>> virgl_renderer_resource_map_fixed() succeeds. The remaining path is
>>> also being improved with:
>>>
>>> https://lore.kernel.org/qemu-devel/20260424-force_rcu-v4-0-
>>> feccfaca0568@rsg.ci.i.u-tokyo.ac.jp/
>>> ("[PATCH v4 0/6] virtio-gpu: Force RCU when unmapping blob")
>>
>> Thanks. force_rcu is a clean fix for the RCU-reclamation part, but it
>> still keeps the unmap synchronous and serial.
>>
>>>
>>>>
>>>>>
>>>>> c) QEMU dispatches the control queue and calls into virglrenderer from
>>>>> its main-loop thread along with display work and many other
>>>>> things.
>>>>> Venus's render server can offload renderer work, but control-queue
>>>>> dispatch remains in QEMU's main loop.
>>>>
>>>> Yes, we did some async optimization in ROCm context, but its
>>>> effectiveness is limited, see bellow.
>>>>
>>>>>
>>>>> In any case, I think you need to do some experiments to track down
>>>>> the real cause. a) is easy to check: just comment out all
>>>>> qemu_console_hw_gl_block() calls; it may corrupt display but
>>>>> removes the blocking. b) can also be tested by leaking the mappings
>>>>> instead of blocking the whole queue. Using a different display
>>>>> device like qxl tells whether c) is causing contention.
>>>>
>>>> Yes, totally agreed. following is my findings. In short words:
>>>>
>>>> Optimization can reduce queue pressure, but it can't withstand
>>>> absolute overload because each command has some overhead. Making all
>>>> commands asynchronous would lead to a debugging hell about
>>>> asynchronous issues.
>>>> And we have high load applications rocmprofiler that continuously
>>>> catch information need virtio queue to handle. But create a new
>>>> backend can not solve it simply, we are trying to find a way. like
>>>> shmem between guest and host, then use cpu polling, bypass the
>>>> virtqueue.
>>>
>>> Most commands are fast on the CPU side, while heavy processing
>>> happens asynchronously on the GPU. Cases (a) and (b) are exceptions.
>>>
>>>>
>>>> The load is mostly memory management. Running an AI model allocates
>>>> and frees a large number of blobs. We already did some optimization
>>>> release them asynchronously, but the host processing is a single
>>>> queue one
>>>> process_cmdq, this is where the main bottleneck in my debugging
>>>> work / my understanding so far. I'm not certain it's the whole
>>>> picture, so please correct if I am wrong.
>>>>
>>>> A model load or unload frees a large batch of BOs and allocates
>>>> another. Some of those commands are async in the virtio-gpu guest
>>>> driver, but QEMU still has to work through them on the one queue,
>>>> which takes time; so even though any single command is quick, there
>>>> are simply too many of them, the single queue backs up, and
>>>> everything behind it, gets delayed.
>>>>
>>>> real work load (a few downstream customisations): loading one 16 GB
>>>> model (gemm4 e4b), drives ~1200 blob creates, a burst of
>>>> ~1400 resource frees at teardown, ~3700 submits and ~6000 virtqueue
>>>> notifies, caused a 22 s guest soft lockup. And the behavior of
>>>> memory operations are controlled by upper layer like pytorch / HIP /
>>>> runtime,
>>>> we can not control it.
>>>>
>>>> To be honest, a separate backend won't fix this. But the real
>>>> solution maybe is compute specific. That logic is only useful to the
>>>> compute path, and folding it into the shared display device /
>>>> renderer would mean churning code that is mature and stable for
>>>> graphics, with regression risk. Keeping compute on its own instance
>>>> and backend lets us iterate on these compute only optimisations.
>>>
>>> A 22-second lockup is too long for those command counts.
>>>
>>> The most probable explanation I have is that the ROCm integration
>>> blocks QEMU's main loop thread while synchronously waiting for GPU
>>> execution. Creating separate devices won't resolve this because the
>>> main loop thread is shared, and synchronously waiting on the GPU
>>> should be avoided in the first place.
>>
>> No synchronously waiting in ROCm backend, we are using user queue, and
>> event waiting, no sync operation in CMD wait. all the resource release
>> in ROCm are all async now.
>> Only the sync thing is memory thing mapping/unmapping in qemu, as long
>> as it remains synchronous, it will be overwhelmed by the massive
>> number of requests.
>
> Mapping and unmapping should not block QEMU's main-loop thread for that
> long. The command counts you reported are relatively small. That is why
> I suspect something else went wrong, such as the main-loop thread being
> inadvertently blocked while waiting for the GPU.
Will investigate it.
>
>>
>>>
>>> In any case, profiling is necessary before touching the implementation.
>>>
>>>>
>>>>>
>>>>>> - Compute contexts need far more blob / shared memory than a
>>>>>> display one.
>>>>>
>>>>> It is not a problem by itself. Frequent mapping and unmapping might
>>>>> amplify the second issue above, but that needs to be measured.
>>>>
>>>> Yes, agreed. I can give more detailed information.
>>>>
>>>>>
>>>>>> - Maybe needs a wider ROCm / compute stack, cause the render
>>>>>> model fits poorly:
>>>>>> rocprofiler (PC sampling, SQTT/SPM, counters, high bandwidth
>>>>>> streams)
>>>>>> and ROCgdb (wave control, address watch, async exceptions an
>>>>>> out of band channel that must not block display).
>>>>>
>>>>> virglrenderer does not impose a particular render model. That's why
>>>>> Vulkan Compute just works with Venus.
>>>>
>>>> Yes but vulkan is used for GFX initally. And can not support many AI
>>>> application.>
>>>>>> - Events, faults and GPU reset/SMI are async and don't map onto
>>>>>> fences.> - All of this is hard to extend cleanly inside a
>>>>>> display capset.
>>>>> Capset is not about display but determines the protocol of the
>>>>> VIRTIO_GPU_CMD_SUBMIT_3D command stream. You have described events,
>>>>> faults and GPU reset/SMI are async don't map onto fences that may
>>>>> be associated with VIRTIO_GPU_CMD_SUBMIT_3D which is dictated by
>>>>> capset. An additional feature may be necessary, and it may or may
>>>>> not be dictated by capset. The other things are irrelevant with the
>>>>> protocol capset represents; they are either behavioral or about
>>>>> different commands.
>>>>
>>>> A fence is the one shot, but event is stateful and repeatable.
>>>> That may or may not be tied to capset. Agreed.
>>>>
>>>>>
>>>>>>
>>>>>> On the QEMU/host side, would something like this be OK? One step,
>>>>>> two parts:
>>>>>>
>>>>>> - a dedicated headless virtio gpu instance for compute.
>>>>>
>>>>> A second device would isolate its virtqueues and device-wide
>>>>> renderer_blocked state. That may be useful if measurements show that
>>>>> these are the bottlenecks, but it is not yet clear that they are or
>>>>> that
>>>>> a second device is the appropriate solution.
>>>>>
>>>>>> - that instance served by a separate ROCm backend library loaded
>>>>>> in-process by QEMU.
>>>>>
>>>>> First, I think we need to establish why ROCm cannot or should not
>>>>> remain
>>>>> in virglrenderer. The virglrenderer, Venus, and VCL maintainers are
>>>>> likely better placed to advise on that boundary. Once the protocol
>>>>> requirements and performance measurements are clear, we can assess the
>>>>> appropriate QEMU integration.
>>>>
>>>> venus is borned for GFX.
>>>> virCL not merged.
>>>>
>>>> To be clear, I'm not saying virglrenderer can't host a ROCm native
>>>> context it clearly can. My hesitation is more about fit and
>>>> direction: virglrenderer has grown up around GL/graphics, and I
>>>> haven't yet found compute oriented plumbing there to build on, while
>>>> ROCm moves very fast and I need something I can keep current with
>>>> low friction.
>>>
>>> Whether keeping ROCm in virglrenderer would create extra friction is
>>> primarily a question for the virglrenderer maintainers. Its graphics
>>> origins do not by themselves motivate adding a separate backend
>>> interface to QEMU.
>>
>> Fair. The first draft version in virglrenderer was in May 2024, and
>> ROCm has gone 5.7 → 7.14 in that window.
>
> One point to note is that virtio-gpu development in QEMU is somewhat
> less active. crosvm is the most active user of virglrenderer, and QEMU
> sometimes lags behind it. If you are considering moving the ROCm
> integration from virglrenderer to QEMU solely because ROCm evolves
> rapidly, I do not think that would be a good idea. A rapidly evolving
> component is better kept in virglrenderer unless there is another reason
> to place it in QEMU.
Actually didn't see something new about compute merged in to
virglrenderer this recently 2 years.
Regards,
Honglei
>
> Regards,
> Akihiko Odaki
prev parent reply other threads:[~2026-08-18 4:26 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-17 3:19 About new backend for GPU compute ROCm in qemu Huang, Honglei
2026-08-17 9:06 ` Alex Bennée
2026-08-17 12:46 ` Huang, Honglei
2026-08-17 11:44 ` Akihiko Odaki
2026-08-17 13:44 ` Huang, Honglei
2026-08-17 14:24 ` Alex Bennée
2026-08-18 2:27 ` Huang, Honglei
2026-08-17 16:29 ` Akihiko Odaki
2026-08-18 2:50 ` Huang, Honglei
2026-08-18 4:05 ` Akihiko Odaki
2026-08-18 4:26 ` Huang, Honglei [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=63d485c3-a418-493d-aad0-26b3c115f5b8@amd.com \
--to=honghuan@amd.com \
--cc=Ray.Huang@amd.com \
--cc=airlied@redhat.com \
--cc=alex.bennee@linaro.org \
--cc=dmitry.osipenko@collabora.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=honglei1.huang@amd.com \
--cc=kraxel@redhat.com \
--cc=marcandre.lureau@redhat.com \
--cc=mst@redhat.com \
--cc=odaki@rsg.ci.i.u-tokyo.ac.jp \
--cc=peter.maydell@linaro.org \
--cc=qemu-devel@nongnu.org \
--cc=sgarzare@redhat.com \
--cc=virtio-comment@lists.oasis-open.org \
--cc=virtualization@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox