Linux virtualization list
 help / color / mirror / Atom feed
From: "Huang, Honglei" <honghuan@amd.com>
To: Akihiko Odaki <odaki@rsg.ci.i.u-tokyo.ac.jp>
Cc: qemu-devel@nongnu.org, virtio-comment@lists.oasis-open.org,
	dri-devel@lists.freedesktop.org, virtualization@lists.linux.dev,
	"Honglei Huang" <honglei1.huang@amd.com>,
	"Huang Rui" <Ray.Huang@amd.com>,
	"Michael S. Tsirkin" <mst@redhat.com>,
	"Alex Bennée" <alex.bennee@linaro.org>,
	"Dmitry Osipenko" <dmitry.osipenko@collabora.com>,
	"Marc-André Lureau" <marcandre.lureau@redhat.com>,
	"Stefano Garzarella" <sgarzare@redhat.com>,
	"Gerd Hoffmann" <kraxel@redhat.com>,
	"David Airlie" <airlied@redhat.com>,
	"Peter Maydell" <peter.maydell@linaro.org>
Subject: Re: About new backend for GPU compute ROCm in qemu
Date: Mon, 17 Aug 2026 21:44:21 +0800	[thread overview]
Message-ID: <78f0583f-93c0-4374-ba37-fd36f6388f0e@amd.com> (raw)
In-Reply-To: <ea899a6b-0a59-4a9c-9e81-3ed70711da59@rsg.ci.i.u-tokyo.ac.jp>



On 8/17/2026 7:44 PM, Akihiko Odaki wrote:
> On 2026/08/17 12:19, Huang, Honglei wrote:
>>
>> Hi Michael, Alex, Dmitry, Akihiko,
> 
> Hi Honglei,
> 
>>
>> I'm bringing AMD GPU compute ROCm based on virtio. I posted a ROCm 
>> over virtio
>> implementation to virglrenderer nine months ago (MR !1568 [1]). The 
>> ROCm side has
>> been supportted by ROCm offical.
>>
>> Current implementation is a virtio gpu context type capset handled inside
>> virglrenderer, sharing the display path. That's an awkward fit, many
>> compute GPUs have no display engine at all.
> 
> I think "sharing the display path" conflates several layers and makes 
> the problem difficult to assess. It would help to identify the concrete 
> constraint behind "awkward fit."
> 
> End-to-end, there are four relevant layers:
> 
> 1. Host GPU stack: hardware, host kernel, and host userspace
> 2. Paravirtualization stack: virglrenderer and QEMU

Yes we are asking can we add a new file like virtio-gpu specific for 
compute, but maybe we can only add a new backend like virglrenderer 
specific for compute.

> 3. Host/guest interface: virtio and the capset-specific command stream

In this plan we may need just add a capset id.

> 4. Guest GPU stack: guest kernel and guest userspace

Won't modify the guest kernel in this plan, this email list.

> 
> Orthogonally, acceleration is separate from display and scanout. A
> physical device may provide both, but acceleration does not require a
> display engine. Linux likewise exposes render and compute interfaces
> separately from modesetting. The userspace interface virglrenderer uses 
> is messy; there is Vulkan, EGL, OpenGL, and now you are adding ROCm. But 
> there is one thing I must note is that acceleration and display is 
> decoupled, and acceleration does not require display.

Yes totally agreed.

> 
> At the protocol layer, context command buffers are carried by 
> VIRTIO_GPU_CMD_SUBMIT_3D. Scanout uses separate core virtio-gpu 
> commands, and VIRTIO_GPU_CMD_GET_DISPLAY_INFO may report no enabled 
> displays. At the implementation layer, QEMU handles scanout 
> presentation. virgl_cmd_set_scanout() obtains resource information 
> through virgl_renderer_resource_get_info() or 
> virgl_renderer_resource_get_info_ext(). That does not make scanout a 
> virglrenderer-owned display path.

Yes, agreed.

> 
> Therefore, if "sharing the display path" means sharing the same device, 
> control queue, and QEMU execution context, that identifies a possible 
> source of contention. If it means that capsets or virglrenderer are 
> inherently tied to display, I do not think that is accurate. Vulkan 
> compute is already used through Venus with libkrun [2], and VCL proposes 
> OpenCL support through virglrenderer [3].

Yes,but the vulkan is for GFX originally, and for some formal AI frame 
work like pytorch, it's support is limited, and it performance is lower 
than ROCm, and vulkan also lacks many AI infrastructure, like composable 
kernel.
And for virCL, actually it is came from same project with ROCm native 
context, but the original author didn't continue to support it, they 
handed it over to someone else to take over. And in the first version of
virCL, it didn't pass the test of actual projects.

And it seems like virCL didn't upstream into virglrenderer also, correct 
me if I am wrong.

>> Beyond that, sharing the display path is increasingly painful:
>>
>>    - Compute hammers the queues more than graphics, so sharing
>>      virtio gpu's single control queue with display/virgl causes 
>> contention
>>      and display stutter.
> 
> All non-cursor commands do share one control queue, but a fence avoids 
> serialization.

yes, agreed.

> 
> There may still be implementation-level contention, and it is not
> necessarily specific to compute. A sufficiently busy graphics workload
> could expose the same bottlenecks. Possible contributors in current QEMU
> include:
> 
> a) qemu_console_hw_gl_block() blocks the entire queue when QEMU only
>     needs to fence scanout commands.

Yes, agreed.

> 
> b) virtio_gpu_virgl_unmap_resource_blob() may also block the entire
>     queue just to delay one command.

Yes, but it is seems like it is must, someone else in AMD tried to use 
async method to relase blob, but it failed to consistency issue, then
reverted to sync version.

> 
> c) QEMU dispatches the control queue and calls into virglrenderer from
>     its main-loop thread along with display work and many other things.
>     Venus's render server can offload renderer work, but control-queue
>     dispatch remains in QEMU's main loop.

Yes, we did some async optimization in ROCm context, but its 
effectiveness is limited, see bellow.

> 
> In any case, I think you need to do some experiments to track down the 
> real cause. a) is easy to check: just comment out all 
> qemu_console_hw_gl_block() calls; it may corrupt display but removes the 
> blocking. b) can also be tested by leaking the mappings instead of 
> blocking the whole queue. Using a different display device like qxl 
> tells whether c) is causing contention.

Yes, totally agreed. following is my findings. In short words:

Optimization can reduce queue pressure, but it can't withstand absolute 
overload because each command has some overhead. Making all commands 
asynchronous would lead to a debugging hell about asynchronous issues.
And we have high load applications rocmprofiler  that continuously catch 
information need virtio queue to handle. But create a new backend can 
not solve it simply, we are trying to find a way. like shmem between 
guest and host, then use cpu polling, bypass the virtqueue.

The load is mostly memory management. Running an AI model allocates and 
frees a large number of blobs. We already did some optimization release 
them asynchronously, but the host processing is a single queue one
process_cmdq, this is where the main bottleneck in my debugging work / 
my understanding so far. I'm not certain it's the whole picture, so 
please correct if I am wrong.

A model load or unload frees a large batch of BOs and allocates another. 
Some of those commands are async in the virtio-gpu guest driver, but 
QEMU still has to work through them on the one queue, which takes time; 
so even though any single command is quick, there are simply too many of 
them, the single queue backs up, and everything behind it, gets delayed.

real work load (a few downstream customisations): loading one 16 GB 
model (gemm4 e4b), drives ~1200 blob creates, a burst of
~1400 resource frees at teardown, ~3700 submits and ~6000 virtqueue 
notifies, caused a 22 s guest soft lockup. And the behavior of memory 
operations are controlled by upper layer like pytorch / HIP / runtime,
we can not control it.

To be honest, a separate backend won't fix this. But the real solution 
maybe is compute specific. That logic is only useful to the compute 
path, and folding it into the shared display device / renderer would 
mean churning code that is mature and stable for graphics, with 
regression risk. Keeping compute on its own instance and backend lets us 
iterate on these compute only optimisations.

> 
>>    - Compute contexts need far more blob / shared memory than a 
>> display one.
> 
> It is not a problem by itself. Frequent mapping and unmapping might 
> amplify the second issue above, but that needs to be measured.

Yes, agreed. I can give more detailed information.

> 
>>    - Maybe needs a wider ROCm / compute stack, cause the render model 
>> fits poorly:
>>      rocprofiler (PC sampling, SQTT/SPM, counters, high bandwidth 
>> streams)
>>      and ROCgdb (wave control, address watch, async exceptions an
>>      out of band channel that must not block display).
> 
> virglrenderer does not impose a particular render model. That's why 
> Vulkan Compute just works with Venus.

Yes but vulkan is used for GFX initally. And can not support many AI 
application.>
>>    - Events, faults and GPU reset/SMI are async and don't map onto 
>> fences.>    - All of this is hard to extend cleanly inside a display 
>> capset.
> Capset is not about display but determines the protocol of the 
> VIRTIO_GPU_CMD_SUBMIT_3D command stream. You have described events, 
> faults and GPU reset/SMI are async don't map onto fences that may be 
> associated with VIRTIO_GPU_CMD_SUBMIT_3D which is dictated by capset. An 
> additional feature may be necessary, and it may or may not be dictated 
> by capset. The other things are irrelevant with the protocol capset 
> represents; they are either behavioral or about different commands.

A fence is the one shot, but event is stateful and repeatable.
That may or may not be tied to capset. Agreed.

> 
>>
>> On the QEMU/host side, would something like this be OK? One step, two 
>> parts:
>>
>>    - a dedicated headless virtio gpu instance for compute.
> 
> A second device would isolate its virtqueues and device-wide
> renderer_blocked state. That may be useful if measurements show that
> these are the bottlenecks, but it is not yet clear that they are or that
> a second device is the appropriate solution.
> 
>>    - that instance served by a separate ROCm backend library loaded
>>      in-process by QEMU.
> 
> First, I think we need to establish why ROCm cannot or should not remain
> in virglrenderer. The virglrenderer, Venus, and VCL maintainers are
> likely better placed to advise on that boundary. Once the protocol
> requirements and performance measurements are clear, we can assess the
> appropriate QEMU integration.

venus is borned for GFX.
virCL not merged.

To be clear, I'm not saying virglrenderer can't host a ROCm native 
context it clearly can. My hesitation is more about fit and direction: 
virglrenderer has grown up around GL/graphics, and I haven't yet found 
compute oriented plumbing there to build on, while ROCm moves very fast 
and I need something I can keep current with low friction.

Regards,
Honglei


> 
> [2] https://developers.redhat.com/articles/2025/06/05/how-we-improved- 
> ai-inference-macos-podman-containers
> [3] https://www.qualcomm.com/developer/blog/2024/10/vcl-virtio-gpu- 
> opencl-driver
> 
> Regards,
> Akihiko Odaki
> 
>>
>> That reuses the existing pluggable backend model, a second virtio gpu + a
>> backend library. It doesn't add dedicated queues for debug/profiling 
>> currently.
>>
>> Waiting for reply and  happy to share more detail. Thanks!
>>
>> [1] https://gitlab.freedesktop.org/virgl/virglrenderer/-/ 
>> merge_requests/1568
>>
>> Regards,
>> Honglei
> 


  reply	other threads:[~2026-08-17 13:44 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-17  3:19 About new backend for GPU compute ROCm in qemu Huang, Honglei
2026-08-17  9:06 ` Alex Bennée
2026-08-17 12:46   ` Huang, Honglei
2026-08-17 11:44 ` Akihiko Odaki
2026-08-17 13:44   ` Huang, Honglei [this message]
2026-08-17 14:24     ` Alex Bennée
2026-08-18  2:27       ` Huang, Honglei
2026-08-17 16:29     ` Akihiko Odaki
2026-08-18  2:50       ` Huang, Honglei
2026-08-18  4:05         ` Akihiko Odaki
2026-08-18  4:26           ` Huang, Honglei

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=78f0583f-93c0-4374-ba37-fd36f6388f0e@amd.com \
    --to=honghuan@amd.com \
    --cc=Ray.Huang@amd.com \
    --cc=airlied@redhat.com \
    --cc=alex.bennee@linaro.org \
    --cc=dmitry.osipenko@collabora.com \
    --cc=dri-devel@lists.freedesktop.org \
    --cc=honglei1.huang@amd.com \
    --cc=kraxel@redhat.com \
    --cc=marcandre.lureau@redhat.com \
    --cc=mst@redhat.com \
    --cc=odaki@rsg.ci.i.u-tokyo.ac.jp \
    --cc=peter.maydell@linaro.org \
    --cc=qemu-devel@nongnu.org \
    --cc=sgarzare@redhat.com \
    --cc=virtio-comment@lists.oasis-open.org \
    --cc=virtualization@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox