From: Andrey Drobyshev <andrey.drobyshev@virtuozzo.com>
To: Vladimir Sementsov-Ogievskiy <vsementsov@yandex-team.ru>,
qemu-devel@nongnu.org
Cc: mst@redhat.com, sgarzare@redhat.com, farosas@suse.de,
peterx@redhat.com, dongli.zhang@oracle.com,
maciej.szmigiero@oracle.com, bchaney@akamai.com,
mark.kanda@oracle.com, den@openvz.org
Subject: Re: [PATCH v3 0/7] migration/cpr: support vhost-vsock devices
Date: Wed, 15 Jul 2026 15:11:48 +0300 [thread overview]
Message-ID: <a2f3622d-acd0-4d31-8111-ae34643e614c@virtuozzo.com> (raw)
In-Reply-To: <db23c15d-cfb2-49a8-b748-5195a36a50d9@yandex-team.ru>
On 7/15/26 12:48 PM, Vladimir Sementsov-Ogievskiy wrote:
> On 15.07.26 13:25, Andrey Drobyshev wrote:
>> On 7/14/26 6:03 PM, Vladimir Sementsov-Ogievskiy wrote:
>>> On 26.06.26 19:46, Andrey Drobyshev wrote:
>>>> v2 -> v3:
>>>>
>>>> Re-design device ownership hand-off (suggested by Dongli). Do not rely
>>>> on SETUP migration notifiers, as this approach forces us to reorder
>>>> generic migration code (see v2 discussion). Instead:
>>>>
>>>> * .pre_save() releases ownership on the source (RESET_OWNER) once VM
>>>> is stopped;
>>>> * .realize() is adjusted to defer device ownership acquisition for an
>>>> incoming CPR;
>>>> * .post_load() actually claims the device (VHOST_SET_OWNER);
>>>> * FAILED migration event callback re-aquires the ownership on the
>>>> source after re_save released it.
>>>>
>>>> v2:https://lore.kernel.org/qemu-devel/57c9c9b3-d758-489b-95b8-d16c258b9b0c@virtuozzo.com
>>>
>>> Hi!
>>>
>>
>> Hello Vladimir!
>>
>>> Two notes:
>>>
>>> 1. Did you consider migrating needed FDs through main migration channel, without
>>> use of CPR, like I do in (not yet landed) "[PATCH v19 00/15] virtio-net: live-TAP local migration" [1]
>>>
>>
>> AFAIU you're doing local same-host migration as a way to update QEMU
>> binary. With this approach transferring FDs through the main channel is
>> indeed cleaner. However in our own downstream we use enhanced in-place
>> memory preservation which is based entirely on cpr-exec migration mode.
>>
>> So, to answer your question: yes, we considered it, and it might work
>> for cpr-transfer where we're also doing same-host migration. But in
>> cpr-exec case there's only one QEMU process involved => we can't use
>> UNIX socket for migration transport => no SCM_RIGHTS FDs passing => FDs
>> should be passed via additional transport, no way around it.
>
>
> Hmm. Looking at code for fd-passing through migration channel:
>
> static bool load_fd(QEMUFile *f, void *pv, size_t size,
> const VMStateField *field, Error **errp)
> {
> int32_t *v = pv;
>
> if (migrate_mode() == MIG_MODE_CPR_EXEC) {
> qemu_get_sbe32s(f, v);
> return true;
> }
> ...
>
>
> Seems like passing FDs through main channel (like in my TAP-series) would
> work both for local migration (with two processes and UNIX socket) and for
> CPR_EXEC mode. So the FD is passed as simple number in case of cpr-exec.
>
> So, if simply pass FDs through migration state, it should work both for
> CPR_EXEC and local migration.
>
> (hmm, still, keeping in mind that TAP-series API still not negotiated,
> it's simpler to start from pure CPR approach, and add local-migration
> support in future)
>
Yes, my claim was too harsh - the FD passing mechanism itself probably
can be implemented via the main migration channel as well, even for
cpr-exec. E.g. if it's not a UNIX socket but a plain file on file
system - we should be good.
But then the issue is timing: VMSTATE_FD delivers the FD during the
device's vmstate load, i.e. around post_load. And cpr_state is loaded
early, before device realization, so the FD is available at realize
time. And for instance our vhost-vsock case does need FD already during
realize to read the backend feature set:
vhost_vsock_device_realize() ->
vhost_dev_init_backend() ->
vhost_dev_init_features()
So for me the exact transport used underneath for FDs passing is not
that important. I'm just exploiting the CPR API to get the passed FD at
an early .realize() stage. If we can somehow guarantee an early FDs
delivery through the main channel - that'll work.
Andrey
>>
>>
>>> 2. I don't know how much vhost-vsock differs from vhost-net, but for vhost-net I remember
>>> that RESET_OWNER + SET_OWNER is effectively equal to simply recreating vhost device on target,
>>> because in kernel vhost device is hardly bound to the process itself, and can't be passed to
>>> another process. So, we can pass only "empty" FD, not the initialized vhost device.
>>> That's why in [1] I only pass tap-fds (and some additional state), but vhost fds are simply
>>> reopened on target (or passed by mgmt app).
>>> Does vhost-vsock work in a different way? Is there real sense in passing vhost fds here?
>>>
>>>
>>> [1] https://lore.kernel.org/qemu-devel/20260714154246.1242856-1-vsementsov@yandex-team.ru/
>>>
>>
>> I'm also not extremely keen on vhost-net internals, but as I understand
>> with vhost-net we have 2 FDs: TAP FD for the network endpoint and vhost
>> FD for the memory tables, vring addresses etc. The TAP FD is
>> process-independent state - and we pass it via migration channel. The
>> vhost FD is entirely per owning process - and we're able to close it on
>> source, reopen on target and bind to the passed TAP FD.
>
> Yes, my understanding matches.
>
>>
>> However with vhost-vsock there's only one FD. It's used both for
>> plumbing and the endpoint. At the very least, vhost device state holds
>> guest CID registration in the global vhost_vsock_hash. If we go the
>> "recreate/reopen the device" path - CID is removed from the hash in
>> vhost_vsock_dev_release(), and then we're gonna get ECONNRESET for any
>> incoming packets => connections are dead. We can't afford that.
>>
>
> OK, thanks for explanation!
>
>
prev parent reply other threads:[~2026-07-15 12:12 UTC|newest]
Thread overview: 26+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-06-26 16:46 [PATCH v3 0/7] migration/cpr: support vhost-vsock devices Andrey Drobyshev
2026-06-26 16:46 ` [PATCH v3 1/7] vhost: add vhost_reset_owner op Andrey Drobyshev
2026-06-26 16:46 ` [PATCH v3 2/7] vhost-vsock: don't reset connections during CPR Andrey Drobyshev
2026-06-26 16:46 ` [PATCH v3 3/7] vhost-vsock: fix FD leak in realize() Andrey Drobyshev
2026-06-26 16:46 ` [PATCH v3 4/7] vhost-vsock: preserve vhost FD during CPR Andrey Drobyshev
2026-08-18 15:29 ` Stefano Garzarella
2026-08-19 14:38 ` Andrey Drobyshev
2026-08-19 15:11 ` Michael S. Tsirkin
2026-08-19 15:22 ` Andrey Drobyshev
2026-08-19 18:53 ` Michael S. Tsirkin
2026-06-26 16:46 ` [PATCH v3 5/7] vhost: factor out vhost_dev_init_backend() Andrey Drobyshev
2026-08-18 15:29 ` Stefano Garzarella
2026-08-19 14:38 ` Andrey Drobyshev
2026-06-26 16:46 ` [PATCH v3 6/7] vhost: add vhost_dev_set_owner() / vhost_dev_reset_owner() helpers Andrey Drobyshev
2026-06-26 16:46 ` [PATCH v3 7/7] vhost-vsock: hand off device ownership across CPR Andrey Drobyshev
2026-08-18 15:30 ` Stefano Garzarella
2026-08-19 14:38 ` Andrey Drobyshev
2026-06-30 18:35 ` [PATCH v3 0/7] migration/cpr: support vhost-vsock devices Maciej S. Szmigiero
2026-08-11 14:12 ` Andrey Drobyshev
2026-08-18 11:02 ` Andrey Drobyshev
2026-08-18 15:33 ` Stefano Garzarella
2026-08-19 14:39 ` Andrey Drobyshev
2026-07-14 16:03 ` Vladimir Sementsov-Ogievskiy
2026-07-15 10:25 ` Andrey Drobyshev
2026-07-15 10:48 ` Vladimir Sementsov-Ogievskiy
2026-07-15 12:11 ` Andrey Drobyshev [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=a2f3622d-acd0-4d31-8111-ae34643e614c@virtuozzo.com \
--to=andrey.drobyshev@virtuozzo.com \
--cc=bchaney@akamai.com \
--cc=den@openvz.org \
--cc=dongli.zhang@oracle.com \
--cc=farosas@suse.de \
--cc=maciej.szmigiero@oracle.com \
--cc=mark.kanda@oracle.com \
--cc=mst@redhat.com \
--cc=peterx@redhat.com \
--cc=qemu-devel@nongnu.org \
--cc=sgarzare@redhat.com \
--cc=vsementsov@yandex-team.ru \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.