All of lore.kernel.org
 help / color / mirror / Atom feed
From: "Jürgen Groß" <jgross@suse.com>
To: Jan Beulich <jbeulich@suse.com>
Cc: dfaggioli@suse.com, gwd@xenproject.org, roger@xenproject.org,
	anthony.perard@vates.tech, julien@xen.org,
	bertrand.marquis@arm.com, michal.orzel@amd.com,
	Volodymyr_Babchuk@epam.com, teddy.astie@vates.tech,
	Furkan Caliskan <frn1furkan10@gmail.com>,
	xen-devel@lists.xenproject.org,
	Andrew Cooper <andrew.cooper3@citrix.com>
Subject: Re: [PATCH v3 1/2] xen/common: add vcpus_create() and keep max_vcpus in sync
Date: Tue, 1 Sep 2026 11:03:49 +0200	[thread overview]
Message-ID: <4a8c362a-9aac-4e0c-aea6-bc5841143ce6@suse.com> (raw)
In-Reply-To: <844b82eb-5a8b-4e49-8a97-f1311120b61f@suse.com>


[-- Attachment #1.1.1: Type: text/plain, Size: 3987 bytes --]

On 01.09.26 10:36, Jan Beulich wrote:
> On 01.09.2026 10:24, Jürgen Groß wrote:
>> On 01.09.26 10:15, Jan Beulich wrote:
>>> On 01.09.2026 10:07, Jürgen Groß wrote:
>>>> On 01.09.26 09:22, Jan Beulich wrote:
>>>>> On 31.08.2026 14:59, Jürgen Groß wrote:
>>>>>> On 31.08.26 11:45, Andrew Cooper wrote:
>>>>>>> On 31/08/2026 6:16 am, Furkan Caliskan wrote:
>>>>>>>> Every vcpu_create() call site that builds more than one vcpu loops
>>>>>>>> over ids up to d->max_vcpus and stops on the first failure, but none
>>>>>>>> of them roll max_vcpus back to match. This leaves d->vcpu[i] == NULL
>>>>>>>> for ids below max_vcpus
>>>>>>>
>>>>>>> As I told you before, you must cope with this property in non-error
>>>>>>> scenarios.
>>>>>>>
>>>>>>>> , which anything walking d->vcpu[] can then
>>>>>>>> dereference. This is what caused the crash: sched_move_domain()
>>>>>>>> walks every vcpu slot up to max_vcpus without checking for empty
>>>>>>>> ones, so when a domain built in a non-default cpupool had vcpu
>>>>>>>> creation fail partway through, domain_kill() later moving it back
>>>>>>>> to the default cpupool handed one of its empty slots straight to
>>>>>>>> the new cpupool's scheduler, causing a NULL-pointer dereference
>>>>>>>> inside sched_alloc_udata().
>>>>>>>>
>>>>>>>> Add vcpus_create(d): creates every vcpu of d up to max_vcpus and
>>>>>>>> rolls max_vcpus back to the failed id on error. This keeps
>>>>>>>> d->vcpu[i] is non-NULL for all i < d->max_vcpus
>>>>>>>
>>>>>>> No, it really doesn't.
>>>>>>>
>>>>>>>> , instead of guarding
>>>>>>>> every reader of d->vcpu[] agains holes individually.
>>>>>>>>
>>>>>>>> Convert every site that builds vcpus in a loop to call this function
>>>>>>>> instead.
>>>>>>>>
>>>>>>>> Fixes: 61649709421a ("xen/domain: Allocate d->vcpu[] in domain_create()")
>>>>>>>> Suggested-by: Juergen Gross <jgross@suse.com>
>>>>>>>> Signed-off-by: Furkan Caliskan <frn1furkan10@gmail.com>
>>>>>>>
>>>>>>> For the avoidance of a long drawn-out argument, nack.  Under no
>>>>>>> circumstances are you editing d->max_cpus after it's put into the domain
>>>>>>> list.
>>>>>>>
>>>>>>> You've chosen to do so at a point where the domain object is live,
>>>>>>> visible in the system and able to be the target of other hypercalls.
>>>>>>>
>>>>>>> Furthermore you have not fixed what your commit message claims.
>>>>>>> d->vcpu[...] is still NULL for an arbitrary period of time, including
>>>>>>> being able to be the target of hypercalls, before vCPUs are created.
>>>>>>>
>>>>>>> All code MUST be able to cope with d->vcpu[...] being NULL.  It's how
>>>>>>> the object lifecycles must work, because creating vCPUs is not atomic
>>>>>>> with respect to creating domains.
>>>>>>
>>>>>> Would you be fine with me creating a patch series moving vcpu creation into
>>>>>> domain_create()?
>>>>>
>>>>> This was discussed before, and however nice it would be for the issue at hand,
>>>>> it would get in the way of us wanting to have CPU policy for domains put in
>>>>> place before vCPU-s are created, such that on x86 the XSAVE area can be sized
>>>>> once and for all.
>>>>
>>>> This could be done when unpausing the domain initially.
>>>
>>> Imo unpausing shouldn't fail because of memory shortage.
>>
>> As long as the domain hasn't started running I don't see why this would be
>> different to the case where not all vcpus could be created.
> 
> I do. One could create a domain ready to be unpaused, but being kept paused
> until whatever event triggers its launching. That better wouldn't fail, except
> in extraordinary situations.

Okay, then we could add XEN_DOMCTL_finalize doing the final allocations AND
doing the creation_finished handling. This would even catch today's domain
crashing in the VMX specific domain_creation_finished() callback.

At the same time XEN_DOMCTL_max_vcpus would be dropped, so there isn't more
work for the toolstack.


Juergen

[-- Attachment #1.1.2: OpenPGP public key --]
[-- Type: application/pgp-keys, Size: 3743 bytes --]

[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 495 bytes --]

  reply	other threads:[~2026-09-01  9:23 UTC|newest]

Thread overview: 15+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-31  5:16 [PATCH v3 0/2] xen/sched: fix crashes when vcpu creation fails Furkan Caliskan
2026-08-31  5:16 ` [PATCH v3 1/2] xen/common: add vcpus_create() and keep max_vcpus in sync Furkan Caliskan
2026-08-31  8:13   ` Jürgen Groß
2026-08-31  9:13     ` Furkan Çalışkan
2026-08-31  9:32   ` [PATCH v4] " Furkan Caliskan
2026-08-31  9:45   ` [PATCH v3 1/2] " Andrew Cooper
2026-08-31 12:59     ` Jürgen Groß
2026-09-01  7:22       ` Jan Beulich
2026-09-01  8:07         ` Jürgen Groß
2026-09-01  8:15           ` Jan Beulich
2026-09-01  8:24             ` Jürgen Groß
2026-09-01  8:36               ` Jan Beulich
2026-09-01  9:03                 ` Jürgen Groß [this message]
2026-08-31  5:16 ` [PATCH v3 2/2] xen/sched: core: kill unarmed timers on sched_init_vcpu() failure Furkan Caliskan
2026-08-31  8:15   ` Jürgen Groß

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=4a8c362a-9aac-4e0c-aea6-bc5841143ce6@suse.com \
    --to=jgross@suse.com \
    --cc=Volodymyr_Babchuk@epam.com \
    --cc=andrew.cooper3@citrix.com \
    --cc=anthony.perard@vates.tech \
    --cc=bertrand.marquis@arm.com \
    --cc=dfaggioli@suse.com \
    --cc=frn1furkan10@gmail.com \
    --cc=gwd@xenproject.org \
    --cc=jbeulich@suse.com \
    --cc=julien@xen.org \
    --cc=michal.orzel@amd.com \
    --cc=roger@xenproject.org \
    --cc=teddy.astie@vates.tech \
    --cc=xen-devel@lists.xenproject.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.