All of lore.kernel.org
 help / color / mirror / Atom feed
From: Jan Beulich <jbeulich@suse.com>
To: "Jürgen Groß" <jgross@suse.com>,
	"Andrew Cooper" <andrew.cooper3@citrix.com>
Cc: dfaggioli@suse.com, gwd@xenproject.org, roger@xenproject.org,
	anthony.perard@vates.tech, julien@xen.org,
	bertrand.marquis@arm.com, michal.orzel@amd.com,
	Volodymyr_Babchuk@epam.com, teddy.astie@vates.tech,
	Furkan Caliskan <frn1furkan10@gmail.com>,
	xen-devel@lists.xenproject.org
Subject: Re: [PATCH v3 1/2] xen/common: add vcpus_create() and keep max_vcpus in sync
Date: Tue, 1 Sep 2026 10:15:02 +0200	[thread overview]
Message-ID: <fcbb592a-b321-467a-ba3f-16aef00803af@suse.com> (raw)
In-Reply-To: <258aa564-4856-4d28-8b90-aa858b25fdd2@suse.com>

On 01.09.2026 10:07, Jürgen Groß wrote:
> On 01.09.26 09:22, Jan Beulich wrote:
>> On 31.08.2026 14:59, Jürgen Groß wrote:
>>> On 31.08.26 11:45, Andrew Cooper wrote:
>>>> On 31/08/2026 6:16 am, Furkan Caliskan wrote:
>>>>> Every vcpu_create() call site that builds more than one vcpu loops
>>>>> over ids up to d->max_vcpus and stops on the first failure, but none
>>>>> of them roll max_vcpus back to match. This leaves d->vcpu[i] == NULL
>>>>> for ids below max_vcpus
>>>>
>>>> As I told you before, you must cope with this property in non-error
>>>> scenarios.
>>>>
>>>>> , which anything walking d->vcpu[] can then
>>>>> dereference. This is what caused the crash: sched_move_domain()
>>>>> walks every vcpu slot up to max_vcpus without checking for empty
>>>>> ones, so when a domain built in a non-default cpupool had vcpu
>>>>> creation fail partway through, domain_kill() later moving it back
>>>>> to the default cpupool handed one of its empty slots straight to
>>>>> the new cpupool's scheduler, causing a NULL-pointer dereference
>>>>> inside sched_alloc_udata().
>>>>>
>>>>> Add vcpus_create(d): creates every vcpu of d up to max_vcpus and
>>>>> rolls max_vcpus back to the failed id on error. This keeps
>>>>> d->vcpu[i] is non-NULL for all i < d->max_vcpus
>>>>
>>>> No, it really doesn't.
>>>>
>>>>> , instead of guarding
>>>>> every reader of d->vcpu[] agains holes individually.
>>>>>
>>>>> Convert every site that builds vcpus in a loop to call this function
>>>>> instead.
>>>>>
>>>>> Fixes: 61649709421a ("xen/domain: Allocate d->vcpu[] in domain_create()")
>>>>> Suggested-by: Juergen Gross <jgross@suse.com>
>>>>> Signed-off-by: Furkan Caliskan <frn1furkan10@gmail.com>
>>>>
>>>> For the avoidance of a long drawn-out argument, nack.  Under no
>>>> circumstances are you editing d->max_cpus after it's put into the domain
>>>> list.
>>>>
>>>> You've chosen to do so at a point where the domain object is live,
>>>> visible in the system and able to be the target of other hypercalls.
>>>>
>>>> Furthermore you have not fixed what your commit message claims.
>>>> d->vcpu[...] is still NULL for an arbitrary period of time, including
>>>> being able to be the target of hypercalls, before vCPUs are created.
>>>>
>>>> All code MUST be able to cope with d->vcpu[...] being NULL.  It's how
>>>> the object lifecycles must work, because creating vCPUs is not atomic
>>>> with respect to creating domains.
>>>
>>> Would you be fine with me creating a patch series moving vcpu creation into
>>> domain_create()?
>>
>> This was discussed before, and however nice it would be for the issue at hand,
>> it would get in the way of us wanting to have CPU policy for domains put in
>> place before vCPU-s are created, such that on x86 the XSAVE area can be sized
>> once and for all.
> 
> This could be done when unpausing the domain initially.

Imo unpausing shouldn't fail because of memory shortage.

> OTOH I don't see xstate_alloc_save_area() looking at the domain's cpu policy
> at all. Is this a plan for the future?

This is to better accommodate the AMX series (which has been pending for years),
and potentially also for architectural-LBR work (which has been posted once, but
was apparently abandoned).

> And additionally there is no guard for avoiding the vcpus being created before
> the policy is being set.

Addressing that is part of Andrew's plan, aiui.

Jan


  reply	other threads:[~2026-09-01  8:15 UTC|newest]

Thread overview: 15+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-31  5:16 [PATCH v3 0/2] xen/sched: fix crashes when vcpu creation fails Furkan Caliskan
2026-08-31  5:16 ` [PATCH v3 1/2] xen/common: add vcpus_create() and keep max_vcpus in sync Furkan Caliskan
2026-08-31  8:13   ` Jürgen Groß
2026-08-31  9:13     ` Furkan Çalışkan
2026-08-31  9:32   ` [PATCH v4] " Furkan Caliskan
2026-08-31  9:45   ` [PATCH v3 1/2] " Andrew Cooper
2026-08-31 12:59     ` Jürgen Groß
2026-09-01  7:22       ` Jan Beulich
2026-09-01  8:07         ` Jürgen Groß
2026-09-01  8:15           ` Jan Beulich [this message]
2026-09-01  8:24             ` Jürgen Groß
2026-09-01  8:36               ` Jan Beulich
2026-09-01  9:03                 ` Jürgen Groß
2026-08-31  5:16 ` [PATCH v3 2/2] xen/sched: core: kill unarmed timers on sched_init_vcpu() failure Furkan Caliskan
2026-08-31  8:15   ` Jürgen Groß

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=fcbb592a-b321-467a-ba3f-16aef00803af@suse.com \
    --to=jbeulich@suse.com \
    --cc=Volodymyr_Babchuk@epam.com \
    --cc=andrew.cooper3@citrix.com \
    --cc=anthony.perard@vates.tech \
    --cc=bertrand.marquis@arm.com \
    --cc=dfaggioli@suse.com \
    --cc=frn1furkan10@gmail.com \
    --cc=gwd@xenproject.org \
    --cc=jgross@suse.com \
    --cc=julien@xen.org \
    --cc=michal.orzel@amd.com \
    --cc=roger@xenproject.org \
    --cc=teddy.astie@vates.tech \
    --cc=xen-devel@lists.xenproject.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.