From: Heiko Carstens <hca@linux.ibm.com>
To: Alexander Gordeev <agordeev@linux.ibm.com>
Cc: Sven Schnelle <svens@linux.ibm.com>,
Vasily Gorbik <gor@linux.ibm.com>,
Christian Borntraeger <borntraeger@linux.ibm.com>,
Janosch Frank <frankja@linux.ibm.com>,
Claudio Imbrenda <imbrenda@linux.ibm.com>,
David Hildenbrand <david@kernel.org>,
linux-s390@vger.kernel.org, kvm@vger.kernel.org,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH v4 1/8] KVM: s390: pv: Use VM_SPARSE area for guest variable storage area
Date: Thu, 3 Sep 2026 19:24:17 +0200 [thread overview]
Message-ID: <20260903172417.12158Bf0-hca@linux.ibm.com> (raw)
In-Reply-To: <e9449c63-b8d2-4ed9-a184-7dcab3e8dfce-agordeev@linux.ibm.com>
On Thu, Sep 03, 2026 at 02:28:35PM +0200, Alexander Gordeev wrote:
> On Mon, Jul 20, 2026 at 10:58:27AM +0200, Heiko Carstens wrote:
> ...
> > This assumes that s390 will gain full support for lazy_mmu_mode_enable()
> > and lazy_mmu_mode_disable() in the future, since as of now the used
> > ptep_get_and_clear() in vunmap_pte_range() does indeed invalidate and
> > flush every single pte entry, but only for s390.
> ...
> > +static int uv_alloc_range_cb(pte_t *ptep, unsigned long addr, void *data)
> > +{
> > + struct page *page;
> > + pte_t pte;
> > +
> > + page = alloc_page(GFP_KERNEL_ACCOUNT | __GFP_ZERO);
>
> In lazy mode this callback is called with preemption disabled, so it leads to:
>
> [ 138.709287] BUG: sleeping function called from invalid context at ./include/linux/sched/mm.h:322
> [ 138.709535] in_atomic(): 1, irqs_disabled(): 0, non_block: 0, pid: 6092, name: qemu-kvm
...
Any chance you missed to provide an important detail: this can only
happen with your not-yet upstream series, but as of now there is no
upstream problem?
Is this correct?
> This could be solved using the below fixup:
>
> @@ -247,7 +247,9 @@ static int uv_alloc_range_cb(pte_t *ptep, unsigned long addr, void *data)
> struct page *page;
> pte_t pte;
>
> + lazy_mmu_mode_pause();
> page = alloc_page(GFP_KERNEL_ACCOUNT | __GFP_ZERO);
> + lazy_mmu_mode_resume();
> if (!page)
> return -ENOMEM;
> pte = __pte(page_to_phys(page) | pgprot_val(PAGE_KERNEL));
>
> The downside is lazy_mmu_mode_resume() does not really re-enable
> the caching and the performance will stay the same even when the
> lazy mode is supported on s390.
>
> This is the same pattern as kasan_populate_vmalloc_pte() - which
> was the only occurrence so far.
>
> Alternatively, the page could be allocated atomically, but I think
> that is less preferrable.
I guess the real fix is to pre-allocate all pages, keep the pointers
to the struct pages in an array, and populate with a different
mechanism. Similar like the vmalloc code is doing.
I wanted to keep this code as simple as possible, but...
> > + if (apply_to_page_range(&init_mm, addr, size, uv_alloc_range_cb, NULL))
>
> lazy_mmu_mode_enable_with_ptes() called from apply_to_pte_range()
> disables preemption.
There is no lazy_mmu_mode_enable_with_ptes() upstream, nor does s390
select ARCH_HAS_LAZY_MMU_MODE until now. Therefore my question above.
I guess my assumption above is correct and there is real urgency to
fix this.
next prev parent reply other threads:[~2026-09-03 17:24 UTC|newest]
Thread overview: 32+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-20 8:58 [PATCH v4 0/8] s390: Reintroduce support for DCACHE_WORD_ACCESS Heiko Carstens
2026-07-20 8:58 ` [PATCH v4 1/8] KVM: s390: pv: Use VM_SPARSE area for guest variable storage area Heiko Carstens
2026-07-20 9:14 ` sashiko-bot
2026-07-20 9:56 ` Christian Borntraeger
2026-07-20 10:15 ` Heiko Carstens
2026-09-03 12:28 ` Alexander Gordeev
2026-09-03 17:24 ` Heiko Carstens [this message]
2026-09-04 18:53 ` Heiko Carstens
2026-07-20 8:58 ` [PATCH v4 2/8] s390/mm: Add missing mm check to do_secure_storage_access() Heiko Carstens
2026-07-20 9:12 ` sashiko-bot
2026-07-20 10:44 ` Christian Borntraeger
2026-07-20 8:58 ` [PATCH v4 3/8] s390/mm: Use lock_mm_and_find_vma() in do_secure_storage_access() Heiko Carstens
2026-07-20 9:19 ` sashiko-bot
2026-07-20 10:45 ` Christian Borntraeger
2026-07-20 8:58 ` [PATCH v4 4/8] s390/mm: Fix handling of vmalloc area " Heiko Carstens
2026-07-20 9:23 ` sashiko-bot
2026-07-20 10:22 ` Christian Borntraeger
2026-07-20 8:58 ` [PATCH v4 5/8] s390/mm: Remove folio handling for kernel faults " Heiko Carstens
2026-07-20 9:30 ` sashiko-bot
2026-07-20 10:53 ` Christian Borntraeger
2026-07-24 8:17 ` Claudio Imbrenda
2026-07-20 8:58 ` [PATCH v4 6/8] s390/mm: Use handle_fault_error() " Heiko Carstens
2026-07-20 9:26 ` sashiko-bot
2026-07-20 8:58 ` [PATCH v4 7/8] s390/mm: Use goto statement " Heiko Carstens
2026-07-20 9:36 ` sashiko-bot
2026-07-20 10:36 ` Christian Borntraeger
2026-07-20 8:58 ` [PATCH v4 8/8] s390: Add support for DCACHE_WORD_ACCESS (again) Heiko Carstens
2026-07-20 9:48 ` sashiko-bot
2026-07-21 9:59 ` Sven Schnelle
2026-07-20 9:03 ` [PATCH v4 0/8] s390: Reintroduce support for DCACHE_WORD_ACCESS Christian Borntraeger
2026-07-20 9:40 ` Heiko Carstens
2026-07-27 10:37 ` Vasily Gorbik
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260903172417.12158Bf0-hca@linux.ibm.com \
--to=hca@linux.ibm.com \
--cc=agordeev@linux.ibm.com \
--cc=borntraeger@linux.ibm.com \
--cc=david@kernel.org \
--cc=frankja@linux.ibm.com \
--cc=gor@linux.ibm.com \
--cc=imbrenda@linux.ibm.com \
--cc=kvm@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-s390@vger.kernel.org \
--cc=svens@linux.ibm.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox