Linux s390 Architecture development
 help / color / mirror / Atom feed
From: sashiko-bot@kernel.org
To: "Alexander Gordeev" <agordeev@linux.ibm.com>
Cc: "Vasily Gorbik" <gor@linux.ibm.com>,
	"Christian Borntraeger" <borntraeger@linux.ibm.com>,
	"Heiko Carstens" <hca@linux.ibm.com>,
	linux-s390@vger.kernel.org,
	"Alexander Gordeev" <agordeev@linux.ibm.com>
Subject: Re: [PATCH v8 13/15] s390/mm: Batch PTE updates in lazy MMU mode
Date: Wed, 07 Oct 2026 12:06:03 +0000	[thread overview]
Message-ID: <sashiko-outbox-162972@kernel.org> (raw)
In-Reply-To: <277e24d315e55592c367ac0c347d44cac5e429f4.1791365932.git.agordeev@linux.ibm.com>

Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
- [High] s390/mm: NULL pointer deref on CPU hotplug lazy MMU rollback failure
- [High] s390/mm: Hardirq torn read of ipte_range bounds causes OOB memory write
- [High] s390/mm: ptep_get_and_clear_full() bypasses fullmm optimization in lazy MMU mode
- [Medium] s390/mm: ptep_get_and_clear_full() bypasses fullmm TLB optimization
--

commit 93b412133f9ff3548a7d0d31445ca8ff8b51da29
Author: Alexander Gordeev <agordeev@linux.ibm.com>

s390/mm: Batch PTE updates in lazy MMU mode

This patch introduces a per-CPU cache to batch PTE updates (via the IPTE
instruction) during lazy MMU mode. This reduces CEC-wide quiesce traffic
for memory-management system calls like mprotect(), munmap(), and mremap(),
yielding significant performance improvements.

> diff --git a/arch/s390/include/asm/pgtable.h b/arch/s390/include/asm/pgtable.h
> index c47264f3abf21..decce76fac743 100644
> --- a/arch/s390/include/asm/pgtable.h
> +++ b/arch/s390/include/asm/pgtable.h
[ ... ]
> @@ -1263,11 +1383,13 @@ static inline pte_t ptep_get_and_clear_full(struct mm_struct *mm,
>  {
>  	pte_t res;
>  
> -	if (full) {
> -		res = ptep_get(ptep);
> -		set_pte(ptep, __pte(_PAGE_INVALID));
> -	} else {
> -		res = ptep_xchg_lazy(mm, addr, ptep, __pte(_PAGE_INVALID));
> +	if (!lazy_mmu_ptep_get_and_clear(addr, ptep, &res)) {
> +		if (full) {
> +			res = __ptep_get(ptep);
> +			__set_pte(ptep, __pte(_PAGE_INVALID));
> +		} else {
> +			res = ptep_xchg_lazy(mm, addr, ptep, __pte(_PAGE_INVALID));
> +		}
>  	}

[Severity: High]
Does intercepting the clear operation via lazy_mmu_ptep_get_and_clear() before
checking the full parameter bypass the fullmm optimization?

During process teardown (e.g. exit_mmap()), fullmm is set to 1 to skip
individual IPTE (TLB flush) instructions because a full MM TLB flush will be
performed at the end. By unconditionally caching the clear operation in the
lazy MMU cache here, leave_ipte_range() will later issue batch IPTE
instructions for these entries when lazy MMU mode is disabled. This appears
to degrade the O(1) TLB teardown into O(N) CEC-wide synchronizing IPTE
instructions.

[ ... ]
> diff --git a/arch/s390/mm/lazy_mmu.c b/arch/s390/mm/lazy_mmu.c
> new file mode 100644
> index 0000000000000..8c1a62f237362
> --- /dev/null
> +++ b/arch/s390/mm/lazy_mmu.c
[ ... ]
> +static void enter_ipte_range(struct mm_struct *mm,
> +			     unsigned long addr, unsigned long end, hw_pte_t *pte)
> +{
> +	struct ipte_range *range;
> +
> +	if (!static_branch_likely(&lazy_mmu_enabled))
> +		return;
> +
> +	range = get_cpu_var(ipte_range);

[Severity: High]
If lazy_mmu_alloc() fails during CPU hotplug, could this retrieve a NULL
ipte_range pointer?

When a lazy_mmu_alloc() failure triggers a CPU hotplug rollback,
lazy_mmu_cpu_offline() frees the memory but does not disable the globally
active lazy_mmu_enabled static branch. Subsequent lazy MMU entries on that
CPU would bypass the static branch check above, retrieve a NULL ipte_range,
and dereference it below.

> +	local_bh_disable();
> +	get_lowcore()->lazy_mmu_count++;
> +
> +	if (mm_is_protected(mm)) {
> +		local_bh_enable();
> +		return;
> +	}
> +
> +	range->mm = mm;
> +	range->base_addr = addr;
> +	range->base_end = end;

[Severity: High]
Can a hardirq observe a torn state here?

The base_addr and base_end fields are updated sequentially. If a hardirq or
NMI strikes after base_addr is set but while base_end is still 0, a
concurrent execution of this_ipte_range() would observe base_end == 0.

This would cause the calculation:
nr_ptes = (range->base_end - range->base_addr) / PAGE_SIZE;

to underflow to a massive value, bypassing the bounds validation and allowing
an out-of-bounds memory write to range->cache[idx].

> +	range->base_pte = pte;
> +
> +	local_bh_enable();
> +}

-- 
Sashiko AI review · https://sashiko.dev/#/patchset/cover.1791365932.git.agordeev@linux.ibm.com?part=13

  reply	other threads:[~2026-10-07 12:06 UTC|newest]

Thread overview: 37+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-07 11:41 [PATCH v8 00/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 01/15] mm: introduce hw_pte_t for PTE table storage Alexander Gordeev
2026-10-07 11:55   ` sashiko-bot
2026-10-07 11:41 ` [PATCH v8 02/15] mm: rename pointers to software PTE values as ptentp Alexander Gordeev
2026-10-07 11:54   ` sashiko-bot
2026-10-07 11:41 ` [PATCH v8 03/15] mm: use hw_pte_t for generic PTE table storage Alexander Gordeev
2026-10-07 12:02   ` sashiko-bot
2026-10-07 11:41 ` [PATCH v8 04/15] mm: convert PTE table entries in ptep_get() Alexander Gordeev
2026-10-07 11:55   ` sashiko-bot
2026-10-07 11:41 ` [PATCH v8 05/15] mm: convert PTE table entry to pte Alexander Gordeev
2026-10-07 11:50   ` sashiko-bot
2026-10-07 11:41 ` [PATCH v8 06/15] mm: add hw_pte_val for HW PTE storage Alexander Gordeev
2026-10-07 11:51   ` sashiko-bot
2026-10-07 11:41 ` [PATCH v8 07/15] mm/kasan: use hw_pte_t for the early shadow PTE table Alexander Gordeev
2026-10-07 11:49   ` sashiko-bot
2026-10-07 11:41 ` [PATCH v8 08/15] drm/i915: use hw_pte_t for PTE range callbacks Alexander Gordeev
2026-10-07 11:51   ` sashiko-bot
2026-10-07 11:41 ` [PATCH v8 09/15] xen: " Alexander Gordeev
2026-10-07 11:52   ` sashiko-bot
2026-10-07 11:41 ` [PATCH v8 10/15] s390/mm: Cleanup pXXp_flush_lazy() routines Alexander Gordeev
2026-10-07 11:52   ` sashiko-bot
2026-10-07 11:41 ` [PATCH v8 11/15] s390: Distinguish hardware and software PTEs Alexander Gordeev
2026-10-07 11:57   ` sashiko-bot
2026-10-07 12:01     ` Alexander Gordeev
2026-10-08 10:32   ` kernel test robot
2026-10-07 11:41 ` [PATCH v8 12/15] mm: Make lazy MMU mode context-aware Alexander Gordeev
2026-10-07 11:53   ` sashiko-bot
2026-10-07 11:41 ` [PATCH v8 13/15] s390/mm: Batch PTE updates in lazy MMU mode Alexander Gordeev
2026-10-07 12:06   ` sashiko-bot [this message]
2026-10-08  6:32     ` Alexander Gordeev
2026-10-07 11:41 ` [PATCH v8 14/15] mm/kasan: Introduce helpers for lazy MMU mode sanitizer Alexander Gordeev
2026-10-07 11:56   ` sashiko-bot
2026-10-07 20:24   ` kernel test robot
2026-10-07 20:35   ` kernel test robot
2026-10-08  1:11   ` kernel test robot
2026-10-07 11:41 ` [PATCH v8 15/15] s390/mm: Lazy " Alexander Gordeev
2026-10-07 11:53   ` sashiko-bot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=sashiko-outbox-162972@kernel.org \
    --to=sashiko-bot@kernel.org \
    --cc=agordeev@linux.ibm.com \
    --cc=borntraeger@linux.ibm.com \
    --cc=gor@linux.ibm.com \
    --cc=hca@linux.ibm.com \
    --cc=linux-s390@vger.kernel.org \
    --cc=sashiko-reviews@lists.linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox