From: "Christophe Leroy (CS GROUP)" <chleroy@kernel.org>
To: Shrikanth Hegde <sshegde@linux.ibm.com>,
Jirka Hladky <jhladky@redhat.com>
Cc: maddy@linux.ibm.com, linuxppc-dev@lists.ozlabs.org,
mpe@ellerman.id.au, npiggin@gmail.com, bigeasy@linutronix.de,
will@kernel.org, linux-kernel@vger.kernel.org,
"Paul E . McKenney" <paulmck@kernel.org>
Subject: Re: [PATCH v4 1/1] powerpc: enable dynamic preemption
Date: Mon, 27 Jul 2026 12:35:32 +0200 [thread overview]
Message-ID: <330ca432-ddb6-4eeb-8415-697e16e73c43@kernel.org> (raw)
In-Reply-To: <85cb5e0e-feb5-4a54-b7b6-a5f6a053dcda@linux.ibm.com>
Hi Shrikanth,
Le 27/07/2026 à 12:28, Shrikanth Hegde a écrit :
> Hi Jirka,
> Please avoid top-posting.
>
> On 7/27/26 3:43 PM, Jirka Hladky wrote:
>> Hi Shrikanth,
>>
>> Thanks for the quick response and for looping in Paul.
>>
>>> This is true only if user selected PREEMPT_DYNAMIC option i think.
>>
>> Yes, exactly. The issue is that CONFIG_PREEMPT_DYNAMIC=y has been in
>> the Fedora/RHEL kernel config since kernel 5.16 (2021). It was
>> silently ignored on ppc64le until your 6.16 patch added
>> HAVE_PREEMPT_DYNAMIC_KEY, so this is the first time PREEMPT_RCU
>> actually takes effect for all Fedora/RHEL ppc64le users.
>>
>>> Does your preemption mode remain the same in two cases?
>>
>> Verified. On 6.15-rc6 (before PREEMPT_DYNAMIC takes effect):
>>
>> CONFIG_PREEMPT_VOLUNTARY=y
>> (no CONFIG_PREEMPT_DYNAMIC, no /sys/kernel/debug/sched/preempt)
>>
>> On both 7.1 and 7.2-rc4 (with PREEMPT_DYNAMIC active):
>>
>> CONFIG_PREEMPT_LAZY=y
>> CONFIG_PREEMPT_DYNAMIC=y
>> CONFIG_PREEMPT_RCU=y
>> cat /sys/kernel/debug/sched/preempt: "full (lazy)"
>>
>> So the runtime preemption mode is the same on 7.1 and 7.2. The
>> difference between 6.15 and 7.1+ is that the preemption model changed
>> from voluntary (static) to full/lazy (dynamic), which is what pulls
>> in PREEMPT_RCU.
>
> That means comparison is between preempt=voluntary vs preempt=lazy.
As far as I can see CONFIG_PREEMPT_LAZY doesn't pull CONFIG_PREEMPT_RCU.
CONFIG_PREEMPT_DYNAMIC does:
config PREEMPT_RCU
bool
default y if (PREEMPT || PREEMPT_RT || PREEMPT_DYNAMIC)
select TREE_RCU
help
This option selects the RCU implementation that is
designed for very large SMP systems with hundreds or
thousands of CPUs, but for which real-time response
is also required. It also scales down nicely to
smaller systems.
Select this option if you are unsure.
Christophe
>
> If you make it full preemption in 6.15 you will likely see similar data
> as 7.1+.
> Only on 7.0/7.1 there is force switch to lazy/full. Can you give that a
> try?
>
> If it shows same data, that implies the regression is mainly due to change
> of preemption modes, rather than the static key stuff.
>
>>
>>> Weak memory model would need barriers irrespective of
>>> HAVE_PREEMPT_DYNAMIC_CALL or HAVE_PREEMPT_DYNAMIC_KEY.
>>> Static key too is expected to minimal cost.
>>
>> You're right, I should clarify -- the overhead is not in the static
>> key mechanism itself. The cost comes from __rcu_read_lock() and
>> __rcu_read_unlock() which are called when CONFIG_PREEMPT_RCU=y.
>> These need real memory barriers (lwsync/isync) on ppc64le regardless
>> of whether the dynamic mechanism uses keys or calls.
>>
>> The real question is: why does enabling PREEMPT_RCU cost ~33% on
>> ppc64le but only ~3% on x86_64? The answer is that x86_64's TSO
>> memory model makes the barriers in __rcu_read_lock/__rcu_read_unlock
>> essentially free, while ppc64le's weak ordering requires explicit
>> lwsync/isync instructions, which are expensive when called thousands
>> of times per second in the SELinux AVC hot path.
>
> Plus, it may call schedule in lazy/pull preemption.
>
>>
>> I also have new data from SELinux isolation testing that helps
>> quantify this. On kernel 7.1 (which already has PREEMPT_RCU=y):
>>
>> SELinux mode kill bogo-ops/sec vs enforcing
>> ------------ ----------------- ------------
>> Enforcing 69,107 baseline
>> Permissive 70,248 +1.6%
>> Disabled 93,566 +35.4%
>>
>> Permissive ~ enforcing confirms the overhead is not in SELinux policy
>> evaluation. Disabling SELinux removes the rcu_read_lock/unlock call
>> sites in the AVC path and recovers most of the performance -- but
>> still leaves a ~13% gap vs 6.12 (no PREEMPT_RCU), which is the base
>> cost of PREEMPT_RCU in the non-SELinux parts of the kill() path.
>>
>
> This seems strange. How come rcu lock/unlock depends on SELinux policy?
> One should call rcu lock/unlock if they are working with rcu updated
> fields.
>
> Does the policy change itself protected with rcu lock/unlock?
> I will check it up.
>
>> So the core issue is: on ppc64le, CONFIG_PREEMPT_RCU makes
>> rcu_read_lock/unlock significantly more expensive, and the kill()
>> syscall path hits them very heavily through SELinux AVC lookups.
>>
>> Thanks,
>> Jirka
>>
>> On Mon, Jul 27, 2026 at 6:19 AM Shrikanth Hegde
>> <sshegde@linux.ibm.com> wrote:
>>>
>>> +cc paul for any further/RCU insights.
>>>
>>> On 7/27/26 12:03 AM, Jirka Hladky wrote:
>>>> Hi Shrikanth, Christophe,
>>>>
>>>
>>> Hi Jirka, thanks for the report.
>>>
>>>> I'm seeing a significant performance regression on ppc64le after
>>>> this
>>>> patch landed in 6.16, caused by CONFIG_PREEMPT_RCU becoming active
>>>> once HAVE_PREEMPT_DYNAMIC_KEY is selected.
>>>>
>>>
>>> This is true only if user selected PREEMPT_DYNAMIC option i think.
>>>
>>> config PREEMPT_RCU
>>> bool
>>> default y if (PREEMPT || PREEMPT_RT || PREEMPT_DYNAMIC)
>>> select TREE_RCU
>>>
>>>
>>>> Benchmark: stress-ng kill stressor (tight kill() syscall loop),
>>>> single thread, POWER10 LPAR (8 vCPUs, 1 core SMT-8).
>>>>
>>>> Bisected across Fedora ELN kernel builds on ppc64le:
>>>>
>>>> kernel CONFIG_PREEMPT_RCU kill bogo-ops/sec
>>>> --- 6.15-rc6 (eln148) no 103,207
>>>> 6.16 (eln150) yes 70,281 (-32%)
>>>> 6.18 (eln154) yes 72,552 (-30%)
>>>>
>>>
>>> Does your preemption mode remain the same in two cases?
>>>
>>>> For comparison, x86_64 (AMD EPYC 9355P) with the same config change
>>>> shows only a 2.8% regression:
>>>>
>>>> 6.12 x86_64 37,436
>>>> 7.2 x86_64 36,392 (-2.8%)
>>>>
>>>> perf report shows the overhead comes from rcu_read_lock/unlock in
>>>> the
>>>> SELinux AVC path (check_kill_permission -> security_task_kill ->
>>>> selinux_task_kill -> avc_has_perm -> avc_lookup):
>>>>
>>>> Function 6.15 (no PREEMPT_RCU) 6.16 (PREEMPT_RCU)
>>>> --- avc_lookup 15.23% 24.79%
>>>> __rcu_read_lock ~0% 4.52%
>>>> __rcu_read_unlock ~0% 4.17%
>>>> selinux_task_kill 6.23% 7.35%
>>>> audit_signal_info* 0.94% 3.59%
>>>>
>>>> On x86_64, rcu_read_lock/unlock are cheap thanks to static calls
>>>> (HAVE_PREEMPT_DYNAMIC_CALL). On ppc64le with the KEY-based
>>>> implementation, the weak memory model requires real barriers
>>>> (lwsync/isync) making each RCU read-side critical section
>>>> significantly more expensive.
>>>
>>> Weak memory model would need barriers irrespective of
>>> HAVE_PREEMPT_DYNAMIC_CALL
>>> or HAVE_PREEMPT_DYNAMIC_KEY. That's my assumption. I will look
>>> more into it. Also i don't know much about PREEMPT_RCU. So might take
>>> a while.
>>>
>>>>
>>>> This aligns with Christophe's earlier review comment that
>>>> HAVE_PREEMPT_DYNAMIC_CALL should be more performant. Would
>>>> implementing static calls for ppc64 be feasible to close this gap?
>>>>
>>>
>>> Static key too is expected to minimal cost. There maybe more into this.
>>>
>>>> Test details:
>>>> - Machine: IBM POWER10 (pvr 0080 0200), pHyp virtualization
>>>> - stress-ng 0.21.03, gcc 14.3.1, glibc 2.39
>>>> - Tuned profile: virtual-guest
>>>> - SELinux: enforcing (permissive recovers only ~7%)
>>>>
>>>> Happy to run additional tests if needed.
>>>>
>>>>
>>>
>>
>>
>
next prev parent reply other threads:[~2026-07-27 10:35 UTC|newest]
Thread overview: 16+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-02-10 18:43 [PATCH v4 0/1] powerpc: Enable dynamic preemption Shrikanth Hegde
2025-02-10 18:43 ` [PATCH v4 1/1] powerpc: enable " Shrikanth Hegde
2026-07-26 18:33 ` Jirka Hladky
2026-07-27 4:18 ` Shrikanth Hegde
2026-07-27 10:13 ` Jirka Hladky
2026-07-27 10:28 ` Shrikanth Hegde
2026-07-27 10:35 ` Christophe Leroy (CS GROUP) [this message]
2026-07-27 16:12 ` Paul E. McKenney
2026-07-27 10:59 ` Jirka Hladky
2026-07-27 11:03 ` Shrikanth Hegde
2026-07-27 12:34 ` Shrikanth Hegde
2026-07-27 16:07 ` Paul E. McKenney
2026-07-27 16:50 ` Jirka Hladky
2026-07-27 17:10 ` Shrikanth Hegde
2026-07-27 18:29 ` Christophe Leroy (CS GROUP)
2026-07-28 0:26 ` Jirka Hladky
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=330ca432-ddb6-4eeb-8415-697e16e73c43@kernel.org \
--to=chleroy@kernel.org \
--cc=bigeasy@linutronix.de \
--cc=jhladky@redhat.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linuxppc-dev@lists.ozlabs.org \
--cc=maddy@linux.ibm.com \
--cc=mpe@ellerman.id.au \
--cc=npiggin@gmail.com \
--cc=paulmck@kernel.org \
--cc=sshegde@linux.ibm.com \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.