All of lore.kernel.org
 help / color / mirror / Atom feed
From: Jarkko Sakkinen <jarkko@kernel.org>
To: Jun Miao <jun.miao@intel.com>
Cc: dave.hansen@linux.intel.com, tglx@linutronix.de,
	mingo@redhat.com, bp@alien8.de, hpa@zytor.com,
	kai.huang@intel.com, linux-sgx@vger.kernel.org,
	linux-kernel@vger.kernel.org, x86@kernel.org, fan.du@intel.com,
	challvy.tee@gmail.com
Subject: Re: [PATCH v6] x86/sgx: Report RCU-Tasks quiescent state in EPC sanitization loop
Date: Fri, 24 Jul 2026 10:07:28 +0300	[thread overview]
Message-ID: <amMPMLAzE0nUiH9C@kernel.org> (raw)
In-Reply-To: <20260724025514.525117-1-jun.miao@intel.com>

On Fri, Jul 24, 2026 at 10:55:14AM +0800, Jun Miao wrote:
> When the kernel boots from kexec, the EPC pages may have a stale state.
> The kernel sanitizes all EPC pages to reset them to a clean state before
> their first use in any enclave.  The EPC size could be several GBs and
> resetting them could take a significant amount of time.  Because of that,
> the kernel performs the reset in a loop through a kernel thread ksgxd() at
> early boot, and there's a cond_resched() after resetting each EPC page.
> 
> This is fine in most cases, but becomes a problem when there's other kernel
> code waiting for an RCU-Tasks grace period but the cond_resched() in
> ksgxd() never triggers rescheduling.  Because cond_resched() doesn't report
> a quiescent state when it doesn't trigger rescheduling, the thread that is
> waiting for an RCU-Tasks grace period will wait until all EPC pages are
> reset.
> 
> For instance, BPF LSM subsystem can invoke synchronize_rcu_tasks() at
> kernel boot time.  A VM with a large EPC assigned and BPF LSM enabled can
> take a long time to boot, with a call trace triggered:
> 
>     rcu_tasks_wait_gp: rcu_tasks grace period number 1 (since boot) is
> 	130631 jiffies old.
>     INFO: task systemd:1 blocked for more than 122 seconds.
>     ...
>     task:systemd  state:D stack:0  pid:1  tpid:1  ppid:0  flags:0x00000002
>     Call Trace:
>     ...
>     schedule_timeout+0x157/0x170
>     wait_for_completion+0x88/0x150
>     __wait_rcu_gp+0x17e/0x190
>     synchronize_rcu_tasks_generic+0x64/0x60
>     ...
>     synchronize_rcu_tasks+0x15/0x20
>     register_ftrace_direct+0x31f/0x350
>     ...
>     bpf_trampoline_link_prog+0x33/0x60
>     bpf_tracing_prog_attach+0x3c5/0x5f0
> 
> Replace cond_resched() with cond_resched_tasks_rcu_qs() which explicitly
> reports quiescent state regardless of whether actual rescheduling is
> triggered.  Resetting all EPC pages in ksgxd() isn't performance critical
> so the extra cost of cond_resched_tasks_rcu_qs() isn't a problem.
> 
> Tests showed this reduced the VM kernel boot time from ~50s to ~700ms.
> 
> Fixes: e7e0545299d8 ("x86/sgx: Initialize metadata for Enclave Page Cache (EPC) sections")
> Suggested-by: Kai Huang <kai.huang@intel.com>
> Co-developed-by: Fan Du <fan.du@intel.com>
> Signed-off-by: Fan Du <fan.du@intel.com>
> Signed-off-by: Jun Miao <jun.miao@intel.com>
> Tested-by: Challvy Tee <challvy.tee@gmail.com>
> Reviewed-by: Kai Huang <kai.huang@intel.com>
> Link: https://github.com/systemd/systemd/issues/40423
> ---
> v1 -> v2:
>  - Clarify the RCU Tasks stall root cause.
>  - Use cond_resched_rcu_qs() following the Kai`s suggestion.
> 
> v2 -> v3:
>  - cee439398933 ("rcu: Rename cond_resched_rcu_qs() to cond_resched_tasks_rcu_qs()")
> 
> v3 -> v4:
>  - Trim down/rewrite changelog following Kai`s suggestion.
> 
> v4 -> v5:
>  - Change the title, not state the problem directly
>  - Corrected spelling and grammatical errors by Kai
>  - Add "Reviewed-by: Kai Huang"
> 
> v5 -> v6:
>  - Add the exactly kind of reminder that why cond_resched() does 
>    not trigger rescheduling.
> 
> ---
>  arch/x86/kernel/cpu/sgx/main.c | 8 +++++++-
>  1 file changed, 7 insertions(+), 1 deletion(-)
> 
> diff --git a/arch/x86/kernel/cpu/sgx/main.c b/arch/x86/kernel/cpu/sgx/main.c
> index 4505f808af5e..207e680732a4 100644
> --- a/arch/x86/kernel/cpu/sgx/main.c
> +++ b/arch/x86/kernel/cpu/sgx/main.c
> @@ -106,7 +106,13 @@ static unsigned long __sgx_sanitize_pages(struct list_head *dirty_page_list)
>  			left_dirty++;
>  		}
> 
> -		cond_resched();
> +		/*
> +		 * cond_resched() only schedules when TIF_NEED_RESCHED is set.
> +		 * During this boot-time loop that condition may not happen for a
> +		 * long time, so report an RCU-Tasks quiescent state explicitly.
> +		 * Therefore, change cond_resched() to cond_resched_tasks_rcu_qs().
> +		 */
> +		cond_resched_tasks_rcu_qs();
>  	}
> 
>  	list_splice(&dirty, dirty_page_list);
> -- 
> 2.43.0
> 

Reviewed-by: Jarkko Sakkinen <jarkko@kernel.org>

BR, Jarkko

      reply	other threads:[~2026-07-24  7:07 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-24  2:55 [PATCH v6] x86/sgx: Report RCU-Tasks quiescent state in EPC sanitization loop Jun Miao
2026-07-24  7:07 ` Jarkko Sakkinen [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=amMPMLAzE0nUiH9C@kernel.org \
    --to=jarkko@kernel.org \
    --cc=bp@alien8.de \
    --cc=challvy.tee@gmail.com \
    --cc=dave.hansen@linux.intel.com \
    --cc=fan.du@intel.com \
    --cc=hpa@zytor.com \
    --cc=jun.miao@intel.com \
    --cc=kai.huang@intel.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-sgx@vger.kernel.org \
    --cc=mingo@redhat.com \
    --cc=tglx@linutronix.de \
    --cc=x86@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.