From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-ej1-f54.google.com (mail-ej1-f54.google.com [209.85.218.54]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A39CD3EDAA3 for ; Tue, 25 Aug 2026 09:28:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.218.54 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787650142; cv=none; b=DG2aBWz1wbBMDsV15fYbbBZh5aVWdIUxWx7xSGZUZkdFtB8sVGeQAnWQ7PTN/96kH06lPo98NEuDWoMo3CBVwoZDryw/VHxNaQGWpbg9EcFT473CZ1PCKFyLtY7gAEBUmGZml97IxBzwdmp/Yiov2NmYczVA0iDQAnXsD+xMGnU= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787650142; c=relaxed/simple; bh=it/Wm0hwwDMU5ZnynexBLplwy6F+rvfkLfXizNRji/k=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=B2++ieVOkQdjRF8PDmHgkJTrknVPCUdja58pewRtPzaNFVpus0PsEgFEgoQ+k5pBCP1KQST2jUa6Ls7Gc443/zYzEUbYg91Day6GhcTbWKSI+C8jDhIo7W7wqGZY+3bKnvg3cNtmBMg4Nb15Q1Y5Yx9QFFhL9tsVBlFS52VHh60= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com; spf=pass smtp.mailfrom=suse.com; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b=KXecCzB9; arc=none smtp.client-ip=209.85.218.54 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=suse.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b="KXecCzB9" Received: by mail-ej1-f54.google.com with SMTP id a640c23a62f3a-c207cb16cf5so711995566b.1 for ; Tue, 25 Aug 2026 02:28:59 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=suse.com; s=google; t=1787650138; x=1788254938; darn=vger.kernel.org; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:from:to:cc:subject :date:message-id:reply-to:content-type; bh=syExKixb5ojbug7gf/PFSmcFjNNr9pCGcOpMliPCGJ0=; b=KXecCzB9GWbUwivOtVwi4yAPfGmIWLVHryQUaFOepMWsHii8s1wzz9Vk33+C5w89Mv Elt1+SCDfHUtGMRiQqQmjMtoZfa46TIaXVRWKUuTB+LB9OXccy7LRr60/nyTryrWHRzT B02CyiOVQMQ6IGb5dA9I/J3ax6et4EF4kcvxDx2k966mHYM5wG61a7S+8WWQHodDIxlv WdNYzzz+XgqFh1hqQUb82dVbYUQOuXXFCfI3LkgbO0mzoIcBASG+y0uc5dJkUwTC8TNz HyvAJsOMLqkWBwFnmNK0OiopNB25J2/M2aytPfQ1oo/mWMGIumzV6D+w8pf4yUhF5uQ2 E7iQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787650138; x=1788254938; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=syExKixb5ojbug7gf/PFSmcFjNNr9pCGcOpMliPCGJ0=; b=Uc8Y/cUds/+UflhngX35Ai35HSJ9Jd9xTI4xsJqIo5k8Z9wPH36nFynn8tfEbqoftK xF0EN/nlmpUxTkulGUvBONlwz6/uMFK2l+k4UJc9S/wcsRGk/f5/9xj9rIDsa48iet8K 89Yg1BvC97ZRfa0dl4iG9I/x6HJY9OfUvMnVaSz0JGAlqEweYmlqVspzt01pMuPj8OL5 FUZ15RqeTzica16uVDCIttUTWVHoPT+X9OG6H3KavQ3C+DUtsUb40hKwgi2jZVMF6xy1 5cy5lgVnsvt5rsYv1kDO67yf8XcwvBrbNWLFKmZqoG7IGj4kz66Ah3D9LaBNDz4E69pT 9DAQ== X-Forwarded-Encrypted: i=1; AHgh+RojxfNuQwg47wCpa3IygJVr4R1gJS1V9VOPiA0KrOkVlnEIugcQLTrjrhF6hya5jiAW1LJMbp4AAz9SNH8=@vger.kernel.org X-Gm-Message-State: AFuF++kIV9wG+DBoVNxj6Iweu4RmUaqQoibi8EiZGssv/ju+6Fwc5v87 ilyEFeE+zqxe0jYEHC1e02tw65zEfjGYsjBu6ju5Fjgl38xDk9/j7hLmWDiPBVABZc4= X-Gm-Gg: AR+sD12eIifxbY4++1zlstcik7ro4SdRrVrl+8YoJh5iSylgtfSCduXH7Pd29llPH/l KpuSqDfQ68GjzijMVirvkPmy1iq+jGiw37o0aZ8wOdGlVhuUFiBDKAvcdMfOKDEZa2K6Dm0jAuD v/xiLkA40R7FLClycXB6bpL/A1Pw+4I+TNHszByD2LnhNBkPoJTyUrmvKXJ4h32lg6bu12QcdRx 9YxHRP+iBQXLAp3m7C8o4gpw3ZR1ZipIZd4FyCN/TR9oIFUE69EjwkQsIsxeM6r7lrC8OAdqv0K qN60dfGKJdKE1SAMbc0PBJkDvUn1KSYjV+0v71A+XC9w78soW7JhqxGTnfLtB8xrYZHE6bMatRf kwLc7FVmixwhHq3SZYn7M+RDvcbfl4CDWqQPQxfo5NN3nOGJxYsm+Okz2lcdSO0KIHx7XwzV3/N Jjo8+Y64wTIv8sfpn0QSYzU/XTLOKGaPMo+XhBKiKE+KCpfUt3YxATUEVcyBSjgE9SFgm+lqu9 X-Received: by 2002:a17:907:3e2a:b0:c21:7feb:4574 with SMTP id a640c23a62f3a-c246a2cece6mr3358491066b.1.1787650137856; Tue, 25 Aug 2026 02:28:57 -0700 (PDT) Received: from pathway.suse.cz ([176.114.240.130]) by smtp.gmail.com with ESMTPSA id a640c23a62f3a-c24dfd91dd4sm399826766b.16.2026.08.25.02.28.56 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 25 Aug 2026 02:28:57 -0700 (PDT) Date: Tue, 25 Aug 2026 11:28:55 +0200 From: Petr Mladek To: Bradley Morgan Cc: Andrew Morton , Jinchao Wang , Feng Tang , Rio , Pnina Feder , Petr Pavlu , Sergey Senozhatsky , linux-kernel@vger.kernel.org, Sashiko , stable@vger.kernel.org Subject: Re: [PATCH v6 5/6] panic: allow force_cpu redirect from an NMI Message-ID: References: <20260818163806.17460-1-include@grrlz.net> <20260818163806.17460-6-include@grrlz.net> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260818163806.17460-6-include@grrlz.net> On Tue 2026-08-18 16:38:05, Bradley Morgan wrote: > nmi_panic() claims panic_cpu via panic_try_start() before calling > panic(). When the panic later reaches panic_try_force_cpu(), the > panic_in_progress() check sees panic_cpu set and refuses to redirect. > The crash kernel runs on the CPU that took the NMI instead of the CPU > requested with panic_force_cpu=: > > nmi_panic() > panic_try_start() wins, panic_cpu = X > panic("%s", msg) > vpanic() > panic_try_force_cpu() > panic_in_progress() true, panic_cpu is X > return false redirect bypassed > panic_try_start() already won > __crash_kexec() on X, not the requested CPU > > Try the redirect before claiming panic_cpu instead, as suggested by > Petr Mladek. nmi_panic() now calls panic_try_force_cpu() first and > claims panic_cpu only when no redirect happened. The requested CPU > claims panic_cpu itself when it runs panic(), so panic_cpu does not > need to be handed off. panic_try_force_cpu() copies the arguments > before formatting (patch 3), so nmi_panic() can pass them to vpanic() > again when no redirect happens. > > The redirect IPI is sent with smp_call_function_single_async(), which > is not guaranteed to work from NMI context. Treat it as best effort. > It is worth the risk because the redirection is only used when the > crash kernel would not work on the panicking CPU anyway. > > Keep returning when the panic is already running on this CPU. A > nested NMI, for example with unknown_nmi_panic while this CPU is > inside panic(), must return and let the interrupted panic() continue > instead of parking the CPU in nmi_panic_self_stop(). > > Mark the redirecting CPU offline before stopping it, like vpanic() > does, so that panic_other_cpus_shutdown() on the target CPU does not > wait for it. > > Fixes: 2e171ab29f91 ("panic: add panic_force_cpu= parameter to redirect panic to a specific CPU") > Reported-by: Sashiko > Closes: https://sashiko.dev/#/patchset/20260708164312.19044-1-include@grrlz.net > Cc: stable@vger.kernel.org > Signed-off-by: Bradley Morgan The patch looks good to me: Reviewed-by: Petr Mladek See some comments below. > --- a/kernel/panic.c > +++ b/kernel/panic.c > @@ -513,10 +513,11 @@ bool panic_on_other_cpu(void) > EXPORT_SYMBOL(panic_on_other_cpu); > > /* > - * A variant of panic() called from NMI context. We return if we've already > - * panicked on this CPU. If another CPU already panicked, loop in > - * nmi_panic_self_stop() which can provide architecture dependent code such > - * as saving register state for crash dump. > + * A variant of panic() called from NMI context. The panic is first > + * redirected to the CPU requested via panic_force_cpu=, when configured. > + * We return if we've already panicked on this CPU. If another CPU already > + * panicked, loop in nmi_panic_self_stop() which can provide architecture > + * dependent code such as saving register state for crash dump. > */ > __printf(2, 3) > void nmi_panic(struct pt_regs *regs, const char *fmt, ...) > @@ -525,6 +526,16 @@ void nmi_panic(struct pt_regs *regs, const char *fmt, ...) > > va_start(args, fmt); > > + /* Try to redirect to the requested CPU before claiming panic_cpu. */ > + if (panic_try_force_cpu(fmt, args)) { > + /* > + * Mark ourselves offline so panic_other_cpus_shutdown() won't > + * wait for us on architectures that check num_online_cpus(). > + */ > + set_cpu_online(raw_smp_processor_id(), false); It is a bit strange that we set this CPU offline in this code path but not when the below panic_try_start() fails. nmi_panic_self_stop(regs) is called in both situations. I guess that we should mark the CPU offline in the other case as well. But it is an indepent change. We could fix this later. Feel free to keep this patch as is. > + nmi_panic_self_stop(regs); > + } > + > if (panic_try_start()) > vpanic(fmt, args); Sashiko AI complains here: | Does this sequence result in printing the 'target CPU is offline' warning | twice? | | If panic_force_cpu is set to an offline CPU, the newly added call to | panic_try_force_cpu() evaluates this logic: | | /* Target CPU is offline, can't redirect */ | if (!cpu_online(panic_force_cpu)) { | pr_warn("panic: target CPU %d is offline, continuing on CPU %d\n", | panic_force_cpu, this_cpu); | return false; | } | | Because the redirect fails, nmi_panic() continues and calls vpanic(). Since | vpanic() unconditionally calls panic_try_force_cpu() as well, the same | offline check will execute and the warning might be printed a second time | before the panic_in_progress() check can abort the redundant redirect attempt. It is right. We might fix this by checking that panic_try_force_cpu() has been called twice, for example, by checking the panic_redirect_cpu value at the beginning. But I would personally ignore this problem. It is a corner case. The fix would make the code more hairy. IMHO, it is not worth it. Best Regards, Petr